#Working With EPUB, DOCX, and PDF
Different file formats behave differently during translation. Choosing the right workflow helps preserve structure and reduces cleanup time. This page covers what File Translation actually does with each format — import behavior, export options, and the limits you will hit.
#Supported Formats And Limits
File Translation accepts EPUB, DOCX, and PDF, up to 50 MB per document ("Documents up to 50MB · Subtitles up to 5MB · Videos up to 500MB."). Subtitle files (SRT/VTT) are handled by the separate Subtitle Translation screen, not File Translation.
| Format | Best for | Watch for |
|---|---|---|
| EPUB | Books and reflowable long-form text | Fixed-layout/manga EPUBs are not supported |
| DOCX | Manuscripts, reports, editable documents | Complex custom styling — verify a sample export |
| Fixed-layout documents and reference files | Text extraction quality, columns, scanned pages |
#Output Formats
The "Output Format" selector offers:
- "Same as input" — the default, and the only way to get PDF output (from a PDF input). There is no standalone PDF output option.
- "Plain Text (.txt)"
- "Word Document (.doc/.docx)"
- "EPUB (.epub)"
A useful pattern: translate a difficult PDF with DOCX or EPUB output to get a clean, editable result instead of fighting layout reconstruction.
#DOCX Workflow
For DOCX, CatScribe's preferred build path patches the original document package, so images, styles, and headers from the source file carry into the output.
Recommended steps:
- Finalize tracked changes and remove comments that are not part of the content before importing.
- Translate a small section first and verify headings, lists, tables, and images in the export.
- Review the translated document in a word processor; reapply any exotic custom styling manually if needed.
Legacy .doc files are not accepted by the uploader — save them as .docx first.
#PDF Workflow
PDF is the most configurable format because PDFs vary the most. After upload, the "Document Advisor" card analyzes the file (searchable vs scanned, page count, columns, image dominance) and recommends an approach — read it before starting.
#PDF Translation Mode
The "PDF Translation Mode" selector controls how the PDF is processed:
| Mode | Use when |
|---|---|
| "Auto (recommended)" | Default — CatScribe picks based on the document analysis |
| "Simple text" | Simple text PDFs; ignores layout and creates a clean linear PDF |
| "High Fidelity (preserves layout and tables)" | Layout and tables matter |
| "Fast (OCR for scanned PDFs)" | Scanned/image-only PDFs — CatScribe runs OCR itself; no external conversion needed |
| "Via DOCX (structure + images)" | Structure plus images; requires LibreOffice installed on your system |
If CatScribe detects a scanned PDF, it tells you and routes you toward the OCR path — you do not need to convert scanned PDFs to searchable text in another tool first.
#What PDF output preserves
- Fonts: bundled Noto fonts cover Latin, Cyrillic, Greek, and Vietnamese with bold/italic variants. CJK text is supported with a dedicated font, but CJK renders without bold or italic (bold degrades to regular weight).
- Page furniture: a localized "Page X of Y" footer, plus an optional header title, is added to the output.
- Images: preserved with deduplication and compression; never upscaled.
- Links: URLs become real clickable link annotations.
- Layout: two-column flow is applied automatically for documents detected as multi-column scientific layouts; otherwise output is single-column.
#PDF limits
Very large or complex PDFs can be rejected with "This PDF is too large or complex to translate in a single pass." — the fix is to split the PDF into smaller files (fewer pages) and translate them separately. Damaged files report "This PDF appears to be damaged or invalid."
Recommended steps:
- Check the Document Advisor verdict and analysis chips ("Searchable PDF" / "Scanned PDF", page count, "OCR required").
- Test a few pages (or a split excerpt) before the full file.
- In review, watch for merged words, broken line breaks, and headers absorbed into paragraphs — extraction artifacts concentrate around headers, footers, and column boundaries. See Common Issues & Fixes.
- If layout reconstruction disappoints, re-run the job with DOCX or EPUB output instead of PDF output.
#When To Convert First
Use a cleaner source instead of PDF when:
- The same content exists as EPUB or DOCX — always prefer those.
- The PDF's text order comes out wrong even in "High Fidelity (preserves layout and tables)" mode.
- Tables, footnotes, or columns extract badly and the content matters more than the layout.
If you must convert externally, review the converted source before importing. A clean source file produces a cleaner translation than any amount of post-editing.
#Export Verification Checklist
- Open the export in the app your readers will use (EPUB reader, word processor, PDF viewer).
- Chapter/section count and order match the source.
- Images present; tables readable.
- Headings and table of contents intact (EPUB/DOCX).
- Page numbering sane and no clipped lines (PDF).
- Spot-check accented characters and any CJK passages.