CatScribe Docs

#How Translation Works

CatScribe is an offline-first translation workspace for long documents. A document is not translated in one giant request — it moves through a pipeline designed so that big books finish reliably, failures stay small and retryable, and every chunk remains individually reviewable.

#The Pipeline

Import -> Analyze -> Chunk -> Translate (+ AI layers) -> Review -> Export
           |                        |                       |
    Document Advisor         retries / pause /        CAT Editor, QA,
    format detection         resume / idempotent      quality scores
                             uploads

#1. Import

File Translation accepts PDF, DOCX, and EPUB files up to 50 MB ("Drag and drop a file here, or click to select"). Files are validated by content, not just extension — a renamed file that is not really a PDF/DOCX/EPUB is rejected. SRT/VTT subtitles are a separate pipeline with their own screen; see Translating Subtitles.

Practical prep still matters: confirm source and target languages, keep a backup of the original, and for very large books consider working by part — see Working With EPUB, DOCX, and PDF.

#2. Analyze

After upload, CatScribe analyzes the document and shows what it found: page count, whether a PDF is searchable or scanned, layout complexity, and estimated work. The "Document Advisor" card rates workflow fit (with reasons like "The document appears to be scanned."), estimates quality/time/memory, and suggests engines. For PDFs you also get the "PDF Translation Mode" selector: "Auto (recommended)", "Simple text", "High Fidelity (preserves layout and tables)", "Fast (OCR for scanned PDFs)", or "Via DOCX (structure + images)".

#4. Translate

Each chunk goes to your selected engine — see Translation Engines. Optionally, AI refinement layers then post-process each chunk (validation, rewrite, technical review) using Ollama, LM Studio, or Claude; the Smart Pipeline Quality Threshold (default 85) lets high-scoring chunks skip the expensive rewrite. Auto Mode chooses all of this from a quality preset; manual mode exposes every control.

While running, the "Translation Progress" panel shows the current chunk number, active layer, provider, model, elapsed time, and retry count. Statuses per chunk: Pending, Translating, Completed, Failed, Skipped, Partial, Paused.

#Stability Features (Why Big Jobs Finish)

  • Idempotent uploads — every upload carries a unique idempotency key; if the app retries an upload after a network hiccup, the server recognizes the duplicate and continues the same job instead of creating a second one.
  • Per-chunk retries — a failed chunk is retried up to 3 times automatically with increasing back-off (network errors wait ~5 s/15 s/45 s; rate limits pause up to 5 minutes). "Auto-Retry Failed Chunks" is on by default.
  • Pause / Resume — pause a job from the Translations screen and resume later; settings changed while paused apply on resume ("New settings will be applied when translation is resumed"). "Auto-Resume Pending Translations" (default on) picks up interrupted jobs after an app restart, once the local engines are ready.
  • Manual recovery — "Retry Failed" and "Retry All" on the Translations screen re-run failed chunks; individual chunks can be re-run with a different engine via Chunk Overrides.
  • Failure of some chunks does not sink the job — it completes as Partial, and you fix the failed chunks afterwards.
  • Optional: "Prevent system sleep while translating" in Settings for long overnight runs.

#5. Review

Review happens in the CAT Editor: source and target side by side, per-chunk quality scores (heuristics + optional COMET), QA Flags for concrete issues (tag mismatches, number differences, glossary violations), version history, and the approve loop (Ctrl+Enter). Glossary terms attached to the job are enforced during translation and re-verified during QA — see Glossaries.

For long projects, review in passes: meaning first, terminology second, style last — see Reviewing Translations.

#6. Export

The "Output Format" selector offers: "Same as input", "Plain Text (.txt)", "Word Document (.doc/.docx)", and "EPUB (.epub)". PDF output is produced by choosing "Same as input" on a PDF source. You can regenerate the output any time after edits with "Generate Document" on the Translations screen.

Format fidelity in brief:

  • EPUB exports re-inject translated text into the original book's structure, preserving the publisher's formatting, styles, and images.
  • DOCX exports patch the original document package where possible, keeping images, styles, and headers.
  • PDF exports re-typeset the text with embedded fonts (including CJK support), page numbers, links, images, and table rendering; layout modes handle single- and multi-column documents.

Always open the exported file and check structure, headings, tables, and images before shipping — see Editing Without Breaking Formatting.

#CatScribe Compared With Simple Translation Tools

Web translators are fine for snippets (CatScribe's own Quick Translate covers that, up to 8,192 characters). The pipeline above is for projects where translation must stay consistent and survive scale:

  • A repeatable chunk-based workflow instead of one-off translation.
  • Glossary enforcement for names and terminology.
  • Chunk-level review, scoring, and correction.
  • Local, offline-first processing — see Offline Translation.
  • Retries, resume, and partial-failure recovery for large books.