#Offline AI Translation
Offline AI translation combines privacy and reliability for professional document workflows. This guide covers why local processing matters, what "offline" honestly requires, and how CatScribe's local engine stack fits together.
#Why Offline Matters
- Privacy. Unpublished manuscripts, client documents, and confidential material never leave your machine. There is no third-party API to trust, no retention policy to read.
- Reliability. No rate limits, no quota exhaustion mid-book, no service deprecations. A pipeline that works today works identically next year.
- Cost. Local engines have no per-character fees. A million-word novel costs electricity.
- Predictability. Performance depends on your hardware, which you control, rather than on a shared service's load.
The trade-off is equally honest: local neural MT on consumer hardware generally trails the best cloud LLMs on raw fluency, and you pay in RAM, disk, and time instead of dollars. The workflows in this handbook — glossaries, refinement layers, human review — exist precisely to close that gap.
#What "Offline" Honestly Requires
No offline system is offline on day one. Models are large, so every local stack has a provisioning phase that needs internet:
| Component | Download size | When |
|---|---|---|
| Argos language packs | ~50–100 MB per language pair | First use of each pair |
| MarianMT / NLLB models | ~200–600 MB (from Hugging Face) | First use, then cached locally |
| COMET quality scoring (optional) | ~1–2 GB | When enabled in Settings |
| AI subtitle detection (optional) | ~250 MB | If installed |
| Ollama models (optional, external) | ~2 GB and up per model | Per model |
Depending on your installer, an English→Portuguese Argos pack may be pre-seeded, but plan on at least one online session to download the language pairs and optional components you need. After provisioning, translation runs fully local — you can translate on a plane, and your documents are never uploaded anywhere regardless of connectivity.
Practical rule: while you have internet, install everything you might want — engine, language packs, AI models, COMET — then verify by translating a test snippet with networking disabled.
#The Local Engine Stack
CatScribe ships several offline engines; in Simple mode they appear as one "Offline Translation" option, with the actual engine set by the "Preferred Offline Provider" setting. See Translation Engines and Offline Translation.
| Engine | Character | Notes |
|---|---|---|
| "Argos (Local)" | Fast, light, the recommended default | 30+ languages via downloadable packs |
| "MarianMT (Local)" | Neural MT built into the app | Models cached after first download; no extra services |
| "NLLB (Local)" | Quality-first neural MT, 200+ languages | Strongest on complex or literary text; uses more RAM/CPU |
| "LibreTranslate (Local)" | Self-hosted server option | Requires Docker on your machine |
On top of base MT, AI refinement stages (validation, rewrite, technical check) can run on local LLMs — this is where local quality catches up with cloud output while keeping everything on-device.
#Ollama And LM Studio Are External Programs
Two integrations matter for local AI refinement, and both are separate applications, not CatScribe components:
- Ollama — a local LLM runtime. CatScribe detects it, offers an automatic installer ("Install Ollama Automatically"), and manages model downloads from its Providers screen, with hardware-tier recommendations (a 16 GB+ machine with a GPU is recommended very different models than an 8 GB laptop).
- LM Studio — a local LLM desktop app with an OpenAI-compatible server. CatScribe auto-detects it and uses its loaded models in the same AI model selectors as Ollama. One gotcha the app itself warns about: "Opening the LM Studio app does not start its Local Server — start it manually under the Developer tab."
Neither is required for offline translation — the MT engines above work alone — but one of them is required for the AI refinement layers and "Improve with AI".
#Hardware Reality
- Base MT engines (Argos) run comfortably on modest CPUs.
- MarianMT/NLLB load models into system RAM and run on CPU; expect a few GB of RAM in use during translation, and heed the in-app "High RAM Usage" warning when selecting them on small machines.
- Local LLM refinement is the expensive layer: model size must fit your RAM/VRAM. CatScribe's Auto Mode and the Providers screen's hardware profile pick appropriately sized models for your machine.
- More quality = more compute. "Fast" presets skip refinement entirely; "Maximum" runs multiple AI stages. Choose per job, not once.
#Offline Workflow Pattern
- Provision online: install the engine, download language packs/models for your pairs, plus any optional components.
- Verify offline: disconnect and run a Quick Translate snippet and a small file end to end.
- Translate in batches: chapter-sized jobs are recoverable and reviewable; see Translating Books.
- Review locally: glossary enforcement, QA flags, and (if installed) COMET quality scores all run on-device.