Private AI
Blog/Private AI document translation for sensitive documents
For AI assistants

Private AI document translation for sensitive documents

By Samuel Seidel · September 9, 2026 · 6 min read

Google Translate and DeepL are good enough for a menu or an email. They're the wrong tool for a merger agreement, a patient record, or a set of financial statements headed for a foreign regulator, because the moment you paste that text in, it has left your organization and gone to a third party whose terms of service you probably haven't read closely enough to bet a client relationship on.

The legal and medical translation industry has known this for years, which is why certified translators still exist and why hospitals sign business associate agreements with translation vendors instead of using whatever's free. The gap is that a proper managed translation vendor is slow and expensive for routine internal work, and a consumer tool is fast but leaks. A self-hosted model splits the difference: translation quality that's genuinely usable, running on hardware that never sends the document anywhere.

Who actually needs this

What open-weight models can actually do here

Translation quality from general-purpose open models has closed most of the gap with dedicated translation engines for the language pairs that matter to most businesses (the major European languages, Chinese, Japanese, Arabic). Models like Qwen2.5 and Llama 3.3 handle document-length translation with reasonable fidelity to register and terminology, particularly if you give them domain context in the prompt. For genuinely specialized text, dense legal or medical terminology, idiom-heavy source material, a human reviewer should still check the output before it goes anywhere official. Treat the model as a first-pass translator that removes most of the manual work, not a certified translation service.

If your volume justifies it, a dedicated translation model like NLLB-200 (Meta's No Language Left Behind, open-weight, covering 200 languages) is worth running alongside a general model, since it's purpose-built for translation rather than adapted to it. For most business use, though, a general instruction-tuned model with a well-written prompt does the job and saves you running two model servers.

Setting it up

The simplest version is a local model served through an OpenAI-compatible API, called from a small script that reads a document, translates it in chunks, and writes the result back out. For PDFs and Word documents, extract the text first, translate, then re-insert it, since asking a model to also preserve exact document formatting tends to produce worse translations than treating format handling separately.

# serve a general-purpose model with strong multilingual coverage
vllm serve Qwen/Qwen2.5-72B-Instruct --host 0.0.0.0 --port 8000

# translate.py - chunk-and-translate a text file, called per document
import sys, requests

def translate(text, target_lang, domain_hint):
    prompt = (
        f"Translate the following {domain_hint} text into {target_lang}. "
        "Preserve legal/technical terminology precisely. Do not summarize or omit anything. "
        "Output only the translation.\n\n" + text
    )
    r = requests.post("http://spark.internal:8000/v1/chat/completions", json={
        "model": "Qwen2.5-72B-Instruct",
        "messages": [{"role": "user", "content": prompt}],
        "temperature": 0.2,
    })
    return r.json()["choices"][0]["message"]["content"]

if __name__ == "__main__":
    src, target, domain = sys.argv[1], sys.argv[2], sys.argv[3]
    text = open(src, encoding="utf-8").read()
    print(translate(text, target, domain))

Run as python translate.py contract.txt German legal. The domain_hint argument matters more than it looks: telling the model it's translating a legal contract versus a medical chart changes how it handles ambiguous terms and register, and it costs nothing to specify.

For long documents, chunk by section rather than by a fixed character count, so the model always sees complete sentences and paragraph context. A contract clause split mid-sentence across two API calls translates worse than the same clause sent whole.

Where the hardware fits

A 72B-class model at Q4 quantization needs roughly 45GB of weights, which fits inside a single Spark's 128GB of unified memory with room for long documents in context. If you're translating at volume across a legal or medical team, that's one machine handling the whole department's throughput, at a flat $0.79/hour rather than a per-page or per-token bill that scales with exactly the volume you're trying to serve.

The setup cost here is genuinely small: stand up the model server, write the chunking script, have someone who speaks both languages review the first batch of output against the source. After that it's a tool anyone on the team can point a document at, and the document never leaves the building.

Related pages

Translate sensitive documents without sending them anywhere.

A dedicated DGX Spark in EU-Central, $0.79/hour, 128GB of unified memory for large multilingual models.

Deploy a Spark Private LLM hosting