BETA This site is in beta. Information is still being added and reviewed.
Practitioner hub
Practitioner guide · AI literacy

Write source text that translates well

Most agencies think about translation as a problem that starts when the English is finished: you hand a document to a translator or paste it into a translation tool, and whatever comes back is what your clients read. This guide moves one step upstream. The single cheapest way to improve your multilingual materials is to write the original English well in the first place, because a clear source produces a clearer translation, whether a human or a machine does the work, and a tangled source produces a tangled translation in every language you publish, silently, since no one on your staff can read the output to catch it.

Why it matters

The source decides the quality, for machines and for humans

Translation, human or machine, cannot recover meaning that the source never made explicit. If your English sentence can be read two ways, the translator has to guess which reading you meant, and a wrong guess becomes a confident, fluent error in the target language. Ambiguity, buried logic, and unfamiliar phrasing are not stylistic problems a good translator smooths over. They are decisions you are pushing downstream to someone with less context than you have.

For machine translation this is measurable. Researchers tested the same content written two ways, once violating a set of plain, controlled-writing rules and once conforming to them, then ran both versions through a translation system. Source text that followed the rules produced higher-quality machine translation across four typologically different target languages: Dutch, Chinese, Arabic, and French (Aikawa et al., 2007). Later work isolated the effect of individual rules and confirmed that constraints on the source, such as shorter sentences and reduced ambiguity, change the machine’s output in identifiable ways (Marzouk, 2021). The professional localization field calls this practice “pre-editing”: clean up the English before translation instead of only fixing the output afterward.

Human translators benefit for the same reasons, plus one more. A translator working from a short, direct, unambiguous sentence spends their effort finding the right words in the target language, not decoding what you meant. Clear source text lowers the cost of translation and the error rate at the same time. Source quality also matters unevenly across languages: machine translation systems learn from large collections of text, and even the largest modern systems, built to cover 200 languages, report systematically lower quality for lower-resourced languages because the underlying data is scarce (NLLB Team et al., 2024). You cannot fix the amount of training data behind a tool, but you can control how much work you hand it. Writing well upstream is a form of equity for the languages your clients speak that the technology serves least.

Plain-language practices that measurably help

These come from two traditions that agree with each other: health-literacy guidance for human readers and controlled-language guidance for machine translation.

Short sentences, one idea each

A long sentence forces a reader or a translator to hold several clauses in mind and work out how they relate. Split compound instructions into separate sentences; this is the single most consistent lever for both comprehension and translation quality (Marzouk, 2021).

Active voice, name who does what

“Submit the form” is clearer than “The form will be reviewed.” A hidden actor is exactly the kind of gap a translator has to fill by guessing (Federal Plain Language Guidelines; CDC Clear Communication Index).

One consistent term per concept

Calling the same thing an “appointment,” a “visit,” and a “session” invites a machine to translate it three different ways. Pick one term for each concept and match it to your glossary.

No idioms, metaphors, or cultural references

“Reach out” and “down the road” carry almost no meaning a plain verb cannot carry better, and the literal translation is often what you get. Write “contact us” and “later.”

No ambiguous pronouns

“Take it after meals; if it upsets your stomach, stop it” leaves “it” doing too much work. Different languages resolve pronouns differently. Name the noun.

No double negatives

“You are not ineligible” is a puzzle even in English, and it often inverts in translation. Say the positive: “You may qualify.”

Spell out abbreviations

Acronyms rarely carry across languages; a translator or a machine may transliterate them, expand them wrongly, or leave them untouched. Write the full term at least once before any short form.

Unambiguous numbers and dates

“6/3/2026” is June 3 to a U.S. reader and March 6 to much of the rest of the world, and translation will not fix that. Write the month as a word and state quantities concretely.

Before and after

A missed-appointment notice

Dense source

Should you be unable to attend your scheduled appointment, kindly notify our office no less than 24 hours in advance so as to avoid being assessed a missed-visit fee, which could otherwise be applied to your account.

Plain source

Please call us at least 24 hours before your appointment if you cannot come. If you do not call, we may charge a fee for the missed visit.

The rewrite splits one long sentence into two, uses the active voice, drops the inverted “should you be unable” opening, and states the deadline as “at least 24 hours” instead of “no less than 24 hours.”
Before and after

A benefits eligibility message

Dense source

You are not ineligible for these benefits, but don’t count your chickens; nothing is set in stone until we’ve gone over your paperwork with a fine-tooth comb.

Plain source

You may qualify for these benefits. We cannot confirm until we review all of your documents.

The rewrite removes the double negative (“not ineligible” becomes “may qualify”), cuts three idioms that do not translate (“count your chickens,” “set in stone,” “fine-tooth comb”), and states the condition plainly in the active voice.
The health-literacy grounding

Plain language helps every reader, not only those reading in translation

Tens of millions of U.S. adults have limited health literacy, and a large share of routine health and benefits materials is written above the level many adults can comfortably read (Institute of Medicine, 2004). Plain language is the recognized response to that gap, and its benefits reach every reader, not only those with the lowest literacy (Stableford and Mettger, 2007).

The effect shows up in controlled tests. In a randomized trial, patients given a plain-language version of self-injection instructions correctly described more of the preparation and pre-injection steps and performed more of the injection steps correctly than patients given the standard version, completing on average 13.1 of 15 steps correctly versus 10.8 (Smith and Wallace, 2013). A systematic review of 49 randomized trials of prescription-medication materials likewise identified plain-language principles among the design practices that improve patient comprehension (Mullen et al., 2018).

Writing the source in plain language is not a concession you make for translation. It improves the English document your English-reading clients receive too.

A caution

Readability scores: a useful signal, a poor target

Many agencies run their drafts through a readability formula, such as Flesch-Kincaid or SMOG, and aim for a grade level. Use these as a rough smoke alarm, not a goal. Readability formulas were largely developed for children’s educational materials, they count surface features like word length and sentence length, and they ignore whether the reader actually understands the content, how the page is laid out, and who the reader is (Redish, 2000).

You can lower a grade-level score by chopping sentences in half while leaving the meaning just as murky, and you can write a genuinely clear sentence that a formula penalizes for a long but common word. Treat a high score as a prompt to look at the writing, then judge it by the practices above and, when you can, by testing it with real readers. A number is a hint, not a verdict.

Your workflow

How this connects to your glossary

Think of two documents working together. The community glossary is the agreed list of terms and their approved translations. The source English is where those terms get used. The practices above keep your English consistent enough that the glossary can do its job: one term per concept, spelled out, no synonyms drifting in, no idioms smuggling in meaning the glossary never approved. When a new term appears in a draft that is not in the glossary, that is the signal to add it deliberately, decide its translation once, and reuse it, instead of letting each translation invent its own.

A practical order of operations: draft the material, run it against the checklist below, reconcile every term against the glossary, and only then send it to translation, human or machine. If you use machine translation, this upstream cleanup is the pre-editing step the research shows reduces downstream errors (Aikawa et al., 2007; Marzouk, 2021). If you use human translators, it is the courtesy that lets them do their real work.

Before a document goes to translation

Sources

  1. Aikawa, T., Schwartz, L., King, R., Corston-Oliver, M., & Lozano, C. (2007). Impact of controlled language on translation quality and post-editing in a statistical machine translation environment. Proceedings of Machine Translation Summit XI, Copenhagen. link
  2. Institute of Medicine, Committee on Health Literacy (Nielsen-Bohlman, L., Panzer, A. M., & Kindig, D. A., eds.) (2004). Health Literacy: A Prescription to End Confusion. National Academies Press. doi:10.17226/10883
  3. Marzouk, S. (2021). An in-depth analysis of the individual impact of controlled language rules on machine translation output: a mixed-methods approach. Machine Translation, 35(2), 167-203. doi:10.1007/s10590-021-09266-0
  4. Mullen, R. J., Duhig, J., Russell, A., Scarazzini, L., Lievano, F., & Wolf, M. S. (2018). Best-practices for the design and development of prescription medication information: A systematic review. Patient Education and Counseling, 101(8), 1351-1367. doi:10.1016/j.pec.2018.03.012
  5. NLLB Team, Costa-jussà, M. R., Cross, J., Çelebi, O., et al. (2024). Scaling neural machine translation to 200 languages. Nature, 630. doi:10.1038/s41586-024-07335-x
  6. Redish, J. (2000). Readability formulas have even more limitations than Klare discusses. ACM Journal of Computer Documentation, 24(3), 132-137. doi:10.1145/344599.344637
  7. Smith, M. Y., & Wallace, L. S. (2013). Reducing drug self-injection errors: a randomized trial comparing a "standard" versus "plain language" version of Patient Instructions for Use. Research in Social and Administrative Pharmacy, 9(5), 621-625. doi:10.1016/j.sapharm.2012.10.007
  8. Stableford, S., & Mettger, W. (2007). Plain language: a strategic response to the health literacy challenge. Journal of Public Health Policy, 28(1), 71-93. doi:10.1057/palgrave.jphp.3200102
  9. U.S. Centers for Disease Control and Prevention (n.d.). The CDC Clear Communication Index and Everyday Words for Public Health Communication. CDC guidance. link
  10. U.S. General Services Administration, Plain Language Action and Information Network (n.d.). Federal Plain Language Guidelines. plainlanguage.gov. link

This guide is the companion to the community glossary guide: that one checks translations after they exist, this one is about the writing you control before anything gets translated at all. Adapt it to your agency’s policies and the languages you serve.