BETA This site is in beta. Information is still being added and reviewed.
Workshop overview
4

The language gap and the interpreter right

The languages your clients most need are exactly where AI is least reliable. This is the most important module for immigrant and refugee services.

Start activity

Back-translate it, then start the glossary

  1. Take one sentence you would actually say to a client that carries a consequence: a deadline, a dose, a right.
  2. Machine-translate it into a client language, then use a second tool to translate the result back, without looking at the original.
  3. Whatever came back wrong, write the term and the agreed translation on a card. Those cards are the start of your community glossary.

Would you have noticed the change if you did not speak the language?

  • AI answers get less accurate as a language has less text online to learn from. The languages your clients speak are often the ones the models have seen least, so a question that works in English can come back wrong in languages like Amharic, Dari, or Haitian Creole.
  • In a health-question benchmark, AI answers were several times more likely to be incorrect in a non-English language than in English, and the gap widened for the lowest-resource languages.
  • Safety guardrails are weaker outside English. A prompt a model would refuse in English can pass in another language, and the risk runs highest in the lower-resource languages your clients use.
  • Machine translation is not safe for consequential documents. Accuracy of emergency-department discharge instructions dropped below 90 percent for several languages, and translation errors give no warning that they are there.
  • The same message costs more and degrades more in your clients' languages. Many non-English scripts take far more tokens to say the same thing, so answers come back shorter, slower, and more likely to be cut off.
  • A free qualified interpreter is a legal right for consequential encounters, and language-concordant care is linked to better outcomes. Keep AI translation for low-stakes text, and get a qualified interpreter whenever care, money, consent, or a legal right is on the line.
Why this matters (evidence)

The languages your clients most need are often the ones these models have seen least online, and answer quality falls along that gradient. In one health-question benchmark, AI answers were about 5.82 times more likely to be incorrect in a non-English language than in English, and the answers grew less consistent as the language got lower-resource (Jin and colleagues, 2024).

Safety guardrails are weaker outside English. Across a large multilingual safety test, every model was significantly less safe when prompted in a non-English language, and the lowest-resource languages were the least safe (Wang and colleagues, 2024). A separate study found low-resource languages were roughly three times more likely to surface harmful content even with no attempt to trick the model (Deng and colleagues, 2023).

Machine translation of consequential documents carries real risk. Automated translation of emergency-department discharge instructions fell below 90 percent accuracy for several languages, including Amharic, and the errors were not predictable from the language (Carreras Tartak and colleagues, 2025). A review of AI-based communication with migrant patients found the most common tool had never been validated for patient safety and warned that a wrong translation can do more harm than no translation (Cronin and colleagues, 2025). When an AI tool rewrote consent forms to read more easily, the simpler version quietly dropped risk information (Oh and colleagues, 2025).

The cost of the gap is built into how these tools read text. Many non-English scripts take far more tokens to express the same message, so speakers of those languages can pay more, wait longer, and get answers that run short or cut off (Petrov and colleagues, 2023). The languages that already get the least reliable answers are the ones the tools handle least efficiently.

The alternative has evidence behind it. In a review of language-concordant care, most studies found at least one improved outcome for clients served in their own language (Diamond and colleagues, 2019). In a pediatric emergency department, using a professional interpreter was strongly associated with complete discharge education for families served in another language (Gutman and colleagues, 2018). A qualified interpreter is more reliable than any AI translation for encounters that carry weight.

The pattern

The languages your clients need most are where AI is weakest

Higher-resourceEnglish
Mid-resourceSpanish, Chinese, Arabic
Lower-resourceAmharic, Dari, Haitian Creole, Somali
AI learns from text, and there is far more of it online in some languages than others. Reliability falls as a language gets lower-resource, and that lower end is exactly where many of your clients are.
The line to hold

AI translation, or a qualified interpreter?

Low-stakes text

A flyer, an event notice, or the general gist of an email. A rough AI translation is acceptable, with a person finishing the job.

Anything consequential

Care, consent, a diagnosis, benefits, a legal right, or safety. Use a qualified interpreter. AI translation does not meet this obligation.

When care, money, consent, or a legal right is on the line, the tool is a qualified interpreter, not an app.
One more layer

The bias is cultural, not only linguistic

AI tools do not just struggle with languages other than English. They also carry a specific cultural point of view, and that shapes what they say about the communities you serve.

Cultural bias, not only language gaps

A model can translate a message correctly and still get a community’s values, family roles, or daily practices wrong. These systems default to Western, English-language assumptions even when asked about other cultures (Atari and colleagues, 2023; Cao and colleagues, 2023). A fluent answer in someone’s own language can still describe the wrong culture.

Stereotyped pictures, stereotyped facts

Ask an image generator for an everyday thing like a house or a family, and for non-Western regions it tends to return a narrow, often poverty-coded or exoticized image (Basu and colleagues, 2023; Bianchi and colleagues, 2023). The same flattening shows up in text, where a model repeats an outsider’s view of a culture and the way people inside it actually live goes missing (Naous and colleagues, 2024).

Long-tail knowledge: rare facts get generalized

A model learns a fact well only when that fact appears often in its training data, so accuracy falls sharply for facts that are rare or thinly documented online (Kandpal and colleagues, 2023; Mallen and colleagues, 2023). Communities with a small digital footprint sit in this long tail, so questions about them come back vague, generalized, or simply wrong.

For advocates, this is the stake

If you ask an AI a general question about a community and take its answer or recommendation at face value, you can reinforce the very stigma and prejudice you work against, because these systems reproduce the patterns in their training data, including its blind spots (Bender and colleagues, 2021). Treat any claim an AI makes about a specific community’s beliefs, customs, or needs as unverified, and confirm it with someone from that community or a real source. Holding that critical, verifying stance is part of advocacy.

The tool can draft and translate. The people you serve are the authority on their own culture.

Seen in the research

What this looks like in practice

  1. When researchers asked image generators for an everyday “house” across 27 countries, the pictures defaulted to a United States and then an Indian look, and for most countries they scored under 3 out of 5 for actually representing the place (Basu and colleagues, 2023).
  2. Prompted in Arabic about everyday life, models filled in Western-defaulted details that clash with the cultural context, such as following a mention of prayer with going for a beer (Naous and colleagues, 2024).
Case study

The discharge instructions

  1. A client who is more comfortable in Amharic comes to your desk after a hospital visit. She hands you a page of discharge instructions in English about a new medication and a follow-up appointment, and asks what she is supposed to do.
  2. The in-person interpreter has left and the phone line has a wait. It is tempting to paste the page into a translation app and read the result aloud so she can get home.
  3. This is a consequential encounter. A wrong dose or a missed warning could send her back to the hospital, and machine translation of discharge instructions has been shown to fall below 90 percent accuracy for languages including Amharic.
  4. Request a qualified interpreter and wait for the line, or set up a callback, before you go through the medication and the follow-up steps.
  5. Use teach-back through the interpreter. Ask her to say in her own words when she will take the medication and when the follow-up is, so you can confirm she understood.
  6. If you use an AI tool at all, hold it to a rough gist so you know the general topic while you wait, and never let the app's wording be the instruction she acts on.
Try this

Ask the same question in two languages

If bilingual staff are present, ask the same factual question in English and in another language, then compare the two answers side by side. Discuss which one you would trust and why, and where an interpreter would be the safer choice. If no bilingual staff are present, discuss a time an AI or machine translation missed something that mattered.

Practice together

Sort your documents by risk

In small groups, list the kinds of text your team translates in a typical week, from a flyer about an event to a consent form or a benefits denial notice. Sort each item into two piles: low-stakes text where a rough machine translation is acceptable, and consequential text where a wrong word could cost someone care, money, or a legal right. For everything in the second pile, name the interpreter service or the professionally translated version your agency will use instead. Post the two lists where staff can see them, and add a line for who to call when the right pile shows up at the desk.