What AI can and cannot do in your daily work
Set realistic expectations for what AI does in day-to-day agency work, and learn why a chatbot can sound completely sure and still be wrong.
Map what you already use it for
- Everyone writes the five things they asked an AI tool to do at work in the last week, one per line.
- Next to each, mark whether you checked the answer against a source, or did not.
- Put the lines on the wall, group the ones that are the same task, and keep the wall as your agency's first map of what AI is actually being used for.
Which tasks on the wall could reach a client with nobody checking them?
What AI helps with, and what to check hard
- Drafting a first version of a letter, email, or notice
- Shortening a long document to its main points
- Translating the general gist of a note to get started
- A plain explanation of how a rule usually works
- A phone number, an address, office hours, or a deadline
- A dollar amount or an eligibility rule
- Anything recent, local, rare, or about a specific person
- Anything a client will act on, confirmed against the official source
- Name the everyday tasks AI genuinely helps with: drafting a first version of a letter or email, shortening a long document to its main points, translating the general meaning of a note, and getting a plain explanation of how a rule usually works.
- Explain in plain words why a chatbot invents things, a mistake often called a hallucination, meaning confident text with no real source behind it. It builds fluent sentences by predicting likely words, so it can produce a clinic name, a phone number, a filing deadline, or an eligibility rule that does not exist.
- Treat made-up answers as common enough to expect, not rare accidents. On factual lookups such as legal citations and references, studies have measured wrong or fabricated answers in anywhere from a quarter to more than half of cases.
- Recognize that an invented answer reads exactly like a correct one, in the same calm and confident tone, so smooth wording is never a sign that a detail is true.
- Build the one habit that protects clients: check anything that would change what a client does, such as a number, a date, an amount, an address, or an eligibility rule, against the official source before anyone acts on it.
- Notice a quiet risk in yourself and your coworkers. People tend to accept a fluent, confident answer without checking it, a pattern researchers call automation bias, meaning trusting the machine more than the evidence warrants.
- See that clients, including immigrant clients, already use these tools in daily life, so the practical goal is guiding safe use early instead of trying to prevent it.
When a chatbot is used for factual lookup, invented answers are common. Testing leading models on legal questions, Dahl and colleagues, 2024 found false or fabricated information in most answers, from 58 percent to 82 percent depending on the model. Asked to supply references, chatbots fabricate those too. In one comparison Chelli and colleagues, 2024 found one tool invented citations more than 90 percent of the time and two others between about 29 and 40 percent.
The harder problem is us. People using computer-generated advice tend to accept it even when it is wrong and other information in front of them contradicts it, a pattern a systematic review named automation bias, meaning trusting the tool more than the evidence supports (Goddard and colleagues, 2012). A small amount of friction helps. Making people form their own answer before they see the AI's suggestion reduced this overreliance (Bucinca and colleagues, 2021).
The strengths are real when the task is drafting or condensing instead of supplying facts. In a study where physicians compared machine-written and expert-written clinical summaries, summaries from an adapted model were rated equal to or better than the experts' in most cases (Van Veen and colleagues, 2024). Even there the authors still found occasional errors, so a person reads the summary against the original before relying on it.
Clients are already using these tools. About 1 in 3 US adults report using AI chatbots to look up health information, and many are not confident the answers are accurate (KFF Health Information and Trust polls, 2024-25). Immigrant clients may be especially open. In a study of immigrant and general-population adults, immigrants reported higher willingness to use AI mental-health chatbots (Yoo and Jang, 2026). That openness is a reason to guide safe use early.
Where its knowledge and its blind spots come from
A chatbot is not connected to the truth. It is a compressed picture of the text it was trained on. Knowing what went into that picture tells you where to trust it and where to check.
The model's trained memory is a fixed picture of text collected up to a cutoff date, so on its own it does not know what came later, and it does not always look things up by itself. Whenever you need something time-sensitive, such as a current rule, a deadline, an income limit, or today's phone number, tell the tool to search the web for it and to show the sources it used. A live search helps, and it still carries risk, because a reply is only as good as the pages it finds and it can misread them, so confirm anything recent or local against the official source either way.
The model is strongest where its training text was abundant and repeated, such as general health explanations in English. It is weakest where the text was thin, such as local procedures, small programs, your clients' languages, and anything specific to one county or clinic. The gaps are not random. They track what is underrepresented online.
The training text carries human decisions and inequities, so the model can reproduce them. In one well-documented case, a widely used health algorithm looked fair but was biased because it used cost as a stand-in for illness, which understated the needs of Black patients (Obermeyer and colleagues, 2019). AI can be confidently wrong in ways that track who the client is.
The model repeats what is common in its data whether or not it is correct, so a popular myth can come back as confidently as a fact. This is why a source you trust, and not the model's tone, is what settles whether something is true.
Treat the tool as a mirror of a large amount of past human text. That is why it is useful for general explanations and drafting, and why anything recent, local, rare, or about a specific person needs a check against a real source.
A made-up answer looks exactly like a correct one
The county clinic on Main Street takes walk-ins until 5 pm. Sliding-scale fees start at $15. Call 734-555-0143.
The county clinic on Main Street takes walk-ins until 7 pm. Sliding-scale fees start at $10. Call 734-555-0198.
A confident answer with a wrong number
- A benefits navigator asks a chatbot for the local Medicaid office's phone number and the current income limit for a family of four, and the chatbot answers right away with a specific number and a specific dollar figure.
- The answer is fluent and sounds certain, so it is tempting to read it straight to the client who is waiting.
- Following the checking habit, the navigator treats both details as unverified and looks them up on the state agency's official page.
- The phone number turns out to be out of date and the income figure is for a different household size, both invented in a way that looked exactly like a correct answer.
- The navigator gives the client the verified number and amount, and uses the chatbot only to draft a plain-language summary of the next steps, which she reads over before sharing.
Find the made-up detail in an AI answer
Read an AI answer aloud that names a clinic, gives a phone number, and states an eligibility rule, exactly as a client might receive it. Ask the group what they would check before acting on it. Then reveal that the phone number and the age rule were invented and sound identical to a correct answer. Each person says one sentence they would use to check a specific detail with a client.
Sort your own tasks by whether AI can be trusted with them
Ask each person to write down three tasks from their past week where they might use a chatbot. As a group, sort every task into two piles. The first pile is tasks where AI helps with wording or shortening, such as drafting a reminder email, turning a long policy into plain language, or translating the general meaning of a note for a colleague. The second pile is tasks where AI would supply a fact a client acts on, such as an office phone number, a filing deadline, an income limit, or an eligibility rule. For each task in the second pile, name the official source you would check it against. Close by noticing that the first pile is where AI saves real time, and the second pile is exactly where a smooth, confident answer is most likely to mislead.