Before your team adopts an AI tool, run it through a short checklist and get clear answers before any client information goes near it.
How to choose
Match the task to the tool, then to the data rule
The taskA fitting toolA quick rough draft or summaryA free, lightweight chatbot is fineCareful reasoning or a long documentA stronger paid model earns its placeAnything with sensitive client dataNo consumer tool at all
The right tool follows the task and the sensitivity of the data, not the brand name; nothing that could identify a client goes into a consumer tool at any tier.
Data policy: ask what the tool collects, where it stores it, and who it shares it with, and get the answer in writing.
Training: ask whether it trains on your inputs and whether you can turn that off, and look for a business agreement that says it will not.
Organization approval: confirm the tool is allowed under your agency's policy before anyone puts client information near it.
Cost: work out the real price at your scale, including per-user and per-use fees that stay small in a demo and grow with a whole team.
Languages and accessibility: check that it works for the languages and access needs of your clients and staff, since a tool strong in English can be weak in the languages you serve.
Portability: make sure you can export your data and take it with you if you leave the tool.
A slick demo or a high benchmark score is a sales signal, so pilot the tool on low-stakes work and judge it on your own tasks before you trust it.
How it works, and where it breaks
1A benchmark score measures one test, and the test may not be your job
Tools get ranked on benchmarks, which are fixed sets of questions with known answers. A high score means the tool did well on that set. Two things keep a score from telling you what you need. The benchmark questions can end up inside the tool's training data, so the tool has effectively seen the answers already and the score is inflated. And the benchmark task can look nothing like your intake letters or your eligibility questions, so a strong score on it says little about your work. Read a score as a hint, and test the tool on tasks you actually do.
2Fluent output can be confidently wrong, so a demo proves little
A tool can produce smooth, professional writing and still be wrong, because reading well and being right are separate things. A demo is staged to show the tool at its best on material the vendor chose. The way to learn whether a tool fits is to run your own real, low-stakes tasks through it, with client details removed, and check the results against the source. Ask the data and training questions in the same round, because a tool that will not answer them plainly should not receive client information whatever the demo looked like.
3A tool can look fine and still cause harm at scale
Surface performance can hide who a tool serves poorly. A widely used health-care algorithm was found to under-refer Black patients for extra care because it used past spending as a stand-in for need, and it appeared to be working correctly the whole time it was in use (Obermeyer and colleagues, 2019). The lesson for buying an AI tool is that a clean-looking output does not reveal how the tool treats the people you serve. Ask about data handling and training, get the promises in a business agreement, and pilot on low-stakes work where a mistake is cheap to catch.
A tool your team is weighingA filled-in adoption checklist to review
Turn the decision into an adoption checklist
I am checking whether our agency should adopt the AI tool I name. Build a short checklist with a row for data handling, training on inputs, whether there is a business agreement, organization approval, real cost at our size, language and accessibility coverage, and data export. For each row, write the plain question I should ask the vendor and leave a blank for their answer. [name the tool and roughly how many staff would use it]
Guardrail. The tool cannot confirm the vendor's answers for you, so get each one in writing before client data goes in. See Module 4.
A vendor's data or privacy pageA plain read of what it does and does not promise
Get a plain read of a vendor's data terms
Here is a vendor's data or privacy page. Put it in plain language: what does it say happens to text I paste, does it train on my inputs, can I turn that off, and does it promise any of this in a business agreement. List anything the page is vague about or does not answer, so I can ask the vendor directly. [paste the vendor's data or privacy text]
Guardrail. A plain read is a starting point; confirm the promises in a signed agreement before client information goes near the tool. See Module 4.
Worked example
The demo that fell apart on a real task
A vendor shows your team a polished demo where the tool rewrites a benefits notice into clean plain language, and it looks ready to buy.
Before deciding, you pilot it on ten of your own real notices with the client details removed, and you check each output against the source.
On your own material the tool reads just as smoothly, and on three of the ten it drops a deadline or changes an eligibility rule while sounding completely sure of itself.
You also ask for the data terms and learn the plan you were shown trains on your inputs with no business agreement to turn that off. You pass on it, or move to a tier that answers the data questions in writing. The demo showed the tool at its best, and your own low-stakes pilot is what showed how it behaves on the work you actually do.
Guardrail. If a tool cannot answer the data and training questions clearly, do not put client information into it. See Module 4.
Sources
Obermeyer, Z., Powers, B., Vogeli, C., & Mullainathan, S. (2019). Dissecting racial bias in an algorithm used to manage the health of populations. Science, 366(6464), 447-453. https://doi.org/10.1126/science.aax2342