BETA This site is in beta. Information is still being added and reviewed.
Practitioner hub
Practitioner guide · AI literacy

Client data and AI: a privacy quick-guide

Consumer AI tools were not built to hold sensitive client information. Typing a client’s details into one, or letting an AI assistant open and read a client’s file, is not the same as making a private note. It is a data-sharing act, and for an immigrant or refugee client the consequences can be serious and sometimes impossible to undo. This guide is a deeper look at the workshop module on what never to put into an AI tool. Print it, keep it near your desk, and use the checklist before you open a chatbot.

Why the stakes are higher

Immigrant and refugee clients already weigh whether contact is safe

That caution is not paranoia; it is documented. Friedman and Venkataramani (2021) matched national survey data to Immigration and Customs Enforcement activity and found that Hispanic adults were less likely to report a usual provider or an annual checkup after ICE activity rose in their state, while non-Hispanic adults showed no such change. Young and colleagues (2023) found that among Latinx and Asian immigrants in California, each additional direct encounter with the enforcement system raised the odds of delaying care (odds ratio 1.30, 95% CI 1.10 to 1.50). Herring and Barnow (2025) estimated that the Secure Communities program was followed by a 16.9 percent decline in provider visits among lawfully present older Hispanic immigrants, a group not personally at risk of deportation; that pattern points to fear and spillover as the driver, not direct removal.

The lesson for AI use is direct. A client’s immigration status, A-number, or case detail is not ordinary contact information. If it leaks, is subpoenaed, or surfaces in a breach, it can expose the client or their household to enforcement, denial of relief, or harm in a country of origin, and those outcomes cannot always be reversed. Treat every entry of client information into an outside tool as carrying that weight.

What actually happens

Typing it in is sending it out

When you type text into a chatbot, paste in a document, or let an AI assistant open and read a client file, that content leaves your computer and travels to the vendor’s data center, where it is processed on their servers. This is true whether you type it yourself or an assistant reads the file for you: having the tool read a file is itself a transfer of that file’s contents to the company. It is not a private note. It is a data-sharing act with a company you do not control.

Two different things can happen next, and they are worth separating. “Stored” means the vendor keeps a copy of your input for some retention period, during which it can be read by the company’s systems, reviewed by staff or contractors under some terms, exposed in a breach, or produced in response to a subpoena. “Trained on” means your input is used to improve the model itself, so it becomes part of what the model learns from. A tool can promise not to train on your data and still store it, so you need to know both answers, not just one.

Why memorization matters

Models can memorize and repeat their training data

This is not hypothetical. Carlini and colleagues (2021) showed that a language model can be made to emit verbatim chunks of its training data, and among what they extracted from GPT-2 were real personally identifying details, including names, phone numbers, and email addresses. Carlini and colleagues (2023) then showed that this memorization grows with the size of the model, with how many times a piece of text was duplicated in training, and with how much surrounding context an attacker supplies.

The takeaway for practitioners is that information fed into a tool that trains on inputs is not only seen by the vendor now; it can later be surfaced to someone else. That is the extraction risk, and it is one reason a written no-training clause is worth insisting on before you use a tool for anything sensitive.

What protects you, and what does not

Encryption protects you from outsiders, not from the vendor

Reputable tools encrypt data in transit, so it cannot be read while crossing the network, and at rest, so stored files are not readable if a disk is stolen. Encryption protects against outsiders. It does not protect you from the vendor, because the vendor decrypts the content in order to process it. Encryption is not confidentiality from the company holding the data.

Most AI tools you reach through a browser or app are cloud tools: the data leaves your control and goes to a third party. A local or on-device model runs on your own machine and does not send the content out, a fundamentally different and generally safer risk profile. The default also depends on the product tier. Free consumer tools are the riskiest by default; OpenAI states that consumer ChatGPT conversations are used to train its models by default unless the user turns training off, while its Business, Enterprise, Team, Edu, and API products are not used to train models by default (OpenAI, accessed 2026). Terms change and every vendor differs, so never assume, and remember that the free version you reach by default is usually the one with the weakest protections.

Where your data goes

The path is the same whether you typed the message or asked an assistant to read a file.

  1. You type or uploadYou type a message, paste a document, or point an AI assistant at a client file to summarize or read.
  2. It leaves your deviceThe content travels over the internet to the vendor, whether you typed it or an assistant read the file for you.
  3. It reaches the vendor’s data centerThe vendor’s servers process it. Staff or contractors may be able to access it under some terms.
  4. It is stored, and maybe used to trainThe vendor may keep a copy for a retention period, and separately may use it to improve future models unless training is turned off.

Encryption protects this trip from outsiders. It does not stop the vendor at the other end from reading what you sent.

The “never enter” list

Do not put any of the following into a consumer or non-approved AI tool, whether by typing it or by asking the tool to read a file that contains it.

Name, date of birth, address, phone, or email

Any one of these can point to a single person, and combined with anything else in the message it narrows further.

Social Security number

A single leaked SSN enables identity theft and fraud that can follow a client for years.

A-number, USCIS receipt or case number, court docket number

These tie a message directly to an immigration or legal case file and to enforcement systems.

Immigration or citizenship status

Status and history are exactly the details enforcement action turns on; they should never sit on a vendor’s server.

Health, mental health, substance use, or disability information

This is protected clinical information, and a breach or subpoena can expose a client’s most private history.

Legal case, arrest, or criminal history details

Case details can surface in litigation or be used against a client in ways they never agreed to.

De-identification and its limits

Removing the name does not make a record anonymous

De-identification means stripping information down so it can no longer be traced to a person: removing not just the direct identifiers such as names, numbers, and addresses, but the quasi-identifiers that combine to point at someone. Replacing "Maria, DOB 3/12/1990, A# 087..." with "a client" is a start, not the finish.

The limits are the important part. Rocher, Hendrickx, and de Montjoye (2019), using a method that holds up even when a data set is incomplete, estimated that 99.98 percent of Americans could be correctly re-identified from any data set using just 15 demographic attributes. Small sets of ordinary attributes are enough to single someone out, so de-identification is never guaranteed, and it fails fastest exactly where your clients live: rare languages, specific countries of origin, small towns, and uncommon case profiles.

Free-text narratives are especially leaky because they carry incidental details a checklist would never think to remove. The safe practice is not to de-identify a real record and paste it in. Work in generic, hypothetical, or aggregated terms instead, the way you would ask a question in a training, without attaching it to an actual person. If you would not read the sentence aloud in a crowded public place with the client present, do not enter it.

Can this go into this tool?

Walk these gates in order. A stop at any gate means stop.

  1. Is the tool on your agency’s approved list, with a signed data agreement?If no, stop. Use an approved tool or none.
  2. Does the content contain anything from the “never enter” list, including inside an attached file?If yes, stop, or remove and generalize it first before continuing.
  3. Could the remaining details, in combination, identify one specific client, family, or small community?If yes or unsure, stop. Assume small-population details are identifying.
  4. Do you know this tool does not train on your inputs and has a stated retention limit?If no, stop until you can confirm it.
  5. If this exact text became public tomorrow, could it harm the client or their household?If yes or unsure, stop.

Only if you clear every gate should you proceed, and even then enter the minimum necessary.

What to require before your agency adopts a tool

Adopting an AI tool for work near client information is an organizational decision, not an individual one.

A written data agreement (BAA or DPA)

HHS requires a Business Associate Agreement with any vendor that creates, receives, maintains, or transmits protected health information on your behalf; a cloud service that stores or processes it counts. For other sensitive data, require the equivalent data processing agreement.

A documented training opt-out

The vendor’s terms should state in writing that your inputs are not used to train their models, and someone should confirm the setting is actually configured that way, not merely available.

Clear retention terms

Know how long inputs are kept, whether you can request deletion, and whether a zero-retention or no-log option exists. Prefer the shortest retention you can get.

Organizational approval and an approved-tools list

Individual staff should not decide tool by tool. There should be a maintained list of what is allowed for what kind of content, and a named person to ask.

Vendor security controls

Encryption in transit and at rest, access controls, and breach-notification commitments are necessary, but they protect against outsiders and do not substitute for the training and retention terms above.

Quick checklist

Sources

  1. Friedman, A. S., & Venkataramani, A. S. (2021). Chilling Effects: US Immigration Enforcement and Health Care Seeking Among Hispanic Adults. Health Affairs, 40(7), 1056–1065. doi:10.1377/hlthaff.2020.02356
  2. Young, M.-E. D. T., Tafolla, S., Saadi, A., Sudhinaraset, M., Chen, L., & Pourat, N. (2023). Beyond “Chilling Effects”: Latinx and Asian Immigrants’ Experiences With Enforcement and Barriers to Health Care. Medical Care, 61(5), 306–313. doi:10.1097/MLR.0000000000001839
  3. Herring, J., & Barnow, B. (2025). Indirect effects of immigration enforcement on health care utilization among lawfully present older Hispanics. Social Science & Medicine, 384, 118540. doi:10.1016/j.socscimed.2025.118540
  4. Carlini, N., Tramèr, F., Wallace, E., Jagielski, M., Herbert-Voss, A., Lee, K., Roberts, A., Brown, T., Song, D., Erlingsson, Ú., Oprea, A., & Raffel, C. (2021). Extracting Training Data from Large Language Models. 30th USENIX Security Symposium (USENIX Security 21). link
  5. Carlini, N., Ippolito, D., Jagielski, M., Lee, K., Tramèr, F., & Zhang, C. (2023). Quantifying Memorization Across Neural Language Models. The Eleventh International Conference on Learning Representations (ICLR 2023). link
  6. Rocher, L., Hendrickx, J. M., & de Montjoye, Y.-A. (2019). Estimating the success of re-identifications in incomplete datasets using generative models. Nature Communications, 10, 3069. doi:10.1038/s41467-019-10933-3
  7. OpenAI (2026). How your data is used to improve model performance; Enterprise privacy. OpenAI help center and enterprise privacy page (accessed 2026). link
  8. U.S. Department of Health and Human Services (2026). Business Associate Contracts; Guidance on HIPAA & Cloud Computing. HHS.gov (accessed 2026). link

This guide describes practice grounded in the studies listed above. Adapt it to your agency’s policies and to the languages and communities you serve.