Fidelity kit
A simple, layered way to keep delivery consistent: a checklist every session, a small sample of sessions rated more closely, a monthly supervision review, and a booster that targets the skills that fade fastest.
The accepted approach to keeping lay-delivered programs on track is layered: a self-report checklist every session, a sampled share of sessions rated more closely, and trends reviewed in supervision (Breitenstein and colleagues). Community health worker programs for immigrant and Latina populations have used exactly this session-checklist plus sampled-rating design (the ALMA and PARTNER-MH studies).
Boundary and "do-not-do" skills decay faster than anything else (Gobezayehu and colleagues, 2014), and inoculation-style spot-the-error skills fade within about two months without a booster (Maertens and colleagues). That is why the booster below targets those two things specifically.
The checklist and sampled ratings run at every step. The booster, first at six to eight weeks and then each quarter, is where the skills that fade fastest get refreshed.
1. Per-session activity checklist
The CHW completes this after every client session. It takes a couple of minutes and becomes the base layer of the fidelity record.
- Consent obtained; the session was voluntary and the client understood that.
- Working language confirmed; interpreter arranged if needed.
- Explained what AI can and cannot do, including the spot-the-error example.
- Ran hands-on practice on the client's own phone (find care and understand a document).
- Coached calibrated trust: guess first, drafts, expect "I'm not sure."
- Used teach-back at least once to confirm understanding.
- Covered the privacy rule and the interpreter right.
- Client left with a written action plan in their language and the follow-up contact.
- Stayed in scope: no diagnosis, medication, legal, or medical advice given.
- Any red flags handled per the escalation card; escalations logged with time.
- AI tools used this session logged.
2. Audio-rating sampling protocol
A small, random share of sessions is recorded and rated more closely than self-report allows. Keep it simple.
- Consent and recording. With the client's informed consent, and only where your site permits, audio-record the session.
- Sample size. Rate 10 to 20 percent of each CHW's sessions, chosen at random rather than hand-picked, so the sample reflects typical delivery.
- Rate against the checklist. A trainer or peer rates the recording using the observation checklist, with attention to the safety items.
- Reliability check. Have two raters score a subset (for example a quarter of the sampled recordings) and compare, so ratings stay consistent between reviewers.
- Store safely. Keep recordings secure, use them only for supervision and quality, and delete them per your site's retention policy.
3. Monthly supervision review
Once a month, the supervisor reviews the record with the CHW. Use this template.
- Checklist completion trend. Are all activities happening consistently, or are some slipping?
- Audio-rating scores. How did sampled sessions score on the observation checklist this month?
- Do-not-do drift. Look specifically for boundary items sliding: giving medication or legal opinions, treating AI output as advice, letting AI stand in for an interpreter. These fade first.
- Escalations. Were red flags recognized and routed correctly and on time? Review each one.
- Teach-back and comprehension. Are clients able to show understanding, or is teach-back being skipped?
- Client feedback and access barriers. Language, childcare, transportation, timing.
- Action items. One or two concrete things to practice or fix before next month.
4. Booster schedule
Skills fade, and the ones that fade fastest are exactly the safety-critical ones. The booster is short and focused, not a full re-training.
- Timing. First booster at 6 to 8 weeks after a CHW begins solo delivery, then about every quarter.
- Length. 60 to 90 minutes.
- Focus 1: do-not-do items. Re-run one boundary role-play (medication out of scope, or crisis disclosure) and rescore the safety items.
- Focus 2: spot-the-error practice. Work a fresh weakened-dose example so the checking habit stays sharp.
- Review the month. Bring any escalations or drift from supervision into the booster discussion.
- Is a session checklist actually completed for every session, not just the observed ones?
- Is the audio sample random rather than hand-picked, so it reflects real delivery?
- Did this month's review look specifically for do-not-do drift, not just overall scores?