01What must never happen
Ask a clinic manager what a scheduling assistant must never do and you get a short list:
- Book a time that doesn’t exist, or one another patient already has.
- Read or change the wrong patient’s appointment.
- Book something the caller didn’t agree to.
- Keep talking about scheduling while someone describes chest pain.
Left alone, a language model will eventually do all four. It offers a time that sounds right. It matches “Lopes” to the nearest “Lopez”. It hears “yes” at the start of “yes, wait, actually”. A better prompt won’t fix these. Each one marks a decision that has to move out of the model and into code.
02Where the model sits
The model is one step in the middle of the call. Everything on either side of it is ordinary code.
- Caller speaksSpeech-to-text turns it into words.
- Red-flag checkcodeRuns on every turn before the model. On a match it reads an emergency script and alerts staff. The model is never called.
- ModellanguageWorks out what the caller wants and picks a tool to call.
- GuardrailscodeRun the four checks below. A refusal goes back to the model with the reason, and the model tries again.
- Calendarsource of truthThe only thing that knows what’s free and what’s booked.
- Agent repliesText-to-speech reads the model’s answer to the caller.
Because the agent works in text, voice is just an adapter at each end. Swapping in a different speech service touches those two ends and nothing else.
03Four checks, all in code
- Who is calling
Nothing about a patient is read or changed until a name and date of birth match a record. The birthday must match exactly. The name may be slightly off, because speech-to-text misspells names.
- Only times the calendar returned
The agent can propose a time only if a calendar search returned it during this call, and the calendar checks it again at the moment of writing. The model can’t invent one.
- A clear yes, on a later turn
Each change is staged and read back. It is written only after the caller answers. Code reads the whole answer, so “yes, wait, actually” doesn’t count as a yes.
- Ask when it’s ambiguous
If more than one patient could match what was heard, the agent asks for a spelling. It doesn’t pick the closest name, and it never says who else matched.
The red-flag check runs before all four. It is a keyword list, tuned to catch everything at the cost of some false alarms. A false alarm costs one transfer to staff. A miss can cost a life.
04Five calls, replayed
Pick a call and step through it. Each turn shows what the caller said, what code checked, what the agent said back, and whether the calendar changed. Keep an eye on the calendar line.
The model’s side of each call is scripted, so the replay is the same every time. Everything else is the agent’s real code. Every tool call ran through the real guardrails and calendar, and the refusals and calendar changes shown are exactly what that code returned.
05Grade the calendar, not the transcript
A transcript can sound perfect while the calendar is wrong. So after each simulated call, the evaluation reads the calendar itself and tracks five numbers:
- Booking completion. The caller wanted a change, and the calendar shows it.
- Wrong or double bookings. Must be zero. This is the hard gate.
- Escalation recall on urgent calls. A miss is the expensive error, so catching every urgent call matters more than avoiding false alarms.
- Turns, and p95 latency per turn. It’s a phone call, and dead air is a bug.
- Handoff rate. Too high and the agent is useless. Too low and it’s overreaching.
Beneath the evaluation sit 105 tests that run offline with no API key. They replace the model with a script, so a test can make the model misbehave on purpose (invent a slot, confirm too early, loop forever) and check that code catches it.
06What it doesn’t do yet
- One process, in memory. Two callers can be offered the same 9:30. The calendar re-checks at write time, but a real deployment needs a database constraint on doctor and start time, because only the database can settle that race.
- English keywords for red flags. The list is tuned to catch everything, so “my dad had a stroke last year” also gets a transfer.
- One clinic, one time zone. No interpreters, insurance checks or multi-slot visits.
Takeaways
- Let the model handle the language. Any decision that is expensive to get wrong goes into code the model calls.
- A time exists only if the calendar returned it during this call.
- Write on a later turn, after a yes that code has judged clear.
- When two records fit, ask. A confident guess opens the wrong patient’s chart.
- Catch emergencies before the model runs.
- Grade the calendar, not the transcript.
Related reading
Your agent isn’t dumb. Its tools are. · We gave agents tools but no undo button · The last click
On the replays. Sunrise Family Clinic, its doctors and its patients are invented. The clinic clock in every replay is fixed at Tuesday, September 22, 10:00 AM. The model’s turns are scripted. Every tool result, refusal and calendar change was produced by running the agent’s code, and the export script asserts each behaviour this page describes, so the export fails if the code stops behaving that way. The name scores in the twins call (0.86 for Ana, 0.91 for Anna) are the same Python difflib ratios the calendar’s matcher uses. This page doesn’t report how a live model performs.