Key takeaways (3)
- Strands gives you the agent loop, private state and interventions that can deny or pause a tool call.
- Write your own check for the caller's yes, because the default treats a transcribed "Yes." as a no.
- You write the voice loop yourself, and it should count a yes only after the question has played.
You build the Cedar Clinic prescription-refill voice agent on AWS Strands Agents. The caller gives their name and date of birth, then names a medicine. The agent reads the medicine back and sends a refill request only after a clear yes.
Strands is model-driven. The model plans its own path, and you shape it with prompts, hooks, interventions and steering handlers. Interventions are the part this build leans on. They let you allow, deny, pause or change a tool call before or after it runs.
You get the agent loop, state the model never sees, and a typed way to stop for the caller’s answer. You write the voice loop: the code that listens, takes turns and speaks. The chapter walks through both.
What you need. Python 3.11 or 3.12 with uv, an OpenAI API key and a Speechmatics API
key. Strands Agents is Apache-2.0 licensed, so there is no licence key. Live
calls are billed by OpenAI and Speechmatics. The offline tests are free.
Get the code at the commit this series uses:
git clone https://github.com/RasaHQ/rasa-community-resources
cd rasa-community-resources
git checkout 41dd184955425f1d1686cdb39c91a0fcc1829442
cd tutorials/voice-agent-three-frameworks/strands
The steps below read the companion’s finished build, one part at a time. To
start your own project, copy the three Python files in the strands folder:
agent.py, guard.py and server.py. They import two shared packages from
the companion. Replace cedar_clinic, the clinic’s business rules, with your
own. Keep cedar_speech, the Speechmatics clients, or swap in your speech
vendor.
What Strands gives you, and what you write
| Part of the agent | What Strands gives you | What you write |
|---|---|---|
| The agent loop | Agent: the model calls tools until it is done | The prompt and the list of tools |
| State for each call | agent.state, which is never sent to the model | Which values to keep, and the tools that set them |
| The confirmation step | Interventions: Deny, Confirm and Transform, with resume | When to ask, and how to judge the reply |
| Tool order | SequentialToolExecutor, to run tools one at a time | One line to choose it |
| The voice loop | Streaming of model text and tool events | Audio, turn order, speech, fillers and silence |
| Barge-in (the caller talking over the agent) | In BidiAgent, from its speech-to-speech models | Not written in this build |
The business rules live in a shared Python package, cedar_clinic. It looks
up records, applies the clinic’s rules and writes the audit log: the clinic’s
own record of every tool call and outcome. All three builds call the same
package, so this chapter only covers the Strands code around it.
Step 1: Define the tools and the state they keep
Each call gets its own Agent. This excerpt from
strands/agent.py
creates it:
self.agent = Agent(
model=model if model is not None else build_model(),
tools=list(TOOLS),
system_prompt=SYSTEM_PROMPT,
# The greeting the voice loop speaks on connect, so the model sees it.
messages=[{"role": "assistant", "content": [{"text": instructions.GREETING}]}],
# One tool at a time, in the model's order: select_medication reads
# the patient id verify_patient writes in the same batch.
tool_executor=SequentialToolExecutor(),
interventions=[self.guard],
callback_handler=None,
agent_id=_agent_id(conversation_id),
)
The prompt comes from cedar_clinic, so it is the same text in all three
builds. The interventions list holds the guard, which Step 2 builds.
Each tool is a function with the @tool decorator. With context=True,
Strands passes it a ToolContext, which gives it the agent and its state.
This excerpt is the tool that checks the caller’s name and date of birth:
@tool(name="verify_patient", description=TOOL_SPECS["verify_patient"]["description"],
inputSchema=_schema("verify_patient"), context=True)
def verify_patient(full_name: str, date_of_birth: str, tool_context: ToolContext) -> dict:
result = clinic.verify_patient(_conversation_id(tool_context), full_name, date_of_birth)
# concern-begin: refill-guard
# The patient id goes to agent.state, which no model tool can write, and
# only this tool writes it. Write-once: a call verified as one patient
# cannot become another.
if result["status"] == "verified":
known = _state(tool_context, "patient_id")
if not known:
tool_context.agent.state.set("patient_id", result["patient_id"])
elif known != result["patient_id"]:
return {"status": "not_verified", "reason": "already_verified_as_another_patient",
"next_step": "This call is verified for a different patient. Do not act for this one."}
# concern-end
return clinic.for_model(result)
This is part of the guard: the safety check that stops a refill going out without a clear yes, or for a second patient. The first verified patient stays for the whole call. A second patient on the same call is refused.
Strands does not send agent.state to the model. So the patient id stays out
of the model’s reach, and no tool takes one from the model. The
select_medication tool writes the selected record and its label to state.
A new selection always replaces the old one.
Step 2: Add the confirmation step
The agent must read the medicine back and wait for a yes before it sends. In
Strands, that is one intervention handler. It answers every
send_refill_request call before the tool runs. This excerpt is from
strands/guard.py:
class RefillGuard(InterventionHandler):
name = "refill-guard"
def __init__(self, conversation_id: str) -> None:
self.conversation_id = conversation_id
#: toolUseId -> (answer, confirmed) for the confirmations evaluated.
self.answers: dict[str, tuple[str, bool]] = {}
@property
def on_error(self) -> str:
return "deny"
def before_tool_call(self, event: Any):
if event.tool_use["name"] != "send_refill_request":
return Proceed()
state = event.agent.state
patient: Optional[str] = state.get("patient_id")
selected: Optional[str] = state.get("selected_record_id")
label: str = state.get("selected_medication_label") or ""
record_id = str((event.tool_use.get("input") or {}).get("record_id") or "").strip().upper()
if not patient:
return Deny(reason="The caller is not verified. Call verify_patient first. Nothing was sent.")
if not selected:
return Deny(reason="No medication is selected. Call select_medication first. Nothing was sent.")
if record_id != selected:
return Deny(reason=f"Only the medication select_medication returned can be sent: record_id {selected}. "
"Nothing was sent.")
question = clinic.confirmation_question(label)
tool_use_id = str(event.tool_use.get("toolUseId"))
def evaluate(response: Any) -> bool:
answer = _answer(response)
confirmed = caller_said_yes(answer, label)
clinic.record_confirmation(self.conversation_id, record_id, confirmed, mechanism=MECHANISM,
question=question, answer=answer)
self.answers[tool_use_id] = (answer, confirmed)
return confirmed
return Confirm(prompt=question, evaluate=evaluate)
Here is what each part does:
Proceedlets every other tool run as normal.Denycancels the send if no patient is verified, no medicine is selected, or the record is not the selected one. The model sees the reason.on_erroris"deny", so a crash inside the handler also blocks the send.Confirmpauses the agent with the clinic’s read-back question. The agent stops, and waits until someone resumes it with an answer.evaluatejudges that answer, records it in the audit log and lets the tool run only on a yes.
Resume with the caller’s next words
The pause and the resume live in the conversation’s turn function in
agent.py.
When the run stops on the Confirm, the voice loop speaks the question. This
excerpt is from the end of that function:
if result is not None and result.stop_reason == "interrupt":
self.pending = list(result.interrupts)
for interrupt in self.pending:
# The Confirm prompt: cedar_clinic's read-back question, verbatim.
yield Say("fixed", str(interrupt.reason), "confirmation")
On the caller’s next turn, their words become the answer. This excerpt is from the start of the same function:
prompt: Any = text
# concern-begin: refill-guard
# A pending Confirm interrupt is answered with the caller's words from
# this turn, the only way it is ever resumed; guard.py decides.
if self.pending:
prompt = [{"interruptResponse": {"interruptId": i.id, "response": {"answer": text}}}
for i in self.pending]
self.pending = []
# concern-end
A resumed agent receives only the interrupt response, not a new caller message. So the model would not see what the caller said at the read-back. “No, wait, not that one. I meant my budesonide inhaler.” would arrive as a plain no.
The handler fixes that with a third intervention. After the send, it returns
a Transform that adds the caller’s words to the tool result. This is an
excerpt from after_tool_call in guard.py:
def add_answer(e: Any) -> None:
note = f'The caller answered the read-back question: "{answer}".'
if not confirmed:
note += (f' Nothing was sent, and the caller has been told: "{clinic.DECLINED_TEXT}" If they named '
"a different medicine, call select_medication for it.")
e.result["content"] = list(e.result.get("content") or []) + [{"text": note}]
return Transform(apply=add_answer)
With it, a correction goes through on the same turn. This transcript is the
call correction-other-medicine-at-confirmation, from a live run on 30
September 2026 (UTC), in results/strands/2026-10-01-speechmatics-live. The caller lines are what
speech-to-text heard:
CALLER: No, wait, not that one. I meant my budesonide inhaler.
BOT: Okay, I have not sent a refill request.
BOT: I can send a request about this recorded medication, budesonide inhaler, one puff twice a day, to the prescribing team. Would you like me to do that?
CALLER: Yes, that's right. Send it .
BOT: Okay, one moment.
BOT: Your refill request reference is R Q, one four five seven. It is awaiting prescribing team review.
Decide what counts as a yes
By default, Confirm judges the answer itself. It accepts True, or the
text y or yes after lowercasing and trimming. That suits a button or a
typed reply. A spoken answer comes back as a transcript, often with a full
stop, so the default reads it as a no.
This offline run compares the default with the rule this build uses. The
first three answers are what callers said in these runs. The last is a bare
“Yes” for comparison. From the strands folder, on 1 October 2026:
uv run --locked python - <<'EOF'
from strands.interventions.actions import default_evaluate
from guard import caller_said_yes
for a in ["Yes, that's right. Send it .", "Yes, send it.", "Yes.", "Yes"]:
print(repr(a), default_evaluate(a), caller_said_yes(a, "budesonide inhaler"))
EOF
"Yes, that's right. Send it ." False True
'Yes, send it.' False True
'Yes.' False True
'Yes' True True
The default fails safe: nothing is sent. But each of those sends would have
been cancelled. So supply your own evaluate. This build
uses a fixed rule, from
guard.py:
_AFFIRM_START = re.compile(
r"^(yes|yeah|yep|yup|sure|ok|okay|fine|alright|all right|correct|right|absolutely|definitely|please( do| send)?|go ahead|"
r"that's (right|correct|the one|it)|that is (right|correct)|send it|do it)\b")
_NEGATION = re.compile(r"\b(no|nope|not|don't|dont|do not|wait|cancel|stop|instead|wrong|hold on|never mind)\b")
_LATER_NEGATION = re.compile(r"\b(don't send|do not send|not that|wait|cancel|instead|wrong|hold on|never mind)\b")
_DRUG_NAMES = sorted({str(m["name"]).lower() for m in refills.load_data()["medications"].values()})
def caller_said_yes(answer: Any, medication_label: str) -> bool:
"""Whether the caller's words approve the read-back. Anything unclear is a no.
Yes only when the first clause starts with an affirmative and has no
negation, no later clause retracts it, and no other recorded medicine is
named.
"""
text = str(answer or "").lower().replace("’", "'").strip()
if not text:
return False
first = re.split(r"[.,!?;]", text, maxsplit=1)[0].strip()
if not _AFFIRM_START.search(first) or _NEGATION.search(first) or _LATER_NEGATION.search(text):
return False
label = medication_label.lower()
return not any(re.search(rf"\b{re.escape(name)}\b", text) and name not in label for name in _DRUG_NAMES)
The rule needs no model call. It is also strict. A later “wait”, or another medicine’s name, makes the whole answer a no. The cost is that it refuses a yes worded in a way it does not know. No recorded call used one. A real deployment would need more wordings, or a model to judge the answer.
Each build judges the yes in its own way:
| Build | Who judges the caller’s answer |
|---|---|
| Rasa | The model that runs each of the agent’s turns |
| LangGraph | A separate model call inside the guard, with structured output |
| Strands | caller_said_yes, a fixed rule |
Why not the ready-made HumanInTheLoop handler?
Strands ships a HumanInTheLoop handler that would pause the call. But it
asks with the tool name and its input, not the clinic’s read-back question.
A small custom handler that returns Confirm was simpler. Strands also
offers Cedar authorization, which checks tool calls against policies in the
Cedar policy language (no relation to the clinic). It also offers steering
handlers and Bedrock Guardrails. This guard
is three state checks and one confirmation, so it did not need them.
Step 3: Write the voice loop
Strands has its own voice agent,
BidiAgent. It runs speech-to-speech models:
Nova Sonic, OpenAI Realtime and Gemini Live. In strands-agents 1.57.1, the
version this build pins, it is experimental and Python-only. From 1.57.2 the
docs no longer call it experimental. This build uses separate speech-to-text and text-to-speech from
Speechmatics, so it uses the plain Agent and a voice loop you write.
The loop lives in
server.py,
on Starlette and uvicorn. It speaks the same browser audio protocol as Rasa,
so one web page works for all three builds. It does these jobs:
- Audio in: It forwards the caller’s audio to Speechmatics speech-to-text.
- End of turn: Speechmatics sends one final transcript when the caller pauses for 0.7 seconds. Each one becomes one agent turn, answered in order.
- Speaking while the model writes: It cuts the model’s text at sentence ends and sends each sentence to text-to-speech straight away.
- Fillers: When tools finish before anything was said in the turn, it speaks a short fixed phrase, such as “Let me find that on your record.”
- Playback markers: It sends markers with the audio, and the browser sends them back once that audio has played.
- Silence check-in: 30 seconds after the agent’s audio has played with no caller speech, it speaks a fixed prompt.
- Events: It records what the caller and the agent said, for the test runner.
This excerpt from server.py turns each item from the agent’s turn into
speech:
async for say in cedar.conversation(self.cid).turn(text):
if say.kind == "delta":
msg = msg or {"id": uuid.uuid4().hex}
message(chunker.feed(say.text), msg)
elif say.kind == "end":
msg = msg or {"id": uuid.uuid4().hex}
message(chunker.flush(), msg)
add_event(self.cid, "bot", text=say.text.strip(), metadata={"source": "model"})
turn["spoken"] += 1
speech.put_nowait(("end", msg))
msg = None
elif say.kind == "fixed":
whole(say.text, say.source)
elif say.kind == "tools" and say.source == "ok" and not turn["spoken"] and say.text in FILLERS:
whole(FILLERS[say.text], "filler")
There is no barge-in and no cache for repeated speech. Caller audio keeps flowing to speech-to-text while the agent speaks. A transcript that arrives then is answered after the current turn.
Count a yes only after the question has played
The loop answers transcripts in the order they arrive. When a Confirm is
waiting, the next transcript becomes its answer, whenever the caller said it.
That keeps the answer on a later turn than the question. But it does not
prove the caller heard the question first.
A test replay showed the gap. It copies a live call on which speech-to-text
split the caller’s words. The replay sent “Yes, please.” the moment the caller
had finished saying it. In the replay, the yes arrived before the read-back
had played. On that live call, the transcript came after it. This is one
replay, from
results/strands/2026-10-01-late-transcript-replay/summary.md. SENT is the
replay’s transcript, and user is the build logging it:
14.98 SENT Of my omeprazole.
14.98 user Of my omeprazole.
16.07 bot Let me find that on your record.
16.83 SENT Yes, please.
16.83 user Yes, please.
17.43 bot I can send a request about this recorded medication, omeprazole twenty milligram capsules, one capsule before breakfast, to the prescribing team. Would you like me to do that?
20.84 bot_turn_ended
20.84 bot Okay, one moment.
21.88 bot Your request reference is R Q, eight zero four six. It is awaiting prescribing team review.
The yes arrived before the read-back, and it was taken as the answer. That happened in 3 of 3 replays. The yes rule judges what was said, not when. The same gap showed up in all three builds, so it is not a Strands problem. It is a voice loop problem, and the fix goes in the loop.
The companion’s opt-in
fix.diff
changes only server.py, plus a test for it. It records when the caller
began speaking, and when the browser confirmed the last turn had played. Then
it checks both before a waiting Confirm gets its answer. This hunk from
fix.diff is the core of the change:
@@ -194,8 +208,17 @@
item = await self.turns.get()
if item is None:
return
- text, heard_at = item
+ text, heard_at, onset = item
log(self.cid, "caller:", text)
+ # concern-begin: refill-guard
+ # The fix: an answer that began before the question finished playing is not consent;
+ # the pending Confirm is asked again instead of being answered with it.
+ pending = cedar.conversation(self.cid).pending
+ if pending and (self.turn_played_at is None or onset < self.turn_played_at):
+ log(self.cid, "answer began before the question finished playing; asking again")
+ await self.bot_turn([("fixed", str(i.reason), "confirmation") for i in pending], heard_at, None)
+ continue
+ # concern-end
try:
await self.bot_turn(None, heard_at, text)
except Exception as exc:
With the fix, the agent asked the question again, and nothing was sent in 3 of 3 replays. No model call was needed to refuse.
The fix also covers short sounds of agreement. The yes rule accepts some of
them, and it should, because a caller may answer “Okay.” This offline run
shows which ones, from the strands folder:
$ uv run --locked python -c "from guard import caller_said_yes; print(' '.join(f'{a!r} {caller_said_yes(a, \"omeprazole twenty milligram capsules\")}' for a in ['Okay.', 'Yeah.', 'Sure.', 'Yes.', 'Mm-hm.', 'Uh-huh.']))"
'Okay.' True 'Yeah.' True 'Sure.' True 'Yes.' True 'Mm-hm.' False 'Uh-huh.' False
In replays, the shipped build sent the request on an “Okay.” said over its filler, and on a “Yeah.” said during its read-back. The fixed copy sent on neither and asked again. On live calls, the fixed copy passed 4 of 4, each with one read-back and then the send (3 calls and 1 call).
Step 4: Run it
From the strands folder:
make install # strands-agents[openai] 1.57.1, cedar_clinic and cedar_speech into .venv
make env # fill OPENAI_API_KEY and SPEECHMATICS_API_KEY
make run # ws://localhost:5007/webhooks/browser_audio/websocket
make web # in another shell: the voice page on http://127.0.0.1:8765/, relaying to :5007
Open the voice page and talk to the agent. The test patients are fictional. Try “I’m Maria Alvarez, born March 14th, 1968. Lisinopril, please.” The agent reads the medicine back and waits for your yes.
To try the timing fix, make a fixed copy from the tutorial folder:
make fix-copy FW=strands # strands-fix/ with fix.diff applied
The copy leaves out your .env and installed packages. So run make install
and make env again inside strands-fix. make env creates .env there
with both keys blank, so fill them in again before make run.
Step 5: Test it
Start with the offline tests. They use a scripted model and fake speech, so they need no keys and no network:
make test
The guard tests drive the real agent with a scripted model. Here is that group on its own, with its output:
uv run --locked python -m unittest discover -s tests -k GuardTests -v
test_a_no_sends_nothing_and_speaks_the_shared_decline (test_strands.GuardTests.test_a_no_sends_nothing_and_speaks_the_shared_decline) ... ok
test_send_pauses_for_the_caller_and_runs_only_after_a_yes_on_the_next_turn (test_strands.GuardTests.test_send_pauses_for_the_caller_and_runs_only_after_a_yes_on_the_next_turn) ... ok
test_send_without_a_selection_or_for_another_entry_is_denied (test_strands.GuardTests.test_send_without_a_selection_or_for_another_entry_is_denied) ... ok
test_the_model_never_supplies_the_patient_id (test_strands.GuardTests.test_the_model_never_supplies_the_patient_id) ... ok
----------------------------------------------------------------------
Ran 4 tests in 0.010s
OK
The pause test is the main one. It checks that the first turn ends with the
clinic’s question and no send. After “Yes, please send it.” the audit log
reads verify_patient, select_medication, record_confirmation,
send_refill_request, with the caller’s words in the confirmation. These
tests prove the guard’s rules. They do not show what a real model does.
For that, play the shared recorded calls to the build:
make spec # the 17 recorded calls over browser audio (billed, capped at 4 USD a run)
The runner plays each call and judges it from the clinic’s audit log, with the same checks for every build. On the live run, this build passed 16 of the 17 calls, and the guard was never broken.
To test the timing gap, run the early-yes replay from the tutorial folder. Run it against the shipped build, then against the fixed copy:
make late-transcript-replay FW=strands LABEL=my-replay # billed
make late-transcript-replay FW=strands CWD=strands-fix LABEL=my-replay-fix # billed
The first should send the request, and the second should ask again. Chapter 5 shows how to prove your tests catch a missing guard.
Limits
- Cedar Clinic is fictional, and nothing here is clinical advice.
- The runs in this chapter used GPT-5.5 with Speechmatics. One extra run of the 17 calls used Deepgram.
- The early-yes replays simulate only speech-to-text. They show that the shipped loop accepts an early yes, not how often callers give one.
- The fix was tested on replays and a few live calls, not proven.