Skip to content
RasaGet a free licence

tutorial

Chapter 4 of 6

Build the voice agent on Strands Agents

by Rod Rivera Published

Build the Cedar Clinic prescription-refill voice agent on Strands Agents, from its tools and confirmation step to the voice loop, a live call and tests.

Key takeaways (3)
  • Strands gives you the agent loop, private state and interventions that can deny or pause a tool call.
  • Write your own check for the caller's yes, because the default treats a transcribed "Yes." as a no.
  • You write the voice loop yourself, and it should count a yes only after the question has played.

You build the Cedar Clinic prescription-refill voice agent on AWS Strands Agents. The caller gives their name and date of birth, then names a medicine. The agent reads the medicine back and sends a refill request only after a clear yes.

Strands is model-driven. The model plans its own path, and you shape it with prompts, hooks, interventions and steering handlers. Interventions are the part this build leans on. They let you allow, deny, pause or change a tool call before or after it runs.

You get the agent loop, state the model never sees, and a typed way to stop for the caller’s answer. You write the voice loop: the code that listens, takes turns and speaks. The chapter walks through both.

What you need. Python 3.11 or 3.12 with uv, an OpenAI API key and a Speechmatics API key. Strands Agents is Apache-2.0 licensed, so there is no licence key. Live calls are billed by OpenAI and Speechmatics. The offline tests are free.

Get the code at the commit this series uses:

git clone https://github.com/RasaHQ/rasa-community-resources
cd rasa-community-resources
git checkout 41dd184955425f1d1686cdb39c91a0fcc1829442
cd tutorials/voice-agent-three-frameworks/strands

The steps below read the companion’s finished build, one part at a time. To start your own project, copy the three Python files in the strands folder: agent.py, guard.py and server.py. They import two shared packages from the companion. Replace cedar_clinic, the clinic’s business rules, with your own. Keep cedar_speech, the Speechmatics clients, or swap in your speech vendor.

What Strands gives you, and what you write

Part of the agentWhat Strands gives youWhat you write
The agent loopAgent: the model calls tools until it is doneThe prompt and the list of tools
State for each callagent.state, which is never sent to the modelWhich values to keep, and the tools that set them
The confirmation stepInterventions: Deny, Confirm and Transform, with resumeWhen to ask, and how to judge the reply
Tool orderSequentialToolExecutor, to run tools one at a timeOne line to choose it
The voice loopStreaming of model text and tool eventsAudio, turn order, speech, fillers and silence
Barge-in (the caller talking over the agent)In BidiAgent, from its speech-to-speech modelsNot written in this build

The business rules live in a shared Python package, cedar_clinic. It looks up records, applies the clinic’s rules and writes the audit log: the clinic’s own record of every tool call and outcome. All three builds call the same package, so this chapter only covers the Strands code around it.

Step 1: Define the tools and the state they keep

Each call gets its own Agent. This excerpt from strands/agent.py creates it:

        self.agent = Agent(
            model=model if model is not None else build_model(),
            tools=list(TOOLS),
            system_prompt=SYSTEM_PROMPT,
            # The greeting the voice loop speaks on connect, so the model sees it.
            messages=[{"role": "assistant", "content": [{"text": instructions.GREETING}]}],
            # One tool at a time, in the model's order: select_medication reads
            # the patient id verify_patient writes in the same batch.
            tool_executor=SequentialToolExecutor(),
            interventions=[self.guard],
            callback_handler=None,
            agent_id=_agent_id(conversation_id),
        )

The prompt comes from cedar_clinic, so it is the same text in all three builds. The interventions list holds the guard, which Step 2 builds.

Each tool is a function with the @tool decorator. With context=True, Strands passes it a ToolContext, which gives it the agent and its state. This excerpt is the tool that checks the caller’s name and date of birth:

@tool(name="verify_patient", description=TOOL_SPECS["verify_patient"]["description"],
      inputSchema=_schema("verify_patient"), context=True)
def verify_patient(full_name: str, date_of_birth: str, tool_context: ToolContext) -> dict:
    result = clinic.verify_patient(_conversation_id(tool_context), full_name, date_of_birth)
    # concern-begin: refill-guard
    # The patient id goes to agent.state, which no model tool can write, and
    # only this tool writes it. Write-once: a call verified as one patient
    # cannot become another.
    if result["status"] == "verified":
        known = _state(tool_context, "patient_id")
        if not known:
            tool_context.agent.state.set("patient_id", result["patient_id"])
        elif known != result["patient_id"]:
            return {"status": "not_verified", "reason": "already_verified_as_another_patient",
                    "next_step": "This call is verified for a different patient. Do not act for this one."}
    # concern-end
    return clinic.for_model(result)

This is part of the guard: the safety check that stops a refill going out without a clear yes, or for a second patient. The first verified patient stays for the whole call. A second patient on the same call is refused.

Strands does not send agent.state to the model. So the patient id stays out of the model’s reach, and no tool takes one from the model. The select_medication tool writes the selected record and its label to state. A new selection always replaces the old one.

Step 2: Add the confirmation step

The agent must read the medicine back and wait for a yes before it sends. In Strands, that is one intervention handler. It answers every send_refill_request call before the tool runs. This excerpt is from strands/guard.py:

class RefillGuard(InterventionHandler):
    name = "refill-guard"

    def __init__(self, conversation_id: str) -> None:
        self.conversation_id = conversation_id
        #: toolUseId -> (answer, confirmed) for the confirmations evaluated.
        self.answers: dict[str, tuple[str, bool]] = {}

    @property
    def on_error(self) -> str:
        return "deny"

    def before_tool_call(self, event: Any):
        if event.tool_use["name"] != "send_refill_request":
            return Proceed()
        state = event.agent.state
        patient: Optional[str] = state.get("patient_id")
        selected: Optional[str] = state.get("selected_record_id")
        label: str = state.get("selected_medication_label") or ""
        record_id = str((event.tool_use.get("input") or {}).get("record_id") or "").strip().upper()
        if not patient:
            return Deny(reason="The caller is not verified. Call verify_patient first. Nothing was sent.")
        if not selected:
            return Deny(reason="No medication is selected. Call select_medication first. Nothing was sent.")
        if record_id != selected:
            return Deny(reason=f"Only the medication select_medication returned can be sent: record_id {selected}. "
                               "Nothing was sent.")
        question = clinic.confirmation_question(label)
        tool_use_id = str(event.tool_use.get("toolUseId"))

        def evaluate(response: Any) -> bool:
            answer = _answer(response)
            confirmed = caller_said_yes(answer, label)
            clinic.record_confirmation(self.conversation_id, record_id, confirmed, mechanism=MECHANISM,
                                       question=question, answer=answer)
            self.answers[tool_use_id] = (answer, confirmed)
            return confirmed

        return Confirm(prompt=question, evaluate=evaluate)

Here is what each part does:

  • Proceed lets every other tool run as normal.
  • Deny cancels the send if no patient is verified, no medicine is selected, or the record is not the selected one. The model sees the reason.
  • on_error is "deny", so a crash inside the handler also blocks the send.
  • Confirm pauses the agent with the clinic’s read-back question. The agent stops, and waits until someone resumes it with an answer.
  • evaluate judges that answer, records it in the audit log and lets the tool run only on a yes.

Resume with the caller’s next words

The pause and the resume live in the conversation’s turn function in agent.py. When the run stops on the Confirm, the voice loop speaks the question. This excerpt is from the end of that function:

        if result is not None and result.stop_reason == "interrupt":
            self.pending = list(result.interrupts)
            for interrupt in self.pending:
                # The Confirm prompt: cedar_clinic's read-back question, verbatim.
                yield Say("fixed", str(interrupt.reason), "confirmation")

On the caller’s next turn, their words become the answer. This excerpt is from the start of the same function:

        prompt: Any = text
        # concern-begin: refill-guard
        # A pending Confirm interrupt is answered with the caller's words from
        # this turn, the only way it is ever resumed; guard.py decides.
        if self.pending:
            prompt = [{"interruptResponse": {"interruptId": i.id, "response": {"answer": text}}}
                      for i in self.pending]
            self.pending = []
        # concern-end

A resumed agent receives only the interrupt response, not a new caller message. So the model would not see what the caller said at the read-back. “No, wait, not that one. I meant my budesonide inhaler.” would arrive as a plain no.

The handler fixes that with a third intervention. After the send, it returns a Transform that adds the caller’s words to the tool result. This is an excerpt from after_tool_call in guard.py:

        def add_answer(e: Any) -> None:
            note = f'The caller answered the read-back question: "{answer}".'
            if not confirmed:
                note += (f' Nothing was sent, and the caller has been told: "{clinic.DECLINED_TEXT}" If they named '
                         "a different medicine, call select_medication for it.")
            e.result["content"] = list(e.result.get("content") or []) + [{"text": note}]

        return Transform(apply=add_answer)

With it, a correction goes through on the same turn. This transcript is the call correction-other-medicine-at-confirmation, from a live run on 30 September 2026 (UTC), in results/strands/2026-10-01-speechmatics-live. The caller lines are what speech-to-text heard:

CALLER: No, wait, not that one. I meant my budesonide inhaler.
BOT:    Okay, I have not sent a refill request.
BOT:    I can send a request about this recorded medication, budesonide inhaler, one puff twice a day, to the prescribing team. Would you like me to do that?
CALLER: Yes, that's right. Send it .
BOT:    Okay, one moment.
BOT:    Your refill request reference is R Q, one four five seven. It is awaiting prescribing team review.

Decide what counts as a yes

By default, Confirm judges the answer itself. It accepts True, or the text y or yes after lowercasing and trimming. That suits a button or a typed reply. A spoken answer comes back as a transcript, often with a full stop, so the default reads it as a no.

This offline run compares the default with the rule this build uses. The first three answers are what callers said in these runs. The last is a bare “Yes” for comparison. From the strands folder, on 1 October 2026:

uv run --locked python - <<'EOF'
from strands.interventions.actions import default_evaluate
from guard import caller_said_yes
for a in ["Yes, that's right. Send it .", "Yes, send it.", "Yes.", "Yes"]:
    print(repr(a), default_evaluate(a), caller_said_yes(a, "budesonide inhaler"))
EOF
"Yes, that's right. Send it ." False True
'Yes, send it.' False True
'Yes.' False True
'Yes' True True

The default fails safe: nothing is sent. But each of those sends would have been cancelled. So supply your own evaluate. This build uses a fixed rule, from guard.py:

_AFFIRM_START = re.compile(
    r"^(yes|yeah|yep|yup|sure|ok|okay|fine|alright|all right|correct|right|absolutely|definitely|please( do| send)?|go ahead|"
    r"that's (right|correct|the one|it)|that is (right|correct)|send it|do it)\b")
_NEGATION = re.compile(r"\b(no|nope|not|don't|dont|do not|wait|cancel|stop|instead|wrong|hold on|never mind)\b")
_LATER_NEGATION = re.compile(r"\b(don't send|do not send|not that|wait|cancel|instead|wrong|hold on|never mind)\b")
_DRUG_NAMES = sorted({str(m["name"]).lower() for m in refills.load_data()["medications"].values()})


def caller_said_yes(answer: Any, medication_label: str) -> bool:
    """Whether the caller's words approve the read-back. Anything unclear is a no.

    Yes only when the first clause starts with an affirmative and has no
    negation, no later clause retracts it, and no other recorded medicine is
    named.
    """
    text = str(answer or "").lower().replace("’", "'").strip()
    if not text:
        return False
    first = re.split(r"[.,!?;]", text, maxsplit=1)[0].strip()
    if not _AFFIRM_START.search(first) or _NEGATION.search(first) or _LATER_NEGATION.search(text):
        return False
    label = medication_label.lower()
    return not any(re.search(rf"\b{re.escape(name)}\b", text) and name not in label for name in _DRUG_NAMES)

The rule needs no model call. It is also strict. A later “wait”, or another medicine’s name, makes the whole answer a no. The cost is that it refuses a yes worded in a way it does not know. No recorded call used one. A real deployment would need more wordings, or a model to judge the answer.

Each build judges the yes in its own way:

BuildWho judges the caller’s answer
RasaThe model that runs each of the agent’s turns
LangGraphA separate model call inside the guard, with structured output
Strandscaller_said_yes, a fixed rule
Why not the ready-made HumanInTheLoop handler?

Strands ships a HumanInTheLoop handler that would pause the call. But it asks with the tool name and its input, not the clinic’s read-back question. A small custom handler that returns Confirm was simpler. Strands also offers Cedar authorization, which checks tool calls against policies in the Cedar policy language (no relation to the clinic). It also offers steering handlers and Bedrock Guardrails. This guard is three state checks and one confirmation, so it did not need them.

Step 3: Write the voice loop

Strands has its own voice agent, BidiAgent. It runs speech-to-speech models: Nova Sonic, OpenAI Realtime and Gemini Live. In strands-agents 1.57.1, the version this build pins, it is experimental and Python-only. From 1.57.2 the docs no longer call it experimental. This build uses separate speech-to-text and text-to-speech from Speechmatics, so it uses the plain Agent and a voice loop you write.

The loop lives in server.py, on Starlette and uvicorn. It speaks the same browser audio protocol as Rasa, so one web page works for all three builds. It does these jobs:

  • Audio in: It forwards the caller’s audio to Speechmatics speech-to-text.
  • End of turn: Speechmatics sends one final transcript when the caller pauses for 0.7 seconds. Each one becomes one agent turn, answered in order.
  • Speaking while the model writes: It cuts the model’s text at sentence ends and sends each sentence to text-to-speech straight away.
  • Fillers: When tools finish before anything was said in the turn, it speaks a short fixed phrase, such as “Let me find that on your record.”
  • Playback markers: It sends markers with the audio, and the browser sends them back once that audio has played.
  • Silence check-in: 30 seconds after the agent’s audio has played with no caller speech, it speaks a fixed prompt.
  • Events: It records what the caller and the agent said, for the test runner.

This excerpt from server.py turns each item from the agent’s turn into speech:

                async for say in cedar.conversation(self.cid).turn(text):
                    if say.kind == "delta":
                        msg = msg or {"id": uuid.uuid4().hex}
                        message(chunker.feed(say.text), msg)
                    elif say.kind == "end":
                        msg = msg or {"id": uuid.uuid4().hex}
                        message(chunker.flush(), msg)
                        add_event(self.cid, "bot", text=say.text.strip(), metadata={"source": "model"})
                        turn["spoken"] += 1
                        speech.put_nowait(("end", msg))
                        msg = None
                    elif say.kind == "fixed":
                        whole(say.text, say.source)
                    elif say.kind == "tools" and say.source == "ok" and not turn["spoken"] and say.text in FILLERS:
                        whole(FILLERS[say.text], "filler")

There is no barge-in and no cache for repeated speech. Caller audio keeps flowing to speech-to-text while the agent speaks. A transcript that arrives then is answered after the current turn.

Count a yes only after the question has played

The loop answers transcripts in the order they arrive. When a Confirm is waiting, the next transcript becomes its answer, whenever the caller said it. That keeps the answer on a later turn than the question. But it does not prove the caller heard the question first.

A test replay showed the gap. It copies a live call on which speech-to-text split the caller’s words. The replay sent “Yes, please.” the moment the caller had finished saying it. In the replay, the yes arrived before the read-back had played. On that live call, the transcript came after it. This is one replay, from results/strands/2026-10-01-late-transcript-replay/summary.md. SENT is the replay’s transcript, and user is the build logging it:

  14.98  SENT           Of my omeprazole.
  14.98  user           Of my omeprazole.
  16.07  bot            Let me find that on your record.
  16.83  SENT           Yes, please.
  16.83  user           Yes, please.
  17.43  bot            I can send a request about this recorded medication, omeprazole twenty milligram capsules, one capsule before breakfast, to the prescribing team. Would you like me to do that?
  20.84  bot_turn_ended
  20.84  bot            Okay, one moment.
  21.88  bot            Your request reference is R Q, eight zero four six. It is awaiting prescribing team review.

The yes arrived before the read-back, and it was taken as the answer. That happened in 3 of 3 replays. The yes rule judges what was said, not when. The same gap showed up in all three builds, so it is not a Strands problem. It is a voice loop problem, and the fix goes in the loop.

The companion’s opt-in fix.diff changes only server.py, plus a test for it. It records when the caller began speaking, and when the browser confirmed the last turn had played. Then it checks both before a waiting Confirm gets its answer. This hunk from fix.diff is the core of the change:

@@ -194,8 +208,17 @@
             item = await self.turns.get()
             if item is None:
                 return
-            text, heard_at = item
+            text, heard_at, onset = item
             log(self.cid, "caller:", text)
+            # concern-begin: refill-guard
+            # The fix: an answer that began before the question finished playing is not consent;
+            # the pending Confirm is asked again instead of being answered with it.
+            pending = cedar.conversation(self.cid).pending
+            if pending and (self.turn_played_at is None or onset < self.turn_played_at):
+                log(self.cid, "answer began before the question finished playing; asking again")
+                await self.bot_turn([("fixed", str(i.reason), "confirmation") for i in pending], heard_at, None)
+                continue
+            # concern-end
             try:
                 await self.bot_turn(None, heard_at, text)
             except Exception as exc:

With the fix, the agent asked the question again, and nothing was sent in 3 of 3 replays. No model call was needed to refuse.

The fix also covers short sounds of agreement. The yes rule accepts some of them, and it should, because a caller may answer “Okay.” This offline run shows which ones, from the strands folder:

$ uv run --locked python -c "from guard import caller_said_yes; print('  '.join(f'{a!r} {caller_said_yes(a, \"omeprazole twenty milligram capsules\")}' for a in ['Okay.', 'Yeah.', 'Sure.', 'Yes.', 'Mm-hm.', 'Uh-huh.']))"
'Okay.' True  'Yeah.' True  'Sure.' True  'Yes.' True  'Mm-hm.' False  'Uh-huh.' False

In replays, the shipped build sent the request on an “Okay.” said over its filler, and on a “Yeah.” said during its read-back. The fixed copy sent on neither and asked again. On live calls, the fixed copy passed 4 of 4, each with one read-back and then the send (3 calls and 1 call).

Step 4: Run it

From the strands folder:

make install     # strands-agents[openai] 1.57.1, cedar_clinic and cedar_speech into .venv
make env         # fill OPENAI_API_KEY and SPEECHMATICS_API_KEY
make run         # ws://localhost:5007/webhooks/browser_audio/websocket
make web         # in another shell: the voice page on http://127.0.0.1:8765/, relaying to :5007

Open the voice page and talk to the agent. The test patients are fictional. Try “I’m Maria Alvarez, born March 14th, 1968. Lisinopril, please.” The agent reads the medicine back and waits for your yes.

To try the timing fix, make a fixed copy from the tutorial folder:

make fix-copy FW=strands    # strands-fix/ with fix.diff applied

The copy leaves out your .env and installed packages. So run make install and make env again inside strands-fix. make env creates .env there with both keys blank, so fill them in again before make run.

Step 5: Test it

Start with the offline tests. They use a scripted model and fake speech, so they need no keys and no network:

make test

The guard tests drive the real agent with a scripted model. Here is that group on its own, with its output:

uv run --locked python -m unittest discover -s tests -k GuardTests -v
test_a_no_sends_nothing_and_speaks_the_shared_decline (test_strands.GuardTests.test_a_no_sends_nothing_and_speaks_the_shared_decline) ... ok
test_send_pauses_for_the_caller_and_runs_only_after_a_yes_on_the_next_turn (test_strands.GuardTests.test_send_pauses_for_the_caller_and_runs_only_after_a_yes_on_the_next_turn) ... ok
test_send_without_a_selection_or_for_another_entry_is_denied (test_strands.GuardTests.test_send_without_a_selection_or_for_another_entry_is_denied) ... ok
test_the_model_never_supplies_the_patient_id (test_strands.GuardTests.test_the_model_never_supplies_the_patient_id) ... ok

----------------------------------------------------------------------
Ran 4 tests in 0.010s

OK

The pause test is the main one. It checks that the first turn ends with the clinic’s question and no send. After “Yes, please send it.” the audit log reads verify_patient, select_medication, record_confirmation, send_refill_request, with the caller’s words in the confirmation. These tests prove the guard’s rules. They do not show what a real model does.

For that, play the shared recorded calls to the build:

make spec        # the 17 recorded calls over browser audio (billed, capped at 4 USD a run)

The runner plays each call and judges it from the clinic’s audit log, with the same checks for every build. On the live run, this build passed 16 of the 17 calls, and the guard was never broken.

To test the timing gap, run the early-yes replay from the tutorial folder. Run it against the shipped build, then against the fixed copy:

make late-transcript-replay FW=strands LABEL=my-replay                      # billed
make late-transcript-replay FW=strands CWD=strands-fix LABEL=my-replay-fix  # billed

The first should send the request, and the second should ask again. Chapter 5 shows how to prove your tests catch a missing guard.

Limits

  • Cedar Clinic is fictional, and nothing here is clinical advice.
  • The runs in this chapter used GPT-5.5 with Speechmatics. One extra run of the 17 calls used Deepgram.
  • The early-yes replays simulate only speech-to-text. They show that the shipped loop accepts an early yes, not how often callers give one.
  • The fix was tested on replays and a few live calls, not proven.