Skip to content
RasaGet a free licence

tutorial

Chapter 3 of 6

Build the voice agent on LangGraph

by Rod Rivera Published

Build the Cedar Clinic prescription-refill voice agent on LangGraph, from its tools and confirmation step to the voice loop, a live call and tests.

Key takeaways (3)
  • LangGraph gives you the agent loop, state for each call, and a pause that waits for the caller's answer.
  • You write the voice loop yourself: audio in and out, turn order, fillers and the silence check-in.
  • Count the caller's yes only after the read-back question has finished playing.

You build the Cedar Clinic prescription-refill voice agent on LangGraph. The caller gives their name and date of birth, then names a medicine. The agent reads the medicine back and sends a refill request only after a clear yes.

LangGraph is LangChain’s low-level orchestration runtime for long-running, stateful agents. LangGraph’s docs recommend starting one level up, with LangChain’s create_agent, which runs on LangGraph. This build does that.

You get a working agent loop, saved state for each call and a built-in way to pause for the caller. You write the voice loop: the code that listens, takes turns and speaks. The chapter walks through both.

What you need. Python 3.11 or 3.12 with uv, an OpenAI API key and a Speechmatics API key. LangGraph and LangChain are MIT-licensed, so there is no licence key. Live calls are billed by OpenAI and Speechmatics. The offline tests are free.

Get the code at the commit this series uses:

git clone https://github.com/RasaHQ/rasa-community-resources
cd rasa-community-resources
git checkout 41dd184955425f1d1686cdb39c91a0fcc1829442
cd tutorials/voice-agent-three-frameworks/langgraph

The steps below read the companion’s finished build, one part at a time. To start your own project, copy the four Python files in the langgraph folder: agent.py, guard.py, voice_loop.py and server.py. They import two shared packages from the companion. Replace cedar_clinic, the clinic’s business rules, with your own. Keep cedar_speech, the Speechmatics clients, or swap in your speech vendor.

What LangGraph gives you, and what you write

Part of the agentWhat LangGraph and LangChain give youWhat you write
The agent loopcreate_agent: the model calls tools until it is doneThe prompt and the list of tools
State for each callA checkpointer, and state fields the model cannot setWhich fields to keep, and the tools that set them
The confirmation stepMiddleware hooks, interrupt() and resumeWhen to ask, and how to judge the reply
The voice loopStreaming of model text, tool calls and custom eventsAudio, turn order, speech, fillers and silence
Barge-in (the caller talking over the agent)No audio parts, so nothing built inNot written in this build

The business rules live in a shared Python package, cedar_clinic. It looks up records, applies the clinic’s rules and writes the audit log: the clinic’s own record of every tool call and outcome. All three builds call the same package, so this chapter only covers the LangGraph code around it.

Step 1: Define the tools and the state they keep

The agent has five tools. Each one calls cedar_clinic and returns its result to the model. This excerpt from langgraph/agent.py builds the agent:

def build_agent(model: Any = None, *, classify: Any = None, checkpointer: Any = None):
    """The compiled graph. One checkpointer per process; thread_id is the conversation id."""
    model = model if model is not None else make_model(streaming=True)
    return create_agent(
        model=model,
        tools=TOOLS,
        system_prompt=SYSTEM_PROMPT,
        # concern-begin: refill-guard
        middleware=[RefillGuard(classify or llm_classifier(make_model()))],
        # concern-end
        checkpointer=checkpointer or InMemorySaver(),
    )

The checkpointer keeps each call’s history under a thread id. The voice loop sets that id to the conversation id, so every call has its own history. The prompt comes from cedar_clinic, so it is the same text in all three builds.

Each tool gets a ToolRuntime, which gives it the call’s state. It returns a Command that updates that state. This excerpt is the tool that checks the caller’s name and date of birth:

@tool("verify_patient", **_spec("verify_patient"))
async def verify_patient(full_name: str, date_of_birth: str, runtime: ToolRuntime) -> Command:
    result = clinic.verify_patient(_conversation_id(runtime), full_name, date_of_birth)
    update: dict = {}
    # concern-begin: refill-guard
    # The patient id goes to state, never to the model. The first verified
    # patient stays: a call verified as one patient cannot become another.
    if result["status"] == "verified":
        current = runtime.state.get("patient_id")
        if not current:
            update = {"patient_id": result["patient_id"], "patient_first_name": result["first_name"]}
        elif current != result["patient_id"]:
            result = {"status": "not_verified", "reason": "already_verified_as_another_patient",
                      "next_step": "This call is verified for a different patient. Do not act for this one."}
    # concern-end
    return _reply(runtime, clinic.for_model(result), **update)

This is part of the guard: the safety check that stops a refill going out without a clear yes, or for a second patient. The first verified patient stays for the whole call. A second patient on the same call is refused.

The select_medication tool works the same way. It writes the selected record and its label to state, and a new selection always replaces the old one.

Keep these fields out of the model’s reach. This excerpt from langgraph/guard.py declares them:

class RefillState(AgentState):
    """Agent state plus what only tools write. Private: not settable from the graph's input."""

    patient_id: NotRequired[Annotated[str, PrivateStateAttr]]
    patient_first_name: NotRequired[Annotated[str, PrivateStateAttr]]
    selected_record_id: NotRequired[Annotated[str, PrivateStateAttr]]
    selected_label: NotRequired[Annotated[str, PrivateStateAttr]]

PrivateStateAttr leaves a field out of the graph’s input and output. So nothing outside the tools can set the patient or the selected medicine. No tool takes a patient id from the model either.

Step 2: Add the confirmation step

The agent must read the medicine back and wait for a yes before it sends. In this build, that is middleware. The middleware hooks belong to LangChain’s create_agent. The guard uses two of them.

The first hook hides the send tool until a medicine is selected. This is an excerpt from guard.py:

    async def awrap_model_call(self, request: ModelRequest, handler: Callable) -> Any:
        """Hide send_refill_request until a record entry is selected."""
        if not request.state.get("selected_record_id"):
            request = request.override(tools=[t for t in request.tools if getattr(t, "name", None) != SEND])
        return await handler(request)

The second hook runs around every tool call. For the send tool, it checks the record, then pauses the run with interrupt(). This excerpt is the rest of the guard class (guard.py):

    async def awrap_tool_call(self, request: ToolCallRequest, handler: Callable) -> Any:
        call = request.tool_call
        if call["name"] != SEND:
            return await handler(request)
        state = request.state
        selected = state.get("selected_record_id") or ""
        record_id = str(call["args"].get("record_id") or "").strip().upper()
        if not selected or record_id != selected:
            return _blocked(call, "medication_not_resolved",
                            "Call select_medication first and send only the record_id it returned.")
        question = clinic.confirmation_question(state.get("selected_label") or "")
        # Everything above runs again on resume and has no side effects.
        resumed = interrupt({"kind": "confirm_refill", "question": question, "record_id": record_id})
        answer = str(resumed.get("text") if isinstance(resumed, dict) else resumed or "").strip()
        confirmed = bool(answer) and await self.classify(question, answer)
        conversation_id = request.runtime.config["configurable"]["thread_id"]
        clinic.record_confirmation(conversation_id, record_id, confirmed, mechanism=MECHANISM,
                                   question=question, answer=answer)
        if not confirmed:
            request.runtime.stream_writer({"say": clinic.DECLINED_TEXT})
            return ToolMessage(json.dumps({
                "status": "declined", "effects": 0, "caller_answer": answer,
                "caller_was_told": clinic.DECLINED_TEXT,
                "next_step": ("Nothing was sent and the caller has been told so; do not repeat it. If the caller "
                              "named a different medicine, call select_medication with it. Otherwise answer "
                              "what they said or ask what else they need."),
            }), tool_call_id=call["id"], name=SEND)
        # A progress event for the voice loop, which may say a filler while the request is sent.
        request.runtime.stream_writer({"confirmed": record_id})
        result = await handler(request)
        return _with_answer(result, answer)

Here is what happens, in order:

  1. A send for any record other than the selected one is refused before the tool runs.
  2. interrupt() stops the run with the clinic’s read-back question. The voice loop speaks it.
  3. The caller’s next turn resumes the run. A separate model call judges whether the words are a yes.
  4. The answer goes to the clinic’s audit log. Only a yes runs the tool.
  5. A no speaks the clinic’s decline and hands the caller’s words back to the model. So “No, wait, not that one. I meant my budesonide inhaler.” gets the inhaler selected and read back on the same turn.

The yes judge uses structured output, so the model must return true or false. This excerpt from guard.py builds it:

def llm_classifier(model: Any) -> Classifier:
    """Yes or no from the model, with structured output: the judgement Rasa's engine also leaves to the model."""
    structured = model.with_structured_output(CallerAnswer, method="json_schema")

    async def classify(question: str, answer: str) -> bool:
        result = await structured.ainvoke(CLASSIFY_PROMPT.format(question=question, answer=answer))
        return bool(result.confirmed)

    return classify

Its prompt says that a no, a different medicine, a question without a yes, or anything unclear is not a yes.

Why not LangChain's HumanInTheLoopMiddleware?

It is built for a reviewer who approves, edits, rejects or responds to a proposed tool call. A caller answers in words, so something still has to turn the words into a decision. The question also has to read the medicine back from state, and the answer has to reach the clinic’s audit log. LangGraph’s interrupts page says interrupts “can be placed anywhere in your code and can be conditional”. So this guard calls interrupt() from its own middleware, and the tool body stays simple.

Step 3: Write the voice loop

LangGraph and LangChain do not handle audio. So you write the voice loop, in voice_loop.py and server.py. This build uses Starlette and uvicorn for the server. It speaks the same browser audio protocol as Rasa, so one web page works for all three builds.

The loop does these jobs:

  • Audio in: It forwards the caller’s audio to Speechmatics speech-to-text.
  • End of turn: Speechmatics sends one final transcript when the caller pauses for 0.7 seconds. The loop answers those turns one at a time, in order.
  • Speaking while the model writes: It cuts the model’s text at sentence ends and sends each sentence to text-to-speech straight away.
  • Fillers: If the model starts a tool call before saying anything, it speaks a short fixed phrase, such as “Let me look at your record.”
  • Playback markers: It sends markers with the audio, and the browser sends them back once that audio has played.
  • Silence check-in: After 30 seconds with no caller speech, it asks “Are you still there? I can help when you’re ready.”
  • Events: It records what the caller and the agent said, for the test runner.

This excerpt from voice_loop.py streams the agent and speaks a filler when a tool call starts:

        async for mode, data in self.agent.astream(payload, self.config,
                                                   stream_mode=["messages", "custom", "updates"]):
            if mode == "messages":
                chunk, meta = data
                if meta.get("langgraph_node") != "model":
                    continue
                if chunk.id != current_id:
                    self._finish(current)
                    current, current_id = None, chunk.id
                delta = chunk.text
                if delta:
                    current = current or Message(clock)
                    self._feed(current, delta)
                    spoke = True
                calls = getattr(chunk, "tool_call_chunks", None) or []
                if calls and not spoke:
                    self._speak_whole(FILLERS.get(calls[0].get("name") or "", "One moment."), clock)
                    spoke = True

The loop also has to know about the guard’s pause. When the run is paused, the caller’s next turn is the answer, so the loop sends it as a resume. This excerpt is from voice_loop.py:

        # concern-begin: refill-guard
        # The guard asked the confirmation question on an earlier caller turn and the run is paused
        # in interrupt(). This caller turn is the answer: it is the resume value, and only a caller
        # turn ever resumes it, so the answer always comes from a later turn than the question.
        if state.interrupts:
            payload = Command(resume={"text": text, "turn": self.turn_count})
        # concern-end

Another part of the loop speaks the interrupt’s question, and the decline the guard writes to the custom stream.

This build does not do barge-in, and it has no cache for repeated speech. Audio that arrives while the agent speaks is answered after the current turn. To add barge-in, you would cancel the task running astream, stop the audio, and decide what the saved state keeps of a half-spoken reply.

To use Deepgram for speech instead, the companion’s launcher starts the same server with Deepgram clients. No file in the LangGraph folder changes.

Count a yes only after the question has played

The loop answers transcripts in the order they arrive. That keeps the answer on a later turn than the question. But it does not prove the caller heard the question first.

A test replay showed the gap. It copies a live call on which speech-to-text split the caller’s words. The replay sent “Yes, please.” the moment the caller had finished saying it. In the replay, the yes arrived before the read-back had played. On that live call, the transcript came after it. This is one replay, from results/langgraph/2026-10-01-late-transcript-replay/summary.md. SENT is the replay’s transcript, and user is the build logging it:

  16.07  bot_turn_ended
  18.85  SENT           Of my omeprazole.
  18.85  user           Of my omeprazole.
  20.58  bot            Let me look at your record.
  20.70  SENT           Yes, please.
  20.70  user           Yes, please.
  26.81  bot            I can send a request about this recorded medication, omeprazole twenty milligram capsules, one capsule before breakfast, to the prescribing team. Would you like me to do that?
  26.81  bot_turn_ended
  29.81  bot            Right, I'll send that request for review.
  32.54  bot            Your request reference is R Q, seven six zero four. It is awaiting prescribing team review, and they will contact you with the outcome.

The early yes resumed the run, and the request went out. That happened in 3 of 3 replays. The same gap showed up in all three builds, so it is not a LangGraph problem. It is a voice loop problem, and the fix goes in the loop.

The companion’s opt-in fix.diff changes only voice_loop.py, plus a test for it. It records when the caller began speaking, and when the browser confirmed the last turn had played. Then it checks both before it resumes. This hunk from fix.diff is the core of the change:

@@ -244,6 +263,15 @@
         # in interrupt(). This caller turn is the answer: it is the resume value, and only a caller
         # turn ever resumes it, so the answer always comes from a later turn than the question.
         if state.interrupts:
+            # The fix: an answer that began before the question finished playing is not consent.
+            # Ask the question again instead of resuming.
+            if self.turn_played_at is None or (onset or clock.started) < self.turn_played_at:
+                log.info("%s: answer began before the question finished playing; asking again", self.id)
+                for item in state.interrupts:
+                    value = getattr(item, "value", None)
+                    if isinstance(value, dict) and value.get("question"):
+                        self._speak_whole(value["question"], clock)
+                return
             payload = Command(resume={"text": text, "turn": self.turn_count})
         # concern-end
         current: Optional[Message] = None

With the fix, the early yes no longer counted. The agent asked the question again, and nothing was sent in 3 of 3 replays. The same held for an “Okay.” said over the filler and a “Yeah.” said during the read-back. On live calls, the fixed copy passed 4 of 4, each with one read-back and then the send (3 calls and 1 call).

Step 4: Run it

From the langgraph folder:

make install     # langgraph, langchain, langchain-openai, starlette, uvicorn, cedar_clinic, cedar_speech
make env         # fill OPENAI_API_KEY and SPEECHMATICS_API_KEY
make run         # ws://localhost:5006/webhooks/browser_audio/websocket
make web         # in another shell: the voice page on http://127.0.0.1:8765/

Open the voice page and talk to the agent. The test patients are fictional. Try “I’m Maria Alvarez, born March 14th, 1968. Lisinopril, please.” The agent reads the medicine back and waits for your yes.

To try the timing fix, make a fixed copy from the tutorial folder:

make fix-copy FW=langgraph    # langgraph-fix/ with fix.diff applied

The copy leaves out your .env and installed packages. So run make install and make env again inside langgraph-fix. make env creates .env there with both keys blank, so fill them in again before make run.

Step 5: Test it

Start with the offline tests. They use a scripted model and fake speech, so they need no keys and no network:

make test

The guard tests drive the real graph with a scripted model. Here is that group on its own, with its output:

uv run --locked python -m unittest discover -s tests -k GuardTests -v
test_a_no_sends_nothing_and_the_model_hears_the_answer (test_guard.GuardTests.test_a_no_sends_nothing_and_the_model_hears_the_answer) ... ok
test_a_send_for_another_record_than_the_selected_one_is_refused (test_guard.GuardTests.test_a_send_for_another_record_than_the_selected_one_is_refused) ... ok
test_a_send_without_a_selection_is_refused_before_the_tool_runs (test_guard.GuardTests.test_a_send_without_a_selection_is_refused_before_the_tool_runs) ... ok
test_a_tool_called_beside_the_paused_send_is_not_run_again_on_resume (test_guard.GuardTests.test_a_tool_called_beside_the_paused_send_is_not_run_again_on_resume) ... ok
test_guard_state_cannot_be_set_from_the_graph_input (test_guard.GuardTests.test_guard_state_cannot_be_set_from_the_graph_input) ... ok
test_send_is_not_offered_until_a_medicine_is_selected (test_guard.GuardTests.test_send_is_not_offered_until_a_medicine_is_selected) ... ok
test_the_model_never_sees_the_patient_id (test_guard.GuardTests.test_the_model_never_sees_the_patient_id) ... ok
test_the_question_pauses_the_run_and_only_a_later_yes_sends (test_guard.GuardTests.test_the_question_pauses_the_run_and_only_a_later_yes_sends) ... ok

----------------------------------------------------------------------
Ran 8 tests in 0.050s

OK

The last test is the main one. It runs one turn and checks that the run paused with the clinic’s question and sent nothing. Then it resumes with “Yes, please send it.” and checks the audit log. These tests prove the guard’s rules. They do not show what a real model does.

For that, play the shared recorded calls to the build:

make spec        # the 17 recorded calls over browser audio (billed, capped at 4 USD a run)

The runner plays each call and judges it from the clinic’s audit log, with the same checks for every build. On the live run, this build passed 16 of the 17 calls, and the guard was never broken.

To test the timing gap, run the early-yes replay from the tutorial folder. Run it against the shipped build, then against the fixed copy:

make late-transcript-replay FW=langgraph LABEL=my-replay                        # billed
make late-transcript-replay FW=langgraph CWD=langgraph-fix LABEL=my-replay-fix  # billed

The first should send the request, and the second should ask again. Chapter 5 shows how to prove your tests catch a missing guard.

Limits

  • Cedar Clinic is fictional, and nothing here is clinical advice.
  • The runs in this chapter used GPT-5.5 with Speechmatics. One extra run of the 17 calls used Deepgram.
  • The early-yes replays simulate only speech-to-text. They show that the shipped loop accepts an early yes, not how often callers give one.
  • The fix was tested on replays and a few live calls, not proven.