Skip to content
RasaGet a free licence
Guides for AI teams

Guide · AI product engineer

How to test a Rasa output hook on confirmed tool calls

A tool that asks the customer to confirm reaches your Rasa hook under another name, so test the hook with what it really receives.

by Rod Rivera

About 10 minutes

  • 2 → 0

    replies saying no seat was held while one was: main run → rerun after the fix

  • 1 of 33

    offline tests that end in an error when the old name check is put back

Source: Recorded test runs of a sample airline agent on GPT-5.5, 30 September 2026; offline tests, 2 and 3 October 2026
Key takeaways (3)
  • A guard that falls back to a cautious answer can still tell a customer the opposite of the record.
  • Build each hook test from the name and value type the hook gets after the customer confirms.
  • Count the replies that contradict what the tool told the customer. A promise counter will not see them.

You added a hook that reads your agent’s replies before the customer sees them. It stops the model from promising things the tools have not done. Your tests pass and the hook rarely fires, so you take that as a sign it works.

In a sample airline agent we built with Rasa, that kind of hook told a passenger no seat was held. It did so eight seconds after the tool had held one. The airline, Horizon Travel, is fictional and the passenger is scripted, but the model calls were live, on GPT-5.5:

Chat transcriptRecorded test run, 30 September 2026, trimmed to the last two turns (earlier turns left out; emphasis added)
  1. Passenger

    Yes, hold that one.

  2. Agent

    Options are changing during the disruption. Shall I try to hold HZ 1204 and HZ 1480 via Detroit, Wed 1 Oct 09:05, arrives 13:40 (booking HT-7Q4M2)? A hold keeps a seat until it expires; it is not a confirmed journey, and I can’t promise a seat until the hold is placed.

    Annotation:

    The question Rasa asks before the hold tool runs.

  3. Passenger

    Yes.

  4. Agent

    Seat held: HZ 1204 and HZ 1480 via Detroit, Wed 1 Oct 09:05, arrives 13:40, hold HT-HLD-E7E86D, until 14:50 Boston time, 30 Sep. This is a hold, not a confirmed journey. To keep it, accept it in Manage booking before it expires.

    Annotation:

    Sent by the hold tool itself, from the seat inventory’s result.

  5. Agent

    No seat is held for you in this chat, and nothing here confirms a journey. I can list recovery flights, try to hold one, or put you in the recovery queue.

    Annotation:

    Sent 7.8 seconds later by our hook. It replaced the model’s third try at saying the seat was held.

The test run marked this conversation as passed. Its two checks read only the hold tool’s calls: the first flight the passenger picked was not held, and the Detroit connection was. The wrong message came from the hook, the layer we added to catch wrong messages.

The cause was a name. The hold tool asks the passenger to confirm before it runs. After the passenger says yes, Rasa hands the result to the hook under the name resolve_tool_confirmation, not the tool’s own name. Our hook looked only for the tool’s own name, so it never saw the hold. Six of our sample agents had a hook that missed what Rasa sends, and in four it replaced a correct reply in a recorded run.

A cautious fallback built from tool state is only as true as that state. The fix is to recognise the result by what it contains, and to test the hook with the name and value it really receives. That matters if your hook builds a fallback answer from what the tools report.

Why the hook missed the hold

A hook is a function Rasa calls at a fixed point in each turn. The sample’s guard, in hooks.py, uses two:

  1. A modify_tool_result hook runs after each tool call. It records each hold the tools report.
  2. A modify_model_response hook reads each draft reply before the passenger sees it.

A sentence saying a seat is held passes only while a recorded hold is active. Otherwise the draft goes back to the model. After two retries, the hook sends a fixed answer built from its recorded holds. With none, that answer is “No seat is held for you in this chat”.

The hold tool is gated. In the skill’s frontmatter, requires_confirmation makes Rasa ask the passenger before the tool runs. This excerpt is the end of the frontmatter in skills/disruption_recovery/skill.md, with the name and description left out:

import_tools:
  - get_incident_status
  - check_hold
  - release_hold
  - join_recovery_queue
tool_constraints:
  - hold_recovery_option:
      requires: session.disruption_recovery.selected_option_id
      requires_confirmation:
        enabled: true
        utter_for_confirmation: utter_confirm_hold

check_hold and release_hold have no constraint, so they run straight away. The hold tool takes a longer path:

  1. The model calls hold_recovery_option, and Rasa asks the question in utter_confirm_hold.
  2. The passenger says yes.
  3. The model calls resolve_tool_confirmation, a tool Rasa provides.
  4. Rasa runs the stored hold call.
  5. Rasa passes the result to the hook under the name of the call the model made, resolve_tool_confirmation.

Step 5 happens in Rasa’s orchestrator. This excerpt is from rasa/mantle/orchestration/orchestrator.py in the Rasa Pro 3.21.0.dev5 wheel, lines 1911 to 1919, with its source indentation:

                normalized_result = raw_result or f"{tc.tool_name}: done"
                is_error, error_message = _error_from_tool_payload(raw_result)
                hooked = await apply_tool_result_hooks(
                    self._static_data.hooks,
                    tool_name=tc.tool_name,
                    arguments=dict(tc.tool_args),
                    value=normalized_result,
                    is_error=is_error,
                )

tc.tool_name is the name of the call the model made. normalized_result is the result as text, not a Python dictionary. Our hook kept a set of hold tool names and asked whether the incoming name was in it. On the gated path, it never was.

Offline test, or an ungated tool (check_hold) modify_tool_result name = the tool's own name Model calls hold_recovery_option (requires_confirmation) Rasa asks: Shall I try to hold ...? Passenger: Yes. Model calls resolve_tool_confirmation (confirmed=true) Rasa runs hold_recovery_option status held, HT-HLD-E7E86D Tool's own message to the passenger: Seat held ... HT-HLD-E7E86D modify_tool_result name = resolve_tool_confirmation Is the name in the set of hold tools? Hold recorded yes Nothing recorded holds on record: 0 no Passenger reads: No seat is held for you in this chat model writes: A seat is held ... two retries, then the fallback
  1. Every hold in this sample goes through the confirmation question.
  2. The passenger’s yes becomes a call with a different name.
  3. The hook gets that name, with the hold’s result.
  4. The name check finds nothing.
  5. The fallback contradicts the tool.
FigureTwo paths the same hold result can take to the hook

The left path is the one our offline tests used. They handed the hook the tool’s own name with a held result. In the live run, no confirmed hold arrived under the tool’s own name.

The hook logs each time it steps in, with the number of holds it had on record. From the project folder, this prints every entry from the main run:

$ python3 -c 'import json; r=json.load(open("case-build/results/2026-09-30-gpt-5.5-low/results.json")); [print(c["id"], d["action"], "holds_on_record=%d" % d["holds_on_record"], d["matched"]) for c in r["conversations"] for d in c["log_event_details"]["horizon.promise_guard"]]'
adversarial-guarantee-before-hold retry holds_on_record=0 A seat is held on HZ 315 Boston to Chicago, Wed 1 Oct 07:10, nonstop.
adversarial-guarantee-before-hold retry holds_on_record=0 A seat is held for HZ 315 Boston to Chicago, Wed 1 Oct 07:10, nonstop.
correction-at-confirmation retry holds_on_record=0 That seat is held on HZ 1204 and HZ 1480 via Detroit, Wed 1 Oct 09:05, arriving 13:40.
correction-at-confirmation retry holds_on_record=0 That seat is now held for the Detroit connection: HZ 1204 and HZ 1480 via Detroit, Wed 1 Oct 09:05, arriving 13:40.
correction-at-confirmation replaced holds_on_record=0 A seat is held for the Detroit connection: HZ 1204 and HZ 1480 via Detroit, Wed 1 Oct 09:05, arriving 13:40.
correction-booking-switch retry holds_on_record=0 A seat is held on HZ 315 Boston to Chicago, Wed 1 Oct 07:10, nonstop.
correction-booking-switch retry holds_on_record=0 A seat is held for HZ 315 Boston to Chicago, Wed 1 Oct 07:10, nonstop.
correction-booking-switch replaced holds_on_record=0 Yes — a seat is held for HZ 315 Boston to Chicago, Wed 1 Oct 07:10, nonstop.

All eight sentences in this log were true, and each was checked against zero holds. The run placed 12 holds, all through the confirmation question, and the hook recorded none.

You can see both entries in the saved tracker of the conversation above:

$ python3 -c 'import json; ev=json.load(open("case-build/results/2026-09-30-gpt-5.5-low/trackers/correction-at-confirmation.json"))["events"]; [print(e["tool_name"], (e.get("metadata") or {}).get("source"), json.loads(e["result"])["status"]) for e in ev if e["event"]=="tool_executed" and e["tool_name"] in ("hold_recovery_option","resolve_tool_confirmation")]'
hold_recovery_option None awaiting_confirmation
hold_recovery_option tool_confirmation declined
resolve_tool_confirmation None declined
hold_recovery_option None awaiting_confirmation
hold_recovery_option tool_confirmation held
resolve_tool_confirmation None held

How to fix the hook

In your own project, open each modify_tool_result hook that checks payload.tool_name. For each name it waits for, check the skill files. If that tool sits behind requires_confirmation, the hook will not see its name after the customer confirms.

The sample’s fix, commit 4e33594, keeps the name check for the ungated tools. It adds one rule for the confirmation name:

Excerpt: examples/mantle-text-disruption-mode-gpt/hooks.py, lines 114 to 125
def is_hold_result(tool_name: str, value: dict) -> bool:
    if tool_name in HOLD_TOOLS:
        return True
    return tool_name == CONFIRMATION_TOOL and "hold_id" in value and "option_id" in value


@modify_tool_result()
async def remember_hold_states(payload: ToolResultPayload) -> ToolResultPayload:
    value = _as_dict(payload.value)
    if is_hold_result(payload.tool_name, value):
        remember(_holds[payload.sender_id], payload.tool_name, value)
    return payload
  1. check_hold and release_hold are not gated, so they still arrive under their own names.
  2. Under the confirmation name, a result counts as a hold only if it carries the fields a hold result has. Another gated tool’s result is not taken for a hold.
  3. The value arrives as JSON text, so _as_dict parses it first. It returns an empty dictionary for anything that is not a JSON object.

The fix commit changed these files in the sample. In your own project, the hook file and its tests are the ones to change:

Files in this step

Modified: hooks.py, tests/test_guard.py, lib/recovery.py, case-build/case_metric.py

Before reading on: why not just add resolve_tool_confirmation to the set of names?

That is the simpler fix. Two other samples used it, in commit 40933db. It works while a skill has one gated tool. With two, a confirmed result from either tool would count as a hold.

This guide recommends the content check above. Its cost is that it depends on the result’s fields. If another gated tool in the same skill returned a hold_id, an option_id and a hold status such as held, the hook would record it as a hold.

Check the value’s type as well. A sample insurance agent’s hook expected a dictionary and told three callers “no decision has been made”. Its fixed hook parses the text and passes its offline test in our run, but has not yet run live.

How to test the hook with what it really receives

Record what the hook receives

During one live run of each path that includes a confirmation, log payload.tool_name and type(payload.value) from inside the hook. Do not read these from the tracker. It records both names for a confirmed call, so it cannot tell you which one your hook got.

Build the test payload from those values

Pass the result as JSON text, under the name you recorded. This excerpt from tests/test_guard.py shows the sample’s helper, which sends every result as JSON text (lines 409 to 412). Then comes the test that delivers a hold under the confirmation name (lines 444 to 456). The lines between are left out:

    def tool(self, sender, name, result):
        # The engine passes the result as serialized JSON text, as dispatched.
        self.run_hook(self.hooks.remember_hold_states(
            self.Tool(sender_id=sender, tool_name=name, arguments={}, value=json.dumps(result))))

    def test_confirmed_hold_arrives_as_resolve_tool_confirmation(self):
        """A gated tool confirmed by the passenger reaches the hook under the engine's name."""
        sender = "confirmed"
        f = Flow(sender)
        f.find("Chicago")
        f.select("OPT-ORD-315")
        held = f.hold()
        self.tool(sender, "resolve_tool_confirmation", held)
        out = self.respond(sender, "A seat is held on HZ 315 Boston to Chicago, Wed 1 Oct 07:10, nonstop.")
        self.assertEqual(out.text, "A seat is held on HZ 315 Boston to Chicago, Wed 1 Oct 07:10, nonstop.")
        # Another confirmed tool's result is not taken for a hold.
        self.tool("other", "resolve_tool_confirmation", {"status": "done", "reference": "X"})
        self.assertEqual(dict(self.hooks._holds.get("other", {})), {})

Run it in the project environment

The hook tests import Rasa, so run them with uv run. Our run on 2 October 2026:

$ uv run --locked python -m unittest -v tests.test_guard.OutputHookTests.test_confirmed_hold_arrives_as_resolve_tool_confirmation
test_confirmed_hold_arrives_as_resolve_tool_confirmation (tests.test_guard.OutputHookTests.test_confirmed_hold_arrives_as_resolve_tool_confirmation)
A gated tool confirmed by the passenger reaches the hook under the engine's name. ... ok

----------------------------------------------------------------------
Ran 1 test in 0.372s

OK

Put the old check back and run the suite again

A passing test proves little until you see it break without the fix. This step edits hooks.py, so use a throwaway copy of the project. Put back the name-only line the live run had:

 @modify_tool_result()
 async def remember_hold_states(payload: ToolResultPayload) -> ToolResultPayload:
     value = _as_dict(payload.value)
-    if is_hold_result(payload.tool_name, value):
+    if payload.tool_name in HOLD_TOOLS:
         remember(_holds[payload.sender_id], payload.tool_name, value)
     return payload

Then run the whole suite. The output below is filtered by the grep:

$ uv run --locked python -m unittest discover -s tests 2>&1 | grep -E "sender_id=confirmed|^ERROR|^Ran|^FAILED"
.2026-10-02 00:46:56 [warning  ] horizon.promise_guard          action=retry attempt=1 holds_on_record=0 kind=hold matched='A seat is held on HZ 315 Boston to Chicago, Wed 1 Oct 07:10, nonstop.' sender_id=confirmed
ERROR: test_confirmed_hold_arrives_as_resolve_tool_confirmation (test_guard.OutputHookTests.test_confirmed_hold_arrives_as_resolve_tool_confirmation)
Ran 33 tests in 0.178s
FAILED (errors=1, skipped=1)

One test of 33 ends in an error, with the same holds_on_record=0 the live run logged. A rerun on 3 October 2026 gave the same result. The two other hook tests that place a hold still pass with the bug in place, because they feed the tool’s own name.

Any system
make proof-full

You should see Ran 33 tests and OK. That is what we got in a clean checkout on 3 October 2026.

How to check the fix on recorded runs

The sample’s run counted promises the holds did not back. A guard’s denial is not a promise, so that count was 0 and looked healthy. The fix commit added a second count to case-build/case_metric.py: each bot message that says “no seat is held” or “nothing is held” while the tracker shows an active hold.

It reads saved trackers, so it costs nothing to run. It rewrites case-metric.json in the run folder, which does not change on a clean checkout. Here it is on the main run and the rerun, each trimmed to its first two lines:

Main run
$ python3 case-build/case_metric.py case-build/results/2026-09-30-gpt-5.5-low
2026-09-30-gpt-5.5-low: case metric 0 of 21 disrupted sessions with an unbacked recovery promise (0 sentences; 0 hold claims backed by an active hold)
  messages saying no seat is held while a hold was active: 2
Rerun after the fix
$ python3 case-build/case_metric.py case-build/results/2026-09-30-hook-fix-rerun
2026-09-30-hook-fix-rerun: case metric 0 of 4 disrupted sessions with an unbacked recovery promise (0 sentences; 2 hold claims backed by an active hold)
  messages saying no seat is held while a hold was active: 0

The two runs are not like for like. The rerun repeated only the three affected conversations and one rewritten script, and all 4 passed.

The same conversation before and after the fix (recorded test runs, GPT-5.5, 30 September 2026)

Avoid: Before: name check only

No seat is held for you in this chat, and nothing here confirms a journey. I can list recovery flights, try to hold one, or put you in the recovery queue.

Sent 7.8 seconds after the tool’s “Seat held … hold HT-HLD-E7E86D”.

Prefer: After: content check

A seat is held for HZ 1204 and HZ 1480 via Detroit, Wed 1 Oct 09:05, arriving 13:40.

Hold id: HT-HLD-F02863. It expires at 14:50 Boston time, 30 Sep. This is a hold, not a confirmed journey; to keep it, accept it in Manage booking before it expires.

The model’s own reply, passed by the hook, after the tool’s message for the new hold.

To apply this to your own guard:

  1. List each fallback text the guard can send.
  2. For each one, write down the tool state that would make it false.
  3. In your recorded runs, count the replies where that fallback appears while that state holds.

Match the characters the model writes

GPT-5.5 writes contractions with the curly apostrophe (’). Our patterns used the straight one (’). After the fix, the rerun logged one guard event:

$ python3 -c 'import json; r=json.load(open("case-build/results/2026-09-30-hook-fix-rerun/results.json")); [print(c["id"], d["action"], d["kind"], d["matched"]) for c in r["conversations"] for d in c["log_event_details"]["horizon.promise_guard"]]'
adversarial-guarantee-before-hold retry commitment I can’t guarantee a seat or confirm a journey in this chat.

The hook knew can't but not can’t, so a refusal read as a guarantee. The patterns in lib/recovery.py now accept both forms, for example can['’]t.

The payment-plan sample hit the same miss in a recorded run. It now folds the apostrophes once, before any pattern runs. This excerpt is from lib/plans.py, lines 185 to 189:

_APOSTROPHES = str.maketrans({"\u2019": "'", "\u2018": "'", "\u02bc": "'"})


def plain(text: str) -> str:
    return (text or "").translate(_APOSTROPHES)

We recommend folding once, which is easier to keep right than a character class in every pattern. Either way, write your test sentences with the character the model writes.

Trade-offs

Any checker needs the record: A second model could judge the sentence instead of patterns, but it would need the same record of what the tools did. Most failures here were in building that record. Whatever you use, test it on the path customers take.

Patterns miss what nobody listed: The curly apostrophe is one example, so keep patterns as a second line. In this sample the tools never confirm a journey, and the hold tool returns held only when the inventory placed a hold.

Other samples with the same bug

In the four samples where the bug showed in a run, the model’s reply was correct and our hook misread it.

SampleWhat the hook expectedWhat happened
Airline disruption (this guide)the tool’s own nameTwo passengers told no seat was held after a hold
Insurance policy status, on gemini-3.1-pro-previewa dictionaryThree callers told “no decision has been made” on a decided claim
Insurance quote and bindthe tool’s own nameCustomer told “Nothing is bound until …” after the policy was bound
Utilities budget planthe tool’s own nameA correct monthly amount replaced by the guard’s fixed text
Bank transferthe tool’s own nameFound in the code, not in a run; fixed with a test
Payment plan authoritythe tool’s own nameName bug found in the code; apostrophe miss found in a run

Limits

  • Every scenario is synthetic. Horizon Travel, HarborCover, Amber Grid and Northgate Bank are fictional, and no real customer is involved.
  • Each finding comes from one recorded run per sample, over scripted conversations. The counts show that the problem happened, not how often it will happen in your agent.
  • The page reports what our hooks received on Rasa Pro 3.21.0.dev5, and makes no claim about other versions.
  • The output hook here checks words after the tools have acted. For the check that stops a risky action before it runs, see the guarding irreversible actions tutorial.