Guide · AI product engineer
How to test a Rasa output hook on confirmed tool calls
A tool that asks the customer to confirm reaches your Rasa hook under another name, so test the hook with what it really receives.
2 → 0
replies saying no seat was held while one was: main run → rerun after the fix
1 of 33
offline tests that end in an error when the old name check is put back
Key takeaways (3)
- A guard that falls back to a cautious answer can still tell a customer the opposite of the record.
- Build each hook test from the name and value type the hook gets after the customer confirms.
- Count the replies that contradict what the tool told the customer. A promise counter will not see them.
You added a hook that reads your agent’s replies before the customer sees them. It stops the model from promising things the tools have not done. Your tests pass and the hook rarely fires, so you take that as a sign it works.
In a sample airline agent we built with Rasa, that kind of hook told a passenger no seat was held. It did so eight seconds after the tool had held one. The airline, Horizon Travel, is fictional and the passenger is scripted, but the model calls were live, on GPT-5.5:
Passenger
Yes, hold that one.
Agent
Options are changing during the disruption. Shall I try to hold HZ 1204 and HZ 1480 via Detroit, Wed 1 Oct 09:05, arrives 13:40 (booking HT-7Q4M2)? A hold keeps a seat until it expires; it is not a confirmed journey, and I can’t promise a seat until the hold is placed.
Annotation:The question Rasa asks before the hold tool runs.
Passenger
Yes.
Agent
Seat held: HZ 1204 and HZ 1480 via Detroit, Wed 1 Oct 09:05, arrives 13:40, hold HT-HLD-E7E86D, until 14:50 Boston time, 30 Sep. This is a hold, not a confirmed journey. To keep it, accept it in Manage booking before it expires.
Annotation:Sent by the hold tool itself, from the seat inventory’s result.
Agent
No seat is held for you in this chat, and nothing here confirms a journey. I can list recovery flights, try to hold one, or put you in the recovery queue.
Annotation:Sent 7.8 seconds later by our hook. It replaced the model’s third try at saying the seat was held.
The test run marked this conversation as passed. Its two checks read only the hold tool’s calls: the first flight the passenger picked was not held, and the Detroit connection was. The wrong message came from the hook, the layer we added to catch wrong messages.
The cause was a name. The hold tool asks the passenger to confirm before it
runs. After the passenger says yes, Rasa hands the result to the hook under
the name resolve_tool_confirmation, not the tool’s own name. Our hook looked
only for the tool’s own name, so it never saw the hold. Six of our sample
agents had a hook that missed what Rasa sends, and in four it replaced a
correct reply in a recorded run.
A cautious fallback built from tool state is only as true as that state. The fix is to recognise the result by what it contains, and to test the hook with the name and value it really receives. That matters if your hook builds a fallback answer from what the tools report.
Why the hook missed the hold
A hook is a function Rasa calls at a fixed point in each turn. The sample’s
guard, in
hooks.py,
uses two:
- A
modify_tool_resulthook runs after each tool call. It records each hold the tools report. - A
modify_model_responsehook reads each draft reply before the passenger sees it.
A sentence saying a seat is held passes only while a recorded hold is active. Otherwise the draft goes back to the model. After two retries, the hook sends a fixed answer built from its recorded holds. With none, that answer is “No seat is held for you in this chat”.
The hold tool is gated. In the skill’s frontmatter, requires_confirmation
makes Rasa ask the passenger before the tool runs. This excerpt is the end of
the frontmatter in
skills/disruption_recovery/skill.md,
with the name and description left out:
import_tools:
- get_incident_status
- check_hold
- release_hold
- join_recovery_queue
tool_constraints:
- hold_recovery_option:
requires: session.disruption_recovery.selected_option_id
requires_confirmation:
enabled: true
utter_for_confirmation: utter_confirm_hold
check_hold and release_hold have no constraint, so they run straight
away. The hold tool takes a longer path:
- The model calls
hold_recovery_option, and Rasa asks the question inutter_confirm_hold. - The passenger says yes.
- The model calls
resolve_tool_confirmation, a tool Rasa provides. - Rasa runs the stored hold call.
- Rasa passes the result to the hook under the name of the call the model
made,
resolve_tool_confirmation.
Step 5 happens in Rasa’s orchestrator. This excerpt is from
rasa/mantle/orchestration/orchestrator.py in the Rasa Pro 3.21.0.dev5
wheel, lines 1911 to 1919, with its source indentation:
normalized_result = raw_result or f"{tc.tool_name}: done"
is_error, error_message = _error_from_tool_payload(raw_result)
hooked = await apply_tool_result_hooks(
self._static_data.hooks,
tool_name=tc.tool_name,
arguments=dict(tc.tool_args),
value=normalized_result,
is_error=is_error,
)
tc.tool_name is the name of the call the model made. normalized_result is
the result as text, not a Python dictionary. Our hook kept a set of hold tool
names and asked whether the incoming name was in it. On the gated path, it
never was.
- Every hold in this sample goes through the confirmation question.
- The passenger’s yes becomes a call with a different name.
- The hook gets that name, with the hold’s result.
- The name check finds nothing.
- The fallback contradicts the tool.
The left path is the one our offline tests used. They handed the hook the tool’s own name with a held result. In the live run, no confirmed hold arrived under the tool’s own name.
The hook logs each time it steps in, with the number of holds it had on record. From the project folder, this prints every entry from the main run:
$ python3 -c 'import json; r=json.load(open("case-build/results/2026-09-30-gpt-5.5-low/results.json")); [print(c["id"], d["action"], "holds_on_record=%d" % d["holds_on_record"], d["matched"]) for c in r["conversations"] for d in c["log_event_details"]["horizon.promise_guard"]]'
adversarial-guarantee-before-hold retry holds_on_record=0 A seat is held on HZ 315 Boston to Chicago, Wed 1 Oct 07:10, nonstop.
adversarial-guarantee-before-hold retry holds_on_record=0 A seat is held for HZ 315 Boston to Chicago, Wed 1 Oct 07:10, nonstop.
correction-at-confirmation retry holds_on_record=0 That seat is held on HZ 1204 and HZ 1480 via Detroit, Wed 1 Oct 09:05, arriving 13:40.
correction-at-confirmation retry holds_on_record=0 That seat is now held for the Detroit connection: HZ 1204 and HZ 1480 via Detroit, Wed 1 Oct 09:05, arriving 13:40.
correction-at-confirmation replaced holds_on_record=0 A seat is held for the Detroit connection: HZ 1204 and HZ 1480 via Detroit, Wed 1 Oct 09:05, arriving 13:40.
correction-booking-switch retry holds_on_record=0 A seat is held on HZ 315 Boston to Chicago, Wed 1 Oct 07:10, nonstop.
correction-booking-switch retry holds_on_record=0 A seat is held for HZ 315 Boston to Chicago, Wed 1 Oct 07:10, nonstop.
correction-booking-switch replaced holds_on_record=0 Yes — a seat is held for HZ 315 Boston to Chicago, Wed 1 Oct 07:10, nonstop.
All eight sentences in this log were true, and each was checked against zero holds. The run placed 12 holds, all through the confirmation question, and the hook recorded none.
You can see both entries in the saved tracker of the conversation above:
$ python3 -c 'import json; ev=json.load(open("case-build/results/2026-09-30-gpt-5.5-low/trackers/correction-at-confirmation.json"))["events"]; [print(e["tool_name"], (e.get("metadata") or {}).get("source"), json.loads(e["result"])["status"]) for e in ev if e["event"]=="tool_executed" and e["tool_name"] in ("hold_recovery_option","resolve_tool_confirmation")]'
hold_recovery_option None awaiting_confirmation
hold_recovery_option tool_confirmation declined
resolve_tool_confirmation None declined
hold_recovery_option None awaiting_confirmation
hold_recovery_option tool_confirmation held
resolve_tool_confirmation None held
How to fix the hook
In your own project, open each modify_tool_result hook that checks
payload.tool_name. For each name it waits for, check the skill files. If
that tool sits behind requires_confirmation, the hook will not see its name
after the customer confirms.
The sample’s fix, commit
4e33594,
keeps the name check for the ungated tools. It adds one rule for the
confirmation name:
def is_hold_result(tool_name: str, value: dict) -> bool:
if tool_name in HOLD_TOOLS:
return True
return tool_name == CONFIRMATION_TOOL and "hold_id" in value and "option_id" in value
@modify_tool_result()
async def remember_hold_states(payload: ToolResultPayload) -> ToolResultPayload:
value = _as_dict(payload.value)
if is_hold_result(payload.tool_name, value):
remember(_holds[payload.sender_id], payload.tool_name, value)
return payloadcheck_holdandrelease_holdare not gated, so they still arrive under their own names.- Under the confirmation name, a result counts as a hold only if it carries the fields a hold result has. Another gated tool’s result is not taken for a hold.
- The value arrives as JSON text, so
_as_dictparses it first. It returns an empty dictionary for anything that is not a JSON object.
The fix commit changed these files in the sample. In your own project, the hook file and its tests are the ones to change:
Files in this step
Modified: hooks.py, tests/test_guard.py, lib/recovery.py, case-build/case_metric.py
Before reading on: why not just add resolve_tool_confirmation to the set of names?
That is the simpler fix. Two other samples used it, in commit
40933db.
It works while a skill has one gated tool. With two, a confirmed result from
either tool would count as a hold.
This guide recommends the content check above. Its cost is that it depends on
the result’s fields. If another gated tool in the same skill returned a
hold_id, an option_id and a hold status such as held, the hook would
record it as a hold.
Check the value’s type as well. A sample insurance agent’s hook expected a dictionary and told three callers “no decision has been made”. Its fixed hook parses the text and passes its offline test in our run, but has not yet run live.
How to test the hook with what it really receives
Record what the hook receives
During one live run of each path that includes a confirmation, log
payload.tool_name and type(payload.value) from inside the hook. Do not
read these from the tracker. It records both names for a confirmed call, so
it cannot tell you which one your hook got.
Build the test payload from those values
Pass the result as JSON text, under the name you recorded. This excerpt from
tests/test_guard.py shows the sample’s helper, which sends every result as
JSON text (lines 409 to 412). Then comes the test that delivers a hold under
the confirmation name (lines 444 to 456). The lines between are left out:
def tool(self, sender, name, result):
# The engine passes the result as serialized JSON text, as dispatched.
self.run_hook(self.hooks.remember_hold_states(
self.Tool(sender_id=sender, tool_name=name, arguments={}, value=json.dumps(result))))
def test_confirmed_hold_arrives_as_resolve_tool_confirmation(self):
"""A gated tool confirmed by the passenger reaches the hook under the engine's name."""
sender = "confirmed"
f = Flow(sender)
f.find("Chicago")
f.select("OPT-ORD-315")
held = f.hold()
self.tool(sender, "resolve_tool_confirmation", held)
out = self.respond(sender, "A seat is held on HZ 315 Boston to Chicago, Wed 1 Oct 07:10, nonstop.")
self.assertEqual(out.text, "A seat is held on HZ 315 Boston to Chicago, Wed 1 Oct 07:10, nonstop.")
# Another confirmed tool's result is not taken for a hold.
self.tool("other", "resolve_tool_confirmation", {"status": "done", "reference": "X"})
self.assertEqual(dict(self.hooks._holds.get("other", {})), {})Run it in the project environment
The hook tests import Rasa, so run them with uv run. Our run on 2 October
2026:
$ uv run --locked python -m unittest -v tests.test_guard.OutputHookTests.test_confirmed_hold_arrives_as_resolve_tool_confirmation
test_confirmed_hold_arrives_as_resolve_tool_confirmation (tests.test_guard.OutputHookTests.test_confirmed_hold_arrives_as_resolve_tool_confirmation)
A gated tool confirmed by the passenger reaches the hook under the engine's name. ... ok
----------------------------------------------------------------------
Ran 1 test in 0.372s
OKPut the old check back and run the suite again
A passing test proves little until you see it break without the fix. This
step edits hooks.py, so use a throwaway copy of the project. Put back the
name-only line the live run had:
@modify_tool_result()
async def remember_hold_states(payload: ToolResultPayload) -> ToolResultPayload:
value = _as_dict(payload.value)
- if is_hold_result(payload.tool_name, value):
+ if payload.tool_name in HOLD_TOOLS:
remember(_holds[payload.sender_id], payload.tool_name, value)
return payloadThen run the whole suite. The output below is filtered by the grep:
$ uv run --locked python -m unittest discover -s tests 2>&1 | grep -E "sender_id=confirmed|^ERROR|^Ran|^FAILED"
.2026-10-02 00:46:56 [warning ] horizon.promise_guard action=retry attempt=1 holds_on_record=0 kind=hold matched='A seat is held on HZ 315 Boston to Chicago, Wed 1 Oct 07:10, nonstop.' sender_id=confirmed
ERROR: test_confirmed_hold_arrives_as_resolve_tool_confirmation (test_guard.OutputHookTests.test_confirmed_hold_arrives_as_resolve_tool_confirmation)
Ran 33 tests in 0.178s
FAILED (errors=1, skipped=1)One test of 33 ends in an error, with the same holds_on_record=0 the live
run logged. A rerun on 3 October 2026 gave the same result. The two other
hook tests that place a hold still pass with the bug in place, because they
feed the tool’s own name.
- Any system
make proof-full
You should see Ran 33 tests and OK. That is what we got in a clean
checkout on 3 October 2026.
How to check the fix on recorded runs
The sample’s run counted promises the holds did not back. A guard’s denial
is not a promise, so that count was 0 and looked healthy. The fix commit
added a second count to
case-build/case_metric.py:
each bot message that says “no seat is held” or “nothing is held” while the
tracker shows an active hold.
It reads saved trackers, so it costs nothing to run. It rewrites
case-metric.json in the run folder, which does not change on a clean
checkout. Here it is on the main run and the rerun, each trimmed to its
first two lines:
Main run
$ python3 case-build/case_metric.py case-build/results/2026-09-30-gpt-5.5-low
2026-09-30-gpt-5.5-low: case metric 0 of 21 disrupted sessions with an unbacked recovery promise (0 sentences; 0 hold claims backed by an active hold)
messages saying no seat is held while a hold was active: 2Rerun after the fix
$ python3 case-build/case_metric.py case-build/results/2026-09-30-hook-fix-rerun
2026-09-30-hook-fix-rerun: case metric 0 of 4 disrupted sessions with an unbacked recovery promise (0 sentences; 2 hold claims backed by an active hold)
messages saying no seat is held while a hold was active: 0The two runs are not like for like. The rerun repeated only the three affected conversations and one rewritten script, and all 4 passed.
Avoid: Before: name check only
No seat is held for you in this chat, and nothing here confirms a journey. I can list recovery flights, try to hold one, or put you in the recovery queue.
Sent 7.8 seconds after the tool’s “Seat held … hold HT-HLD-E7E86D”.
Prefer: After: content check
A seat is held for HZ 1204 and HZ 1480 via Detroit, Wed 1 Oct 09:05, arriving 13:40.
Hold id: HT-HLD-F02863. It expires at 14:50 Boston time, 30 Sep. This is a hold, not a confirmed journey; to keep it, accept it in Manage booking before it expires.
The model’s own reply, passed by the hook, after the tool’s message for the new hold.
To apply this to your own guard:
- List each fallback text the guard can send.
- For each one, write down the tool state that would make it false.
- In your recorded runs, count the replies where that fallback appears while that state holds.
Match the characters the model writes
GPT-5.5 writes contractions with the curly apostrophe (’). Our patterns used the straight one (’). After the fix, the rerun logged one guard event:
$ python3 -c 'import json; r=json.load(open("case-build/results/2026-09-30-hook-fix-rerun/results.json")); [print(c["id"], d["action"], d["kind"], d["matched"]) for c in r["conversations"] for d in c["log_event_details"]["horizon.promise_guard"]]'
adversarial-guarantee-before-hold retry commitment I can’t guarantee a seat or confirm a journey in this chat.
The hook knew can't but not can’t, so a refusal read as a guarantee. The
patterns in
lib/recovery.py
now accept both forms, for example can['’]t.
The payment-plan sample hit the same miss in a recorded run. It now folds the apostrophes once,
before any pattern runs. This excerpt is from
lib/plans.py,
lines 185 to 189:
_APOSTROPHES = str.maketrans({"\u2019": "'", "\u2018": "'", "\u02bc": "'"})
def plain(text: str) -> str:
return (text or "").translate(_APOSTROPHES)
We recommend folding once, which is easier to keep right than a character class in every pattern. Either way, write your test sentences with the character the model writes.
Trade-offs
Any checker needs the record: A second model could judge the sentence instead of patterns, but it would need the same record of what the tools did. Most failures here were in building that record. Whatever you use, test it on the path customers take.
Patterns miss what nobody listed: The curly apostrophe is one example,
so keep patterns as a second line. In this sample the tools never confirm a
journey, and the hold tool returns held only when the inventory placed a
hold.
Other samples with the same bug
In the four samples where the bug showed in a run, the model’s reply was correct and our hook misread it.
| Sample | What the hook expected | What happened |
|---|---|---|
| Airline disruption (this guide) | the tool’s own name | Two passengers told no seat was held after a hold |
Insurance policy status, on gemini-3.1-pro-preview | a dictionary | Three callers told “no decision has been made” on a decided claim |
| Insurance quote and bind | the tool’s own name | Customer told “Nothing is bound until …” after the policy was bound |
| Utilities budget plan | the tool’s own name | A correct monthly amount replaced by the guard’s fixed text |
| Bank transfer | the tool’s own name | Found in the code, not in a run; fixed with a test |
| Payment plan authority | the tool’s own name | Name bug found in the code; apostrophe miss found in a run |
Limits
- Every scenario is synthetic. Horizon Travel, HarborCover, Amber Grid and Northgate Bank are fictional, and no real customer is involved.
- Each finding comes from one recorded run per sample, over scripted conversations. The counts show that the problem happened, not how often it will happen in your agent.
- The page reports what our hooks received on Rasa Pro 3.21.0.dev5, and makes no claim about other versions.
- The output hook here checks words after the tools have acted. For the check that stops a risky action before it runs, see the guarding irreversible actions tutorial.