1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
Your agent reports "Done! The record has been updated," but the record was never updated. How do you detect and prevent this?
30-second answerSay your answer out loud first, then reveal.
Root causes to check in traces
- No tool call at all: the model "imagined" doing it. Common when instructions push for helpfulness, or when the tool wasn't available.
- Tool error ignored: the result said
{"status":"error"}or was empty, and the model glossed over it. - Ambiguous success: the tool returned
200 OKbut the write didn't persist (async queue, validation dropped it). - Truncated context: the error was cut off during compaction.
Prevention
- Explicit tool results:
{"updated": true, "record_id": "...", "new_status": "closed"}or a clear error, never an empty body. - Post-condition checks in code: after
update_record, read it back. The tool returns success only if the post-condition holds. - Ground the final response: the user-facing confirmation is generated from the verified tool result (template or structured), not free-form model text.
- Action ledger: the UI shows actions actually performed, from logs. The model's narration is secondary.
- Prompt: "Only report actions as completed if the tool confirmed success; otherwise say what failed."
- Evals: compare claimed outcomes with environment state; track a "false success rate" metric. This is one of the most damaging failure modes for user trust.
Related
You understood something today that you didn't yesterday.