Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q47HardScenario

Your agent reports "Done! The record has been updated," but the record was never updated. How do you detect and prevent this?

30-second answerSay your answer out loud first, then reveal.

Root causes to check in traces

  1. No tool call at all: the model "imagined" doing it. Common when instructions push for helpfulness, or when the tool wasn't available.
  2. Tool error ignored: the result said {"status":"error"} or was empty, and the model glossed over it.
  3. Ambiguous success: the tool returned 200 OK but the write didn't persist (async queue, validation dropped it).
  4. Truncated context: the error was cut off during compaction.

Prevention

  1. Explicit tool results: {"updated": true, "record_id": "...", "new_status": "closed"} or a clear error, never an empty body.
  2. Post-condition checks in code: after update_record, read it back. The tool returns success only if the post-condition holds.
  3. Ground the final response: the user-facing confirmation is generated from the verified tool result (template or structured), not free-form model text.
  4. Action ledger: the UI shows actions actually performed, from logs. The model's narration is secondary.
  5. Prompt: "Only report actions as completed if the tool confirmed success; otherwise say what failed."
  6. Evals: compare claimed outcomes with environment state; track a "false success rate" metric. This is one of the most damaging failure modes for user trust.

You understood something today that you didn't yesterday.