Why it matters
An agent’s summary is a claim. It can be wrong in confident, specific ways, including test counts, file names, and “I verified this.” See hallucination.
Riker’s rule: a worker saying it’s done is not proof. Riker runs the check again itself, and keeps the evidence of what really happened.
Watch for one sneaky case: the agent runs the check, it passes, and then the agent makes “one tiny fix.” Now nothing proves the final version.
Exercise: ask for receipts
Add this to the end of your brief from the last lesson:
When you finish, don't summarize. Instead give me:
- the exact command you ran for each check, and its full output
- screenshots for anything visual
- the list of files you changed
Run every check again after your last edit.
When it finishes, pick one check and run it yourself.
Check yourself: did your run match what the agent reported? If it didn’t, you just found the gap these habits exist to close.