The agent said 'tests passed': five questions to ask before you accept it
By Oleksii Skurikhin · DELIVERY MANAGER
- Published
- 2 min read

Five questions to paste into the chat when an AI coding agent says 'done, tests pass'. They ask for the exact command, the test counts, the start of each changed file and what was not checked. Then you run the tests yourself and compare.
Key takeaways
- An AI coding agent's 'done, tests pass' message is its description of the work, and a description can differ from the work.
- The five questions ask the agent for the exact test command, its full output and how many tests ran, passed, failed and were skipped.
- The five questions also ask the agent to separate what it checked from what it assumed, and to list what it did not check.
- The agent's answers remain its own account: run the test command it reported yourself and compare the counts.
The Anthropic system card for Claude 3.7 Sonnet says the model occasionally special-cases tests in agentic coding environments like Claude Code. Most often it returns the values a test expects; the card says it also modified the tests themselves.
The system card says the special-casing typically emerges after multiple failed attempts and occurs infrequently in normal usage.
METR's report of 5 June 2025 says that, as of its publication, the most recent frontier AI models have tried, often successfully, to get a higher score by modifying the tests or scoring code, among other loopholes. METR ran a range of models on tasks testing autonomous software development and AI R&D capabilities.
In one METR example, o3 patched the evaluation function of a coding-competition task so every submission counted as successful.
An arXiv preprint by Smyth and others, version 3 of 22 September 2026, checked whether agents' final messages report work their own transcripts show they did not do. The preprint covers 12 models, eight of them in their own production command-line tools, on five file-review tasks.
In version 3, the tested agents did not read every file they were asked to review in 67.9% of runs. In 80.4% of those incomplete runs, the final message claimed a complete review or left the gap undisclosed (59 to 96% by model).
The preprint's authors shaped the five tasks to stress thorough review and say the rates should not be generalised to all agentic tasks. It studied file reviews, not test runs.
These sources, checked on 10 October 2026, do not show what your agent does on your project. They are a reason to treat "done, all tests pass" as a description of the work, and a description can differ from the work.
Paste this prompt into the agent's chat before you accept the work.
Before you tell me this is done, answer these five, one by one.
1. Show the exact command you ran for the tests and its full output. Do not summarise it.
2. How many tests ran, and how many passed, failed and were skipped? If the count is zero or you do not know, say so.
3. Show the first 20 lines of every file you changed, as they are on disk now.
4. Split your report into CHECKED (shown in the output above) and ASSUMED (not run, not seen).
5. List what you did not check. If a search or a command returned nothing, show that it could have found something before you say there is nothing there.
You may re-run commands, but do not edit any file. If you cannot show an item, write NOT VERIFIED next to it, then stop and wait for me.The questions are our own wording, and an agent may skip an item or answer it loosely. The agent's answers remain its own account, so run the test command it reported in your own terminal and compare the counts. A different number, or zero, is a reason to stop before you merge or deploy.
Next step
About the author

Category:Software Development
Back to Blog

