Four illustrative scenarios of green tests and a broken product, and why no test could see them.
Picture a small team shipping a checkout feature for an online marketplace. They are using an AI coding agent, and it is going well. The agent writes the code, writes the tests, runs them, and reports that everything passes.
It is a very good feeling, the kind where you almost don't want to look at the thing, in case looking changes the answer. So someone looks. They use the feature by hand for an afternoon and find four problems, and not one of them had been caught by a single test.
The four scenarios below are illustrations. They are built from failure patterns that turn up again and again in software, not accounts of one particular company. They are here because each one hides in a different place, and that is the whole point.
The code was correct according to its own tests. That was exactly the problem.
None of this is a story about a careless tool. The tests are doing their job. Open the scenarios one at a time and look at what the tests said next to what a person found.
Tap a scenario to see what the tests were happy about, what a person ran into, and why the two never met.
One was a gap between two correct parts, one a misleading error, one a quirk of the data, and one a matter of feel. A test suite only checks the cases someone thought to write down. Every one of these needed a person to run the real thing.
Testers have a name for what went wrong in all four. A test needs an oracle, some way of knowing what the right answer is. The tests an AI agent writes use an oracle too, and it is the agent's own idea of what the feature should do. When that idea has a gap, the tests carry the same gap, and they pass with complete confidence.
Deciding what "correct" means for a real person using a real screen is the part that stays human. It is judgment, and judgment is not on the list of things a test can do for itself.
Green tells you the code agrees with the tests. It says nothing about whether the tests agree with the world.
This is not an argument against AI coding agents. They are fast, patient, and thorough in a way most of us are not at the end of a long day. Asked for a broad set of checks on a backend, an agent will produce them quickly: wrong passwords, missing fields, oversized input, the cases a tired person means to write and never gets around to.
That is where these tools are strong. They are wide. They cover a lot of ground quickly, including the dull corners. The places they struggle are the deep ones, where someone has to know what the thing is for and how a person will really use it.
There is a practical side to this as well. When a person reports a problem precisely, with the exact error text and the exact steps, an agent can usually fix it fast. A vague "it's broken" gets a guess. A precise report gets a repair. How well you describe a problem decides how quickly it goes away.
People often ask where "human review" goes when an agent writes the code. A simple answer that works: every change lands on its own branch and in a pull request, someone runs it and clicks through it like a real user would, and only then does it merge, and only on a person's say-so.
That last step is the one that matters. It is a gate where a person, not a script, decides that this is good enough to go live.
If you are working with AI-written tests, three questions are worth carrying into the room.
AI makes writing code and tests much faster. It does not remove the need for someone to run the real flow and ask whether it actually works.
Working with AI in your QA team and not sure where the human checks belong? Let us ngobrol about it.
Written by Aditya Mirza Bahari, 2026. The scenarios are illustrative composites of common failure patterns, not reports of any specific company.