One boring task, minimal access, and an honest way to tell whether it helped.
"We're going to start using AI in the QA team." It is a sentence that gets said in a lot of rooms, and it tends to end one of two ways.
Either the experiment quietly dies a few weeks later, because nobody could tell whether it was helping. Or somebody gives an agent far too much access far too soon, and the first real lesson arrives the hard way.
There is a calmer route between those two, and it starts smaller than most people expect.
The goal of week one is not to prove AI is amazing. It is to find out, honestly, whether it helped.
Instead of asking what AI can do in general, start from how testing already works in your team and ask where help would make sense at each step. It gives the whole conversation something solid to stand on.
Choose a stage of the testing process. On the left, where an agent can assist. On the right, what still belongs to a person.
At every stage the accountability stays with the tester. AI can make the work faster, but it cannot carry the responsibility for it.
Choose a task that is repetitive and low in risk. Drafting test cases from one user story is a good candidate. So is summarizing the log of a single failed test. Stay away from anything that touches production while you are still learning.
Give the agent as little access as it can work with. At the start that means it can read, and nothing more. Write access and a place in real environments can come later, after you have evidence that the results deserve that trust.
A week is enough to learn a great deal, and it can look like this.
Take one user story and ask for draft test cases.
Rate every test case honestly and keep the tally.
Repeat on a second story, with a better request.
Repeat on a third story and notice what is improving.
Look at the numbers before deciding what comes next.
Without a count, you cannot tell whether AI is helping or only feeling helpful. Try the scorecard below with the outputs from your own week. Rate each one as right as it came, usable after a small fix, or wrong.
Rate ten outputs and see what the tally suggests. The thresholds are my own rule of thumb, not research, so treat the result as a prompt for discussion.
The most common mistakes are easy to see once someone names them.
Start small, measure, then widen. It is slower than the hype and a lot faster than repairing the damage.
Before an agent touches your systems, work out what the worst day could look like.
Written by Aditya Mirza Bahari, 2026. The scorecard thresholds are a rule of thumb, not research.