Home Framework Assessment Learning Insights Consult CTFL Exam Free Assessment
NgobrolQA Insight

What Can Your AI Agent Break?

Before an agent touches your code or your systems, ask what the worst day could look like.

NgobrolQA • 4 min read

Once an AI agent can run commands, read your repository and call your APIs, the question changes. It stops being how clever the agent is and becomes how much it is allowed to touch.

Security people have had an answer to this for a very long time. It is called least privilege, and it fits in one sentence: give something only the access it genuinely needs, and nothing more.

It sounds almost too simple to need an article. The reason it does is that real access tends to grow in the cracks of a busy week, one convenient shortcut at a time. Two illustrations show how that looks.

• • •

Two ways access quietly grows

These are illustrations built from patterns that turn up in many teams, not accounts of one particular company.

First, picture a small startup moving fast. A developer needs the payment integration to work before a demo, so they put the provider's API key and an admin password into a configuration file and commit it "just for now". The repository is public. Nobody notices, because everything works.

Anyone who finds that file now holds the keys. They do not need to open the admin page or log in at all. They can call the backend directly, change or delete content, read data meant to be private, or spend money on the company's account. Nothing was hacked. The door was simply left open, and nothing in the daily routine would ever show it.

Second, picture a team that gives an AI agent a database credential so it can "clean up test data". The credential works on production as well as on the test environment, because that was the credential already lying around. One misread instruction later, the agent runs a delete against the wrong place. It followed its instructions exactly. The access was the mistake.

AI agents raise the stakes in two ways. They read repositories, and they often carry what they read into prompts, logs and output. And because people now build faster, the habit of "I'll tidy it up later" has less and less time to be tidied.

The question worth asking about any agent: if it gets something wrong, or is steered wrong, what is the worst it could do with the access it has right now?

• • •

What it looks like to widen access

Try giving an imaginary agent some access and watch the worst-case damage move. Then switch on a few guardrails and see how much of it they take back.

Interactive

How big is the blast radius?

An illustration, not a standard. The weights are my own judgment, meant to make the trade-off visible.

What the agent can do
What you put around it
ContainedSeriousSevere

• • •

What a team can change afterwards

The fixes are rarely exotic. Secrets move out of the code and into settings that live only on the server, or into a proper secret manager. If a secret is missing, the system refuses every request instead of letting everyone in. Pages stop carrying passwords at all, so the visitor types one and the server checks it. Wrong guesses are slowed down so that guessing is not worth the effort.

A leaked secret is treated as leaked for good and replaced, even after the file is cleaned, because version control remembers everything. And any content the site displays from users or admins is rebuilt from a short list of allowed tags, so that even if a password leaks again, an attacker still cannot run a script on the site.

Before you hand an agent the keys

Make sure there are no secrets in the repository, in prompts, or in any configuration file the agent can read. Keep them in a secret manager or in environment settings. Give the least access that works: reading first, writing later, production last, with one credential per purpose and the narrowest scope you can manage.

Let the agent work in a sandbox or on staging rather than directly on production. Ask a person to approve anything that is hard to undo, like deploying, deleting, sending something outside, or changing permissions. Keep a record of what the agent did and when. Treat what comes back from tools as untrusted, since outside text can hide instructions. And have a plan for changing a leaked secret quickly, ideally without a redeploy.

  1. Where could a secret be sitting that this agent can read?
  2. What is the smallest access that still lets it do the job?
  3. Who has to say yes before anything irreversible happens?

Capability is what makes an agent useful. Access is what makes it dangerous. Decide on the second as carefully as you admire the first.

Test it before you trust it

See why green tests are not the same as a working feature, with four real bugs from an AI-built tool.

Written by Aditya Mirza Bahari, 2026. The scenarios are illustrations of common patterns, and the risk weights in the demo are a judgment call, not a standard.