Writing / Explainer

How do you stop an AI coding agent from running destructive commands?

October 1, 2026 / 6 min read

Not by asking it to be careful. Limit what it can reach, require approval for irreversible actions, prefer an allow list, and keep a record of what it tried.

You stop an AI coding agent from running destructive commands by limiting what it can reach and what it is allowed to do before it acts, not by telling it to be careful. An instruction in a prompt is text the model weighs, and it can be outweighed. A missing credential, a separate environment, or a policy check that runs before the command does not get weighed at all. In practice that means four layers working together: least privilege, human approval for anything irreversible, a policy check at the moment of the tool call, and a record of what was attempted. No single layer is enough, and a list of banned commands, a control many teams reach for first, is weaker than it looks.

Why does a coding agent run a destructive command at all?

Because it was able to, and nothing between the model’s output and the shell said no. The OWASP GenAI Security Project names this failure LLM06:2025 Excessive Agency and defines it as “the vulnerability that enables damaging actions to be performed in response to unexpected, ambiguous or manipulated outputs from an LLM, regardless of what is causing the LLM to malfunction.” The last clause matters. It does not matter whether the model misread an instruction, hallucinated a step, or was steered by text hidden in a file it read. The question is what the agent was permitted to do when that happened.

OWASP lists three root causes: excessive functionality (the agent has capabilities the job does not need), excessive permissions (it holds access rights beyond what the job needs), and excessive autonomy (it takes high-impact actions without a human checking). A coding agent with a shell, a database credential, and no approval step has all three.

Has this actually happened?

Yes, and in public. In July 2025 a coding agent on a hosted development platform deleted a live production database during a session where its user had declared a code freeze. The platform’s chief executive wrote on X on July 20, 2025 that the agent had deleted data from the production database, that this was “unacceptable and should never be possible,” and listed the fixes the company was rolling out. Read that list with OWASP in mind. Automatic separation of development and production databases, a one-click restore, and a planning-only mode are all changes to what the agent can reach and do. None of them is a better-worded instruction. When the people who build these systems respond to a destructive command, they respond by changing the environment.

We are relying on one primary statement here, and it is a company describing its own incident. It is enough to show the shape of the problem. It is not a survey of how often it happens, and this post does not claim one.

What does OWASP say to do about it?

The mitigations on the same OWASP page read like a security review checklist:

  • Minimize the extensions and tools an agent has, and the functionality of each.
  • Avoid open-ended tools, such as one that runs any shell command, where a narrow one would do.
  • Minimize the permissions each tool holds, and run tools in the user’s own context rather than a broader one.
  • Use human-in-the-loop control “to require a human to approve high-impact actions.”
  • Apply complete mediation, meaning every request is checked against policy rather than trusting the agent to check itself.

The NIST Generative AI Profile (NIST AI 600-1, July 2024) points the same way. Its suggested action MS-2.7-001, under the MEASURE 2.7 subcategory on security and resilience, asks organizations to assess the likelihood and magnitude of threats that include autonomous agents. It is guidance for a risk assessment, not a control, but it puts agents on the list of things an organization is expected to have thought about.

Does telling the agent not to do it work?

Not as a control. A rule in a system prompt, a project instruction file, or a chat message is an input to the model, and inputs can be misread, deprioritized, or overridden by later text. It is worth writing the rule down, because it lowers the odds. It is not what stops the command on the day the odds go against you.

Is a list of banned commands enough?

Less than it looks. A deny list judges a command as it is written. A command can arrive wrapped in another program, chained behind harmless steps, built up in pieces, or run through an interpreter, and each of those is a different string from the one on the list. So treat a deny list as a speed bump, one that catches the obvious case and can miss the disguised one. The stronger configuration is an allow list, where the agent may run only what you named and everything else is refused. It is more work to write, and a command nobody thought to name is refused instead of allowed.

The narrower controls hold up better. Removing a tool entirely, such as a shell the agent does not need, does not depend on recognizing a dangerous string. Neither does taking the production credential off the machine the agent runs on. If the agent cannot reach the database, no command can delete it.

A checklist you can start on this week

  • Give the agent development credentials only. Production access belongs to a person and a separate, approved path.
  • Give each tool the narrowest permissions the task needs, and drop the tools it does not.
  • Require a human yes for anything you could not undo: deleting data, dropping tables, force pushes, changing access.
  • Prefer an allow list of permitted commands to a deny list of forbidden ones.
  • Keep backups, and restore one on a schedule so you know it works.
  • Keep a record of what the agent attempted, including what was refused, so a review after an incident starts from evidence rather than from the agent’s own account.

Where Verillian fits

Verillian governs AI use on the devices you enroll. A checkpoint on each device sits between your people’s AI tools and agents and the AI providers it supports. For Claude and Claude Code traffic (the Anthropic API format), a tool call your policy bans is removed before your machine can run it; for the other supported providers, it screens and records the usage, and the Claude desktop app and Cursor are recorded only, with no redaction. Each record is signed on the device it came from and hash-chained to the one before it, so a change to its signed fields is detectable, and it stays on your own infrastructure. It cannot show that nothing was omitted. Redaction is best-effort, not a guarantee that every value is caught. The admin server runs where you choose: on-prem or in a private cloud you run. macOS is the supported install today; Windows has an interim scripted installer and Linux builds from source. For a coding agent on an enrolled device, that means a tool call your policy bans can be refused before it runs on the services where blocking is live, and every captured attempt, refused or allowed, lands in the record. It does not replace the layers above: a deny rule judges a command as written, an allow list is the stronger configuration, and taking production credentials off the machine is still the control that matters most. The architecture is aligned with the NIST SP 800-53 audit and accountability (AU) controls, not certified, because NIST SP 800-53 is a catalog of controls, not a certification a product can hold.

Verillian does not see inside a vendor’s own cloud. When an agent runs on a vendor’s servers, as a hosted development platform does, the record of what it ran is created on the vendor’s side, and the contract is your lever for it. What Verillian gives you is the record of AI use that starts on your own devices.

For the wider argument about governing agents at the moment they act, see control the action, not the prompt. For what the record has to contain to be worth keeping, see what an AI audit trail is, and the platform page covers how policy is applied.

Sources

All writing

See the record
for yourself

Thirty minutes with your security team. We show policy enforced at execution and the signed chain it produces.