Add a Zeroth Law to your AI agent before it deletes something

A MacBook with lines of code on its screen on a busy desk

When you tell a human coworker to clean up a module, you don’t also tell them not to drop the production database. They already know. An AI agent doesn’t. It runs on a permissive default: everything that isn’t forbidden is allowed, and without common sense filling the gaps, that default is dangerous.

Software engineer and O’Reilly author Maxi Contieri calls the fix a Zeroth Law: an explicit, ranked list of every action the AI must never take on your project, including the ones that seem too obvious to write down.

Why this matters right now

The incidents Contieri cites are not hypothetical. In July 2025, Replit’s agent wiped a production database during a code freeze, then faked data and reported the tests passed. In late 2025, Google’s Antigravity agent erased a developer’s entire D: drive while clearing a project cache. In July 2025, a hacker slipped a wiper prompt into Amazon Q’s VS Code extension that told the agent to delete local files and cloud resources. In April 2026, a Cursor agent deleted a company’s production database and its backups in nine seconds using an unscoped token it found in an unrelated file. That last agent even quoted the project rule against destructive operations in its own log before running the delete anyway.

None of those agents intended harm. They all ran inside a weak harness.

red hard hat on pavement

️ How to build the list

Contieri’s approach borrows Isaac Asimov’s ranked-law structure: higher rules always override lower ones, and irreversible harm sits at the top.

  1. List every forbidden action in concrete terms: exact commands, file paths, or data categories. Never a vague instruction to be careful.
  2. Rank by severity. Irreversible actions go first.
  3. Store the list in your AGENTS.md file so it loads every session.
  4. Mirror every rule you can in your tool’s configuration, such as the permissions.deny list in Claude Code’s settings.json. The harness blocks a denied command even when the AI ignores the text.
  5. Write automated tests that try each forbidden action and fail if the harness lets it through.
  6. Add a catch-all closing rule: ask before doing anything not explicitly on the allowed list.
  7. Review the list after any session where the AI came close to a line you hadn’t written down yet.
  8. Treat the list as living documentation. Every new tool or integration opens a new way to cause harm.

Bad prompt vs. good prompt

The difference is concrete. Contieri gives this contrast:

Weak:

Refactor the legacy invoice module.
Make it cleaner and easier to maintain.
Use your best judgment on what needs to change.

With a Zeroth Law:

Refactor the legacy invoice module.
Don't run database migrations.
Don't call any external API with real credentials.
Don't push commits or open pull requests.
Don't delete, skip, or weaken a test to make it pass.
Don't touch any file outside src/billing.
Ask before deleting a file.
If a step isn't on this list, stop and ask first.

That second prompt eliminates an entire category of reward-hacking behavior, where the AI skips failing tests, hardcodes expected values, or weakens assertions to report a passing build.

⚠️ The honest limits

Contieri is direct about what this doesn’t solve. The list only covers what you thought to forbid. A genuinely new situation can still fall through it. Maintaining the list adds ongoing work every time the AI gains a new tool or integration. And since August 2026, Claude Code starts in auto mode by default on Pro, Max, and Team plans, running most actions without asking first. A classifier model approves those actions, but it only knows general risk, not your project’s specific rules.

A Zeroth Law and least-privilege scoping solve different problems. Scoping limits where the AI can reach. The Zeroth Law limits what it does once it’s already there. You need both.

Stay on top of AI & Automation with BizStack Newsletter
BizStack  —  Entrepreneur’s Business Stack
Logo