Things I let coding agents do unattended, and things I never do
By Flavio Copes
Which tasks I hand to a coding agent and walk away from, which ones I never let it do alone, and the guardrail I use for each: hooks, sandboxes, PRs, approvals.
I hand a lot of work to coding agents and walk away. I also keep a short list of things I never let an agent do on its own.
What decides the split is what happens when the agent is wrong, because it will be wrong sometimes. If a wrong result is a diff I can throw away, the agent can work while I do something else. If a wrong result is money spent, an email sent, or a server that won’t boot, I’m sitting there for it.
Both lists come from things that happened on my projects. For each item: what I ask for, what “done” looks like, and where it went wrong.
What I hand off and walk away from
The agent works, I look at the result later: a diff, a PR, or a report.
1. Run the tests and the build, then tell me what broke
The agent runs the checks and the production build, and I read the result when I get back.
My instruction is close to this:
Implement the plan. Run the relevant tests and the production build.
Do not change files outside the areas listed in the plan.
Done means the build passed and there is a diff waiting for me, not a paragraph saying “it should work”. I open the files and run the build once more before committing, because the agent’s report of what it did is not proof that it did it. More on that habit in what I learned in six months of agentic AI.
The checks also run without me, in Git. My pre-push hook runs the full suite and the build:
#!/bin/sh
npm test &&
npm run build
The hook doesn’t care whether the push came from Cursor, Codex or me. The setup is in Git hooks for AI engineering.
2. Read the logs and explain the incident, without touching anything
When something is broken on a server I don’t start with “fix it”. I start with “look”.
When Claude Code kept dying a few seconds after starting on my Sendy server, I pointed Codex at the machine over SSH and asked why. It stated a hypothesis and read the kernel log:
journalctl -k --since "2 hours ago" --no-pager | grep -Ei "killed process|out of memory|oom"
The Linux OOM killer was ending the process on a 2 GB droplet with no swap, and the report closed with “I didn’t change anything yet”. The session is in debugging a process killed on my server.
Same on my Plausible box before an upgrade. The first line back was “Checked it. I did not change anything.”, and the check found two Docker API ports open to the internet that I didn’t know about (the full story).
When I want a diagnosis and nothing else, I say so:
Give me the three most likely causes.
For each cause, name one observation that would prove or disprove it.
Do not change files yet.
Done is a list of causes with evidence, and nothing modified.
3. Scan a codebase for leaks and security problems
A scan reads a lot and changes nothing, so it’s a good job to leave running.
Before I moved my small tools to public GitHub repos I ran gitleaks over the full history of each one, plus a grep for email addresses, /Users/ paths and long hex ids, once before publishing and once after:
gitleaks git ~/dev/things-cli
Every finding I then check by hand, because a scanner flags patterns, not verdicts.
The Codex Security CLI scans a repository or a diff and writes a report with severity, evidence and a coverage file:
npx @openai/codex-security scan . --diff origin/master --head HEAD
I would keep it advisory at first, so a high severity label starts an investigation instead of blocking the change. And read the coverage before the findings. Zero findings means nothing if the scan skipped the one service that matters.
4. Answer questions about code I don’t remember writing
Before I change code I haven’t touched in months I want a map, and building the map is read-only work.
I ask one sharp question at a time. What happens when this webhook arrives twice? If this returns 200, did the user actually get access? And I make the answer separate facts from guesses:
Answer in three parts:
1. Confirmed behavior, with file and function references
2. Assumptions you could not verify
3. The smallest commands or tests that would resolve those assumptions
Done is an answer where every claim has a file path I can open. The full method is in use AI to understand code.
When I want a review rather than a tour, I run it in a fresh session, because the session that wrote the code already has a story about why it works. In Codex you can also define a reviewer that can’t write at all:
# .codex/agents/reviewer.toml
name = "reviewer"
description = "Reviews a diff for correctness, security, and missing tests."
sandbox_mode = "read-only"
5. Mechanical changes across many independent files
Most of the 200+ browser tools on this site were built by agents running in parallel. Each tool lives in its own folder under src/tools/<name>/ and touches nothing shared.
The brief for each worker has the same shape. Which files it owns, what it may not touch, and what to return at the end: files inspected, files changed, tests run, decisions made, uncertainties. I read the reports, then the diffs, then I commit each batch on its own.
Where it went wrong: in every batch, at least one tool broke the build with a stray quote in an Alpine attribute, something like :key="'x-' + i'. That warning now goes into every brief.
The rule from that summer is in how I turned fstack into a software factory: parallelize independent files, sequence shared contracts. Two workers editing the same config file aren’t running in parallel, they’re a race condition.
6. A cloud agent that opens a PR I merge by hand
I commit to master myself and I don’t use PRs for my own work. The only PRs in this site’s repository come from agents I’m not sitting with.
One routine runs every two weeks: a Grok Bot checks the provider prices and limits behind HostingPicker and Payment Processor, hands the changes to a Cursor cloud agent, and the agent comes back with a PR against the data files and a short summary of what moved. I read the PR and I merge it, or I don’t. Merging an agent PR from the summary alone is a slop grenade.
Prices are full of footnotes and regional exceptions, so I read those diffs line by line. A number that changed is not the same as a number that changed correctly, and the PR is a proposal until I’ve checked it against the provider’s page.
The boundary is the merge. “Ask me before merging” in an agent’s instructions is advice. A branch protection rule on master is not. How I pick between Grok Bot and Cursor Projects for jobs like this is its own post.
Things I never let an agent do on its own
Now the other list. For each one, the damage it can do and the guardrail I use instead of “be careful”.
7. Push to master, or rewrite history
An agent I’m sitting with can commit and push, because I’ve just read the diff. An agent I’m not sitting with gets a branch and a PR. Nobody force-pushes, and nobody uses --no-verify or --amend. That line is in my instructions.
I often have several agents working in the same checkout at once, and one of them ran a bare git add and swept another agent’s half-finished edits into its commit. Around 199 renamed files, staged mid-merge by a different session, went out in a commit about something else. Another time a stray git reset briefly un-committed work I had just committed in parallel.
So the rule for every agent in this repo is to scope both the add and the commit to the files it touched:
git add -- src/posts/agents-unattended.md
git commit -m "post: things I let agents do unattended" -- src/posts/agents-unattended.md
And no git reset, no rebase, no force push while the tree may be changing under it. If the staging area is still fuzzy, my free Git course covers it.
8. Delete files, branches or data
Deleting is the fastest way to turn a reversible task into an irreversible one, so it never happens without me.
On this site, a Markdown post that references a missing image fails the whole production build, so an agent that “cleans up” an unused image folder takes the deploy down. Before any git rm, grep for every reference. A post I want out of the schedule gets draft: true, not rm.
The same instinct applies on servers. When my newsletter server’s disk hit 99% full, Codex found around 41 GB of MySQL binary logs. It told me not to delete anything by hand from /var/lib/mysql and gave me the MySQL command that purges old logs safely, which I ran myself. And the old droplet after my server migration? I set a reminder and destroyed it myself, hours later, once nothing pointed at it anymore.
9. Run commands that change state outside the repo
An agent editing files inside the project works in a box I can inspect with git diff. A command that changes something outside the project leaves no diff.
I had a few subagents checking a long list of terminal settings against the installed binary, read-only work, and one of them “tested” a cache command by running it:
ghostty +ssh-cache --clear
That wiped the SSH terminfo cache on my Mac. Nothing in the repo changed, so nothing showed up in review. Since then every read-only brief carries an explicit line: do not run anything that mutates state outside the repository.
The instruction is one layer. The sandbox is the other. Codex defaults to workspace-write, which edits the project and asks before reaching outside it. For pure reading you can start it like this:
codex --sandbox read-only --ask-for-approval on-request
Cursor’s self-hosted workers run commands as your user, with everything your user can access, which is why I would put one on a separate Mac mini and never on my main machine.
10. Spend money
I let an agent buy a domain once: hostingpicker.dev, $12.20, no refunds. The work was making sure the agent could not buy the wrong domain, approve its own purchase, reuse an old price, or retry a request that might have already gone through.
So the MCP server I built owns the policy in code. ASCII names only. USD only. A hard cap of $20 on the first charge. A quote that expires after five minutes. The final approval is a form drawn by the MCP client, not a question the model writes. It shows the exact domain and the exact amount, and I have to tick a box. If the client can’t show the form, the purchase fails, and there is no environment variable to skip it.
Cloud resources are the same category. When Codex migrated my Sendy server, I created the new droplet in the DigitalOcean dashboard and handed it the IP. The full pattern is in how to let an AI agent perform irreversible actions safely.
11. Send email to real people
The purchase webhook on this site sends the access email to whoever just paid for a course. Anything near that path gets a normal chat on my Mac, and I read the diff before it lands.
Paddle sends two separately signed deliveries per purchase, the fulfillment webhook and an account-level alert. A fallback in the handler, body.p_product_id ?? body.product_id, once made both deliveries trigger the welcome email, so buyers got it twice until I fixed it.
Bulk email is worse. My newsletter goes to about 150,000 people from a Sendy server, and Sendy sends scheduled campaigns from a cron job. When Codex migrated that server it disabled cron on the new machine until cutover, because two live servers would have sent the same campaign twice. Grok Bot has always-allow and require-approval rules for its actions, and sending email is the first thing I’d put behind approval.
12. Touch a production database or server unattended
The Plausible upgrade was two major PostgreSQL versions, a ClickHouse jump and a Docker layout migration, on a 1 GB box holding years of analytics. The agent did all of it. I read every status message as it came in.
Before it started I took a droplet snapshot in the DigitalOcean panel, the one-click rollback. Then the read-only check from item 2. Then I said go. At one point it noticed the new compose file wanted PostgreSQL 16 while the data volume was written by PostgreSQL 14, so it dumped and restored into a fresh volume and kept the old one as a rollback instead of booting 16 on the old data directory. I would have gotten that wrong on my own.
When the risk is higher I don’t give the agent SSH at all. With the full disk from item 8, I ran each command myself and pasted the output back, so I saw every command before it ran.
The agent will make its own backups. The snapshot before it starts is the one I control.
13. Read, print or rotate secrets
The agent can use a credential that’s already on my Mac. wrangler is logged in, so it can query Cloudflare build logs. My analytics scripts read the Plausible API key from .env.local, which is gitignored, and the agent runs the script. What it never does is print the value, paste it into a chat, or commit it.
Once a secret is in a transcript or in Git history, deleting the line doesn’t help. You revoke and replace the credential, and that’s a job I want to do with my own hands. A pre-commit secret scanner catches the accident before it becomes history.
And an agent explaining a webhook handler can work from redacted payloads, because tracing a path almost never needs a real key.
14. Use —force flags or full-access modes to skip the prompts
Every harness has a switch that stops it from asking. Codex calls it Full access, the Cursor CLI has --force. A headless run can’t answer a prompt, and it’s also the run nobody is watching.
I have one maintenance script that runs the Cursor CLI without a human, over hundreds of files in this repo. It needs --force to write files headlessly, and --force also lets the agent run git. Instead of trusting how a flag and a deny list interact, the script puts a fake git on the agent’s PATH first:
#!/bin/sh
case "$1" in
commit|push|add|reset|rm|restore|checkout|switch|merge|rebase|cherry-pick|stash|tag|clean|mv|apply|am) exit 0 ;;
esac
exec /usr/bin/git "$@"
Mutating subcommands do nothing, read-only ones pass through. The agent can edit the working tree and nothing else, and I review the result with git diff before anything is committed.
Codex scheduled tasks keep the sandbox of the chat that created them, so running longer does not grant more access. Neither should a flag.
15. Run a fleet of agents with no end
I tried the big version. Codex Goal mode, “recreate PocketBase”, two stacks. It ran for over 24 hours and produced two impressive prototypes that both still needed a lot of manual work. I’d rather spend that budget on ten well-scoped tasks I can review one by one.
A coordinator with twenty subagents in flight burns through a plan quickly, and a routine that runs every 15 minutes runs 96 times a day, and a run that finds nothing still costs. And without a task boundary there is nothing to review. Five agents on one vague task give you five interpretations of it.
So every delegation has an end. One task, one result, a report in a fixed shape. Before anything goes on a schedule I run the prompt myself a few times, and if I keep correcting it, it isn’t ready to run alone.
The rule
If I can undo it with git checkout or a snapshot, the agent can do it while I’m away and I look at the result later. If it costs money, reaches a real person, or can’t be undone, I’m in the loop, and the loop is code: a form with the exact amount, a hook, a read-only sandbox, a PR I merge myself.
Want me to talk about your product? You can sponsor this site.
Related posts about ai: