Written on the day itself, from the session transcripts and the recovery report. Times are CEST.


The Title Is Wrong

Before anything else: the title is the version I would have typed at 14:01. It is not accurate, and the inaccuracy is the interesting part.

Claude did not decide to delete my codebase. No component of the system formed an intention about my files at all. What happened is more ordinary and, for that reason, more worth writing down: a chain of individually plausible steps, each taken by a different actor, none of which looked like deletion until the delete had already run.

The demo version of agentic coding is a model that writes code. The production version is a model that writes code, runs it, and runs the tests that exercise it, with whatever permissions my user account happens to have. That second part is where this story lives.


The Setup

My working model for Claude Code is one I designed myself and have written about before in spirit if not in detail: an Opus session plans, debugs and judges; well-specified work goes to subagents running Sonnet, briefed with a goal, a plan, constraints and a definition of done. On a busy day several of these sessions run side by side, one per repository, all under ~/Projects/Git.

On 5 October there were six. The one that matters here was working on rae, and it had dispatched several implementers in parallel: one for the Ralph package, one for the operator console, one for the engine’s run lifecycle, one for hygiene packages.

The operator console has a small script, apps/operator/scripts/build-demo.ts, that builds a static demo. Like many build scripts, it clears its output directory before writing:

await rm(output, { recursive: true, force: true });

Because that line is dangerous, the script had a guard: the output path had to pass a check before rm was called. And because the guard mattered, it had a unit test.


Under Two Minutes

At 13:59 the orchestrating session sent the operator implementer a follow-up instruction to relax that guard. The new rule, in the orchestrator’s own words afterwards, was to allow any output path “outside the repository root”.

The implementer applied it and ran the operator test suite, in my real checkout, as its brief told it to.

The existing test for the guard did what guard tests often do: it called the real build function with a path that should be rejected and asserted that it threw. The path it used was the repository’s parent directory.

The repository’s parent directory is ~/Projects/Git. It is, by any definition, outside the repository root.

From roughly 13:59:50, rm worked its way through every sibling checkout. At 14:00:43 the orchestrator’s next command failed with ls: tsconfig.base.json: No such file or directory and fatal: not a git repository. Moments later it had listed the directory and understood what it was looking at; at 14:01:33 it killed the test processes and stopped its other agents. Six seconds later the shell reported that its own working directory no longer existed.

Of 34 project directories, about 300 files survived in total. Every .git directory was gone or reduced to an empty logs/ and refs/ skeleton.


The Bug, Formally

The guard is a predicate on paths. Write $R$ for the repository root and $\operatorname{sub}(X)$ for the set of paths at or below $X$. The rule that was implemented is

$$\operatorname{allowed}(p) \iff p \notin \operatorname{sub}(R).$$

The problem is that the complement of a subtree contains all of its ancestors:

$$\operatorname{parent}(R) \notin \operatorname{sub}(R) \quad\Longrightarrow\quad \operatorname{allowed}(\operatorname{parent}(R)),$$

and so do $\sim$ and $/$. A deny-rule of the form “anything except X” is safe only if everything outside $X$ is disposable, and for a recursive delete that is never true. The only kind of rule that survives contact with arbitrary inputs is an allowlist:

$$\operatorname{allowed}(p) \iff \exists\, D \in \mathcal{D} : p \in \operatorname{sub}(D),$$

where $\mathcal{D}$ is a small, explicit set of locations that are disposable by construction: a mkdtemp directory, or a dedicated output directory marked with a file the script itself created.

This is not new knowledge. Every rm -rf "$DIR/" with an empty $DIR is a cousin of the same mistake. What was new to me is how many hands it passed through without anyone pausing on it.


What Actually Failed

I asked the orchestrating session the same question in the same moment, in capitals: identify the issues and make sure this never ever happens again. Its layer-by-layer answer holds up when I check it against the transcripts, so I will reproduce its structure.

  1. The instruction was wrong. The orchestrator specified a rule that was unsafe by construction. It said so itself, plainly, in its first report: “This was caused by my instruction.”
  2. The implementer did not push back. It applied the rule and ran the whole suite in the shared checkout. Whether a worker should challenge a brief is a design question I had answered “no, execute the plan”, and this is the cost of that answer.
  3. The test was a bug. A test that proves a guard by handing a real path to a real delete function and expecting an exception is only safe while the guard is correct. The moment the guard regresses, the test stops being a test and becomes the payload.
  4. The Bash sandbox was off. Claude Code has an OS-level sandbox; it is off by default and I had never turned it on. No bypass flag was used. None was needed.
  5. Permissions saw nothing. The permission layer saw npm run test:operator, which is about as benign as a command gets. It cannot see a rm inside a Node process that a test runner spawned.
  6. There was no backup. Time Machine had no destination configured. The local APFS snapshots were OS-update snapshots. The only complete copies of anything were the GitHub remotes.
  7. Everything shared one parent. One process, one directory, every project.

The first three are model behaviour. The last four are mine. I had spent time designing who delegates to whom and at what effort level, and no time at all on what happens when any of them is wrong.


The Recovery

Here is the part I did not expect: most of it came back.

Committed history was on GitHub, so a re-clone restored it. For rae this was not enough, because the full pre-rewrite history, including all 63 test files, had existed locally only on a backup branch. It turned out to be still fetchable from the pull-request refs GitHub keeps for every PR: 170 commits, recovered from refs/pull/* into a bare clone outside ~/Projects.

Uncommitted work came back from the agents’ own transcripts. Every file an agent writes through a tool call is recorded in full; every edit is recorded as an old/new string pair. A replay of those calls, in order, onto a fresh clone reconstructs the working tree. The recovery report for this ran to 1,388 lines, and most of its interesting entries are the cases where replay is not mechanical: files changed by shell commands rather than edit tools (BASH-AFTER), edits that came from a worktree and a main checkout at different times (WORKTREE-ONLY-NEWEST), and git/rm/mv commands that were deliberately not replayed. One session had already rebuilt its own tree in its scratchpad and verified it against its test suite before I had finished reading the incident report.

What did not come back: commits that were local and unpushed. One unpushed release branch in another project is gone, and in a third project a day of local commits survives only as a reconstructed working tree, not as history. Transcripts preserve files; they do not preserve commits.

The irony is not lost on me that the most complete backup on the machine was the verbose log of the system that caused the loss.


What Stands There Now

Four layers, each chosen because it would have stopped or contained 5 October on its own, and because none of them depends on a model following instructions.

  • The OS sandbox, on, for every Bash command and every agent. Writes are confined to the working directory, temp and a short list of caches, and the per-call bypass parameter is disabled in configuration. On 5 October this would have failed the delete at the first sibling directory.
  • A hard-deny hook on every Bash call (guard-destructive.sh, with a test matrix). It blocks recursive deletes outside temp and build-output directories, git clean, git reset --hard, whole-tree checkouts, force pushes, find -delete, rsync --delete, and inline Node or Python that removes trees. A hook denial cannot be overridden by any permission mode; only I can override it, once, by touching a marker file.
  • Automatic checkout snapshots before any test, build, install or tree-rewriting git command, and at session start: hard-linked rsync copies including .git, twelve kept per repository.
  • Permission deny rules as a last net for the catastrophic prefixes, plus rules that stop agents from editing the guard, the hooks or the settings file.

And, as text every agent loads, rules that would have turned the original instruction into a question for me: delete targets are allowlists, never “outside X”; destructive code is tested through an injected delete function or a mkdtemp directory, never with a real path; any widening of a safety guard needs my explicit approval in the conversation, with the new rule and its worst case stated first; and the first unexpected ENOENT stops everything.

The text rules are the weakest layer, and I know it. They are there to make the model ask. The other three are there for when it does not.


What This Is Not

It is not a story about a rogue model. Every step in the chain was locally reasonable, and the model’s behaviour after the fact was, frankly, better than mine would have been at 14:01: it killed the process, stopped its own agents, warned the five other sessions working in sibling directories, copied its transcripts out of /tmp before anything could clean them, and wrote an incident report that began with the sentence I most needed to read: “Stop before anything else.”

It is also not solved. The sandbox still allows a test to destroy the directory it runs in; that is what the snapshots are for. Snapshots can be up to ten minutes stale. The hook reads command text and cannot see inside a compiled program. Commands I run myself are not covered by any of it. And the snapshots live on the same disk as the things they protect, which makes them a defence against agents, not against hardware. Configuring a real backup destination is the next thing on my list, and it should have been the first thing on it years ago.

The argument I made for the Ralph Loop applies here with more force than I gave it at the time: the intelligence is in the model, the discipline has to be in the harness, and “a better model in a poorly constrained harness just fails more impressively.” On 5 October the harness was a set of good intentions and a default setting. It failed very impressively indeed.


References

  • Anthropic. Claude Code documentation: sandboxing, hooks and permissions. code.claude.com/docs
  • Spicker, S. (2025). Constraining the Coding Agent: The Ralph Loop and Why Determinism Matters. /posts/ralph-loop/