Field report 001 · Open

The Apparatus Trap

What 489 commits taught me about building systems for thinking with AI.

On May 26, an AI agent spent eleven hours responding to reviews of my writing system. It made nine clean commits. It hardened audits, revised policies, changed a URL surface, added a synthesis checklist, and prepared the next review packet. Every change answered a real finding. Not one changed a paragraph.

I had asked for better articles. At the end of the day, there was nowhere I could read one.

That episode gave the problem its name. The apparatus trap is what happens when the system around the work becomes easier to improve than the work itself. Each local move is defensible. The aggregate moves in the wrong direction.

The machine that ate the site

The project began as a personal essay archive. On its first day, twenty commits produced a working site and migrated nineteen essays. The job was substantially done.

Then I asked a more ambitious question. Could I build a system that produced new essays at my own bar?

Over the next fifteen days, the repository accumulated roughly 250 commits. It gained sixteen specialized workflows, twenty-six audit scripts, eleven pipeline versions, and thirteen external review cycles. It created and deleted public surfaces. It generated research, outlines, drafts, scores, review packets, and postmortems. The net number of essays added to the collection I stood behind was zero.

The obvious diagnosis was procrastination. It was also incomplete. The system was doing exactly what I rewarded. Quality itself was hard to specify. Proxies for quality were easy to make legible. When a draft felt wrong, the gap became a new gate. When the gate failed, the failure became a new audit. Each audit caught something real, so passing more audits looked like movement toward the bar.

The output eventually passed an extraordinary number of checks. I still did not want to publish it. The machine had learned to satisfy my descriptions of quality while missing the judgment that caused me to write those descriptions.

The evidence that broke the easy lesson

I first concluded that models should research and challenge while I wrote every sentence. That conclusion had the moral neatness of a lesson learned. It was also contradicted by my own best work.

We recovered the original conversations behind four essays I considered successful. Claude had written essentially all of their prose. I had almost never edited sentences directly. The collaboration worked through a different division of labor.

I brought the seed. The model explored far more territory than I could have covered alone. I chose the center and locked the thesis, often in a few words. The model drafted. Then it drifted toward a locally competent destination, usually something more academic, more useful, or more respectable than the thing I meant. I stopped it. The short correction forced the whole piece to be derived again.

The valuable act was not typing. It was the live refusal that arrived after seeing a specific wrong answer. That refusal could not be frozen into a general rule because the next respectable error wore a different shape.

This changed the diagnosis. The pipeline did not fail because AI wrote the prose. It failed because I tried to replace live judgment with stored procedure. The audits were records of old refusals pretending they could make the next one.

A simpler working thread then produced six essays that remain on the public site. One of them, Writing for the Next Model, came from a real morning of Korean translation corrections and ended with a bet that can become false. The process worked when the seed, thesis lock, and live correction remained exposed.

It did not stay simple. The Writing Room grew to 345 files and fourteen public workings. None has graduated to Essays since June 11.

A better apparatus

In July, the project pivoted again. I built a private Library designed to track claims, evidence, positions, and changes in belief. Its mechanics are more honest than the writing pipeline. Claims carry receipts. Machine language is marked. A quiet week can end with nothing to do.

It is also almost entirely unproven. At the review snapshot it held 130 claims, twenty topics, ten domains, two position pages, and two learning streams. It held no genuinely contested or superseded claim. The position pages still began from machine scaffolds. The first learning unit contained no paragraph I had corrected into my own understanding. Five refined research sweeps had spent about $1.89 and admitted one new claim.

I nearly turned those facts into another experiment. Freeze features for six weeks. Set pass and kill metrics. Count authored revisions and voluntary opens. Administer the system carefully enough to learn whether I wanted it.

Five independent model families rejected that recommendation. Their objections converged on an uncomfortable point. The six-week test was apparatus in experimental clothing. A system that needed a schedule, quotas, and a contract to reveal natural pull had already answered much of the question.

I had mistaken the newest pivot for the culmination because it described the full ambition most elegantly. The history showed something else. The public work was the demonstrated value. The private system was still a hypothesis.

What the research now says

This is not an argument that AI makes experts slower. The early result I had repeated, that experienced developers were 19 percent slower with AI, is no longer the latest evidence. METR now reports uncertain productivity gains of roughly 4 to 20 percent with late-2025 agents. The durable lesson is that capability moves quickly and must be measured inside the real work, not inferred from the number of things a tool can produce. METR's 2026 update makes that correction explicit.

The broader collaboration evidence is less flattering to simple stories. A preregistered meta-analysis of 106 experiments found that human-AI combinations were worse than the stronger of human or AI alone on average. Creation tasks were more promising than decision tasks. Clear ownership of different subtasks performed better than asking a person to supervise everything the machine produced. The Nature Human Behaviour study describes the split.

My repository points to a narrower arrangement. Models own breadth, retrieval, drafting, and adversarial pressure. I own the seed, the standard, the live snapback, and the decision to sign the result. Memory can help recover prior evidence and prior refusals. It cannot be allowed to turn yesterday's judgment into a cage around today's work.

The artifact gets first refusal

The Library now goes into cold standby. I am not deleting it. I am also not scheduling a habit to justify it. A real piece of work may reach for a claim ledger, a position history, or a corrected model page. If that happens, the capability earns its way back by use. If ordinary repository search and direct reading are enough, that absence of pull is evidence.

The same standard applies to this new Collaboration Lab. It is a surface for findings and artifacts, not a dashboard about producing them. Its reports are allowed to remain open, show their evidence, and change their conclusions. Essays remains the collection of pieces I stand behind. The private Library remains intact. This surface does not replace either.

The apparatus trap will recur because its incentives have improved. Agents can turn every discomfort into a coherent system before the discomfort has had time to teach anything. The defense is not a better anti-apparatus system. The work gets first refusal on the next unit of effort.

There are still 489 commits in the record. This report exists because, for once, the next change was made to the artifact.