Harness engineering: agent environments
Episode participants
A solo episode without invited guests.
What we discussed on the recording
In episode 35, Alexander Polomodov reviews Ryan Lopopolo’s article “Harness engineering: leveraging Codex in an agent-first world.” An OpenAI team starts with an empty repository and chooses a constraint: Codex writes all the code. People set priorities and acceptance criteria, then improve the environment when an agent gets stuck. The case concerns an internal product built from scratch; transferring it to an existing system requires separate validation.
Engineering attention becomes the bottleneck for reviewing changes. Separate Git worktrees and application instances, Chrome DevTools, logs, metrics, and traces let agents check their own results. Knowledge held in conversations and documents moves into the repository. AGENTS.md becomes a map to architecture documents, plans, and specifications, while linters explain violations and how to fix them.
Strict boundaries and verifiable invariants leave implementation choices open within modules. An agent can reproduce a bug, record behavior before and after a fix, open a PR, and go through review, escalating decisions to a person when necessary. The product’s minimal blocking checks reflect the relative costs of waiting and correcting errors; the host stresses that more costly failures may require different rules. Background cleanup guided by shared principles helps prevent poor patterns from accumulating.
A million lines, around 1,500 PRs, and a reported tenfold speedup do not establish quality or cost. The article provides no control group, breakdown of human time, token costs, or long-term outcomes. The host proposes testing one mechanism against recurring failures, comparing repeated runs with a baseline, and tracking accepted tasks and both human and agent time. The epilogue examines Symphony: coordination moves from individual sessions to tasks, with a public draft specification and an Elixir implementation available for study.