A ban on writing code changes engineering work
Following the discussions of SWE-agent's interface and work handoffs between sessions, the host moves to the whole development environment. The source is Ryan Lopopolo's engineering post, “Harness engineering: leveraging Codex in an agent-first world,” published by OpenAI in February 2026. The team deliberately stopped writing code by hand. When Codex struggled, people identified a missing capability and changed its surroundings; the agent also implemented those environment changes. Asking it to try harder could not replace a tool or a clear rule.
The project was a new internal product, started from an empty repository in late August 2025. The team reports that over five months it grew from three to seven engineers and produced about a million lines of code and 1,500 PRs, averaging 3.5 PRs per engineer per day. Hundreds of employees used the product, joined by external alpha testers. People set priorities, understood user needs, and defined acceptance criteria, while Codex created code, tests, configuration, and documentation. The authors estimated development took roughly a tenth of the time manual coding would have required.
Verification requires an application the agent can inspect
The stream of changes soon made human attention the bottleneck: agents produced more work than engineers could review. Delegating some verification required a way to observe the application. Each agent worked in a separate Git worktree and started its own application instance, keeping parallel tasks from sharing state. Chrome DevTools let it inspect the DOM, navigate, click buttons, and capture screenshots to compare behavior before and after a change.
Each instance also received its own observability stack, with logs, metrics, and traces accessible through LogQL, PromQL, and TraceQL. Requirements such as service startup within 800 milliseconds or trace spans below two seconds in critical scenarios became measurable. Codex could check them and continue fixing failures. Browser and telemetry access answered a concrete question: how does the implementer know its change works and respects the stated constraints?
Knowledge must be accessible to the implementer
Documents in Google Docs, decisions in Slack, and knowledge in employees' heads cannot help an agent that cannot reach them. The host compares this with onboarding an engineer: the difficulty appears before implementation, when the newcomer does not know where to find an explanation of the system. Explicit agreements and a clear place to store them benefit both. Moving knowledge into the repository expanded what Codex could read, check, and change.
One enormous AGENTS.md, however, worked poorly. It displaced the task, code, and relevant documents from context; equally important instructions lost their priority, while the file became stale and was difficult to check mechanically. The team separated architecture, design decisions, execution plans, and product specifications. AGENTS.md became a navigation map. The agent could retrieve relevant information as needed, and individual documents could be checked against understandable rules.
Architecture becomes enforceable feedback
The team chose strict boundaries and a predictable structure for autonomous implementation: fixed layers within business domains and rules for their dependencies. The host stresses that teams do not need to copy this particular architecture. They need to choose their own explicitly, explain it, and check new changes against it. Custom linters did more than report an error: they explained the violation and how to fix it. The agent received a next step instead of having to guess the engineer's intention.
Control focused on invariants: boundaries, correctness, and reproducibility. For example, incoming data shapes had to be checked at the application boundary. Alexander illustrates the opposite with an Android example he had encountered: an unchecked object reached the internals, and a missing field caused a failure later. The agent could choose its internal implementation more freely. Its code might differ from a person's taste; the more consequential questions were whether it worked and whether the next agent could understand it and continue.
Autonomy depends on the consequences of failure
In the described workflow, Codex examines the project, reproduces a bug and records a video, implements a fix, tests the application, and records the result. It then opens a PR, receives agent review, addresses comments, and fixes failures. A person can join or make a decision when escalation is needed, but human participation in every review is optional. This workflow followed an initial investment in the environment; throughput did not rise immediately.
The team also reduced blocking checks. PRs are short-lived, flaky tests can be rerun, and some problems can be corrected in the next iteration. For this internal product, waiting could cost more than a later fix. The host limits that conclusion: the same tradeoff may be unsuitable when an error has financial consequences or changes infrastructure. Autonomy and merge policies depend on the consequences of the particular work.
Drift becomes recurring cleanup
Codex reproduces existing patterns, including poor ones, so the system's quality can drift. Initially, engineers spent Fridays removing accumulated problems, but this did not scale. They then defined shared principles and mechanical rules, such as using common utility packages instead of multiplying custom helpers. Background Codex tasks searched for deviations and gradually repaired the code. Recurring review comments became input to regular technical-debt work, with checkable rules guiding the cleanup.
Transferability must be tested on your own tasks
A million lines and the PR count show the scale of the case without establishing the effect of each harness component. The post does not report PR sizes, defect and rollback rates, the distribution of human hours, or token costs. The tenfold speedup is the authors' estimate. Architectural coherence over years, the most useful role for human judgment, and the effects of stronger models remain open questions. Experience with a new product cannot automatically transfer to a twenty-year-old monolith.
The host then offers a test that the article does not describe: find a recurring agent failure, add the missing capability, and compare the result with the original environment. Probabilistic behavior requires repeated runs. Track accepted tasks, human time, recurring defects and failures, tokens, and elapsed time. An agent may free a person while spending five hours on something the person could finish in two. A negative result provides a reason to remove the change, rather than treating every addition to the harness as an improvement.
The next bottleneck is coordinating agents
The epilogue introduces the same team's continuation, OpenAI Symphony. Even after implementation became autonomous, people still launched sessions, switched between them, and recovered stuck attempts. The next step puts a task tracker between engineers and an orchestrator that coordinates agents in separate worktrees. Engineers prepare tasks. The host mentions the open repository released in March, an April 27 follow-up linked to the February case, a draft specification, and an Elixir reference implementation. This moves the organization of work up another level.
What to take away
- 01Treat an agent failure as a potentially missing environment capability: application observation, access to knowledge, or a checkable rule.
- 02A concise knowledge map, explicit architecture, and linters that explain fixes support autonomous continuation.
- 03Freedom inside an implementation requires boundary controls; making human review optional depends on the cost of a possible mistake.
- 04The OpenAI team's account demonstrates feasibility in its own product. Measure value for your team through accepted work, human time, quality, and the full cost of execution.
Sources
- Audio transcript
- Episode slides
- Episode recording