Skip to content
back to the archive page
#AI4SDLC

Harness engineering: how an experimental team at OpenAI reorganized development around agents for an internal product (Series #AI4SDLC)

While preparing Research Insights #35, I studied OpenAI’s article on harness engineering. It came out on February 11, 2026, and for that moment it contains many interesting and groundbreaking ideas about an engineer’s work. Especially when read through this question: what knowledge about a system needs to be preserved as code itself becomes cheaper?

Ryan Lopopolo’s team built an internal product while deliberately forbidding themselves from writing code by hand. According to their account, in five months they produced around a million lines, including documentation and infrastructure, and roughly 1,500 merged PRs. The engineers worked on the environment in which Codex could do the work.

I would highlight four things from the article:

1️⃣ Knowledge must be accessible to the agent. A short AGENTS.md acts as a map of the documentation. Requirements and the reasons behind decisions live in the repository. 2️⃣ Architecture needs verification. Allowed dependencies and layer boundaries are enforced by linters and structural tests. 3️⃣ The agent needs eyes. Access to the interface, logs, and metrics lets it reproduce errors and verify fixes. 4️⃣ Reducing complexity takes regular work. Background tasks look for deviations and propose small refactorings. Errors become reasons to improve rules and tools.

I like the shift of engineering judgment into the environment: a stated requirement can be applied and checked repeatedly. It strongly echoes the subject of my DotNext keynote: knowledge about a system should outlive its implementation.

The story had an interesting continuation. In an April interview, Lopopolo said he had worked in Frontier Product Exploration, a team building enterprise agent products. By his estimate, the first month and a half of this kind of development was about ten times slower than manual work. They had to invest in tools and context before the experiment began to pay off.

The next bottleneck was switching between agent sessions, a problem we discussed in Research Insights #34. That led to Symphony: it takes tasks from Linear, launches agents in separate workspaces, and carries the work through to acceptance. The human defines the work and evaluates the result. Interestingly, the Symphony repository contains a SPEC.md specification and an experimental Elixir implementation. You can ask an agent to build your own version from that specification. It is a concrete step toward reconstructing an implementation from preserved knowledge.

By September, Lopopolo was working at Google Cloud. In a new conversation, he advised investing first in tools and context: move corrections from one-off hints into documentation, linters, and tests. The original experiment did begin with an empty repository, though. In April, the author clarified that a human smoke test still preceded the app’s release. Bringing this process into a mature system requires an understanding of its risks.

Join Research Insights #35 this Friday. We’ll discuss the article in more detail: which parts are reproducible engineering practice, and what you need to be able to verify before giving agents more work.

P.S. I couldn’t attach a PDF because simply pressing Ctrl + P on OpenAI’s website saves a PDF with some of the text missing. I had to read the article the old-fashioned way, in the browser :)

#AI4SDLC #AI #Agents #Engineering #Architecture #PlatformEngineering

Open video on YouTube