[2/2] The New SDLC: Harnesses, the Factory Model, and the Economics of Agentic Engineering (Category AI4SDLC)
Continuing my discussion of Google’s May whitepaper: in part one, we covered the shift from syntax to intent, the spectrum from vibe coding to agentic engineering, and context engineering. Now let us look at what surrounds a model and turns it into a working agent.
The document’s favorite formula is Agent = Model + Harness. When starting with agents, it is tempting to equate the system with its model: a new model makes the agent smarter; an old one makes it dumber. The authors argue that this view leads to poor investment decisions. The model is only one input. Everything else—prompts, tools, context policies, hooks, sandboxes, sub-agents, and observability—is the harness that enables it to finish the job. They estimate that the model accounts for roughly 10%, and the harness roughly 90%, of the experience of using Claude Code, Cursor, Codex, or Gemini CLI.
They offer concrete evidence. On Terminal Bench 2.0, one team moved a coding agent from outside the top 30 into the top 5 by changing only its harness, keeping the model fixed. A separate LangChain study improved the score by 13.7 points by tweaking only the system prompt, tools, and middleware around a fixed model. Their conclusion deserves a place on the wall: an honest investigation often reveals that agent failures are configuration failures rather than model failures.
That leads to the factory model. The developer’s main product is no longer code, but the system that produces code: specifications and context, executing agents, tests and quality gates, feedback loops that return errors to agents, and guardrails. A factory manager does not machine every part by hand; they design the production line and quality controls. The modern developer gives agents success criteria rather than step-by-step instructions and lets them iterate.
The developer’s role splits into two modes: 1️⃣ Conductor: working in real time in the IDE, with control at the keystroke level. Useful for exploration, prototypes, and learning a new API. 2️⃣ Orchestrator: working asynchronously at the level of goals, delegating to several agents and reviewing the result rather than every line. Useful for features, migrations, and test generation. Most people switch between these modes during the day.
There is also the “80% problem.” An agent quickly generates around 80% of a feature’s code, but the remaining 20%—edge cases, error handling, integration points, and subtle correctness issues—requires deep context that models often lack. Errors have shifted from syntax to concepts: wrong assumptions about business logic and overlooked edge cases. These are harder to spot precisely because the code looks right and passes basic tests.
The authors’ economic perspective is interesting. For leaders, total cost of ownership matters more than velocity. Vibe coding looks cheap because of low CapEx, but hides high OpEx: token consumption, repeated prompting, the maintenance burden of spaghetti code, and security patches. The authors estimate that at the crossover point it costs 3–10× more per feature. Agentic engineering reverses this: high CapEx for specifications, tests, and structured context, followed by low marginal OpEx. Context engineering and intelligent model routing—large models for difficult work, cheaper ones for deterministic tasks—become direct financial tools.
The practical recommendations cover three levels: 1️⃣ Developers: create AGENTS.md, install a set of skills, turn a recurring workflow into a first agent, write tests and evals before generating code, and review every line going into production. 2️⃣ Leaders: establish context engineering as a practice, treating AGENTS.md, prompts, evals, and skills as versioned, reviewed code. Judge by evals rather than demos. 3️⃣ Organizations: invest in production foundations before scaling, adopt open standards such as MCP and A2A, and plan for hybrid human–agent teams.
The conclusion, “Intent as the New Interface,” sets out three principles: 1️⃣ Structure scales, vibes don't. 2️⃣ AI amplifies your engineering culture, including its strengths and weaknesses. 3️⃣ The human role evolves rather than disappears. It ends with:
Generation is solved. Verification, judgment, and direction are the new craft
For me, this is exactly the logic I discussed in my article on agent-first IDPs: Agent = Model + Harness, evals as quality control, and production foundations before scale. They are the same ideas at the platform level. Both texts reach the same conclusion: discipline—specifications, tests, evals, and harnesses—is a condition for speed.
#AI #AI4SDLC #VibeCoding #Engineering #Architecture #Management