Warp: you corrected the agent. What did the system learn? (Series #AI4SDLC)
An interesting story from Warp’s creator, continuing the Factory discussion, where we explored organizing the execution and verification of agent tasks, and the Conductor discussion, where we looked at how a person manages several executors. Both leave a question: if you had to correct an agent today, what will change in its work tomorrow?
Zach Lloyd, founder and CEO of Warp, proposes using those corrections to improve the whole process. His talk “Software Engineering Is Becoming Factory Engineering” took place at AI Engineer World’s Fair on June 30, 2026; the standalone recording appeared on September 27.
In Lloyd’s view, engineers will increasingly work on the system that produces changes to a product. Warp sells infrastructure for this process, so the company’s interest is clear. The process itself is familiar: examine the task, prepare a specification, write code, review, release, and observe the result. In Lloyd’s account, people continue to review specifications, code, and product behavior.
I am more interested in the second cycle: working on the process itself. An agent performs a review, and an experienced engineer corrects one of its comments. An observer agent analyzes that feedback and proposes an update to a skill: an instruction for future reviews. A correction can then outlive the current PR and help other team members. This follows nicely from the Vercel d0 story: there, repeated requests became skills, and I wondered who would check and update that accumulated knowledge. Warp proposes using human feedback as one input.
The idea became more concrete after the talk. In Warp’s July guide, an agent already opens a PR to change a review skill. In the documentation updated on September 24, proposed fixes are tied to failed runs and require human review. Here, “self-improvement” means changing instructions and processes; it does not imply training model weights.
I would question the way success is assessed. In his message to the team, Lloyd suggests treating manual, interactive work with an agent as a failure to learn from. The goal is to increase the share of automated changes.
Repeated explanations of an already known rule are something we want to eliminate. But jointly exploring a solution to a new problem can be useful work, even if it lowers the share of automated changes. This again brings to mind Dex Horthy from HumanLayer, with his attention to design and understanding the system. I would assess which interventions were eliminated and what happened to the quality of the result. The percentage of autonomous tasks alone will not tell us. Human feedback can be wrong too, of course. Turn it straight into a shared instruction and tomorrow the agent will diligently repeat our mistake.
I would take three practical steps from this:
1️⃣ Find one recurring correction Look at recent reviews, formulate a rule, and keep an example of the mistake. Start with a small skill whose value can be tested. 2️⃣ Test the new instruction on other tasks Compare the old and new versions, including cases that already worked. Accept the change after verification and retain the ability to roll it back. 3️⃣ Calculate the full cost of the result Include time spent defining the task, reviewing, and reworking it, model costs, and the maintenance of the automation. Then check whether repeated corrections have decreased at comparable quality.
#AI4SDLC #AI #Agents #PlatformEngineering #Evals