Skip to content
back to the archive page
#AI4SDLC

Factory and HumanLayer: how much autonomy should agents have? (Series #AI4SDLC)

I watched an interesting talk by Tereza Tížková, who works on growth at Factory. She spoke at AI Engineer World’s Fair on June 30; the standalone recording came out on September 27. This time the discussion is about organizing the entire development process around agents. It continues a thread I explored in May when I covered Luke Alvoeiro’s Factory talk about their Droid agent and Missions mode. Work is split among agents: some plan, others write code, and others check the result. The completion criteria are written before implementation, and the reviewer must not be the code’s author.

Returning to Tereza’s more recent talk, her idea is that a software factory should collect feedback, select tasks, carry them out, verify the result, and start the next cycle. That requires three things: the ability to change models and tools, sustained autonomous agent work, and accumulated knowledge of the project. The human sets the direction and decides what is worth building.

The familiar scheme gains a few useful details here. 🔸 Money Factory matches the model to the task: it picks the cheapest model that the system estimates can do the job. It can move to a stronger one if needed. For long tasks, it matters how much the entire path to an accepted result costs, including checks and retries. 🔸 Context Descriptions of every connected tool can fill an agent’s memory before work even begins. Factory first shows a short catalog, loading details as needed. According to the company’s figures, this reduced input tokens by roughly 15% on average across the MCP sessions it studied. In the group with 100+ tools whose descriptions were loaded on demand, average savings were around 51%. The effect therefore depends heavily on how much is connected. 🔸 Team readiness Agents also need documentation, tests, a clear way to run the project, and written working rules. If only one person knows how to run the build and important agreements live in chat, the agent will have to figure it all out again. Factory calls this readiness check Agent Readiness.

This is a good point to recall Dex Horthy and his Why Software Factories Fail. Dex also builds tools for agentic development. His HumanLayer is an environment where teams discuss agent plans and review code changes.

His emphasis is different: passing tests do not tell you how easy the system will be to change six months from now. Dex therefore suggests discussing requirements, architecture, and code structure upfront, then moving in small, verifiable steps. He also acknowledges that he cannot yet convincingly prove that agents make code harder to maintain over time.

Factory shows how to organize autonomous agent work and verify its results. Dex examines in more detail what tests struggle to check: code structure and the team’s ability to keep changing it. Both have their own products and interests in this discussion. I would combine these approaches: automate execution, while discussing important structural decisions before the agent writes the code.

I would take away three practical steps: 1️⃣ For engineers Before starting an agent, write down what should work and how to check it. Afterwards, make sure you can explain the key decisions in the resulting code. 2️⃣ For platform teams Take one repository and check whether an agent can start the environment, find the project rules, and run tests by itself. Examine every manual hint: what is missing from the documentation or tools? 3️⃣ For managers For a small task, calculate the cost of the finished change: models, checks, rework, and people’s time. Compare it with the previous process, and assess the effect on quality separately.

#AI4SDLC #AI #Agents #Engineering #Management

Open video on YouTube