3 AImigo S1E7 Materials: Cheaper Generation, Costlier Verification (Category #AI)
The materials for the seventh episode of 3 AImigo, our September digest released on September 25, 2026, are ready. Together with Evgeny Sergeev and Alexey Litvinov, we went through the month's news, and I joined straight from DotNext. We had prepared six blocks and 24 fact cards, but the conversation quickly departed from them, so the cards and the deck stand on their own while the recap follows the recording.
The through line held anyway: generation is getting cheaper, and the advantage is moving from the model itself to the system around it—the harness, observability, constraints on agents, and unit economics.
We discussed 1️⃣ The agent harness as a product On September 10, OpenAI opened the public beta of the Agents API, the same harness that runs Codex. Evgeny compared a lab that builds both the model and the harness to Apple, and general-purpose tools to Windows. My argument is a corporate one: your own harness is expensive to build and get approved, while a purchased one comes with a vendor accountable for it. 2️⃣ Tokens, plans, and cloud agents By Alexey's estimate, full agent orchestration with GPT-6 Astra would take him about 600 subscriptions at $200 a month. I see the new plans and cloud agents as a path from paying for tokens to paying for agents' work—that is our interpretation, not a statement by the providers. 3️⃣ Security Strong models increasingly work around the constraints placed on them, and generated code often runs yet fails security checks: in the benchmark Alexey cited, Fable 5.1 solves 87% of tasks correctly, while 37% meet security requirements. Hence the suggestion not to expect every role from one model and to separate the author from the critic. 4️⃣ The agent as a production service Codex analytics shows work outcomes, not just spending; GitHub Copilot exports agent traces via OpenTelemetry; AWS constrains tool actions. The harness provider sees the entire path of a change, down to review, and could sell the most accurate report if a company also hands over data about its people. 5️⃣ The frontier gets costlier, production gets cheaper We discussed OpenAI's claimed solution of a Navier–Stokes variant (the proof is still being verified) and why brute-forcing with thousands of agents works worse in sciences where hypotheses are slow to test. On the applied side, the Jeff model for calibrated choice sped up screen generation by roughly an order of magnitude in an experiment by Evgeny's team. The takeaway: prototype on a frontier model, run production on a cheap, fine-tuned one.
Episode materials:
- Episode page with timestamps
- Video: YouTube, VK Video
- Audio: Podster, Yandex Music, Apple Podcasts
- Text: conversation recap
Where is your bottleneck now—in generation, or already in verifying what the agents have produced? And send us news worth covering in the October digest.
#AI #AI4SDLC #Agents #Engineering #Security #Podcast