
AI development is a co-evolving stack
Hardware, models, harnesses, tools, and traces—as one production system
Slide contents
1. AI development is a co-evolving stack
Hardware, models, harnesses, tools, and traces—as one production system
2. A strong model is not a system
The same checkpoint produces different outcomes
Different context packaging
Different action primitives
Different authority boundaries
Different release evidence
3. An agent lives with consequences
Capability becomes a product through action and verification
can — Capability — Model behavior
may — Authority — Tools and identity
did — Outcome — Result and gates
4. The advantage lives between layers
A failure becomes a safe improvement across the stack
5. 01. Two clocks
Local adaptation runs in weeks; foundations evolve over years
6. Not every failure needs a new model
Local failures get fast fixes; repeated classes move upstream
7. Hardware changes the feasible architecture
The cost of behavior depends on more than FLOPS
8. Constraints can trigger co-design too
Architecture adapts to available hardware and partnerships
Adapt inward
DeepSeek-V3 on H800
FP8 and MoE co-adapt
Training stack follows constraints
Partner outward
Anthropic × Trainium
Shared model-hardware roadmap
No hyperscaler ownership required
9. 02. Action environment
API compatibility does not create behavioral compatibility
10. A model learns a particular world
Names, errors, and action primitives change the trajectory
11. Autonomy needs a protocol
Long tasks collapse without state and verification
State across context windows
Small verifiable iterations
Bounded execution environment
Recovery after failure
12. The harness covers the next weakness
A mature capability becomes a tool; a new risk gets scaffolding
13. 03. Action contracts
A tool is a behavioral interface, not merely a schema
14. A call standard does not guarantee action
Discovery and schema are only the start
MCP standardizes
Tool discovery
Input schema
Call semantics
System must ensure
Correct identity
Bounded response
Verifiable effect
15. A good tool bounds consequences
Identity, policy, response, and evidence are designed together
16. The gateway governs choice and recovery
Value comes from governed action, not server count
01 — Choose — Right tool and arguments
02 — Act — Right identity and scope
03 — Recover — Legible failure and retry
17. Observing does not mean training
Telemetry, eval, and training require different rights
18. A convincing path cannot prove outcomes
Verify the end state in the environment and reliability across runs
Trajectory
Actions and tools
Errors and recovery
Cost and policy events
Outcome
Environment end state
Hidden executable checks
Repeated-run reliability
19. Every failure must be replayable
Change one layer and rerun the gate
Capture trace and final state
Replay in a frozen environment
Change one controlled layer
Rerun release gates
Release with measured evidence
Cycle speed is an operating capability
20. 04. Evidence and power
Testing harness half-life and the effect of integration
21. Same score, different mechanisms
Codex 3/10 · Gemini CLI 3/10 · OpenCode 3/10 under a strict criterion
22. Half-life is a metaphor, not a metric
Git shows changes, not the full rollout
Observed
Mechanisms are repeatedly retuned
Some implementations get replaced
Compatibility work stays active
Not established
A consistent replacement threshold
Production rollout state
Universal 180-day constant
23. Several integration loops are already viable
Vertical, partnership, open-weight, routing, enterprise
24. Rent what changes; own what verifies
Adapt where the market meets your environment
25. Five questions test the business case
The pattern matters more than control appetite
26. Build release authority first
Six steps close the part of the loop you control
Capture
Choose real episodes
Separate trace data modes
Inventory action contracts
Control
Build portable replay
Run repeated release gates
Test one layer at once
27. Your moat is evidence-to-change speed
Components get cheaper; learning capability remains
Evaluate the production system
Run fast and slow loops
Design tools for behavior
Own episodes, policy, outcomes
Change one layer with evidence
Failure → replay → change → gate → release
28. Evidence lives across different layers
The longread contains the full list and caveats
Hardware · NVIDIA · Google · DeepSeek
Harnesses · OpenAI · Anthropic · Cursor
Tools · MCP · evals · telemetry
Code history · routing · policy
29. Thank you!
AI development as a co-evolving stack
The full longread, every source, and future reviews are available at polomodov.tech and in Knizhny kub
Alexander Polomodov, Technical Director & Fellow, T-Technologies
@book_cube