
AI development is a co-evolving stack
Hardware, models, harnesses, tools, and traces—as one production system

Hardware, models, harnesses, tools, and traces—as one production system
Hardware, models, harnesses, tools, and traces—as one production system
The same checkpoint produces different outcomes
Different context packaging
Different action primitives
Different authority boundaries
Different release evidence
Capability becomes a product through action and verification
can — Capability — Model behavior
may — Authority — Tools and identity
did — Evidence — Outcome and gates
A failure becomes a safe improvement across the stack
Local adaptation runs in weeks; foundations evolve over years
Local failures get fast fixes; repeated classes move upstream
The cost of behavior depends on more than FLOPS
Architecture adapts to available hardware and partnerships
Adapt inward
DeepSeek-V3 on H800
FP8 and MoE co-adapt
Training stack follows constraints
Partner outward
Anthropic × Trainium
Shared model-hardware roadmap
No hyperscaler ownership required
API compatibility does not create behavioral compatibility
Names, errors, and action primitives change the trajectory
Long tasks collapse without state and verification
Persistent state across turns
Small verifiable iterations
Bounded execution environment
Recovery after failure
A mature capability becomes a tool; a new risk gets scaffolding
A tool is a behavioral interface, not merely a schema
Discovery and schema are only the start
MCP standardizes
Tool discovery
Input schema
Call semantics
System must ensure
Correct identity
Bounded response
Verifiable effect
Identity, policy, response, and evidence are designed together
Value comes from governed action, not server count
01 — Choose — Right tool and arguments
02 — Act — Right identity and scope
03 — Recover — Legible failure and retry
Telemetry, eval, and training require different rights
Verify the end state in the environment and reliability across runs
Trajectory
Actions and tools
Errors and recovery
Cost and policy events
Outcome
Environment end state
Hidden executable checks
Repeated-run reliability
Change one layer and rerun the gate
Capture trace and final state
Replay in a frozen environment
Change one controlled layer
Rerun quality and safety gates
Release with measured evidence
Cycle speed is an operating capability
Testing harness half-life and the effect of integration
Most mechanisms change; half do not necessarily disappear
Git shows code changes, not the complete rollout
Observed
≥7 of 10 mechanisms retuned
About 3–4 clearly replaced
Fast compatibility program
Not established
Two projects at ≥5 replacements
Production rollout state
Universal 180-day constant
Several fast loops emerge, with routing between them
Adapt where the market meets your environment
Five yes answers beat a desire to control everything
Six steps close the part of the loop you control
Capture
Choose real episodes
Separate trace data modes
Inventory action contracts
Control
Build portable replay
Run repeated release gates
Test one layer at once
Components get cheaper; learning capability remains
Evaluate the production system
Run fast and slow loops
Design tools for behavior
Own episodes, policy, outcomes
Change one layer with evidence
Failure → replay → change → gate → release
The longread contains the full list and caveats
Hardware · NVIDIA · Google · DeepSeek
Harnesses · OpenAI · Anthropic · Cursor
Tools · MCP · evals · telemetry
Code history · routing · policy
AI development as a co-evolving stack
The full longread, every source, and future reviews are available at polomodov.tech and in Knizhny kub
Alexander Polomodov, Technical Director & Fellow, T-Technologies
@book_cube