Skip to content
Research Insights Made Simple logo
Podcast · August 7, 2026

AI development is a co-evolving stack

Hardware, models, harnesses, tools, and traces—as one production system

/ Research Insights Made Simple #27 · Co-evolving stack

Slide contents

  1. 1. AI development is a co-evolving stack

    Hardware, models, harnesses, tools, and traces—as one production system

  2. 2. A strong model does not make a system

    The same checkpoint produces different outcomes

    Different context packaging

    Different action primitives

    Different authority boundaries

    Different release evidence

  3. 3. An agent lives in a world of consequences

    Capability becomes a product through action and verification

    can — Capability — Model behavior

    may — Authority — Tools and identity

    did — Evidence — Outcome and gates

  4. 4. The advantage lives between layers

    A failure becomes a safe improvement across the stack

  5. 5. 01. Two clocks

    Local adaptation runs in weeks; foundations evolve over years

  6. 6. Not every failure needs a new model

    Local failures get fast fixes; repeated classes move upstream

  7. 7. Hardware changes the feasible architecture

    The cost of behavior depends on more than FLOPS

  8. 8. Constraints can trigger co-design too

    Architecture adapts to available hardware and partnerships

    Adapt inward

    DeepSeek-V3 on H800

    FP8 and MoE co-adapt

    Training stack follows constraints

    Partner outward

    Anthropic × Trainium

    Shared model-hardware roadmap

    No hyperscaler ownership required

  9. 9. 02. Action environment

    API compatibility does not create behavioral compatibility

  10. 10. A model learns a particular world

    Names, errors, and action primitives change the trajectory

  11. 11. Autonomy needs a protocol

    Long tasks collapse without state and verification

    Persistent state across turns

    Small verifiable iterations

    Bounded execution environment

    Recovery after failure

  12. 12. The harness covers the next weakness

    A mature capability becomes a tool; a new risk gets scaffolding

  13. 13. 03. Action contracts

    A tool is a behavioral interface, not merely a schema

  14. 14. A call standard does not guarantee action

    Discovery and schema are only the start

    MCP standardizes

    Tool discovery

    Input schema

    Call semantics

    System must ensure

    Correct identity

    Bounded response

    Verifiable effect

  15. 15. A good tool bounds consequences

    Identity, policy, response, and evidence are designed together

  16. 16. The gateway learns from selection failures

    Value comes from governed action, not server count

    01 — Choose — Right tool and arguments

    02 — Act — Right identity and scope

    03 — Recover — Legible failure and retry

  17. 17. Observing does not mean training

    Telemetry, eval, and training require different rights

  18. 18. A convincing path does not prove the result

    Verify the end state in the environment and reliability across runs

    Trajectory

    Actions and tools

    Errors and recovery

    Cost and policy events

    Outcome

    Environment end state

    Hidden executable checks

    Repeated-run reliability

  19. 19. Every failure must be replayable

    Change one layer and rerun the gate

    Capture trace and final state

    Replay in a frozen environment

    Change one controlled layer

    Rerun quality and safety gates

    Release with measured evidence

    Cycle speed is an operating capability

  20. 20. 04. Evidence and power

    Testing harness half-life and the effect of integration

  21. 21. A harness is retuned within six months

    Most mechanisms change; half do not necessarily disappear

  22. 22. Half-life is a metaphor, not a metric

    Git shows code changes, not the complete rollout

    Observed

    ≥7 of 10 mechanisms retuned

    About 3–4 clearly replaced

    Fast compatibility program

    Not established

    Two projects at ≥5 replacements

    Production rollout state

    Universal 180-day constant

  23. 23. One universal vertical will not win

    Several fast loops emerge, with routing between them

  24. 24. Rent what changes; own what proves

    Adapt where the market meets your environment

  25. 25. An owned harness must be earned

    Five yes answers beat a desire to control everything

  26. 26. Build release authority first

    Six steps close the part of the loop you control

    Capture

    Choose real episodes

    Separate trace data modes

    Inventory action contracts

    Control

    Build portable replay

    Run repeated release gates

    Test one layer at once

  27. 27. Your moat is evidence-to-change speed

    Components get cheaper; learning capability remains

    Evaluate the production system

    Run fast and slow loops

    Design tools for behavior

    Own episodes, policy, outcomes

    Change one layer with evidence

    Failure → replay → change → gate → release

  28. 28. Evidence lives across different layers

    The longread contains the full list and caveats

    Hardware · NVIDIA · Google · DeepSeek

    Harnesses · OpenAI · Anthropic · Cursor

    Tools · MCP · evals · telemetry

    Code history · routing · policy

  29. 29. Thank you!

    AI development as a co-evolving stack

    The full longread, every source, and future reviews are available at polomodov.tech and in Knizhny kub

    Alexander Polomodov, Technical Director & Fellow, T-Technologies

    @book_cube