Skip to content
Research Insights Made Simple logo
Podcast · 29 July 2026

Why the Architect AI Copilot Still Hasn’t Arrived

A review of 51 studies with Sergey Baranov

/ Research Insights Made Simple #24 · AI × Architecture

Slide contents

  1. 1. Why the Architect AI Copilot Still Hasn’t Arrived

    A review of 51 studies with Sergey Baranov

  2. 2. AI Sees a Snapshot. Architecture Lives as History

  3. 3. A Diagram Captures Form. A Decision Holds a Trade-off

    Artifact

    Diagram

    ADR

    Pattern list

    Architecture

    Why now

    Cost of change

    Consequences later

  4. 4. The Main Paper Is a Field Map, Not a Copilot Experiment

  5. 5. The 51-Study Corpus Combines Search, Screening, and Extraction

  6. 6. Application Breadth Is Wide; Integration Depth Is Not

  7. 7. The Wider the Decision, the Harder It Is to Verify

  8. 8. Three Experiments Show Value — and Its Limits

  9. 9. An LLM Generates a Decision Candidate, Not an Approved ADR

  10. 10. Runtime Studies at Least Expose a Measurable Outcome

  11. 11. A Result Holds Only Within Its Unit of Analysis

    What was measured

    Responsibility recall · 4 projects

    Pattern accuracy · 3 classes

    ADR-candidate completeness · n=95

    Deployment time · AKS testbed

    What does not follow

    All requirements are covered

    Conflicting qualities are balanced

    The best system boundary was chosen

    The decision survived multiple releases

  12. 12. The Authors Classified 21 Tools on the ISL Scale

  13. 13. Tools Create Islands of Intelligence

  14. 14. Artifacts Exist. The Architecture Chain Does Not

  15. 15. Fifteen Architecture Problems Lead to Six Requirements for AI

  16. 16. Six Gaps Remain Before a Real Copilot

  17. 17. Adaptation Requires Memory of the Long Horizon

    AICH1 · update the recommendation

    Fact: a requirement version changes

    Trigger: code/runtime drift

    Operation: find dependent decisions

    Output: revised option + evidence

    AICH6 · retain consequences

    Observe: debt and smells

    Align: versions and incidents

    Assess: erosion across releases

    Return: signal into the next decision

  18. 18. Traceability Without Context Remains Formal

    AICH2 · continuous traceability

    Requirement ↔ ADR · type and version

    ADR ↔ component · rationale

    Component ↔ code · ownership

    Decision ↔ runtime · evidence

    AICH3 · local context

    Domain rules and data

    Team topology and ownership

    Regulation and security boundaries

    Migration cost and legacy constraints

  19. 19. Expert Review Requires Evidence, Not Confidence

    AICH4 · expert review

    Rule: industry standard

    Constraint: security/regulation

    Exception: local domain rule

    Decision: who accepts residual risk

    AICH5 · evidence-based measure

    Quality attribute: what to protect

    Signal: what to observe

    Baseline: what to compare

    Threshold: when the decision changes

  20. 20. ArchBench Standardizes the Pipeline, Not Metric Validity

  21. 21. R2ABench Separates Form, Graph, Meaning, and Evidence

  22. 22. Knowing the Answer Is Not Owning the Trade-off

  23. 23. A Benchmark Measures Capability, Not Industrial Impact

    Benchmark

    Fixed unit of analysis

    Replayable input and reference

    Defined metric + baseline

    Controlled scope and cost

    Production

    Incomplete and changing facts

    Organizational constraints

    Conflicting quality attributes

    Consequences and ownership over time

  24. 24. Form Improves Faster Than Architectural Coherence

  25. 25. The Roadmap Begins with Knowledge Infrastructure

  26. 26. A Decision Must Trace to Runtime and Back

  27. 27. AI Analyzes. The Architect Owns the Trade-off

  28. 28. Architecture Knowledge Is Not a Large Prompt

    Prompt

    A one-run context snapshot

    Implicit version and freshness

    Weak provenance

    No owner for a claim

    Living knowledge

    Schema + typed relations

    Versions + freshness policy

    Evidence + provenance

    Owner + review workflow

  29. 29. Walk One Change in Both Directions

  30. 30. Start with a Workflow, Not a Product Category

  31. 31. What the Evidence Supports — and Where It Ends

    Main paper: 51 studies → 14 areas → 6 gaps → 5 pillars

    Local results: 70% accuracy · ≈75% recall · n=95 ADRs · up to 34%

    A 21-tool snapshot shows local intelligence, not one shared loop

    ArchBench / R2ABench / CAKE / SAKE improve evaluation, not production-impact evidence

    Practical start: living knowledge + bridge evaluation + human accountability

    AI makes form cheaper—so coherence, evidence, and ownership become more valuable

  32. 32. Architecture Memory Matters More Than Another Answer

    The longread, paper, and replication package are linked from the first slide

    Book Cube

    Continue the discussion on architecture, AI4SDLC, and engineering research in the channel.

    Alexander Polomodov, Technical Director & Fellow, T-Technologies

    @Book_Cube