
Why the Architect AI Copilot Still Hasn’t Arrived
Unpacking a systematic review of 51 studies with Sergey Baranov
Slide contents
1. Why the Architect AI Copilot Still Hasn’t Arrived
Unpacking a systematic review of 51 studies with Sergey Baranov
2. AI Sees Snapshots; Architecture Keeps History
3. Diagrams Capture Form; Decisions Hold Trade-offs
Artifact
Diagram
ADR
Pattern list
Architecture
Why now
Cost of change
Consequences later
4. Review Maps Evidence, Not Copilot Performance
5. 51 Studies Survived Multi-stage Screening
6. Broad Coverage, Shallow Integration
7. Wider Decisions Are Harder to Verify
8. Three Experiments Show Value — and Its Limits
9. LLM Produces Candidates, Not Approved ADRs
10. Runtime Gives AI a Measurable Outcome
11. Results Hold Within Their Measurement Boundary
What was measured
Extraction / patterns · 4 projects / 3 classes
ADR candidate · n=95
Deployment time · AKS testbed
What does not follow
Requirements and trade-offs are covered
The system boundary is optimal
The decision survives releases
12. Twenty-one Tools Span Four ISL Levels
13. Tools Create Islands of Intelligence
14. Artifacts Exist. The Architecture Chain Does Not
15. Fifteen Problems Become Six AI Requirements
16. Six Gaps Remain Before a Real Copilot
17. Adaptation Requires Memory of the Long Horizon
AICH1 · update the recommendation
A requirement version changed
Find dependent decisions
Rebuild the option + evidence
AICH6 · retain consequences
Observe debt and smells
Link versions, incidents, and erosion
Return the signal to the decision
18. Traceability Without Context Remains Formal
AICH2 · continuous traceability
Requirement ↔ ADR · type/version
ADR ↔ component · rationale
Component/code ↔ runtime · evidence
AICH3 · local context
Domain rules and data
Teams · regulation · security boundaries
Migration cost + legacy constraints
19. Expert Review Requires Evidence, Not Confidence
AICH4 · expert review
Standards · security · regulation
Local domain exceptions
Owner accepts residual risk
AICH5 · evidence-based measure
Quality attribute
Signal + baseline
Threshold changes the decision
20. ArchBench Standardizes the Pipeline, Not Metric Validity
21. R2ABench Separates Form, Graph, Meaning, and Evidence
22. Knowledge Is Not Decision Ownership
23. Benchmarks Measure Capability, Not Impact
Benchmark
Fixed evaluation object
Replayable input + baseline
Defined metric, scope, and cost
Production
Incomplete, changing facts
Constraints + conflicting qualities
Consequences and ownership over time
24. Form Improves Faster Than Architectural Coherence
25. The Roadmap Begins with Knowledge Infrastructure
26. Decisions Must Trace to Runtime and Back
27. AI Analyzes. The Architect Owns the Trade-off
28. Architecture Knowledge Is Not a Large Prompt
Prompt
One-run context snapshot
Implicit version + freshness
Weak provenance; no claim owner
Living knowledge
Schema + typed relations
Versions + freshness policy
Evidence + provenance + owner
29. Walk One Change in Both Directions
30. Start with a Workflow, Not Products
31. What the Evidence Supports
51 studies map the field, not ROI
Narrow tasks show bounded value
21 tools automate separate islands
Benchmarks stop before production impact
Start with living knowledge + human accountability
Cheaper form raises the value of coherence, evidence, and ownership
32. Architecture Memory Matters More Than Another Answer
The longread, paper, and replication package are linked from the first slide
Book Cube
Continue the discussion on architecture, AI4SDLC, and engineering research in the channel.
Alexander Polomodov, Technical Director & Fellow, T-Technologies
@Book_Cube