
Why the Architect AI Copilot Still Hasn’t Arrived
A review of 51 studies with Sergey Baranov

A review of 51 studies with Sergey Baranov
A review of 51 studies with Sergey Baranov
Artifact
Diagram
ADR
Pattern list
Architecture
Why now
Cost of change
Consequences later
What was measured
Responsibility recall · 4 projects
Pattern accuracy · 3 classes
ADR-candidate completeness · n=95
Deployment time · AKS testbed
What does not follow
All requirements are covered
Conflicting qualities are balanced
The best system boundary was chosen
The decision survived multiple releases
AICH1 · update the recommendation
Fact: a requirement version changes
Trigger: code/runtime drift
Operation: find dependent decisions
Output: revised option + evidence
AICH6 · retain consequences
Observe: debt and smells
Align: versions and incidents
Assess: erosion across releases
Return: signal into the next decision
AICH2 · continuous traceability
Requirement ↔ ADR · type and version
ADR ↔ component · rationale
Component ↔ code · ownership
Decision ↔ runtime · evidence
AICH3 · local context
Domain rules and data
Team topology and ownership
Regulation and security boundaries
Migration cost and legacy constraints
AICH4 · expert review
Rule: industry standard
Constraint: security/regulation
Exception: local domain rule
Decision: who accepts residual risk
AICH5 · evidence-based measure
Quality attribute: what to protect
Signal: what to observe
Baseline: what to compare
Threshold: when the decision changes
Benchmark
Fixed unit of analysis
Replayable input and reference
Defined metric + baseline
Controlled scope and cost
Production
Incomplete and changing facts
Organizational constraints
Conflicting quality attributes
Consequences and ownership over time
Prompt
A one-run context snapshot
Implicit version and freshness
Weak provenance
No owner for a claim
Living knowledge
Schema + typed relations
Versions + freshness policy
Evidence + provenance
Owner + review workflow
Main paper: 51 studies → 14 areas → 6 gaps → 5 pillars
Local results: 70% accuracy · ≈75% recall · n=95 ADRs · up to 34%
A 21-tool snapshot shows local intelligence, not one shared loop
ArchBench / R2ABench / CAKE / SAKE improve evaluation, not production-impact evidence
Practical start: living knowledge + bridge evaluation + human accountability
AI makes form cheaper—so coherence, evidence, and ownership become more valuable
The longread, paper, and replication package are linked from the first slide
Book Cube
Continue the discussion on architecture, AI4SDLC, and engineering research in the channel.
Alexander Polomodov, Technical Director & Fellow, T-Technologies
@Book_Cube