
Building evals that work
Replayable episodes, hidden judges, and a production scorecard. With Evgeny Sergeev, Engineering Director at Flo Health

Replayable episodes, hidden judges, and a production scorecard. With Evgeny Sergeev, Engineering Director at Flo Health
Replayable episodes, hidden judges, and a production scorecard. With Evgeny Sergeev, Engineering Director at Flo Health
Correct outcome, unsafe path
Unsafe diff passes tests
Future solution leaks into context
Rerun takes another path
Correct SQL, wrong grain
The model is only one part of an agent loop
what — Outcome — Actual task result
how — Trace — Tools and actions
how much — Operations — Cost and stability
Freeze the work, constraints, judge, and release criteria
Start state → contract → hidden judge → trace → release gate
Otherwise, different environments and privileges get compared
Frozen start state
One commit or snapshot
Same data and services
Future solution stays hidden
Agent contract
Allowed tools and network
Read, write, approve
Attempts, time, budget
Persuasive language is not evidence
Code: hidden tests and policies
Workflow: tool calls and arguments
Data: state and invariants
Review: findings become checks
Correct output cannot excuse harm
The right tool was selected
Arguments match the contract
No unsafe side effects
No needless loops or retries
A run series matters more than the best run
k runs — Pass^k — Success across runs
σ — Variance — Trajectory dispersion
$ / task — Budget — Cost of stable success
Regression needs a holdout; reality needs fresh tasks
Static holdout
Stable regression gate
Agent-version comparison
Solutions remain hidden
Rolling live set
Fresh PRs and incidents
New data questions
Changing work environment
From a business brief to incidents and architecture decisions
Requirements → code → review → tests → operations
A PRD must survive implementation
Traceability from brief to AC
Functional requirements stay complete
NFRs and constraints survive
Hallucinations stay outside scope
The same pass rate can hide different risks
Feature coding
Spec + pre-PR commit
Hidden integration tests
Full suite and policy
Bug fixing
Issue + buggy commit
Fail-to-pass regression
Adjacent scenarios stay safe
Comments and coverage alone establish very little
Code review
Severity-weighted recall
False positives are bounded
Finding accepted or executed
Test generation
Fail before, pass after
Branch and mutation coverage
Stable without flakiness
“What happened?” is not an eval by itself
Alert, topology, and telemetry
Runbooks and allowed actions
Triage, RCA, mitigation
Safe handoff to a human
Architecture and data platforms have no single gold answer
Hard constraints → completeness → downstream executability
Analytical answers and pipelines need different evidence
Outcome, trace, operations, safety, and human acceptance
Weights change by domain; safety remains a gate
Unsafe action stops release
Safety stop
Excess privilege and PII
Unsafe write or remediation
No rollback path
Human acceptance
The PR was accepted
The RCA helped on-call
The team can use it
Scenario family × input × judge × metrics × gate
Input bundle reconstructs the work
Hidden judge establishes outcome
Scorecard explains behavior
Release gate makes the decision
Artifact → frozen state → contract → judge → release gate
What to carry into your engineering system
Evaluate replayable episodes
Check outcome and trace
Repeat the runs
Keep safety as a stop
Refresh the rolling live set
Models change. Your episode catalog remains.
Code, tools, SRE, architecture, and data platforms
SWE-bench · FEA-Bench · τ-bench
SWE-PRBench · TestExplora · AIOpsLab
R2ABench · ArchBench · DAB · BLADE
ELT-Bench · DataGovBench · coSTAR
Production-grade agent evals
The full longread, sources, and future reviews are available at polomodov.tech and in Knizhny kub
Alexander Polomodov, Technical Director & Fellow, T-Technologies
@book_cube