Operations and failure testing
Which observations establish correct service recovery?
Slide contents
1. Operations and failure testing
Which observations establish correct service recovery?
2. A running process may not serve users
Check a completed operation through the real path.
3. Correctness and speed need separate checks
A fast response can violate an invariant.
4. An SLI defines numerator and denominator
The user contract defines a good event.
5. An SLO adds a target and a window
A percentage needs a time window.
6. The error budget uses the same events
At 99%, 1,000 attempts allow 10 failures.
7. Mean latency hides the tail
Threshold ratios and distributions answer different questions.
8. Task acceptance differs from completion
A fast API can hide a growing queue.
9. An experiment needs a comparable baseline
A fault changes one explicitly named factor.
10. Metrics, logs and traces complement one another
Each signal preserves different details.
11. A trace connects intervals of one operation
A span records work boundaries and context.
12. Context travels with the message
Trace ID and eventId have different responsibilities.
13. Sampled traces are not a complete counter
A missing trace does not mean no operation occurred.
14. Workload describes arrivals, not just threads
Two generators can create different experiments.
15. The load generator also has a capacity limit
Check that the source can produce the workload.
16. Write the experiment protocol before running
Define stopping and checking criteria in advance.
17. Fault injection respects the declared model
Stopping, partitioning and disk loss differ.
18. The history preserves unknown outcomes
A timeout does not become a proven abort.
19. The invariant checker is part of the test
A restarted process does not imply recovered results.
20. A planned handoff differs from sudden failure
The same stopped-node count does not equalize scenarios.
21. Recovery has several time boundaries
Each boundary answers a different question.
22. A better tail can cost extra work
Measure p99 with extra requests and cancellation.
23. A report connects each claim to an artifact
Expected behavior is not observed evidence.
24. Recalculate availability and recovery
Use the handout’s synthetic history.
Define denominator
Separate intervals
Find the unsupported claim
25. A fault test ends by checking meaning
Measurement, invariant and limits form one conclusion.
26. Decisions for our service
Define an SLI, window and error budget.
Preserve a reproducible fault-experiment history.
Separate election, client progress and restored redundancy.