Start with failure
What must a service preserve when part of the system is unavailable?
Slide contents
1. Start with failure
What must a service preserve when part of the system is unavailable?
2. API success creates a system obligation
An acknowledged task must survive the stated failure model.
3. A service has independent failure points
API, storage, and broker can fail differently.
4. Parts of a system keep working during a failure
Client and server can observe different outcomes.
5. Crash-stop removes all future process steps
A stopped participant never returns in that execution.
6. Recovery depends on memory that outlives the process
RAM disappears; acknowledged durable records must remain.
7. A message may disappear, repeat, or arrive late
Network arrival need not match application call order.
8. Silence does not prove that a participant died
Slow computation and a broken link may look identical.
9. A false reply requires a different protection model
A crash protocol need not tolerate arbitrary rule violations.
10. The lab model bounds its promises
Losing all durable copies is not an ordinary node failure.
11. Safety rules out a bad state
One key must not create two different tasks.
12. Progress needs environmental conditions
Processing should resume after recovery.
13. Durable data can be temporarily unavailable
No response does not imply a lost record.
14. An invariant describes what state means
Verification needs keys, tasks, and results.
15. The basic effect stays inside storage
The lab effect is a stored task result.
16. Acknowledgement follows the durable decision
Replying before persistence breaks the promise.
17. A reply can disappear after commit
Synthetic trace of one accepted task.
18. Unknown is a distinct client outcome
It requires result recovery rather than guessing.
19. Creation links state with delivery intent
Task, idempotency, and outbox commit in one txn.
20. ACK follows result commit
Result, dedup, and completed change together.
21. Completion needs several live dependencies
etcd, broker, and worker serve different roles.
22. A lost ACK does not require a second result
Synthetic trace: task T7, two message deliveries.
23. A failure test checks recovered state
“The process stayed up” is not an invariant check.
24. Separate broken invariants from stalled progress
Three histories require different conclusions.
25. A report starts with promises and assumptions
Then come scenario, observations, and conclusion.
26. Decisions for our service
Separate observations from failure assumptions.
State safety and the conditions for progress.
Check an acknowledged task after a failure.