Skip to content
HSE · Lecture 01

Start with failure

What must a service preserve when part of the system is unavailable?

Distributed Systems · HSE · 01

Slide contents

  1. 1. Start with failure

    What must a service preserve when part of the system is unavailable?

  2. 2. API success creates a system obligation

    An acknowledged task must survive the stated failure model.

  3. 3. A service has independent failure points

    API, storage, and broker can fail differently.

  4. 4. Parts of a system keep working during a failure

    Client and server can observe different outcomes.

  5. 5. Crash-stop removes all future process steps

    A stopped participant never returns in that execution.

  6. 6. Recovery depends on memory that outlives the process

    RAM disappears; acknowledged durable records must remain.

  7. 7. A message may disappear, repeat, or arrive late

    Network arrival need not match application call order.

  8. 8. Silence does not prove that a participant died

    Slow computation and a broken link may look identical.

  9. 9. A false reply requires a different protection model

    A crash protocol need not tolerate arbitrary rule violations.

  10. 10. The lab model bounds its promises

    Losing all durable copies is not an ordinary node failure.

  11. 11. Safety rules out a bad state

    One key must not create two different tasks.

  12. 12. Progress needs environmental conditions

    Processing should resume after recovery.

  13. 13. Durable data can be temporarily unavailable

    No response does not imply a lost record.

  14. 14. An invariant describes what state means

    Verification needs keys, tasks, and results.

  15. 15. The basic effect stays inside storage

    The lab effect is a stored task result.

  16. 16. Acknowledgement follows the durable decision

    Replying before persistence breaks the promise.

  17. 17. A reply can disappear after commit

    Synthetic trace of one accepted task.

  18. 18. Unknown is a distinct client outcome

    It requires result recovery rather than guessing.

  19. 19. Creation links state with delivery intent

    Task, idempotency, and outbox commit in one txn.

  20. 20. ACK follows result commit

    Result, dedup, and completed change together.

  21. 21. Completion needs several live dependencies

    etcd, broker, and worker serve different roles.

  22. 22. A lost ACK does not require a second result

    Synthetic trace: task T7, two message deliveries.

  23. 23. A failure test checks recovered state

    “The process stayed up” is not an invariant check.

  24. 24. Separate broken invariants from stalled progress

    Three histories require different conclusions.

  25. 25. A report starts with promises and assumptions

    Then come scenario, observations, and conclusion.

  26. 26. Decisions for our service

    Separate observations from failure assumptions.

    State safety and the conditions for progress.

    Check an acknowledged task after a failure.