Skip to content
HSE · Lecture 09

Operations and failure testing

Which observations establish correct service recovery?

Distributed Systems · HSE · 09

Slide contents

  1. 1. Operations and failure testing

    Which observations establish correct service recovery?

  2. 2. A running process may not serve users

    Check a completed operation through the real path.

  3. 3. Correctness and speed need separate checks

    A fast response can violate an invariant.

  4. 4. An SLI defines numerator and denominator

    The user contract defines a good event.

  5. 5. An SLO adds a target and a window

    A percentage needs a time window.

  6. 6. The error budget uses the same events

    At 99%, 1,000 attempts allow 10 failures.

  7. 7. Mean latency hides the tail

    Threshold ratios and distributions answer different questions.

  8. 8. Task acceptance differs from completion

    A fast API can hide a growing queue.

  9. 9. An experiment needs a comparable baseline

    A fault changes one explicitly named factor.

  10. 10. Metrics, logs and traces complement one another

    Each signal preserves different details.

  11. 11. A trace connects intervals of one operation

    A span records work boundaries and context.

  12. 12. Context travels with the message

    Trace ID and eventId have different responsibilities.

  13. 13. Sampled traces are not a complete counter

    A missing trace does not mean no operation occurred.

  14. 14. Workload describes arrivals, not just threads

    Two generators can create different experiments.

  15. 15. The load generator also has a capacity limit

    Check that the source can produce the workload.

  16. 16. Write the experiment protocol before running

    Define stopping and checking criteria in advance.

  17. 17. Fault injection respects the declared model

    Stopping, partitioning and disk loss differ.

  18. 18. The history preserves unknown outcomes

    A timeout does not become a proven abort.

  19. 19. The invariant checker is part of the test

    A restarted process does not imply recovered results.

  20. 20. A planned handoff differs from sudden failure

    The same stopped-node count does not equalize scenarios.

  21. 21. Recovery has several time boundaries

    Each boundary answers a different question.

  22. 22. A better tail can cost extra work

    Measure p99 with extra requests and cancellation.

  23. 23. A report connects each claim to an artifact

    Expected behavior is not observed evidence.

  24. 24. Recalculate availability and recovery

    Use the handout’s synthetic history.

    Define denominator

    Separate intervals

    Find the unsupported claim

  25. 25. A fault test ends by checking meaning

    Measurement, invariant and limits form one conclusion.

  26. 26. Decisions for our service

    Define an SLI, window and error budget.

    Preserve a reproducible fault-experiment history.

    Separate election, client progress and restored redundancy.