Skip to content
Course syllabus

Lab 04: Put the system to the test

Integrate the service and add metrics and tracing. Run a load experiment and a failure scenario, then defend your design.

Lab source, setup, and report template

Lab 4. User-visible outcomes under load and failure

Validate the complete project while preserving the invariants of Labs 1–3. Measure task completion, backlog, errors and recovery alongside HTTP responses. Prerequisites: corrected earlier labs and the lectures on tail latency, observability and operational failure. Allow about 90 minutes plus defense preparation.

Run

After README setup, run from the repository root:

PYTHONDONTWRITEBYTECODE=1 /tmp/ds-course-venv/bin/python labs/distributed-systems/runner.py --lab 4 --solution starter --report /tmp/lab04.json
PYTHONDONTWRITEBYTECODE=1 /tmp/ds-course-venv/bin/python labs/distributed-systems/runner.py --lab all --solution starter --report /tmp/project-final.json

Use --solution reference for the reference implementation. Load generation and assertions live in tests/scenarios.py. Baseline counts are 24, 24 and 32 unique tasks, one sequential client, and an intentional pause of about 15 ms between initial requests. This is a reproducible small experiment, not a saturated throughput benchmark.

Assignment

  1. Write hypotheses for admission latency and time to completion before running. Define denominators: initial requests, all attempts including retries, and unique tasks.
  2. Connect JSON operations with API/relay/worker JSONL using task_id, event_id and trace_id. Identify where “HTTP succeeded” still does not mean “the effect completed.”
  3. Change one parameter: offered rate, worker delay or timeout. Retain baseline and experiment; avoid changing everything at once. Explain the cost in backlog, errors, client waiting and duplicate work.
  4. Prepare an architecture defense covering atomicity, reads, retry/retention, capacity assumptions, recovery and the excluded external effect.

Three scenarios

  • Slow worker. The worker waits 120 ms per delivery while the client offers 24 tasks. A real backlog must accumulate and later drain. Distinguish admission latency, completion latency and backlog drain time.
  • Broker stopped. The NATS container stops while HTTP accepts 24 tasks into etcd. Pending outbox grows. After the broker returns with the same volume, all tasks must complete without loss or repeated effects.
  • Leader change under load. The actual etcd leader stops during a stream of 32 tasks. Errors and uncertain outcomes are retried with their original keys. After processing and return of the third member, compare every identity, result and the counter.

Final baseline invariant: 80 unique amount=1 tasks produce 80 results and counter {effects:80,value:80} after recovery, regardless of attempts and deliveries. Errors during faults are allowed but must remain visible. Zero observed errors in a short run does not prove continuous availability.

Deliverables

  • Machine reports for all four suites, the baseline and one modified experiment. Include command, versions, resources, counts and observation window.
  • HTTP admission p50/p95/p99 with sample counts and errors; separate completion latency, backlog and defined recovery intervals. State percentile method and polling uncertainty.
  • A backlog plot or several timestamped snapshots, one request’s causal chain, and an explanation of its longest delay. There is no single synchronized wall clock across a trace; runner durations use its monotonic clock.
  • A report, ADR, diff and reproduction instructions. Proposed defense criteria are correctness, reproducibility, measurement and honest limitations; the instructor sets the official grading scale.

Reading and defense

The Tail at Scale frames latency tails and retry cost; Simple Testing Can Prevent Most Critical Failures supports targeted fault testing; MapReduce discusses repeated execution and slow tasks. Supporting research decks: Tail at Scale, Simple Testing, MapReduce and Kafka. For user outcomes, read SRE Workbook: Implementing SLOs.

Explain why a fast 201 can coexist with slow completion; why p99 from 24 samples is weak production evidence; where retries amplify load; and why an external payment needs a new protocol. Propose an alternative to the shared counter, stating the changed invariant and new tests it requires.

Recovery and limits

The runner cleans only its own project; manual cleanup parameters after SIGKILL are saved beside the report, as described in README. One physical host, one broker, short local runs, retained disks and bounded payloads are part of the model. Do not extrapolate numbers to WAN or different hardware. Distinguish observing a leader, client availability and restored redundancy; these are not interchangeable timings.

The baseline ThreadingHTTPServer has no application-level HTTP concurrency limit. This experiment studies work backlog under bounded offered load, not server stability under unbounded HTTP overload. See the final project for deliverables and proposed defense weights.

Proposed defense criteria

This is an author proposal, not HSE policy. Correctness — 40%: mandatory invariants across all four suites; reproducibility — 25%: commands, versions, real faults and owned-resource cleanup; measurements — 20%: counts/errors, admission/completion, backlog and defined intervals; reasoning — 15%: two ADRs, boundaries and a justified alternative. Full credit requires verified evidence; partial credit identifies a concrete disclosed gap; zero means missing or substituted evidence. A lost or repeated effect prevents technical completion. The instructor announces the passing threshold and exact scores. Full rubric with levels.