Skip to content
Course syllabus

Lab 01: The network is unreliable

Build a task-service API. Add timeouts, retries, and idempotency keys. Simulate a lost response.

Lab source, setup, and report template

Lab 1. A lost response does not imply a lost operation

Build a reliable HTTP contract for one task service. You should be able to distinguish retries from new operations, justify the atomic boundary, and reproduce an uncertain outcome. Prerequisites: HTTP, Python, Docker, and the lectures on failure models and service communication. Allow about 90 minutes plus independent completion; this is an author proposal, not an official HSE timetable.

Setup and execution

Use Python 3.12+, Docker and Compose. The shared guide covers installation, the API contract, isolation and cleanup. Run from the repository root:

python3 -m venv /tmp/ds-course-venv
/tmp/ds-course-venv/bin/python -m pip install -r labs/distributed-systems/requirements.txt
PYTHONDONTWRITEBYTECODE=1 /tmp/ds-course-venv/bin/python labs/distributed-systems/runner.py --lab 1 --solution starter --report /tmp/lab01.json

The starter runs a real API and etcd cluster, but intentionally fails atomicity checks. Edit starter/protocol.py and rerun the same command. Instructors can select --solution reference. Both implementations run the assertions in tests/scenarios.py; weakening assertions is not a solution.

Assignment

  1. Declare the Idempotency-Key scope: one run namespace, 1–128 ASCII characters, retained for the entire run. The body is {"amount": 1} with an optional label; see README for bounds. The same key and JSON have one identity. Changed content returns 409.
  2. Replace check-then-write with a conditional etcd transaction. Task, idempotency and outbox must have the same revision: preserve publication intent now so Lab3 can extend the same protocol. Lab1 does not require a broker to complete work.
  3. Specify client behavior after a timeout: reuse the original key and do not infer that the write never happened. GET /tasks/<id> uses linearizable reads.

Three scenarios

  • Response lost after commit. The runner sends X-Lab-Drop-Response: true; the API commits and closes the real TCP connection. The first response is unavailable. Retrying returns 200 and the original task_id; a different amount returns 409.
  • Concurrent retries. Twelve threads send an identical request together. Exactly one receives 201, eleven receive 200, and all share one task_id. Two concurrent requests with one key but different amounts produce one 201 and one 409.
  • API crash. After an acknowledged creation, the API receives SIGKILL and restarts. Retrying returns the previous task_id, read from etcd rather than process memory.

Invariant: at most one accepted operation per key during retention; task, idempotency and outbox commit atomically. An acknowledged task survives an API crash under this failure model. This does not yet guarantee worker completion or global “exactly once.”

Deliverables

  • Starter diff explaining the linearization point and the losing compare operation.
  • /tmp/lab01.json and a small JSONL excerpt with task_id/trace_id; include versions, command and attempt count.
  • A sequence showing commit → lost response → retry, explaining why two competitors cannot both receive 201.
  • A report retaining the failed baseline and the corrected run. Do not commit virtual environments, container volumes or full diagnostic logs.

Discussion and reading

Relate operation boundaries to Lamport’s Time, Clocks, and the Ordering of Events, and local atomicity to Helland’s Life beyond Distributed Transactions. The instructor’s Lamport and Helland research decks support the exercise; this public handout links primary sources. The implementation mechanism is etcd transactions.

Explain what changes when idempotency expires after an hour, JSON serialization differs, or an external payment API replaces the local effect. Extension: use two keys for one business intention and show why transport idempotency does not recognize the relationship automatically.

Recovery and limitations

The runner cleans only its unique Compose project, including volumes. Ctrl+C triggers cleanup; after SIGKILL use the saved compose.env and README recovery command. Ports bind to loopback. Never use a global prune. Whole-host loss, majority-disk loss and Byzantine behavior are not tested. A timeout can mean an unknown outcome rather than a rejected write.