Skip to content
All education
HSE University / proposed syllabus

Distributed Systems

Learn to design systems where the network is unreliable, data lives on several nodes, and failure is part of normal operation.

One service, several independent nodesWhat if a node fails?clientAPIstatequeueworkersdelays · retries · failures
months
06
lectures
10
labs
04
end-to-end project
01

A proposed six-month syllabus. Dates, the timetable, and grading rules will be confirmed before launch.

01 / learning path

From failure models to a working service

Two lectures per month for the first five months, with practice between sessions. Month six is for integration, experiments, and the final project defense.

  1. 01 / Month

    Models

    Failures and time

    Lectures 01–02

  2. 02 / Month

    Communication

    Networks and replicas

    Lectures 03–04 · lab 01

  3. 03 / Month

    Coordination

    Consensus and data

    Lectures 05–06 · lab 02

  4. 04 / Month

    Interaction

    Transactions and events

    Lectures 07–08 · lab 03

  5. 05 / Month

    Reliability

    Testing and architecture

    Lectures 09–10

  6. 06 / Month

    Integration

    Experiments and defense

    Lab 04 · final project

02 / theory

10 lectures. One coherent system.

Each topic starts with an engineering question and leads to a concrete outcome. Expand a lecture to see its content.

01Start with failureMonth 1

Failure models, network latency, partial failures, and invariants.

After the lecture: Define what a service must preserve when a node fails.

02Time and event orderingMonth 1

Physical and logical clocks, causality, Lamport clocks, and vector clocks.

After the lecture: Reconstruct causal order from the logs of several nodes.

03The network is part of the applicationMonth 2

RPC, timeouts, retries, idempotency, and backpressure.

After the lecture: Design an API that handles lost responses and duplicate delivery.

04Replication and consistencyMonth 2

Leaders and replicas, quorums, consistency models, CAP, and PACELC.

After the lecture: Explain the trade-offs between availability, latency, and consistency.

05Consensus and coordinationMonth 3

Raft, leader election, replicated logs, and fault-tolerance limits.

After the lecture: Trace a leader change and the minority partition during a network split.

06Partitioning data and loadMonth 3

Sharding, key selection, consistent hashing, hot keys, and rebalancing.

After the lecture: Choose a partitioning scheme and identify its bottlenecks.

07Transactions across servicesMonth 4

Isolation, two-phase commit, sagas, compensations, and the outbox pattern.

After the lecture: Preserve a business invariant without a shared cross-service transaction.

08Events, queues, and streamsMonth 4

Message ordering, delivery guarantees, deduplication, and log replay.

After the lecture: Design a consumer that safely processes redelivered events.

09Operations and failure testingMonth 5

SLIs/SLOs, tracing, load experiments, fault injection, and recovery.

After the lecture: Test invariants and measure recovery after a failure.

10Architecture as a set of decisionsMonth 5

Putting a system together: requirements, latency budgets, constraints, and reliability costs.

After the lecture: Defend a service architecture with measurements and explicit trade-offs.

03 / practice

Four labs. One project.

Build a distributed task service: an API accepts requests, storage preserves state, and a queue dispatches work to workers. Each lab adds a new layer and a failure scenario.

BuildBuildBreakBreakMeasureMeasure
The learning loop in every lab: build, break, measure. Experiment findings → the next version of the system.
LAB 01Month 2 · after lecture 4

The network is unreliable

Build a task-service API. Add timeouts, retries, and idempotency keys. Simulate a lost response.

What to verify

A repeated request creates no duplicate task; a test reproduces the lost response.

What to submit: API + failure tests

LAB 02Month 3 · after lecture 6

One node is not a system

Add replicated storage using an existing consensus implementation. Stop the leader and partition the network.

What to verify

Record read/write availability, durability of acknowledged data, and recovery time.

What to submit: Cluster + experiment report

LAB 03Month 4 · after lecture 8

The event arrives twice

Add a queue and task workers. Implement an outbox and deduplication; test a crash between saving a result and acknowledging a message.

What to verify

Redelivery creates no duplicate business effect; unfinished work resumes after a restart.

What to submit: Workers + invariant checks

LAB 04Month 6 · after lecture 10

Put the system to the test

Integrate the service and add metrics and tracing. Run a load experiment and a failure scenario, then defend your design.

What to verify

A reproducible setup, measured latency and recovery, and documented system limitations.

What to submit: Service + report + defense

04 / outcomes

An architecture needs evidence

By the end of the course, students will have a working service and experiments that demonstrate its guarantees and limitations.

  • Define invariants and choose a consistency model for the problem.
  • Explain system behavior under delays, retries, and partial failures.
  • Test design decisions through experiments and defend architectural trade-offs.

Who it is for

Students who can already build applications and want to understand what changes when a single process becomes a system of several nodes.

Before you start

One programming language, HTTP and database basics, Git, and basic container skills. No prior distributed systems experience is required.