Choose by the work you want to do
A topic list says little about a course’s depth. Discussing Raft, implementing it, and testing a service built on etcd are three different learning outcomes. We therefore compare the practical work and how it is checked, rather than counting mentions of an algorithm.
Entries in the table: 8 / 8
| Course and version | Main emphasis | What students do |
|---|---|---|
| Our proposed HSE courseDraft syllabus · October 2026 | Service guarantees under partial failure | 10 lectures, 4 labs: a task service using etcd and NATS |
| CambridgeConcurrent and Distributed Systems · 2024/25 | Models, time, ordering, and consistency | Exercises and reasoning about algorithms; 8 distributed systems lectures |
| MIT 6.5840Spring 2026 | Systems papers and protocol implementation | Go: from MapReduce to a student-built Raft and sharded store |
| UW CSE452Spring 2025 | Protocol design and invariant checking | Java / DSLabs: 4 labs, design documents, model checking |
| Yandex SDADeveloping Distributed Systems · undated syllabus | Broad foundations: from ABD to TLA+ and CRDTs | Public topic list; assignments and assessment are not described in the reviewed catalog |
| ITMO-MPPAssignment repositories · 2026 snapshots | Individual algorithms with explicit contracts | Kotlin / Java: implementing Raft and consistent hashing |
| CMU 15-440Pittsburgh · Fall 2025 | Systems programming and substantial projects | Go; projects, homework, and exams |
| Gossip GlomersOpen workshop using Maelstrom | Protocols under network delays and partitions | Local processes, network simulation, history checking |
The filter shows our selection of the most relevant examples for each goal, rather than complete topic coverage. We do not combine hours, credits, and assignment counts into a ranking: prerequisites and formats vary too much.
Six approaches and their connections to our course
Cambridge: a language for reasoning about systems
2024/25 · Part IBDistributed systems account for 8 of the combined course’s 16 lectures. Failure models and clocks lead into causality, ordered broadcast, replication, consensus, transactions, and CRDTs: data structures that support consistent merging of independent changes. Students are expected to know object-oriented programming and operating systems.
Connection: our lectures on failure models, time and causality, consistency, and consensus. Cambridge offers a useful example of a progression from message ordering to replicated state machines.
What to borrow: short pseudocode exercises and explanations of counterexamples. Dedicated material on broadcast and CRDTs would extend our theoretical coverage. The reviewed pages do not list a separate assessed programming lab; this does not imply an absence of practical exercises. Lecture recordings for this particular version are available through university services.
MIT 6.5840: build a system from a paper
Spring 2026 · graduate levelThe Go assignments progress through MapReduce, a key/value server, a student-built Raft implementation, and a store built on Raft, before distributing data across replica groups. An approved project can replace the final lab. Paper readings come with short written responses; students have their code checked and discuss it with course staff.
Connection: lecture 5 covers consensus, lecture 6 covers data partitioning and migration, and lab 2 investigates the existing Raft implementation in etcd. MIT requires substantially more independent implementation of these mechanisms.
What to borrow: a separate advanced path for implementing a protocol and migrating partitions. A full Raft implementation is a substantial additional commitment. MIT’s assignment includes persistence and snapshots, but Raft membership changes are outside its scope. The schedule URL has no year and changes over time; this comparison uses the version labeled Spring 2026.
UW CSE452: describe the protocol before writing code
Spring 2025 · senior undergraduate levelFour labs use Java and DSLabs to cover RPC retries, primary/backup replication, Paxos, sharded storage, and transactions. Before implementing the more complex assignments, students describe node states, messages, timers, and invariants. Afterwards, they review the results of their work.
Connection: our first lab also starts with retries and idempotency. Lecture 7 explains atomicity and transactions. The testing methods differ: DSLabs explores event orderings in a model, while our lab environment reproduces specified failures in real components.
What to borrow: a one-page protocol design before coding and an analysis of a discovered bug afterwards. Model checking is limited to the states explored and does not prove correctness for every possible execution. DSLabs itself is open; this does not mean that all GitLab materials for a particular offering are publicly accessible.
Yandex SDA: broad foundations and verification tools
Undated syllabusThe official catalog lists failure models, ABD, gossip, atomic broadcast, Dynamo, Paxos and Raft, transactions, data processing, Jepsen, TLA+, and CRDTs. A background in multithreading is recommended. The fourteen syllabus entries are topics, not a confirmed number of classes.
Connection: the foundations overlap with our lectures on replication and consensus, while failure testing connects to lecture 9. Formal methods and distributed data processing are more prominent in Yandex SDA’s published topic list.
What to borrow: exercises on protocol models and merging independent changes. This page does not establish the assignment language, difficulty, grading rules, or required paper list.
ITMO-MPP: an algorithm with a testable contract
Pinned assignment versions · 2026The open Kotlin/Java assignments ask students to implement Raft and consistent hashing. In Raft, students implement leader election and log replication using the provided storage and state machine. The basic assignment does not include log compaction. The hashing assignment checks key ownership and the ranges transferred when nodes are added or removed.
Connection: our lecture 5 explains the protocol, while lecture 6 covers data placement and movement. A small hashing implementation could connect a lecture diagram to working code.
Scope: this comparison concerns open assignments at pinned repository versions. They do not establish ITMO’s full curriculum, semester workload, or formal admission requirements.
CMU 15-440: a foundation in systems engineering
Pittsburgh · Fall 2025The course combines distributed systems with substantial Go projects. It expects a systems programming background at the level of 15-213/15-513. Assessment includes projects, homework, and exams; the syllabus emphasizes process interaction, unreliable communication, debugging, and handling failures.
Connection: our coverage of RPC and retries, operations, experiments under load, and defending an architecture. The balance between programming and explaining design decisions is a useful reference.
Scope: we compare the Pittsburgh offering. The public pages did not let us reconstruct the Fall 2025 project details, and the syllabus retains passages from earlier years. We therefore do not attribute assignments from archives or the Qatar campus to this offering.
Gossip Glomers: a small protocol under controlled failures
The Fly.io workshop uses Maelstrom for six groups of challenges: echo, unique IDs, broadcast, a counter, a Kafka-style log, and transactions. This is a standalone exercise set, rather than a university semester. Examples use Go; the input/output protocol supports other languages.
Nodes run as local processes exchanging JSON through standard streams. The network is simulated: delays and partitions can be introduced, histories recorded, and their properties checked. Our lab 3 and lab 4 investigate real etcd and NATS instances. These methods complement each other: the first helps uncover protocol bugs, while the second reveals integration behavior.
Fault-tolerant broadcast would make a useful short extension: first achieve delivery after connectivity returns, then measure the message count. A history checker is not an exhaustive correctness proof, and the counter challenge does not, by itself, require students to implement a CRDT.
Our focus: guarantees across component boundaries
In our syllabus, 10 lectures and 4 labs share a single task service. An API accepts requests, etcd stores state, and NATS delivers work to a worker. Students complete two key operations: creating and completing a task. The starter provides the API, message delivery, and experiment runner.
- An uncertain request outcome. In lab 1, the response is lost after a successful write. Students must retry the request while preserving a single logical result.
- Leader changes and network partitions. Lab 2 demonstrates the behavior of three real etcd nodes. Students investigate an existing consensus implementation and explain its limits.
- Storage, delivery, and effects. In lab 3, the task and a record for later publication are stored atomically. After redelivery, the worker must not update the local counter a second time; message acknowledgement follows the result commit.
What is checked, and what is outside the scope. The environment contains 12 specified failure scenarios, rather than a general linearizability checker. All processes run on one physical host; there is one broker, and its disk is retained. The guarantee of a single effect applies to a counter in etcd while deduplication records are retained, not to an external payment. Implementing Raft, actual migration between shards, and loss of the broker’s disk are outside the core labs.
This gives the course an applied focus: the central task is to state a guarantee for the whole operation and test where it breaks. MIT and UW go deeper into independent algorithm implementation; our running example emphasizes the boundaries between HTTP, storage, and a queue.
How the comparison could shape the next revision
The following are the author’s proposals for a future revision. These assignments have not yet been added to the current labs.
- Describe the protocol before coding. Following UW’s example: states, messages, timers, an invariant, and the boundary of the atomic operation. Follow the experiment with a short explanation of what caused a failure. This could apply to all four labs without changing the project.
- Broadcast and merging changes. Add a progression from message ordering to replication in lectures 2, 4, and 8, followed by a small CRDT exercise. Cambridge and Yandex SDA offer theoretical references; Gossip Glomers provides broadcast practice.
- Check histories and search for counterexamples. Extend the specified failures with randomized event sequences and operation history checking. Separately explore a bounded model of a small protocol, as in DSLabs. Retain the integration experiments with real components.
- Offer optional advanced practice. Start with consistent hashing or primary/backup replication. Put partition migration and a full Raft implementation in a separate, substantial block with additional workload, drawing on MIT and ITMO-MPP.
Papers connect the curricula
MIT’s reading list includes MapReduce, GFS, Paxos, Raft, Linearizability, and Spanner; UW’s includes Lamport, Paxos, Bigtable, MapReduce, and Dynamo. In our course, these works connect to lectures and standalone Research Insights Made Simple decks. For example, the reading for lecture 2 explores time and causality, lecture 5 leads into consensus, and lecture 7 into transactions.
A useful addition, following MIT’s example, would be one short response before discussion: “Which assumption does the system rely on, and what breaks if it no longer holds?” A deck helps explain a paper; the response and experiment show whether the student understands its limits.