Skip to content
back to the archive page
#AI

Stanford CME295: Getting into the Foundations of LLMs (Category #AI)

I’ve started studying Stanford CME295 — Transformers & Large Language Models, a course I’d like to recommend to everyone. It’s an excellent way to get into the foundations: the instructors explain complex concepts in very concrete, accessible terms. They first show the problem that needs solving, work through a simple example, and then move on to the formulas. That makes it much easier to understand why a particular component exists in the model at all.

The course is taught by twin brothers Afshine and Shervine Amidi, both Adjunct Professors at Stanford ICME. Both studied at École Centrale Paris; Afshine then went to MIT, and Shervine to Stanford. In the first lecture, they also describe their industry careers: Uber, Google, and now Netflix.

The website already has all nine lectures from 2025, so you can work through the entire course right away. Meanwhile, the 2026 course started just a week ago, on September 25. According to the instructors, the content will change substantially: the past year has brought new approaches to model training, agents, and text generation itself. The syllabus already shows where those changes are heading.

Nine lectures are planned, following a fairly clear progression: 🔸 How the model works: tokens, vector representations, attention, and the Transformer. Then LLM families, context, and strategies for choosing the next token during generation. 🔸 Training: from pretraining to SFT and RLHF, LoRA, reinforcement learning, and distillation. 🔸 Systems and agents: speeding up models, caching intermediate computations, connecting tools, organizing memory, and managing context. 🔸 Evaluation and new directions: assessing model and agent quality, understanding where LLM-as-a-judge goes wrong, and exploring diffusion language models and multimodality.

Incidentally, the 2026 edition gives LLM systems and reinforcement learning their own lectures, adds a lecture on diffusion LLMs, and introduces MCP, context compaction, skills, and plugins in the agents syllabus. So there will be plenty to follow even if you’ve already watched last year’s recordings.

The explanations are accessible, though familiarity with linear algebra, calculus, and basic machine learning concepts will help. There are formulas here, and they are worth understanding. Slides and a concise cheatsheet are available on the website; the course has no homework, but it does have two exams.

In the next post, I’ll discuss the first lecture of the new course, whose recording was published this week, on September 28. It builds the foundations: from splitting text into tokens to the Transformer and an end-to-end sentence translation example.

#AI #LLM #Learning #Engineering #Agents

Open video on YouTube

Public sources