Skip to content
back to the archive page
#AI

Stanford CME295, Lecture 1: How Text Becomes Computation (Category #AI)

The first lecture of Stanford CME295, a course I’ve discussed before, turned out to be an interesting one—in 1 hour and 44 minutes, Afshine and Shervine Amidi explain how to turn text into numbers, connect those numbers with context, and put it all together into a working translation model. The recording’s strength is that every new component appears as an answer to a concrete problem. So even in the rather imposing Transformer diagram, the purpose of the individual blocks gradually becomes clear. Below, I’ll go over the main points of the lecture.

1️⃣ Tokens: what pieces should we split text into? You can use whole words, but then the vocabulary grows large and unfamiliar words cause problems. You can use individual characters, but the sequences become long and expensive to process. A common middle ground is parts of words. BPE builds a vocabulary by repeatedly merging frequently occurring pairs. A token therefore does not have to correspond to a word, and the way text is split depends on the data used to train the tokenizer.

2️⃣ Embeddings: how do we give those pieces a useful numerical representation? Vocabulary indices merely distinguish tokens from one another. For the model to work, each token needs a vector—a set of numbers learned along with the model. Using Word2vec as an example, the instructors show how training on text can produce useful relationships between words. But a fixed vector is not enough: the English word bank can mean a financial institution or the side of a river. The context changes, while the word’s initial representation stays the same.

3️⃣ Attention: how do we account for the surrounding text? RNNs and LSTMs process a sequence step by step, updating an internal state. It is a bit like retelling a book using one note that you keep updating: distant details are hard to retain, and each computation depends on the previous step.

Attention creates direct connections between token representations. The model computes weights for those connections and combines information according to those weights. Q, K, and V are conveniently understood as “what we’re looking for,” “what we match against,” and “what information we take.” This is an explanatory metaphor: internally, everything happens through learned transformations of vectors. Self-attention lets a token’s representation be updated using other tokens in the same sequence.

4️⃣ Transformer: how do we put the pieces together? In the original architecture, the encoder builds context-aware representations of the input text, while the decoder generates translation tokens one at a time, consulting those representations and the text already generated. Positional information helps account for order. A mask prevents the model from peeking at future tokens during training. Neural network layers transform the information gathered, while additional connections and normalization help train a deep model.

At the end, the instructors work through all of this using a single sentence about a teddy bear: from tokenization to next-token probabilities and the finished translation. This makes it particularly easy to check how the individual mechanisms work together. They specifically examine the Transformer from the 2017 paper; other model families are announced for the next lecture.

The lecture will be useful to people who already use LLMs and want to understand how they work, developers of AI applications, and anyone beginning to study the subject systematically. The explanations are accessible, but familiarity with vectors, matrix multiplication, and basic neural network training concepts will help: the instructors work through formulas and dimensions. I’d pause as I watched, and after the end-to-end example, try drawing the entire path from text to the next token myself.

#AI #LLM #Learning #Engineering #Architecture

Open video on YouTube