Skip to content
back to the archive page
#AI

Everything I Learned Training Frontier Small Models — Maxime Labonne, Liquid AI (Category AI)

I watched this short talk by Maxime Labonne, Head of Post-Training at Liquid AI. He explains how the company designs and trains edge models, from architecture to reinforcement learning. The central point is that small models are not simply shrunken large models. Their product constraints differ: they run on phones, laptops, cars, IoT and embedded devices, where memory, latency and performance on specific tasks matter more than general chat ability.

The main ideas:

1️⃣ A small model needs a specialisation A large frontier model can be reasonably good at everything: code, text, reasoning, search, creativity and long context. Small models have fewer parameters, less capacity to retain facts and less room for self-correction. They suit narrow, repeatable workflows: extraction, structured outputs, function calling, tool use and private on-device processing. Liquid AI explicitly describes the limits of LFM2.5-350M.

2️⃣ Parameters can go to the wrong places Labonne’s slides show how much of a small model can be consumed by the embedding layer: 63% for Gemma 3 270M and 29% for Qwen3.5-0.8B. That is expensive: the headline parameter count can overstate useful capacity. LFM2.5-350M devotes 19% to embeddings, with an effective size of 287M, directly affecting latency, memory and quality.

3️⃣ FLOPs are not enough: profile on target hardware Liquid AI checks which operations are actually fast on devices such as the Galaxy S24 Ultra and Ryzen HX 370. If you are building for a phone, browser, laptop or embedded device, a leaderboard is not enough. Measure prefill, decoding, memory footprint, quantisation effects, cold starts, battery use and tool-calling reliability on your actual runtime.

4️⃣ More pre-training can help even at small scale It seems intuitive that a 350M model would quickly saturate with data — there are even the Chinchilla scaling laws. Liquid AI presents another side: LFM2.5-350M received additional pre-training, going from 10T to 28T tokens, followed by large-scale reinforcement learning. Its official post says it outperforms models more than twice its size on several knowledge, instruction-following and tool-use benchmarks.

5️⃣ Doom looping is a distinct failure mode A particularly interesting section covers models repeating text instead of reaching an answer. This is especially troublesome for small reasoning models and agent workflows: the user expects an answer, tool call or action, while the model repeats its reasoning. Liquid AI describes an intervention that reduced doom looping from 15.74% to 0.36%. Evaluation therefore needs to cover behavioural failures, including loops and empty answers, as well as accuracy.

6️⃣ Small models are especially interesting in agent systems The future need not be one huge model answering everything. A more plausible architecture is a router, specialised models, tools, evaluations and a fallback to a large model. The large model can plan, reason and handle difficult cases; the small local model can execute plans quickly and cheaply.

#AI #LLM #SmallModels #EdgeAI #Architecture #Engineering #ML #Agents #DevEx #Management

Open video on YouTube