Review of Book "AI Engineering" 2 - Chapter 2. Understanding Foundation Models (AI column)
Out. podcast with the analysis of the cool book "AI Engineering", which gives an idea of the creation of gen AI applications. The book is disassembled by Alexander Polomodov, Technical Director of T-Bank, and Evgeny Sergeev, Engineering Director at Flo. In the second series, we discussed the second chapter of the book, which deals with foundational models. The chapter was complicated, but it seems that Zhenya and I coped and discussed the following topics in large strokes.
Introduction, chapter structure and model learning stages Data and languages: the impact of language representation in datasets Domain knowledge and the need for specialized models
- Specific models (example of MedPaLM)
- Multimodal models (text + images) Transition from RNN/Seq2Seq to Transformers Transformer structure and attention mechanism History of transformer development and their distribution Transformer parameters and components Context window and MLP blocks Resource constraints and learning optimization Post-training models: SFT and RLHF Training models through RFHF (examples of prompts and responses)
- Sampling and their strategies Hallucinations of models and their nature Cost of errors and model scenarios
The podcast release is available in Youtube, VK Video, Podster.fm, Ya Music.
#Architecture #Software #AI #Engineering #ML #Data #SystemDesign #DistributedSystems