Skip to content
#AI

Review of Book "AI Engineering" 3 - Chapter 3 & 4Evaluation Methodology and Evaluate AI Systems(AI column)

#AI #Architecture #Software #Engineering #ML #Data #SystemDesign #DistributedSystems

Out. podcast with the analysis of the cool book "AI Engineering", which gives an idea of the evaluation of both the foundation models themselves and applications based on them. The book is disassembled by Alexander Polomodov, Technical Director of T-Bank, and Evgeny Sergeev, Engineering Director at Flo. In fact, we discussed two chapters in this series: Chapter. 3Evaluation Methodology" and "Chapter 4: Evaluate AI Systems”. And if you put them down, they are presented below.

Introduction and subject matter Why evaluating AI applications is difficult; the growing importance of validation Validation in pipelines and the complexity of domains Benchmark restrictions and transition to product validation Risks of uncontrolled generation Information theory: entropy as a base of metrics Cross-entropy and KL-divergence for model evaluation Perplexy and the impact of context on model confidence Functional correctness vs non-functional requirements From lexical to semantic closeness; embeddings Validation Patterns and AI as a Judge Pair comparisons and model rankings; transitivity and voting System frame: criteria → model selection → Pipeline assembly Fact check and reference check; trusted sources; human baseline Pipeline design: independent tests, guidelines, markup; final conclusions

The podcast release is available in Youtube, VK Video, Podster.fm, Ya Music.

#Architecture #Software #AI #Engineering #ML #Data #SystemDesign #DistributedSystems