Skip to content
all episodes
Research Insights Made Simple · episode 28

The Data Platform in 2026: From DWH to Lakehouse and AI Agents

1:37:08

Episode participants

Conversation

What we discussed on the recording

Alexander Polomodov, Nikolay Golov, and Alexander Filatov examine data-platform workloads. OLTP serves short operations; analytics scans large sets. In a Postgres TPC-H run, two of twenty queries did not finish within an hour on tables up to 100 million rows. That growth needs a specialized engine, not just a larger server.

MPP couples storage and compute: a slow node delays queries, concurrency harms isolation, and expansion requires repartitioning; cost rises beyond 20–30 nodes. A lakehouse separates layers, but beyond S3 and Iceberg it needs a catalog and access control; workloads add retries, compaction, and orphan files.

Tengri uses isolated compute per session, a query-aware cache, and on-demand parallelism. Its Vertica migration began in 2022, yet four years later procedures remain; one presale involved two million lines. The platform creates value when tied to forecasting, pricing, recommendations, or another measurable outcome.

Analytics agents need semantic context and checks for bad filters and multiplying joins. They generate five times more queries, including Cartesian products and destructive SQL. Permissions block danger, isolation protects neighbors, and the harness verifies effects; a human remains accountable.

Data platformsPlatform engineeringDatabasesAI in SDLC