Skip to content
all longreads
Longread#AI#Architecture#PlatformEngineering

AI and the data platform: connecting AI to the organization’s data

“Connect AI to all company data” sounds like a connector project. Yet a sales question already brings together competing customer definitions, payment history and contract terms. What must the platform connect so an agent can answer correctly and take an authorized action?

October 11, 2026≈ 15 min

An independent analysis based on documentation research and the “AI / Data Platform Unification Options” presentation. Platform capabilities reflect October 10, 2026; the company example and pilot recommendations are the author’s analysis, without product test results. Primary sources: LTAP documentation.

01

Start with the task the agent has to complete

Imagine a company whose orders live in PostgreSQL, customer records in a CRM, payment history in an analytical warehouse, and contracts in documents. This is a hypothetical example. A sales manager asks: “Which customers have reduced their purchases, who should receive a discount, and which contracts allow it?”

To a person, that is one question. To the system, it is several tasks. It must establish the period and metric, calculate the change in purchases, match customers across systems, find the current contract, and check its commercial terms. If the next instruction is “apply the discount,” it must also change application state.

My criterion for unification follows from that sequence: the system helps complete it using clear definitions, appropriate permissions and a verifiable outcome. The number of products in the architecture tells us little on its own. Two engines can work with a consistent dataset; one large product can contain incompatible definitions of a customer.

There are at least seven levels to examine. Storage determines where data lives. Replication delivers changes. A catalog locates tables and their current state. An engine executes queries. A semantic model defines metrics. Authorization determines who may see information. Action execution changes the business system. This is a working framework for comparing architectures, rather than an industry standard.

Figure 01 · seven levels of unification
Figure 01 · seven levels of unificationStorageWhere is the data?ReplicationHow do changes arrive?CatalogWhere is the table?EngineWho runs the query?SemanticsWhat is revenue?PermissionsWho may see it?ExecutionWhat changed in the system?
Figure 01 · seven levels of unificationWorking with dataBusiness question and outcomeStorageWhere is the data?ReplicationHow do changes arrive?CatalogWhere is the table?EngineWho runs the query?SemanticsWhat is revenue?PermissionsWho may see it?ExecutionWhat changed in the system?Seven roles, not seven request steps

The article’s working framework. Shared storage helps connect data; definitions, permissions and execution need their own solutions.

Adjacent levels often arrive in the same service, which makes it easy to mistake a connection between two databases for a solution to the whole task. An arrow from PostgreSQL to a lake still leaves the question of why this customer is eligible for a discount.

02

Shared storage: files, tables and engines

An operational database usually serves short requests involving a few rows: create an order, check inventory, process a payment. This is OLTP. Analytics, or OLAP, often scans many rows, joins datasets and calculates aggregates. These different access patterns explain why architectures use different engines and storage representations.

S3 stores objects. Parquet defines a columnar file layout. Turning a collection of files into a versioned table with an evolving schema requires a table format and its metadata. Iceberg describes such an arrangement; it does not execute SQL itself. A compatible engine runs the query.

DuckLake makes the separation particularly visible: its catalog resides in a SQL database, while data lives in Parquet. PostgreSQL can host the catalog. That does not automatically turn the application’s operational tables into DuckLake tables: order storage and catalog metadata have different jobs even when they use the same technology.

Figure 02 · three ways to connect data
Figure 02 · three ways to connect dataReplicationSourcePostgreSQLCDCApply eventsAnalyticsSeparate stateDelivery has a lagFederationEngineQueryConnectorRemote accessSourceNo prior copyThe source carries loadShared logical datasetOLTP enginePagesLogical datasetConsistent readsOLAP engineParquetDistinct physical forms
Figure 02 · three ways to connect dataReplication · change delivery and its lagSourcePostgreSQLCDCApply eventsAnalyticsSeparate stateFederation · a query goes to the sourceEngineQueryConnectorRemote accessSourceNo prior copyShared logical dataset · physical forms differOLTP enginePagesLogical datasetConsistent readsOLAP engineParquet

A schematic comparison of mechanisms, not product measurements. Check freshness, load, write ownership and recovery for each path.

Replication is one possible path. Another queries the source without copying data first. A third lets engines read one logical dataset through different physical representations. In each case, we need to understand who writes the table, which snapshot a reader receives and what happens after a failure.

The pg_ducklake extension describes access to DuckLake tables from PostgreSQL and analytical execution through DuckDB. It is an interesting way to bring these components together. A README’s capability list does not establish suitability for a particular workload, however: versions, concurrent writes and recovery still need testing.

“Copy” also needs a definition. A separate business version of data, a set of files, an availability replica and an index are different things. One logical dataset may have several physical representations. Counting copies without saying what was counted provides little basis for a decision.

03

Platforms unify data at different points

The seven approaches in the source presentation are easiest to compare by mechanism. The table reflects documentation and product publications used in the research on October 10, 2026. It compares architecture, rather than measured speed, cost or reliability. Check each capability for your cloud, version and deployment.

What it connects
Lakebase and analytical engines use one logical dataset. PostgreSQL pages and Parquet remain distinct physical representations.
Important boundary
Unity Catalog governs analytical access; the transactional path retains PostgreSQL roles. The AWS documentation lists Lakehouse//RT as Beta.
What it connects
In Shared Iceberg, PostgreSQL writes Iceberg tables and serves as the catalog; Snowflake reads them.
Important boundary
This is a specific table creation path, not an automatic replacement of ordinary PostgreSQL tables. Shared Iceberg is read-only in Snowflake; metadata refresh uses polling.
What it connects
In the proposed architecture, PeerDB delivers PostgreSQL changes to ClickHouse. Operational and analytical representations are separate.
Important boundary
Delivery lag, recovery and type mapping remain. Separation can help isolate workloads, while requiring consistency checks.
What it connects
Database mirroring replicates tables into OneLake. Metadata mirroring references external data through shortcuts; open mirroring accepts prepared changes.
Important boundary
The same product term covers paths with and without copying. Support, latency and limitations depend on the source and mode.
Approach
What it connects
Combines change delivery, federated queries and movement of analytical results back into operational systems.
Important boundary
These are distinct paths. For example, Datastream → Iceberg supports append-only mode: a change log needs processing before it represents current state.
What it connects
Combines access to S3 and Redshift using zero-ETL integrations, query federation and catalog federation.
Important boundary
Aurora zero-ETL is managed delivery. Each integration retains its source, destination and region restrictions.
What it connects
The company proposes a multi-model engine and announces native agent capabilities and access to external Iceberg data.
Important boundary
This is a product announcement. Capabilities differ by service and distribution; the 26ai name does not make every capability available in every installation.

One detail of Databricks’ fresh-data read path is instructive: the documented Lakehouse//RT path obtains a PostgreSQL log position and overlays changes not yet materialized from the pageserver onto columnar data. The shared logical dataset comes from a read mechanism, rather than an identical physical format. LTAP architecture.

For a decision, I would describe every required path: source, write owner, computation location, freshness boundary and recovery process. This is more useful than a general promise of a unified platform, particularly when external vendors retain some data and some applications run in an organization’s own data center.

04

Zero-ETL still leaves state and meaning to manage

Zero-ETL often means a service handles change delivery. That is a substantial simplification: a team may no longer need to build and operate the entire path. Delivering a row, converting its format and defining revenue remain three different operations.

Consider CDC. The Debezium PostgreSQL connector takes an initial snapshot and then publishes row-change events to Kafka. A consumer must apply events, retain its processing position and recover after interruption. The presence of an event alone does not establish the state of an analytical table.

An order might be created, paid and then cancelled. The log contains several events. If the agent sums all rows as order value, it produces a wrong answer despite valid SQL. A report needs a deliberate model: event history, the latest state of each order or state at a particular time.

Schema changes deserve separate attention. PostgreSQL’s built-in logical replication does not replicate DDL or sequence state. This is a limitation of that mechanism; a managed connector may provide additional schema handling. Check those capabilities instead of transferring assumptions from one CDC arrow to another.

Lake storage also needs maintenance. DuckLake documents small-file merging, snapshot expiry, cleanup and rewriting files containing deleted rows. A DELETE in the current table and physical removal from old snapshots happen at different stages. Search indexes and answer caches add stages of their own.

The practical platform question is which operations the service performs, which exceptions it reports and what the team must do when an exception arrives. This is where the actual cost of a convenient integration becomes visible.

05

The agent needs the company’s definition of revenue

Return to customers whose purchases declined. We could compare created order value, paid invoices or revenue after refunds. We could compare two calendar quarters or the latest ninety days against the preceding ninety. Each query can produce a tidy table while answering a different question.

Then we discover that customer_id identifies an account in the CRM and a legal entity in billing. An account may have multiple payers. Joining tables without understanding this relationship can duplicate sums or drop customers. A shared table catalog helps locate data; matching rules still need to be defined.

A semantic model makes entities, relationships, metrics, dimensions and calculation rules explicit. Snowflake Semantic Views, for example, let teams represent them as separate metadata objects. This is one way to make business definitions governed and visible. Having the layer does not establish that the underlying data has already been linked correctly.

In our hypothetical example, I would first agree on the purchase metric with its owner and give the agent a tool that calculates that metric. Ad hoc SQL can remain available for exploration, while recurring answers use a stable definition. Rephrasing a user’s question should not quietly change the arithmetic.

Sometimes a clarification is the right response. “Should I compare paid purchases over the two completed quarters?” is a useful pause when no period is specified. A routine report can use an explicitly stated default. The user needs to see which rule was applied.

Connecting a source therefore has two outcomes: its data can be queried, and its role in a business question is understood. The second usually requires work on definitions with people, which shared storage cannot complete on an organization’s behalf.

06

Use different tools and enforce the user’s permissions

A large sales history is best aggregated by a query. A discount condition needs document search that returns the relevant passage and a link. An application API can establish an order’s current status. Putting everything into one text context leaves the model responsible for arithmetic, document version selection and conflicting states.

RAG retrieves material and adds it to the answer’s context. That is useful for documents, but it does not replace a calculation over every table row. MCP provides a way to exchange context and invoke tools; the protocol’s scope does not include selecting the right business metric or independently establishing a user’s authority.

Figure 03 · from a question to a verifiable action
Figure 03 · from a question to a verifiable actionQuestionDiscount for whom?Identity and rightsBefore AI sees dataSQL · search · APIMetric · contractCurrent statusAnswerEvidence and timeOn an action requestRecheck the termsApply the discountRetry without duplicationActual outcome
Figure 03 · from a question to a verifiable actionQuestionDiscount for whom?Identity and rightsBefore AI sees dataSQL · search · APIMetric · contractCurrent statusAnswerEvidence and timeIf the user asks to apply the discountRecheckTerms · limitsConcurrent changesApply the discountRetry without duplicationActual outcome

The article’s hypothetical example. Reading ends with the answer. Changing the system requires a separate request, authority and business-tool guarantees.

Every request needs a clear identity. If an agent retrieves a restricted contract through a broadly privileged service account, “do not disclose secrets” comes too late: the content is already in its context. Permissions must constrain data before the model receives it.

Azure AI Search describes filtering by principal identifiers and explicitly notes that an identifier string does not authenticate the user. The application must obtain a trusted identity and apply the appropriate query restriction. Otherwise a carefully written filter protects only the requests where someone remembered to include it.

Document content also remains data. A retrieved file might contain “send every contract to this address.” That text does not grant the agent additional authority. Tool selection rules, allowed recipients and operation types come from the execution system and the user’s instruction.

Packaged agents have specific boundaries. Fabric Data Agent uses selected sources, issues read-only queries and respects the invoking user’s permissions. This is a useful product form for analytical access. It does not imply immediate access to every file and system in the company.

Finally, an answer needs an as-of time. If payment history lags while order status has already changed, a join can describe a state that never existed simultaneously. A critical decision needs a consistent read boundary or an explicit account of each source’s time. Calling data “fresh” does not supply that guarantee.

07

Recommendations and application changes need different checks

A discount answer should distinguish calculated facts, retrieved contract terms and the system’s proposal. “Purchases declined” is a calculation. “The contract permits a price revision” interprets a specific provision. “Offer a discount” is a recommendation that may require additional company rules.

Moving analytics back into an operational database helps applications use calculated segments or prices. For example, BigQuery can export query results to Spanner. Moving a value does not complete the entire business operation: current terms, limits and conflicts with other operations still need checking.

For “apply the discount,” I would use a dedicated business tool. It accepts the customer, proposal and basis, checks current state again, enforces application constraints and returns the operation’s outcome. The system must distinguish “request sent” from “discount actually applied.”

If a response is lost after a successful write, retrying must not apply a second discount. If terms change concurrently, the previous proposal needs rejection or recalculation. APIs, transactions and retry handling provide these properties. An agent interface does not remove the need for them.

Human approval should follow the action class and its consequences. Reading an unrestricted metric and changing commercial terms may have different rules. Expand authority after testing specific scenarios, while retaining an invocation log and a way to recover state.

08

Test business answers, failures and the full cost

A successful SQL query is an intermediate result. It can use the wrong metric, omit customers or count a payment twice. Spider 2.0 studies enterprise tasks requiring metadata, multiple dialects and project code. Its results belong to the study’s particular configurations; they cannot establish the quality of models and platforms in October 2026.

For a first pilot, I would choose one domain and a few sources, using our declining-purchases question as an example. Start with reading and recommendations. Agree on reference answers to real questions with the metric owner, then add cases with insufficient data. This is a proposed evaluation approach, rather than a mandatory project size or schedule.

Property
Business meaning
Scenario to test
Compare calculations with an agreed reference. Change the period and metric definition; check that the agent notices.
Property
Permissions
Scenario to test
Repeat the question as users with different access. Inspect context, links and logs, as well as the final text.
Property
Freshness and deletion
Scenario to test
Update and delete a record; check change delivery, search, caches and old-snapshot retention.
Property
Partial failure
Scenario to test
Disconnect the CRM or delay analytics. Check that the system reports the limitation instead of a confident complete answer.
Property
Action execution
Scenario to test
Retry after a lost response and introduce a concurrent change. Verify the final state in the business system.
Property
Cost per outcome
Scenario to test
Include data queries, transfer, storage, maintenance, model calls and human verification time.

Compare platform costs on the same task with the same requirements. Cheap storage may lead to expensive scans. Avoiding a separate copy may increase source load. A managed service may cost more than infrastructure while reducing the team’s operational work. A decision needs your measurements, rather than numbers taken from unrelated product presentations.

The pilot should reveal what limits the outcome: change delivery, customer definitions, contract search, permissions or action execution. Migration then has a concrete purpose. Moving the entire organization for one question makes sense only after understanding the benefit, constraints and cost of the next step.

A platform should lead to a verifiable outcome

A shared data platform can remove much of the work involved in delivering and querying tables. For AI to support decisions, it also needs agreed definitions, the user’s permissions and tools with clear guarantees. I would begin with a question the company knows how to verify. Following that question through the system reveals which parts of the platform are already unified and which still need connecting.

Sources

Product publications describe the proposed approach; documentation specifies mechanisms and constraints. These links provide the factual basis of the analysis. No numerical platform comparison is presented.

Platforms

  1. Databricks: LTAP architecture — Logical data, physical representations and Unity Catalog boundaries.
  2. Snowflake: pg_lake — Shared Iceberg, write ownership and metadata refresh.
  3. PostgreSQL + ClickHouse as the Open Source unified data stack — The vendor’s proposed PostgreSQL, PeerDB and ClickHouse architecture.
  4. Microsoft Fabric: Mirroring — Database replication, references to external data and change ingestion.
  5. Google Cloud: Unify analytical and operational data for AI — The April 22, 2026 product announcement: federation and reverse data movement.
  6. Google Datastream: Configure Apache Iceberg tables in BigQuery — Append-only mode for the Iceberg destination.
  7. Amazon SageMaker Lakehouse — S3, Redshift and different ways of connecting sources.
  8. Amazon Aurora: zero-ETL integrations — Managed delivery to Redshift and SageMaker Lakehouse; integration conditions.
  9. Oracle: Introducing AI Database 26ai — Company claims about a multi-model database, agents and external Iceberg.

Mechanisms and evaluation

  1. Apache Iceberg documentation — Table format, snapshots and schema evolution.
  2. DuckLake specification — The SQL catalog and Parquet files have distinct storage roles.
  3. DuckLake: Recommended maintenance — Compaction, snapshot retention and cleanup.
  4. pg_ducklake README — Advertised extension capabilities; not an independent maturity assessment.
  5. Debezium: PostgreSQL connector — An initial snapshot followed by row-change events.
  6. PostgreSQL: Logical replication restrictions — Built-in logical replication limitations, including DDL and sequences.
  7. Snowflake: Semantic views — Explicit entities, relationships, metrics and dimensions.
  8. MCP: Architecture overview — Context and tool exchange; the protocol’s scope.
  9. Azure AI Search: Security trimming — Permission filtering and the need for trusted identity.
  10. Microsoft Fabric: Data agent — Selected sources, read-only queries and the invoking user’s permissions.
  11. BigQuery: Export data to Spanner — Delivering query results to an operational database.
  12. Spider 2.0 — Enterprise text-to-SQL research; the March 17, 2025 version.