DataOps / MLOps for AI Business
How managers govern data, models, release flow and economics for production AI
How managers govern data, models, release flow and economics for production AI
How managers govern data, models, release flow and economics for production AI
Technical Director & Fellow, T-Technologies
Architecture and engineering R&D.
AI adoption in SDLC.
Focus: governed production AI.
Releasing models safely on top of a data platform
We skip the basics — No data-platform recap.
We look as managers — Decisions, owners, loop control.
We count economics — Data, features, inference, review.
Governed path from data to decision
Experiment → pilot → production boundary.
Standard release path and checks.
Platform vs domain ownership.
Outcome governance metrics.
Responsibility matters more than tool choice
Managed / self-service — Pipelines, registry, serving, observability; platform standards.
Domain ML teams — Quality, hypothesis, error cost and retrain stay with teams.
Hybrid — Shared standard path; dedicated runtime for critical scenarios.
What speeds up results
Hypothesis → live validation.
Repeatable release path.
Self-service without losing control.
What contains risk
Reproducible data, features, evals.
Stop rules, rollback, blast radius.
Unit cost and path reversibility.
Which data and features are safe for models
Data platform
Data delivery, catalogs, lineage, access.
Reliable path to data products.
DataOps/MLOps
Training/inference/release data fitness.
Features, evals, release, runtime, feedback.
Each handoff should leave an artifact: data, features, evaluation, release decision and rollback plan
A feature becomes a product contract: schema, owner, freshness, logic version and production readiness
Unreproducible training slice.
Features without owner/version/SLA.
Offline metrics hide production skew.
Retrain hides data/policy debt.
Model release must be governed
Different release modes answer different questions: what changes, where risk sits and how fast rollback can happen
Decision participants
Product: error cost, outcome.
ML: quality, calibration, risk.
Platform: release path, rollback, SLO.
What breaks without a matrix
Quality separate from risk.
Release without consequence owner.
Rollback after degradation.
AUC misses cost, latency, review.
Threshold/policy/fallback change outcome.
Release card records what changed.
Decision combines quality, risk, cost.
A production model costs money every minute: in features, inference, review and operational support
Choose the inference mode by latency, error cost, capacity cost and freshness requirements
When managed is enough
Fast launch, standard load profile.
Baseline SLO and monitoring are enough.
Error cost < platform complexity.
When to own
Custom routing, fallback, strict SLA.
Predictable volume; unit cost matters.
Special access/audit/environment policies.
Features: backfill, materialization, freshness, storage.
Training/evals: experiments, retrain, baseline, replay.
Serving: CPU/GPU, peak, batch/online, cache.
Review: labeling, queues, escalation, corrections.
How to allocate cost
Scenario: antifraud, recommendations, support.
Model/endpoint: expensive routes.
Domain/product: decision owner.
How to use it in governance
Showback reveals behavior first.
Chargeback after metrics and rules.
Optimize architecture, not quarter budget.
After release you need a system that sees degradation, collects feedback and scales without heroics
Human-in-the-loop should be a controlled path for learning, review and escalation, not a permanent workaround
DataOps/MLOps maturity shows up in repeatable release, risk ownership, observability and unit economics
Managed service: common path, fast launch.
Platform: release, observability, SLO, FinOps.
Domains: hypothesis, error cost, feedback.
Own runtime: control, scale, economics.
A production model is a business loop.
Google Cloud — MLOps pipelines.
Sculley et al. — Hidden Technical Debt.
Feast Feature Store documentation: docs.feast.dev
NIST AI Risk Management Framework 1.0: nist.gov/itl/ai-risk-management-framework