Money appears when a decision changes
In conversations about monetisation, it is easy to confuse a resource, an activity, and an outcome. A billion events are a resource. A thousand reports are an activity. Lower fraud losses or additional margin from a new search experience are outcomes. Between them sits a decision: what did the company start doing differently, and why was that more valuable than the old behaviour?
I count four kinds of value from data. First, incremental revenue, which must then be converted into margin. Second, costs that no longer have to be incurred. Third, a reduction in expected losses. Fourth, the value of deciding sooner: a product reaches the market before the season, cash returns to circulation faster, or a dangerous experiment is stopped earlier. Any of the last three can outweigh a standalone data-selling business.
Profit = incremental margin + avoided costs + prevented losses + a distinct acceleration effect − total cost − expected harm.
This is a reasoning model, not an accounting identity. Its terms must not overlap: the value of hours released and the entire outcome created during those same hours cannot both be booked as savings. If faster delivery is already reflected in incremental margin, it cannot be added again. If ₽500 of contribution margin per order already includes delivery, the same delivery cost cannot be subtracted a second time.
Hours saved first create spare productive capacity. They become money only when the need for an additional hire disappears, overtime falls, or that capacity produces a measurable useful outcome. Similarly, a rejected fraudulent transaction creates an avoided loss only after accounting for the probability of loss, any recovery, and the cost of wrongly rejecting an honest customer.
Some risks cannot be made acceptable by averaging them into an expected-loss figure. A deal with positive expected value may still be unacceptable if a rare outcome could threaten the core business. A legal prohibition does not become permission because a provision was added to the financial model.01A favourite presentation trap: revenue follows the optimistic scenario, costs stop at the cloud bill, and user trust is declared free.
One dataset can support three different businesses
Imagine a marketplace with searches, clicks, orders, prices, and returns. The same dataset can be licensed to an external buyer, used to investigate the company’s own product, or fed into an algorithm that changes the customer experience every day. The recipient of value, the way the effect is proved, and the team’s obligations are different in each case.
| Route | What changes | Where to look for money |
|---|---|---|
| Licensing | A marketplace seller buys a category index, forecast, or API and changes purchasing | Payment for information; the cost of delivering it repeatedly |
| Internal decisions | A team investigates the funnel, tests a hypothesis, and chooses a change | Incremental margin, avoided costs, and failed launches stopped in time |
| Product feature | Search ranks items, fraud prevention decides, or pricing adapts | Retention, conversion, losses, and the value of a subscription or fee |
The boundary follows the specific operation. A demand report for which a marketplace seller pays separately belongs to the first route. If the service itself uses that report to calculate a supplier order and stands behind the recommendation, it has become an information product with a broader promise. An experiment on ranking is an internal decision process; the selected ranking running in the app is a product feature. A single initiative can move through several routes at once.
There is therefore no universal ladder from “lower risk” to “higher margin”. Licensing can have clear rights and steady demand. Internal analytics can run for years without changing how the company works. A product may spend more on predictions, support, and compensation than the feature creates in customer value. Technical sophistication alone says nothing about the commercial quality of an idea.
Possessing a copy does not give you every right over it
The server is inside the company, the storage bill has been paid, and an employee has access. All of that describes factual control over a copy. The permitted uses remain an open question. A single dataset may contain copyrighted text, personal data, a partner’s contractual restrictions, trade secrets, and rights in the selection or arrangement of the database.
A useful practice is to replace the word “ours” with a register of permissions: where did the material come from, who authorised its collection, may it be passed on, may derivative outputs be created, may a model be trained on it, may it power search, and may it be kept after the contract ends? A licence to display a book to a reader does not necessarily authorise every machine use of its contents. Rights in a database do not cancel the rights of the people and authors within it. The EU Database Directive distinguishes protection of a database from rights in its contents.
In the EU, processing personal data requires a lawful basis under Article 6 of the GDPR. Consent is one basis; others include necessity for performing a contract with the data subject, or legitimate interests after assessing necessity and balancing those interests against the individual’s interests and rights. A new data-trading business cannot automatically be declared necessary for an old service merely because it was mentioned in the terms of use. The European Commission’s guidance explains lawful bases separately from the commercial form of a deal.
A separate question is whether the new purpose is compatible with the purpose for which the data were collected. Under Articles 5(1)(b) and 6(4) of the GDPR, the assessment considers the link between the purposes, the context of the relationship with the person, the nature of the data, possible consequences, and safeguards. Where further processing rests on new consent or a basis laid down in law, the corresponding regime applies. Turning a service history into a product is therefore a separate design decision, not a silent extension of old analytics. The purpose-limitation principle is set out in the GDPR itself.
What informed consent looks like
When a company relies specifically on consent, a person must understand who will use the data and for what purpose. Under the GDPR, consent must be freely given, specific, informed, and unambiguous; the company must be able to demonstrate it, and withdrawal must be as easy as giving consent. Withdrawal does not retrospectively invalidate processing that was lawful before it. These are regulatory conditions, not a question of font size.
The engineering consequence is that purpose and consent version must be connected to the record and the operation that uses it. A checkbox in a separate database is little help if an export comes from an old snapshot while permissions are checked against the account’s current state. The company needs to reconstruct why a particular record was included in a particular delivery.
A copyable asset is unusual because the seller does not physically part with it. Multiple licences may be issued, while the scarcity of the seller’s own advantage declines. Exclusivity, term, territory, permitted tasks, and the ability to pass data to third parties become part of the product itself. The cost of the file on disk says almost nothing about the value of those rights.
Accountability continues after the transfer
The most dangerous sentence in such a deal sounds harmless: “We delivered the data; what the buyer does next is their problem.” Start by identifying the parties’ actual roles. Under the GDPR, a controller determines the purposes and essential means of processing; a processor acts on the controller’s behalf. A buyer that chooses its own purpose may be a separate controller. Calling it a subprocessor in the contract is not enough. The EDPB guidelines determine roles from the parties’ actual conduct.
When a processor engages another processor, Article 28 imposes rules on the controller’s authorisation, equivalent obligations, and the original processor’s responsibility to the controller for the other processor’s performance. A buyer acting as an independent controller requires a different arrangement. Roles shape the contract, audit rights, assistance with people’s requests, and incident notification. At the end of the services, Article 28(3)(g) requires the processor, at the controller’s choice, to return or delete the personal data and delete existing copies unless the law requires their retention. See GDPR Articles 4 and 28.
“What if we anonymise it?”
Removing a name does not guarantee anonymity. A persistent identifier, precise timestamps, and a rare sequence of actions may allow a record to be linked to a person through an external dataset. Pseudonymisation reduces some risks, but the data may remain personal. Anonymity has to be assessed in light of the means reasonably available for identification. That is how the EDPB distinguishes the concepts.
Aggregates also need testing. In a small group, one person may account for nearly the whole measure; a series of similar queries can expose a hidden contribution. Practical safeguards depend on the threat: larger groups, coarser timestamps, query restrictions, and the removal of rare combinations. “We only provide aggregates” describes the form of an answer, not the result of a privacy assessment.
Deletion is a system-wide operation
The GDPR right to erasure has grounds and exceptions; it is not an unconditional command to erase every record in every circumstance. When the duty does apply, deleting one primary table is not enough: the company must know the recipients, copies, indexes, and processing outputs that themselves remain personal data. Not every derived metric is automatically subject to erasure. Article 19 regulates notification of recipients, subject to impossibility or disproportionate effort. The European Commission’s explanation connects a request to onward disclosure.
For a model, the question is harder than for a row in an index. The presence of a training record does not automatically require every model to be destroyed, but calling an artefact “weights” does not guarantee anonymity either. The EDPB calls for a case-specific assessment of whether a model is anonymous; individual regulatory orders may also reach derivative products. The useful boundary is in EDPB Opinion 28/2024.
No payment does not mean no processing
Under the GDPR, disclosure remains processing whether or not an invoice is issued. Under the CCPA, “sale” includes monetary or other valuable consideration, while disclosure for cross-context behavioural advertising is defined separately and may occur without payment. Exceptions and statutory scope still apply; they are set out in the CCPA definitions. The CPPA guidance explains sale, sharing, and opt-out rights. For a management decision, the point is straightforward: barter, a “partner exchange”, and free access require the same careful map of purposes and roles.
The buyer pays for a useful signal and reliable delivery
“Selling data” may mean a one-off export, a recurring dataset, an industry index, an API, a content licence, or a score without the raw records. These models incur different costs after the first sale. An export has to be documented and transferred; an API has to stay available; a prediction has to be checked as conditions change. Profit cannot be estimated from the cost of the last database query.
| Form | What the buyer pays for | What must be maintained |
|---|---|---|
| Export | Coverage and rare observations | Documentation, rights to the copy, and an agreed period |
| Index or report | A market comparison | Methodology, a stable sample, and historical restatement |
| Continuously updated API | Freshness and easy integration | Schema versions, availability, limits, corrections, and deletions |
| Corpus licence | Content and the agreed uses | Provenance of rights, permitted derivatives, and term limits |
| Prediction or score | A better decision | Calibration, cost of error, monitoring, and support |
Scale is an advantage only when paired with relevance. Five years of clicks are useless to a buyer who needs to know whether an item is in stock today. Millions of observations from one audience need not represent the whole market. A change in the mix of sellers can move a price index even when prices themselves have not changed. Methodology and representativeness can therefore cost more than collection itself.
The full economics include preparing and correcting data, negotiations, integration, infrastructure, security, legal work, customer support, and terminating access. The more each buyer asks for a “slightly different” schema, the closer the business comes to custom analytics. That may be a profitable service, but its margin scales differently from a single product with a repeatable contract.
A large supplier faces another cost: the competitive advantage it has handed away. Selling a signal may help partners and improve the market—or it may help them bypass the platform or reproduce its function. The entire business, including that possibility, belongs in the calculation, not just the new unit’s revenue.
Reddit: a scarce corpus can produce meaningful revenue
Reddit combines topic-specific communities, living language, discussion, and constant updates. For an AI partner, the archive is not the only attraction: new questions, answers, and corrections help keep a product current. In February 2024, Reddit announced an expanded partnership with Google and access to structured content through its Data API. The primary announcement describes the access mechanism itself.
In its IPO documents, Reddit reported data-licensing arrangements entered into in January 2024 with an aggregate contract value of $203 million and terms of two to three years. At the time, it expected to recognise at least $66.4 million in 2024. That is the combined value of several contracts and a then-current revenue-recognition forecast—not realised profit and not the price of a single Google agreement. The IPO prospectus preserves that distinction.
The case shows that direct monetisation can be material. It does not show that every corporate event log has comparable demand. Corpus scarcity, freshness, and a limited number of substitutes create bargaining power; cleaning and repeated delivery turn that advantage into obligations to the buyer.
In its 2025 annual report, Reddit said that substantially all of the contract value associated with licensing revenue came from two partners. The entire “other revenue” category was about $140 million, or 6.4% of total revenue. Because that category also includes other products, it cannot be treated as licensing revenue alone. The company also discloses risks from non-renewal, API reliability, and unlicensed access. Reddit’s annual report makes concentration a concrete limit of the model, not an abstract concern.
Finally, the fact that a discussion is public does not erase its participants’ expectations. People come to talk within a community; commercial use of their contributions changes their relationship with the platform. Even a legally supportable deal needs a clear answer to three questions: what do participants receive, how are deletions honoured, and does the deal damage the environment in which valuable new content is created? This is my conclusion about the model’s durability, not a measured price for any particular dispute between Reddit and its users.
In its published policy, Reddit requires partners to honour deletions and prohibits the use of licensed content to identify people or target advertising. These are stated restrictions, not an independent audit of their enforcement. The Public Content Policy shows that licensed access comes with downstream-use rules even when the original discussions are public.
Avast/Jumpshot: side revenue can put the core product at risk
A user installs security software expecting it to protect them from tracking. In its complaint, the FTC alleged that Avast collected detailed browsing information through browser extensions and antivirus software, and that Jumpshot sold it to more than one hundred third-party customers. The regulator challenged both the disclosure of the practice and claims that the data had been adequately anonymised. In agreeing to settle, Avast neither admitted nor denied the allegations except for jurisdictional facts. The FTC complaint describes the collection of browsing data from 2014 onwards and its subsequent sale.
A detailed clickstream is valuable for studying behaviour: where a person arrived from, what they searched for, which product they viewed, and how the visit ended. The same detail makes linkage to another source easier. Removing a name from a row is not enough when a buyer already knows the exact time of a particular event and can attach a long browsing history to it.
The chronology matters. Avast announced that it would wind down Jumpshot in January 2020. The final FTC order arrived in June 2024: it barred Avast from selling or licensing browsing data for advertising purposes and required a $16.5 million payment. “The regulator shut down Avast” is therefore wrong. Avast decided to close the Jumpshot business; the later prohibition had a defined scope. See Avast’s wind-down announcement and the FTC’s final-order announcement.
The order also covered browsing data transferred to Jumpshot and the models, algorithms, and software Jumpshot developed from it. Avast had to delete them, instruct recipients to delete their copies and derivatives, and notify affected users. The text of the order is not a direct order to every buyer. These are the terms of this case; they cannot be projected automatically onto every processing activity or every AI model. The economic lesson remains substantial: risk can materialise years later and reach not only a source dataset but also what was built from it.
The useful question here is asymmetry, not whether the revenue was “pennies”. How much incremental profit must the unit earn to justify a threat to trust in the core product? Who inside the company accounts for that harm? If the unit receives credit for all of the revenue while the whole brand bears the reputational loss, its local financial model systematically rewards decisions that are too risky.
Cambridge Analytica: control matters even without a direct sale
In that case, an app collected information about its users and their friends through the Facebook API available at the time; the data was then transferred for political profiling. The FTC describes deceptive collection and use. This was a case about purpose and access control, not an ordinary purchase of a database from Facebook or a breach of its servers. The FTC’s analysis shows why a partner contract does not replace oversight of downstream use.
O’Reilly: rights and provenance become product properties
In O’Reilly Answers, a user asks a question of a professional corpus and receives an answer with sources. In its account of the 2024 version, the company explains retrieval and generation from library material, tracking the contribution of sources, and paying royalties to authors. That is importantly different from assuming the entire corpus was simply sold to train a general-purpose model. The creators’ description of Answers explains the architecture.
O’Reilly also says that authors generally retain copyright while the publisher operates under licensed rights. Its account of those rights makes this a useful extension of the discussion about possessing a copy. Content, permission to use it, and customer access are separate elements of one design.
Value for the reader appears when the path from a working question to a verifiable answer becomes shorter. Provenance lets the reader check a claim, return to its context, and recognise an author’s contribution. In my classification, this sits at the intersection of licensing and a product feature: the quality corpus provides the foundation, while convenient access to knowledge creates the customer proposition.
None of this establishes a specific profit figure for Answers: the sources cited here do not contain a public causal estimate of its financial effect. They do show a constructive alternative to selling a raw dataset. The company keeps the connection to the source, controls access, and improves the task for which the customer already uses the product.
Statist and Hippo: an internal product must change how people work
Statist and Hippo are T-Bank’s internal products: product analytics and an experimentation platform. Their value is easiest to examine as two tasks: first notice and investigate a problem, then test the effect of a change. This is an explanatory model of the process, not a claim that the two systems have a particular technical integration.
In its public description of Statist, T-Bank explains that it built an internal system because of the constraints, cost, and data requirements of external tools. An event catalogue, schemas, typed SDKs, and data validation help teams agree on what their measurements mean. Usage figures on the page illustrate the platform’s scale, not its financial return. The Statist team’s account is useful precisely as a description of the organisational and engineering problem it solved.
Suppose a team sees fewer completed purchases. It must check whether checkout starts are counted consistently across devices, whether an event disappeared after an update, and whether the audience changed. Only then does it make sense to look for a product cause. A catalogue, data quality, and common definitions save time in every such investigation—if teams actually use them. The financial effect comes later, when the inquiry leads to a useful action.
The second task is running an experiment. A team needs valid comparison groups, agreed measures, a sufficient sample, and a decision rule written in advance. Hippo represents that role in this article. The T-Bank article confirms that the bank has its own experimentation platform, Hippo. The source establishes the product’s role, not its financial effect or detailed architecture.
An internal platform has users, switching costs, and competitors of its own: a familiar spreadsheet, hand-written SQL, or a colleague’s opinion. If the path to an answer is too difficult, the team will work around the system. Measure not only active accounts, therefore, but time from question to decision, repeated use of agreed metrics, and the share of decisions whose outcome is actually known.02A platform can be popular and still lack a demonstrated ROI. Popularity confirms a need; financial impact requires the next step.Code of Leadership · a conversation about Statist
How to evaluate the platform itself
I would build a portfolio of decisions, not add up every “winning” experiment. For each decision, record the observation, the accountable leader, the action taken, the comparison with no change, and the financial outcome. Then ask a second question: how much of that outcome was created specifically by the platform? The team might have run the same test manually, later, or with another tool. The full effect of a feature cannot be credited to the product team, analytics, and the experimentation system as three independent gains.
A practical yardstick is the conservative share of decision gains, prevented losses, and genuinely used released capacity that can be attributed to the platform, less the full cost of development, operation, migration, and training. This is the author’s evaluation model, not a published calculation by T-Bank. When the platform’s distinct contribution cannot be isolated reliably, it is more honest to show a range and the cost of the alternative process. Event and experiment counts explain the workload; they do not demonstrate payback.
A counterfactual turns a beautiful chart into testable economics
Sales rose after the new search experience launched. But the holidays also ended, a paid acquisition campaign started, and a competitor raised its price. A before-and-after comparison does not tell us how much of the growth search caused. We need a counterfactual: an estimate of what would have happened without the change. A randomised experiment often gives the most direct estimate, provided its design matches the real structure of the product.
Before the test, choose a primary metric, the size of effect worth detecting, the unit of randomisation, and guardrail metrics. A guide to controlled online experiments covers these foundations. In a marketplace, buyers compete for the same inventory while sellers change their behaviour across buyers, so treatment and control can affect one another. Simple randomisation can then produce a biased estimate. The market’s structure and the experiment’s assumptions need to be tested; one design studied in the literature randomises both sides of the marketplace. Work by Johari and co-authors shows how the bias depends on the balance of supply and demand.
Consider a wholly hypothetical calculation. Across 10 million eligible visits per year, baseline conversion is 10%. A new version produces 10.2%: an increase of 0.2 percentage points, or 2% relative to baseline. At full rollout, that means 20,000 additional orders. If each contributes ₽500 after the ordinary variable costs, the result is ₽10 million in incremental contribution.
Suppose the change requires another ₽2 million in costs not included in that ₽500, while the team and platform cost allocated to this use case is ₽3 million for the same year. That leaves ₽5 million in incremental impact over twelve months after the costs listed above. This does not mean the whole platform costs ₽3 million, nor does it count investment twice: capital and one-off expenditure must be allocated explicitly across the chosen horizon. A decision to continue the use case also has to distinguish the costs that would actually disappear if it stopped. An allocated share of a common platform may remain in the budget even after a particular feature is switched off.
This is a point estimate. An investment decision also needs the range of uncertainty, the actual share of traffic reached, and the duration of the effect. If the feature reaches half the eligible flow, the full projected margin cannot be booked. If returns rise or a discount merely shifts existing orders, part of the uplift disappears. An annual extrapolation from a short test must be checked against seasonality and long-term behaviour.
Why a negative result can still be valuable
A good experimentation platform allows a harmful change to be rejected before full release. The avoided damage must, however, be estimated against a plausible decision without the experiment. If the team would never have released the variant anyway, the platform cannot claim the loss it supposedly prevented.
This is where Goodhart’s law appears: when a measure becomes a target in its own right, the team may optimise it at the expense of the original goal. Clicks rise because of promises the item cannot fulfil; revenue rises through discounts that consume the margin; the share of “winning” experiments rises because the team tries enough metrics. Decision criteria chosen in advance, measurement-quality controls, and a willingness to accept no gain are all necessary.
Netflix and Uber: the customer experiences the quality of the decision
A company can show data to a customer, use it to improve one feature, or build the entire product promise around a continuously updated decision. This is a useful spectrum, not an official standard. The test is simple: if the data flow disappeared tomorrow, would the feature become worse, or would the product cease to make sense?
Netflix reduces the effort of choosing
Netflix describes recommendations in terms of interaction history, ratings, people with similar tastes, and content attributes. Signals from new sessions update the system. The subscription remains the commercial product; personalisation helps a member find something worth watching. Netflix’s explanation establishes the mechanism, but does not by itself give the incremental revenue attributable to a particular model.
A ranking metric is an intermediate engineering measure. The business cares about a satisfying choice, return visits, and willingness to keep the subscription. More viewing time need not mean greater long-term satisfaction. The feature’s effect must be measured at several horizons while preserving the member’s ability to search and choose independently.
Michelangelo and an arrival-time prediction operate at different levels
Michelangelo is Uber’s internal ML platform for data preparation, training, deployment, and model monitoring. In the original account, an Uber Eats example shows a delivery-time prediction becoming part of the customer-facing application. Uber’s publication neatly separates the internal ability to ship models from the external value of predictable delivery.
For the customer, forecast quality means being able to plan. For the company, it creates a hypothesis about ordering, cancellation, support contacts, and repeat use. A more accurate model does not yet prove profit; those transitions still need to be tested. Sometimes a simpler, low-latency model wins overall because a complex model cannot answer in time for the decision.
Uber separately describes historical-data checks, shadow testing, staged rollout, rollback, and fallback paths. These mechanisms protect the product when new models are released. Its account of ML deployment safety shows why training is only one part of the lifecycle. Teams also need input and latency monitoring, degradation detection, an incident owner, and a way to keep serving customers when the model fails.
Fraud prevention makes the cost of error especially clear. Letting a fraudster through is bad, but blocking a legitimate customer is costly too. A model threshold expresses the trade-off among losses, declines, manual review, and customer experience. It is selected for a specific operation and revisited as the environment changes. There is no universal level of “sufficient accuracy”.
Volume alone cannot start a feedback loop
The promise “more users → more data → a better product → more users” works only when every transition is present. Use must create a useful signal; the company must be allowed to apply it; the signal must improve a decision; the customer must notice the improvement; and that improvement must lead to further use. More storage cannot repair a broken link.
Recommendation systems also create the reverse problem: the system itself decides what a user sees. The absence of a click on an item that was never shown does not prove a lack of interest. Training only on the consequences of earlier decisions can cement an old error. Suitable experiments, coverage controls, and segment-level quality checks are needed, not just a global average.
For a small company, access to the outcome of a decision is often more valuable than a large archive. A modest dataset with a clear operational outcome can beat millions of events unconnected to a purchase or loss. But when decisions are rare, the effect is small, and feedback arrives a year later, a sophisticated in-house platform may not pay for itself. Buying an industry signal and testing a simple rule may be cheaper.
ClickHouse sells the ability to work with data
ClickHouse grew out of the analytics needs of Yandex Metrica, became open source in 2016, and formed the foundation of a separate company in 2021. The project’s history illustrates another transition: an internal technical solution can become an external infrastructure product. Here, the commercial object is the technology and its operation, not someone else’s user records.
Internal success does not prove readiness for an external business. Inside the company, users and operating conditions are known and informal arrangements may be acceptable. Outside, customers expect self-service onboarding, tenant isolation, support, version management, sales, and contractual accountability. That is a new product with its own costs. Releasing open-source software does not mean accumulating a shared database of customer data either: community feedback and a customer’s data are different assets.
Six conditions that make a deal worth pursuing
My working hypothesis is that a company not founded as a data supplier should first test internal value and a customer-facing feature, while evaluating direct transfer as a separate proposition. This is an order of exploration, not a universal maturity ladder: licensing may be the natural starting point for a specialised information business.
| Condition | Evidence required |
|---|---|
| Rights | A map of provenance, purposes, and permitted actions; clear roles and a process for handling people’s requests |
| Scarcity | The buyer can explain why the signal beats an available alternative and cannot be replaced cheaply |
| Repeatability | The reason a customer will renew is known, as is what happens to value when updates stop |
| Decision | The buyer’s costly action, current method, and expected improvement are named |
| Economics | A positive result after preparation, sales, support, infrastructure, and contract termination |
| Bounded harm | Leakage, re-identification, customer dependence, and loss of proprietary advantage are understood |
These conditions primarily describe a sustainable business with recurring delivery. A one-off deal can also be profitable: instead of renewal, prove that its single payment covers preparation and every continuing obligation. The other checks still apply.
This is not a scorecard in which points are added up. Excellent demand does not compensate for absent rights, and high margin does not make unbounded harm acceptable. Each condition benefits from a stop rule written in advance: for example, the buyer refuses to state a purpose, demands unrestricted onward transfer, or rejects a deletion procedure.
Imagine an offer to buy access to a behavioural dataset. It looks attractive at first. Then it emerges that the identifiers are persistent, the buyer may enrich the records, the purpose is broad, exclusivity is required, and the buyer rules out audits. Those are not five minor contract edits: the risk and the value of the deal to the seller have both changed. The right response is to reprice or walk away, even after a long negotiation.
Product, finance, the data owner, security, and legal all need to assess the proposal together. Each holds a different part of the result: demand, full economics, quality and provenance, access controls, legal basis, and contract. Residual risk must be accepted by someone with authority over the entire business affected.
Start with one decision card
On Monday, I would not begin by inventorying every terabyte. I would first choose several costly recurring decisions: releasing a change, replenishing stock, setting a price, accepting a payment, or choosing an offer for a customer. For each, record its frequency, the current error, the cost of that error, and the person able to change the action.
Then I would complete a one-page card for one use case. Who consumes the result? What do they do today? Which signal would change their action? Are the rights clear and the data good enough? What will the result be compared against? What must not get worse? What is the total cost over the chosen horizon? Who stops the system, and what happens when it fails? The answers should fit on one page and make sense to both product and finance.
The next step is a small test. For an external buyer, that means evidence of willingness to pay for a specific decision and an agreed sample. For an internal process, it means investigation and a controlled change. For a customer-facing feature, it means a limited rollout with a useful outcome measured. None requires the company to build its own warehouse, experimentation platform, and universal model first.
Scale the mechanism whose effect can be observed: improve its quality, automate repeatable operations, lower its cost, and widen its reach. A failed test also leaves something valuable when it is clear in advance which expensive bet it allowed the company to avoid. Every new use, however, still needs its purpose and economics checked.
A good team should eventually answer “where is the profit?” in concrete terms: here is the decision, here is the comparison, here is the action, here is the incremental outcome, and here are the costs. Only then is it useful to discuss platform size and the next model. Until then, talk of data wealth remains a hypothesis.
Five conclusions I would take into practice
- 01Data becomes an economic asset when the rights, the use case, and the mechanism that produces money are known. The size of a warehouse proves none of this on its own.
- 02Licensing can generate meaningful revenue. But the buyer pays for a scarce, useful signal, freshness, and reliable delivery, while the seller takes on long-term obligations.
- 03Possessing a copy does not confer every right to use it. Consent, purposes, participant roles, and deletion have to be designed alongside the product architecture.
- 04Internal value passes through a changed action and a causal test. Statist and Hippo can organise that process; event volume and experiment count remain intermediate measures.
- 05A data-powered feature pays off through a customer outcome at an acceptable cost of error and operation. Start with one decision and testable economics, then expand the mechanism that works.
Continue the analysis
The “Where’s the Profit in Data, Lebowski?” podcast begins with this question and continues the conversation about data. For the engineering context, see data platform fundamentals and DataOps and MLOps. For the total cost of a useful outcome, see the economics of AI in software development.
Sources and limits of the evidence
Company documents support the mechanisms and historical figures those companies describe. They do not replace an independent ROI assessment. The legal sources below concern the EU and California; requirements in other jurisdictions must be checked separately.
Rights and processing
- EU · Database Directive 96/9/ECDatabase protection and rights in its contents.
- European Commission · Legal grounds for processing dataLawful bases under the GDPR.
- EDPB · Guidelines 05/2020 on consentRequirements for valid consent.
- EDPB · Guidelines 07/2020The functional roles of controllers and processors.
- EU · General Data Protection RegulationArticles 4, 6, 7, 17, 19 and 28.
- EDPB · Anonymisation and pseudonymisationAnonymisation and pseudonymisation are different.
- European Commission · Dealing with requests from individualsGrounds, exceptions and downstream notification.
- EDPB · Opinion 28/2024Case-by-case assessment of AI model anonymity.
- CPPA · Frequently asked questionsSale, advertising-related sharing and opt-out rights.
- California Civil Code · § 1798.140Statutory definitions of sale and sharing.
Licensing and trust
- Reddit · Google partnership, 22 February 2024Structured access through the Data API.
- Reddit · IPO prospectus, March 2024Contract value, terms and expected revenue recognition.
- Reddit · Annual report 2025Other revenue, concentration and licensing risks.
- Reddit · Public Content PolicyStated partner restrictions, not an independent audit.
- FTC · Complaint against Avast, 2024The regulator’s allegations on collection, de-identification and sale.
- Avast · Wind-down of Jumpshot, 30 January 2020The company’s announcement distributed through PR Newswire.
- FTC · Final decision and order against Avast, 2024Scope of the ban, $16.5 million payment and deletion obligations.
- FTC · Final order announcement, June 2024Timing and scope of the final decision.
- FTC · Cambridge Analytica opinion, 2019Deceptive collection and subsequent data use.
- O’Reilly · The R in RAG Stands for Royalties, 2024Corpus-grounded answers, citations and author royalties.
- O’Reilly · Approach to generative AICopyright holders and licensed uses.
Decisions and products
- T‑Bank · StatistA public description of the internal analytics platform.
- T‑Bank · Hippo teamAn official description of the experimentation platform’s purpose.
- Kohavi et al. · Practical Guide to Controlled Experiments on the Web, 2007Causal measurement and limitations of online experimentation.
- Johari et al. · Experimental Design in Two-Sided Platforms: An Analysis of BiasMarket interference, estimation bias and two-sided randomisation.
- Netflix · How recommendations workSignals and personalisation, not the monetary impact of one model.
- Uber · Meet Michelangelo, 2017The internal platform and delivery-time prediction.
- Uber · ML Model Deployment Safety, 2025Shadow testing, staged rollout and rollback.
- Alexey Milovidov · Introducing ClickHouse, Inc.From Metrica’s analytics needs to open source and an independent company.