Skip to content
#AI4SDLC

Token Burn as Proof of Work: Why AI Metrics Easily Leave to Work (Category AI4SDLC)

#AI4SDLC #AI #Engineering #Management #Metrics #DevEx

Increasingly, I think about token maximizing and metrics like how many tokens an employee has burned. There is a funny, but unpleasant rhyme with proof of work in crypt: a person does not present a result, but a proof of a calculation spent. Look, I burned enough tokens, so I was working right. Just yesterday on the sidelines of Highload++ discussed this with other experts. (debated after State of AI4SDLC on HighLoad + +"). So I decided to write a post with my thoughts on the subject.

I understand why cognate tokens per employee are a seductive metric: it’s simple, technically assembled and looks good on a dashboard for top managers. It is still easy to count licenses, MAUs, the number of requests, tokens cobbled by an employee – all this quickly adds up to the adoption schedule. One problem is that adoption is not directly related to value. (Value or outcome in English.). In general, activity is not a result, and token burn is not engineering performance.

A similar logic already existed in the old management culture, where “work is measured by fatigue” as in the song Nautilus Pompilius. Those who left later tried harder. Those who have been to meetings more often are more involved. Those who send more emails are more productive. Now the same frame easily moves to AI: who burned more tokens, he seems to use the new technology better.

But the tokens themselves say nothing about the development system. Has cycle time decreased? Has the review become faster and better? Has the share of rework decreased? Has recovery been faster and the frequency of incidents decreased? Has it reached sales, customers and money? If these questions are not answered, token burn becomes corporate proof of work. We prove that the calculation was spent. But we do not prove that value was created.

The right framework is not proof of work, but proof of value. Not "how many tokens were burned," but "what value did they get for this computing work." For AI4SDLC, this means looking at the entire contour: task statement, context, generation, review, tests, merge, rollout, operation, quality and cost of the result.

And here begins the adult and unpleasant part. Value needs to be able to measure.

At the level of the entire company, you can look at the financial statements: revenue, margin, costs, growth, customer retention. But then the attribution begins. How can you honestly break down contributions into units, teams, people, platforms, features, and changes to SDLCs? What is the effect of AI, what is the effect of good management, what is seasonality, and what is just a successful release?

This requires a baseline prior to implementation, a normal process map, ownership metrics, a DORA/SPACE/DevEx link to the economy, data quality and agreement on what decisions we make on these numbers. You need managerial maturity, not just a new counter in the billing model.

This is why token maximizing is so attractive. It's cheaper organizationally. You don’t need to know what value is. Don't clean the dirty data. Do not associate engineering metrics with product results. Do not admit that the effect of AI in one command can be excellent, in the second zero, and in the third negative.

You can simply set a goal: increase the consumption of tokens per employee. Avos will do it anyway.

Token burn is useful as a technical and financial telemetry, but dangerous as a managerial goal. It is normal to watch alongside adoption, cost per task, lead time, quality/risk and developer experience. (DevEx). But as soon as the burned tokens themselves become proof of good work, the organization begins to optimize not the result, but the visibility of the activity. AI metrics become adults only when a company is willing to answer an uncomfortable question: not how much computing we’ve spent, but what’s gotten better in a real production system.

P.S. You can read about our approach. My analysis of Anna Gromova's speech From T-Bank at the AI Dev Conf just a month ago, where she talked about what and how we measure and why we think it's the right approach.

#AI #AI4SDLC #Engineering #Management #Metrics #DevEx