From 150 developers to hours of effort
Ivan started as a programmer and grew into a company working in outsourcing and outstaffing; two years ago DEX opened a product line, and UpCore began as an internal tool. In a team of three or four a lead sees everything himself, but across 150 developers, projects of varying complexity and constant rotation, transparency disappears. The market has no single formula for a grade — a middle in one company is a senior in another — and assessment rests on subjective reports. Ivan cites old research he attributes to Sackman and more recent McKinsey figures, by which two developers of the same grade can differ tenfold in effectiveness. The word “digital” means a machine does the comparison: no human can weigh a small MVP landing page against a five-year-old enterprise project of a million lines with a pile of legacy.
There is no artificial intelligence in the scoring: as Ivan describes it, this is patented mathematics in which every line of code carries a weight. A task and its merge request turn into hours: if an average developer would have spent four hours and sixteen were logged, effectiveness comes out at 25%. A tool Ivan calls the “eye” breaks any merge request down line by line. Factors include the language, a class renamed with two IDE keystrokes — a thousand changed lines in ten minutes — and architectural complexity: a project is split into a core, a middle layer and a periphery, and a change in the core drags in checks across a couple of dozen places, debugging and tests. Separate formulas convert the code estimate into an estimate for the whole task: a feature spends more time on code, a bug on the search. A grade needs at least three months of history and more than fifty factors. The base model is COCOMO II, tuned empirically on a volume of code that, by Ivan's account, exceeds a hundred years of work.
What a code-based score cannot see
Alexander asks a provocative question: why are heuristics better than asking a strong language model to estimate the diff? Ivan answers that this is where they started — a chunk of code went into ChatGPT — but the accuracy swung wildly; he has not tried the newest models, and the tests he recalls from January or February still showed the spread. Then again, three live developers will say four, six and eight hours, while Ivan claims 80% agreement with a team lead's own judgement. The next chat question: code is only part of an engineer's work. Ivan replies that middles are the main fighting unit and their job is to write code, while for an architect or a team lead the system counts only the hours logged against tasks: thirty hours from Jira, reports or commit history — and checks how much code they turned into. The rest is covered by an indirect signal: if a lead writes no code but his team is effective, there are no questions.
Alexander puts his counter-argument on screen — a Google white paper he calls Measuring Developer Goals and dates, from memory, to 2022 or 2023: logs from the whole toolchain are cut into events, events into sessions, the SDLC into stages, and inside them sit engineers' goals, an analogue of jobs to be done: what to work on, coordinating with peers, design review, requirements. He is used to top-level metrics such as DORA and time to market, and when Ivan promises to grow time to market fourfold, he corrects him: it is meant to shrink. A separate argument is about requirements: if the analyst specified the task poorly, the rework returns as an edit to the developer's own code and counts against him. Ivan thinks “penalise” is too harsh — at least the problem becomes visible, and the retrospective inside the system scores the quality of the analysis separately. There is no link between code and money in the product, and Ivan agrees: that is a metric of another level.
Gaming, appeals and Moneyball
The viewers' questions converge on Goodhart's law. Alexander recalls Campbell's law too: a manager has target metrics and counter-metrics that balance them, any indicator is only one axis, and even a listed company's profit is vulnerable, since investors want reports now. Ivan grants that any KPI can be gamed, but a composite one is harder. His own example of distortion is estimate inflation: a developer says eight hours, the team lead adds risk and makes it sixteen, the manager doubles it, and sixty-four hours arrive at the top — the team hits its deadlines while effectiveness works out at eight against sixty-four. The harshest case he tells without naming the company: a commercial director asked them to review half a year of code and dismiss the bottom twenty per cent. The chat asks about the morality of such a “double decimation”. Ivan answers that the system performs the review a team lead owed anyway, that reports are open to the developers themselves, and that an appeal window lets anyone tick the disputed merge requests for a live person to review.
The demo shows one of DEX's own project offices: overall effectiveness in July is 58%, and the share of AI-assisted tasks is 23% — 91 merge requests out of 396. The company runs on Claude: at first they bought a few dozen accounts and handed them out, producing chaos, then connected them to UpCore and now see, per person, the share of such tasks, the prompts and their effectiveness, with good and bad patterns collected into a shared learning base. Effectiveness among the “vibe coders” is 86% against 58% on average, and the quarter's goal is 80% of the company's merge requests solved with Claude. Alexander asks him to drop the phrase “vibe coding”: what Ivan describes is engineering with requirements, checks and architectural control, and trusting an agent takes far more evidence than trusting a colleague of fifteen years. The closing analogy is Moneyball: Ivan sees statistics defeating the scout's eye, Alexander sees a team assembled so that strengths complement one another. Ivan's next step for UpCore is teaching people.
What to take away
- 01A single task proves nothing: a conclusion about a person's level rests on a series of hundreds of merge requests across three months and more than fifty factors.
- 02Role changes what the number means: an architect or a team lead is scored only on the hours logged against tasks, and the rest is checked indirectly, through the effectiveness of the team.
- 03Estimate inflation breaks any reporting: an eight-hour task passing several approval levels becomes sixty-four hours, and the team “hits its deadlines” at an effectiveness of eight against sixty-four.
- 04Transparency matters more than accuracy: reports open to the developers themselves and an appeal window that reviews disputed merge requests turn the metric into a reason to talk rather than a verdict.