A tech lead builds the team’s capacity to make decisions
My central argument is that AI increases the value of a tech lead who can turn an ambiguous problem into a verifiable outcome and spread that ability across the team. A role built around exclusive knowledge and mandatory personal involvement in every change becomes increasingly fragile.
This extends my earlier position. In “Tech leads and architects,” I described a tech lead as an engineer working within a domain with one or two teams, helping improve their services: reliability, recovery, test automation and deployment. Architects more often work across team boundaries. That describes my experience rather than a universal job specification. Yet the core criterion is already there: the condition of the system and the team’s results.
A tech lead may previously have spent much of the day writing the first version of a solution. Some of that work can now be delegated. The time released is worth spending on questions that otherwise lack an owner: are we solving the right problem, who is affected, what must the team understand for itself, and who has authority to accept the risk?
My argument does not depend on AI being permanently unable to reason about architecture. It can assist with framing, risk discovery and verification too. But the ability to suggest an answer does not establish commitments to users and neighboring teams. A tech lead organizes decision-making around those commitments, including decisions the team should be able to make without the lead.
Problem → constraints → alternatives → verification → a decision with an owner → an observable outcome
My practical question is whether the team can keep changing its system safely, and knows when to stop, if the tech lead takes a week off. If it cannot, adding agents is likely to deepen its dependence on one person.
Faster implementation changes the queue a tech lead needs to see
In my positions for Deep Tech Night and DotNext talk, I connected cheaper generation with work shifting toward problem framing and verification. For this roundtable, the management implication matters: making each contributor faster is insufficient. We need to understand where their output goes next.
Consider a hypothetical team that used to prepare ten changes a week but could properly review and release eight. Agents now prepare thirty. If the subsequent stages remain unchanged, the tech lead inherits a growing queue, more context switching and aging branches. This illustrates a capacity constraint; it is not a measurement from a company.
The diagram represents my working hypothesis about shifting queues. It does not establish that implementation has ceased to constrain every team. Some wait on requirements, some lack a usable test environment, and some struggle to identify a problem customers need solved. A tech lead has to locate the team’s actual queue by following real changes.
In Book Cube, I discussed the transfer of work: an author submits faster while a reviewer receives more material to understand. My management conclusion is to include both people’s time when assessing the benefit. An engineer reporting review overload has provided useful evidence about the process.
The first intervention may be straightforward: smaller changes, less work in progress and an assigned reviewer before work begins. This is consistent with the broad framing of DORA 2025: AI amplifies an organization’s existing strengths and weaknesses. That framing does not guarantee a speedup for any particular team.
Architecture and RFCs: preserve the reasons behind a decision
AI is useful for exploring alternatives, analyzing dependencies, preparing an experiment and drafting an RFC. Yet a polished document can conceal an unanswered question. We can generate ten pages about queues and microservices without identifying which behavior must remain intact or who uses the system in ways its developers never anticipated.
In my account of adopting a culture of writing, I explained why I began actively writing RFCs, ADRs and instructions: information could stop depending on me. The criterion remains the same with agents. A useful document lets the next engineer understand and challenge a decision. More generated pages do not guarantee that.
One idea I particularly valued in my review of Chad Fowler’s Regenerative Software is that implementations can change while contracts, checks and decision rationale survive replacement. That translates into concrete work for a tech lead: identify the system’s promises, extract important knowledge from incidental code details, and ensure verification concerns behavior users actually need.
At DotNext, I recalled an episode from my work as a systems analyst. After layoffs in 2008, a system failed to close the next month: an experienced user had previously resolved conflicts and compensated for its shortcomings by hand. Part of the functioning system existed in someone’s actions and had never entered its description. This story predates AI. It explains why I do not consider repository analysis a sufficient investigation of a system.
What I want to see in an RFC
| Question | What the decision should contain |
|---|---|
| What problem are we solving? | An observable problem, affected users and a sign of improvement. |
| What must remain intact? | Contracts, data, constraints, dependent teams and actual usage patterns. |
| What did we compare? | Alternatives, including making no change, and the reasons for the choice. |
| How will we establish success? | Behavior checks, an experiment, stopping conditions and a recovery method. |
| Who made the decision? | An owner, material objections and conditions for revisiting the decision. |
This is my practical checklist, not a requirement to write an RFC for every edit. A short description can be sufficient for a small local task. The scope of approval should follow the consequences, rather than the model’s ability to produce a lengthy document.
Code review must establish reasons to trust the change
AI can search for defects, spot inconsistencies, explain a difficult section and suggest additional tests. I would use those capabilities before handing a change to a colleague. But “another agent reviewed and approved it” does not establish independence: both agents may have accepted the same mistaken interpretation of a requirement.
Imagine a change to payment refunds. An agent writes an implementation and a test in which a repeated request creates another refund. The test is green and the code matches it, yet the requirement is wrong. We need a criterion grounded in the domain: repeating a request with the same key must not repeat its financial effect. We then need checks for concurrent requests, partial failure and recovery. This teaching example distinguishes consistency among artifacts from correct behavior.
I would ask the author to bring a compact evidence package: the goal, material constraints, what changed, which checks actually ran, what remains unknown and how to undo the change. An agent’s test report should be connected to execution results for the version under review. Once behavior changes, an earlier green run does not automatically settle the question.
The tech lead need not become the final reviewer of every change. Component ownership can be distributed, recurring checks automated, and change classes requiring additional approval defined in advance. People spend more attention on meaning, interactions and gaps in verification; tools handle repeatable work.
Feedback from production remains essential. When an error reaches users, ask which signal was missing and how to add it. Another promise to read code more carefully seldom changes the system. A reproducible case, a check at the component boundary and a clear owner give the team a way to learn from the failure.
Accountability needs the authority to stop an action
“A human remains responsible” is an easy statement to make when nothing follows from it. That person needs authority to constrain the agent, inspect the change, reject it and organize recovery. Assigning responsibility without time, information or a right to stop creates a formality.
My position for Deep Tech Night set a strict boundary: a human approves production releases. In “When agents write code,” I separately discussed conditions for potentially extending autonomy. These are distinct points: a direction of development does not repeal the current release policy.
| Situation | What to delegate | What to assign to people |
|---|---|---|
| Research and drafts | Search approved sources, explore alternatives and run local experiments. | Set the goal, permitted data, standards of evidence and choice of solution. |
| A bounded change | Make changes in an isolated environment and run automated checks. | Define acceptance criteria, component ownership and acceptance of the change. |
| Data, access and external contracts | Prepare the plan, analyze consequences and rehearse. | Explicitly approve risk and authority, with a tested recovery plan. |
| Production release | Prepare the release and execute permitted steps through the established process. | Authorize release and manage consequences under the team’s policy. |
This is a proposed starting framework. Product, operations and security owners refine it within their organization. A tech lead answers for technical soundness within their scope; a manager provides resources and work allocation; a product owner sets goals and priorities. Role names vary. The agreements need to be explicit.
Access policy must still work when an agent misunderstands an instruction. A prohibition on a dangerous operation therefore needs support from tool and environment restrictions. The ability to devise a useful action does not authorize the agent to use any data or change any service.
I am willing to expand autonomy as we accumulate verifiable results for a particular task class: bounded potential damage, independent checks, observability and recovery. General confidence in a new model is insufficient. Nor should one tech lead quietly accept risks beyond their mandate on behalf of the organization.
Which skills carry less weight, and which become critical?
I would talk about a redistribution of value rather than skills disappearing. Syntax knowledge remains useful but distinguishes a strong engineer less when a draft is easy to obtain. Reading program behavior and identifying a mistaken assumption become especially valuable precisely because plausible implementations are more plentiful.
| Less distinctive on its own | Increasingly important |
|---|---|
| Quickly writing standard code from a familiar pattern. | Clarifying the problem and recognizing unnecessary work before implementation. |
| Remembering every API detail without consulting documentation. | Checking semantics, constraints and compatibility in the actual system. |
| Personally writing every architecture document. | Comparing trade-offs and organizing a substantive discussion. |
| Being the mandatory reviewer for every change. | Distributing ownership and establishing reliable checks. |
| Demonstrating a large number of completed tasks. | Showing useful outcomes alongside quality, cost and team growth. |
In everyday work, these become three connected abilities. First, express intent and constraints. Second, choose a verification method capable of disproving a convincing answer. Third, reach agreement with people about the consequences. AI can help at every step, while the tech lead must understand the grounds on which the team acts.
Should a tech lead still code? I would maintain regular contact with code, debugging, tests and operations, especially at difficult system boundaries. I have no universal quota for time spent writing code by hand. The practical criterion is whether the lead can detect an incorrect model of behavior and help the team test it.
Developing engineers requires a plan of its own
In “Juniors after code,” I distinguished two outcomes: what someone accomplished with an agent and what they learned. A quickly completed task may deliver both, one or neither. A polished finished project is weaker evidence of its author’s independence.
In Anthropic’s experiment with 52 engineers learning the Trio library, the AI-assisted group averaged 50% on an immediate knowledge assessment, compared with 67% for the group without AI. That is a difference of 17 percentage points; the time improvement was not statistically significant. A small experiment on learning a new tool does not determine the long-term future of the profession. It gives us a reason to assess learning separately from completion speed.
My review of Stanislas Dehaene’s How We Learn offers a useful educational framework: attention, active engagement, feedback and consolidation. Applying that framework to engineering mentorship is my interpretation. If an agent always proposes the first hypothesis and fixes the first error, we should examine what thinking remains for the learner.
I would structure a learning episode as follows. First, the engineer predicts the behavior: what happens when a request repeats or a dependency fails? They then use the agent, test their assumption and explain any discrepancy. Next, they receive a changed condition and solve it with less assistance. The mentor looks for transfer of understanding, as well as working code.
This requires real, bounded tasks, mentor time and permission to make mistakes in a safe environment. Some exercises should exclude generation; others should teach delegation. A blanket AI ban and handing all learning work to AI both fail to account for the purpose of a particular exercise.
This applies to experienced engineers too. If one senior engineer constantly repairs agent output for everyone, that person becomes overloaded while the team gains little independence. A tech lead should share case discussions, ask people to explain decision rationale and turn significant mistakes into shared learning examples.
What I would change over the next month
I would begin with one repeatable task class in one service: for example, small API changes that leave the access model intact and avoid irreversible migrations. That scope makes it easier to observe outcomes and correct the process before expanding. The following plan is a practical heuristic, not a promise of transformation in four weeks.
| Period | Action | Verifiable result |
|---|---|---|
| Week 1 | Examine recent comparable changes: waiting, review, rework and failures. Agree on the goal and owners. | A map of the actual queue and initial quality and cost measures. |
| Week 2 | Introduce a concise task description, tool restrictions and an evidence package for review. | The team understands agent permissions and what makes a change ready. |
| Week 3 | Try the process on a bounded stream of tasks. Retain failed attempts and include a learning discussion. | Visible failure causes, reviewer workload and contributor independence. |
| Week 4 | Compare outcomes with the starting point and decide whether to expand, adjust or stop. | A decision with supporting evidence, an owner and a date for reassessment. |
I would immediately change assignment size, the handoff to reviewers and the explicitness of permissions. I would retain trade-off discussions, boundary checks, gradual rollout, observability, recovery and incident review. Those practices can be automated; their functions still need to be performed.
Not every service needs a new platform built in-house. When existing isolation, build and release tools already provide the required constraints and visibility, start with those. Add infrastructure to address a specific observed failure or recurring cost.
Look at accepted changes and the condition of the team
The share of AI-written code, query counts and agent counts measure tool use. To assess a tech lead’s work, I care more about useful outcomes across the team. In my article on development economics, I proposed measuring the full cost of an accepted task, including people’s time and failed attempts.
| What to assess | What evidence to look for |
|---|---|
| Work flowing through the team | Time from framing to acceptance and release, review waiting time and work in progress. |
| Quality | Rework, failures after release and the ability to restore service. |
| Full cost | Models, tools, infrastructure and everyone’s labor, including reviewers. |
| Value | A change in user outcomes or a justified reduction in risk and future costs. |
| Independence | Who can explain a decision, diagnose a failure and complete the next similar task. |
Compare similar tasks with acceptance defined in advance. Otherwise, splitting work or selecting only easy examples can create an illusion of progress. Four weeks may be insufficient to assess rare failures or long-term skill growth; no incidents in a small sample does not prove safety.
I would agree on a stopping signal in advance. If authors spend less time but total review and rework costs rise, postpone expansion. If releases are faster but only one person can explain the behavior, fix the distribution of knowledge. The goal is a sustained improvement the team can maintain.
The position I am bringing to the roundtable
An opening statement — about a minute
I would start by saying that a tech lead has always been needed for the team’s outcomes and the system’s quality. AI makes the difference between producing artifacts and making engineering decisions more visible. Code, RFCs and review comments arrive faster; confidence that we are solving the right problem and preserving the system’s commitments requires work of its own.
For me, the role therefore shifts toward organizing that work: clear tasks, verifiable criteria, explicit authority and people’s development. The most dangerous arrangement is many agents and one exhausted tech lead expected to approve everything. A good arrangement is a team that can delegate, explain decisions and detect mistakes. I would assess the transition through useful accepted changes, their full cost and the team’s independence.
Three objections worth discussing
“A strong model will design and verify better than a human.” That may well happen for particular tasks. I am willing to change human involvement based on evidence. But we need to establish who sets the criteria, authorizes action and detects gaps in the criteria themselves. Automating those functions requires evidence for a specific task class.
“All these checks will consume the speedup.” Sometimes they will. That is an experimental result too. We should then change task size, tools or verification methods. If the only way to demonstrate value is to exclude reviewer costs, the economics are incomplete.
“You have just described a good tech lead.” At the level of principles, yes. What changes is the balance of volume: more alternatives and changes arrive, while attention and the team’s capacity to understand consequences do not automatically increase. Existing responsibilities require a different allocation of time, authority and automation.
Questions for the other participants
- Which decision can an agent already carry through to an outcome without your tech lead, and on what grounds?
- Where did the queue grow after introducing AI, and who inherited the extra work?
- How do you establish that an engineer learned something beyond obtaining a correct answer?
- Which check or access boundary actually allowed you to reduce manual oversight?
- What observable result would make you reconsider your current autonomy boundaries?
Five takeaways
- 01A tech lead’s value lies in decision quality and team independence. Rapid artifact production becomes a smaller part of the role.
- 02Faster implementation helps when the team can accept its output. Account for work across the process, including reviewers.
- 03Architecture contracts, behavior checks and decision rationale should survive replacement of code and tools.
- 04Accountability requires authority, observability and recovery. Expand autonomy for task classes supported by evidence.
- 05Shipping work and developing expertise are separate outcomes. A tech lead needs to deliberately support both.
Sources behind the position
The roundtable announcement and date come from the event description supplied for preparation. The main arguments develop my materials below. The tables, first-month plan and teaching example are management proposals made here; external research retains its own limitations.
Book Cube
- Tech leads and architects — Domain responsibility, engineering metrics and discussion of major decisions.
- How I came to a culture of writing — A personal account of using RFCs and ADRs to keep knowledge from depending on one leader.
- Regenerative Software — a review of Chad Fowler’s book — The author’s review: replaceable implementations, preserved behavior and decision rationale.
- Four tensions of AI in development — A note on work shifting between people and stages; research findings retain the limits of the original sample.
- How We Learn — a review of Stanislas Dehaene’s book — The basis for an educational interpretation: attention, engagement, feedback and consolidation.
Talks and articles
- State of AI4SDLC · DotNext, September 25, 2026 — Slides and speaker notes: system boundaries, behavior checks, decision history and the month-end closing episode.
- When code became cheap · positions for Deep Tech Night — The earlier position on queues, expertise and human approval for production releases.
- When agents write code: what remains engineering — Verifiable outcomes, delegation boundaries and conditions for revisiting autonomy.
- Juniors after code — The distinction between a completed task and acquired independence.
- The economics of AI in development — The full cost of an accepted outcome rather than generation cost.
External primary sources
- DORA 2025 — AI amplifies existing organizational strengths and weaknesses. Primary source checked September 30, 2026.
- How AI assistance impacts the formation of coding skills — An experiment with 52 engineers learning Trio; an immediate knowledge assessment. Checked September 30, 2026.