BK
← Writing
engineering managementAIorg designdeveloper productivity

Engineering team design in the agentic coding era

AI can make code cheaper to produce without making software delivery proportionally cheaper. That changes how I think about team size, seniority, review, and ownership.

There is a version of the AI coding story where engineering teams get dramatically smaller. If one engineer can run several coding agents at once, the obvious conclusion is that a four-person team should eventually be able to produce what an eight- or ten-person team used to produce. I think some version of that will be true in some environments, but I would be careful about designing an org around the assumption that code generation translates directly into engineering productivity.

The data is already messier than the demos. METR's 2025 randomized study had experienced open-source developers work on repositories they knew well and found that they took 19% longer when AI tools were available, even though the developers believed AI had made them faster. METR tried to repeat the experiment with newer tools and said in February 2026 that the newer data was too affected by selection effects to give a reliable estimate, although it did suggest that tools had improved. Google's 2025 DORA research is more positive, finding a relationship between AI adoption and higher delivery throughput and product performance, but it also still finds a negative relationship with delivery stability.

That does not mean AI coding is useless. I use it constantly and it obviously changes what one engineer can attempt. It means that writing code is only one part of the production function. A model can generate a patch quickly, but someone still has to decide whether the patch solves the right problem, fits the architecture, introduces weird failure modes, can be operated in production, and should be merged at all. METR recently had maintainers review AI-generated patches that had already passed SWE-bench's automated grader and found that roughly half still would not have been merged. A 2026 longitudinal study of professional engineers also found a shift away from code creation and toward what the authors call supervisory engineering work: directing, evaluating, and correcting AI output.

That is the org-design problem I find more interesting. If implementation gets cheaper but verification, architecture, product judgment, and coordination do not get cheaper at the same rate, then the scarce resources on a team move. The team should be designed around those scarce resources rather than around how many people can physically type code.

In the pre-AI model, it was normal to think about capacity in engineers. A squad with six engineers could take on more implementation than a squad with three, even if the relationship was never perfectly linear. With agents, the unit gets fuzzier because one engineer can have several implementation threads moving at once. But that does not mean the team can safely accept several times as much change. Review queues can get longer, test failures can multiply, architectural inconsistency can creep in, and the number of things somebody has to keep in their head can increase faster than the useful output does.

I think this argues for smaller teams only when the surrounding system is strong enough to support them. A mature codebase with good tests, clear service boundaries, safe deployment, useful observability, and an internal platform that makes the right thing easy is a much better candidate for a small agent-heavy team than a tightly coupled system with slow feedback and a lot of tribal knowledge. DORA's 2025 work makes a similar point: high-quality internal platforms are strongly associated with organizations getting more value from AI.

It also changes what I would optimize for in the humans on the team. I want engineers who can own a problem end to end, decompose it well, give an agent enough context to do useful work, recognize when the output is subtly wrong, and make good tradeoffs when the requirements are ambiguous. Technical depth still matters, but the value of judgment goes up when producing another implementation is cheap. The person who can generate five patches is less useful than the person who knows which two should exist and can tell whether either is actually safe.

That does not mean every team should become a collection of senior engineers. I already think senior-heavy teams have their own problems, and AI can make the opportunity problem worse if all of the interesting judgment work gets concentrated in the most experienced people while everyone else becomes an agent operator. Junior and mid-level engineers still need real ownership, debugging experience, design work, and chances to build taste. If we automate away every piece of work that used to teach those skills, we eventually create a team full of people who are very good at supervising systems they never learned how to build.

The manager and technical lead roles change too. If an engineer can have three or four agents producing changes in parallel, review capacity and change coordination become things you have to manage deliberately. I would rather have fewer concurrent changes with clear ownership than maximize the number of agents running. The goal is not to keep the model busy. The goal is to move a coherent product or system forward without creating a pile of verification debt for the humans.

This also changes how I would measure whether a smaller AI-native team is actually replacing the output of a larger pre-AI team. Lines of code, pull request count, and agent task completion are almost useless for this. I would look at how long it takes to get a meaningful change from decision to healthy production, how much review time it consumes, how often it gets rolled back, what happens to incident rates and support load, whether the system stays maintainable, and whether users actually get more value. If those things improve with fewer people, then the team really is more productive.

My guess is that the near-term win from agentic coding is not going to be "every engineering team can immediately be half the size." It is more likely that good teams can hold headcount flatter while taking on more scope, or that a small team can own something that previously would have required a larger one if the architecture and tooling are already good. In messy systems, AI may just let us produce mess faster.

So I would not start team design with a target headcount reduction. I would start by asking what work remains scarce after code generation gets cheap. Right now that still looks a lot like judgment, context, review, architecture, product understanding, and operational responsibility. The smallest useful engineering team is not the smallest group of people that can generate the code. It is the smallest group that can carry all of those things without quality collapsing.