The LLM Will Not Come Back Tomorrow
An LLM can flag a risk. Following it up tomorrow takes stored context, a trigger, and someone accountable for the out...
10 min read
26.09.2026, By Stephan Schwab
When delivery stalls, two fixes get offered: hire more developers or replace them with autonomous AI agents. Opposite headcount plans, same assumption. Both treat execution as the missing piece. Yet neither can decide whether a sales promise belongs in the product, who owns the resulting workflow, or what operations must support after release. A new hire and an agent can both build the wrong thing faster. The scarce capability is judgment with the access and time to connect business intent, technical choices, and operational consequences.
Suppose sales promises customers live order status, but operations updates the source system once a day. Another developer can build a dashboard. An AI agent can draft one by lunch. Neither makes the promise true. Someone has to bring sales, operations, product, and development together, decide what “live” can honestly mean, and change either the promise or the workflow.
That is a judgment problem. Treat it as a capacity problem and the company gets a polished display of yesterday’s data.
The recurring dream of replacing developers has worn plenty of costumes. Readable languages, code generators, visual tools, low-code platforms, and now AI each removed real friction. None removed the need to reason through the business rules, exceptions, and consequences that make software difficult.
The opposite response is just as tempting: when delivery slows, hire another developer. One story subtracts the human; the other adds one. Both can avoid the awkward question of why the current team is stuck.
If the work is understood and the team can absorb another contributor, more capacity helps. If a bounded task has clear checks, AI can make that work faster. Use both where they fit.
But neither is a diagnosis. Ask first:
If those answers are missing, switching between people and agents changes nothing. It only changes how quickly the ambiguity becomes code.
An AI agent can inspect a codebase, plan a change, edit files, run tests, and try again. That is a serious increase in technical reach. The word autonomous makes it sound as if the system can also take over the judgment that gave the task its meaning.
It cannot. As the earlier account of what an AI agent actually is explains, an agent is a model inside a software loop with tools, permissions, and a stop condition. Autonomy describes how many steps it may take between human decisions. It does not turn the loop into a product owner.
The fantasy that AI will run the company by itself has the same flaw at a larger scale: action can be delegated, but accountability stays with the people who designed and authorized the system.
Give the agent the dashboard request and it may build the dashboard. Ask it to challenge the request and it may notice that yesterday’s data cannot support a live-status promise. Good. Sales, operations, and product still have to decide which promise to keep. The agent cannot own that choice.
Local speed does not prove system coherence. The company can now produce five plausible implementations before lunch and still have no idea which one belongs in the product.
Netflix Chief Product and Technology Officer Elizabeth Stone’s conversation with Lenny Rachitsky makes the same distinction: AI-generated output is abundant, while systems thinking is what keeps quality and direction from disappearing in the flood.
The faster the parts move, the more damaging bad connections become:
The bottleneck moved. Typing was never the whole job, but now even the pretense is getting difficult to maintain. The scarce capability is deciding what deserves to exist, how it fits, who must own it, and what the organization will have to live with after release.
That is systems thinking.
The return of the generalist is often described as a technical expansion. Frontend developers learn the backend. Backend developers work with infrastructure. Everyone learns enough AI tooling to be dangerous before breakfast.
Useful, but incomplete.
The difficult boundaries sit between customer behavior, commercial promises, domain rules, product choices, architecture, operations, and the way decisions get made.
Consider a supposedly simple workflow automation. The ticket asks for an approval step.
A technical executor can implement the state transition.
A systems thinker asks different questions:
Those are not interruptions before the real work.
They are the real work.
The code makes the answers executable. It does not make the answers correct. Delivery often fails in the translation between boxes: a domain assumption loses its caveat, a deadline becomes an architectural decision, or a developer solves the literal ticket while the workflow stays broken. Adaptable generalists follow that decision across boundaries and notice where its meaning changes.
That breadth is not a lack of specialization. Connecting domains is the specialization.
The CEO sees delayed initiatives, rising costs, and promises that keep moving through summaries that edit out the awkward connections. The CTO sees more detail, but an incident interrupts an architecture decision, a hiring problem interrupts a product discussion, and an urgent customer issue interrupts the work needed to prevent the next one.
The team sees the local causes: an unstable requirement, a fragile deployment, disagreeing stakeholders, a workaround becoming permanent. Nobody has the mandate and time to follow the pattern across them.
The organization is not short of intelligence. Its intelligence is distributed, busy, and trapped inside roles.
Useful judgment stays close enough to inspect reality while retaining enough independence to question routine explanations. Its advantage is dedicated attention: following a problem across decisions, code, handoffs, and releases until the constraint becomes visible.
The replacement story flatters a familiar management wish: separate thinking from doing, make the doing cheap, and leave the existing decisions undisturbed. If the software still fails to fit the business, blame execution and shop for a better executor.
Adding a developer can serve the same avoidance. It looks constructive without requiring anyone to admit that the brief is wrong, the product has three incompatible definitions, or the commercial promise has no operational owner.
The framing determines what the organization permits the people and tools to do. Ask them to execute the ticket and they will. Ask them to inspect the assumptions behind it and they may find the reason delivery keeps circling back.
The second request is less comfortable. It crosses sales, product, development, and operations. It may reveal that a team waited days for a decision and was then blamed for taking weeks to deliver.
So the company asks for a faster implementation instead.
Then it buys a workshop on alignment.
Apparently irony has a budget code.
The person making these connections may be a senior developer, a product leader, or the CTO. The title matters less than the mandate: follow a problem across the domains that shape it and challenge a brief that does not survive contact with reality.
That judgment looks like:
Technical ability lets that person inspect the code rather than rely on a presentation and test whether an explanation survives contact with the system. AI can speed up the inspection. Neither code familiarity nor AI output alone supplies the business context or decision authority.
Hands-on proximity keeps the judgment honest. Independence and protected attention make the pattern visible.
A CEO should be suspicious of a plan that reports agent throughput or headcount savings while rework and delivery times stay put. Make the operating result concrete.
The company should see:
Headcount reductions and agent throughput are easy to count. Neither proves that lead time, release confidence, or customer outcomes improved.
The economic distinction is straightforward. More people or more automation can increase what the existing delivery system attempts. Better judgment can change why that system keeps losing time in the first place.
The fiction of the autonomous agent tempts leaders to treat a long tool-running sequence as transferred responsibility. It is not. The organization chose the task, attached the tools, granted the permissions, and decided where the loop must stop for review.
The CTO remains accountable for the technical organization and its direction. Domain experts remain accountable for domain truth. Product leaders remain accountable for product choices. Developers remain accountable for the quality of what they create.
The CTO need not approve every tool call. They do need to set permissions, checkpoints, and evidence requirements, and give the team room to surface contradictions. When an agent flags a broken assumption, the relevant people must decide what changes. The point is to make that decision visible and act on it inside the work.
Hire another developer when the work is understood and the team can turn added capacity into useful output. Use an agent when the task is bounded, its results are verifiable, and its permissions and stop conditions are explicit. Those are real uses, not universal answers.
Dedicated cross-domain judgment becomes more useful when:
People and agents can work together. Sometimes the sensible sequence is to expose the real constraint, repair the delivery path, and then add people or automation where they can finally help.
When you do not understand why delivery is stuck, neither a hiring requisition nor an agent demo is the conservative choice.
Either choice can become a way of paying to postpone the hard question.
If you want only technical execution, define the task, its checks, and its owner. Then choose the people and tools that suit it. There is no shame in naming the need honestly.
If you want someone to improve delivery across business, domain, product, technology, and operations, do not lock that person inside the technical box. Give them access to the assumptions, the subject-matter experts, the conflicting incentives, and the decisions that shape the work before it reaches the backlog.
Then measure the result at system level:
AI will keep making more output available. That is not the same as making the organization better at deciding, connecting, and owning.
The companies that understand the difference will not merely ship more artifacts. They will build systems that still make sense after everyone has moved quickly.
That requires technical ability.
It also requires the judgment to know when technology is not the only domain in the room.
Tell me what is happening. I listen, ask a few practical questions, and reflect back what I see: where the risk may sit, what may be blocking delivery, and what looks worth checking next. No pitch, no obligation. Confidential and direct.
Talk it through. Practical reflection, no pitch.
Start a ConversationVisibility and hands-on delivery
Navigator gives your leadership clear insight into patterns, blockers, and capacity. Our Embedded Delivery Partner writes production code with your team and gets delivery moving.