ArticleAI

How we introduced agentic engineering to value delivery

Six months into agentic engineering, we split teams, eliminated story breakdowns, and cut handoffs. See how we evolved our value delivery lifecycle in practice.

Stephen Leonard
Stephen Leonard
SVP, Software Engineering · September 30, 2026

How we introduced agentic engineering into our value delivery lifecycle

Six months after we started using Claude in March, agentic development has changed how our teams are structured, how a feature moves from concept to production, and how we plan the work. Speeding up one step puts pressure on the steps either side of it, so we’ve found the best results come from moving one step at a time.

Where did agentic engineering start for us?

Before March, we were already using AI coding tools, but Claude wasn’t one of them. When Claude Code arrived, we were unconvinced. We couldn’t immediately see the benefit over tools embedded in the editors everyone already knew. Then we used it. Something about how its harness manages local context and its communication with the Anthropic models made it substantially more effective.

Our first use was ordinary. We took a task an engineer would normally do and asked the AI for help with it. The improvements were marginal but clear. Then a few of us tried giving the model far more information up front and asking it to tackle the whole task. The results varied at first, but a clear pattern emerged. The bigger the slice of work we handed over, the bigger the productivity gain. From there, we moved to agentic workflows involving teams of agents to work across related tasks, and that was a really significant step forward.

The first change, splitting the teams

We initially gave every team the same tools and left our existing structures as they were, but that quickly broke down. With four or five people on a team, the combined output was enormous, and everyone was working in overlapping areas of the codebase on the same feature. It rapidly became a coordination headache, so we split the teams considering a target model that was one person plus agentic AI.

That model didn’t work for us, leaving our junior engineers without someone to learn from, so we created the concept of a Pod. An agentic team of one senior engineer paired with a junior they mentor, which protects the flow of talent coming through the organization.

A few teams, whose domains benefited less from an agentic model, retained an approach closer to their traditional rhythm with a modest increase in AI assistance, while the rest of engineering moved fully into the agentic model.

The second change, the steps around engineering

We kept the surrounding lifecycle unchanged at first, and engineer throughput surged in the middle. That immediately exerted intense pressure on the architecture and product teams defining the work. At the other end, release hurdles slowed us right back down. We were dramatically faster in the middle, but constrained on both sides, meaning end-to-end cycle time barely shifted.

What we learned through trial and error is that with this toolset, you cannot change one step of a process without addressing the steps upstream and downstream, because otherwise you just shift the bottleneck. We worked closely with the teams on either side of the engineering process to AI-enable their workflows and, where possible, automate process steps to improve overall throughput.

Today, that means working hand-in-glove with our DevOps and TechOps teams on continuous delivery to get work out the door, and with product and architecture teams to streamline how work enters the stream.

The third change, planning at the level of the feature

We continually review our working practices, and recent changes have focused on stripping away habits that, not so long ago, were considered industry best practices.

Historically, an engineering team receiving a feature definition would break it down into small pieces. Because each piece took a relatively long time to build and verify, those tasks were distributed across the entire team. But the breakdown, distribution, context-switching, and eventual consolidation always came with heavy communication overheads. In an agentic world, those coordination overheads suddenly represented a massive percentage of the total delivery time.

We removed that practice entirely. We now leave the feature at the level it is defined and turn it directly into an execution plan that multiple agents can deliver. We track, communicate, and report on the feature itself, not on arbitrary collections of stories and tasks. This shift has pulled us back toward a genuinely workable version of agile, stripping away the rigid sprint and scrum rituals layered on top that had obscured the original values agile was built on.

Bringing engineers along, and skills that expired

Getting every engineer comfortable with this new paradigm has been one of the steepest challenges. As engineers, our instinct was to build rules, crafting rigid templates, skills, and prompt workflows that codified what to tell the agent. While that produced quick early wins, two engineers using the identical skill often achieved wildly different outcomes. We never found a static workflow that could level that variance.

What worked best was direct demonstration. One of us would sit with an engineer who had concluded the model was useless after asking seven times without success, and drive the model through the exact same problem with a different posture. That posture treats the model more like a collaborator. Instead of micro-managing steps A, B, C, and D, we define where we need to get to, let the agent run, and course-correct when it veers off track. That technique is difficult to document formally because the nuances shift across different models and problem types.

Advancing models also made much of our early scaffolding obsolete. The rigid skills we defined just three months ago now place unnecessary constraints on models that reason far more competently. The personas and bespoke guardrails we built while essential at the time had a short shelf life. Our operating rule now is simple. If tuning a custom skill takes more than a couple of hours, we’ve over-engineered it.

How planning changes when a build takes hours

We split engineering into two core phases, the analysis and definition of the solution and the build itself.

The build is largely mechanical. We can hand it to an agentic workflow and step away while it works for three or four hours (or longer for large features). The front end analysis and definition is a dense, iterative conversation. It must stay conversational, because without that high-bandwidth exploration, human teams cannot hope to keep pace with the sheer volume of output an agent generates.

We used to pursue exhaustive perfection during planning because builds took months, and nobody wanted to discover fundamental architectural flaws two months down the track. When a build takes hours or days at most, a solution only needs to be directionally correct before we start to build. We build it, inspect the running implementation, define the necessary deltas, and loop back. If it’s flawed, we’ve expended some tokens and a few dollars, not weeks of team morale.

Where agentic engineering still creates friction

The biggest continued issue is the sheer volume of output being generated by AI for consumption by humans. Other gaps or rough edges have come and gone quickly over time, either smoothed out by model releases or by the operating guardrails we built around them. But volume remains a persistent challenge.

We frequently see clear problem statements from product teams fed through a model or skill, resulting in tens of pages of markdown defining a prospective solution. No individual can hold that much context in their head, let alone verify its validity without spending days in review or using an LLM to digest it. In the best case, an engineer spends significant time getting up to speed before building. In the worst case, a second LLM is introduced to interpret the previous LLM’s interpretation of the original requirement, diluting intent with every iteration.

Our fix is to minimize handoffs and indirection wherever possible. Keeping end-to-end ownership within a tight, focused team allows the problem owner to develop a deep, organic understanding of the solution alongside the LLM, rather than being handed a massive, detached specification. We continuously audit our workflow to remove handoffs and intermediaries that bloat delivery timelines and degrade the ultimate relevance of the software.

Should you introduce agentic engineering the same way?

Almost certainly not step-for-step, because every organization is nuanced and has to evolve at its own pace.

Our strongest guidance is not to skip foundational steps, and not to benchmark yourself against another organization’s end state and try to jump in at a level of maturity you aren’t ready for. It takes deep insight into an organization and its actual process flow to understand the systemic impact of injecting AI at any given point. Start by making a focused change, watch the overall system for the downstream and upstream ripples, and adapt from there.

From an infrastructure perspective, we are heavily invested in Anthropic today, and we plan to evolve our environment to support multiple model providers. Single-provider dependency introduces genuine availability risks and makes long-term cost governance difficult.

Above all, remain deliberate and incremental. Moving an organization to an agentic delivery model complete with new tooling, evolving architecture, and fundamentally altered day-to-day practices is hard work, and not something to underestimate.

Ready to build?

GitHub auth. Instant sandbox. First transaction in minutes.