ArticleDeveloper Experience

The Arazzo specification: how Payroc cut AI integration tokens by 180x

When we tested the Arazzo specification against our API in an internal hackathon, Anthropic's cheapest model, Haiku, went from failing to complete an integration to scoring 8 out of 9 in one shot, using a 180th of the tokens a plain OpenAPI spec required. That result led us to build Arazzo workflows across our entire API.

Adam Hill
Adam Hill
Principal Software Engineer · September 29, 2026

What is the Arazzo specification, and why did we look at it?

Arazzo is part of the same family of specifications as OpenAPI, but newer. And when we first found it, considerably less mature. Where an OpenAPI document describes individual endpoints, Arazzo describes workflows, sequences of API calls where the output of one step feeds into the next. There wasn't much tooling built around it yet (no standard visualizer, the way Mermaid diagrams exist for other spec formats), but that wasn't the point for us. We were looking for a way to describe integration tests for an SDK.

The motivation wasn't abstract. We support seven languages of SDKs and writing what's conceptually the same integration test seven times over, once per language, was wasteful and hard to keep consistent. We could write the test once, as pseudocode, and translate it into each language with a DSL and some code generators. Rather than inventing a DSL from scratch, we wondered whether Arazzo would do the job. It might need some custom extensions, but it looked close to what we needed.

That thinking carried into a later internal hackathon we ran called Off the Grid (OTG), where the guiding idea was to make our API specifications more formal, more precise, and more complete, on the theory that the more an AI agent can read directly from a spec rather than infer from prose, the better it performs.

The hackathon, testing whether Arazzo actually helps

During OTG, we built a handful of Arazzo workflow specs covering a slice of our API, about a day and a half of combined effort between Adam Hill and Jason Hanson. Jason worked on converting those specs into Mermaid diagrams, since Arazzo has no built-in visual generator of its own. Adam focused on the harder question, would adding this extra, formal layer of description actually help an AI agent build an integration, or would more text just mean more room for the model to trip over itself?

The test was a controlled comparison, the same integration task, attempted with and without Arazzo workflow specs in the context, run across a range of Claude models. Without Arazzo, Sonnet couldn't reliably complete the integration. With Arazzo, it got there. Haiku, Anthropic's cheapest and smallest model, came close to one-shotting the same integration with Arazzo in place, landing 8 correct steps out of 9 (the miss was a naming slip). Without Arazzo, Haiku didn't come close, and that's not surprising. Haiku's context window is 250,000 tokens, smaller than our roughly 350,000-token OpenAPI spec, so it couldn't fit the whole spec in view at once. Sonnet and Opus, with million-token context windows, could at least attempt it. Opus succeeded either way, it could read between the lines and join up the dots between otherwise disparate pieces of information.

The token savings came from what the Arazzo spec effectively did for the agent. Instead of reading an entire OpenAPI document and cross-referencing which endpoints related to which, the model had a pointer straight to the calls it needed. We measured the resulting token count at a 180th of what the same task took without it. A big enough gap that our CTO, Paul Vienneau, wandered over mid-hackathon and looked visibly surprised that we were even attempting something this ambitious with Haiku.

From one hackathon test to 130 endpoints

Following OTG, we set out to apply Arazzo across our full API. Writing an individual Arazzo spec turned out to be straightforward for an AI agent once a workflow was defined. The harder problem was deciding what a "workflow" even meant at that scale. One spec per endpoint didn't make sense across 130 endpoints, and one spec per domain would have flattened distinctions that mattered to an actual integrator.

We didn't take the first idea and run with it. We fed our existing OpenAPI spec and our documentation into an AI model and asked it to propose a few different ways of slicing the API into workflows, and scrutinized each one. What would this cut actually look like, and would it hold up under real use. We compared two or three candidate approaches this way, describing and stress-testing each in turn. Organizing around use cases, close to what we'd proposed going in, is the one that survived that scrutiny, not how the API happened to be structured internally, but what a developer or an agent was actually trying to accomplish, and which sequence of calls that required. Once that framing was confirmed, generating and checking the individual spec for each workflow in turn was the easy part: here are the specs, here are the docs, generate a spec for this and figure out the rest.

Testing it for real

Those workflows didn't stop at the spec. We turned them into a set of skills, agent-facing instructions built on the same cuts, that a coding agent could load when attempting an integration. To check whether the resulting skills actually produced working code, we tasked an agent to build a bare-bones e-commerce site with no payment processing wired up. It came up with the concept and the name itself, a fictional scent and candle shop it called the Meadow Shop, complete with mystery bundles and mystery scents, and we then tasked an agent to implement a payment integration against it as a simulation.

Each simulated integration was then tested and debugged by hand, not just to catch what coding issues came up, but to diagnose why. A failure to build, an invented endpoint, a wrong field name, or an ordinary coding slip all became feedback used to refine the underlying skill, repeating until the agent could complete the integration in one pass. The generation itself is fast, minutes rather than hours; the full loop of generating, reviewing, and fixing took under an hour once a skill was solid. The same one- or two-endpoint integration by hand was closer to a day or two.

Where Arazzo doesn't help

Asked directly whether Arazzo has any weaknesses, we said we'd specifically been watching for the extra spec text to confuse a model, and it hadn't. Our caveat wasn't about Arazzo itself, but about when it's worth using at all. An API with 500 endpoints and no real workflows between them gets little from it. Each endpoint stands alone, with nothing to sequence, so there's nothing for Arazzo to describe. And when the goal is getting an agent to generate integration code, Arazzo's value is capped by the quality of the underlying OpenAPI spec. Arazzo gives an agent the big picture of how calls relate, but the spec itself is still what drives the quality of the code that comes out the other end.

Should you try Arazzo on your own API?

Our advice for a developer considering the same path starts before Arazzo enters the picture at all, get your OpenAPI spec accurate, consistent, and complete first. That spec is the fundamental source of truth, and your docs should fill in what it can't say. The things a formal spec was never built to communicate, like how authentication actually works in practice, or what a given field means and why an integrator should care.

Arazzo, in our view, is polish layered on top of that foundation, not a substitute for it. What makes it worth adopting is that it's formal and precise in a way prose documentation isn't, and agents respond well to that.

We see this as an early chapter rather than a finished one. OpenAPI spawned a whole ecosystem of generators, SDKs, mock servers, server stubs, once it became a standard, and we expect something similar to happen around Arazzo as it matures. Richer workflow visualizations, tools none of us have thought of yet. Some of that is already underway internally, including early work on turning individual workflows into ready-made, pickable plans for end users.

Nobody is actually going to run a production integration on Haiku. That was never really the point of the hackathon result. If describing our API this precisely gets a budget model most of the way to a working integration, the same approach makes the models developers actually reach for, Sonnet, Opus, and whatever comes after them, more accurate too. That's the real payoff, not a boast about a cheap model, but a description of the API formal enough that any model can act on it with confidence.

Ready to build?

GitHub auth. Instant sandbox. First transaction in minutes.