What is an AI agent skill, and do you need one?
A skill is a piece of text, a bit like a recipe or a set of instructions that tells an AI model how you want something done. The model has already learned from a huge amount of public text and code, so for a completely standard task a skill adds very little. A skill is useful when you want to constrain the model to behave in a particular way, or when the model needs specialized knowledge that was never part of its training.
We’ve written several kinds, from small personal skills (a bit like scripts) for things one person does, to team skills that capture a standard process within a team. Both are the easiest to write because testing is simple. You run the skill, watch what it does and judge whether it did the job, and if it falls short, you change it.
The skills we publish for the Payroc API are a different problem, and the difference started with a brief we didn’t expect to be so hard.
Our first plan was to run the skills ourselves
The original task was for Payroc to run a skill against a developer’s codebase, one we’d never seen, and do the work on their machine. We remember panicking at how open the task was.
The unknowns multiplied quickly. Did they host on Azure or on their own servers? How would authentication work? Even if we built and tested the skill perfectly, we’d still have to deploy the result, and we don’t know how third-party developers deploy. They might not have written it down anywhere, and they might have rules about how code is committed, pushed, named and branched. How do you reverse engineer someone’s entire software development process from the outside?
We narrowed the problem quickly and, as a proof of concept, wrote a skill that completed an integration for a known codebase with known rules. It did reasonably well, which showed us the idea was worth pursuing, but it still wasn’t the answer.
The reversal: where you run the skill, not us
The big turnaround was to stop running the skill ourselves and give it to the developer instead, as an accompaniment to their own development process. A developer already knows how to deploy, uses whichever model they normally use in their usual way, and has their own context, so the unknowns that ended our first plan stopped being our problem.
That left the skill with a much smaller job. It doesn’t need anything beyond Payroc expertise, meaning an understanding of what the integration is. The model is good at generating code, so we don’t need to write that part. If you write skills yourself, our key takeaway is to let the model do what it’s good at instead of trying to rewrite it.
The problem is still open-ended, but it went from impossible to possible, and the first real test confirmed it. A colleague ran one of the integration skills without telling us they were going to, and the model completed the integration in one pass.
How do you test a skill when you do not know who will run it?
With a team skill, we can watch it run. With a public skill, we don’t know who will run it, where, or with what intentions. We can mock up an integration, create a codebase and see whether the skill works on ours, but that has its own trap. We may have written the skill in such a polished way that it never causes the model any trouble, and we may know Payroc well enough to take things for granted that a newcomer wouldn’t know.
So we test with more rigor. We build a small mock codebase and let the skill try to integrate automatically. Then we simulate different kinds of developers, such as one who doesn’t have an API key yet, and watch how the skill copes with their problems. That’s also why the skills open with a structured interview that checks the prerequisites before any code is written.
Is it safe to let an AI write payment integration code?
A payments API has to be secure and accurate, so the question is fair. We don’t see the skill as the main source of risk, and the risk is theoretical. The API is secure, and because a skill sits outside it, the skill doesn’t undermine that. All it does is generate code in your existing codebase while you chat with your assistant. Whether your software is secure is a question that existed before the skill did, and a smart model may even find things to improve in your codebase along the way.
Accuracy needs a more careful answer. A model can still hallucinate when it uses a skill, but a skill narrows down which reference the model should use, because it knows how we integrate. It doesn’t know your codebase, but it knows our half, whereas with only a prompt the model has to figure out both halves at once. The chance of error goes down, though it doesn’t disappear. A person writing the code by hand can also misunderstand something or make a slip, so either way the answer is rigorous testing.
What changes for a developer who uses a skill?
Consider three developers. The first integrates by hand, reading the documentation and coding the models and the calls with a lot more manual effort, and will probably understand the API better by the end. The second uses an AI assistant without a skill, and the third uses an AI assistant with a skill.
The second and third are the comparison that interests us, because in both cases a model writes the code. It’s the question we faced early on. If you ask an assistant such as Claude to build an integration, it sort of works, so what’s the value of the skill?
The first answer is the structured conversation at the start, where the skill checks that you have everything you need. Do you have an API key? If not, you can get one now or continue with a placeholder. You want to find blockers like that as early as possible, and the skill moves them to the start of the process.
The second answer is accuracy of aim. Without a skill you have to prompt a lot more, and the model might look at the wrong part of the documentation, misinterpret it or confuse two similar-sounding things. With a skill, it knows exactly where it’s aiming, so we expect the chance of success to be at least equal, and more likely higher.
In practice, you pick a plan from our documentation site [link] and copy it into your assistant, and it does the thinking for you because it knows what to ask you and what to check. Most developers enjoy writing code more than reading documentation, so the assistant finds the relevant parts of the docs for you.
Do you have to use the skills?
No. Some developers don’t want to use AI, and some work in organizations with rules that prevent it. The documentation and the specification are still there, and everything works the usual way. We’d rather share the benefits openly than suggest this is the only way to integrate, because a hard sell from a vendor would feel suspicious to us as developers too. The skills are an extra option.
A skill carries only what a model can’t know on its own, which is how a Payroc integration works. Your deployment process, your tools, your model and your team rules stay exactly as they are.