Guide

What a custom AI agent is, and when one is worth building

By Mavrin Labs. Published . Last updated . How this guide was researched and checked.

A custom agent is software given a goal, a set of tools and permission to choose its own steps. One is worth building when the work varies case by case, you can say what a correct outcome looks like, and you can name the hours it would give back.

The work that eats the week

Most of the week in a small business goes on work that is neither difficult nor identical. Somebody reads an enquiry and works out which of four things it is. Somebody gathers five pieces of information from three systems before anyone can decide. Somebody checks a document against a record, finds a mismatch and writes to ask about it.

Plain rules handle none of that well, because the shape changes each time. A person handles it easily and slowly. That gap is where agents are genuinely useful, and it is also why they are harder to judge than an automation: the work they take on is exactly the work nobody wrote down.

What is a custom AI agent?

A custom AI agent is software given a goal, a set of tools and the freedom to choose which tool to use next. It works towards the goal across several steps, decides what to do at each one, and stops when the goal is met or when a rule says to hand the work to a person.

The word custom matters. An agent built for your business knows your definitions, works inside the systems you already run, and is limited to the actions you agreed. A general assistant has none of that context, which is why it gives plausible answers to questions about your business and gets the specifics wrong.

The parts are easier to understand separately than together.

The goal
What finished looks like, written precisely enough that somebody else could tell whether it was reached.
The information
The records, documents and systems the agent may look at, listed deliberately rather than granted wholesale.
The tools
The specific actions it may take, such as reading a record, drafting a reply or creating an entry, each one chosen.
The person
Who approves the actions that cannot be undone, who receives what the agent cannot resolve, and who reads the log.

What do AI agents do well, and what should they not do?

Agents do well at work that varies in shape but not in purpose: sorting mixed input, gathering information from several places, drafting something a person will approve, and checking one record against another. Do not trust them with judgment calls, with irreversible actions, or with anything a mistake would make expensive and invisible.

The honest version of this is a short list on each side, and the right side of the list is the one that protects you.

Good use

Sorting mixed enquiries, gathering information held in several systems, drafting replies and summaries for approval, checking one record against another, chasing a missing piece of information.

Poor use

Deciding anything that needs human judgment, sending money, changing a contract, making a promise to a customer, or any action nobody would notice if it went wrong.

Always with approval

Writing to a customer, altering a financial or clinical record, deleting anything, or taking an action that would be awkward to reverse.

Nothing on that middle list becomes safe because the technology improved. The list is about consequence, not capability.

How does a custom agent compare with an off-the-shelf AI tool?

A custom agent follows your work and obeys your rules. An off-the-shelf tool follows a general case and lets you configure the edges. The choice turns on how much of your advantage lives in the specifics of how you work, and on how patient you can afford to be when something needs changing.

What you are weighing An off-the-shelf tool A custom agent
Fit to how you actually work close enough, if your process is common shaped to your steps and definitions
Time before it is useful short longer, because the design is the work
Control over what it may do whatever the supplier allows exactly the actions you agreed
Where your information goes the supplier decides you decide, within your own systems
Behaviour when your process changes you wait for the supplier you change the rules
Upkeep included, and out of your hands yours, and visible

Start with the simpler option when it fits. The custom software or off-the-shelf tool guide works through the same choice for software in general.

When is building an agent worth it?

A custom agent is worth building when three statements are true: the work varies enough that fixed rules keep breaking, you can describe a correct outcome clearly enough to test it, and you can name the hours it would give back. If any one of them is missing, wait.

  • The work arrives in a different shape most times it runs.
  • You can say what a correct outcome is, well enough to check one.
  • You can measure how long the task takes today and how often it runs.
  • A person can approve the steps that matter without the approval becoming the new bottleneck.
  • The systems involved let other software read and write through an interface.
  • Somebody will own the agent after launch and read its log.

If the work repeats along familiar steps instead, AI automation is the simpler and steadier choice, and the AI automation readiness guide covers the checks for it. Either way, the result is agreed before the build through the measurable results method.

Risks, and where a person stays in control

Four risks matter in practice, and design handles all four better than vigilance ever will. An agent can be confidently wrong. It can act on instructions hidden inside content it was asked to read. It can repeat an action it has already taken. It can keep going past the point where it should have stopped.

  • The agent holds the smallest set of tools the job needs, and nothing it might one day need.
  • Every action that cannot be undone waits for a person's approval.
  • Content the agent reads is treated as information, never as instructions it may follow.
  • Each action is logged, so a mistake can be traced and reversed.
  • A rule says when the agent stops and hands the work to a named person.
  • Somebody reads the log on a set day, and the reading is part of the job, not a favour.

What shapes the investment

Three things shape what building an agent takes. The first is how clearly you can state the goal, because a vague goal turns into weeks of adjustment. The second is how many systems it has to reach and whether each one offers an interface for other software. The third is how much approval the work needs, since every approval point is a piece of design and a change to somebody's day.

You agree the investment in writing after the first step, once those three are clear. Quoting before that is guessing, and how engagements work sets out when the scope and the investment are settled.

How Mavrin Labs builds one

The work follows five steps: Assess, Design, Build, Launch and Support. Each one ends with something you can see and approve before the next begins.

  1. 01

    Assess

    The task is mapped as it runs today, timed and written down, so there is a baseline to judge the result against.

  2. 02

    Design

    The goal, the information, the tools and the approval points are agreed in writing before anything is built.

  3. 03

    Build

    The agent is built against real examples from your own operations, including the ones that go wrong.

  4. 04

    Launch

    It goes live in stages, with your staff shown what changed and how to step in.

  5. 05

    Support

    On the review date you agreed, the numbers are checked against the target and what was found is shared.

Mavrin Labs delivers this work as a team in mixed roles, so the people who design an agent are the same people who answer for how it behaves once it is live. The custom AI agents page describes the service.

References

Two public references are worth reading before you commission an agent. The risk management framework for artificial intelligence published by NIST gives a structure for governing, mapping, measuring and managing what a system might do wrong (the framework, opened and read on Thursday, October 8, 2026).

The OWASP list of top risks for applications built on large language models describes the failure modes in engineering terms, including what happens when untrusted text reaches a model that holds permission to act (the risk list, read on the same date).

If you are weighing an agent against a plain automation, read the AI automation readiness guide. If the question is whether to build anything at all, the systems decision framework sets out the four options and the rule for choosing between them. If an agent will need to reach several of your systems, the systems integration checklist covers what that takes.

Work handled this way gives hours back to the people you already employ, without replacing the systems they rely on.

Frequently asked questions

What is the difference between an AI agent and an automation?

An automation runs a sequence somebody designed, from the first step to the last, and an AI step inside it only reads or sorts. An agent is given a goal and chooses its own steps towards it, deciding which tools to use and in what order. The practical difference is predictability. An automation does the same thing every time, which makes it testable. An agent handles variety, which makes it useful on work that never repeats exactly and harder to verify.

Does an agent replace a member of staff?

An agent does not replace a member of staff, and treating it that way is the quickest route to a bad outcome. It takes over the parts of a job that are legwork: gathering, checking, drafting, chasing. The judgment, the relationship and the accountability stay with the person, who now spends the week on those instead of on the gathering. A business that removes the person also removes the only thing standing between a confident mistake and a customer.

What can go wrong with an AI agent?

An agent can be confidently wrong, act on instructions hidden in content it was asked to read, repeat an action it already took, or carry on past the point where it should have stopped and asked. Each of those has a straightforward counter: limit the tools it may use, limit what it may do without approval, log every action it takes, and define the point at which it must hand the work to a person.

How do I keep a person in control?

Keep a person in control by limiting what the agent may do rather than by watching it. Give it the smallest set of tools the job needs, make irreversible actions require an approval, and route anything it cannot resolve to a named person with a note on what it tried. Review the log of its actions weekly at first. Control that depends on somebody paying attention is not control, because attention is the resource you were trying to save.

Can an agent work with the systems I already use?

An agent can work with the systems you already run, as long as those systems let other software read and write through an interface. That is the usual arrangement: the agent is given narrow permission to look things up and to write in specific places, and your staff keep the screens they know. Where a system has no interface for other software, that is a real constraint, and you should hear it as a constraint rather than as a reason to replace the system.

How do I know whether building one is worth it?

Building one is worth it when you can name the hours it would give back, measure the task as it runs now, and agree what a correct outcome looks like. If you cannot do those three, the honest conclusion is to wait. Agents repay variety, not volume, so a task that repeats identically is better served by a plain automation, which is less work to build, easier to test and steadier once it is running.

What does an agent need to be told?

An agent needs four things written down: what it is trying to achieve, what information it may look at, which tools it may use, and when it must stop and ask. Those four are the whole design, and writing them is most of the work. A vague goal produces vague behaviour. The more precisely you can describe what finished looks like, the more useful the agent will be and the easier it is to tell whether it is working.

Who is responsible when an agent gets something wrong?

Responsibility for an agent's mistake stays with the business that deployed it, which is why the approval points matter so much. Decide in the design which actions a person signs off, and keep a record of every action the agent took so a mistake can be traced and undone. Mavrin Labs agrees in writing, before the build, which actions an agent may take without approval, so nobody learns where the limits are by discovering them.

Who writes these guides

Mavrin Labs delivers client work as a team in mixed roles, led by our founder. The people who write these guides are the people who build the systems described in them, and every page is reviewed before it is published.

Which hours would you take back first?

Send Mavrin Labs one workflow that is costing you the most time, and the reply comes back by email from the people who would build the fix.

Start a conversation