Building AI agents: the hard part is not the model

AI agents are easy to demo and hard to put in front of customers. What we are building, and the four things that decide whether an agent survives real work.

An agent demo takes an afternoon. Give a model some tools, let it decide what to call, and watch it book the meeting. It is genuinely impressive, and it is almost no signal about whether the thing can be put in front of a customer.

We are building AI agents — systems that take an instruction, decide what to do, and act inside a real product rather than answering in a chat window. Most of the engineering has nothing to do with the model.

What an agent actually is

Strip the marketing off and an agent is three things: a model that decides, a set of tools it may call, and a loop that keeps going until it is done or stopped.

Everything hard lives in the last two. The model is a component you can swap. The tools are your product, and the loop is where an agent either does useful work or spends your money going in circles.

Four things that decide whether it survives

It must be allowed to do less than it can

A demo agent has access to everything. A production agent needs a boundary: what it may read, what it may change, and what it must ask about first.

That boundary is not a prompt instruction. Prompts are a request, not a permission system. If an agent must not delete a customer record, it must not hold a tool that can — the model's judgement is the wrong place to enforce it.

It must be allowed to fail visibly

The failure case is not "the model said something wrong". It is "the model did something wrong and nobody noticed for a week".

So the design starts from the failure: what happens when confidence is low, when a tool errors, when the loop runs longer than expected. In LawManager, the AI-powered legal platform we built and run, nothing the model extracts becomes a fact in the system until a person accepts it, and every extracted value points back at the source it came from. In legal work a confident wrong answer is worse than no answer. That is true of more domains than people admit.

It must be affordable at the hundredth run

An agent that loops is an agent with a variable cost per task, and that variable is decided by a model rather than by you. A task that costs pennies in testing can cost a great deal more on a bad input.

That means caps on steps, caps on spend, caching what is stable, and using a smaller model where a smaller model is enough. Cost per completed task is a design constraint from the first week, not something to look at after the invoice.

It must be testable

The thing that makes agents hard to ship is that the same input can produce a different path twice. Ordinary tests do not cover it.

What works is an evaluation set: a fixed collection of real tasks with known-good outcomes, re-run on every change to the prompt, the tools or the model. Without it, you cannot tell an improvement from a regression, and you will eventually ship one believing it was the other.

Where agents are worth it

Not everywhere. An agent earns its place when the work is genuinely multi-step, the steps vary by input, and a person would otherwise do it by hand across several systems.

Where the steps are fixed, you do not want an agent. You want automation — a scheduled job that does the same thing every time, for a fraction of the cost and none of the uncertainty. We tell clients this regularly, and it is usually the cheaper answer.

What we are building

We are building agents that live inside products rather than beside them: taking an instruction, doing the multi-step work across the systems a business already runs, and handing back something a person can check.

The judgement we bring is mostly about restraint — which parts should be an agent, which should be ordinary software, and where a human has to stay in the loop because the cost of being wrong is real.

If you are trying to work out which of those your problem is, that is the conversation.

More write-ups