Loading...
×
AIQON

AI Agent Development

An agent is software that decides what to do next. Given a goal and a set of tools it can call, it plans, acts, checks the result and continues. That is a genuinely different thing from a chatbot, and it fails in genuinely different ways.

We build agents that carry out multi-step work against real systems, with the guardrails and audit trail that implies. We are equally willing to tell you when a deterministic script would do the job better, because for a great many tasks it would.

What We Build

Agents that operate your systems rather than talk about them. Reading from and writing to your databases, calling your internal APIs, moving work through a queue, producing a document, raising the exception when something does not fit the pattern.

Retrieval over your own material, so answers are grounded in your documents, your contracts and your data rather than in whatever the model absorbed during training. This is the single most common requirement we see, and the one most often implemented badly.

Tool and API integration, including Model Context Protocol servers, so an agent has a defined set of things it can do rather than a general licence to act.

Workflows where several agents or steps hand off to each other, with the state held somewhere durable so a failure halfway through is recoverable rather than a silent loss.

Human-in-the-loop interfaces: the queue where a person reviews, edits or approves what the agent produced, because for most commercially serious tasks that is the design that actually ships.

When an Agent Is the Right Tool

Agents earn their cost where the work is genuinely variable: the input arrives in an unpredictable shape, the sensible next step depends on what was found, and the rules are too numerous or too fuzzy to enumerate. Triaging inbound requests, reconciling records that almost match, pulling a defensible answer out of a pile of unstructured documents.

They are the wrong tool for work that is already deterministic. If you can write the rules down, write them down. A script that always does the same thing is cheaper, faster, testable, and does not need a person checking its output.

They are also the wrong tool where a wrong answer is expensive and undetectable. An agent that is right ninety-five per cent of the time is excellent for drafting and unacceptable for posting entries to a ledger, unless something downstream catches the other five per cent.

The first question we ask is what happens when it is wrong. If nobody can answer that, the design is not finished.

Guardrails, Permissions and Approval

An agent should hold its own credentials, scoped to exactly what it needs, never a borrowed set belonging to an administrator. If it can read one table, it should not be able to drop another.

We separate actions that are reversible from actions that are not. Drafting, proposing and flagging can run unattended. Sending, paying, publishing and deleting go through a person, or through a check that a person defined.

Anything an agent touches, it should have written down: what it was asked, what it decided, which tools it called with what arguments, and what came back. Without that trail you cannot debug a bad outcome, and you cannot answer an auditor.

Content reaching the model from outside your organisation, an email, a web page, an uploaded document, is data and never instructions. Prompt injection is a real attack, not a theoretical one, and the defence is architectural: constrain what the agent is able to do, rather than trusting it to ignore the instruction.

Knowing Whether It Actually Works

A demo proves an agent can succeed once. It says nothing about how often it succeeds, and that is the only number that matters commercially.

We build an evaluation set from your real cases, including the awkward ones, and measure against it. That turns "it seems good" into a figure you can decide with, and it tells you whether a change to the prompt or the model made things better or quietly worse.

Non-determinism is the part teams underestimate. The same input can produce different output, so testing has to be statistical rather than a single assertion, and regressions appear as a shifted rate rather than a red build.

Once it is live, the same measurement continues. Model versions change underneath you, your data drifts, and an agent that was accurate in March can be mediocre by September without anyone having changed a line of code.

Cost, Latency and Running It in Production

Agent economics are unusual: the cost is per use rather than per server, and a loop that retries can multiply it without warning. We design with a ceiling, and we monitor spend per task rather than only in total.

Latency is a design constraint, not a detail. An agent that thinks for forty seconds is fine in a nightly batch and unusable in a live chat. Deciding which one you are building changes the architecture.

Model choice is a trade rather than a ranking. The largest model is not automatically correct: plenty of steps run perfectly well on a smaller, faster, cheaper one, and the useful design routes the hard step to the capable model and the routine step to the cheap one.

We avoid designs that cannot be moved. Providers change pricing, deprecate versions and alter behaviour, so the model sits behind an interface you can swap rather than threaded through the whole codebase.

How We Start

With one task, chosen because it is annoying, frequent and currently done by hand. Not a platform, and not a strategy document.

We establish how it is done today, what a good outcome looks like, and what happens when it goes wrong. Then we build the narrow version, measure it against real cases, and put it in front of the people who do the work now.

If it holds up, it widens. If it does not, you have learned that for the cost of a short engagement rather than a programme, which is a good outcome as well.

Talk to us about AI Agents

Tell us what you are trying to do and we will tell you honestly whether we are the right people for it.

Get in touch