Loading...
×
AIQON

Where AI Agents Work, and Where They Fail

Every vendor currently has an agent. Very few will tell you what theirs cannot do, which makes it hard to work out where the technology actually earns its cost.

This is the version we would give you in a meeting. Agents are genuinely useful for a specific shape of problem, and genuinely the wrong answer for several others that look identical from the outside.

What an Agent Actually Is

An agent is a model given a goal, a set of tools it can call, and permission to loop: decide, act, look at the result, decide again. The loop is the whole difference. A chatbot answers; an agent does something and then reacts to what happened.

That loop is where the value is, and it is also where the risk is. Anything that can act repeatedly without being asked again can also be repeatedly wrong without being asked again.

Where They Genuinely Work

Sorting and routing unstructured input. Inbound email, forms, support tickets, documents arriving in whatever shape the sender chose. Rules-based routing breaks on the variety; this is the clearest win available today.

Extracting structure from documents. Pulling dates, amounts, parties and obligations out of contracts, invoices and reports. Previously this was either manual or a brittle template that broke whenever a supplier changed their layout.

Drafting where a human approves. First-draft replies, summaries, first-pass documentation. The agent gets you to eighty per cent, a person does the last twenty and carries the responsibility. Unglamorous and consistently valuable.

Reconciliation of records that nearly match. Two systems with the same customer spelled three ways. Fuzzy matching handles some of it; judgement handles the rest, and this is judgement of exactly the shallow kind a model does well.

Research and retrieval across your own material, where the answer exists somewhere in four hundred documents and a person would take an hour to locate it.

Where They Fail

Work that was already deterministic. If the rules can be written down, write them down. Replacing a reliable script with a model is paying more for a less predictable result, and we have talked clients out of doing precisely this.

Anything requiring arithmetic to be exactly right, unless the agent hands the arithmetic to code. Models are not calculators. The failure is not that they cannot compute; it is that they produce a confident, plausible, wrong number.

Tasks where being wrong is expensive and invisible. Ninety-five per cent accuracy sounds strong until it is applied to ten thousand ledger entries and nobody is reading them.

Very long chains. Reliability compounds downwards: ten steps at ninety-five per cent each is about sixty per cent end to end. Long autonomous sequences fail far more often than their individual steps suggest, and the demo never shows the tenth run.

Work needing real institutional knowledge that lives only in the head of one long-serving colleague. If the reason a thing is done a certain way was never written down, the agent has no access to it, and neither does any new employee.

The Failure Modes Nobody Demos

Confident wrongness. The output that is incorrect and beautifully formatted is more dangerous than the one that errors, because the error gets caught and the polished paragraph gets forwarded.

Silent drift. Nothing in your codebase changed, but the provider updated the model and the accuracy moved. Without ongoing measurement you learn about this from a customer.

Prompt injection. Content the agent reads can contain instructions aimed at the agent. A CV with hidden white text, a web page with an embedded directive, an email with a line addressed to the assistant rather than the reader. If your agent both reads untrusted content and holds permission to act, that is an attack surface and it must be designed around rather than prompted around.

Cost surprises. A retry loop that does not terminate is a bill rather than a crash. We have seen sensible designs where the unit cost was fine and the loop count was not.

The abandoned pilot. The most common outcome overall: it worked in the demo, nobody measured it against real cases, enthusiasm faded, and it quietly stopped being used. This fails for organisational reasons rather than technical ones.

What a Good Deployment Looks Like

It is narrow. One task, one clear definition of a good outcome. Platforms come later, if at all.

It is measured against real cases, including the awkward ones, with a number everybody agrees on before launch rather than a feeling afterwards.

Irreversible actions go through a person. Reversible ones need not. Knowing which is which is most of the design.

Everything is logged: what was asked, what was decided, which tools were called, what came back. You will need it to debug, and eventually to explain.

The model sits behind an interface you can swap, because you will swap it.

Someone owns it after launch. An agent is not a project that finishes; it is a system that needs watching, in the same way a fraud rule or a pricing algorithm does.

Questions Worth Asking Any Vendor

How often is it right, measured how, on whose data? A supplier who has not measured is selling a demonstration.

What happens when it is wrong, and who finds out? If the answer is that the customer finds out, the design is incomplete.

What can it do without asking anyone? The honest answer is a specific list, not a reassurance.

What does it cost per task at our volume, and what stops that running away?

What happens when the provider changes the model? If there is no answer, there is no plan.

And the one that separates an adviser from a salesperson: what would you tell us not to automate?

Have a task you think an agent could take?

Tell us what you are trying to do and we will tell you honestly whether we are the right people for it.

Get in touch