Skip to content
Xplore
← Cambridge Tech Week hub
What we do · two minutes

We record what an AI agent actually did, and check what it claims against that record.

No prior knowledge assumed. Four paragraphs, then the numbers.

1 · What an agent is

An AI agent is a program that lets a language model do a job on its own: read documents, search the web, query a database, call tools, and write a result. Companies are handing agents real work — support tickets, research, reporting, clinical screening. Soon a bank or a hospital group will run thousands of them, built by dozens of vendors.

2 · The problem

What an agent says it did and what it actually did are two different things. It can cite a document it never opened, or quote a number that appears nowhere in what it read. Today the only way to read the full record of a run is to ask another language model — a guess about a guess. And for agents you did not build, there is no record at all: only the vendor's description.

3 · What we built

Software that sits next to the agent, not inside it. It records what the agent saw and what it said, keeps the two apart, and turns every trace — a query, an opened page, a tool call, a claim, a figure — into one typed graph. Each type carries its own rule for how it is checked. So the same graph shows what happened and computes whether it was correct, with fixed rules, not another model's opinion.

The output is a passport for every run: what was verified, what was not, and how much of the run was visible at all. Nothing changes in the agent, its model or its runtime.

4 · What it lets you do
  • aRun agents you did not build, and still know what they did.
  • bFind the exact step where an agent went wrong — and train it until it stops.
  • cShow a regulator, an auditor or a customer a verdict that can be re-run a year later with the same result.
  • dCheck every run, not a sample — at a flat cost per check, whether there are ten agents or ten thousand.
Recorded from runs, not estimated
TRL 4independent assessment · Edinburgh Napier University · 5,078 cases
79 %verified on a production agent we did not build, on a runtime we do not control
96external agents, five model providers, scored on the public testbed
0.15 → 0.91a cardiac-screening agent, 30 training iterations, run by the customer's own team

Sources and definitions on key facts. Client names under NDA.

What would you like?

Xplore Intelligence Ltd · Edinburgh · hello@xploreintelligence.co.uk · Stand: Innovation Alley, Tue 15 – Wed 16 Sep