Skip to content
Xplore
← Cambridge Tech Week hub
Cases · recorded from runs

Five things that happened, and what each one proves.

Client names are under NDA. Every figure below is recorded from runs, not estimated; sources on key facts.

A regulated medical deviceIn production · multi-year · first contract >£1M

A medical-technology company runs a cardiac-screening agent inside a Holter product: 52 patients at 24 hours each, seven event types, ten analysis tools. Over thirty training iterations the agent's weighted score went from 0.15 to 0.91. The customer's own team trains, tests and operates it. No Xplore engineer is in the loop.

Proves: a third party operates the platform in a regulated domain, over time.

An agent we did not build, on a runtime we do not controlProduction · Aug 2026

A web-research agent filled a six-widget monitoring board in about 27 tool calls. We captured it from the outside, as a bridge over the runtime's own log, without touching the agent. The run scored 79 % verified (92 % raw, capped at 0.85 because the capture level leaves some things unseen). Its four risks were real: sixteen claims were attached to only two of the six widgets — four published figures had no evidence behind them. A capture fault in the same run showed up at once as a collapse in reference resolution.

Proves: verification of an uninstrumented runtime, and a passport that diagnoses its own blind spots.

VolumeJul 2026

A research console on an external agent runtime produced nineteen corporate-intelligence dossiers — 398 pages from 1,018 sources — onto a board where every figure links to its source. All nineteen passed quality review. Human review of provenance at that rate would have been the bottleneck; the board was verifiable by construction.

Proves: an agent working the live web at volume, feeding a verifiable board.

Somebody else's taskDesign-partner benchmark · 2026

A partner defined a supply-chain shock scenario. External agents across several model families ran it — 31 on the public board — with a best impact accuracy of 0.70. Training an internal agent on the same scenario lifted it from 0.48 to 0.92.

Proves: the configuration and the training loop transfer to a task we did not write.

Independent assessmentEdinburgh Napier University · Jul 2026

The platform, including Node Resolution, was assessed at TRL 4 over 5,078 scored cases (letter ref. IV-XPLORE-2026). On a public testbed, 96 externally built agents from five model providers have been scored on nine cases.

Proves: the method holds when someone else does the scoring.

Want one of these on your agent?

Xplore Intelligence Ltd · Edinburgh · hello@xploreintelligence.co.uk · Stand: Innovation Alley, Tue 15 – Wed 16 Sep