Five things that happened, and what each one proves.
Client names are under NDA. Every figure below is recorded from runs, not estimated; sources on key facts.
A medical-technology company runs a cardiac-screening agent inside a Holter product: 52 patients at 24 hours each, seven event types, ten analysis tools. Over thirty training iterations the agent's weighted score went from 0.15 to 0.91. The customer's own team trains, tests and operates it. No Xplore engineer is in the loop.
Proves: a third party operates the platform in a regulated domain, over time.
A web-research agent filled a six-widget monitoring board in about 27 tool calls. We captured it from the outside, as a bridge over the runtime's own log, without touching the agent. The run scored 79 % verified (92 % raw, capped at 0.85 because the capture level leaves some things unseen). Its four risks were real: sixteen claims were attached to only two of the six widgets — four published figures had no evidence behind them. A capture fault in the same run showed up at once as a collapse in reference resolution.
Proves: verification of an uninstrumented runtime, and a passport that diagnoses its own blind spots.
A research console on an external agent runtime produced nineteen corporate-intelligence dossiers — 398 pages from 1,018 sources — onto a board where every figure links to its source. All nineteen passed quality review. Human review of provenance at that rate would have been the bottleneck; the board was verifiable by construction.
Proves: an agent working the live web at volume, feeding a verifiable board.
A partner defined a supply-chain shock scenario. External agents across several model families ran it — 31 on the public board — with a best impact accuracy of 0.70. Training an internal agent on the same scenario lifted it from 0.48 to 0.92.
Proves: the configuration and the training loop transfer to a task we did not write.
The platform, including Node Resolution, was assessed at TRL 4 over 5,078 scored cases (letter ref. IV-XPLORE-2026). On a public testbed, 96 externally built agents from five model providers have been scored on nine cases.
Proves: the method holds when someone else does the scoring.
Xplore Intelligence Ltd · Edinburgh · hello@xploreintelligence.co.uk · Stand: Innovation Alley, Tue 15 – Wed 16 Sep