CADRIA — Consensus-Anchored Distributed Review of Intelligent Agents — exists to make the deployment of increasingly capable AI systems safe, trustworthy, and accountable.
The Problem
As AI systems begin to exceed the capability of any individual reviewer, traditional oversight breaks down. A single human cannot reliably check the work of a system that operates faster, broader, and deeper than they can. CADRIA's answer: a network of many weaker, specialised Human Digital Representatives that collectively oversee stronger AI — anchored in consensus, bounded by worst-case risk, and ultimately governed by humans.
Key Principles
Many weaker HDRs collectively oversee stronger AI. Oversight capacity grows with the network, not with any single reviewer's ability.
Real-time gating, bounding, and escalation of actions — control is applied at the moment it matters, before outputs reach the world.
Randomised HDR selection, diverse evaluation criteria, and full auditability make the system resistant to gaming, collusion, and manipulation — including by the AI under review.
Humans retain ultimate authority for high-impact decisions. CADRIA augments human judgment; it never replaces it.
Technology Enablers
Machine learning powers routing, scoring, and consensus aggregation.
Secure architecture and data privacy built in from the ground up.
Easy integration into existing systems and workflows.
Optional distributed ledger for tamper-evident audit trails.
Modular, technology-agnostic design for long-term flexibility.
Long-horizon monitoring infrastructure for delayed and indirect harms.
Reliable, scalable oversight of increasingly capable AI systems — ensuring safer, more trustworthy and accountable autonomous AI in the real world.