CADRIA — Consensus-Anchored Distributed Review of Intelligent Agents — enables reliable oversight of increasingly capable AI systems through networks of weaker, specialised Human Digital Representatives (HDRs) and human-in-the-loop control.
The Pipeline
Every output from a strong AI agent passes through a distributed review pipeline before it reaches the real world.
A capable AI system acts autonomously and produces outputs — content, decisions, plans, actions.
The output is decomposed into verifiable claims, steps, or assertions.
Claims are routed to a dynamically selected network of HDRs with domain expertise — legal, medical, ethics, engineering, finance, safety.
Scores are aggregated to reach consensus with probabilistic confidence and worst-case risk bounds.
Approve, revise, reject, or escalate to human experts. Humans retain ultimate authority.
Why CADRIA
As AI systems become more capable than any single reviewer, oversight must become distributed, adaptive, and anchored in consensus.
Many weaker HDRs collectively oversee stronger AI — no single point of judgment failure.
Real-time gating, bounding, and escalation of actions before they take effect.
Randomised selection, diverse criteria, and full auditability resist gaming and manipulation.
Humans retain ultimate authority for high-impact decisions.
Continuous Learning
Feedback from real-world outcomes improves HDRs and consensus models over time.
New failure modes are detected via continuous red-teaming and monitoring.
Review thresholds and routing adjust dynamically based on risk and context.
Every decision is recorded for transparency and accountability.
Delayed or indirect harms are tracked long after deployment.
Each cycle strengthens the network — oversight quality grows with usage.
CADRIA enables reliable, scalable oversight of increasingly capable AI systems — ensuring safer, more trustworthy and accountable autonomous AI in the real world.