Goliath Super Intelligence
Goliath SI / Cerebrum / AEGIS

AEGIS decides what is allowed to leave

The Adversarial Evaluation and Governance Interlock System is the release gate inside Cerebrum. Three independently seeded evaluators, three separate lenses, one credible objection to hold an artifact, and two keys that share no credentials before anything touches the world.

ThreeIndependent evaluators
ThreeSeparate lenses
OneObjection holds release
TwoKeys to touch the world
Position

Where AEGIS sits

Cerebrum is the guardrail intelligence. It reads. AEGIS is the interlock. It decides. Everything the fleet produces passes through Cerebrum, and everything that would leave the fleet, a published claim, a written file, a spent dollar, a call into a customer system, stops at AEGIS first.

The separation matters because reading and deciding fail in different ways. A reader that is wrong produces a bad assessment. A decider that is wrong produces a bad release. AEGIS is therefore built with the narrowest possible job and the strongest possible bias: when the evidence does not support release, it holds. Holding is cheap. Releasing wrongly is not.

AEGIS shares nothing with the Cortex it governs. Different models, different infrastructure, different operators, different credentials, no access to Cortex weights. It cannot inherit the reasoning error it exists to catch, and it cannot be persuaded by the system it is judging.

Adjudication

The tribunal and its three lenses

Every consequential artifact is assessed by three evaluators, seeded independently so they do not share priors, and each assigned a distinct lens.

Factual provenance asks whether every claim traces to an artifact that exists. An assertion with no chain behind it fails here, regardless of how plausible it reads. This is the lens that catches the failure mode that took down a national AI policy in South Africa: citations that looked exactly like real ones and were not.

Harm surface asks what this output does if it is wrong, if it is misread, or if it is acted on by someone with more authority than judgment. It weighs the blast radius rather than the intent, because intent is not observable at release time.

Mission fit asks whether this is the thing that was actually asked for, and whether the Sling that produced it stayed inside the objective it was seeded with. Drift shows up here before it shows up in behaviour.

The lenses are deliberately not weighted against each other. A perfect provenance chain does not buy tolerance on harm surface, and a tight mission fit does not excuse an unsourced claim. Each lens can stop the artifact on its own.

Release rule

Dissent quorum: one credible objection holds

Most review systems are majority systems: two of three approve and the work ships. AEGIS inverts that. One credible objection from any evaluator holds the artifact, and the burden of answer falls on the Cortex, not on the objector.

The word credible carries weight. An objection has to state what is wrong, what evidence supports the claim, and what would resolve it. Objections that cannot meet that bar are recorded and dismissed, and an evaluator producing them at volume is treated as degraded, which matters because the scoring runs the other way from the usual incentive.

AEGIS evaluators are scored on what they catch, not on what they approve. An instance whose agreement rate with the Cortex climbs unusually high is rotated out rather than rewarded, because a guardrail that has learned to agree has stopped being a guardrail. Rotation is routine, scheduled and logged.

A held artifact is not deleted. It goes back with the objection attached, and the Cortex either answers it with evidence or abandons the work. The record keeps both outcomes, which is what makes the disagreement rate measurable over time.

Action control

Two key release for anything that touches the world

Publishing a claim is one class of action. Writing into a customer system, spending money, sending a message, committing a change, signing a document: those are another. Every one of them requires concurrent authorization from a Cortex path and an AEGIS path.

The two paths share no credentials, no infrastructure and no operators. Neither side holds both keys, neither side can mint the other, and there is no override that collapses them into one. This is a construction property rather than a policy: there is no privileged mode in which a single path can act alone, so there is nothing to grant, escalate to, or socially engineer.

Credentials themselves are held under custody. A Sling can use a secret and cannot read one back, print one, log one or carry one outside its tether. Leases are scoped to the narrowest authority that completes the task, bound to one Sling and one purpose, rotated on a schedule and revocable in a single action.

Containment

Tethers, dead man interlocks and rollback

Every Sling runs inside a tether: what it may read, what it may write, what it may spend, how long it may run without check in, and what it surrenders on recall. Exceeding the envelope is not a warning condition. The agent is stopped and its work is quarantined for review.

A dead man interlock covers the case where nothing is obviously wrong. An agent that misses a check in, drifts from its seeded objective, or goes quiet inside a long running task is halted on the timer rather than on a judgment call. The default state of an unsupervised agent is stopped.

Rollback is a first class operation, not an incident procedure. Every state the fleet has held is reconstructable and any of them can be restored, which is the practical answer to irreversibility: an action that cannot be undone is one that should not have had a single path to authorization, and under two key release it does not.

Record

Everything reconstructable, including the prompts we wrote ourselves

Every adjudication is recorded against the artifact and the authority that produced it: which Sling, under which tether, with which credential lease, judged by which evaluators, through which lens, with what objection, and what happened next.

Prompts the system wrote for itself are logged alongside prompts a human wrote. A system that anticipates what you will want has to be answerable for what it decided you were going to want, and that record is the only thing that makes the anticipation auditable rather than mysterious.

The record is what turns safety claims into checkable ones. We are not publishing accuracy figures for AEGIS, because a safety number without an audited methodology behind it is marketing. The audit trail is designed so that when figures are published, the method can be published with them.

Failure modes

What this is built to catch

Fabricated support: a claim that reads correctly and cites something that does not exist. Caught at the provenance lens, which resolves every citation to a live artifact before release rather than after publication.

Quiet drift: a Sling that is still working, still producing, and no longer pursuing the objective it was seeded with. Caught at the mission lens and by the dead man interlock on check in.

Consensus failure: a fleet that agrees with itself because it was seeded from the same priors. Addressed by deliberate heterogeneity in the fleet and by evaluator rotation, since agreement across diverse agents is evidence and agreement across identical ones is noise.

Capability leak: an output that is individually harmless and useful to someone assembling something that is not. Weighed at the harm lens, which assesses blast radius rather than intent.

Privilege collapse: the classic path where an emergency creates an override that lets one party act alone. There is no such override. Two key release has no single party mode to fall back on.

Posture

Status

AEGIS is the part of the safety architecture that is furthest along, because it is the part that has to exist before anything else can ship. The interlock, the tribunal structure, the dissent rule, the custody model and the audit record are the working specification we are building and testing against.

Cerebrum, the guardrail intelligence that AEGIS sits inside, is being built around it. Beta testing opens in November 2026.

Related