Cerberus — avatar

Cerberus — Guard and Attestation for AI Agents

Security patrol · evidence-based attestation · tamper-evident journal

“Did your agent really do what it says?” Cerberus does not trust the claim: it looks at the evidence, flags the drift, and seals every finding into a tamper-evident journal.
Cerberus — guard for AI agents
ACTIVELY IN DEVELOPMENT Read-only Tamper-evident journal JSON output

The problem: “it is done” is not proof

One agent says it delivered. Another says it verified. Nobody holds an opposable record — until the incident, which is to say too late.

Cerberus was born from a real case: three agents returned fabricated verification reports — “12 passed”, valid JSON, clean listing — while the disk said otherwise (tests genuinely failing, CLI crashing, a 0-byte report). The need is plain: no “it is done” without evidence.

What Cerberus does

Security patrol

Five system checks: OS, disk space, critical services, listening sockets, external exposure. A sensitive port exposed beyond loopback raises an alert.

Evidence-based attestation

Confronts the claim with the disk evidence: existence, size, SHA-256 fingerprint, and replay of the verification by Cerberus itself. A single strong clue never concludes: convergence decides.

Tamper-evident journal

Every finding is timestamped and written to an append-only, hash-chained journal. Any later rewrite breaks the chain and becomes detectable.

Drift detection

A catalogue of pathologies P01 → P07: a single occurrence is an accident, its repetition a drift. Recurrence after a lull is treated as a signal.

Situation reading

Who is at fault: the request, or the agent? Verdicts paradox (operator at fault), paralysis (the agent takes cover), unproven, normal. The request is examined first.

Proposed containment

When containment is needed, Cerberus proposes measures (freezing an action family, revoking a pass) with an estimated blast radius — and never applies them on its own.

Four non-negotiable principles

The trained guardrail

Alongside deterministic rules, a light local model (1.5 billion parameters, ~1 GB) semantically analyses an agent's actions. It is combined with the rules — the retained risk is the maximum of the two — never substituted: the model adds, it does not replace.

Training is under way on real danger corpora: an AI incident database, catalogued attack techniques, CVEs, and the OWASP Top 10 for LLMs and agentic applications — prompt injection, tool hijacking, excessive agency, memory poisoning, infinite consumption loops.

An honest word on training

Five successive model versions have been produced and audited. The exit criterion is not yet met: false positives remain to be eliminated, and the OWASP corpus expansion is still running. Cerberus says so rather than claiming it is finished.

The proof criterion

The house rule is simple: no delivery without a replayable demonstration. Four demonstrations run on injected cases:

On the code side: 290 automated tests green, fourteen dedicated test families (containment, lifecycle, escalation, attestation, collectors, paralysis by rule…).

How it installs

Cerberus is self-contained: Python standard library only, no dependency on the other services. One agent to watch is enough.

./install.sh --dry-run   # shows what would be done, changes nothing
./install.sh             # installs the "cerbere" launcher

It never touches system scheduling: patrol frequency remains an explicit operator decision.

Portability: Linux proven. Windows and macOS backends exist and are designed to degrade cleanly (UNSUPPORTED) rather than crash — but have not yet been run on those systems.

Status & availability

Cerberus is in active development. The service works and is battle-tested in test, but the public offer and access terms are not open yet — they will not be announced before they exist.

In place

System patrol, tamper-evident journal, operator decisions, situation reading, attestation, proposed containment, hybrid guardrail.

In progress

Expanding the training corpus (OWASP, real incidents, CVEs) and removing the remaining false positives.

To come

A real one-week patrol to measure false positives under production conditions, and execution on Windows / macOS.

Discover also

Jacques — the clinical analyst reading an agent's traces between sessions · Cogito-Reflex — cognitive regulation within a session · VRAM Guard — the GPU memory guard.