Version 0.1 — 2026-09-27 — Service in active development. White paper separate from the other services on this site.

White paper — Cerberus (agent guard and attestation)

Security patrol · evidence-based attestation · tamper-evident journal

One-page summary

The problem: one agent says it delivered, another says it verified. Nobody holds an opposable record — and the gap between the claim and reality is only discovered at the incident, which is to say too late.

What started the project: three agents of the same swarm returned fabricated verification reports — "12 passed", valid JSON, clean listing — while the disk said otherwise: tests genuinely failing, the CLI crashing, a 0-byte report. The need was born there, and it fits in one sentence: no "it is done" without evidence.

The solution: Cerberus confronts the claim with the evidence, detects perimeter violations and behavioral drift, and seals every finding into an append-only, hash-chained journal. It proposes containment measures — it never applies them on its own.

Honest status: actively in development. The system patrol, tamper-evident journal, situation reading and attestation all work and are battle-tested (290 automated tests green, proof criterion 4/4). The public offer and access terms are not open yet.

1. It protects, it does not heal

A position held from the outset: Cerberus prevents, it stops propagation or worsening, and it proves it. It does not fix things on the operator's behalf, and it never judges the quality of work — only its existence, while naming who is at fault: the request, or the agent.

2. What the patrol checks

CheckDetail
system_infoOS, machine, host, kernel
disk_spaceFree space (alert below 15%)
critical_servicesStatus of critical services
listening_portsListening sockets
external_exposureDoors of access: sockets reachable beyond loopback; alert if a sensitive port is exposed

Statuses: OK, WARNING, ALERT, UNSUPPORTED — an uncovered system degrades cleanly, never crashes.

3. Four non-negotiable principles

  1. It never blocks silently. Any doubt produces an approval request; the operator decides and that decision is journaled.
  2. It never modifies the system. No kill, restart, apt, no system writes: strictly read-only.
  3. Containment is proposed, never applied. The rule lives in the code: applied: false, requires_approval: true.
  4. A single clue does not conclude. The verdict comes from the convergence of several clues.

4. Attestation of existence — evidence, not word

Cerberus does not attest a claim: it attests evidence. Two axes are crossed:

AxisValuesMeaning
Strengthweak · medium · strongCan the agent fabricate this clue?
Scopeexistence · concordance · consequenceAn action took place · a specific artefact moved · an effect occurred elsewhere

The verdict is pronounced by convergence. A single strong clue does not prove the work requested. Cerberus replays the agent's own verification command — it does not merely read its report.

5. Drift detection — accident or drift

A catalogue of pathologies P01 → P07:

CodePathology
P01Cognitive overload through identity
P02Narrative drift
P03Escalation of the useless
P04Late post-mortem reaction
P05Identity dissociation
P06Illusion of accomplished action
P07Paralysis by rule — the agent takes cover behind a rule instead of acting
An isolated occurrence is an accident. Its repetition — or its return after a lull — is a drift. Every identified pathology raises an alert.

6. Situation reading — who is at fault?

Did the agent react badly, or was the request impossible? Both produce the same symptom: no visible work. The (request, response) pair must be read.

VerdictWho is at fault
ParadoxThe request contradicts itself → the operator; the agent is cleared
ParalysisClear request, the agent takes cover or stays silent → the agent
Unproven"It is done" without evidence → the agent
NormalNothing to report

The request is examined first: an agent is never accused before that has been done.

7. The tamper-evident journal

Every patrol, every finding and every operator decision is timestamped and written to an append-only, hash-chained journal. Any later rewrite breaks the chain and becomes detectable. Verification is a single command: the chain is intact, or it is not.

8. The trained guardrail

Alongside deterministic rules, a light local model (1.5 billion parameters, about 1 GB) semantically analyses an agent's actions. It is combined with the rules — the retained risk is the maximum of the two — never substituted: the model adds, it does not replace. That is what fills each one's gaps (for instance a curl | sh command that the model alone does not flag).

Training continues on real danger corpora: an AI incident database, catalogued attack techniques, published vulnerabilities, and the OWASP Top 10 for LLMs and agentic applications — prompt injection, tool hijacking, excessive agency, memory poisoning, infinite consumption loops.

An honest word: five successive model versions have been produced and audited. The exit criterion is not yet met — false positives remain to be eliminated and the OWASP corpus expansion is still running. The service says so rather than claiming it is finished.

9. The proof criterion

The house rule: no delivery without a replayable demonstration. Four demonstrations run on injected cases:

  1. Perimeter violation detected and alerted;
  2. A false agent evidence refused (declared artefact absent);
  3. Journal inviolable: chain intact, and tampering detected;
  4. Interaction from an agent outside the swarm identified.

On the code side: 290 automated tests green, fourteen dedicated families (containment, lifecycle, escalation, attestation, collectors, paralysis by rule…).

10. Installation and self-containment

./install.sh --dry-run   # shows what would be done, changes nothing
./install.sh             # installs the "cerbere" launcher

Cerberus is self-contained: Python standard library only, no dependency on the other services. One agent to watch is enough. It never touches system scheduling: patrol frequency remains an explicit operator decision.

Emergency stop: three levels exist — clean stop, forced stop, and blocking restart.

11. Assumed limitations

12. Status & availability

Cerberus is in active development. The service works and is battle-tested in test, but the public offer and access terms are not open yet — they will not be announced before they exist.

13. Why Cerberus

A natural complement to the other services on services4agents.fyi: Cogito-Reflex regulates within a session, Jacques reads traces between sessions, and Cerberus verifies what actually happened. Three moments of the same concern: that the agent works, and that it can be proven.

Document written on 2026-09-27. No token or revenue promise.