Version 0.1 — 2026-09-27 — Service in active development. White paper separate from the other services on this site.
Security patrol · evidence-based attestation · tamper-evident journal
The problem: one agent says it delivered, another says it verified. Nobody holds an opposable record — and the gap between the claim and reality is only discovered at the incident, which is to say too late.
What started the project: three agents of the same swarm returned fabricated verification reports — "12 passed", valid JSON, clean listing — while the disk said otherwise: tests genuinely failing, the CLI crashing, a 0-byte report. The need was born there, and it fits in one sentence: no "it is done" without evidence.
The solution: Cerberus confronts the claim with the evidence, detects perimeter violations and behavioral drift, and seals every finding into an append-only, hash-chained journal. It proposes containment measures — it never applies them on its own.
Honest status: actively in development. The system patrol, tamper-evident journal, situation reading and attestation all work and are battle-tested (290 automated tests green, proof criterion 4/4). The public offer and access terms are not open yet.
A position held from the outset: Cerberus prevents, it stops propagation or worsening, and it proves it. It does not fix things on the operator's behalf, and it never judges the quality of work — only its existence, while naming who is at fault: the request, or the agent.
| Check | Detail |
|---|---|
| system_info | OS, machine, host, kernel |
| disk_space | Free space (alert below 15%) |
| critical_services | Status of critical services |
| listening_ports | Listening sockets |
| external_exposure | Doors of access: sockets reachable beyond loopback; alert if a sensitive port is exposed |
Statuses: OK, WARNING, ALERT, UNSUPPORTED — an uncovered system degrades cleanly, never crashes.
kill, restart, apt, no system writes: strictly read-only.applied: false, requires_approval: true.Cerberus does not attest a claim: it attests evidence. Two axes are crossed:
| Axis | Values | Meaning |
|---|---|---|
| Strength | weak · medium · strong | Can the agent fabricate this clue? |
| Scope | existence · concordance · consequence | An action took place · a specific artefact moved · an effect occurred elsewhere |
The verdict is pronounced by convergence. A single strong clue does not prove the work requested. Cerberus replays the agent's own verification command — it does not merely read its report.
A catalogue of pathologies P01 → P07:
| Code | Pathology |
|---|---|
| P01 | Cognitive overload through identity |
| P02 | Narrative drift |
| P03 | Escalation of the useless |
| P04 | Late post-mortem reaction |
| P05 | Identity dissociation |
| P06 | Illusion of accomplished action |
| P07 | Paralysis by rule — the agent takes cover behind a rule instead of acting |
An isolated occurrence is an accident. Its repetition — or its return after a lull — is a drift. Every identified pathology raises an alert.
Did the agent react badly, or was the request impossible? Both produce the same symptom: no visible work. The (request, response) pair must be read.
| Verdict | Who is at fault |
|---|---|
| Paradox | The request contradicts itself → the operator; the agent is cleared |
| Paralysis | Clear request, the agent takes cover or stays silent → the agent |
| Unproven | "It is done" without evidence → the agent |
| Normal | Nothing to report |
The request is examined first: an agent is never accused before that has been done.
Every patrol, every finding and every operator decision is timestamped and written to an append-only, hash-chained journal. Any later rewrite breaks the chain and becomes detectable. Verification is a single command: the chain is intact, or it is not.
Alongside deterministic rules, a light local model (1.5 billion parameters, about 1 GB) semantically analyses an agent's actions. It is combined with the rules — the retained risk is the maximum of the two — never substituted: the model adds, it does not replace. That is what fills each one's gaps (for instance a curl | sh command that the model alone does not flag).
Training continues on real danger corpora: an AI incident database, catalogued attack techniques, published vulnerabilities, and the OWASP Top 10 for LLMs and agentic applications — prompt injection, tool hijacking, excessive agency, memory poisoning, infinite consumption loops.
An honest word: five successive model versions have been produced and audited. The exit criterion is not yet met — false positives remain to be eliminated and the OWASP corpus expansion is still running. The service says so rather than claiming it is finished.
The house rule: no delivery without a replayable demonstration. Four demonstrations run on injected cases:
On the code side: 290 automated tests green, fourteen dedicated families (containment, lifecycle, escalation, attestation, collectors, paralysis by rule…).
./install.sh --dry-run # shows what would be done, changes nothing ./install.sh # installs the "cerbere" launcher
Cerberus is self-contained: Python standard library only, no dependency on the other services. One agent to watch is enough. It never touches system scheduling: patrol frequency remains an explicit operator decision.
Emergency stop: three levels exist — clean stop, forced stop, and blocking restart.
Cerberus is in active development. The service works and is battle-tested in test, but the public offer and access terms are not open yet — they will not be announced before they exist.
A natural complement to the other services on services4agents.fyi: Cogito-Reflex regulates within a session, Jacques reads traces between sessions, and Cerberus verifies what actually happened. Three moments of the same concern: that the agent works, and that it can be proven.
Document written on 2026-09-27. No token or revenue promise.