<!-- GENERATED by tools/gen-md-twins.py from argus.html. The HTML page is authoritative; if this file disagrees with it, this file is stale. -->
# How OneDroid Argus Works — Holdout, Load and Continuous Verification

> The scenario set is sealed and hash-anchored before the build, run against a pinned artefact the coding agent never tested it on, revealed in full after the verdict, and replayable by an open-source runner you control.

Canonical: https://onedroid.ai/argus

OneDroid

OneDroid Argus

# How Argus verifies what your agent built.

Argus is the verifier your coding agent cannot pass by guessing, because it never
sees the test.

Ask us for early access
The evidence contract

The runner is open source, Apache-2.0, and runs inside your environment.

The evidence contract

## Sealed before the build.

A tester writes the scenario set: a person or an agent who is not the builder, working
from the ticket, the acceptance criteria and the shape of your system. Each scenario
says what to trigger, where to look and what a healthy answer is. The tester seals the
set, and its hash goes on a cryptographic ledger before the work is judged, so anyone
can check that the scenarios existed first and never changed. The hash proves priority.
Keeping them unseen is a separate mechanism: the builder holds a credential that can
ask for a run and read the verdict, and cannot read a scenario or an expected value.
Both halves matter; we say both. Decoy probes that would detect a leak over time are
planned, not built.

The run

## Run unseen, then adjudicated.

The runner sits next to the system under test, inside your environment, and nothing is
mocked. It drives HTTP, the message broker, the database, external delivery, MCP tool
calls, web pages in a real browser, permissions and rate limits, and gathers evidence
by correlation id. During the build it reports only what happened — failing ids and
observed reality, so your agent can fix things — never what was expected. The final
run happens once, against a pinned artefact, and the control plane adjudicates it.

### The reveal

The verdict binds to a specific build, so it cannot be re-pointed at a later one. It
comes with a certificate: proof without the tests, which anyone can check offline with
the CLI and no account. When the tester chooses to reveal the set, every scenario is
published with its trigger, expectation, observed evidence and verdict, and any
scenario can be replayed by the open-source runner against the same artefact — by you,
by your auditor, by anyone. The ledger proves the order of events and that nothing was
edited. Running it again is what proves the result is true.

Every scenario ends as passed, failed or errored, and an error is never a pass. A
broken driver, a missing fixture or an unreachable layer is reported as exactly that,
so a run cannot go green because nothing ran. A load run has a fourth outcome,
degraded. The verdict is a decision about one build, with the evidence attached. It is
not a score: there is no number to optimise for and no “safe” label.

Mode two

## Under load.

A scenario can declare a load: a number of users held for a duration, with the p95
latency and the error rate it must stay under. For message brokers there is more: a
ramp of sessions in steps, up to 2,000 sessions a step by default, that finds the limit. Each
step records the rates offered, sent, confirmed and delivered, the time from publish
to confirm and to delivery, the errors by class, and any time the broker blocked. A
ramp that crosses the limit at its top step has measured the limit, which is not a
failure. Under load, expectations are tolerances, and the report says so.

The limits, stated plainly: the stepped ramp exists for AMQP brokers only, and one for
web interfaces is not built. A broker takes load only after its operator lists it as
one that may, and a live broker should never be listed. No load ramp sits under a
certified verdict or on a schedule.

Mode three

## Continuously.

Once a hand-run is green, or every red in it is one you have read and decided to keep,
the reviewed scenario set goes on a monitor schedule: every 15 minutes at the shortest,
with no one driving it. The web app answers one question first, “Is it OK?”, from the
scheduled runs only: a failing run someone started by hand never turns it red, and a
system no monitor has measured reads “not measured”, never green. What the system
cannot fix by itself is listed under “Needs a person”, with who acts next.

A track record per service and decoy probes in live work are planned, not built.

New, in early use

## One sealed set, two systems.

The same sealed checks can run against two systems, or two versions of one, and their
recorded outputs are compared. The result is a table: rows are checks, columns are
systems, and each cell reads same, differs or not measured. What may differ is declared
before the run, with a tolerance for numbers. A difference a person accepts is pinned
to the two outputs they were shown: if either output changes, the approval stops
counting and the cell differs again.

This is the newest part and we say so: it ran end to end for the first time on
4 October 2026, on three systems. Expect it to change.

Surfaces

## Four ways in, three of them working.

A CLI with machine-readable output for agents; skills your coding agent already knows
how to load; an MCP server with separate credentials for the tester and the builder;
and a CI action whose check can block the merge.

### CLI

Machine-readable output, built for an agent to read rather than a human to
eyeball. Built for macOS and Linux.

Released — access by request

### Agent skills

One skill sets Argus up next to your system. Another turns a plain description
of what should happen into valid scenarios.

Released — access by request

### MCP server

The tester's credential writes, seals and reveals. The builder's can start a run
and read the verdict, never a scenario or an expected value.

Released — access by request

### CI action

A check with a real conclusion. It can fail the build — which is the whole
argument.

Not built yet

The release repository is private today, so the first step is to
ask us for access. The steps after that are in the
Argus docs. This page describes release
0.3.59, of 5 October 2026.

`npx onedroid verify init # in the repo your agent is working in`

This command does not work yet. It is what a one-line start will look
like, not something you can paste today. A site that sells verification does not get
to round that up.

Boundaries

## What Argus is not.

### Not a smoke test

It asks what happens under load, at the step where the system stops being
comfortable, and keeps asking on a schedule after the merge.

### Not a code reviewer

It never reads the diff. It exercises the running system and reports what
happened.

### Not a test generator you keep

Scenarios stay hidden from the builder until the tester reveals them, so nothing
can be optimised for them.

### Not a free hand on a live system

A monitor can watch a live system, on its owner's terms. For an app that moves
money, one switch makes Argus refuse every write the owner has not listed, with
limits. Database access is read-only. A load ramp never targets a live broker.

### Not proof the builder never saw the tests

That is custody — a credential that cannot read a test, and an environment that
does not hold one — described plainly above. The anchor proves priority, not that
the builder never looked.

### Not a score

Passed, failed or errored — per scenario, with the evidence. No aggregate
number, no “safe” badge, nothing to game. An error never counts as a pass.

What it installs

## It brings its own run layer.

Argus reaches its control plane through OneDroid Synapse, the
governed MCP gateway, and keeps the scenario corpus and its own track record in
OneDroid Engram, versioned in a Postgres you control. The written
method behind the gate is open source. Reference material
lives at docs.onedroid.ai.

## Find out what your agent actually built.

Talk to us
Bring us your security review

Free runner, Apache-2.0. Email
michal@onedroid.ai —
you'll reach Michal Bacia, the team that builds the product, not a sales queue.
We reply within one business day.
