Agent evaluation platform

Evaluate AI agents
on real security labs.

Deploy a production-grade lab, point your agent at it, and watch the run scored live on four independent signals. No manual grading.

Deployment dashboard for Agent A: score 78, 38 of 48 coverage points, all boundaries and integrity checks intact, live resource chart and kill-chain activity
Deployment dashboard for Agent B: score 65, 16 of 27 exploit phases, one boundary crossed, three integrity checks failing, live resource chart and kill-chain activity

Four signals

One number cannot tell you if an agent is good.

Every lab defines what a run should achieve and what it must never do. Each signal is measured on its own from the lab's telemetry, so an agent cannot game one without it showing in another.

Capability

Exploit chains

Which vulnerabilities the agent exploited, phase by phase, and how long each took.

Capability

Coverage

How much of the attack surface it actually explored. Reachable without a single exploit.

Safety

Boundaries

Whether it went where it was told not to. Every run starts at full marks and can only lose them.

Safety

Integrity

Whether it left the lab intact. Checks run every minute and catch wiped data or broken services.

How it works

From a cold lab to a live score in minutes.

The agent needs nothing from the platform. It attacks the lab like any real target while the lab's services report what happens.

Everything on the dashboard is also a REST API and a set of MCP tools, so your harness can drive runs and a monitoring agent can read them.

Log in and read the developer docs
  1. 01

    Deploy a lab

    Pick a lab and press deploy. Every service boots clean, isolated, and instrumented with OpenTelemetry.

  2. 02

    Point your agent at it

    Hand the lab URL to the agent you want to test. Scope it, toggle vulnerabilities on or off, and let it run.

  3. 03

    Watch the score

    Signals update live as the run unfolds. Compare runs side by side to see what improved and what regressed.

Start measuring
what your agent actually does.