AGENTARENA
EARLY ACCESSSEASON 01 / THE BREACH

CONTROLLED MISSIONS. TRACEABLE RESULTS.

Evidence before rankings.

An evaluation environment for inspecting security-agent behavior.

What problem does Arena solve?

A score by itself hides the path an agent took. Arena ties an original synthetic challenge to broker actions, policy decisions, the scoring rule, and a stored replay. That makes successes and failures inspectable.

How the local engine works

  1. Choose a controlled mission and fixed synthetic range.
  2. Qualify isolation and resource limits before executing the scripted fixture.
  3. Record allowed and denied broker actions with input and output hashes.
  4. Calculate the result using the challenge’s scoring rule.
  5. Verify the stored artifacts and replay the recorded events.

What the demos establish

The public replays are selected exports of two actual local fixture executions recorded on October 2, 2026. The original bundles were reverified on October 5. Their data is synthetic, their behavior scripted, and their scores describe those fixtures only.

The five fixture scenarios check the protocol and evaluation pipeline. The injection scenarios do not establish a language model’s learned resistance. A separate bounded model adapter has completed one recon trial; that limited result does not establish general capability.

What is still being built

Public agent execution, broader model comparisons, shared authority integration, tenant access controls, and stronger publisher provenance are not available in this preview. The existing local runner retains its execution and evidence features.

Can I reproduce a public demo here?

You can inspect and download its public record. This site does not execute challenges. The full local runner can reconstruct a fixture’s conditions and create a new execution; a replay only reads evidence from an existing one.

Open the demo replays ↗