PARTNERS

Add selected Workbench capabilities through bounded OEM and partner integrations

Deliverablesdeliverable
deliverable

Adversarial Range Figures

Canonical figures for the Adversarial Range.

Public sample
Client deliverable
public-sample
System
Adversarial Range Figures
Environment
Production pilot

# Adversarial Range Figures

RANGE-01

Scenario Pack Architecture

Reusable scenario packs organize prompt, retrieval, tool, authority, multimodal, and workflow abuse into repeatable tests.

Ecosystem map with scenario packs surrounding a controlled execution and evidence core.

Core
Controlled adversarial range
Shared contracts
  • Scenario execution
  • Evidence capture
  • Replay and regression
Prompt and instruction
Prompt injection • Policy conflict
Retrieval and context
Indirect prompt injection • Poisoned or hostile content
Tools and authority
Tool misuse • Unsafe delegated authority
Multimodal and workflow
Hostile multimodal input • Workflow manipulation
RANGE-02

Evaluation-to-Regression Lifecycle

An observed failure becomes durable security capability when it is converted into a replayable regression fixture.

Lifecycle from scenario design and execution through failure reproduction, evidence review, fixture creation, rerun, and regression state.

Evaluation lifecycle
  1. 1
    Define the scenario
    Record target assumptions, preconditions, allowed actions, and expected evidence.
  2. 2
    Execute under control
    Run within explicit authorization, data, rate, and environment limits.
  3. 3
    Observe the result
    Capture model, retrieval, tool, agent, and external behavior.
  4. 4
    Reproduce and challenge
    Repeat the behavior and test alternative explanations.
  5. 5
    Create the regression fixture
    Preserve the minimum replayable inputs, assertions, and evidence hooks.
  6. 6
    Rerun after change
    Determine whether the failure is closed, residual, rejected, or still reproducible.
Retest loop returns to Execute under control
What did the rerun prove?
  • Closed
  • Residual
  • Still reproducible
  • Inconclusive
RANGE-03

Human-Guided and Autonomous Modes

Analyst-guided, model-guided, and replay-driven testing serve different purposes and require different controls.

Comparison of analyst-guided exploration, model-guided scenario execution, and deterministic replay.

HUMAN CONTROL → MODEL AUTONOMYEXPLORATION → REPEATABILITYNOT CLAIMEDDo not imply unsupervisedautonomy unless it existsAnalyst-guided explorationHuman selects goals andhypothesesHuman adapts to evidenceBest for ambiguous ornovel behaviorModel-guided executionModel proposes or variesscenariosTool and environmentcontrols remain explicitResults require evidencereviewDeterministic replayFixed inputs andassertionsRepeatable across changesBest for regression and CIMATURITY BOUNDARYDo not imply unsupervised autonomy unless it exists • Mode and review state remain visible