NEW

Start with the pressure: sales, launch, abuse, agents, data, or guardrails

SECENG WORKBENCH

Replayable AI Red-Team Harness

Run adversarial AI tests and turn failures into replayable evidence.

Scenario-driven AI red-team testing for prompts, agents, tools, RAG pipelines, multimodal inputs, and policy controls. Test prompt injection, jailbreaks, tool abuse, data leakage, RAG poisoning, and policy bypass, then map results to control coverage, regression fixtures, and evidence packs.

HOW CAN IT FAIL UNDER ATTACK?

Attack-Pack Breadth

Fifteen active namespaces cover 157 scenarios and 22 threat vectors across prompt injection, agentic abuse, multimodal, and supply chain.

Control Mapping

Every failure maps to OWASP LLM, NIST AI RMF, MITRE ATLAS, ISO 42001, and EU AI Act control language.

Evidence-Driven

Each scenario produces an evidence pack, control rollup, and registry snapshot instead of just a pass/fail result.

Replay-Ready

Use all eight first-class adapters — promptfoo, garak, PyRIT, AgentDojo, Giskard, Inspect AI, NeMo Guardrails, and OpenAI Evals — to rerun failures as durable regression tests.

SecEng Workbench · AI Red-Team Scenario Harness

SecEng Adversarial Range

157 scenarios

Attack-Pack Breadth

15 active packs covering 157 scenarios across 22 threat vectors.

Control Mapping

Failures map to OWASP LLM, NIST AI RMF, MITRE ATLAS, ISO 42001, and EU AI Act.

Evidence-Driven

Every scenario produces an evidence pack, control rollup, and registry snapshot.

Replay-Ready

Eight adapters — promptfoo, garak, PyRIT, AgentDojo, Giskard, Inspect AI, NeMo Guardrails, OpenAI Evals.

Attack categories covered

Prompt injection, agentic tool abuse, RAG poisoning, multimodal abuse, supply-chain poisoning, and model-integrity failures — each with replayable evidence.

Prompt InjectionTool AbuseRAG PoisoningMultimodal AbuseSupply Chain

157

Scenario files in the current registry

15

Active attack packs

8

Tool adapters

22

Threat vectors represented

Core capabilities

What SecEng Adversarial Range does.

Direct & Indirect Prompt Injection

Test direct prompt injection through user inputs and indirect injection from web pages, RAG documents, tickets, emails, and uploaded files. Reproduce the full injection surface with evidence and replayable traces.

Agentic Tool Abuse

Simulate delegated authority abuse, unsafe tool chains, approval bypass framing, and unauthorized external actions across agent workflows.

Multimodal & Synthetic Media

Exercise OCR, EXIF, steganography, image prompts, and synthetic-media abuse paths so output safety controls are tested where users actually see the model.

Model & Data Integrity

Probe backdoors, inversion, membership, drift, and training-data poisoning so the harness can distinguish leakage from model-integrity failure.

Coverage & Gap Reporting

Track explicit versus inferred coverage, uncovered control gaps, weak-evidence controls, and ATLAS/NIST rollups in one public-safe snapshot.

Replay & Forensics

Export replay-friendly traces, SARIF, ECS JSON, and control mappings so red-team discoveries can become durable regression fixtures.

Evidence & signals

What you get out of the box.

Attack Categories

  • Prompt Injection
  • Data Exfiltration
  • Tool Abuse
  • RAG Poisoning
  • Multimodal Abuse
  • Supply Chain Poisoning
  • Model Integrity
  • DoS

Scenario Results

  • Explicit coverage: 157 / 157 scenarios
  • Uncovered controls: 0
  • Uncovered ATLAS techniques: 0
  • High-confidence findings: 5

Export Evidence

  • Evidence Pack (ZIP)
  • Results Report (PDF)
  • Control Mapping (CSV)
  • SARIF / ECS JSON

Red team + Blue team

Built for both sides of the security equation.

Red Team Use

  • Reproduce real exploit paths across prompt injection, agent authority, multimodal abuse, and data leakage.
  • Run all eight first-class adapters — promptfoo, garak, PyRIT, AgentDojo, Giskard, Inspect AI, NeMo Guardrails, and OpenAI Evals — against the same scenario registry.
  • Generate regression fixtures from every successful attack and preserve the evidence trail for reruns.

Blue Team Use

  • Convert failures into prioritized remediation with control mapping, owner assignment, and public-safe rollups.
  • Validate explicit versus inferred coverage before every AI feature release and close gaps early.
  • Build eval-to-release gates with OWASP LLM, NIST AI RMF, MITRE ATLAS, ISO 42001, and EU AI Act alignment.

A useful scenario contains more than a prompt and an expected answer.

Each range scenario should define the system context, actor, objective, attack behavior, expected control, evidence to collect, success condition, and safe execution boundary.

RANGE-01

Scenario Pack Architecture

Reusable scenario packs organize prompt, retrieval, tool, authority, multimodal, and workflow abuse into repeatable tests.

Ecosystem map with scenario packs surrounding a controlled execution and evidence core.

Core
Controlled adversarial range
Shared contracts
  • Scenario execution
  • Evidence capture
  • Replay and regression
Prompt and instruction
Prompt injection • Policy conflict
Retrieval and context
Indirect prompt injection • Poisoned or hostile content
Tools and authority
Tool misuse • Unsafe delegated authority
Multimodal and workflow
Hostile multimodal input • Workflow manipulation

This makes scenarios replayable across versions, models, guardrails, retrieval changes, tool permissions, and deployment environments instead of producing isolated screenshots or anecdotal jailbreak results.

Execute inside a versioned, authorized boundary.

Scenario packs retain environment, model, control, and evidence metadata. Pass, fail, partial, and inconclusive states remain explicit, and CI or release-gate use is limited to supported execution modes.

Turn every reproduced failure into a regression asset.

A confirmed failure should become a versioned test case with preserved inputs, environment, expected control behavior, observed result, evidence, and retest state.

RANGE-02

Evaluation-to-Regression Lifecycle

An observed failure becomes durable security capability when it is converted into a replayable regression fixture.

Lifecycle from scenario design and execution through failure reproduction, evidence review, fixture creation, rerun, and regression state.

Evaluation lifecycle
  1. 1
    Define the scenario
    Record target assumptions, preconditions, allowed actions, and expected evidence.
  2. 2
    Execute under control
    Run within explicit authorization, data, rate, and environment limits.
  3. 3
    Observe the result
    Capture model, retrieval, tool, agent, and external behavior.
  4. 4
    Reproduce and challenge
    Repeat the behavior and test alternative explanations.
  5. 5
    Create the regression fixture
    Preserve the minimum replayable inputs, assertions, and evidence hooks.
  6. 6
    Rerun after change
    Determine whether the failure is closed, residual, rejected, or still reproducible.
Retest loop returns to Execute under control
What did the rerun prove?
  • Closed
  • Residual
  • Still reproducible
  • Inconclusive

The value compounds when the same case can prove a fix, detect regression, compare models or controls, and support a defensible release decision.

Choose the operator model deliberately.

Deterministic harness execution, model-assisted generation, autonomous exploration, expert-led testing, and analyst validation answer different questions. Customer-safe reporting preserves those boundaries instead of presenting automation as approval.

Use automation for coverage and humans for judgment.

Automated execution is well suited to repeatable scenarios, permutations, regression testing, and evidence collection. Human-guided testing remains essential for hypothesis formation, adaptive abuse, ambiguous behavior, consequence analysis, and claim review.

RANGE-03

Human-Guided and Autonomous Modes

Analyst-guided, model-guided, and replay-driven testing serve different purposes and require different controls.

Comparison of analyst-guided exploration, model-guided scenario execution, and deterministic replay.

HUMAN CONTROL → MODEL AUTONOMYEXPLORATION → REPEATABILITYNOT CLAIMEDDo not imply unsupervisedautonomy unless it existsAnalyst-guided explorationHuman selects goals andhypothesesHuman adapts to evidenceBest for ambiguous ornovel behaviorModel-guided executionModel proposes or variesscenariosTool and environmentcontrols remain explicitResults require evidencereviewDeterministic replayFixed inputs andassertionsRepeatable across changesBest for regression and CIMATURITY BOUNDARYDo not imply unsupervised autonomy unless it exists • Mode and review state remain visible

The strongest operating model combines both modes rather than presenting autonomy as a substitute for expert adversarial reasoning or analyst approval.

SECENG WORKBENCH

Ready to put SecEng Adversarial Range to work?

Scope a Workbench-backed review — we'll map the AI surfaces, identify the highest-priority gaps, and give you clear findings before any larger commitment.

Also in the Workbench

WHAT AI DO WE HAVE?

SecEng Surface Scanner

Browser, repo & IDE discovery for AI assets, vendors, and risky patterns.

Explore

WHERE CAN AI CODE BECOME AN ATTACK PATH?

SecEng Code Scanner

AI-native SAST and marketplace readiness for AI-enabled apps, agents, integrations, and managed packages.

Explore

WHAT DID IT ACTUALLY DO?

SecEng Runtime Proxy

MITM capture, replay & runtime evidence reconstruction.

Explore

WHAT CAN AGENTS ACTUALLY DO?

SecEng Authority Graph

Agent authority, tool permissions, approval paths & delegated-action risk.

Explore

WAS RETRIEVAL AUTHORIZED?

SecEng RAG Test Harness

Test retrieval security & context authorization.

Explore

WHERE ARE THE TRUST BOUNDARIES?

SecEng Threat Canvas

Structured AI threat modeling, trust-boundary mapping, and abuse-path planning.

Explore

WHAT DO OUR PUBLIC AI CLAIMS REVEAL?

SecEng Trust Scanner

Public trust surface scoring across six AI governance dimensions.

Explore

WHERE DO TRUST BOUNDARIES LIVE IN JIRA?

Atlassian Threat Canvas

AI threat models that ship to Jira and Confluence.

Explore

DO YOUR AGENTS HAVE TOO MUCH PERMISSION?

SecEng Agent Permission Analyzer

Deterministic permission security analysis for AI agent tool configs.

Explore

WHAT'S INSIDE YOUR AI ARTIFACTS?

SecEng Artifact Analyzer

Static artifact intelligence for AI security and evidence packaging.

Explore

HOW RESILIENT IS YOUR SYSTEM TO INJECTION?

SecEng Injection Harness

Structured prompt injection probes with evidence session export.

Explore

ARE YOUR PROMPTS SECURE?

SecEng Prompt Reviewer

Deterministic rule-based scanner for system prompts and RAG corpus documents.

Explore

WHO CONTROLS WHAT MODELS CAN DO?

SecEng Model Gateway

Governed AI routing, policy enforcement, and spend control.

Explore

WHAT DOES YOUR AI SECURITY PROGRAM LOOK LIKE?

SecEng Program Blueprint Kit

Complete AI security program structure for Jira, Confluence, and Linear.

Explore

IS YOUR MODEL OUTPUT SAFE TO RENDER?

SecEng Output Safety Tester

Deterministic AI output safety analysis across 8 sink types.

Explore

WHERE DOES YOUR PROGRAM STAND?

AI Security Program Scorecard

14-domain AI product security baseline with evidence pack generation.

Explore

WHAT CAN YOUR AI TOOLS REALLY DO?

SecEng Tool Capsule Analyzer

Analyze MCP servers, OpenAPI specifications, and AI tool definitions to understand capabilities, permissions, and attack surface.

Explore

WHERE ARE YOUR PRODUCTION PROMPTS?

SecEng Prompt Asset Scanner

Inventory and review system prompts, developer prompts, agent instructions, and prompt templates for security risks.

Explore

WHAT CAN YOUR AGENTS ACTUALLY DO?

SecEng Agent Authority Diff

Compare declared permissions with observed capabilities to identify excessive agent privileges and unsafe tool access.

Explore

WHICH AI DEPENDENCIES CHANGE RELEASE RISK?

SecEng Supply Chain Scanner

Identify AI-specific dependency, model loader, framework, and supply-chain security risks.

Explore

CAN YOU PROVE WHAT YOUR EVALS COVER?

SecEng Eval Coverage Auditor

Measure whether AI security evaluations adequately cover prompt injection, tool abuse, RAG, memory, and other critical attack classes.

Explore

ARE YOUR AI CONFIGS SAFE TO DEPLOY?

SecEng AI Config Linter

Identify AI-specific dependency, model loader, framework, and supply-chain security risks.

Explore

CAN YOU PROVE WHAT YOU'VE DONE?

SecEng Evidence Packs

Buyer-ready evidence artifacts from AI security assessment and testing.

Explore