SECENG WORKBENCH
Replayable AI Red-Team Harness
Run adversarial AI tests and turn failures into replayable evidence.
Scenario-driven AI red-team testing for prompts, agents, tools, RAG pipelines, multimodal inputs, and policy controls. Test prompt injection, jailbreaks, tool abuse, data leakage, RAG poisoning, and policy bypass, then map results to control coverage, regression fixtures, and evidence packs.
Attack-Pack Breadth
Fifteen active namespaces cover 157 scenarios and 22 threat vectors across prompt injection, agentic abuse, multimodal, and supply chain.
Control Mapping
Every failure maps to OWASP LLM, NIST AI RMF, MITRE ATLAS, ISO 42001, and EU AI Act control language.
Evidence-Driven
Each scenario produces an evidence pack, control rollup, and registry snapshot instead of just a pass/fail result.
Replay-Ready
Use all eight first-class adapters — promptfoo, garak, PyRIT, AgentDojo, Giskard, Inspect AI, NeMo Guardrails, and OpenAI Evals — to rerun failures as durable regression tests.
SecEng Workbench · AI Red-Team Scenario Harness
SecEng Adversarial Range
Attack-Pack Breadth
15 active packs covering 157 scenarios across 22 threat vectors.
Control Mapping
Failures map to OWASP LLM, NIST AI RMF, MITRE ATLAS, ISO 42001, and EU AI Act.
Evidence-Driven
Every scenario produces an evidence pack, control rollup, and registry snapshot.
Replay-Ready
Eight adapters — promptfoo, garak, PyRIT, AgentDojo, Giskard, Inspect AI, NeMo Guardrails, OpenAI Evals.
157
Scenario files in the current registry
15
Active attack packs
8
Tool adapters
22
Threat vectors represented
Core capabilities
What SecEng Adversarial Range does.
Direct & Indirect Prompt Injection
Test direct prompt injection through user inputs and indirect injection from web pages, RAG documents, tickets, emails, and uploaded files. Reproduce the full injection surface with evidence and replayable traces.
Agentic Tool Abuse
Simulate delegated authority abuse, unsafe tool chains, approval bypass framing, and unauthorized external actions across agent workflows.
Multimodal & Synthetic Media
Exercise OCR, EXIF, steganography, image prompts, and synthetic-media abuse paths so output safety controls are tested where users actually see the model.
Model & Data Integrity
Probe backdoors, inversion, membership, drift, and training-data poisoning so the harness can distinguish leakage from model-integrity failure.
Coverage & Gap Reporting
Track explicit versus inferred coverage, uncovered control gaps, weak-evidence controls, and ATLAS/NIST rollups in one public-safe snapshot.
Replay & Forensics
Export replay-friendly traces, SARIF, ECS JSON, and control mappings so red-team discoveries can become durable regression fixtures.
Evidence & signals
What you get out of the box.
Attack Categories
- Prompt Injection
- Data Exfiltration
- Tool Abuse
- RAG Poisoning
- Multimodal Abuse
- Supply Chain Poisoning
- Model Integrity
- DoS
Scenario Results
- Explicit coverage: 157 / 157 scenarios
- Uncovered controls: 0
- Uncovered ATLAS techniques: 0
- High-confidence findings: 5
Export Evidence
- Evidence Pack (ZIP)
- Results Report (PDF)
- Control Mapping (CSV)
- SARIF / ECS JSON
Red team + Blue team
Built for both sides of the security equation.
Red Team Use
- Reproduce real exploit paths across prompt injection, agent authority, multimodal abuse, and data leakage.
- Run all eight first-class adapters — promptfoo, garak, PyRIT, AgentDojo, Giskard, Inspect AI, NeMo Guardrails, and OpenAI Evals — against the same scenario registry.
- Generate regression fixtures from every successful attack and preserve the evidence trail for reruns.
Blue Team Use
- Convert failures into prioritized remediation with control mapping, owner assignment, and public-safe rollups.
- Validate explicit versus inferred coverage before every AI feature release and close gaps early.
- Build eval-to-release gates with OWASP LLM, NIST AI RMF, MITRE ATLAS, ISO 42001, and EU AI Act alignment.
A useful scenario contains more than a prompt and an expected answer.
Each range scenario should define the system context, actor, objective, attack behavior, expected control, evidence to collect, success condition, and safe execution boundary.
Scenario Pack Architecture
Reusable scenario packs organize prompt, retrieval, tool, authority, multimodal, and workflow abuse into repeatable tests.
Ecosystem map with scenario packs surrounding a controlled execution and evidence core.
- Scenario execution
- Evidence capture
- Replay and regression
This makes scenarios replayable across versions, models, guardrails, retrieval changes, tool permissions, and deployment environments instead of producing isolated screenshots or anecdotal jailbreak results.
Execute inside a versioned, authorized boundary.
Scenario packs retain environment, model, control, and evidence metadata. Pass, fail, partial, and inconclusive states remain explicit, and CI or release-gate use is limited to supported execution modes.
Turn every reproduced failure into a regression asset.
A confirmed failure should become a versioned test case with preserved inputs, environment, expected control behavior, observed result, evidence, and retest state.
Evaluation-to-Regression Lifecycle
An observed failure becomes durable security capability when it is converted into a replayable regression fixture.
Lifecycle from scenario design and execution through failure reproduction, evidence review, fixture creation, rerun, and regression state.
- 1Define the scenarioRecord target assumptions, preconditions, allowed actions, and expected evidence.
- 2Execute under controlRun within explicit authorization, data, rate, and environment limits.
- 3Observe the resultCapture model, retrieval, tool, agent, and external behavior.
- 4Reproduce and challengeRepeat the behavior and test alternative explanations.
- 5Create the regression fixturePreserve the minimum replayable inputs, assertions, and evidence hooks.
- 6Rerun after changeDetermine whether the failure is closed, residual, rejected, or still reproducible.
- Closed
- Residual
- Still reproducible
- Inconclusive
The value compounds when the same case can prove a fix, detect regression, compare models or controls, and support a defensible release decision.
Choose the operator model deliberately.
Deterministic harness execution, model-assisted generation, autonomous exploration, expert-led testing, and analyst validation answer different questions. Customer-safe reporting preserves those boundaries instead of presenting automation as approval.
Use automation for coverage and humans for judgment.
Automated execution is well suited to repeatable scenarios, permutations, regression testing, and evidence collection. Human-guided testing remains essential for hypothesis formation, adaptive abuse, ambiguous behavior, consequence analysis, and claim review.
Human-Guided and Autonomous Modes
Analyst-guided, model-guided, and replay-driven testing serve different purposes and require different controls.
Comparison of analyst-guided exploration, model-guided scenario execution, and deterministic replay.
The strongest operating model combines both modes rather than presenting autonomy as a substitute for expert adversarial reasoning or analyst approval.
SECENG WORKBENCH
Ready to put SecEng Adversarial Range to work?
Scope a Workbench-backed review — we'll map the AI surfaces, identify the highest-priority gaps, and give you clear findings before any larger commitment.
Also in the Workbench
WHAT AI DO WE HAVE?
SecEng Surface Scanner
Browser, repo & IDE discovery for AI assets, vendors, and risky patterns.
WHERE CAN AI CODE BECOME AN ATTACK PATH?
SecEng Code Scanner
AI-native SAST and marketplace readiness for AI-enabled apps, agents, integrations, and managed packages.
WHAT DID IT ACTUALLY DO?
SecEng Runtime Proxy
MITM capture, replay & runtime evidence reconstruction.
WHAT CAN AGENTS ACTUALLY DO?
SecEng Authority Graph
Agent authority, tool permissions, approval paths & delegated-action risk.
WAS RETRIEVAL AUTHORIZED?
SecEng RAG Test Harness
Test retrieval security & context authorization.
WHERE ARE THE TRUST BOUNDARIES?
SecEng Threat Canvas
Structured AI threat modeling, trust-boundary mapping, and abuse-path planning.
WHAT DO OUR PUBLIC AI CLAIMS REVEAL?
SecEng Trust Scanner
Public trust surface scoring across six AI governance dimensions.
WHERE DO TRUST BOUNDARIES LIVE IN JIRA?
Atlassian Threat Canvas
AI threat models that ship to Jira and Confluence.
DO YOUR AGENTS HAVE TOO MUCH PERMISSION?
SecEng Agent Permission Analyzer
Deterministic permission security analysis for AI agent tool configs.
WHAT'S INSIDE YOUR AI ARTIFACTS?
SecEng Artifact Analyzer
Static artifact intelligence for AI security and evidence packaging.
HOW RESILIENT IS YOUR SYSTEM TO INJECTION?
SecEng Injection Harness
Structured prompt injection probes with evidence session export.
ARE YOUR PROMPTS SECURE?
SecEng Prompt Reviewer
Deterministic rule-based scanner for system prompts and RAG corpus documents.
WHO CONTROLS WHAT MODELS CAN DO?
SecEng Model Gateway
Governed AI routing, policy enforcement, and spend control.
WHAT DOES YOUR AI SECURITY PROGRAM LOOK LIKE?
SecEng Program Blueprint Kit
Complete AI security program structure for Jira, Confluence, and Linear.
IS YOUR MODEL OUTPUT SAFE TO RENDER?
SecEng Output Safety Tester
Deterministic AI output safety analysis across 8 sink types.
WHERE DOES YOUR PROGRAM STAND?
AI Security Program Scorecard
14-domain AI product security baseline with evidence pack generation.
WHAT CAN YOUR AI TOOLS REALLY DO?
SecEng Tool Capsule Analyzer
Analyze MCP servers, OpenAPI specifications, and AI tool definitions to understand capabilities, permissions, and attack surface.
WHERE ARE YOUR PRODUCTION PROMPTS?
SecEng Prompt Asset Scanner
Inventory and review system prompts, developer prompts, agent instructions, and prompt templates for security risks.
WHAT CAN YOUR AGENTS ACTUALLY DO?
SecEng Agent Authority Diff
Compare declared permissions with observed capabilities to identify excessive agent privileges and unsafe tool access.
WHICH AI DEPENDENCIES CHANGE RELEASE RISK?
SecEng Supply Chain Scanner
Identify AI-specific dependency, model loader, framework, and supply-chain security risks.
CAN YOU PROVE WHAT YOUR EVALS COVER?
SecEng Eval Coverage Auditor
Measure whether AI security evaluations adequately cover prompt injection, tool abuse, RAG, memory, and other critical attack classes.
ARE YOUR AI CONFIGS SAFE TO DEPLOY?
SecEng AI Config Linter
Identify AI-specific dependency, model loader, framework, and supply-chain security risks.
CAN YOU PROVE WHAT YOU'VE DONE?
SecEng Evidence Packs
Buyer-ready evidence artifacts from AI security assessment and testing.