Service · VALIDATE
AI Guardrails & Evals Review
Determine what your guardrails and evals actually cover — and what still gets through.
Decision answered
What do the current guardrails and evals actually cover, and what failure cases still pass?
Duration
Typical duration: 2–5 weeks, depending on scope.
Primary output
Guardrails and Evals Coverage Review with regression and release criteria
Best for
Teams with guardrails or evals already in place that need a coverage review before release.
Engagement type
Scoped review
What is in scope
- Guardrail architecture
- Eval suite and coverage
- Representative failure cases
- Regression design
- Release conditions and ownership
Inputs needed
- Guardrail architecture and configuration
- Current eval suite and results
- Representative failure reports
- Release and monitoring context
What the work actually does
- Guardrail architecture, safety policy, refusal, fallback, and monitoring review
- Eval suite, abuse case, failure mode, and regression coverage review
- Prompt/control regression testing and release quality gate recommendations
- Engineering-ready remediation plan for guardrails, evals, and release criteria
What the work delivers
- Guardrails and Evals Review Memo
- Eval Coverage Review
- Failure Mode Register
- Regression Test Plan
- Release-condition recommendations
Evidence produced
- Reviewed coverage observations
- Reproduced failure cases
- Supported regression candidates
- Release-condition evidence requirements
Boundary
What this engagement does not establish
- A guarantee that guardrails block every failure mode
- A universal model-safety assessment
- A certification or compliance determination
- Automatic use of every Workbench capability
After the engagement
Prioritize the uncovered failure cases, implement approved regression tests and release conditions, assign owners, and retest the changed controls.
Optional deliverables: Retest Record, Program work package. Selected according to engagement scope.
Supporting Workbench capabilities
Selected according to scope.
The engagement outcome and evidence are the deliverable. These AI Security Workbench capabilities support the work where they add value; their presence here does not mean every engagement uses all of them.
Code Scanner
Code-derived validation cases and release-condition candidates.
Runtime Trace
Capture supported runtime observations for failure analysis and regression evidence.
RAG Security Testing
Test retrieval authorization, provenance, context integrity, and indirect-injection cases where retrieval is in scope.
Program Blueprint
Translate approved release conditions, owners, evidence requirements, and regression expectations into program work.
Delivery / subject-matter leads
Delivery leads are confirmed during scoping.
Adjacent services
Choose by the decision you need to make.
HARDEN
Agentic Workflow Security & Hardening
Where can delegated authority, tool use, identities, permissions, approvals, or side effects exceed the intended boundary?
RETRIEVAL
RAG Security Testing
Are retrieval authorization, tenant boundaries, provenance, context integrity, and indirect-injection controls holding?
ATTACK
AI Red Team & Adversarial Testing
Which realistic abuse and adversarial behaviors can be reproduced, and what controls break them?