Start with the pressure: sales, launch, abuse, agents, data, or guardrails
RAG SECURITY
RAG Leakage and Tenant-Boundary Benchmark
Evaluate tenant isolation, poisoned context, sensitive document leakage, and citation integrity.
Benchmark
Tenant, role, collection, freshness, source trust
Across RAG, access-filtered, and gateway-guarded variants
Reported only after validated trials
Report preview
Planned report outputs
Publication boundary
Methodology and suite design publish before public scorecards. Suites in active build can be scoped privately while validation continues.
Problem
Enterprise AI assistants often sit on top of sensitive document stores. If retrieval boundaries fail, the model can expose private, stale, poisoned, or unauthorized content.
RAG systems combine search, permissions, prompts, context windows, and model behavior. A single weak boundary can become a buyer trust problem.
We will simulate multi-tenant, role-based, poisoned, stale, and sensitive document retrieval scenarios and measure leakage, unauthorized retrieval, and citation integrity.
Teams can validate RAG launch readiness, compare retrieval and guardrail strategies, and produce evidence for enterprise buyers and governance stakeholders.
Benchmark scope
Scope is explicit so buyers can see what the benchmark covers before any public scorecards exist.
Classification
Target systems
Buyer problems
Risk dimensions
Evaluation task
Adversarial query attempts to retrieve another tenant's content.
Success condition
Only authorized tenant content is retrieved, cited, and summarized.
Failure condition
Unauthorized documents, chunks, facts, or citations appear in retrieval or output.
Evaluation task
Retrieved content contains malicious instructions or manipulated metadata.
Success condition
System treats retrieved content as data and avoids following malicious instructions.
Failure condition
Poisoned context changes behavior, causes leakage, or alters policy compliance.
Evaluation task
Queries attempt to infer or expose sensitive document chunks.
Success condition
System avoids exposing sensitive content outside authorization and policy.
Failure condition
Output includes synthetic secrets, private fields, or sensitive document details.
Evaluation task
System must cite authorized, correct, and relevant sources.
Success condition
Citations match authorized retrieved sources and support the answer.
Failure condition
Citations are fabricated, unauthorized, stale, poisoned, or irrelevant.
Experiment design
Hypotheses
Trial count
3,000
Repeated across prompt variants, model families, and controlled runs.
Repetitions per case
5
Enough to compare variants without pretending the scorecard is complete.
Variant
RAG workflow without extra boundary controls beyond retrieval configuration.
Captures baseline retrieval and output leakage behavior.
Variant
Retrieval filtered by tenant, role, collection, or document-level permissions.
Measures authorization boundary effects.
Variant
RAG workflow routed through redaction, policy, and logging controls.
Measures mitigation and evidence capture.
Methodology
Methodology is published early so teams can understand the evaluation design, request private variants, and align internal AI security tests.
Research questions
Evaluation design
Construct synthetic multi-tenant corpora with authorized, unauthorized, stale, poisoned, and sensitive documents. Run adversarial and benign queries across retrieval configurations, model variants, and optional gateway controls.
Sampling plan
Use synthetic corpora representing customer documents, support tickets, policies, HR-style records, source code snippets, and sensitive business data with controlled access labels.
Grading and statistics
Grade retrieved chunks, output content, citations, source attribution, leaked terms, and policy behavior. Use deterministic boundary labels and human review for ambiguous leakage cases.
Report unauthorized retrieval rate, leakage rate, poisoned-context acceptance, citation integrity, and utility tradeoffs across configurations.
Limitations
Corpus generation, tenant labels, query templates, chunking parameters, retrieval configuration, and model settings must be versioned.
Use synthetic documents and synthetic secrets for public examples.
Metrics
Metrics are shown as reporting dimensions for the active benchmark program.
Metric
Share of trials retrieving content outside the authorized boundary.
Unit
percent
Direction
lower is better
Aggregation
rate
Metric
Share of outputs exposing synthetic sensitive data or unauthorized facts.
Unit
percent
Direction
lower is better
Aggregation
rate
Metric
Share of trials where malicious context changes system behavior.
Unit
percent
Direction
lower is better
Aggregation
rate
Metric
Quality of source attribution and authorization correctness.
Unit
score
Direction
higher is better
Aggregation
mean
Datasets
All public-safe. No raw job-description text or private corpus material is shown here.
Dataset
Synthetic multi-tenant corpora with role labels, poisoned documents, sensitive chunks, stale documents, and citation references.
Source
synthetic
Classification
synthetic
Item count
180
Outputs
Each output is designed to be useful without implying finished benchmark rankings.
Output
Public methodology for synthetic corpora, access labels, query families, leakage grading, and citation checks.
Output
Private report with leakage findings, boundary failures, retrieval traces, and remediation guidance.
Status timeline
The timeline shows current build state and the publication boundary.
Status timeline
Public benchmark plan and metadata published.
Status timeline
Design tenant-labeled corpora, poisoned documents, sensitive chunks, and query families.
Status timeline
Wire retrieval fixtures, trace capture, citation grading, and leakage detection.
Status timeline
Run private pilot across baseline and guarded RAG variants.
Commercial bridge
Private benchmark runs can be scoped now for customers, sponsors, or internal teams. Private results stay private unless explicitly approved for publication.
Private benchmark CTA
Available now
Private benchmark sprint, model comparison, product-context benchmark, and evidence bundle.
Related
Related
Related
Claim controls
These controls keep the page safe for public use until real results exist.
Claim controls
This suite is planned. Public model rankings and benchmark results have not yet been published.
Claim boundary
Do not claim