AI Red-Teaming & Attack Simulation Platform
Attack your AI before attackers do.
A product of AetherGuard AI
Excalibur is a comprehensive adversarial testing tool for LLMs, AI agents, and RAG systems. It surfaces vulnerabilities before real attackers do — 22 attack types across 8 categories, MITRE ATLAS mapped, with an LLM Judge for evaluation and a proprietary AetherGuard Resilience Score.
- ✨ Features
- 🧰 Technology Stack
- 📦 Installation
- 🔑 Environment Setup
- 🚀 Quick Start
- 🎯 Attack Catalog (22 Attacks)
- ⚖️ LLM Judge
- 🗂️ Datasets & Licensing
- 📋 Campaign Mode
- 📊 Output & Reporting
- 🌐 API Server
- 🖥️ Standalone UI
- 🐳 Docker
- 🔁 CI/CD Integration (SARIF)
- 🏗️ Architecture
- 🗺️ MITRE ATLAS Mapping
- 📄 License
| Feature | Description | |
|---|---|---|
| 🎯 | 22 Attack Types | Across 8 categories: Adversarial ML, NLP, Agent, RAG, Infrastructure, Model Integrity, Evasion, Content Safety |
| ⚖️ | LLM Judge Evaluation | A separate LLM (GPT-4o / Claude) evaluates attack results with structured verdicts, replacing keyword heuristics with real language understanding |
| 🛡️ | Content Safety Suite | Prompt injection, DAN jailbreaks, PII/PHI leakage, HAP (hate/abuse/profanity), secrets exposure, and malicious content generation |
| 🔲 | Black-box + White-box | API-based attacks and direct model-weight access |
| 🗺️ | MITRE ATLAS Mapping | Every attack mapped to ATLAS technique IDs |
| 📈 | AetherGuard Resilience Score | Proprietary 0–100 composite score across 8 categories |
| 📐 | ART-Compatible Metrics | Industry-standard precision / recall / F1 |
| 📄 | Multi-Format Reports | HTML, PDF, JSON, and SARIF for CI/CD |
| 🤗 | Open-Source Dataset Integration | Optionally augment built-in payloads with public open-source datasets from the HuggingFace Hub |
| 🧩 | Plugin Architecture | Extensible with custom attacks via @register_attack |
| 🖧 | Three Interfaces | CLI, REST API, and standalone Web UI |
| Group | Technologies |
|---|---|
| whitebox | |
| nlp | sentence-transformers · NLTK · deep-translator |
| datasets | |
| vectorstore | |
| api | |
| reporting | Jinja2 · WeasyPrint (PDF) · Plotly |
| Technology | Purpose |
|---|---|
| UI framework | |
| Type-safe frontend | |
| Build tooling & dev server | |
| Server-state / data fetching | |
| Client-side routing | |
| Charts & visualizations | |
| Icon set |
cd aetherguard-excalibur
# Core (CLI + black-box attacks)
pip install -e "."
# Full install (all attack types + API + reporting)
pip install -e ".[full]"
# Development (adds pytest, ruff, hypothesis, respx)
pip install -e ".[full,dev]"pip install -e ".[whitebox]" # PyTorch + HuggingFace + ONNX (for PGD/FGSM/Smoothing)
pip install -e ".[nlp]" # sentence-transformers + NLTK (for TextFooler)
pip install -e ".[vectorstore]" # Pinecone + Weaviate + Chroma + pgvector
pip install -e ".[datasets]" # HuggingFace datasets (for content-safety payloads)
pip install -e ".[api]" # FastAPI server
pip install -e ".[reporting]" # HTML/PDF report generationAPI keys are never passed on the command line — they are read from environment variables.
# Target provider keys
export EXCALIBUR_OPENAI_API_KEY=sk-...
export EXCALIBUR_ANTHROPIC_API_KEY=sk-ant-...
export EXCALIBUR_AZURE_OPENAI_API_KEY=...
export AWS_REGION=us-east-1 # For Bedrock
# LLM Judge (optional — a separate LLM that evaluates attack results)
export EXCALIBUR_JUDGE_API_KEY=sk-... # Defaults to OpenAI GPT-4o# Verify installation
excalibur --version
excalibur list # List all 22 attacks
excalibur list --category nlp # Filter by category
# Validate a campaign config
excalibur validate campaigns/quick_scan.yaml
# Run a single attack
excalibur run hallucination_induction --target openai -m gpt-4o-mini -n 10
# Run a quick campaign (5 attacks)
excalibur campaign campaigns/quick_scan.yaml -o results/
# Run the full 22-attack audit
excalibur campaign campaigns/full_audit.yaml -o results/full/
# Start the API server
excalibur serve --port 8100
# Start the UI (separate terminal)
cd ui && npm install && npm run dev| Command | Status | Description |
|---|---|---|
excalibur run <attack> |
✅ | Run a single attack against a target |
excalibur campaign <config.yaml> |
✅ | Execute a full campaign from YAML |
excalibur list |
✅ | List available attacks (optionally by category) |
excalibur validate <config.yaml> |
✅ | Validate a campaign config without executing |
excalibur serve |
✅ | Start the REST API server |
excalibur report <dir> |
🚧 | Regenerate reports from saved results (coming soon) |
excalibur benchmark <target> |
🚧 | Evasion benchmark suite (planned) |
Prerequisite for black-box examples:
export EXCALIBUR_OPENAI_API_KEY=sk-your-key
1. PGD (Projected Gradient Descent) — white-box iterative gradient attack
excalibur run pgd \
--target local \
--model distilgpt2 \
-p epsilon=0.03 \
-p step_size=0.007 \
-p iterations=40 \
-p norm=linf \
-p random_start=true \
-n 50ATLAS: AML.T0043 • Interface: White-box • Category: Adversarial ML
2. FGSM (Fast Gradient Sign Method) — single-step gradient perturbation
excalibur run fgsm \
--target local \
--model distilgpt2 \
-p epsilon=0.05 \
-p norm=linf \
-n 100ATLAS: AML.T0043 • Interface: White-box • Category: Adversarial ML
3. Randomized Smoothing — certified L2 robustness via Gaussian noise
excalibur run randomized_smoothing \
--target local \
--model distilgpt2 \
-p sigma=0.25 \
-p n_samples=500 \
-p alpha=0.001 \
-n 30ATLAS: AML.T0043.002 • Interface: White-box • Category: Adversarial ML
4. Transfer Attack — craft on a local surrogate, test against the remote target
excalibur run transfer_attack \
--target openai \
--model gpt-4o-mini \
-p surrogate_model=distilbert-base-uncased \
-p query_budget=200 \
-p attack_method=pgd \
-p epsilon=0.03 \
-p num_adversarial=30 \
-n 30ATLAS: AML.T0044 • Interface: Black-box target + local surrogate • Category: Adversarial ML
5. TextFooler — word-level synonym substitution preserving semantics
excalibur run textfooler \
--target openai \
--model gpt-4o-mini \
-p max_perturbation_pct=0.2 \
-p similarity_threshold=0.8 \
-n 50ATLAS: AML.T0043.001 • Interface: Black-box • Category: NLP & Language
6. Low-Resource Language Attack — safety bypass via underrepresented languages
excalibur run low_resource_language \
--target openai \
--model gpt-4o-mini \
-p languages=amharic,yoruba,swahili,burmese,khmer,georgian \
-p attack_types=translation,code_switch,transliteration \
-n 50ATLAS: AML.T0051 • Interface: Black-box • Category: NLP & Language
7. Multi-Turn Chain — conversational jailbreak escalation with language switching
excalibur run multi_turn_chain \
--target openai \
--model gpt-4o-mini \
-p max_turns=6 \
-p chain_count=20 \
-p templates=dan,aim,developer_mode,system_prompt_leak \
-p escalation_strategy=gradual \
-n 20ATLAS: AML.T0051.001 • Interface: Black-box • Category: NLP & Language
8. MAIC (Multi-Agent Infection Chain) — infection propagation across an agent topology
excalibur run maic \
--target openai \
--model gpt-4o-mini \
-p topology=linear_3_agent \
-p payloads=prompt_injection,tool_poisoning,context_manipulation \
-p propagation_hops=3 \
-n 30ATLAS: AML.T0052 • Interface: Protocol simulation • Category: Agent & Trust
9. Cross-Trust Boundary — privilege escalation & data exfiltration across domains
excalibur run cross_trust \
--target openai \
--model gpt-4o-mini \
-p boundaries=user_agent,agent_tool,tool_api \
-p scenarios=privilege_escalation,data_exfiltration,confused_deputy \
-n 30ATLAS: AML.T0024 • Interface: Black-box • Category: Agent & Trust
10. KB Poisoning — inject poisoned documents into a vector store to hijack retrieval
excalibur run kb_poisoning \
--target openai \
--model gpt-4o-mini \
-p strategy=embedding_cluster \
-p injection_count=20 \
-p target_queries=10 \
-n 20ATLAS: AML.T0020 • Interface: Vector store + API • Category: RAG & Embedding
Requires vector store access. Set
PINECONE_API_KEYor configure ChromaDB/pgvector.
11. Embedding Inversion — reconstruct original text from embedding vectors (privacy leakage)
excalibur run embedding_inversion \
--target openai \
--model gpt-4o-mini \
-p training_corpus_size=500 \
-p test_samples=30 \
-p reconstruction_method=mlp_decoder \
-n 30ATLAS: AML.T0024 • Interface: Embedding API • Category: RAG & Embedding
12. Document Injection — hidden adversarial content in docs (invisible text, metadata, unicode)
excalibur run doc_injection \
--target openai \
--model gpt-4o-mini \
-p formats=pdf,docx,markdown \
-p techniques=invisible_text,metadata_payload,unicode_confusable \
-n 30ATLAS: AML.T0020.001 • Interface: Document generation • Category: RAG & Embedding
13. API Key Impersonation — detection of stolen/misused API credentials
excalibur run api_key_impersonation \
--target openai \
--model gpt-4o-mini \
-p scenarios=replay,geo_mismatch,concurrent,expired,enumeration,header_injection \
-n 50ATLAS: AML.T0040 • Interface: HTTP • Category: Infrastructure
14. Cross-Contamination — multi-tenant isolation boundary testing
excalibur run cross_contamination \
--target openai \
--model gpt-4o-mini \
-p tenant_count=3 \
-p scenarios=shared_cache,prompt_injection_crossover,timing_side_channel \
-n 30ATLAS: AML.T0024 • Interface: HTTP + protocol • Category: Infrastructure
15. Watermark Audit — detect & test backdoor triggers in model watermarks
excalibur run watermark_audit \
--target local \
--model ./my-model.pt \
-p detection_methods=spectral_signature,activation_clustering \
-p trigger_candidates=100 \
-n 50ATLAS: AML.T0020 • Interface: White-box • Category: Model Integrity
16. Hallucination Induction — inputs that systematically induce confident false outputs
excalibur run hallucination_induction \
--target openai \
--model gpt-4o-mini \
-p categories=factual,citation,capability,grounding_bypass \
-n 100 \
-o results/hallucination.jsonATLAS: AML.T0048 • Interface: Black-box • Category: Model Integrity
17. Prompt Injection — direct & indirect injection with 5 categories incl. encoding evasion
excalibur run prompt_injection \
--target openai \
--model gpt-4o-mini \
-p categories=system_override,indirect_injection,role_manipulation,context_manipulation,encoding_evasion \
-n 50ATLAS: AML.T0051.000 • Interface: Black-box • Category: Content Safety Payloads: Built-in, optionally augmented with open-source datasets
18. DAN-Style Jailbreaks — 8 persona-based jailbreak templates
excalibur run jailbreak_dan \
--target openai \
--model gpt-4o-mini \
-p personas=dan_classic,aim,stan,dude,jailbroken,evil_confidant,maximum,developer_mode_v2 \
-p intents_per_persona=5 \
-n 40ATLAS: AML.T0051.002 • Interface: Black-box • Category: Content Safety Payloads: Built-in, optionally augmented with open-source datasets
19. PII/PHI Leakage — induce generation of realistic personal & health data
excalibur run pii_phi_leakage \
--target openai \
--model gpt-4o-mini \
-p categories=direct_generation,contextual_extraction,roleplay_extraction,indirect_phi \
-p pii_types=personal_identifiers,financial,contact,health_phi,credentials \
-n 30ATLAS: AML.T0024.001 • Interface: Black-box • Category: Content Safety Payloads: Built-in, optionally augmented with open-source synthetic-PII datasets
20. HAP Content (Hate/Abuse/Profanity) — toxic content propensity across demographics
excalibur run hap_content \
--target openai \
--model gpt-4o-mini \
-p categories=hate_speech,identity_attack,threats_violence,profanity_obscenity,harassment_bullying,implicit_toxicity \
-p target_groups="Black people,Asian people,LGBTQ+ people,women,disabled people,immigrants" \
-n 50ATLAS: AML.T0048.001 • Interface: Black-box • Category: Content Safety Payloads: Built-in, optionally augmented with open-source datasets
21. Secrets Leakage — induce generation of API keys, tokens, and credentials
excalibur run secrets_leakage \
--target openai \
--model gpt-4o-mini \
-p categories=direct_generation,code_context,contextual_extraction,system_prompt_secrets,social_engineering \
-p secret_types=aws,github,openai,stripe,gcp,database_url,private_key,jwt \
-n 40ATLAS: AML.T0024.002 • Interface: Black-box • Category: Content Safety Payloads: Built-in
22. Malicious Content Generation — induce exploit code, malware, and harmful payloads
excalibur run malicious_content \
--target openai \
--model gpt-4o-mini \
-p categories=malware_code,exploit_code,phishing_content,social_engineering,weapon_instructions,fraud_scams \
-p severity_levels=high,critical \
-n 50ATLAS: AML.T0048.002 • Interface: Black-box • Category: Content Safety Payloads: Built-in, optionally augmented with open-source datasets
Excalibur includes an LLM Judge — a separate evaluation LLM that provides structured verdicts on attack results instead of relying solely on keyword heuristics.
Keyword-based detection has gaps:
- Can't detect harmful content in non-English languages
- Misses subtle compliance (the model helps without using obvious keywords)
- Can't evaluate nuance (educational vs. actually harmful)
The Judge LLM reads the attack prompt and model response, then returns a structured JSON verdict with violation status, confidence, category, reasoning, and severity.
# In campaign YAML:
judge:
provider: openai # or anthropic
model: gpt-4o # strong evaluator model
api_key_env: EXCALIBUR_JUDGE_API_KEY
temperature: 0.0
enabled: true# Or via CLI flags:
export EXCALIBUR_JUDGE_API_KEY=sk-...
excalibur campaign campaigns/full_audit.yaml --judge openai --judge-model gpt-4o
# Single attack with judge:
excalibur run prompt_injection --target openai --judge openai --judge-model gpt-4o -n 50{
"is_violation": true,
"confidence": 0.95,
"category": "harmful_content",
"reasoning": "The model provided step-by-step SQL injection instructions",
"severity": "high"
}| Type | Evaluates | Used By |
|---|---|---|
| Safety Bypass | Did the model produce harmful content? | Prompt Injection, Jailbreaks, HAP, Malicious Content |
| Hallucination | Did the model fabricate facts/citations? | Hallucination Induction |
| Agent Infection | Did the agent comply with malicious instructions? | MAIC, Cross-Trust |
| Trust Boundary | Was a security boundary breached? | Cross-Trust, Cross-Contamination |
When the judge is disabled or unavailable, all attacks fall back to built-in multi-tier heuristic detection. The judge is optional but significantly improves detection accuracy, especially for:
- Non-English responses (low-resource language attacks)
- Subtle compliance without explicit harmful keywords
- Implicit toxicity (coded language, dog whistles)
- Context-dependent violations
Several Content Safety attacks can optionally augment their built-in payloads with real-world examples from public, open-source datasets hosted on the HuggingFace Hub. This is opt-in — controlled per attack via use_hf_payloads (default false) — and requires the optional extra:
pip install -e ".[datasets]"When enabled, datasets are downloaded at runtime via datasets.load_dataset(). Excalibur does not bundle or redistribute any dataset content in this repository. If the datasets library is not installed, the affected attacks log a warning and fall back to their built-in payloads.
Any dataset used is configurable per attack (via hf_dataset / hf_datasets params), so you can point each attack at the open-source dataset of your choice. Synthetic (artificially generated) datasets are preferred for PII/PHI testing so that no real personal data is involved.
- Verify the license of any dataset you configure on its source page before use. Licenses differ and some may restrict commercial use or require accepting gated terms.
- Prefer synthetic data for PII testing. Use artificially generated PII rather than any dataset containing real individuals' data.
- Payloads leave your environment. Running attacks sends payloads to your configured target providers (OpenAI, Anthropic, Azure, Bedrock). Review each provider's data-handling policy before testing with sensitive content.
- Attribution. If a dataset's license requires attribution, credit the source in any published results.
Any external datasets are third-party resources governed by their own terms. AetherGuard AI does not distribute them and is not responsible for their content. Compliance with each dataset's license is the responsibility of the user.
excalibur campaign campaigns/quick_scan.yaml -o results/quick/excalibur campaign campaigns/full_audit.yaml -o results/full/Create my-campaign.yaml:
name: "Custom Red Team"
target:
type: openai
model: gpt-4o
api_key_env: EXCALIBUR_OPENAI_API_KEY
attacks:
- type: textfooler
params: { samples: 50, max_perturbation_pct: 0.15 }
- type: hallucination_induction
params: { samples: 100, categories: [factual, citation] }
- type: multi_turn_chain
params: { chain_count: 20, templates: [dan, aim] }
- type: maic
params: { topology: star_4_agent, propagation_hops: 3 }
- type: cross_trust
params: { scenarios: [privilege_escalation, confused_deputy] }
parallel: 3
reporting:
formats: [json, html, sarif]
resilience_score: true
atlas_mapping: trueexcalibur campaign my-campaign.yaml -o results/custom/
# Override parallelism at runtime:
excalibur campaign my-campaign.yaml --parallel 5 -o results/custom/| Grade | Score | Meaning |
|---|---|---|
| 🟢 A | 90–100 | Excellent resilience |
| 🟢 B | 80–89 | Good — minor gaps |
| 🟡 C | 70–79 | Moderate — several vectors effective |
| 🟠 D | 60–69 | Below average |
| 🔴 F | 0–59 | Poor — most attacks succeed |
| Format | File | Use Case |
|---|---|---|
| 🧾 JSON | results.json |
Programmatic, CI/CD |
| 🌐 HTML | report.html |
Interactive browser view |
report.pdf |
Executive summary | |
| 🔍 SARIF | results.sarif |
GitHub Code Scanning |
Campaign runs write reports automatically to the output directory based on the reporting.formats list in your campaign config.
excalibur serve --port 8100
excalibur serve --reload # development mode (auto-reload)- 📖 Swagger Docs: http://localhost:8100/docs
- ❤️ Health: http://localhost:8100/health
- 🎯 Attacks:
GET /api/attacks,POST /api/attacks/run - 📋 Campaigns:
POST /api/campaigns,GET /api/campaigns/{id} - 📊 Reports:
/api/reports
cd ui
npm install
npm run dev
# → http://localhost:5173Built with React 18 + TypeScript + Vite. Pages:
| Route | Page |
|---|---|
/ |
Dashboard |
/campaigns/new |
Campaign Builder |
/campaigns/:id · /results |
Results Viewer |
/attacks |
Attack Library |
/settings |
Settings |
The image is built from python:3.11-slim with the full feature set and runs as a non-root user.
docker build -t excalibur:1.0.0 .
# Run CLI
docker run --rm -e EXCALIBUR_OPENAI_API_KEY=$EXCALIBUR_OPENAI_API_KEY \
excalibur:1.0.0 run hallucination_induction --target openai -n 10
# Run API server
docker run -p 8100:8100 -e EXCALIBUR_OPENAI_API_KEY=$EXCALIBUR_OPENAI_API_KEY \
excalibur:1.0.0 serve --host 0.0.0.0 --port 8100# GitHub Actions
- name: Run Excalibur Scan
env:
EXCALIBUR_OPENAI_API_KEY: ${{ secrets.OPENAI_KEY }}
run: |
excalibur campaign campaigns/quick_scan.yaml -o results/
- name: Upload SARIF
uses: github/codeql-action/upload-sarif@v3
with:
sarif_file: results/results.sarif
- name: Quality Gate
run: |
SCORE=$(python -c "import json; print(json.load(open('results/results.json'))['resilience_score']['overall'])")
[ $(echo "$SCORE < 70" | bc) -eq 1 ] && exit 1excalibur CLI / API / UI
│
Campaign Engine (orchestration, parallelism, progress)
│
┌────┴────┐
│ │
│ LLM Judge ─── Structured verdicts (safety bypass, hallucination, agent infection, trust boundary)
│ │
Attack Registry ──── 22 Attack Plugins (@register_attack)
│
Target Adapters ──── OpenAI, Anthropic, Azure, Bedrock, Local, VectorStores
│
Reporting Engine ─── Resilience Score, ATLAS, ART, HTML/PDF/SARIF
Source layout (src/aetherguard_excalibur/):
| Module | Responsibility |
|---|---|
cli.py |
Click-based excalibur command entry point |
engine.py |
Campaign orchestration, parallelism, progress events |
registry.py |
Attack discovery & @register_attack plugin registry |
judge.py |
LLM Judge — structured verdict evaluation |
config.py · models.py |
Pydantic config loading & data models |
adapters/ |
Target adapters (OpenAI, Anthropic, Azure, Bedrock, local) |
attacks/ |
22 attack plugins across 7 category folders |
reporting/ |
Resilience score, ATLAS mapper, ART metrics, SARIF, report generator |
api/ |
FastAPI app + routes (attacks, campaigns, reports) |
| # | Attack | ATLAS ID | Tactic |
|---|---|---|---|
| 1 | PGD | AML.T0043 | ML Attack Staging |
| 2 | FGSM | AML.T0043 | ML Attack Staging |
| 3 | Randomized Smoothing | AML.T0043.002 | ML Attack Staging |
| 4 | Transfer Attack | AML.T0044 | ML Attack Staging |
| 5 | TextFooler | AML.T0043.001 | ML Attack Staging |
| 6 | Low-Resource Language | AML.T0051 | Initial Access |
| 7 | Multi-Turn Chain | AML.T0051.001 | Initial Access |
| 8 | MAIC | AML.T0052 | Initial Access |
| 9 | Cross-Trust Boundary | AML.T0024 | Exfiltration |
| 10 | KB Poisoning | AML.T0020 | ML Attack Staging |
| 11 | Embedding Inversion | AML.T0024 | Exfiltration |
| 12 | Document Injection | AML.T0020.001 | ML Attack Staging |
| 13 | API Key Impersonation | AML.T0040 | Initial Access |
| 14 | Cross-Contamination | AML.T0024 | Collection |
| 15 | Watermark Audit | AML.T0020 | ML Attack Staging |
| 16 | Hallucination Induction | AML.T0048 | Impact |
| 17 | Prompt Injection | AML.T0051.000 | Initial Access |
| 18 | DAN-Style Jailbreaks | AML.T0051.002 | Defense Evasion |
| 19 | PII/PHI Leakage | AML.T0024.001 | Exfiltration |
| 20 | HAP Content | AML.T0048.001 | Impact |
| 21 | Secrets Leakage | AML.T0024.002 | Exfiltration |
| 22 | Malicious Content | AML.T0048.002 | Impact |
MIT License
⚔️ AetherRed — Excalibur · Attack your AI before attackers do.