Open-source, 100% reproducible AI Agent Runtime Security Benchmark & Sandbox Environment (RFC-010 Draft Protocol).
-
Updated
Oct 8, 2026 - Python
Open-source, 100% reproducible AI Agent Runtime Security Benchmark & Sandbox Environment (RFC-010 Draft Protocol).
Adversarial security benchmark for agent authorization: does a compromised agent's policy-violating proposal become an unauthorized external effect? 61 trials, nine families, an independent oracle, per-mechanism ablation, confidence intervals. 0 unauthorized effects in 61 attack trials (95% CI [0.0%, 5.9%]). Reproduction is partial.
Open, privacy-bounded assurance for AI agents: containment provenance, identity passports, authorization twins, OTel evidence, CI gates, and OSCAL.
Agent Memory Integrity: a conformance test suite that measures whether AI agent memory and checkpoint stores notice tampering. Sixteen measured rows across LangGraph (SQLite, Postgres, Redis), OpenAI Agents SDK, LlamaIndex, Letta, Mem0 and six integrity tools. IETF draft-khandelwal-bmwg-agent-memory-integrity.
Open-source benchmark for adversarial evidence attacks on LLM-based cybersecurity auditors, targeting ACM AsiaCCS 2027.
Local-first workbench to run, inspect, compare, report, and gate OpenAI Codex Security scans.
Open deterministic security tests for unsafe multi-agent handoffs and authority escalation.
Vendor-neutral benchmark measuring how MCP security proxies/gateways DEFEND against 22+ attack vectors — crosswalked to NIST AI RMF & OWASP LLM/Agentic Top 10. CI-gated, reproducible, DOI-cited. Submit your tool to the leaderboard.
Deterministic security benchmark for tool-using AI agents
FreightSkillBench is a reproducible benchmark for evaluating document-to-transaction integrity, prompt-injection risk, and security controls in AI-enabled shipping and logistics workflows.
21 passive security evaluation inputs: conventional appsec and exploratory hardware, firmware, PLC, RTOS, database and policy domains. No scans started.
Internal PyPI SCA precision and recall benchmark corpus
Production-grade microservices security benchmark featuring OWASP Top 10 logic exploits, automated remediation, custom Semgrep SAST rules, and CI/CD DevSecOps gates.
Security evaluation input | Database-engine source, SQL parser, storage, transaction and extension boundaries | production-control-no-vulnerability-claim
Commit-pinned cal.diy source corpus; Vybscan ground-truth oracle available in benchmark-results
Security evaluation input | OpenMP C/C++/Fortran data races and race-free controls | labeled-positive-negative
ReplayBench-IoT: reproducible IoT replay-defense benchmark with Monte Carlo sweeps, CI, static demo, and hardware-validation artifacts.
Reproducible intentionally vulnerable JavaScript/TypeScript SAST accuracy benchmark
Internal npm SCA precision and recall benchmark corpus
TURNCOAT: an independent, reproducible prompt-injection benchmark for AI coding agents. Open corpus + methodology, responsible disclosure.
To associate your repository with the security-benchmark topic, visit your repo's landing page and select "manage topics."