A python library built on top of Inspect AI to support Control research and evaluations, by Redwood and EquiStamp
-
Updated
Oct 8, 2026 - Python
A python library built on top of Inspect AI to support Control research and evaluations, by Redwood and EquiStamp
Kernel-enforced authority and runtime security for Linux systems, autonomous software, and AI agents.
Research developing AI control protocols using task decomposition
🚦🗺️ UrbanFlow AI is a web app for generating 3D traffic simulations from real OpenStreetMap areas. It builds SUMO scenarios, runs microscopic vehicles, pedestrians, buses and trams, edits road events, controls real traffic lights with TraCI, trains JSON AI policies, saves models, and shows live metrics, charts, and notebooks.
An intelligent traffic management system that dynamically adjusts highway lane configurations using AI-powered congestion detection and a movable median barrier.
Replayable agent trajectories reconstructed from the July 2026 frontier-lab evaluation-containment failures, plus benign controls for measuring false-positive cost.
Python client for Aegis — stabilize AI systems instantly with a simple API call.
Judge-first framework where LLM outputs must converge under explicit, adversarial oracles.
In-depth exploration of Large Language Models (LLMs), their potential biases, limitations, and the challenges in controlling their outputs. It also includes a Flask application that uses an LLM to perform research on a company and generate a report on its potential for partnership opportunities.
Kho lưu trữ này chứa tài liệu, bài tập, và mã nguồn liên quan đến môn Trí tuệ nhân tạo trong điều khiển. Môn học tập trung vào ứng dụng AI trong các hệ thống điều khiển tự động, bao gồm lý thuyết và thực hành.
Can a monitor catch a hidden goal by reading a model's chain of thought? Eval harness: gpt-oss-20b under vLLM in single-GPU PBS jobs; generated code graded in a time-limited subprocess on public and hidden tests; monitor sampled 8 times (bimodal at temperature 0). Hypothesis (monitoring degrades in lower-resource languages) failed; null published.
AI Governance — human-controlled code injection with oversight
Benchmark for detecting insider threats by AI agents in a simulated frontier AI lab
An attacker × monitor factorial in ControlArena: which model you pick as your trusted monitor matters more than its capability tier.
White-box detection of collusion in an untrusted monitor: model organisms, linear probes, and a control evaluation that prices what they buy.
Independent AI governance and control standard
Lichtarbeit
v0.8 preview · 2,288 multi-step agent trajectories · 513 hand-authored gold + 1,775 provenance-flagged augmented. A per-step benchmark that scores whether a verifier catches drift inside an agent's trajectory — not whether a prompt is harmful.
GG Tank Watch - frozen public-information archive of a resolved May 2026 chemical emergency. Conduit-only design; responsible-AI safety patterns enforced in code and tests.
Slides, drawing code, original character art and sources for a 15-minute introduction to AI control and multi-agent systems, in English and Spanish. Running example: the July 2026 OpenAI–Hugging Face incident. SPAR Fall 2026.
To associate your repository with the ai-control topic, visit your repo's landing page and select "manage topics."