Human-led AI evaluation research system for observing model behaviour through binary evals, signal traces, and retained failures.
python sqlite openai model-evaluation ai-safety human-in-the-loop ai-research openai-api human-ai-interaction ai-evaluation evals llm-evaluation human-ai-collaboration human-ai-alignment ai-reliability research-models binary-evals
-
Updated
Sep 9, 2026 - Python