Before the Arrest: Benchmarking LLMs on Criminal Profiling from Incomplete Evidence
This repository hosts the PIJ benchmark (code, prompts, and dataset access). PIJ evaluates large language models on pre-arrest criminal profiling from incomplete evidence, together with crime-process reconstruction and sentence prediction, using 2,500 real homicide cases from five countries.
Content warning. Case materials describe violent crime.
The final version of the dataset is being prepared and is not released yet. We will update this repository as soon as the release is ready.
If you need access before the public release for a legitimate research purpose, please email the authors.
PIJ is a diagnostic and bias-auditing benchmark. It is not a tool for identifying suspects, ranking investigative leads, allocating police resources, or making charging or sentencing decisions.
Prohibited: operational investigation, commercial use of the dataset, and law-enforcement deployment.
| Material | License |
|---|---|
| Dataset (upon release) | CC BY-NC 4.0 plus a signed data-use agreement. The DUA will additionally prohibit commercial use and law-enforcement use. |
- 2026-08 Repository created. Dataset and code will follow after review.