A Python moderation service that classifies English text as Hateful, Offensive, or Clean and recommends allow, review, or block. It combines a fine-tuned classifier with policy-based LLM review, configurable routing, and evaluation tools.
The service supports two modes:
- Cascade: a classifier makes the initial prediction. Score thresholds, input coverage, and audit sampling determine which cases receive LLM review.
- Policy only: an LLM reviews every item against the written policy, without a classifier checkpoint.
Unresolved cases return review_required with action: review. The API provides
recommendations; enforcement and a human-review queue belong to the integrating
application.
My master's thesis, Modeling Offensive Language as a Distinct Class for Hate Speech Detection (Kim, 2025), compared RoBERTa models under three-class and binary label schemes across Davidson and HateCheck-XR. It examined how distinguishing offensive language from hate speech affects classification and generalization.
This project extends that research into a moderation API, adding written-policy review and application-specific actions. The training code and checkpoint support retain the research workflow; policy alignment explains how the labels carry over.
flowchart TD
A[Text] --> B{Operating mode}
B -->|cascade| C[Classifier and input coverage]
C --> D[Score thresholds and audit sampling]
D -->|automatic| G[Content label]
D -->|review| E[Policy context and LLM adjudication]
B -->|policy only| E
E --> F{Valid supported output?}
F -->|yes| G
F -->|no| H[Review required]
G --> I[Action policy: allow / review / block]
H --> I
ModerationPipeline is shared by the API and evaluation tools. The LLM receives the full policy by default, with optional clause retrieval for comparison.
| Component | Implementation |
|---|---|
| Classifier | PyTorch and Hugging Face Transformers; RoBERTa checkpoint loading, batch inference, label mapping, and token-coverage checks. |
| API | FastAPI and Pydantic; validated requests and responses, configurable routing, and decision statistics. |
| Policy review | OpenAI-compatible LLM client, structured verdict validation, and optional Chroma retrieval with sentence-transformer embeddings. |
| Evaluation | Reproducible checkpoint and pipeline comparisons, per-class error analysis, confidence diagnostics, and red-team regression tests. |
| Delivery | pytest, Ruff, GitHub Actions, and Docker packaging. |
Content labels and application actions are configured separately:
| Action policy | Hateful | Offensive | Clean |
|---|---|---|---|
general |
block | review | allow |
open_discussion |
block | allow | allow |
family |
block | block | allow |
Python 3.12. The default demo runs offline with a stub classifier and mock LLM.
python -m venv .venv
# Windows PowerShell: .\.venv\Scripts\Activate.ps1
# macOS/Linux: source .venv/bin/activate
python -m pip install -r requirements-lock.txt
python -m uvicorn serving.app:app --host 127.0.0.1 --port 8000Open the API documentation or submit a request:
Invoke-RestMethod http://127.0.0.1:8000/moderate -Method Post -ContentType application/json `
-Body '{"text":"I enjoyed the community picnic today."}'POST /moderate returns the label, action, routing reasons, policy citations, and
model usage. GET /health identifies the active components and policy;
GET /stats reports routing and decision counts.
For policy-only mode, connect an OpenAI-compatible chat endpoint:
$env:MODERATION_MODE = 'policy_only'
$env:LLM_BASE_URL = 'http://localhost:11434/v1'
$env:LLM_MODEL = '<installed-model-name>'
$env:REQUIRE_REAL_MODELS = 'true'
python -m uvicorn serving.app:app --host 127.0.0.1 --port 8000Provider authentication uses LLM_API_KEY when required. For cascade mode, install
requirements-models.txt, set MODERATION_MODE=cascade, and supply a ternary
classifier through MODEL_DIR or MODEL_CHECKPOINT. Checkpoints are stored
externally. Checkpoint setup covers formats, label mapping,
and preprocessing.
Settings are loaded from serving/config.json with environment overrides. The environment example provides a starting configuration.
| Setting | Options / default |
|---|---|
MODERATION_MODE |
cascade (default), policy_only |
ACTION_POLICY |
general (default), open_discussion, family |
ROUTING_PROFILE |
balanced (default), strict, lean, or a profile path |
ADJUDICATION_STRATEGY |
direct (default), cot, self_consistency |
RETRIEVAL_MODE, RAG_TOP_K |
all (default) or top_k; 6 clauses |
EMBEDDING_BACKEND |
hashed (default), sentence_transformers, auto |
REQUIRE_REAL_MODELS |
Rejects mock components when true; default false |
Routing profiles review every predicted Hateful item; strict also reviews every
predicted Offensive item. The supplied thresholds are provisional.
Design notes cover routing limits and failure handling.
Checkpoints across five seeds from the thesis were evaluated on all 3,855 HateCheck-XR cases, with mean macro-F1 0.378. The diagnostic results expose generalization errors and limits of confidence-based routing. Verification records the per-seed results and 243 passing tests. Real-LLM cascade effectiveness remains unevaluated.
The end-to-end runner compares classifier, cascade, and policy-only decisions on the same cases. The red-team harness adds text-obfuscation, prompt-injection, and jailbreak tests with benign controls. Default evaluation runs use a stub classifier and mock LLM to check software behavior.
python -m pytest
ruff check .
python -m evals.run_end_to_end --limit 36 --compare-retrieval
python -m redteam.run_redteam --gateEvaluation methods cover configured-model comparisons, confidence analysis, and threshold selection. Research setup covers training and cross-dataset experiments. CI runs automated checks, generates reports, and checks the container build.
docker build -f docker/Dockerfile -t moderation-service .
docker run --rm -p 127.0.0.1:8000:8000 moderation-serviceDeployment notes cover container configuration and an Azure Container Apps setup.
The research builds on Khurana et al.'s defverify and Hugging Face example code. Diagnostic sources: HateCheck (Röttger et al., 2021), Khurana et al.'s extension, and Davidson et al.. The diagnostic dataset contains the original offensive-language examples. Licensing is documented in LICENSE and third_party/APACHE-2.0.txt.