Skip to content

Latest commit

 

History

18 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Policy-aware Content Moderation

CI

A Python moderation service that classifies English text as Hateful, Offensive, or Clean and recommends allow, review, or block. It combines a fine-tuned classifier with policy-based LLM review, configurable routing, and evaluation tools.

The service supports two modes:

  • Cascade: a classifier makes the initial prediction. Score thresholds, input coverage, and audit sampling determine which cases receive LLM review.
  • Policy only: an LLM reviews every item against the written policy, without a classifier checkpoint.

Unresolved cases return review_required with action: review. The API provides recommendations; enforcement and a human-review queue belong to the integrating application.

Thesis connection

My master's thesis, Modeling Offensive Language as a Distinct Class for Hate Speech Detection (Kim, 2025), compared RoBERTa models under three-class and binary label schemes across Davidson and HateCheck-XR. It examined how distinguishing offensive language from hate speech affects classification and generalization.

This project extends that research into a moderation API, adding written-policy review and application-specific actions. The training code and checkpoint support retain the research workflow; policy alignment explains how the labels carry over.

Architecture

flowchart TD
    A[Text] --> B{Operating mode}
    B -->|cascade| C[Classifier and input coverage]
    C --> D[Score thresholds and audit sampling]
    D -->|automatic| G[Content label]
    D -->|review| E[Policy context and LLM adjudication]
    B -->|policy only| E
    E --> F{Valid supported output?}
    F -->|yes| G
    F -->|no| H[Review required]
    G --> I[Action policy: allow / review / block]
    H --> I
Loading

ModerationPipeline is shared by the API and evaluation tools. The LLM receives the full policy by default, with optional clause retrieval for comparison.

Component Implementation
Classifier PyTorch and Hugging Face Transformers; RoBERTa checkpoint loading, batch inference, label mapping, and token-coverage checks.
API FastAPI and Pydantic; validated requests and responses, configurable routing, and decision statistics.
Policy review OpenAI-compatible LLM client, structured verdict validation, and optional Chroma retrieval with sentence-transformer embeddings.
Evaluation Reproducible checkpoint and pipeline comparisons, per-class error analysis, confidence diagnostics, and red-team regression tests.
Delivery pytest, Ruff, GitHub Actions, and Docker packaging.

Content labels and application actions are configured separately:

Action policy Hateful Offensive Clean
general block review allow
open_discussion block allow allow
family block block allow

Quickstart

Python 3.12. The default demo runs offline with a stub classifier and mock LLM.

python -m venv .venv
# Windows PowerShell: .\.venv\Scripts\Activate.ps1
# macOS/Linux: source .venv/bin/activate
python -m pip install -r requirements-lock.txt
python -m uvicorn serving.app:app --host 127.0.0.1 --port 8000

Open the API documentation or submit a request:

Invoke-RestMethod http://127.0.0.1:8000/moderate -Method Post -ContentType application/json `
  -Body '{"text":"I enjoyed the community picnic today."}'

POST /moderate returns the label, action, routing reasons, policy citations, and model usage. GET /health identifies the active components and policy; GET /stats reports routing and decision counts.

Models and configuration

For policy-only mode, connect an OpenAI-compatible chat endpoint:

$env:MODERATION_MODE = 'policy_only'
$env:LLM_BASE_URL = 'http://localhost:11434/v1'
$env:LLM_MODEL = '<installed-model-name>'
$env:REQUIRE_REAL_MODELS = 'true'
python -m uvicorn serving.app:app --host 127.0.0.1 --port 8000

Provider authentication uses LLM_API_KEY when required. For cascade mode, install requirements-models.txt, set MODERATION_MODE=cascade, and supply a ternary classifier through MODEL_DIR or MODEL_CHECKPOINT. Checkpoints are stored externally. Checkpoint setup covers formats, label mapping, and preprocessing.

Settings are loaded from serving/config.json with environment overrides. The environment example provides a starting configuration.

Setting Options / default
MODERATION_MODE cascade (default), policy_only
ACTION_POLICY general (default), open_discussion, family
ROUTING_PROFILE balanced (default), strict, lean, or a profile path
ADJUDICATION_STRATEGY direct (default), cot, self_consistency
RETRIEVAL_MODE, RAG_TOP_K all (default) or top_k; 6 clauses
EMBEDDING_BACKEND hashed (default), sentence_transformers, auto
REQUIRE_REAL_MODELS Rejects mock components when true; default false

Routing profiles review every predicted Hateful item; strict also reviews every predicted Offensive item. The supplied thresholds are provisional. Design notes cover routing limits and failure handling.

Evaluation and results

Checkpoints across five seeds from the thesis were evaluated on all 3,855 HateCheck-XR cases, with mean macro-F1 0.378. The diagnostic results expose generalization errors and limits of confidence-based routing. Verification records the per-seed results and 243 passing tests. Real-LLM cascade effectiveness remains unevaluated.

The end-to-end runner compares classifier, cascade, and policy-only decisions on the same cases. The red-team harness adds text-obfuscation, prompt-injection, and jailbreak tests with benign controls. Default evaluation runs use a stub classifier and mock LLM to check software behavior.

python -m pytest
ruff check .
python -m evals.run_end_to_end --limit 36 --compare-retrieval
python -m redteam.run_redteam --gate

Evaluation methods cover configured-model comparisons, confidence analysis, and threshold selection. Research setup covers training and cross-dataset experiments. CI runs automated checks, generates reports, and checks the container build.

Deployment and credits

docker build -f docker/Dockerfile -t moderation-service .
docker run --rm -p 127.0.0.1:8000:8000 moderation-service

Deployment notes cover container configuration and an Azure Container Apps setup.

The research builds on Khurana et al.'s defverify and Hugging Face example code. Diagnostic sources: HateCheck (Röttger et al., 2021), Khurana et al.'s extension, and Davidson et al.. The diagnostic dataset contains the original offensive-language examples. Licensing is documented in LICENSE and third_party/APACHE-2.0.txt.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages