Skip to content

Repository files navigation

Modeling Offensive Language as a Distinct Class for Hate Speech Detection

This repository is part of my master's thesis project,"Modeling Offensive Language as a Distinct Class for Hate Speech Detection" (Kim, 2025), supervised by Dr. Antske Fokkens and Dr. Hennie van der Vliet. The project explored how modeling offensive (but not hateful) language as a distinct class impacts the task of detection of hate speech. Using a ternary classification scheme (Hateful, Offensive, Clean), I fine-tuned and evaluated a RoBERTa-base model in the full three-class setup and in binary variants where two classes are merged or the offensive class is removed (Hate vs. Non-hate, Non-clean vs. Clean, and Hate vs. Clean). The code used in this study includes my modifications and extensions of Khurana et al. (2025)'s code.

HateCheck-XR

In the project, to probe model behavior beyond set-internal performance, I revised both HateCheck (Röttger et al., 2021) and an existing extension by Khurana et al. (2025), aligning them with the ternary system by re-annotating them and correcting errors present in the extension. The resulting dataset, HateCheck-XR (3,855 cases, 37 functionalities), is available in this repository at datasets/hatecheck-xr/hatecheck-xr.csv (semicolon-separated).

Folder structure

Project
├─ configs/
│  ├─ example.json           # quick smoke-test config
│  ├─ train/example.json     # training config
│  └─ test/example.json      # evaluation config with {seed} placeholders
├─ datasets/
│  ├─ davidson/              # place the prepared Davidson HF dataset here (not distributed)
│  └─ hatecheck-xr/hatecheck-xr.csv
├─ hs_generalization/
│  ├─ __init__.py
│  ├─ modes.py           # class schemes (3class / hate_nonhate / nonclean_clean / hate_clean) & logit projections
│  ├─ train.py           # fine-tuning pipeline
│  ├─ test.py            # unified evaluator (Davidson test split & HateCheck-XR)
│  ├─ run_many.py        # runs test.py across seeds/checkpoints
│  └─ utils.py
├─ scripts/
│  └─ create_hf_dataset.py   # builds the HF dataset 
├─ outputs/                  # create the folder and checkpoints of the fine-tuned model on your own
├─ third_party
│  └─  APACHE-2.0.txt
├─ LICENSE
├─ README.md
├─ Thesis_Areum.pdf          
├─ requirements.txt
└─ setup.py

Label encoding used throughout: 0 = hateful, 1 = offensive, 2 = clean.

Set-up

Set up the environment like the following:

# Create environment.
conda create -n hs-generalization python=3.9
conda activate hs-generalization

# Install packages.
pip install -e .
pip install -r requirements.txt

Data preparation

The Davidson dataset is not redistributed here. Download labeled_data.csv from the official repository and build the HuggingFace dataset. For example:

python scripts/create_hf_dataset.py -n davidson -p path/to/labeled_data.csv -o datasets/davidson -s "[0.8, 0.1, 0.1]"

Training

Create a config file (see configs/train/example.json, which contains the hyperparameters used in the thesis) and run:

python -m hs_generalization.train -c configs/train/example.json

The training mode is set in the config via task.train_mode: one of 3class, hate_nonhate, nonclean_clean, hate_clean. The size of the classification head is derived from the mode automatically. A checkpoint is saved after each epoch as seed{seed}_{model_name}_{epoch}.pt; per the thesis, select the checkpoint with the best validation macro F1 per seed.

Evaluation

Create a config file (see configs/test/example.json) and run like the example:

# running based on a single seed/checkpoint
python -m hs_generalization.test -c configs/test/example.json --dataset davidson --eval-mode 3class --train-mode 3class --seed 5 --checkpoint "outputs/davidson/RoBERTa-base/3class/seed5_RoBERTa-base_7.pt"

# evaluating the same checkpoint on HateCheck-XR
python -m hs_generalization.test -c configs/test/example.json --dataset hatecheck_xr --eval-mode 3class --train-mode 3class --seed 5 --checkpoint "outputs/davidson/RoBERTa-base/3class/seed5_RoBERTa-base_7.pt" --hatecheck-csv datasets/hatecheck-xr/hatecheck-xr.csv

# If you want to run multiple seeds/checkpoints at once, use run_many.py:
python -m hs_generalization.run_many ^
  -c configs/test/example.json ^
  --dataset hatecheck_xr ^
  --eval-mode 3class ^
  --train-mode 3class ^
  --seeds 7,222,550,999,3111 ^
  --ckpt-pattern "outputs/davidson/RoBERTa-base/3class/seed{seed}_*.pt" ^
  --hatecheck-csv datasets/hatecheck-xr/hatecheck-xr.csv

--train-mode describes the label space the checkpoint was trained on; --eval-mode describes how the ground truth is scored.

Update (2026 July):

I extended my Master's thesis on hate speech detection into a production-style content-moderation service: Available at this repository.

About

Study the impact of modeling a distinct class for offensive language in hate speech detection

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages