[NLPCC 2024] Shared Task 10: Regulating Large Language Models
-
Updated
Jun 12, 2024
[NLPCC 2024] Shared Task 10: Regulating Large Language Models
Detoxifying Online Discourse: A Guided Response Generation Approach for Reducing Toxicity in User-Generated Text
This reposity contains the source code of the ACL'25 paper "Contrastive Perplexity for Controlled Generation: An Application in Detoxifying Large Language Models". Paper abstract: "The generation of undesirable and factually incorrect content of large language models poses a significant challenge and remains largely an unsolved issue. This pap...
U of T 2025 Winter MIE1517 Final project
Fine-tuning FLAN-T5 with PPO and PEFT to generate less toxic text summaries. This notebook leverages Meta AI's hate speech reward model and utilizes RLHF techniques for improved safety.
sNeuron-TST adapted as a baseline for the ParaDeHate hate-speech detoxification paper (arXiv:2506.01484), with fixes to run on modern transformers (upstreamed in wenlai-lavine/sNeuron-TST#6).
Exploratory framework for span-guided multilingual text detoxification across English, Chinese, and Korean.
Code for studying when harmful-span guidance helps or hurts the toxicity-meaning trade-off in text detoxification.
PMLDL Assignment 1
This repository contains the source code for the frontend of the HealHub Health website: HealHub Health is a website that provides information about mental healthcare servces.
A multilingual text analysis system that performs sentiment analysis and toxicity detection with detoxified text generation.
Multilingual harmful-span detection and controlled detoxification across English, Chinese, and Korean.
To associate your repository with the detoxification topic, visit your repo's landing page and select "manage topics."