Unitary's unbiased-toxic-roberta Fixes the Bias Flaw Killing AI Moderation
A RoBERTa classifier from Unitary scores comments across seven toxicity and identity-attack labels while actively minimizing demographic bias during training.
- Unitary's unbiased-toxic-roberta trends with 953k downloads for bias-aware toxicity classification.
- RoBERTa-base returns 7 labels including toxicity, threat, insult, identity_attack, sexual_explicit.
- Trained on Jigsaw Unintended Bias (Civil Comments) to minimize identity-mention false positives.
- Scores 93.74 AUC vs 94.73 Kaggle leader, without ensembles.
- Install via
pip install detoxifyor load through Transformers pipeline, Apache-2.0. - Hub weights lag the GitHub repo; use repo for latest checkpoints.
At publication, the model card reports nearly one million downloads and more than 3,500 likes for Unitary’s unbiased-toxic-roberta. The 125-million-parameter English classifier addresses a common moderation failure: neutral identity terms can receive high toxicity scores because those terms correlate with abuse in training data. Developers get an open, locally runnable baseline with category-level outputs and an explicit bias-reduction objective.
The model is the unbiased checkpoint from the Detoxify repository, which provides models for the original, unintended-bias, and multilingual Jigsaw Toxic Comment challenges. Laura Hanu developed the project at Unitary using PyTorch Lightning and Hugging Face Transformers. The repository and model are available under the Apache-2.0 license.
Seven harm scores from one pass
The checkpoint emits independent, probability-like scores between 0 and 1 for seven moderation categories. Applications can inspect every score, apply category-specific thresholds, or combine the outputs with other moderation signals.
| Group | Labels | What they cover |
|---|---|---|
| Overall toxicity | toxicity, severe_toxicity |
Broad harmful-language signals and more extreme cases |
| Abusive language | obscene, threat, insult |
Profanity, threatened harm, and direct abuse |
| Targeted or explicit content | identity_attack, sexual_explicit |
Identity-based attacks and sexually explicit language |
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.