Unitary's unbiased-toxic-roberta Fixes the Bias Flaw Killing AI Moderation

A RoBERTa classifier from Unitary scores comments across seven toxicity and identity-attack labels while actively minimizing demographic bias during training.

·
·
Unitary's unbiased-toxic-roberta Fixes the Bias Flaw Killing AI ModerationPRO
  • Unitary's unbiased-toxic-roberta trends with 953k downloads for bias-aware toxicity classification.
  • RoBERTa-base returns 7 labels including toxicity, threat, insult, identity_attack, sexual_explicit.
  • Trained on Jigsaw Unintended Bias (Civil Comments) to minimize identity-mention false positives.
  • Scores 93.74 AUC vs 94.73 Kaggle leader, without ensembles.
  • Install via pip install detoxify or load through Transformers pipeline, Apache-2.0.
  • Hub weights lag the GitHub repo; use repo for latest checkpoints.

At publication, the model card reports nearly one million downloads and more than 3,500 likes for Unitary’s unbiased-toxic-roberta. The 125-million-parameter English classifier addresses a common moderation failure: neutral identity terms can receive high toxicity scores because those terms correlate with abuse in training data. Developers get an open, locally runnable baseline with category-level outputs and an explicit bias-reduction objective.

The model is the unbiased checkpoint from the Detoxify repository, which provides models for the original, unintended-bias, and multilingual Jigsaw Toxic Comment challenges. Laura Hanu developed the project at Unitary using PyTorch Lightning and Hugging Face Transformers. The repository and model are available under the Apache-2.0 license.

Seven harm scores from one pass

The checkpoint emits independent, probability-like scores between 0 and 1 for seven moderation categories. Applications can inspect every score, apply category-specific thresholds, or combine the outputs with other moderation signals.

Group Labels What they cover
Overall toxicity toxicity, severe_toxicity Broad harmful-language signals and more extreme cases
Abusive language obscene, threat, insult Profanity, threatened harm, and direct abuse
Targeted or explicit content identity_attack, sexual_explicit Identity-based attacks and sexually explicit language

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads