It thinks it can
catch you.
A model watches everything you type and scores it across six categories — PROFANITY · INSULT · HARASSMENT · SEXUAL · THREAT · HATE. Your job: slip something past it.
Pick your angle
Pick a targeted challenge and write your attempt. Watch the model score it live, gauge by gauge.
A real, anonymized submission lands in front of you. Label it. Disagree with the model and you've found something worth retraining on.
Two submissions face off. Pick the more toxic one, then see if the model calls it the same way.
One constrained prompt, every day. Today it might be obfuscation. Tomorrow, code-mixing.
The rules
Every submission carries two separate, explicit consent flags — training use (required) and public display (optional, off by default). The leaderboard never shows raw text, only a score. Don't target real, identifiable people — write about fictional or generic targets only. Full details: how your data is used.