# Discourse AI - Toxicity

**URL:** https://meta.discourse.org/t/discourse-ai-toxicity/259215
**Category:** Site Management
**Tags:** ai, ai-toxicity
**Created:** [April 24, 2023, 7:39pm UTC](https://meta.discourse.org/t/discourse-ai-toxicity/259215 "2023-04-24T19:39:50Z")
**Posts on this page:** 10
**Page:** 1

<div class="post-metadata">

### Author: ![Discourse](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/discourse/32/148734_2.png) [@Discourse](https://meta.discourse.org/u/Discourse)
#### Post date: [April 24, 2023, 7:39pm UTC](https://meta.discourse.org/t/discourse-ai-toxicity/259215/1 "2023-04-24T19:39:50Z")

</div>

> 🔖 This topic covers the configuration of the Toxicity feature of the [Discourse AI](https://meta.discourse.org/t/discourse-ai/259214?slient=true) plugin.
> 
> 🙋 Required user level: Administrator

The Toxicity modules can automatically classify the toxicity score of every new post and chat message in your Discourse instance. You can also enable automatic flagging of content that crosses a threshold.

Classifications are stored in the database, so you can enable the plugin and use [Data Explorer](https://meta.discourse.org/t/32566?silent=true) for reports of the classification happening for new content in Discourse immediately. We will soon ship some default [Data Explorer](https://meta.discourse.org/t/32566?silent=true) queries with the plugin to make this easier.

### Settings

- **ai\_toxicity\_enabled** : Enables or disables the module

- **ai\_toxicity\_inference\_service\_api\_endpoint** : URL where the API is running for the toxicity module. If you are using CDCK hosting this is automatically handled for you. If you are self-hosting check the [self-hosting guide](https://meta.discourse.org/t/discourse-ai-self-hosted-guide/259598).

- **ai\_toxicity\_inference\_service\_api\_key** : API key for the toxicity API configured above. If you are using CDCK hosting this is automatically handled for you. If you are self-hosting check the [self-hosting guide](https://meta.discourse.org/t/discourse-ai-self-hosted-guide/259598).

- **ai\_toxicity\_inference\_service\_api\_model** : ai\_toxicity\_inference\_service\_api\_model: We offer three different models: `original`, `unbiased`, and `multilingual`. `unbiased` is recommended over `original` because it’ll try not to carry over biases introduced by the training material into the classification. For multilingual communities, the last model supports Italian, French, Russian, Portuguese, Spanish, and Turkish.

- **ai\_toxicity\_flag\_automatically** : Automatically flag posts/chat messages when the classification for a specific category surpasses the configured threshold. Available categories are `toxicity`, `severe_toxicity`, `obscene`, `identity_attack`, `insult`, `threat`, and `sexual_explicit`. There’s an `ai_toxicity_flag_threshold_${category}` setting for each one.

- **ai\_toxicity\_groups\_bypass** : Users on those groups will not have their posts classified by the toxicity module. By default includes staff users.

## Additional resources

- [Discourse AI](https://meta.discourse.org/t/discourse-ai/259214?silent=true)
- [Install plugins on a self-hosted site](https://meta.discourse.org/t/install-plugins-in-discourse/19157?silent=true)

> Last edited by @hugh 2024-08-06T05:37:39Z
> 
> Last checked by @hugh 2024-08-06T05:37:44Z
> 
> > **Check document**
> >
> > Perform check on document:

---

<div class="post-metadata">

### Author: ![Hifihedgehog](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/hifihedgehog/32/140207_2.png) [@Hifihedgehog](https://meta.discourse.org/u/Hifihedgehog)
#### Post date: [September 11, 2023, 11:18pm UTC](https://meta.discourse.org/t/discourse-ai-toxicity/259215/3 "2023-09-11T23:18:43Z")

</div>

Tuning this a bit right now, am I correct in assuming that a higher threshold is more stringent and a lower one more lenient?

---

<div class="post-metadata">

### Author: ![JimPas](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/jimpas/32/148179_2.png) [@JimPas](https://meta.discourse.org/u/JimPas)
#### Post date: [September 12, 2023, 5:08am UTC](https://meta.discourse.org/t/discourse-ai-toxicity/259215/4 "2023-09-12T05:08:44Z")

</div>

I would say the higher the threshold, the more lenient it would be. A lower threshold would be more apt to flag a post as being toxic since it would take less to trigger a flag, thus a higher threshold would require more to trigger a flag.  
Low threshold = easy to cross  
High threshold = harder to cross

---

<div class="post-metadata">

### Author: ![nathank](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/nathank/32/290039_2.png) [@nathank](https://meta.discourse.org/u/nathank)
#### Post date: [November 23, 2023, 7:45am UTC](https://meta.discourse.org/t/discourse-ai-toxicity/259215/5 "2023-11-23T07:45:44Z")

</div>

I want to have a mechanism to catch attempts at commercial activity on our site - not toxicity per se, but very damaging to our community.

This is close, but not quite looking for the thing we are interested in.

Have you considered this dimension?

---

<div class="post-metadata">

### Author: ![Falco](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/falco/32/179432_2.png) [@Falco](https://meta.discourse.org/u/Falco)
#### Post date: [November 23, 2023, 12:00pm UTC](https://meta.discourse.org/t/discourse-ai-toxicity/259215/6 "2023-11-23T12:00:12Z")

</div>

That’s covered by [Discourse AI Post Classifier - Automation rule](https://meta.discourse.org/t/discourse-ai-post-classifier-automation-rule/281227). Let me know how it goes.

---

<div class="post-metadata">

### Author: ![Mr.X\_Mr.X](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/mr.x_mr.x/32/126610_2.png) [@Mr.X\_Mr.X](https://meta.discourse.org/u/Mr.X_Mr.X)
#### Post date: [April 17, 2024, 2:09am UTC](https://meta.discourse.org/t/discourse-ai-toxicity/259215/7 "2024-04-17T02:09:25Z")

</div>

Can someone help me set it up with Google Perspective API? I’d put a ad in the market place but i think here is more apropriate.

---

<div class="post-metadata">

### Author: ![Samantha\_Venia\_Logan](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/samantha_venia_logan/32/442576_2.png) [@Samantha\_Venia\_Logan](https://meta.discourse.org/u/Samantha_Venia_Logan)
#### Post date: [August 26, 2024, 5:46am UTC](https://meta.discourse.org/t/discourse-ai-toxicity/259215/8 "2024-08-26T05:46:42Z")

</div>

I know this was a year ago but please let me know how this implementation went! I am personally vested in it ^^ That said, please correct me if I’m wrong @Discourse, but the attributes you mention on this page ARE Perspective’s atomic metrics, as implemented through Detoxify so adding Perspective is a bit of a moot point right?

> - **ai\_toxicity\_flag\_automatically** : Automatically flag posts/chat messages when the classification for a specific category surpasses the configured threshold. Available categories are `toxicity`, `severe_toxicity`, `obscene`, `identity_attack`, `insult`, `threat`, and `sexual_explicit`. There’s an `ai_toxicity_flag_threshold_${category}` setting for each one.

Regardless, [Detoxify](https://github.com/unitaryai/detoxify) can be implemented by the [Kaggle community](https://www.kaggle.com/) community. That’s a great place to find someone to implement it because that’s precisely what Kaggle does 🙂

---

<div class="post-metadata">

### Author: ![Falco](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/falco/32/179432_2.png) [@Falco](https://meta.discourse.org/u/Falco)
#### Post date: [August 26, 2024, 7:21pm UTC](https://meta.discourse.org/t/discourse-ai-toxicity/259215/9 "2024-08-26T19:21:48Z")

</div>

> [@Samantha\_Venia\_Logan](#):
>
> I know this was a year ago but please let me know how this implementation went!

We integrated [GitHub - unitaryai/detoxify: Trained models & code to predict toxic comments on all 3 Jigsaw Toxic Comment Challenges. Built using ⚡ Pytorch Lightning and 🤗 Transformers. For access to our API, please email us at contact@unitary.ai. · GitHub](https://github.com/unitaryai/detoxify) models to handle automatic post toxicity classification and perform automatic flagging when over a configurable threshold.

What we found, is that while it works great if you have a zero tolerance for typical toxicity on your instances, like what more “brand” owned instance are, for other more community oriented Discourse instances, the toxicity models were too strict, generating too much flags in more lenient instances.

Because of that our current plan is to [Depreate Toxicity](https://meta.discourse.org/t/whats-next-for-toxicity-detection-in-discourse-ai/320283) and move this feature to our AI Triage plugin, where we give a customizable prompt for admins to adapt their automatic Toxicity detection to the levels of what are allowed in their instance.

We also plan on offering our customer a hosted moderation LLM, in the likes of [https://ai.google.dev/gemma/docs/shieldgemma](https://ai.google.dev/gemma/docs/shieldgemma) or [[2312.06674] Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations](https://arxiv.org/abs/2312.06674), which perfomed very well in our internal evals against the same dataset used in the original Jigsaw Kaggle competition that spawned Detoxify.

---

<div class="post-metadata">

### Author: ![Saif](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/saif/32/318253_2.png) [@Saif](https://meta.discourse.org/u/Saif)
#### Post date: [December 9, 2024, 6:05pm UTC](https://meta.discourse.org/t/discourse-ai-toxicity/259215/10 "2024-12-09T18:05:14Z")

</div>



---

<div class="post-metadata">

### Author: ![Saif](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/saif/32/318253_2.png) [@Saif](https://meta.discourse.org/u/Saif)
#### Post date: [December 9, 2024, 6:05pm UTC](https://meta.discourse.org/t/discourse-ai-toxicity/259215/11 "2024-12-09T18:05:20Z")

</div>


