# AI Image Captioning Feature in Discourse AI Plugin

**URL:** https://meta.discourse.org/t/ai-image-captioning-feature-in-discourse-ai-plugin/296087
**Category:** Announcements
**Tags:** ai, ai-helper
**Created:** [February 20, 2024, 5:53pm UTC](https://meta.discourse.org/t/ai-image-captioning-feature-in-discourse-ai-plugin/296087 "2024-02-20T17:53:28Z")
**Posts on this page:** 20
**Page:** 1

<div class="post-metadata">

### Author: ![Falco](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/falco/32/179432_2.png) [@Falco](https://meta.discourse.org/u/Falco)
#### Post date: [February 20, 2024, 5:53pm UTC](https://meta.discourse.org/t/ai-image-captioning-feature-in-discourse-ai-plugin/296087/1 "2024-02-20T17:53:28Z")

</div>

We’ve introduced an **AI Image Captioning** feature to the Discourse AI plugin, enabling automatic caption generation for images in posts. This functionality aims to improve content accessibility and enrich visual elements within your community.

### Features and Use

- **Automatic AI Captions** : Upon uploading an image in the editor, you can generate a caption automatically using AI.
- **Editable Captions** : The generated caption can be edited to better suit your content’s context and tone.
- **Enhanced Accessibility** : The feature supports creating more accessible content for users relying on screen readers.

### How to Use

1. Upload an image in the Discourse editor.
2. Click the “Caption with AI” button near the image.
3. A generated caption will appear, which you can modify.
4. Accept the caption to include it in your post.

### Feedback

Your feedback is crucial for refining this feature. It’s enabled here on Meta, so please share your experiences, issues, or suggestions here on this topic.

### AI Model

This feature supports both the open-source model LLaVa 1.6 or with the OpenAI API.

---

<div class="post-metadata">

### Author: ![frold](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/frold/32/82547_2.png) [@frold](https://meta.discourse.org/u/frold)
#### Post date: [February 20, 2024, 5:56pm UTC](https://meta.discourse.org/t/ai-image-captioning-feature-in-discourse-ai-plugin/296087/2 "2024-02-20T17:56:46Z")

</div>

Funny I used it earliner in this post. I was very impressed. It could read the image and tell what it was about in this post

[https://meta.discourse.org/t/discourse-subscriptions/140818/609?u=frold](https://meta.discourse.org/t/discourse-subscriptions/140818/609)

---

<div class="post-metadata">

### Author: ![EricGT](https://avatars.discourse-cdn.com/v4/letter/e/f1d935/32.png) [@EricGT](https://meta.discourse.org/u/EricGT)
#### Post date: [February 20, 2024, 6:10pm UTC](https://meta.discourse.org/t/ai-image-captioning-feature-in-discourse-ai-plugin/296087/3 "2024-02-20T18:10:07Z")

</div>

Noted this on the OpenAI forum

> **[How are blind people using OpenAI technology?](https://community.openai.com/t/how-are-blind-people-using-openai-technology/591018/19?u=ericgt)**
>
> FYI

---

<div class="post-metadata">

### Author: ![Jagster](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/jagster/32/192154_2.png) [@Jagster](https://meta.discourse.org/u/Jagster)
#### Post date: [February 20, 2024, 6:18pm UTC](https://meta.discourse.org/t/ai-image-captioning-feature-in-discourse-ai-plugin/296087/4 "2024-02-20T18:18:36Z")

</div>

I don’t know how we get mobile users remember to use that, because they have to jump away from editor.

Is that caption used as alt-text too?

---

<div class="post-metadata">

### Author: ![Falco](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/falco/32/179432_2.png) [@Falco](https://meta.discourse.org/u/Falco)
#### Post date: [February 20, 2024, 6:21pm UTC](https://meta.discourse.org/t/ai-image-captioning-feature-in-discourse-ai-plugin/296087/5 "2024-02-20T18:21:03Z")

</div>

> [@Jagster](#):
>
> Is that caption used as alt-text too?

Yes.

> [@Jagster](#):
>
> I don’t know how we get mobile users remember to use that, because they have to jump away from editor.

We plan on adding JIT reminders in the near future if the reception is good.

---

<div class="post-metadata">

### Author: ![Falco](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/falco/32/179432_2.png) [@Falco](https://meta.discourse.org/u/Falco)
#### Post date: [February 21, 2024, 5:00pm UTC](https://meta.discourse.org/t/ai-image-captioning-feature-in-discourse-ai-plugin/296087/6 "2024-02-21T17:00:10Z")

</div>

2 posts were split to a new topic: [Support for prompt customization in DiscourseAI](https://meta.discourse.org/t/support-for-prompt-customization-in-discourseai/296204)

---

<div class="post-metadata">

### Author: ![pmusaraj](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/pmusaraj/32/119489_2.png) [@pmusaraj](https://meta.discourse.org/u/pmusaraj)
#### Post date: [February 20, 2024, 10:15pm UTC](https://meta.discourse.org/t/ai-image-captioning-feature-in-discourse-ai-plugin/296087/7 "2024-02-20T22:15:37Z")

</div>

![A person is posing with a thumbs-up gesture, wearing glasses and a plaid shirt.](https://global.discourse-cdn.com/meta/original/4X/9/6/7/9679bc053829a45a30e9f5fe9ef6093506715ef7.jpeg)

_It can see the plaid shirt, but it can’t detect George Costanza. 🤣_

Jokes aside, this is great especially for #accessibility. In previous A11Y reports, missing alt text on images is one of the main items raised, and previously we’ve written all that off since images are user-uploaded content. This now draws a path forward to much, much better accessibility.

---

<div class="post-metadata">

### Author: ![Tris20](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/tris20/32/264639_2.png) [@Tris20](https://meta.discourse.org/u/Tris20)
#### Post date: [February 21, 2024, 8:23am UTC](https://meta.discourse.org/t/ai-image-captioning-feature-in-discourse-ai-plugin/296087/8 "2024-02-21T08:23:24Z")

</div>

![The image shows a computer terminal with multiple command line entries related to Git version control, where the user is attempting to push large files to a repository, but encounters an error because the file sizes exceed GitHub's allowed limit.](https://global.discourse-cdn.com/meta/original/4X/7/2/3/7230236d1e249283169dc112f39e491d2da3f368.png)

In the case of error messages, is there any way to encourage it to caption the main part of the error so the search engine picks up on it?

> **Some other results**
>
> It identifies the third correctly as the IBM EWM tool, but does not recognise 2 as being Rhapsody, and 1 being Vector Davinci. None the less these captions are pretty reasonable.
> 
> ![The image shows a software development environment with diagrams and tables related to ECU (Electronic Control Unit) programming or configuration, featuring panels for tasks, libraries, and output messages.](https://global.discourse-cdn.com/meta/original/4X/7/1/9/71941f2b53bf9b380168f070703d69dbaae88bf8.jpeg)
> 
> ![The image shows a UML (Unified Modeling Language) object diagram on a software development platform, possibly depicting the structure and relationships of components within a dishwasher system.](https://global.discourse-cdn.com/meta/original/4X/4/6/e/46eda5fa6888bd569f03e8ba3fa8609d92e72ff1.jpeg)
> 
> ![The image displays a screenshot of the IBM Engineering Workflow Management interface with various project management tasks and activities organized into sections such as "Incoming Work," "2.0," and "2.0 M1."](https://global.discourse-cdn.com/meta/original/4X/8/2/7/827d13487b7c03358fbfb5face0a9ba5cdae2082.png)

---

<div class="post-metadata">

### Author: ![tpetrov](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/tpetrov/32/164643_2.png) [@tpetrov](https://meta.discourse.org/u/tpetrov)
#### Post date: [February 21, 2024, 9:55am UTC](https://meta.discourse.org/t/ai-image-captioning-feature-in-discourse-ai-plugin/296087/9 "2024-02-21T09:55:09Z")

</div>

This is an awesome feature!

But it’s very hard to find. The user needs to hover over the image to see the button and then click it (and most people wion’t know about that).  
Even though I knew and was looking for the feature, I had the check the video to get that I need to hover.  
IMO it should be “in your face” to be used in the beginning. I’d even make it create the captions by default, without the user having to click anything :drevil:

---

<div class="post-metadata">

### Author: ![Falco](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/falco/32/179432_2.png) [@Falco](https://meta.discourse.org/u/Falco)
#### Post date: [February 21, 2024, 5:04pm UTC](https://meta.discourse.org/t/ai-image-captioning-feature-in-discourse-ai-plugin/296087/10 "2024-02-21T17:04:04Z")

</div>

> [@Tris20](#):
>
> In the case of error messages, is there any way to encourage it to caption the main part of the error so the search engine picks up on it?

We will eventually makes those prompts customizable, so this will then be possible.

> [@tpetrov](#):
>
> But it’s very hard to find.

As a new feature, our idea is to introduce it in a very unobtrusive way to gather feedback, and then make it easier to find and even automatic.

---

<div class="post-metadata">

### Author: ![JammyDodger](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/jammydodger/32/254611_2.png) [@JammyDodger](https://meta.discourse.org/u/JammyDodger)
#### Post date: [March 12, 2024, 9:36am UTC](https://meta.discourse.org/t/ai-image-captioning-feature-in-discourse-ai-plugin/296087/13 "2024-03-12T09:36:00Z")

</div>

6 posts were split to a new topic: [Issues configuring AI image captions](https://meta.discourse.org/t/issues-configuring-ai-image-captions/298980)

---

<div class="post-metadata">

### Author: ![ecki](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/ecki/32/264025_2.png) [@ecki](https://meta.discourse.org/u/ecki)
#### Post date: [March 15, 2024, 12:41pm UTC](https://meta.discourse.org/t/ai-image-captioning-feature-in-discourse-ai-plugin/296087/14 "2024-03-15T12:41:53Z")

</div>

Will that send the (Internet) Image link to the AI Service or upload the Image content or run some “hashing” locally in discourse? Is it server-side or javascript (i.e. exposing the client ip to external service).

---

<div class="post-metadata">

### Author: ![Falco](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/falco/32/179432_2.png) [@Falco](https://meta.discourse.org/u/Falco)
#### Post date: [March 15, 2024, 1:12pm UTC](https://meta.discourse.org/t/ai-image-captioning-feature-in-discourse-ai-plugin/296087/15 "2024-03-15T13:12:17Z")

</div>

It sends a link to the image to the service you selected for the captioning. It happens server-side, as there are credentials involved.

If you want the feature but don’t want to involve third-parties, you can always run LLaVa in your own server.

---

<div class="post-metadata">

### Author: ![ecki](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/ecki/32/264025_2.png) [@ecki](https://meta.discourse.org/u/ecki)
#### Post date: [March 15, 2024, 3:33pm UTC](https://meta.discourse.org/t/ai-image-captioning-feature-in-discourse-ai-plugin/296087/16 "2024-03-15T15:33:28Z")

</div>

> [@Falco](#):
>
> you can always run LLaVa in your own server.

agreed, however the quality might suuffer from hardware limitations. Maybe you could share some recommendation in regards to model-sizes and quatisation or minimum vram from your experience. (not sure if they have quantized models at all, their “zoo” seems to have only full models).

---

<div class="post-metadata">

### Author: ![Falco](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/falco/32/179432_2.png) [@Falco](https://meta.discourse.org/u/Falco)
#### Post date: [March 15, 2024, 3:46pm UTC](https://meta.discourse.org/t/ai-image-captioning-feature-in-discourse-ai-plugin/296087/17 "2024-03-15T15:46:57Z")

</div>

We are running the full model, but the smallest version of it with Mistral 7B. It’s taking 21GB VRAM in our single A100 servers, and it’s ran via `ghcr.io/xfalcox/llava:latest` container image.

Sadly the ecosystem for multi-modal models ain’t as mature as the text2text ones, so we can’t yet leverage inference servers like vLLM or TGI and are left with those one-off microservices. This may change this year, multimodal is on vLLM roadmap, but until then we can at least test the waters with those services.

---

<div class="post-metadata">

### Author: ![seanblue](https://avatars.discourse-cdn.com/v4/letter/s/dc4da7/32.png) [@seanblue](https://meta.discourse.org/u/seanblue)
#### Post date: [March 21, 2024, 10:34pm UTC](https://meta.discourse.org/t/ai-image-captioning-feature-in-discourse-ai-plugin/296087/18 "2024-03-21T22:34:15Z")

</div>

I have some small UX feedback for this. On small images, the “Capture with AI” button blocks not only the image itself but other text in the post, making it hard to review the post when editing.

 ![image](https://global.discourse-cdn.com/meta/original/4X/4/5/1/45166ca942d957b207ecc337c038aa92f8e8309a.png)

---

<div class="post-metadata">

### Author: ![Moin](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/moin/32/554653_2.png) [@Moin](https://meta.discourse.org/u/Moin)
#### Post date: [March 21, 2024, 10:55pm UTC](https://meta.discourse.org/t/ai-image-captioning-feature-in-discourse-ai-plugin/296087/19 "2024-03-21T22:55:04Z")

</div>

> [@100% 75% 50% next to images in preview window](https://meta.discourse.org/t/100-75-50-next-to-images-in-preview-window/299883/2):
>
> Yes, we do need to improve this specifically for small images. Thanks for pointing out the lack of detail on hover, that should be an easy addition.

---

<div class="post-metadata">

### Author: ![mattdm](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/mattdm/32/216484_2.png) [@mattdm](https://meta.discourse.org/u/mattdm)
#### Post date: [April 12, 2024, 1:59pm UTC](https://meta.discourse.org/t/ai-image-captioning-feature-in-discourse-ai-plugin/296087/21 "2024-04-12T13:59:38Z")

</div>

I am seeing all generated captions (both here and on my site) start with “The image contains” or “An image of” or similar. This seems unnecessary and redundant. Could the prompt be updated to tell it that it doesn’t need to explain that the image is an image?

---

<div class="post-metadata">

### Author: ![sam](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/sam/32/102149_2.png) [@sam](https://meta.discourse.org/u/sam)
#### Post date: [April 17, 2024, 3:20am UTC](https://meta.discourse.org/t/ai-image-captioning-feature-in-discourse-ai-plugin/296087/22 "2024-04-17T03:20:55Z")

</div>

> [@mattdm](#):
>
> Could the prompt be updated to tell it that it doesn’t need to explain that the image is an image?

It is so tricky to hone cause different models have different tolerances, but one plan we have is to allow community owners control over the prompts so they can experiment.

---

<div class="post-metadata">

### Author: ![Isambard](https://avatars.discourse-cdn.com/v4/letter/i/858c86/32.png) [@Isambard](https://meta.discourse.org/u/Isambard)
#### Post date: [June 3, 2024, 5:11pm UTC](https://meta.discourse.org/t/ai-image-captioning-feature-in-discourse-ai-plugin/296087/23 "2024-06-03T17:11:00Z")

</div>

@mattdm You can achieve this simply by pre-seeding the generated answer with “An image of”. This way the LLM thinks that it has already generated the introduction and will generate just the remainder.

[Next page](https://meta.discourse.org/t/ai-image-captioning-feature-in-discourse-ai-plugin/296087.md?page=2)
