Discourse AI: a Gemini error event inside a streamed reply is ignored, and the partial text is stored as a complete translation

Description

When Google’s Gemini API fails partway through a streamed reply, Discourse AI treats the reply as complete. With AI translation enabled, the fragment received before the failure is stored as the finished translation of the post or topic title. Nothing is logged as an error, and the backfill never retries the item, because a translation row now exists.

Observed on a self-hosted site running release/2026.7 (2026.7.3, commit f1caa6321287f918644fba9ff8943579fdcf027a) with the bundled discourse-ai plugin. The LLM is gemini-3.8-flash through the Google provider, with the service tier set to flex. Flex is the tier where Google sheds load, so it produces these mid-stream failures regularly; the handling described below does not depend on the tier.

In these cases Google answers with HTTP 200, streams one or more normal events, and then sends an error event in the same stream:

data: {"candidates": [ ... ],"usageMetadata": { ... },"serviceTier": "flex","modelVersion": "gemini-3.8-flash","responseId": "..."}

data: {"error": {"code": 503,"message": "This model is currently experiencing high demand. Spikes in demand are usually temporary. Please try again later.","status": "UNAVAILABLE"}}

The audit log records response_status 200 and a small response_tokens count. Examples of what was stored: a 917-character post saved as 26 characters (its greeting line only), link-only posts saved as the first 8 to 24 characters of the URL, and topic titles cut off mid-word.

Measured on this site: 23 of 394 answered Flex-tier translation calls (5.8%) ended this way. None of 952 standard-tier calls to the same model did.

Root cause

  1. DiscourseAi::Completions::Endpoints::Gemini#decode_chunk reads only candidates from each parsed stream event. An event whose top-level key is error yields no parts and is skipped without any check.
  2. DiscourseAi::Completions::Endpoints::Base decides between success, retry and failure from the HTTP status alone (response.code.to_i != 200). The status is 200 here, so the existing retry logic for 503 is never reached.
  3. The stream then ends normally, and the caller receives the text accumulated so far. DiscourseAi::Translation::PostLocalizer and TopicLocalizer save it.

The same code is present on main as of 2026-10-03.

Suggested fix

Treat a top-level error object in a streamed Gemini event as a failed completion: discard the partial output and raise CompletionFailed, so that the existing retry handling applies to retriable codes such as 503 and callers never receive a fragment as a complete reply.

Reproduce

The failure depends on Google’s load, so it cannot be forced. On a site using a Gemini model on the Flex tier with AI translation, affected calls can be found afterwards in the audit log:

SELECT id, created_at, post_id, topic_id, response_status, response_tokens
FROM ai_api_audit_logs
WHERE feature_name = 'translation'
  AND raw_response_payload LIKE '%data: {"error"%'
  AND response_tokens > 0
ORDER BY created_at

Each row is a call that returned HTTP 200, produced some output, and then carried an error event. The corresponding rows in post_localizations or topic_localizations hold the truncated text.

Workaround

Moving the translation agents to the standard service tier stopped new occurrences. The stored fragments had to be found with the query above, deleted, and translated again.

1 Like