Description
When a Gemini reply is streamed, Discourse AI occasionally processes only the first part of it. The complete reply is received from Google and stored in ai_api_audit_logs.raw_response_payload, the stream ends with finishReason STOP, and nothing is logged as an error. The caller receives the text up to some event and nothing after it. With AI translation, the truncated text is stored as the finished translation.
Observed on a self-hosted site running release/2026.7 (2026.7.3) with the bundled discourse-ai plugin and gemini-3.8-flash on the standard service tier. In 3 of about 26,200 translation calls over three days, response_tokens in the audit log was lower than the final candidatesTokenCount Google reported in the same reply: 196 of 241, 84 of 328 and 131 of 179. In each case the stored count equals Google’s cumulative count at one intermediate event, so decoding stopped after that event.
This is separate from the in-stream error events fixed by #44262: there is no error event here, and the decoder is unchanged by that pull request.
Root cause
DiscourseAi::Completions::Endpoints::Gemini::GeminiStreamingDecoder#decode splits its buffer on /\r?\n\r?\n/ and accepts a segment only if it starts with data: {.
When a network chunk ends inside the blank line that separates two events, for example after the first \r\n of \r\n\r\n, the event before it parses and the buffer is emptied. The next chunk then begins with the remaining \r\n, so the next segment is "\r\ndata: {...}". It does not start with data: {, so it is kept in the buffer as an incomplete line. Every later event is appended to that buffer and is never recognised either. The stream ends normally and the remaining events are discarded with the buffer.
The same happens with LF-only separators when a chunk ends between the two line feeds.
The decoder is identical at v2026.7.3 and on main at b5548f76 (2026-10-05).
Reproduce
In a Rails console:
decoder = DiscourseAi::Completions::Endpoints::Gemini::GeminiStreamingDecoder.new
event = ->(text, tokens) { %(data: {"candidates": [{"content": {"parts": [{"text": "#{text}"}],"role": "model"}}],"usageMetadata": {"candidatesTokenCount": #{tokens}}}) }
chunks = [
event.("A", 11) + "\r\n\r\n",
event.("B", 38) + "\r\n",
"\r\n" + event.("C", 65) + "\r\n\r\n",
event.("D", 95) + "\r\n\r\n",
]
chunks.flat_map { |chunk| decoder.decode(chunk) }.map { |e| e.dig(:candidates, 0, :content, :parts, 0, :text) }
Result: ["A", "B"]. Expected: ["A", "B", "C", "D"]. Moving the chunk boundary to any point outside the separator, including the middle of an event, gives the expected result.
Suggested fix
Ignore leading line breaks on a segment before testing it, for example:
line = line.sub(/\A[\r\n]+/, "")
if line.start_with?("data: {")
With this change the reproduction above returns all four events, as do boundaries after \r\n, after \r\n\r, inside an event, and with LF-only separators.
Finding affected calls
Each row returned by this query is a call that Google completed normally and Discourse decoded only in part:
SELECT l.id, l.created_at, l.post_id, l.topic_id, l.response_tokens AS stored_tokens, g.sent_tokens
FROM ai_api_audit_logs l
CROSS JOIN LATERAL (
SELECT MAX((m)[1]::int) AS sent_tokens
FROM regexp_matches(l.raw_response_payload, '"candidatesTokenCount":\s*(\d+)', 'g') AS m
) g
WHERE l.language_model LIKE 'gemini%'
AND l.raw_response_payload ~ '"finishReason":\s*"STOP"'
AND g.sent_tokens <> l.response_tokens
ORDER BY l.created_at
Impact
The fault depends only on where network chunk boundaries fall, so it is not specific to translation or to a service tier; any streamed Gemini completion can lose its tail this way. For translation the result is a silently shortened post or title that the backfill does not retry, because a localization row exists.