# Discourse AI: parallel tool calls silently dropped with Mistral (streaming)

**URL:** <https://meta.discourse.org/t/discourse-ai-parallel-tool-calls-silently-dropped-with-mistral-streaming/413903>\
**Category:** Bug\
**Tags:** ai\
**Created:** [October 4, 2026, 3:30pm UTC](https://meta.discourse.org/t/discourse-ai-parallel-tool-calls-silently-dropped-with-mistral-streaming/413903 "2026-10-04T15:30:07Z")\
**Posts on this page:** 1\
**Page:** 1

<div class="post-metadata">

**Author:** ![jmx](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/jmx/32/579660_2.png) [@jmx](https://meta.discourse.org/u/jmx)\
**Post date:** [October 4, 2026, 3:30pm UTC](https://meta.discourse.org/t/discourse-ai-parallel-tool-calls-silently-dropped-with-mistral-streaming/413903/1 "2026-10-04T15:30:08Z")

</div>

**Summary:** when a Mistral model asks for several tools at once (for example three searches), Discourse AI runs only the first one. The agent then answers from partial results, and nothing warns the admin or the user.

When an LLM returns several tool calls in a **single streamed chunk** , Discourse AI keeps the first one and silently drops the others. Mistral does exactly this when it makes parallel tool calls, so with a Mistral LLM configured, an agent asks for several tools but only one of them actually runs. No error is raised or logged, and the final answer looks complete.

Tested on `2026.10.0-latest` (`67bc74d0d`, current head of `main`, `latest` and `tests-passed` as of 2026-10-04), provider `mistral`, model `mistral-large-2512`, native tools, streaming (the default). I first saw it with `mistral-medium-2508`.

### Minimal repro (no network, no API key)

```ruby
# frozen_string_literal: true
# Network-free repro. Run with: bin/rails runner repro.rb

Processor = DiscourseAi::Completions::OpenAiMessageProcessor

calls = [
  { index: 0, id: "call_1", type: "function", function: { name: "echo", arguments: '{"text":"one"}' } },
  { index: 1, id: "call_2", type: "function", function: { name: "echo", arguments: '{"text":"two"}' } },
]

# Streaming: both tool calls in ONE chunk, as Mistral sends them
processor = Processor.new
chunk = { choices: [{ index: 0, delta: { tool_calls: calls }, finish_reason: "tool_calls" }] }
streamed = [processor.process_streamed_message(chunk), *processor.finish].compact
puts "streamed: #{streamed.size} of #{calls.size} -> #{streamed.map(&:id).inspect}"

# Control: the same tool calls as a non-streamed response
message = { choices: [{ index: 0, message: { role: "assistant", tool_calls: calls }, finish_reason: "tool_calls" }] }
non_streamed = Processor.new.process_message(message)
puts "non-streamed: #{non_streamed.size} of #{calls.size} -> #{non_streamed.map(&:id).inspect}"

```

Output:

```plaintext
streamed: 1 of 2 -> ["call_1"]
non-streamed: 2 of 2 -> ["call_1", "call_2"]

```

Same result with 3 calls (1 of 3), with `partial_tool_calls: true`, and when the raw SSE goes through `Endpoints::Mistral#decode_chunk`.

> **What Mistral actually streams (raw chunk)**
>
> Direct streaming call to the Mistral API, two dummy tools, prompt _“What is the current weather in Paris and the local time in Tokyo?”_. Every call is complete, all in one chunk, with `finish_reason` in that same chunk:
> 
> ```plaintext
> data: {"id":"5a2166a42dd14374a54944726c759667","object":"chat.completion.chunk","created":1791036485,"model":"mistral-large-2512","choices":[{"index":0,"delta":{"tool_calls":[{"id":"bM0Xl0KSh","type":"function","function":{"name":"get_weather","arguments":"{\"city\": \"Paris\"}"},"index":0},{"id":"kgOYVL46d","type":"function","function":{"name":"get_local_time","arguments":"{\"city\": \"Tokyo\"}"},"index":1}]},"finish_reason":"tool_calls"}],"usage":{"prompt_tokens":172,"total_tokens":195,"completion_tokens":23,"prompt_tokens_details":{"cached_tokens":0},"service_tier":"standard"},"p":"abcdefghijklmnopqrstuvwxyz"}
> 
> ```
> 
> Mistral did this on 6 out of 6 attempts (2 and 3 calls). With `parallel_tool_calls: false` in the request, it returns one call per response (5 out of 5).

### Cause

`process_streamed_message` only reads element `0` of the `tool_calls` array, and the processor tracks a single `@tool`:

> <https://github.com/discourse/discourse/blob/67bc74d0d83f8037ec538c1299b8d8cb59211319/plugins/discourse-ai/lib/completions/open_ai_message_processor.rb#L46-L52>

Multiple calls are handled only when they arrive **one per chunk** : a new `id` in a later chunk closes the current call. That is how OpenAI streams them, and what the spec _“properly handles multiple tool calls”_ covers. When several calls share one chunk, elements 1+ are never read, and `finish` closes the only call it knows about.

The non-streamed path (`process_message`) iterates the whole array, which is why the control in the repro gets everything.

### Impact

- With a real agent (search and read tools, Mistral, streaming), over one run of 6 LLM responses: the first response held **3 searches in one chunk, and only 1 was run**. The next 4 responses held one call each, and all of them ran. The last one was the final answer. In total, 7 tool calls were requested and 5 were run. Two searches were never executed, and the agent answered as if it had all the results.
- The dropped calls are recorded in Discourse AI’s own `AiApiAuditLog` (`raw_response_payload`), so they are easy to check on any affected site. Nothing reaches the logs, Logster or the user.
- `bot.rb` is not involved: it runs every `ToolCall` it receives. The dropped ones never reach it.
- The only lever in the UI is `disable_native_tools`, which switches to XML tools and gives up native tool calling (not tested here). The `mistral` provider does not expose `disable_streaming`, and Discourse AI never sends `parallel_tool_calls`.

**Scope:** the processor is shared by every OpenAI-compatible endpoint (OpenAI, Azure, Groq, Mistral, OpenRouter, vLLM). I have only confirmed the problem with Mistral. Others would be affected whenever they group several calls in one chunk, which I have not tested.

### Workaround

I use a small plugin patch that asks Mistral for one tool call per response. It has worked since early September. It relies on a private method, though, and it does not fix the decoder.

> **Workaround patch**
>
> ```ruby
> # Workaround: OpenAiMessageProcessor#process_streamed_message only reads
> # tool_calls[0] of each streamed chunk. Mistral may return several tool calls
> # in the same chunk (parallel_tool_calls defaults to true), so all but the
> # first are silently dropped. Ask Mistral for one tool call per response.
> module MistralSingleToolCallPerResponse
> private
> 
> def prepare_payload(prompt, model_params, dialect)
> payload = super
> payload[:parallel_tool_calls] = false if payload[:tools].present?
> payload
> end
> end
> 
> # plugin.rb, inside after_initialize
> reloadable_patch do
> if defined?(::DiscourseAi::Completions::Endpoints::Mistral)
> ::DiscourseAi::Completions::Endpoints::Mistral.prepend(MistralSingleToolCallPerResponse)
> end
> end
> 
> ```

### Possible fix

- Read every element of `tool_calls`, keeping one state per `index` (falling back to `id`) instead of the single `@tool`, and let `process_streamed_message` return several objects (`decode_chunk` already flattens).
- Make `finish` close every open call.
- With `partial_tool_calls`, use one streaming parser per call.
- Add a spec for several tool calls in one chunk.

A quick mitigation could be to send `parallel_tool_calls: false` for Mistral, or to expose `disable_streaming` for that provider.

Happy to help test a fix, or to open a PR if that’s useful.
