Summary: when a Mistral model asks for several tools at once (for example three searches), Discourse AI runs only the first one. The agent then answers from partial results, and nothing warns the admin or the user.
When an LLM returns several tool calls in a single streamed chunk, Discourse AI keeps the first one and silently drops the others. Mistral does exactly this when it makes parallel tool calls, so with a Mistral LLM configured, an agent asks for several tools but only one of them actually runs. No error is raised or logged, and the final answer looks complete.
Tested on 2026.10.0-latest (67bc74d0d, current head of main, latest and tests-passed as of 2026-10-04), provider mistral, model mistral-large-2512, native tools, streaming (the default). I first saw it with mistral-medium-2508.
Minimal repro (no network, no API key)
# frozen_string_literal: true
# Network-free repro. Run with: bin/rails runner repro.rb
Processor = DiscourseAi::Completions::OpenAiMessageProcessor
calls = [
{ index: 0, id: "call_1", type: "function", function: { name: "echo", arguments: '{"text":"one"}' } },
{ index: 1, id: "call_2", type: "function", function: { name: "echo", arguments: '{"text":"two"}' } },
]
# Streaming: both tool calls in ONE chunk, as Mistral sends them
processor = Processor.new
chunk = { choices: [{ index: 0, delta: { tool_calls: calls }, finish_reason: "tool_calls" }] }
streamed = [processor.process_streamed_message(chunk), *processor.finish].compact
puts "streamed: #{streamed.size} of #{calls.size} -> #{streamed.map(&:id).inspect}"
# Control: the same tool calls as a non-streamed response
message = { choices: [{ index: 0, message: { role: "assistant", tool_calls: calls }, finish_reason: "tool_calls" }] }
non_streamed = Processor.new.process_message(message)
puts "non-streamed: #{non_streamed.size} of #{calls.size} -> #{non_streamed.map(&:id).inspect}"
Output:
streamed: 1 of 2 -> ["call_1"]
non-streamed: 2 of 2 -> ["call_1", "call_2"]
Same result with 3 calls (1 of 3), with partial_tool_calls: true, and when the raw SSE goes through Endpoints::Mistral#decode_chunk.
What Mistral actually streams (raw chunk)
Direct streaming call to the Mistral API, two dummy tools, prompt “What is the current weather in Paris and the local time in Tokyo?”. Every call is complete, all in one chunk, with finish_reason in that same chunk:
data: {"id":"5a2166a42dd14374a54944726c759667","object":"chat.completion.chunk","created":1791036485,"model":"mistral-large-2512","choices":[{"index":0,"delta":{"tool_calls":[{"id":"bM0Xl0KSh","type":"function","function":{"name":"get_weather","arguments":"{\"city\": \"Paris\"}"},"index":0},{"id":"kgOYVL46d","type":"function","function":{"name":"get_local_time","arguments":"{\"city\": \"Tokyo\"}"},"index":1}]},"finish_reason":"tool_calls"}],"usage":{"prompt_tokens":172,"total_tokens":195,"completion_tokens":23,"prompt_tokens_details":{"cached_tokens":0},"service_tier":"standard"},"p":"abcdefghijklmnopqrstuvwxyz"}
Mistral did this on 6 out of 6 attempts (2 and 3 calls). With parallel_tool_calls: false in the request, it returns one call per response (5 out of 5).
Cause
process_streamed_message only reads element 0 of the tool_calls array, and the processor tracks a single @tool:
Multiple calls are handled only when they arrive one per chunk: a new id in a later chunk closes the current call. That is how OpenAI streams them, and what the spec “properly handles multiple tool calls” covers. When several calls share one chunk, elements 1+ are never read, and finish closes the only call it knows about.
The non-streamed path (process_message) iterates the whole array, which is why the control in the repro gets everything.
Impact
- With a real agent (search and read tools, Mistral, streaming), over one run of 6 LLM responses: the first response held 3 searches in one chunk, and only 1 was run. The next 4 responses held one call each, and all of them ran. The last one was the final answer. In total, 7 tool calls were requested and 5 were run. Two searches were never executed, and the agent answered as if it had all the results.
- The dropped calls are recorded in Discourse AI’s own
AiApiAuditLog(raw_response_payload), so they are easy to check on any affected site. Nothing reaches the logs, Logster or the user. bot.rbis not involved: it runs everyToolCallit receives. The dropped ones never reach it.- The only lever in the UI is
disable_native_tools, which switches to XML tools and gives up native tool calling (not tested here). Themistralprovider does not exposedisable_streaming, and Discourse AI never sendsparallel_tool_calls.
Scope: the processor is shared by every OpenAI-compatible endpoint (OpenAI, Azure, Groq, Mistral, OpenRouter, vLLM). I have only confirmed the problem with Mistral. Others would be affected whenever they group several calls in one chunk, which I have not tested.
Workaround
I use a small plugin patch that asks Mistral for one tool call per response. It has worked since early September. It relies on a private method, though, and it does not fix the decoder.
Workaround patch
# Workaround: OpenAiMessageProcessor#process_streamed_message only reads
# tool_calls[0] of each streamed chunk. Mistral may return several tool calls
# in the same chunk (parallel_tool_calls defaults to true), so all but the
# first are silently dropped. Ask Mistral for one tool call per response.
module MistralSingleToolCallPerResponse
private
def prepare_payload(prompt, model_params, dialect)
payload = super
payload[:parallel_tool_calls] = false if payload[:tools].present?
payload
end
end
# plugin.rb, inside after_initialize
reloadable_patch do
if defined?(::DiscourseAi::Completions::Endpoints::Mistral)
::DiscourseAi::Completions::Endpoints::Mistral.prepend(MistralSingleToolCallPerResponse)
end
end
Possible fix
- Read every element of
tool_calls, keeping one state perindex(falling back toid) instead of the single@tool, and letprocess_streamed_messagereturn several objects (decode_chunkalready flattens). - Make
finishclose every open call. - With
partial_tool_calls, use one streaming parser per call. - Add a spec for several tool calls in one chunk.
A quick mitigation could be to send parallel_tool_calls: false for Mistral, or to expose disable_streaming for that provider.
Happy to help test a fix, or to open a PR if that’s useful.