Discourse AI: parallel tool calls silently dropped with Mistral (streaming)

Summary: when a Mistral model asks for several tools at once (for example three searches), Discourse AI runs only the first one. The agent then answers from partial results, and nothing warns the admin or the user.

When an LLM returns several tool calls in a single streamed chunk, Discourse AI keeps the first one and silently drops the others. Mistral does exactly this when it makes parallel tool calls, so with a Mistral LLM configured, an agent asks for several tools but only one of them actually runs. No error is raised or logged, and the final answer looks complete.

Tested on 2026.10.0-latest (67bc74d0d, current head of main, latest and tests-passed as of 2026-10-04), provider mistral, model mistral-large-2512, native tools, streaming (the default). I first saw it with mistral-medium-2508.

Minimal repro (no network, no API key)

# frozen_string_literal: true
# Network-free repro. Run with: bin/rails runner repro.rb

Processor = DiscourseAi::Completions::OpenAiMessageProcessor

calls = [
  { index: 0, id: "call_1", type: "function", function: { name: "echo", arguments: '{"text":"one"}' } },
  { index: 1, id: "call_2", type: "function", function: { name: "echo", arguments: '{"text":"two"}' } },
]

# Streaming: both tool calls in ONE chunk, as Mistral sends them
processor = Processor.new
chunk = { choices: [{ index: 0, delta: { tool_calls: calls }, finish_reason: "tool_calls" }] }
streamed = [processor.process_streamed_message(chunk), *processor.finish].compact
puts "streamed:     #{streamed.size} of #{calls.size} -> #{streamed.map(&:id).inspect}"

# Control: the same tool calls as a non-streamed response
message = { choices: [{ index: 0, message: { role: "assistant", tool_calls: calls }, finish_reason: "tool_calls" }] }
non_streamed = Processor.new.process_message(message)
puts "non-streamed: #{non_streamed.size} of #{calls.size} -> #{non_streamed.map(&:id).inspect}"

Output:

streamed:     1 of 2 -> ["call_1"]
non-streamed: 2 of 2 -> ["call_1", "call_2"]

Same result with 3 calls (1 of 3), with partial_tool_calls: true, and when the raw SSE goes through Endpoints::Mistral#decode_chunk.

What Mistral actually streams (raw chunk)

Direct streaming call to the Mistral API, two dummy tools, prompt “What is the current weather in Paris and the local time in Tokyo?”. Every call is complete, all in one chunk, with finish_reason in that same chunk:

data: {"id":"5a2166a42dd14374a54944726c759667","object":"chat.completion.chunk","created":1791036485,"model":"mistral-large-2512","choices":[{"index":0,"delta":{"tool_calls":[{"id":"bM0Xl0KSh","type":"function","function":{"name":"get_weather","arguments":"{\"city\": \"Paris\"}"},"index":0},{"id":"kgOYVL46d","type":"function","function":{"name":"get_local_time","arguments":"{\"city\": \"Tokyo\"}"},"index":1}]},"finish_reason":"tool_calls"}],"usage":{"prompt_tokens":172,"total_tokens":195,"completion_tokens":23,"prompt_tokens_details":{"cached_tokens":0},"service_tier":"standard"},"p":"abcdefghijklmnopqrstuvwxyz"}

Mistral did this on 6 out of 6 attempts (2 and 3 calls). With parallel_tool_calls: false in the request, it returns one call per response (5 out of 5).

Cause

process_streamed_message only reads element 0 of the tool_calls array, and the processor tracks a single @tool:

Multiple calls are handled only when they arrive one per chunk: a new id in a later chunk closes the current call. That is how OpenAI streams them, and what the spec “properly handles multiple tool calls” covers. When several calls share one chunk, elements 1+ are never read, and finish closes the only call it knows about.

The non-streamed path (process_message) iterates the whole array, which is why the control in the repro gets everything.

Impact

  • With a real agent (search and read tools, Mistral, streaming), over one run of 6 LLM responses: the first response held 3 searches in one chunk, and only 1 was run. The next 4 responses held one call each, and all of them ran. The last one was the final answer. In total, 7 tool calls were requested and 5 were run. Two searches were never executed, and the agent answered as if it had all the results.
  • The dropped calls are recorded in Discourse AI’s own AiApiAuditLog (raw_response_payload), so they are easy to check on any affected site. Nothing reaches the logs, Logster or the user.
  • bot.rb is not involved: it runs every ToolCall it receives. The dropped ones never reach it.
  • The only lever in the UI is disable_native_tools, which switches to XML tools and gives up native tool calling (not tested here). The mistral provider does not expose disable_streaming, and Discourse AI never sends parallel_tool_calls.

Scope: the processor is shared by every OpenAI-compatible endpoint (OpenAI, Azure, Groq, Mistral, OpenRouter, vLLM). I have only confirmed the problem with Mistral. Others would be affected whenever they group several calls in one chunk, which I have not tested.

Workaround

I use a small plugin patch that asks Mistral for one tool call per response. It has worked since early September. It relies on a private method, though, and it does not fix the decoder.

Workaround patch
# Workaround: OpenAiMessageProcessor#process_streamed_message only reads
# tool_calls[0] of each streamed chunk. Mistral may return several tool calls
# in the same chunk (parallel_tool_calls defaults to true), so all but the
# first are silently dropped. Ask Mistral for one tool call per response.
module MistralSingleToolCallPerResponse
  private

  def prepare_payload(prompt, model_params, dialect)
    payload = super
    payload[:parallel_tool_calls] = false if payload[:tools].present?
    payload
  end
end

# plugin.rb, inside after_initialize
reloadable_patch do
  if defined?(::DiscourseAi::Completions::Endpoints::Mistral)
    ::DiscourseAi::Completions::Endpoints::Mistral.prepend(MistralSingleToolCallPerResponse)
  end
end

Possible fix

  • Read every element of tool_calls, keeping one state per index (falling back to id) instead of the single @tool, and let process_streamed_message return several objects (decode_chunk already flattens).
  • Make finish close every open call.
  • With partial_tool_calls, use one streaming parser per call.
  • Add a spec for several tool calls in one chunk.

A quick mitigation could be to send parallel_tool_calls: false for Mistral, or to expose disable_streaming for that provider.

Happy to help test a fix, or to open a PR if that’s useful.

2 Likes