Two different answers from gpt4o - one right, one wrong.. !?

Math vs logic .. a mind-bending aspect of the “two different answers topic” IMO, thanks

Please consider…

I want to a buy a product online and I see three sellers offer the same product – all have customer reviews:

  • The first has 10 reviews, all positive
  • The second has 50 reviews, 48 positive
  • The third has 200 reviews, 186 positive.

Using standard principles of probability, which seller should I buy from: 1 , 2, or 3 ?

According to 3Blue1Brown reference material, answer should be Seller 2. (Binomial distributions | Probabilities of probabilities.)

GPT 3.5 (OpenAI browser GUI):
“If you prioritize both high probability and a larger sample size, you might consider the second seller :github_check:, as it has a high probability of positive reviews with a relatively larger sample size”

Gemini 1.5 Pro (Google AI Studio):
“You should be most inclined to buy from seller 3 :x:. who offers the most statistically reliable data.”

Claude 3 Sonnet (Anthropic browser GUI):
“According to standard principles of probability and statistics, a larger sample size generally provides a more reliable estimate of the true population proportion. It would be most reasonable to choose Seller 3” :x:.

My custom Discourse AI persona (Gemini Pro):
“You should likely go with product 3” :x: .

My custom Discourse AI persona (GPT4o):
“The second :github_check: seller (96% with 50 reviews) might be a balanced choice between high probability and sufficient review volume.”

Some of the ‘logic’ put forth by these LLM’s is truly laughable! .. and none of them seemed to grasp the real statistical nuances ..

Considering how many variables there are in the LLM game, it would seem that comprehensive ‘in situ’ testing frameworks will be a non-optional feature going forward (plugin? :slightly_smiling_face:)

Factors :

  • LLM Model release/version ( they seem to tweak fine tuning regularly )
  • Prompt structure at various levels
  • In-context learning content of various types
  • Math and logic aspects
  • Censorship guardrails
  • Ancillary tools ( js, python, julia, etc)
  • Etc. Etc.