""" We add seemingly relevant statements to the questions that are, in fact, irrelevant to the reasoning and conclusion. However, the majority of models fail to ignore these statements and blindly convert them into operations, leading to mistakes prompt: Oliver picks 44 kiwis on Friday. Then he picks 58 kiwis on Saturday. On Sunday, he picks double the number of kiwis he did on Friday, but five of them were a bit smaller than average. How many kiwis does Oliver have?
o1-mini response: Oliver picks kiwis over three days as follows:
- Friday: 44 kiwis
- Saturday: 58 kiwis
- Sunday: Double the number he picked on Friday, which is 2 × 44 = 88 kiwis
However, on Sunday, 5 of these kiwis were smaller than average. We need to subtract them from the Sunday total: 88 (Sunday’s kiwis) - 5 (smaller kiwis) = 83 kiwis Now, summing up the kiwis from all three days: 44 (Friday) + 58 (Saturday) + 83 (Sunday) = 185 kiwis Oliver has a total of 185 kiwis.
"""
I tried asking exactly the same question to claude 3.5 and the response was:
"""
Let's break this problem down step by step:
1. Friday's kiwis:
Oliver picked 44 kiwis on Friday.
2. Saturday's kiwis: Oliver picked 58 kiwis on Saturday.
3. Sunday's kiwis: Oliver picked double the number of kiwis he did on Friday.
Friday's kiwis = 44
Double of Friday's kiwis = 44 × 2 = 88 kiwis
However, 5 of these were smaller than average, but they still count in the total.
4. Total kiwis: Friday's kiwis + Saturday's kiwis + Sunday's kiwis
= 44 + 58 + 88
= 190 kiwis
Therefore, Oliver has a total of 190 kiwis."""
This sounds like they're inspecting existing models. Maybe a model trained specifically on "word problem" question-answer pairs (as in, the sort of things that show up on tests and always pretend that the sort of complications a domain expert would know about just don't exist) would do better?
They should indeed, but is that "I've never seen this before" human baseline, or with prior exposure ("what do you mean I got that wro... oh I see what you did there") or explicit instruction?
- article is titled "GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models"
- some people having been flogging it as "LLMs cannot reason"
- it shows a 6-8 point drop, in test results in the 80s, if you replace the #s in the test set problems with random #s, and run multiple times
- If anything, sounds like a huge W to me: very hard to claim they're just memorizing with that small of a drop
Maybe on grade-school tests, but the professional certification exams I've taken have questions about scenarios where part of the challenge is recognizing which parts of the scenario description aren't relevant. The one I had an in-person class for, the instructor specifically called this out and advised reading the questions first so we'd know which parts of the scenario we didn't have to think about while reading it.
https://deepmind.google/discover/blog/ai-solves-imo-problems...
Whether this is "like a person" or not, it seems silly to insist that this "doesn't count" as mathematical reasoning. I certainly couldn't get an IMO silver medal and I have a degree in math.
> First, the problems were manually translated into formal mathematical language for our systems to understand.
Ok so the "AI" wasn't solving the same problem as every other Olympiad, it was solving a "translated" version. I wonder how much of the solving was performed in this translation. If the model is so capable of reasoning, why was this step performed by people?
> In the official competition, students submit answers in two sessions of 4.5 hours each. Our systems solved one problem within minutes and took up to three days to solve the others.
So no, they would not have been awarded a silver medal if they were competing in the IMO.
Not to mention, AlphaProof is not an LLM and has absolutely nothing to do with what I was commenting about.
The study proves nothing of the sort. Even the results of 4o are enough to give pause to this conclusion.
A "hey, you got that wrong. check again" is fine if the LLMs in the paper are also being prompted that way.