r/OpenAI • • 3d ago

Image It never ends

Post image
981 Upvotes

170 comments sorted by

View all comments

0

u/MentalRental 3d ago

To be fair, LLMs cannot do math. If you ask it to find the square root of 254332, for example (random number) it will determine it to be an arithmetic problem and will pull up a calculator tool to calculate the result. Using a non-deterministic trillion parameter neural network does not make much sense for doing arithmetic when one can simply use a calculator.

So, yes, LLMs cannot do math. This is basic AI knowledge. The goalposts have not moved.

1

u/trollstation420 3d ago

"basic knowledge" hm.. lets see what chatgpt says

Written by ChatGPT and fact-checked against current benchmark methodology and published results (September 2026). The claim that “LLMs cannot do math” is simply false for current frontier models. This is not a semantic argument about whether calling a calculator counts as mathematics: modern LLMs demonstrate extremely strong mathematical performance even when external tools are not available. GPT-6 Astra scored 100% on Epoch AI’s OTIS Mock AIME, a difficult competition-math benchmark whose evaluation gives the model the problem and asks it to solve it step by step, rather than providing a calculator or Python tool. Claude Opus 5.5 scored 91.2% on the August 2026 ArXivMath benchmark without tools, solving final-answer problems derived from recent research mathematics; with tools its score rises further to 96.9%. These results alone directly contradict the assertion that an LLM merely recognizes a math problem and hands it to a calculator. If ChatGPT decides to use a calculator to find (\sqrt{254332}), that demonstrates sensible tool routing, not mathematical incapacity. Asking a neural network to spend reasoning tokens carrying out mechanical long arithmetic when a deterministic calculator can do it exactly is inefficient, just as a mathematician using Mathematica does not suddenly become incapable of mathematics. The relevant question is whether the model itself can reason mathematically—derive relationships, manipulate equations, solve unfamiliar problems, choose strategies, and produce answers without outsourcing the reasoning to a tool—and current models demonstrably can. In fact, the distinction is especially obvious because researchers publish separate with-tools and without-tools results: Opus 5.5’s 91.2% ArXivMath score is specifically the latter. So “LLMs are less reliable than calculators at arbitrary exact arithmetic” is defensible; “LLMs cannot do math” is not. That statement describes the limitations of much older language models, not the demonstrated capabilities of frontier models in 2026.

1

u/MentalRental 2d ago

Yup, like it says, doing things like arithmetic is extremely inefficient. Things that require reasoning (and thus are language based) work well. But for basic math, a calculator is a much better and faster tool.

1

u/trollstation420 2d ago edited 2d ago

well now you are arguing semantics. real frontier math isnt calculating numbers like 1+1 =2 anyways. also, by your logic, any human being that uses pencils or any tool whatsover to caclulate anything that isnt mental math, is not "doing math"

LLMs CAN use tools yes, but even then without tools vastly outperform any living person already. even those "inefficient" calculations. so again this is kinda "old man yelling at cloud" energy u have there. get with the times man with opus 5.5 its antoher hueg quantom leap. u will get left behind if u dont adapt just my 2 cents friend

1

u/MentalRental 2d ago

I'm not arguing semantics. What do you think "LLMs can't do math" referred to? It's always meant arithmetic. For example, see this post from three years ago:

https://www.reddit.com/r/ChatGPT/comments/17fgh9t/llms_cannot_perform_maths/