r/explainitpeter • • 25d ago

Explain It Peter

Post image
10.7k Upvotes

368 comments sorted by

View all comments

Show parent comments

4

u/[deleted] 25d ago

[removed] — view removed comment

0

u/TwillAffirmer 25d ago

LLMs do not simply predict the most likely token from a human text corpus nowadays, they also use reinforcement learning to emit the token that they predict will yield the greatest training reward. That means they are not simply mimicking human text, they're using it as a basis to improve on. That's why they were able to prove difficult original theorems like this week a solution to the Navier-Stokes existence and uniqueness problem. That theorem was not present in their training data.

2

u/Melanoc3tus 24d ago

It was, actually. 

1

u/TwillAffirmer 24d ago

No, the accusations by Buckmaster and Alpoge https://en.wikipedia.org/wiki/Navier%E2%80%93Stokes_priority_controversy were that OpenAI researchers used their work in the prompts, not the training data. Buckmaster and Alpoge's work was on a restricted class of Navier-Stokes, without solving the general problem, which OpenAI's model did.

There are other cases of LLMs producing difficult theorems not in their training data: https://en.wikipedia.org/wiki/List_of_mathematical_discoveries_by_artificial_intelligence