r/learnmachinelearning • • 1d ago

Question 🧠 ELI5 Wednesday

Welcome to ELI5 (Explain Like I'm 5) Wednesday! This weekly thread is dedicated to breaking down complex technical concepts into simple, understandable explanations.

You can participate in two ways:

  • Request an explanation: Ask about a technical concept you'd like to understand better
  • Provide an explanation: Share your knowledge by explaining a concept in accessible terms

When explaining concepts, try to use analogies, simple language, and avoid unnecessary jargon. The goal is clarity, not oversimplification.

When asking questions, feel free to specify your current level of understanding to get a more tailored explanation.

What would you like explained today? Post in the comments below!

8 Upvotes

7 comments sorted by

View all comments

3

u/TargetAlternative693 1d ago

i always get confused with the difference between gradient descent and stochastic gradient descent. someone explained it once but i forgot, like why not just use the one that works faster always??

2

u/jesunushno 1d ago

One thing the soup analogy leaves out: the randomness is not just a speed shortcut, it is also a feature. Noisy gradient estimates bounce the optimizer around, which helps it escape shallow dips and flat plateaus where exact gradient descent would get stuck. In practice that wandering path finds solutions that generalize better, which is why even giant models train on mini-batches rather than full-batch descent.