← Back to Homepage

Stochastic Gradient Descent Made Fun

The Math Behind ChatGPT and AI

Guest Speaker - Griffin Rutherford - MS Computer Science - Mountain Runner

Start with slopes and curves. Follow small steps downhill, then see how the same idea helps train AI.

Drag the bowls to change your view. Pause the animations to discuss a frame.

Write your answer, then choose Check answer for feedback. Short formulas are checked mathematically. Written feedback identifies key ideas and suggests what to add; open the example to compare your reasoning. Press Enter in a short-answer field, or Ctrl/⌘ + Enter in a written response.

1.
What expression gives the slope between two points?
2.
What does it mean when the slope is zero?
3.
For a parabola written as y = ax² + bx + c, what is the expression for the x-coordinate of its vertex?
4.
If you connect two dots that are equidistant from the vertex with a line, what is the slope of that line?
5.
For the upward-opening parabola y = x², how does the slope change from the left to right side of the vertex?
6.
Using this quartic equation as an example: y = x⁴ + 2x³ − 3x² + x + 1 We have something that looks like a vertex, but this higher-order polynomial can have several turning points. There is no single parabola-vertex shortcut for this curve. One local minimum is near (-2.22, -13.60).
In real life problem solving, why might finding a vertex or local min/max be useful?
7.
If f(x) represents the amount of water pollution, and x is the amount of a resource used in an engineering project, would it be better to find the local minimum or local maximum of this function? Explain your reasoning.
8.
But the world is a lot more complicated than one number in and one number out! We can actually extend this concept to 3D! Don't worry about all the underlying math, I want to focus on intuition. A 3D plot can show two inputs on the floor and an output as height. Below is the 3D equivalent of a parabola called a paraboloid:

Click and drag to rotate the paraboloid!

Why do you think the real world doesn't normally use 2D functions for science and engineering?
9.
Let's say you could walk inside that paraboloid above, but you must be blindfolded. If you didn't exactly know where the bottom is, how would you figure out where the lowest point is? What concept from the start of the lesson helps us solve this question? Hint: There is no equation for this problem.
The "Blindfolded in a Valley" Analogy

Drag to orbit the valley. The person stays visible as you compare the terrain from different angles.

10.

Yes! The world is messy and we can't just plug stuff into equations. As we have a bunch of data, we can't rely on the quadratic formula, completing the square, etc. I mean it would be insane to solve quartic equations by hand with this nightmare formula:

So what in the world do we do? We use computers!

Computers can crunch a bunch of math insanely fast. We can also use context clues to help computers make a lot of rapid guesses back-to-back, so they get closer to the output we want.

Remember the concept of local minima/maxima and how we talked about slope? Those concepts all exist in 3D too.

So how do you think a computer can find the local minimum of a function if all we know is the input, output, and slope of where we are "standing"? Hint: Imagine you are walking in the paraboloid from the last page trying to find the very bottom.
Paraboloid surface

At each point, we can calculate the slope at that point and use it to decide which downhill step to take.

11.

Spoiler for Q10: Computers can think like someone standing at the edge of a paraboloid, and make small steps downward until it knows the slope is flat. Remember that line we drew through the vertex of the parabola from Q4? Computers can figure that out, but in 3D too.

In Algebra II mode, we can say the slope is flat when the secant plane has no incline, like a perfectly level sheet across the bowl.

Computers really just make a bunch of educated guesses until it gets close enough to the result we want: the lowest point of the paraboloid.

This is the concept of Gradient Descent, a foundational reason why ChatGPT works!!

So instead of walking inside the paraboloid, let's imagine we have a super weird bowl with marbles inside. We drop marbles somewhere randomly in this bowl, and wherever the marbles end up landing decides how much money is saved. The deeper inside the bowl they end up, the better.

We have two holes. Think of the deepest hole as the global minimum and the other hole is a local minimum.

Let's say you could control the bowl with your hands and you want as many marbles as possible to end up in the deeper hole. You are blindfolded once again, so you can't tilt the bowl in the direction you know has the deeper hole.

What could you do to try to get the most marbles in the deepest hole after the marbles are dropped? Could anything be done?

The 3D Marble Bowl

Marbles roll across a surface with a deepest point (global minimum) and shallower dents (local minima). Drag to orbit, then try the shaking slider.

Shaking illustrates wandering and settling. Actual SGD uses randomly selected training examples to estimate a downhill direction; it does not add arbitrary shaking or guarantee a better minimum.

Green ring: deepest (global) minimum. Gold rings: shallower local traps. Drag, use arrow keys, or choose Auto Rotate.

Noise: 28%
Marbles in Global Minimum: 0

Loop running: marbles reset automatically.

12.

You just learned the concept of Stochastic Gradient Descent! You can now connect slopes, small updates, and prediction errors to training AI!

Stochastic is a fancy word for random.

In actual SGD, the computer selects a random group of training examples and uses their prediction errors to estimate a downhill direction. Different groups can suggest different steps; more randomness does not guarantee a better answer.

Some shakiness is good, but too much causes the marbles to fall out, so we need just the right amount.

The moving marbles illustrate wandering and settling. With a fixed dataset and error measure, fitting moves a candidate across a fixed error surface. The shaping bowl later is a visual story about learning patterns. Temperature is different: it matters later, when ChatGPT is actually writing its response.

Instead of marbles in weird bowl holes, instead picture the depth of the hole as reducing pollution, saving resources, etc.

In your own words, how can random movement help in this bowl illustration, and why might too much prevent the marbles from settling?

You think that’s crazy? It’s about to get nuts.

13.
What in the world does this have to do with ChatGPT?
AI like ChatGPT stores how words relate to each other as a whole bunch of numbers organized in a specific way. Check this out:

In this first example, ChatGPT stores the concept of “Royalty” on the x-axis and “Gender Identity” on the y-axis. Word vector example with royalty and gender axes

In this second example, ChatGPT stores the concept of “Activity Intensity” on the x-axis and “Temperature” on the y-axis. Word vector example with activity intensity and temperature axes Word vector example with concept arithmetic

Here, we can see how we can perform math operations between word concepts in the way ChatGPT thinks. If we go by the coordinate pairs, we can calculate: Monarch = Queen - Woman.

Before we jump into four dimensions, it helps to think about Flatland, written by Edwin Abbott Abbott: a world whose inhabitants live on a two-dimensional plane. Its class depictions are also a satire of the rigidity of Victorian-era social hierarchy. A 3D object passing through that plane creates a changing 2D section where the object meets it. A shadow is a different kind of view: a projection can merge points at different depths. The Beyond 3D bonus lets you experiment with both.
Flatland projection illustration showing a higher-dimensional object as a lower-dimensional view Flatland inhabitants discussing how a higher-dimensional object appears in their two-dimensional world
Vector projection is the same core idea in math: we take something with more information and ask, "How much of it points in this direction?" A word vector can be projected onto axes like royalty, temperature, or activity intensity, and each projection gives us one meaningful part of the full concept.

Now let’s add one more dimension. Below, the words sit in 3D concept space, but they only "exist" when the slider matches their hidden 4th dimension. You can choose different "lenses" to see how AI organizes concepts like Flavor, Vibe, or Arcane Power.
4D Word Vector Visualization: The Flavor-Scanner

Drag to rotate. Move the slider to scan the 4th dimension (Edibility). Watch as words morph into view and grow in size when the "Flavor-Scanner" hits their specific coordinate.

Scanning Edibility: 72%
Cosmic / Massive Microscopic / Life Synthetic / Artifacts Culinary Concepts

So ChatGPT does all of this, but instead of pairs of numbers, it is able to perform these operations on lists of thousands of numbers at a time. This is because there are so many more relationships between words outside of “Royalty” and “Gender Identity” or “Activity Intensity” and “Temperature”.

So instead of 2D space (with parabolas) or even 3D space (with paraboloids)...

We are working with functions in THOUSAND-PLUS DIMENSIONAL SPACE!!!! Illustration of high-dimensional space

What does it mean to "train AI"? We are using two analogies. One analogy is words coming into place. Another analogy is a bowl being shaped.

Both visuals are simplified ways to imagine training: the model is gradually learning patterns from data.
What Does it Mean to Train AI?

This animation is a depiction of how ChatGPT is "trained" to know what to say. Words coming into place and a bowl being shaped are both analogies for what it means to train AI.

Words Falling Into Place
The Bowl Being Shaped
What do you think happens to the variety of output choices when we raise or lower temperature? How is choosing an output different from training the model? Use the next animation to test your prediction.
14.
How ChatGPT Generates an Output

After training, the model uses the patterns it learned to choose one word after another. The word space helps represent possible next words, and the shaped bowl is an analogy for the trained model already being ready to guide the output.

Temperature: 38% - coherent with some variety
Words Selected in Order
Trained Bowl Guiding Choices
Next Word Probability
To conclude, we started by talking about how to calculate the slope of a function, and we ended with the insanity of dropping marbles in strange, thousand-plus-dimensional bowls to explain how ChatGPT works. What you are learning in class now aligns with the foundation of not only AI, but how the world works in SO many ways. Below, please connect ONE topic you learned in class to ONE new idea you learned today.

P.S.: What we talked about today is exactly why TikTok and Instagram Reels are so addicting and the Oxford Word of the Year is “Rage Bait”. The “marbles at the bottom of the bowl” for social media have to do with watch time, comments, and shares. Outrage and dishonesty is more effective at this than positivity and the full truth.

One last note: examples like walking down a hill blindfolded, shaping a bowl, and dropping marbles into a bowl are analogies. They are not a perfect representation of how AI works, but they help get the general ideas across.

Thank you so much for listening and participating!!

About the Speaker (Griffin Rutherford)

I hold a Bachelor and Master of Science in Computer Science from Colorado School of Mines. People say I have "golden retriever energy", and I love being physically active. I am an avid weightlifter, trail cyclist, snowboarder, runner, and more. I am now the CTO of Coherascent Labs, an education technology and research company. I was a member of the Alpha Tau Omega fraternity as Philanthropy Chair, and I lived in the frat house basement for two years (at the expense of my GPA). I also love to sing, write, and enjoy long conversations with friends and family.

Alpine Lake trail selfie