A perfect score on the practice questions

Imagine a student who memorises last year's exam, word for word. Give them the same questions and they score 100%. Give them new questions on the same subject and they fall apart, because they never learned the subject. They learned the exam.

Computers that learn from data can fail in exactly the same way. The failure has a name: overfitting. It is one of the first ideas worth knowing if you ever touch machine learning, and you can watch it happen on a single chart.

Learning, in its smallest form, is fitting a curve

Machine learning sounds mysterious, but its smallest version is something you may have drawn at school: points on a graph, and a line through them. The points are the training data, the examples the computer learns from. The line or curve is the model, a rule that guesses an output y for any input x. Training means picking the curve that sits closest to the points.

"Closest" needs a number, so we measure an error: roughly, how far off the curve is from the points on average. (Precisely, it is the root-mean-square error: square each miss, average the squares, take the square root. Smaller is better, and 0 means every point is hit exactly.)

Real measurements are never perfect. The 10 training points below come from a smooth wave plus random noise, the little wobbles you get from imperfect measuring. The wave is the pattern we hope the model finds. The noise is what it should ignore.

How flexible should the model be?

Here the model is a polynomial curve, and its degree controls how flexible it is. Degree 1 is a straight line. Degree 2 is a parabola, with one bend. Each step up allows one more bend: a degree-9 curve can turn up to 8 times, and with 10 points it has enough freedom to pass through every single one of them.

Try it. Slide the degree, then drag the green points up and down and watch the curve react.

training points (drag them)test pointsfitted curve- - - the hidden wave
error on training points0
error on test points0
Green = error on the points it learned from. Orange = error on 20 fresh points it never saw. Tap a pair of bars to jump to that degree.

What to look for

Start at degree 1. A straight line cannot follow a wave, so both errors are high (about 0.48 on the training points and 0.51 on the test points). That is underfitting: the model is too stiff to capture the pattern.

Move to degree 3 and the curve finds the wave. The errors drop to about 0.17 and 0.32. From degree 3 to degree 7 the training error creeps down a little (0.17 to 0.13) while the test error stays near 0.3. Extra flexibility is not hurting yet, but it is not helping either.

Now go to degree 8, then 9. The training error falls to 0.05, then to exactly 0. By that measure the model is perfect. But the test error jumps to 1.16 and then 1.47, about five times worse than at degree 7. The curve now threads every noisy point and swings wildly in the gaps and near the edges. That is overfitting: the model learned the noise, not the wave.

Notice the floor, too. Even a perfect model would score around 0.25 on the test points, because those points carry their own random noise that nothing can predict. A test error of 0.3 is close to as good as this data allows.

One more experiment: set degree 9 and drag a single point up a little. The whole curve lurches. Do the same at degree 3 and the curve barely moves. An overfitted model is not only wrong on new data, it is also jumpy, because it treats every wobble as meaningful.

Why the test points matter

Look at the green numbers alone and degree 9 wins every time. Training error rewards memorising, so it can never warn you about overfitting. The only error that says anything about new data is the one measured on data the model did not see while learning. That is the test set: examples held back on purpose.

The habit is simple and worth keeping: split your data before you start, train on one part, score on the other, and never let the test part influence the training. When the training score is excellent and the test score is poor, you have found overfitting.

The same experiment in code

This is the Python behind the chart, using the NumPy library's polyfit to do the training. It prints both errors for three degrees:

import numpy as np

# 10 training points (a wavy curve plus random noise)
train_x = np.array([-0.89, -0.668, -0.478, -0.322, -0.116, 0.13, 0.26, 0.526, 0.724, 0.897])
train_y = np.array([-0.216, -0.774, -0.971, -1.08, -0.364, 0.571, 0.394, 0.882, 0.288, -0.006])

# 20 test points the model never sees while it learns
test_x = np.linspace(-0.95, 0.95, 20)
test_y = np.array([-0.617, -0.513, -1.024, -0.823, -0.949, -1.034, -1.52, -0.842, -0.466, -0.128,
                   -0.226, 0.335, 0.462, 0.689, 1.253, 0.786, 0.883, 0.928, 0.308, 0.129])

def rmse(guess, truth):
    return np.sqrt(np.mean((guess - truth) ** 2))

for degree in (1, 3, 9):
    coeffs = np.polyfit(train_x, train_y, degree)       # "training"
    train_err = rmse(np.polyval(coeffs, train_x), train_y)
    test_err = rmse(np.polyval(coeffs, test_x), test_y)
    print(f"degree {degree}: train error {train_err:.3f}, test error {test_err:.3f}")

Running it prints:

degree 1: train error 0.478, test error 0.507
degree 3: train error 0.171, test error 0.323
degree 9: train error 0.000, test error 1.466

These match the numbers in the chart. The training error only ever goes down as the degree goes up. The test error goes down, then back up.

So what for your own projects?

A model with more adjustable parts needs more data to keep it honest. The degree-9 curve has 10 adjustable numbers and only 10 points to learn from, so it has no choice but to memorise them. With many more points, the same curve would be held in place by the data.

When data is scarce, a simpler model is often the better one. And a big gap between the training score and the test score is the warning sign to look for, whether the model is a ten-point curve or something much larger. Real systems have far more adjustable numbers than this one, and the teams that build them hold back examples for the same reason you just saw. How they keep overfitting in check is a story for another read, but the question they ask is one you can ask too: how does it do on data it has never seen?