Logistic regression
Bend a line into a probability. Move the boundary and watch the log loss react.
Bend a line into a probability. Move the boundary and watch the log loss react.
Tap empty space to add a Class 1 point; drag points to move them. Drag the square to slide the boundary and the circle to turn it.
Logistic regression is a classifier. It first computes a score z = w₁x₁ + w₂x₂ + b, exactly like a line. Points where z = 0 form the decision boundary; the arrow shows the side where z is positive.
The sigmoid σ squashes any score into a probability between 0 and 1: z = 0 gives 0.5, big positive scores get close to 1, big negative ones close to 0. Making all the weights bigger keeps the same boundary but makes the model more confident: the dashed p = 0.1 and 0.9 lines squeeze in.
Log loss charges −ln(p) for the probability given to the true class. A point the model is unsure about (p = 0.5) costs 0.69; a confident mistake (p = 0.01) costs 4.6. Training is gradient descent on this loss, and the gradient is beautifully simple: each point pushes the weights by its error (p − y) times its features.
Point inspector
Tap a point (or use the arrows) to see its score, probability and loss.
Logistic regression is a classifier. It first computes a score z = w₁x₁ + w₂x₂ + b, exactly like a line. Points where z = 0 form the decision boundary; the arrow shows the side where z is positive.
The sigmoid σ squashes any score into a probability between 0 and 1: z = 0 gives 0.5, big positive scores get close to 1, big negative ones close to 0. Making all the weights bigger keeps the same boundary but makes the model more confident: the dashed p = 0.1 and 0.9 lines squeeze in.
Log loss charges −ln(p) for the probability given to the true class. A point the model is unsure about (p = 0.5) costs 0.69; a confident mistake (p = 0.01) costs 4.6. Training is gradient descent on this loss, and the gradient is beautifully simple: each point pushes the weights by its error (p − y) times its features.
Things to try