Image convolution
Slide a 3×3 kernel over pixels: blur, sharpen and detect edges by hand.
Slide a 3×3 kernel over pixels: blur, sharpen and detect edges by hand.
Output pixel (3, 3) = Σ kernel × input
sum = 510 → shown as 255 (clipped)
Convolution slides a small grid of weights, the kernel, across the image. At each position it multiplies the 9 pixels underneath by the 9 weights and adds them up to make one output pixel. Weights that sum to 1 keep brightness (blur, sharpen); weights that sum to 0 respond only to change, so flat areas go to 0 and edges light up.
Without padding the kernel can’t centre on border pixels, so the output shrinks; zero padding keeps the size and stride 2 skips every other position. CNNs learn these weights instead of hand-picking them (and, like here, don’t flip the kernel). Outputs are clipped to 0–255 for display, so negative responses show as black.
3×3 kernel
(10 + 2·0 − 3) / 1 + 1 = 8
Convolution slides a small grid of weights, the kernel, across the image. At each position it multiplies the 9 pixels underneath by the 9 weights and adds them up to make one output pixel. Weights that sum to 1 keep brightness (blur, sharpen); weights that sum to 0 respond only to change, so flat areas go to 0 and edges light up.
Without padding the kernel can’t centre on border pixels, so the output shrinks; zero padding keeps the size and stride 2 skips every other position. CNNs learn these weights instead of hand-picking them (and, like here, don’t flip the kernel). Outputs are clipped to 0–255 for display, so negative responses show as black.
Things to try