🧠 · Foundations

CNN (vision)

The architecture that taught machines to see, one local pixel patch at a time.

In one line

CNNs learn visual features by sliding small filters across an image instead of looking at every pixel at once.

ConceptWhat it is

A Convolutional Neural Network (CNN) processes images by sliding small learnable filters across the pixel grid, detecting edges, textures, and shapes in early layers and combining them into objects in deeper layers. It exists because treating an image as a flat list of pixels ignores the fact that nearby pixels are related and that a cat's ear looks the same whether it is in the top-left or bottom-right of the frame.

This local connectivity and weight sharing makes CNNs dramatically more parameter-efficient than a fully connected network for grid-like data, which is why they dominated computer vision for a decade before vision transformers arrived.

How it worksThe mechanics

A convolutional filter, typically 3x3 or 5x5 pixels, slides across the input image computing dot products at every position, producing a feature map that highlights where a pattern like an edge appears. Stacking many filters per layer and pooling layers that downsample between them lets the network build up from edges to textures to parts to whole objects, with the final layers feeding a classifier or detector head.

At a glanceSee it

CNN (vision) diagram
CNN (vision) diagram 1

Inside one convolution — multiply the patch by the filter, sum, add bias, activate, then slide reusing the very same weights, which is exactly what makes the response shift-equivariant.

CNN (vision) diagram 2

How depth defeats locality — each deeper unit pools the ones below it, so its receptive field widens layer by layer until far-apart pixels, which never talk directly, finally interact.

When to use itWhere it fits

  • Image classification, object detection, and segmentation where spatial locality matters.
  • Latency-sensitive edge or mobile vision, since CNNs are cheap to run compared to large vision transformers.
  • Medical imaging and industrial defect detection with modest labeled datasets.
  • As a lightweight backbone inside a larger multimodal pipeline.

When NOT to use itLimits & anti-patterns

  • Long-range spatial reasoning across a large image, where limited receptive fields miss global context that attention captures naturally.
  • Sequence or language tasks, where there is no 2D grid structure to exploit.
  • Extremely large-scale pretraining where vision transformers now generally out-scale CNNs given enough data.

Trade-offsAdvantages & costs

Advantages
  • Parameter-efficient thanks to weight sharing across spatial positions.
  • Fast inference and well-optimized hardware kernels on GPUs and mobile NPUs.
  • Strong inductive bias for images means good accuracy on smaller datasets.
  • Mature tooling and decades of production hardening.
Trade-offs & costs
  • Limited receptive field means global context requires many stacked layers.
  • Less flexible than transformers for multimodal fusion with text.
  • Performance plateaus versus vision transformers at very large data and compute scale.
  • Architecture search for the right depth and filter sizes is still somewhat manual.

ExampleIn the real world

Tesla's Autopilot vision stack used CNN backbones like HydraNets to detect lanes, vehicles, and pedestrians from camera feeds in real time on car hardware.

ToolsHow to implement it

  • PyTorch torchvisionready-made CNN backbones like ResNet and EfficientNet.
  • TensorFlow/Kerasproduction-grade training and mobile export via TFLite.
  • ONNX Runtimeportable, fast inference across edge devices.
  • NVIDIA TensorRThardware-accelerated CNN inference for real-time vision.

Cost & effortWhat it takes

Low to moderate training cost on commodity GPUs; inference is cheap and low-latency, even on edge devices; needs labeled image data but far less than a vision transformer trained from scratch.

A living map of modern AI — kept current every morning