How Computers See Images
To a computer, a digital image is nothing more than a 2D or 3D grid of numbers. A grayscale image is a 2D matrix of pixel intensity values ranging from 0 (black) to 255 (white). A color RGB image consists of three stacked matrices corresponding to Red, Green, and Blue color channels.
Image Processing with OpenCV
OpenCV (Open Source Computer Vision Library) provides hundreds of optimized algorithms for basic and intermediate visual tasks: image resizing, color space conversions (RGB to HSV/Grayscale), thresholding, and morphological filtering.
import cv2
# Read an image from disk
img = cv2.imread('sample.jpg')
# Convert color to grayscale
gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY)
# Apply Canny Edge Detection
edges = cv2.Canny(gray, threshold1=100, threshold2=200)
# Save result
cv2.imwrite('edges.jpg', edges)
Convolutional Neural Networks (CNNs)
While traditional computer vision relied on manually engineered filters (e.g., Sobel filters for edge detection), Convolutional Neural Networks automatically learn optimal spatial filters directly from training datasets.
- Convolutional Layers: Slide small weight matrices (kernels) across input images to create feature maps detecting edges, textures, and higher-order shapes.
- Pooling Layers: Reduce the spatial dimensions of feature maps (e.g., Max Pooling) to achieve translation invariance and decrease computational cost.
- Fully Connected Layers: Flatten feature maps into a classification vector that predicts category probabilities.
Conclusion
Computer Vision bridges the gap between raw pixel arrays and semantic understanding, powering applications in medical imaging, autonomous vehicles, face recognition, and industrial robotics.