Image Segmentation
Image segmentation from thresholding and morphology to learned encoder-decoder models, with emphasis on mask quality, annotation noise and runtime cost.
Image segmentation assigns pixels or regions to meaningful classes. Unlike classification, the result preserves spatial location and can feed directly into measurement, inspection, medical imaging and scene-analysis pipelines.
Classical methods
Thresholding separates regions according to intensity or another scalar feature. Global thresholds may work under stable illumination, while local thresholds are more robust under spatial variation. Otsu's method selects a threshold by maximising between-class variance.
Region growing, watershed and connected-component analysis provide more structural alternatives. Morphological opening, closing, erosion and dilation are commonly used to regularise masks.
Learned segmentation
Image segmentation is often implemented with encoder-decoder architectures such as U-Net. Semantic segmentation assigns class labels, instance segmentation separates individual objects, and panoptic segmentation combines both views.
Intersection over Union and Dice score are common overlap metrics. Pixel accuracy alone may be misleading when background pixels dominate.
Computation and annotation quality
High-resolution masks require substantial memory and computation. Tiled processing can preserve small structures but adds merge overhead. Inconsistent annotation boundaries also limit attainable model quality.
A segmentation mask is usually an intermediate representation consumed by tracking, OCR, measurement or another decision stage.