Image Feature Vectors and Matching
Local and global image descriptors, similarity metrics, dimensionality reduction and matching strategies for large visual collections.
Image features transform raw pixels into representations that are easier to compare, index and classify. The objective is not only dimensionality reduction but preservation of task-relevant discriminative information.
Local and global descriptors
Global descriptors summarise the complete image through colour, texture or learned embeddings. Local descriptors represent neighbourhoods around selected keypoints. SIFT, SURF and ORB are classical examples of this approach.
Once an image is mapped to a feature vector, matching becomes a distance or similarity problem in representation space.
Similarity metrics
Euclidean distance, cosine similarity and Hamming distance suit different descriptor families. Binary descriptors naturally use Hamming distance, while normalised dense embeddings often use cosine similarity.
A matching pipeline may include preprocessing, keypoint selection, descriptor computation, candidate retrieval, geometric verification and thresholding. Methods such as RANSAC can reduce false candidates by enforcing geometric consistency.
Large collections
Linear scanning becomes expensive when millions of vectors are stored. Approximate-nearest-neighbour indices, quantisation and hierarchical search reduce query cost. Vector dimension also affects storage, cache behaviour and index size.
Feature extraction is therefore a representation problem designed jointly with the similarity metric and retrieval structure.