Multi-Head Attention
An attention mechanism that runs several learned attention projections in parallel to capture different representation subspaces.
Artificial-Intelligence Context
Multi-head attention applies several learned query, key, and value projections in parallel so different heads can operate in different representation subspaces. Their outputs are concatenated or combined and projected back into the model dimension.
Architecture Boundary
Increasing head count does not automatically improve a model. Head dimension, total hidden size, KV sharing, model depth, and training objective jointly determine capacity and serving cost.
Related AI Concepts
Direct source: The primary paper or official specification for Multi-Head Attention is linked here for verification.