Multi-Head Attention

Turkish equivalent: Çok başlı dikkatDomain: Artificial Intelligence

An attention mechanism that runs several learned attention projections in parallel to capture different representation subspaces.

Artificial-Intelligence Context

Multi-head attention applies several learned query, key, and value projections in parallel so different heads can operate in different representation subspaces. Their outputs are concatenated or combined and projected back into the model dimension.

Architecture Boundary

Increasing head count does not automatically improve a model. Head dimension, total hidden size, KV sharing, model depth, and training objective jointly determine capacity and serving cost.

Direct source: The primary paper or official specification for Multi-Head Attention is linked here for verification.