Grouped-Query Attention
An attention variant in which multiple query heads share fewer key/value heads to reduce memory and inference cost.
Artificial-Intelligence Context
Grouped-query attention lets several query heads share a smaller number of key/value heads. The design aims to preserve much of the representational flexibility of multi-head attention while reducing KV-cache size and memory-bandwidth demand during inference.
Architecture Boundary
GQA and Mixture of Experts optimize different parts of a model. GQA changes sharing among attention heads; MoE routes tokens among feed-forward experts.
Related AI Concepts
Structural Boundary Between MHA and MQA
GQA is easier to reason about as an intermediate point between two attention designs. Let H_q be the number of query heads and H_kv the number of key/value heads. Conventional multi-head attention uses H_kv = H_q; multi-query attention uses H_kv = 1; grouped-query attention operates in the region 1 < H_kv < H_q, so groups of query heads share key/value heads.
This sharing matters directly to KV Cache memory during autoregressive inference: with the other dimensions fixed, stored K/V state scales approximately with H_kv. Fewer KV heads do not guarantee the same quality or throughput on every model because kernels, batch shape, sequence length, and memory bandwidth still matter. Ainslie et al. define this intermediate design and its uptraining method directly: paper.