Self-Attention
An attention operation in which tokens attend to other tokens within the same sequence to build context-dependent representations.
Artificial-Intelligence Context
In self-attention, each token forms a query and compares it with keys from tokens in the same sequence, then combines the corresponding values into a context-dependent representation. A causal mask prevents access to future positions in autoregressive models.
Computational Boundary
Standard full self-attention constructs interactions that grow quadratically with sequence length. Modern kernels can reduce memory traffic and improve implementation efficiency, but they do not remove the underlying all-pairs dependency of full attention.
Related AI Concepts
Direct source: The primary paper or official specification for Self-Attention is linked here for verification.