Self-Attention

Turkish equivalent: Öz-dikkatDomain: Artificial Intelligence

An attention operation in which tokens attend to other tokens within the same sequence to build context-dependent representations.

Artificial-Intelligence Context

In self-attention, each token forms a query and compares it with keys from tokens in the same sequence, then combines the corresponding values into a context-dependent representation. A causal mask prevents access to future positions in autoregressive models.

Computational Boundary

Standard full self-attention constructs interactions that grow quadratically with sequence length. Modern kernels can reduce memory traffic and improve implementation efficiency, but they do not remove the underlying all-pairs dependency of full attention.

Direct source: The primary paper or official specification for Self-Attention is linked here for verification.