Rotary Position Embedding

Turkish equivalent: Döner konum gömmeDomain: Artificial Intelligence

A positional encoding technique that rotates query and key vectors so relative token positions are represented directly in attention geometry.

Artificial-Intelligence Context

Rotary Position Embedding applies position-dependent rotations to query and key vectors so relative position is expressed through attention geometry. RoPE is widely used in long-context transformer models, and context-extension methods often modify the rotation-frequency schedule beyond the original training range.

Architecture Boundary

RoPE is not an attention mechanism by itself. It is a positional component that changes how token position is represented inside attention.

What the Rotation Encodes

RoPE does not merely add a separate position vector to an embedding. It applies position-dependent rotations to pairs of query and key components. The dot product between rotated query/key vectors consequently depends on their relative offset as well as their absolute positions. The important property is therefore not the label "long context" by itself, but the fact that positional information is embedded into attention geometry.

Extrapolating beyond the training context is a separate problem. Context-extension methods that rescale frequencies or positions are not identical to RoPE itself, and they do not guarantee unchanged quality outside the training range. The mathematical formulation and relative-position property are given directly in the RoFormer paper by Su et al.: original source. Related concepts are Self-Attention and Context Window.

Related technical publications

Publications whose title or summary directly references this concept.