Transformer
A sequence-model architecture built primarily around attention, feed-forward blocks, residual connections, and normalization rather than recurrence.
Architecture
A transformer processes token representations through attention, feed-forward blocks, residual connections, and normalization rather than relying on recurrence as the primary sequencing mechanism. Its ability to train many positions in parallel and model long-range interactions helped make it a foundation for language, vision, and multimodal systems.
Conceptual Boundary
A transformer is not synonymous with a large language model. Model scale, training objective, tokenization, data, positional encoding, and decoding strategy are separate design dimensions.
Related Concepts
Source
- https://arxiv.org/abs/1706.03762
Related article: Pattern and Text Recognition from Images
Related technical articles: Visual Pattern and Text Recognition, Artificial Intelligence: Philosophy, Theory and Practice.