Transformer

Turkish equivalent: Transformer mimarisiDomain: Artificial Intelligence

A sequence-model architecture built primarily around attention, feed-forward blocks, residual connections, and normalization rather than recurrence.

Architecture

A transformer processes token representations through attention, feed-forward blocks, residual connections, and normalization rather than relying on recurrence as the primary sequencing mechanism. Its ability to train many positions in parallel and model long-range interactions helped make it a foundation for language, vision, and multimodal systems.

Conceptual Boundary

A transformer is not synonymous with a large language model. Model scale, training objective, tokenization, data, positional encoding, and decoding strategy are separate design dimensions.

Source

  • https://arxiv.org/abs/1706.03762

Related article: Pattern and Text Recognition from Images

Related technical articles: Visual Pattern and Text Recognition, Artificial Intelligence: Philosophy, Theory and Practice.

Related technical publications

Publications whose title or summary directly references this concept.

Large Language Models

A comprehensive course note on large language models from language modeling and Transformer architecture through pre-training, post-training, RAG, fine-tuning, reasoning, tool use, agents, security, evaluation, and production engineering, connected to other AI paradigms.

Natural Language Processing

A Turkish-centered NLP course covering Unicode and normalization, tokenization, morphology, syntax, semantics, statistical NLP, Transformers, multilingual models, speech, LLMs, and production engineering in a comparative language perspective.

Artificial Intelligence: Philosophy, Theory and Practice

A course note connecting the philosophical roots of artificial intelligence with definitions of intelligence, logic and symbolic AI, search, expert systems, probability, neural networks, NLP, biometrics, Transformers, RAG, safety, agents, and production engineering.

Whisper Architecture in Speech Recognition Systems

Whisper converts 16 kHz audio to a log-Mel spectrogram and generates tokens autoregressively with an encoder-decoder Transformer; language, task, and timestamp information share the same token sequence.