Pipeline Parallelism

Turkish equivalent: Pipeline paralelliğiDomain: Machine Learning Systems

A distributed-training or inference strategy that partitions model layers into stages and overlaps processing of multiple microbatches across devices.

ML-Systems Context

Pipeline parallelism partitions model layers into stages placed on different devices and overlaps several microbatches across those stages. It can reduce per-device memory requirements for very large models, but pipeline bubbles, stage imbalance, and cross-stage communication constrain throughput.

Terminology Boundary

Model pipeline parallelism is distinct from the general software pipeline pattern. Here the object being partitioned is the neural-network computation graph.

Direct source: The primary paper or official specification for Pipeline Parallelism is linked here for verification.