Pipeline Parallelism
A distributed-training or inference strategy that partitions model layers into stages and overlaps processing of multiple microbatches across devices.
ML-Systems Context
Pipeline parallelism partitions model layers into stages placed on different devices and overlaps several microbatches across those stages. It can reduce per-device memory requirements for very large models, but pipeline bubbles, stage imbalance, and cross-stage communication constrain throughput.
Terminology Boundary
Model pipeline parallelism is distinct from the general software pipeline pattern. Here the object being partitioned is the neural-network computation graph.
Related ML-Systems Concepts
- Tensor Parallelism
- Data Parallelism
- Throughput
- Load Balancing
Direct source: The primary paper or official specification for Pipeline Parallelism is linked here for verification.