Tensor Parallelism
A model-parallel strategy that splits individual tensor operations and model parameters across multiple devices that cooperate on each layer.
ML-Systems Context
Tensor parallelism partitions individual tensor operations and parameter matrices across cooperating devices. It is used when layers or model weights do not fit efficiently on one accelerator, or when the computation of each layer needs to be distributed.
Scaling Boundary
Tensor parallelism is not data parallelism: devices collaborate on the same sample and layer computation. Collective-communication latency and interconnect bandwidth therefore become direct scaling limits.
Related ML-Systems Concepts
Direct source: The primary paper or official specification for Tensor Parallelism is linked here for verification.