Data Parallelism

Turkish equivalent: Veri paralelliğiDomain: Machine Learning Systems

A distributed-training strategy that replicates a model across workers and partitions input batches while synchronizing parameter gradients or updates.

ML-Systems Context

Data parallelism replicates a model across workers and gives each replica a different part of the input batch. Training then synchronizes gradients or parameter updates, commonly through collective communication such as all-reduce; inference can use independent replicas to distribute requests.

Scaling Boundary

Data parallelism increases aggregate capacity but does not solve the problem of a single model that cannot fit in one device's memory. Communication cost, global batch size, and optimizer state also limit training scalability.