Skip to content

Model Parallelism Research

📊 Data & AI
⏱️ ~1 week 📚 Prerequisites: Distributed Training

When you'd use this

Split a model too big for one device across multiple GPUs.

Split a model too big for one device across several — for very large networks.

What you'll learn

  • Data parallelism vs model parallelism
  • When a model doesn't fit on one GPU
  • Tensor and layer splitting
  • The communication cost
  • Frameworks

Modern models (large LLMs) can be too big to fit in a single GPU's memory. Model parallelism splits the model itself across multiple devices. This contrasts with data parallelism, which splits the data.

This is a GPU/distributed topic

Model parallelism requires multiple GPUs and frameworks (PyTorch, DeepSpeed, Megatron) not available here. This page is conceptual, following documented framework approaches.


Data vs model parallelism

Data vs model parallelism in Model Parallelism — what it is and when to use it.

Two fundamentally different ways to parallelize training:

   DATA PARALLELISM                    MODEL PARALLELISM
   same model on each GPU               model SPLIT across GPUs
   different data batch per GPU         same data flows through the pieces
   [model][model][model]                [layer1-2][layer3-4][layer5-6]
     GPU0    GPU1   GPU2                    GPU0      GPU1      GPU2
   → for models that FIT on one GPU      → for models TOO BIG for one GPU
  • Data parallelism (see Distributed Training) — replicate the whole model on each GPU, feed each a different slice of the batch, average the gradients. Simple and common — but requires the model to fit on one GPU.
  • Model parallelism — when the model is too big to fit, split it across GPUs.

Ways to split a model

By layer (pipeline) or within a layer (tensor) across devices.

  • Tensor parallelism — split individual layers across GPUs. A big matrix multiply is divided so each GPU computes part of it, then results are combined. Fine-grained; heavy communication.
  • Layer (pipeline) parallelism — put different layers on different GPUs. Layer 1-2 on GPU0, 3-4 on GPU1, etc. Data flows through like an assembly line (see Pipeline Parallelism).
  • Expert parallelism — for mixture-of-experts models, put different "expert" sub-networks on different GPUs.

Large-model training often combines all of these ("3D parallelism": data + tensor + pipeline).


The communication cost

Splitting adds cross-device traffic that can dominate — balance carefully.

Communication is the bottleneck

Splitting a model means GPUs must constantly exchange intermediate results over their interconnect. This communication can dominate — a naive split can be slower than one GPU because the devices spend more time talking than computing. The engineering challenge is minimizing and overlapping communication with computation. This is why high-end training uses fast interconnects (NVLink, InfiniBand) — the network, not the math, is often the limit.


Frameworks

DeepSpeed, Megatron, and FSDP that implement these splits.

You don't implement this by hand — specialized frameworks do (documented; not installed here):

Framework Role
PyTorch FSDP Fully Sharded Data Parallel — shards model across GPUs
DeepSpeed (Microsoft) ZeRO optimizer sharding, pipeline + tensor parallelism
Megatron-LM (NVIDIA) Tensor/pipeline parallelism for huge transformers
torch.distributed The lower-level primitives
# PyTorch FSDP (documented API)
from torch.distributed.fsdp import FullyShardedDataParallel as FSDP
model = FSDP(model)     # shards params/gradients/optimizer state across GPUs

These handle the splitting, communication, and gradient synchronization so you configure rather than hand-code it.


When you need it

A core question explored in Model Parallelism: When you need it.

  • You need it when a model + its activations + optimizer state exceed one GPU's memory (large LLMs, huge vision models).
  • You don't for models that fit on one GPU — use simpler data parallelism to go faster.

Most practitioners use data parallelism; model parallelism is for the frontier of large-model training. It connects directly to Pipeline Parallelism (one splitting strategy) and Distributed Training (the broader topic).


Practice exercises

  1. Explain the difference between data and model parallelism and when each applies.
  2. Describe why communication cost can make a naive model split slower than a single GPU.
  3. Compare tensor parallelism (split within a layer) and pipeline parallelism (split across layers).
  4. Estimate roughly: a model with 10B float32 params needs how much memory just for weights? Why might it not fit on one GPU?
  5. Research PyTorch FSDP or DeepSpeed ZeRO and summarize what it shards across GPUs.

💬 Discussion

Have a question about this topic? Found an error? Share your thoughts below.