zerodp. ZeroDP implements an efficient zero-copy data parallel approach for serving Mixture-of-Experts (MoE) models, where expert weights are shared across data parallel ranks via CUDA IPC (Inter-Process Communication) handles. This eliminates the need to duplicate expert weights across GPUs, significantly reducing memory requirements.

github.com/doublewordai/zerodp

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.