Overview

This repository provides the inference code for the DeepSeek-V3 series of large language models. DeepSeek-V3 is a Mixture-of-Experts (MoE) model featuring novel architectures like Multi-head Latent Attention (MLA) and DeepSeekMoE, optimized for efficient inference and training. Notably, it pioneers an auxiliary-loss-free load balancing strategy and utilizes Multi-Token Prediction (MTP) during training.

This documentation provides a technical breakdown of the inference codebase, focusing on its architecture, components, and usage.

On this page