Iniciar sesión

WAP: Weekly AI Papers

WAP: Weekly AI Papers

DeepSeek V3

14 min · 8 de ene de 2025

Portada del episodio DeepSeek V3

Descripción

DeepSeek-V3, a 671B-parameter Mixture-of-Experts large language model. It covers the model's architecture, including Multi-Head Latent Attention and an innovative auxiliary-loss-free load balancing strategy for DeepSeekMoE. The training process, encompassing pre-training on 14.8 trillion tokens and post-training using supervised fine-tuning and reinforcement learning, is described. paper: https://github.com/deepseek-ai/DeepSeek-V3/blob/main/DeepSeek_V3.pdf

Comentarios

0

Sé la primera persona en comentar

¡Regístrate ahora y únete a la comunidad de WAP: Weekly AI Papers!

Todos los episodios

1 episodios

DeepSeek V3

DeepSeek-V3, a 671B-parameter Mixture-of-Experts large language model. It covers the model's architecture, including Multi-Head Latent Attention and an innovative auxiliary-loss-free load balancing strategy for DeepSeekMoE. The training process, encompassing pre-training on 14.8 trillion tokens and post-training using supervised fine-tuning and reinforcement learning, is described. paper: https://github.com/deepseek-ai/DeepSeek-V3/blob/main/DeepSeek_V3.pdf

8 de ene de 202514 min