DeepSeek-V3 Technical Report
2024/12/27 by DeepSeek-AI, Aixin Liu, Liu, Aixin +404 · 39 voices · 6 citations
Computer Science · Engineering · #Distributed and Parallel Computing Systems #Robotics and Automated Systems #cs.AI #cs.CL
paper · pdf · doi:10.48550/arxiv.2412.19437
openalex publication_date 2024/12/27 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/29
Abstract
We present DeepSeek-V3, a strong Mixture-of-Experts (MoE) language model with 671B total parameters with 37B activated for each token. To achieve efficient inference and cost-effective training, DeepSeek-V3 adopts Multi-head Latent Attention (MLA) and DeepSeekMoE architectures, which were thoroughly validated in DeepSeek-V2. Furthermore, DeepSeek-V3 pioneers an auxiliary-loss-free strategy for load balancing and sets a multi-token prediction training objective for stronger performance. We pre-train DeepSeek-V3 on 14.8 trillion diverse and high-quality tokens, followed by Supervised Fine-Tuning and Reinforcement Learning stages to fully harness its capabilities. Comprehensive evaluations reveal that DeepSeek-V3 outperforms other open-source models and achieves performance comparable to leading closed-source models. Despite its excellent performance, DeepSeek-V3 requires only 2.788M H800 GPU hours for its full training. In addition, its training process is remarkably stable. Throughout the entire training process, we did not experience any irrecoverable loss spikes or perform any rollbacks. The model checkpoints are available at https://github.com/deepseek-ai/DeepSeek-V3.
Cited by
Discussions
- DeepSeek-V3 Technical Report [hn, 132 points, 34 comments]
- #Deepseek annonce-t-il la fin de #NVIDIA ? Deepseek affirme avoir égalé #ChatGPT avec un coût hardware/énergie de 5,6 millions USD$ (le coût de location des GPU). Les estimations me paraissent crédib [bsky, 61 points, 3 comments]
- Yeah, I think this is where that number comes from. It’s ONLY compute costs arxiv.org/pdf/2412.19437 [bsky, 6 points, 1 comments]
- R1 is based on DSV3, which had an exhaustive technical report. It's manual memory management and FP8 training. Not complicated stuff. arxiv.org/pdf/2412.19437 [bsky, 5 points, 1 comments]
- Finally some time to read the DeepSeek v3 paper, really nice. It's well written I find, not spending a lot of time on each of the new or derived ideas, as is sometimes the case for articles which have [bsky, 3 points, 1 comments]
- That number comes directly from their technical report. Page 5. [bsky, 3 points, 0 comments]
- DeepSeek-V3 Technical Report [hn, 3 points, 0 comments]
- DeepSeek-R1-zero starts from DeepSeek-V3-base as initial policy. This has MATH 61.6 and GSM8K 89.3 (v3 report Table 3). V2->3 pre-training "[enhances] ratio of mathematical and programming samples." I [bsky, 3 points, 0 comments]
- Here is the technical report - I'm not a specialist but there doesn't seem to be any reason beyond Altman's wounded pride not to take it at face value. arxiv.org/abs/2412.194... [bsky, 3 points, 2 comments]
- For those wondering what I'm on about, this Wikipedia article provides a good overview: en.wikipedia.org/wiki/Expert_... Also, Deepseek has published a paper on arXiv describing their Mixture of Exper [bsky, 3 points, 0 comments]
- arxiv.org/abs/2412.19437 [bsky, 2 points, 0 comments]
- ooffff! the "DeepSeek-V3 Technical Report" is info dense! arxiv.org/abs/2412.19437 [bsky, 2 points, 1 comments]
- I found the technical report for the DeepSeek model and wanted to just link it to this post for those that are interested! 👀 [bsky, 2 points, 0 comments]
- From DeepSeek-V3 Technical Report: “For non-reasoning data, such as creative writing, role-play, and simple question answering, we utilize DeepSeek-V2.5 to generate responses and enlist human annotat [bsky, 2 points, 1 comments]
- Here's the DeepSeek-V3 arXiv paper: arxiv.org/pdf/2412.19437 #deepseek #AI #LLMs #NLP #opensource [bsky, 2 points, 0 comments]
- Det er DeepSeek-V3 (og ikke deres reasoning model DeepSeek-R1, der performer på niveau med OpenAIs o1) som er trænet for $5.576M. Se hvordan budget er fordelt træningsfaser og mere om hvordan de har [bsky, 2 points, 1 comments]
- DeepSeek-V3 Technical Report arxiv.org/abs/2412.1... [bsky, 1 points, 0 comments]
- あとでよむ / DeepSeek V3のテクニカルレポート arxiv.org/abs/2412.19437 [bsky, 1 points, 0 comments]
- The Making Of paper for DeepSeek V3 is wild. OpenAI bragging about how much money they're lighting on fire looks a little out of touch given the competition these folks are bringing. arxiv.org/ab [bsky, 1 points, 1 comments]
- DeepSeek Technical Report in pdf format. arxiv.org/pdf/2412.194... [bsky, 1 points, 0 comments]
- The tech report says 2.788 M h800 hours: arxiv.org/pdf/2412.19437 At $2 an hour that comes out to around $5.5 million. [bsky, 1 points, 1 comments]
- "...we introduce DeepSeek-V3, a large MoE language model with 671B total parameters and 37B activated parameters, trained on 14.8T tokens... achieves performance comparable to leading closed-source mo [bsky, 1 points, 0 comments]
- arxiv.org/abs/2412.19437 [bsky, 0 points, 0 comments]
- DeepSeek-V3 Technical Report https://arxiv.org/abs/2412.19437 (https://news.ycombinator.com/item?id=43490167) [bsky, 0 points, 0 comments]
- DeepSeek-V3 Technical Report https://arxiv.org/abs/2412.19437 (https://news.ycombinator.com/item?id=43490167) [bsky, 0 points, 0 comments]
- arxiv.org/abs/2412.19437 [bsky, 0 points, 0 comments]
- This might explain the - for a lot of the folks - astonishing performance of the current DeepSeek model. Just in case you like to read the technical details of the model: arxiv.org/pdf/2412.19437 [bsky, 0 points, 0 comments]
- DeepSeek-V3 Technical Report Comments https://arxiv.org/abs/2412.19437 Event Attributes [bsky, 0 points, 0 comments]
- DeepSeek-V3 Technical Report arxiv.org/abs/2412.19437 [bsky, 0 points, 0 comments]
- arxiv.org/abs/2412.19437 [bsky, 0 points, 0 comments]
- DeepSeek-V3 Technical Report [bsky, 0 points, 0 comments]
- DeepSeek-V3 Technical Report https://arxiv.org/abs/2412.19437 https://news.ycombinator.com/item?id=43490167 [bsky, 0 points, 0 comments]
- DeepSeek-V3 Technical Report https://arxiv.org/abs/2412.19437 [bsky, 0 points, 0 comments]
- DeepSeek-V3 Technical Report New technical report on advanced AI model demonstrating potential improvements in language model architecture and performance Read here [bsky, 0 points, 0 comments]
- https://arxiv.org/abs/2412.19437 DeepSeek-V3についての技術論文。 671Bのパラメータを持つMoE言語モデルで、各トークンに対して37Bがアクティブになる設計です。 効率的な推論とコスト効率の高いトレーニングを目的としています。 [bsky, 0 points, 0 comments]
- DeepSeek-V3 Technical Report [bsky, 0 points, 0 comments]
- DeepSeek-V3 Technical Report arxiv.org/abs/2412.19437 [bsky, 0 points, 0 comments]
- DeepSeek V3 technical report arxiv.org/abs/2412.19437 [bsky, 0 points, 0 comments]
- ....and this is the December 2024 technical paper from Deepseek on V3 arxiv.org/pdf/2412.19437 [bsky, 0 points, 1 comments]
Related