Introduction
A Chinese AI lab has released what could be the most powerful “open” AI model yet. DeepSeek V3 from DeepSeek is open sourced and can be downloaded and modified for commercial use.
A General AI Model
DeepSeek V3 can do various text based tasks such as coding, translating, writing essays or emails from descriptive prompts. According to DeepSeek’s internal test, it outperforms the open AI models and even surpasses some closed AI systems that can only be accessed through APIs.
Awesome Performance
In Codeforces coding competition, DeepSeek V3 beats Meta’s Llama 3.1 405B, OpenAI’s GPT-4o, and Alibaba’s Qwen 2.5 72B. On Aider Polyglot benchmark, it’s the top one.
Huge Training Dataset and Size
The model is trained on 14.8 trillion tokens dataset, which is about 11.1 trillion words. It has 671 billion parameters, much more than Llama 3.1’s 405 billion parameters. More parameters often means better performance but also demands more powerful hardware to run.
Efficient Training Process
Despite the size, DeepSeek V3 is trained in just 2 months using Nvidia H800 GPUs in a Chinese data center. The training cost is only $5.5 million, a fraction of what OpenAI spend on similar projects. This is even more impressive considering the recent US restrictions on Chinese companies buying advanced GPUs.
Limitations to Address
While DeepSeek V3 is a tech achievement, it’s not perfect. For example, it won’t respond to politically sensitive topics like Tiananmen Square, which means there’s bias in its responses.
A Step Forward for Open AI
It’s a big step for open AI. Its performance, versatility and cost make it a strong competitor to both open and closed AI systems.
Stay tuned as DeepSeek V3 continues to redefine the boundaries of what open AI models can achieve.





