DeepSeek came and it provided this platform for way more Chinese companies that are releasing these fantastic models to kind of have this new
- DeepSeek, Kimi, MiniMax, Z.ai, Moonshot. We're just going Chinese.
DeepSeek is private company, by the way.
DeepSeek is doing fantastic work for disseminating understanding of AI.
deepseek. That's an ominous warning because, you know, let's say you decide, look, it's
company DeepSeek released DeepSeek R1, that I think it's fair to say surprised everyone with near-state-of-the-art performance, with allegedly much less compute for much cheaper. And from then
From DeepSeek, OpenAI, Google xAI, Meta, Anthropic, to Nvidia and DSMC, and to US, China, Taiwan relations,
The DeepSeek-R1 model has a very permissive license.
With DeepSeek's MLA, with this new attention architecture, they need to do some clever things, because they're not set up the same
At DeepSeek, because they have certain limitations around the GPUs that they have access to, the interconnects are limited to some extent
But DeepSeek certainly did it publicly and they may have done it even better because they were gimp on a certain aspect
But DeepSeek's implementation is so complex.
And DeepSeek does the same thing and some of 'em are shared or a lot, we have to take them on face value
You can run DeepSeek on Perplexity. Sitting here, we're like, "We use OpenAI GPT-5 Pro consistently." We're all willing to pay for the marginal
this was based on DeepSeek-V3, which came out the year before in December 2024. There are multiple things on the architecture side. What is fascinating is... I mean, that's what I do with my
And some of DeepSeek's earlier papers, they talk about their training data being distilled for math, I shouldn't use this word yet,
But in DeepSeek's case, which is part of why this was so popular even outside the AI community, is that you can see how the language model
The common one that DeepSeek used is rotary positional impendings, which is called RoPE.
But what DeepSeek did that maybe only the leading labs have only just started recently doing is have such a high sparsity factor.
China’s DeepSeek is able to compete with models like ChatGPT, despite having a fraction of the compute.
I would say you mentioned the DeepSeek moment, and I think DeepSeek is winning the hearts of the people who work on open-weight models because they share
- So some of these models like DeepSeek have the love of the people because they are open-weight. How long do you think the Chinese companies keep
the other ones—it's not that DeepSeek got worse, it's just that the other ones are using the ideas from DeepSeek. For example, you mentioned Kimi—same
And these types of people are at DeepSeek and leading American Frontier Labs, but there are not many places.
What's the difference between DeepSeek-V3 and DeepSeek-R1?
And what to know about these new DeepSeek models is that they do this internet large-scale pre-training once to get what is called DeepSeek-V3 base.
Where this changes is with the DeepSeek-R1, what is called these reasoning models, is when you see tokens coming from these models to start,
This architecture for what is called DeepSeekMoE.
And one of the innovations in DeepSeek's architecture is that they change the routing mechanism in mixture of expert models.
brightly. The new DeepSeek models are still very strong, but that's kind of a... it could look back as a big narrative point where in 2025
- Yes. You mentioned DeepSeek losing its crown. I do think to some extent, yes, but we also have to consider though, they are still, I would say, slightly ahead. And
model development, because DeepSeek famously is built by a hedge fund, Highflyer Capital, and we don't know exactly what they use the
- So I would say the year's almost bookended by both DeepSeek V3 and R1. And then on the other hand, in December, DeepSeek-V3.2.
So, I used the DeepSeek moment that shook the AI world a bit as an opportunity to sit down with them and lay it all out.
But the quote, DeepSeek moment is indeed real.
A lot of people are curious to understand China's DeepSeek AI models, so let's lay it out.
- Yeah, so DeepSeek-V3 is a new mixture of experts, transformer language model from DeepSeek who is based in China.
And then, between the DeepSeek custom license and the Llama license, we could get into this whole rabbit hole.
And that's one of the things that DeepSeek did well is they publish a lot of the details.
- Yeah, especially in the DeepSeek-V3, which is their pre-training paper, they were very clear that they are doing interventions on the technical stack that go at many different levels.
And then, what DeepSeek did is they've done two different post-training regimes to make the models have specific desirable behaviors.
And this is what they did to create the DeepSeek-V3 model.
in order to create the model that is called DeepSeek-R1.
I almost forgot to talk about the difference between DeepSeek-V3 and R1 on the user experience side.
And so, DeepSeek's model is 600 something billion parameters.
our AI platforms because they have something called deepseek. I just want to put a fact out there because we're
Of course, DeepSeek, like the Chinese models, is the exception, especially when it comes to sensitive Chinese topics.
giant MoE model, very similar to DeepSeek architecture in December. And then a startup, RCAI, and both Nemotron and NVIDIA have teased MoE
Nathan, can you describe what DeepSeek-V3 and DeepSeek-R1 are, how they work, how they're trained?
So, we'll get into cost numbers for DeepSeek-V3 on mostly GPU hours and how much you could pay to rent those yourselves.