GPUs . But when pre-training on 10,000 or 100,000 GPUs , you hit very different failures. GPUs break in weird ways, and on a 100,000
GPUs , CPUs, and AI hardware in general.
over 1,100 Rubin GPUs , 60 exaflops, 10 petabytes per second of scale bandwidth.
use our GPUs , and OpenAI can still get distribution out of this," which is another very real thing, because it doesn't cost them anything.
You need all these GPUs at once.
Then we moved to GPUs , because that could parallelize the process much better.
People started using their GPUs -- their graphics processors-- to do this.
where you send your prompt to GPUs run by certain companies, and these companies will have different distributions and policies on how your data is stored,
by the restrictions of the GPUs that were shipped into China legally, not the ones that are smuggled, but legally shipped in,
I just talked about is goes beyond CPUs and GPUs and networking chips and scale up switches and scale out switches.
- Yeah, GPUs and— Architecture, algorithms, design, um— - So, you constantly have an eye on the entire stack, and you're having to, like,
And this is a communication between all the GPUs in the network, whether it's in training or inference.
You could train them on one to eight GPUs whereas, you know, now we train jobs on tens of thousands, soon going to hundreds of thousands of GPUs .
First there is the fact that I have four GPUs , graphics processing units, at my disposal.
One is the actual dreaming all happens on the GPUs .
We're finally in the stage where we have GPUs in this thing so that we can do convolutional filters and start picking up stuff at any frequency they do it at.
Everybody heard of Nvidia, with the GPUs .
they're always talking about like, "Our GPUs are hurting." And I think during one of these gpt-oss-120b release sessions, Sam Altman said, "Oh, we're releasing this because we can use your GPUs .
the GPUs are different but you probably train gpt-oss-20b way faster in wall-clock time than GPT-2 was trained at the time.
At training time, your efficiency with your GPUs is dramatically improved by using this architecture if it is well-implemented.
of a certain set of the GPU resources or a certain set of the GPUs , and then the rest of the training network sits idle,
Walmart-sized building somewhere, but those GPUs , those uh computer stacks,
techniques and then Google buys them and throws a boatload of uh gpus at them and invents tpus for them
It's now at 200,000 GPUs and growing very quickly.
And then, you can also look at how many hours of these GPUs that are used.
At DeepSeek, because they have certain limitations around the GPUs that they have access to, the interconnects are limited to some extent
So they don't know what scaling laws are or GPUs or comput or whatever and let's try and keep this as simple as
came from video games and entertainment in some way, whether it's through GPUs or through early multi-user dungeons, et cetera.
That is a hybrid HPC system that puts together CPUs, GPUs , and mix to achieve very high performance.
Once you have that, and again shortly after maybe, well, you can scale these things up on millions of GPUs , right?
We were already selling millions and millions of GeForce GPUs a year, and we said, "You know, we ought to put CUDA on GeForce
There's also analysis showing that the way the Chinese models are served—you could argue this is due to export controls— is that they use fewer GPUs per replica, which makes them slower
They're at the limits of the GPUs .
And these companies could have millions of GPUs .
To do a primer for language model reinforcement learning, what you're doing is having two sets of GPUs .
Whether that was AI, graphics, physics engines, hardware, even GPUs of course were designed for gaming originally.
So, as I understand, that went below CUDA, so they go super low programming of GPUs .
That uses, you know, these days, you know, tens of thousands, sometimes many tens of thousands of GPUs or TPUs
More recently, we've actually now started to see that we can replicate this on CPUs and then potentially even on TPUs and GPUs .
At Google, some of you will know that we've invested in a variety of different computational platforms-- CPUs, GPUs , now TPUs for the tensor processing