What It Means for One Company to Own 100,000 GPUs
xAI's Colossus runs 100,000 NVIDIA H100s. All of South Korea held 2,000. Inside the AI compute gap.
Start with a number. At the end of 2024, every research institute and company in South Korea, added together, held roughly 2,000 NVIDIA H100s. Not two thousand racks. Two thousand cards.
Now hold that against a single company. In Memphis, xAI built Colossus and ran 100,000 H100s as one machine. The gap is not a gap. It is a category difference. The national inventory of a country with the world's fifth-largest chip industry amounted to about one-fiftieth of what one firm was already operating.
An H100 is not the graphics card in a gaming PC. It is a data center accelerator built to train large models, and at cluster scale it is not even plugged into a slot. It is bolted onto a dedicated board in what NVIDIA calls an SXM module, the way an engine is mounted to a chassis rather than dropped into a socket. That mounting is what allows the power delivery and the NVLink interconnect that let many GPUs behave as one very large GPU. Nearly every cluster above a few thousand chips is built this way.
Buying GPUs and running GPUs are different problems, and the second one is much harder.
xAI learned that at cost. While training Grok 3 on a cluster of about 30,000 GPUs, the team found that the existing software simply did not hold at that scale. They ended up rewriting close to half of the core code of the training system from scratch. This is the part of the AI story that does not appear in procurement announcements. Inside a data center, the network drops for an instant, a GPU throttles on heat, and the timing of computations passing between tens of thousands of chips falls out of step. Any one of those stops the entire run, because every GPU has to arrive at the next step together. Getting 100,000 of them to behave as a single computing machine is an achievement of a different kind than acquiring 100,000 of them.
The detail that stayed with me is the size of the team that owned that training system. Seven people. Musk has said publicly that they produced better results with roughly a tenth of the headcount of their competitors.
That number is the reason this belongs in a book about rockets.
Anyone who has followed how SpaceX builds hardware will recognize the pattern immediately. Redesign the component in-house rather than buying it. Put a small team on it. Treat failure as data and iterate faster than the failure can accumulate. Colossus in Memphis was built by the same method, by the same person, under the same principles. And in February 2026, the structural link became formal: SpaceX absorbed xAI, placing rockets, satellite internet, and AI under one vertically integrated roof. The 100,000-GPU cluster is now an asset inside the SpaceX economy. The GPU story is the SpaceX story wearing a different face.
Return to Korea, because the comparison is instructive rather than decorative.
The Korean government committed roughly 1.46 trillion won in a 2025 supplementary budget to secure about 13,000 advanced GPUs, including H200 and B200 units, and began distributing them to universities, research institutes, and industry in February 2026. The plan calls for roughly 15,000 more in 2026, building toward a national AI infrastructure of about 37,000 cards.
That national target does not reach half of what one company was already running before it began scaling further.
And the raw count understates the problem, because of what this column has been about from the first paragraph. Training a competitive frontier model is estimated to require at least 16,000 H100s operating in synchrony. Not 16,000 cards somewhere in the country. Sixteen thousand cards behaving as one brain, exchanging data in real time, on one fabric, under one scheduler, with software that survives the scale. Allocating a few hundred cards at a time to individual labs produces useful research. It does not produce a foundation model.
A country can buy chips. Whether it can wire them into a single machine, and staff that machine with seven people who will rewrite the stack when it breaks, is a different question. The answer to that question is what one company in Memphis has been demonstrating, and it now sits on the balance sheet of a rocket company.
For research inquiries or collaboration, contact: ceo@technorns.com