Latest / Elon Musk Podcast / Musk's $25 Billion Custom AI Chip Factory
Transcript
- 0:00Elon Musk's Tesla and SpaceX are they're spending 20 to $25
- 0:04billion to build a custom semiconductor fabrication plant
- 0:08down in Texas. Yeah.
- 0:09And this is specifically to manufacture to a nanometer
- 0:13artificial intelligence chips. Yeah, the massive scale of that
- 0:17financial commitment requires, you know, serious pause.
- 0:20I mean, the Terafab project aims to produce 100 to 200 billion
- 0:25custom chips annually, and that level of volume goes way beyond
- 0:31just like outfitting a fleet of vehicles, so.
- 0:33We have this fierce battle forming for control over
- 0:35physical hardware right in front of us.
- 0:38On one side you've got established semiconductor giants
- 0:40pushing these massive new computing platforms, and on the
- 0:44other, a vehicle manufacturer attempting to build the
- 0:46factories themselves. Our conversation focuses on the
- 0:49tension between those two approaches, looking strictly at
- 0:52the physical limitations of computing based on the
- 0:54materials. Which really alters the basic
- 0:56mechanics of how artificial intelligence gets physically
- 0:58built and distributed globally. So what happens to the global
- 1:01tech economy when the biggest buyers of computer chips decide
- 1:06to become the manufacturers? Well, the entire computing
- 1:09ecosystem has to reevaluate how hardware is designed.
- 1:12I mean, the demands of the software are just stretching the
- 1:15physical limits of silicon to the breaking point.
- 1:18Right. And when you look at the
- 1:19established semiconductor giants, NVIDIA is moving in a
- 1:22direction that feels entirely focused on brute force.
- 1:26Oh, they absolutely are. Their Vera Rubin platform
- 1:29utilizes 6 different specialized chips that function as a single
- 1:33supercomputer. 6 chips. Yeah.
- 1:35And doing this drops inference token cost by a factor of 10 and
- 1:39it delivers five times the inference compute of the
- 1:42previous Blackwell generation. Wait, back up a second.
- 1:44Clarify the exact difference between inference and training
- 1:47before we move on, just so we're clear.
- 1:48Sure. So training is the education
- 1:50phase. You feed the artificial
- 1:52intelligence billions of pieces of data over several months, so
- 1:55it learns the patterns. Inference is when the artificial
- 1:58intelligence actually takes a test.
- 2:00It takes a prompt from you. The user applies the rules it
- 2:04learned and generates a response.
- 2:06Got it? Historically, training required
- 2:08massive supercomputers, while inference could just run on your
- 2:11phone or, you know, lighter hardware.
- 2:14But the requirements for inference are exploding right
- 2:16now. Because the models themselves
- 2:18are fundamentally changing, we are moving toward HNTK AI, where
- 2:22systems act and reason autonomously.
- 2:25They don't just generate a paragraph of text.
- 2:27No. They write code, test it,
- 2:29realize it failed, rewrite it. Exactly, and communicate with
- 2:32other AI agents to solve multi step problems for you.
- 2:35Let me give an example to make sure we have the mechanics
- 2:37right. Go for it.
- 2:39Instead of just asking a program to write a recipe, you are
- 2:41asking it to check your fridge, order the missing groceries,
- 2:45track the delivery and Preheat your oven.
- 2:48The system has to constantly pull new data, evaluate its own
- 2:52logic, and make physical or digital actions on your behalf.
- 2:55Exactly. And the raw capability required
- 2:58for that level of continuous reasoning comes with immense
- 3:01power requirements. The new Reuben chip uses 2200
- 3:05watts of power per die. That is exactly double the power
- 3:10draw of the previous. Generation and unbelievable.
- 3:12To put 2200 watts in perspective, that is the
- 3:17equivalent of running a heavy duty industrial microwave
- 3:20continuously on a piece of silicone the size of a drink
- 3:22coaster. I have to push back on the
- 3:24practicality of this. It sounds like dropping a
- 3:27massive freight locomotive engine into a standard sedan
- 3:30chassis, right? How can standard data centers
- 3:33possibly cool or provide enough electricity for racks filled
- 3:37with these 2200 Watt tips? If you have hundreds of
- 3:40thousands of these running simultaneously, the math on the
- 3:43power grid just stops working. Well, standard data centers
- 3:46actually cannot handle it. That extreme power density means
- 3:49if you put a standard air cooling fan on that chip, the
- 3:52heat generation is so intense the silicon would literally melt
- 3:55itself into slag. Wow, I'm built.
- 3:57It yeah, the thermal limits of the materials are being pushed
- 4:00to their absolute breaking point.
- 4:01The power draw requires purpose built infrastructure.
- 4:04Which means the consequence of that shift severely limits who
- 4:07can participate. The massive power requirement of
- 4:10the Reuven platform restricts this highly capable tier of
- 4:13computing to only the wealthiest corporations.
- 4:16We are seeing a transition from regular data centers to
- 4:19specialized AI factories. These are GW scale facilities
- 4:23specifically built to churn out AI tokens.
- 4:26So it pushes everyone else out. Exactly.
- 4:29This concentrates the most powerful artificial intelligence
- 4:31capabilities into the hands of a few massive tech conglomerates
- 4:35who can actually afford the electricity and the concrete
- 4:38infrastructure structure. So you have these massive power
- 4:40hungry processors, but processing power requires
- 4:43matching memory speed. What happens when you pair a
- 4:46blazing fast processor with slow memory?
- 4:49The processor ends up sitting completely idle.
- 4:52The industry is currently transitioning to HBM 4 memory
- 4:55stacks to solve this. Companies like Micron are
- 4:57manufacturing units that offer 36 to 48 gigabytes of capacity
- 5:01and two terabytes per second of bandwidth per stack.
- 5:04Let me make sure I understand the mechanics of this
- 5:06bottleneck. Having a supercomputer brain
- 5:09paired with slow memory feels like having a genius locked in a
- 5:14room, but you are forcing them to learn entirely through a tiny
- 5:18plastic straw. That's a good way to picture.
- 5:20It you are sliding tiny pieces of information under the door
- 5:23one at a time, the genius can solve anything instantly, but
- 5:26they spend 99% of their time just waiting for you to slide
- 5:30the next piece of paper. The data cannot get in or out
- 5:33fast enough to matter, wasting all that electricity.
- 5:36That bottleneck is known across the industry as the memory wall.
- 5:40When an agentic AI is reasoning through a complex problem, it
- 5:44has to remember the entire history of its actions.
- 5:47The context window. Yes, the context window.
- 5:50If the memory bandwidth is too slow, the AI cannot retain its
- 5:53train of thought. HBM 4 solves this by changing
- 5:56the physical geometry of the hardware.
- 5:57How so? Instead of laying memory chips
- 5:59flat next to the processor on a green motherboard, data travels
- 6:03across microscopic highways. It stacks memory chips
- 6:06vertically like a skyscraper. And how does the data get up and
- 6:09down the skyscraper? By punching microscopic holes
- 6:12straight through the silicon to physically connect the floors,
- 6:14they are called through silicon vias.
- 6:17This allows vast amounts of data to flow simultaneously up and
- 6:21down the stack, rather than waiting in a single file line.
- 6:24But pushing that much data through microscopic holes has to
- 6:27generate an unbelievable amount of friction and heat.
- 6:30It does, and the memory requirements do not stop there.
- 6:34Future projections extend all the way to HBM 8, predicting
- 6:38staggering 15,000 Watt GPU. 15,000 watts.
- 6:42This extreme bandwidth is strictly required to hold the
- 6:45context window for massive neural networks.
- 6:47Which goes back to the cooling problem.
- 6:49The consequence of this massive heat generation forces the
- 6:52entire industry to abandoned traditional air cooling and
- 6:56adopt direct to chip liquid cooling and full immersion
- 6:59systems completely. This alters physical data center
- 7:02construction from the ground up. You are looking at facilities
- 7:04where the server racks are literally dunked into vats of
- 7:07non conductive engineered fluids, or where chilled liquid
- 7:11is piped directly over the bare silicon dye.
- 7:14So you can no longer retrofit an old warehouse to be a data
- 7:17center. No, you have to design the
- 7:18complex fluid plumbing before you even pour the concrete.
- 7:22So if these massive liquid cooled data centers are hitting
- 7:26their physical limits just to run artificial intelligence, how
- 7:30do you put that same kind of brain into a car driving down
- 7:33the highway? You cannot put a water cooling
- 7:35tower on the roof of a passenger sedan.
- 7:38That is exactly where the battle shrinks down to fit inside
- 7:41vehicles. Nvidia's Drive Thor delivers
- 7:442000 teraflops using a custom 4 nanometer process.
- 7:49OK, this single chip is powerful enough to run the entire vehicle
- 7:52stack, including all the autonomous driving functions and
- 7:55the dashboard. Infotainment brands like
- 7:57Mercedes and BYD are already adopting it.
- 8:00Hold on. Explain what a nanometer process
- 8:03actually means in this context, and what a teraflop does for the
- 8:06driver sitting behind the wheel. Sure.
- 8:08Nanometers measure the size of the microscopic transistors
- 8:11carved into the chip. Smaller the number, the more
- 8:13switches you can pack onto the silicon.
- 8:16Think of it like shrinking the width of the roads in a city so
- 8:18you can build more highways in the same amount of space.
- 8:21Exactly. Electrons travel faster and less
- 8:24power is wasted as heat. 4 nanometer process is incredibly
- 8:28dense and a teraflop is a trillion calculations per
- 8:32second. So 2000 teraflops means the car
- 8:35can calculate 2000 trillion mathematical operations every
- 8:39single second to understand its surroundings, right?
- 8:42And how does Drive Thor actually manage all those calculations
- 8:45without melting? It achieves this by taking
- 8:47server grade ARM cores, the basic thinking units or managers
- 8:51of the chip, and combining them with an integrated transformer
- 8:54engine that is the exact same hardware logic used in large
- 8:57language models packed right into the car.
- 9:00Meanwhile, you look at Tesla's current hardware AI 4, and it
- 9:03relies on older 7 nanometer Samsung nodes, legacy ARM cores
- 9:08and a memory bandwidth of just 384 gigabytes per second.
- 9:11The contrast in raw specifications is massive when
- 9:14you put them side by side. Which makes me think Tesla is
- 9:17bringing an outdated smartphone processor to a supercomputer
- 9:21fight. I mean they are at a severe
- 9:23hardware disadvantage trying to run massive vision only neural
- 9:28networks on silicon that is generations behind what NVIDIA
- 9:31is offering to the rest of the automotive industry.
- 9:33Well that assumes everyone is trying to run the exact same
- 9:36software Fairpoint. Tesla utilizes strict software
- 9:40hardware Co design because they control the software entirely.
- 9:43They strip out generic tasks. Tesla's upcoming AI 5 chip is
- 9:48projected to operate on just 150 watts, while rivaling the
- 9:51performance of a 700 Watt server GPU.
- 9:54Wait, hold on, how does a 150 Watt chip match the performance
- 9:57of a massive server GPU? You cannot just magically ignore
- 10:01the physics of computing power. It is the difference between a
- 10:04generalist and a specialist drive.
- 10:06Thor has to be compatible with whatever operating system
- 10:09Mercedes, BYD or any other manufacturer wants to use.
- 10:13It carries transistors specifically dedicated to
- 10:15rendering high resolution graphics for the dashboard,
- 10:18running third party voice assistance, and handling
- 10:20generalized safety redundancies. And AI5 doesn't do that.
- 10:23No, AI5IS stripped of all that generalized compatibility.
- 10:27Every single circuit is physically carved into the
- 10:30silicon specifically to process Tesla's proprietary neural
- 10:34network matrix math. Matrix math being the specific
- 10:37type of grid based calculation that neural networks use to
- 10:41understand images from the car's cameras.
- 10:43Yes. Imagine a spreadsheet with a
- 10:45million cells representing pixels from a camera.
- 10:48Neural networks multiply every cell by every other cell to find
- 10:51edges, shapes and objects like pedestrians or stop signs.
- 10:54In general, processors do this clumsily, moving data back and.
- 10:57Forth exactly, AI 5 is physically wired to do only the
- 11:00specific math problem perfectly. That physical specialization
- 11:04saves hundreds of watts of power.
- 11:06The consequence of that extreme vertical integration naturally
- 11:09limits Tesla's speed, though the delays in the AI5 chip force
- 11:13their upcoming cyber cab to launch on older hardware.
- 11:16Does. But it opens up the extreme
- 11:18energy efficiency required for edge computing in cars and
- 11:21battery powered humanoid robots. And edge computing just means
- 11:25the processing happens locally on the physical device itself,
- 11:29rather than sending data back and forth to a cloud server.
- 11:32Right. If you are building millions of
- 11:33Optimus humanoid robots, you cannot strap a 700 Watt liquid
- 11:38cooled server to their backs. The battery would drain in
- 11:41minutes. Exactly.
- 11:42The robot needs to process visual data and balance itself
- 11:45instantly, which requires A specialized low power inference
- 11:49engine right there in its chest. Tesla's response to future
- 11:52supply constraints regarding those specialized engines is the
- 11:55Terrafab project. The facility in Texas is a joint
- 11:58venture with SpaceX. The addition of SpaceX
- 12:01fundamentally shifts the compute allocation. 80% of the compute
- 12:04output is actually earmarked for space applications like
- 12:08satellite logic and rocket navigation.
- 12:10Leaving only 20% for vehicles and robots, right?
- 12:13I'm highly skeptical of this execution.
- 12:15Mastering atomic level physics and securing extreme ultraviolet
- 12:19lithography equipment takes established foundries decades of
- 12:22painful iteration. Let me explain extreme
- 12:25ultraviolet lithography for a second, because it illustrates
- 12:27your point perfectly. To carve A2 nanometer transistor
- 12:32you have to blast droplets of molten tin with a high power
- 12:36laser 50,000 times a second. Wow.
- 12:38And that generates light with a wavelength so tiny it can etch
- 12:42patterns almost at the scale of single atoms.
- 12:46The mirrors required to direct that light are some of the
- 12:48flattest surfaces ever created in the universe.
- 12:51That's insane. If you scaled one of those
- 12:53mirrors to the size of a country, the tallest mountain
- 12:56would be less than a millimeter high.
- 12:58Exactly. You were talking about a car
- 13:00company attempting to run A2 nanometer fab.
- 13:03It is highly improbable they can spin up a pristine clean room
- 13:07environment and start churning out perfectly etched silicon
- 13:10wafers without massive yield issues.
- 13:13They would just be printing expensive pieces of sand.
- 13:15That skepticism is entirely warranted if they were doing it
- 13:18completely alone, but Terrafab relies on a hybrid model.
- 13:21How does that work? Tesla is supplying massive
- 13:23capital and extreme volume demand while partnering with
- 13:26established foundries to utilize their existing process
- 13:29technology. So they're not inventing the
- 13:31manufacturing process from scratch.
- 13:33No, not. They are providing the billions
- 13:35of dollars in the guaranteed demand to accelerate a partner's
- 13:39existing technology road map on a dedicated line.
- 13:42Exactly when you are projecting a need for hundreds of billions
- 13:46of chips for autonomous vehicles, robots and satellites,
- 13:49you simply cannot wait in line behind Apple and NVIDIA for
- 13:52manufacturing capacity. You have to fund your own line.
- 13:56The consequence here alters the core business model of vehicle
- 14:00manufacturing. By funding their own silicon
- 14:02production, they insulate themselves from global supply
- 14:05chain shocks and create a massive dedicated compute
- 14:09foundry for the space economy. Space applications require
- 14:12highly specific radiation hardened hardware.
- 14:15When a computer is operating outside the Earth's atmosphere,
- 14:18stray cosmic rays can actually strike the silicon and flip a
- 14:22bit from A0 to A1. Just a single ray.
- 14:25Yes, and that single microscopic event can cause a rocket
- 14:28navigation system to fail instantly.
- 14:31Specialized physical shielding and redundant logic gates carved
- 14:35right into the chip. When you combine the compute
- 14:38needs of Starlink satellites communicating in orbit, Starship
- 14:41calculating autonomous atmospheric reentry, and Tesla
- 14:45vehicles navigating city streets, the sheer volume
- 14:49justifies building a dedicated fabrication plant.
- 14:52They are treating silicon the same way they treated battery
- 14:55cells, bringing it in house to guarantee supply.
- 14:58So mastering artificial intelligence requires total
- 15:01control over the physical silicon through massive capital
- 15:05investment and intense energy infrastructure.
- 15:07It leaves you wondering what happens to the traditional tech
- 15:10sector when physical industry takes over the silicon.
- 15:13If a vehicle and rocket manufacturer suddenly controls a
- 15:16massive share of the world's most advanced chip production,
- 15:19are they even a car company anymore, or are we just riding
- 15:22inside computers? If you're not subscribed yet,
- 15:24take a second and hit follow on whatever app you're using.
- 15:27It helps us keep making this. We appreciate you being here.