Latest / Elon Musk Podcast / AI Tokens worth $250,000
Transcript
- 0:00NVIDIA CEO Jensen Huang stated that a software engineer earning
- 0:03a $500,000 salary must consume at least $250,000 worth of AI
- 0:09tokens annually or he would be deeply alarmed.
- 0:12Yeah. That is a wild ratio to think
- 0:15about. And just to clarify, a token is
- 0:18basically the absolute basic unit that artificial
- 0:20intelligence uses to process and generate text.
- 0:23So if you take a common word like unbelievable, the AI
- 0:27doesn't actually read it as one whole piece.
- 0:29It breaks it into these smaller chunks, unbelieve and able and
- 0:33those individual chunks, those are tokens.
- 0:36What we're seeing in the industry right now is this
- 0:38massive shift. Computing power is moving from
- 0:40just being this, you know, background infrastructure cost
- 0:43into a direct, measurable unit of human productivity.
- 0:47Brings us to our mission for this analysis.
- 0:49Today we are looking through a stack of research papers,
- 0:51industry surveys, and technical interviews to answer this exact
- 0:55question. Does spending hundreds of
- 0:56thousands of dollars on compute actually make an engineer 10
- 1:00times more productive? Or are companies just funding
- 1:03the illusion of speed? That is the core issue really,
- 1:06because right now token budgets are essentially becoming a
- 1:10direct recruiting tool. Wow, really?
- 1:12Oh, yeah, They are sitting right alongside your salary, your
- 1:15annual bonuses, your equity packages.
- 1:17I mean, we are looking at what is basically the fourth pillar
- 1:20of compensation. So people are actually asking
- 1:23about this in interviews. Exactly.
- 1:24Candidates interviewing at Open AI are actively asking about
- 1:27their compute budget before they even consider accepting a job
- 1:31offer. That is insane.
- 1:33Right, and there are engineers burning through 10 billion
- 1:36tokens in a single week. 10 billion.
- 1:39Yeah, there was this one developer at Purcell who
- 1:42actually racked up a $10,000 bill in one single day just
- 1:45deploying agents to build a a new service.
- 1:48I mean, if you're listening to this and wondering how someone
- 1:50even spends $10,000 in a day, they are basically writing
- 1:55scripts that automatically trigger hundreds of language
- 1:59models to work simultaneously. Precisely.
- 2:01And if those models get stuck in a feedback loop, you know,
- 2:04continuously talking to each other, requesting new data, the
- 2:07billing meter just spins completely out of control.
- 2:10It just runs the credit card right up.
- 2:12Right. It's like think about hiring a
- 2:14top tier biotech scientist. You pay that scientist a massive
- 2:18salary, but you do not just put them in an empty room with a
- 2:21notepad. No, of course not.
- 2:22They require a massive laboratory materials budget.
- 2:25You know, centrifuges, chemical regions, sterile environments
- 2:28just to function. Huang himself actually compared
- 2:32engineers who refused to use tokens to chip designers,
- 2:36insisting on using paper and pencil instead of computer aided
- 2:39design tools. Which is a great analogy,
- 2:41honestly. And this completely alters how
- 2:43organizations have to approach their hiring and their corporate
- 2:46budgets. How so?
- 2:47Well, a candidate's inherent value to a company is now
- 2:50fundamentally tied to their ability to efficiently spend the
- 2:53company's compute resources. Oh, I see.
- 2:56You aren't just hiring a coder anymore, you are hiring a
- 2:59manager of digital resources sources.
- 3:01The resource just happens to be tokens, and the budget required
- 3:04to sustain them is just enormous.
- 3:06Right, so if individual compensation is changing that
- 3:08drastically, the structure of the entire workforce must warp
- 3:12to accommodate it. Right, Absolutely, it has to.
- 3:15Because engineers are transitioning away from writing
- 3:18individual lines of code, they're spending their days
- 3:20writing ideas, defining architectures, and kind of
- 3:24orchestrating massive teams of digital workers.
- 3:26Yeah, the orchestration part is key.
- 3:28NVIDIA actually projects a future where there's 75,000
- 3:31human employees will work alongside 7.5 million AI agents.
- 3:37Which is just a staggering number.
- 3:39That is a 100 to one ratio of machines to humans.
- 3:43And that completely redefines what a software engineer
- 3:45actually does during their 9 to 5 schedule.
- 3:48Wait, back up 100 digital workers for every human?
- 3:53How does a single person even manage that volume of output?
- 3:56I mean logistically. Well, they absolutely do not do
- 3:59it manually. Software tools, code, compilers,
- 4:01relational databases, they are all going to see massive spikes
- 4:04in overall usage because these agents are basically the new
- 4:06power users right through. Especially systems like Open
- 4:09Claw, which is essentially this framework that allows an AI to
- 4:12browse the web, click buttons and type in terminal windows
- 4:15just like you would. These agents are operating
- 4:18around the clock. So they're just running 24/7.
- 4:20Exactly. The human is not reviewing every
- 4:23single action. The human is directing the
- 4:25agents with high level goals and then those agents are
- 4:28interacting directly with the enterprise software.
- 4:32You know, querying the databases and running automated tests
- 4:35continuously. Which creates an environment of
- 4:37extreme parallel workflows. Yes, humans are no longer
- 4:41bottlenecked by their own typing speed, which opens up concurrent
- 4:44project development on just a massive scale.
- 4:47Oh, completely. Like if you are an engineer, you
- 4:50can have 50 agents compiling your code while fifty other
- 4:53agents run complex security tests all happening
- 4:55simultaneously. Your output is completely
- 4:58detached from how fast your fingers move across a physical
- 5:01keyboard. But, and this is a big but, with
- 5:03those agents running continuously in the background,
- 5:06the sheer output seems massive. Which leads directly to the
- 5:09question of actual measurable productivity, because generating
- 5:13text quickly is fundamentally different from producing a
- 5:17finished working application that customers can actually use.
- 5:20See I agree, but there is this severe disconnect between how
- 5:25fast developers feel they are working and how fast they are
- 5:28actually shipping finished products.
- 5:30Oh, for sure. If you are a software engineer
- 5:32listening right now, you have probably felt that intense
- 5:36dopamine rush of generating a massive scaffolding of code in a
- 5:40few seconds. It.
- 5:41Feels great. It does.
- 5:43You type a prompt into your window, and an entire functional
- 5:47module just appears on your screen.
- 5:50It feels incredibly efficient. You feel like you're doing the
- 5:52work of 10 people. I completely see why it feels so
- 5:55fast, but I have to push back a bit because developers are
- 5:57hitting what we call the 70% problem A. 70% problem.
- 6:01Yeah, the artificial intelligence handles the easy
- 6:0370%, instantly laying down the basic syntax and structure.
- 6:07But that final 30% requires painstaking human logic, deep
- 6:11systemic understanding, and intense review to actually work
- 6:15in a production environment. So the last mile is the hardest.
- 6:18Exactly because the machine lacks actual intent, it
- 6:22generates mathematical patterns that look correct to the human
- 6:25eye. But the human has to ensure
- 6:27those patterns actually solve the specific, nuanced business
- 6:31problem the company is facing. Well, backing that up, a
- 6:34randomized trial conducted by MBTR actually showed that
- 6:38experienced developers using these tools were 19% slower at
- 6:42completing their tasks, even though they self reported
- 6:45feeling 20% faster. Which is a wild psychological
- 6:48trick being played on the human brain it.
- 6:50Really is. They felt faster, but they were
- 6:53empirically slower. And there's a completely
- 6:55separate DX study tracking enterprise companies that found
- 6:58a 65% increase in AI tool usage. You did only a 10% gain in
- 7:03actual pull request throughput. Only 10%.
- 7:06Yeah, and a pull request is just the formal process where a
- 7:09developer submits their finished code to be merged into the main
- 7:11company project. So two usage skyrockets, but
- 7:15finished work barely moves at all.
- 7:16Furthermore, a Stack Overflow survey revealed that only 16.3%
- 7:22of developers feel these tools make them highly productive.
- 7:25The empirical numbers simply do not support that emotional
- 7:28feeling of speed. So the focus of engineering work
- 7:31is forced to shift from creative problem solving to heavy
- 7:35validation and editing exactly, which heavily limits the
- 7:38perceived return on investment for senior talent.
- 7:41A senior engineer is spending more time untangling the
- 7:44machines overly complex mess than they would have spent just
- 7:47writing the simple, elegant code correctly from scratch.
- 7:51And because engineers are severely bogged down validating
- 7:54this endless stream of machine output, cognitive fatigue sets
- 7:58in. And that creates massive
- 8:00vulnerabilities when that validation process inevitably
- 8:03fails. Right, because you get tired.
- 8:04You are no longer writing software, you are proofreading a
- 8:07machine that never sleeps and never stops typing.
- 8:10And high speed generation introduces severe technical debt
- 8:13and major security risks into corporate systems.
- 8:16And a PIRO study actually found that machine generated code
- 8:19introduced 322% more privilege escalation paths.
- 8:23Which is terrifying. It is, and privilege escalation
- 8:26is basically when a standard user is accidentally given the
- 8:29keys to the Kingdom, allowing them to access administrative
- 8:33controls they should never have. The study also found 153% more
- 8:39fundamental design flaws. And the worst part is these
- 8:42broken commits were merged into the main projects four times
- 8:45faster than human written code. Four times faster, so it
- 8:49completely bypassed normal review scrutiny.
- 8:51Exactly. People trust the output far too
- 8:54much because it just looks so clean on the surface.
- 8:56There was also a 40% increase in exposed secrets.
- 8:59Exposed secrets like passwords. Yeah, because these models need
- 9:03massive amounts of context to write good code.
- 9:05Developers are just copying and pasting their entire working
- 9:08environments into the prompt window.
- 9:10Oh, so they are accidentally including their company's
- 9:13private passwords and live API keys sending highly classified
- 9:17data directly to external cloud servers?
- 9:19Hold on, so the AI operates like a highly confident intern who
- 9:23writes incredibly fast but routinely leaves the front door
- 9:26completely the unlocked? That is a highly accurate way to
- 9:28put it. Yeah.
- 9:30In fact, pull requests containing machine generated
- 9:33code requires 60% more security comments from human reviewers
- 9:38just to catch those basic mistakes.
- 9:41The digital intern produces an incredible volume of text, but
- 9:44the senior staff pays a heavy price in continuous oversight
- 9:47and correction. Unbridled agent deployment is
- 9:50just too dangerous on its own. Then it's forcing companies to
- 9:53adopt strict control planes and policy based guardrails just to
- 9:57monitor what the agents are doing on the network.
- 10:00You have. To You cannot just let 7 million
- 10:02digital workers loose on your enterprise systems without a
- 10:05massive, dedicated security apparatus tracking their every
- 10:08single move. But despite these glaring
- 10:10security flaws and the obvious workflow bottlenecks we just
- 10:13talked about, the sheer demand for compute continues to
- 10:16skyrocket, and it's fundamentally altering global
- 10:19macroeconomics. Yeah, the physical
- 10:20infrastructure required to power this transition is staggering.
- 10:24It is the Jebens paradox basically takes effect here.
- 10:27As we make scaffolding code cheaper and more efficient to
- 10:30roduce, we consume vastly more of it, which requires
- 10:33unprecedented physical infrastructure.
- 10:35Right, like with steam power. Exactly.
- 10:38When steam edges became more efficient in the 19th century,
- 10:41we did not use less coal. We found entirely new ways to
- 10:44use steam power, driving coal consumption through the roof.
- 10:48As tokens get cheaper, we deploy millions more agents to do
- 10:52thousands of new tasks. And while the Bureau of Labor
- 10:55Statistics shows only a 1.3% aggregate productivity
- 10:59improvement across the entire workforce. 1.3%.
- 11:02Just 1.3. Meanwhile, NVIDIA currently has
- 11:05order visibility of $1 trillion for its next generation
- 11:09Blackwell and Vera Rubin architectures just to meet the
- 11:12global demand for tokens I. Have to ask though, why drop $50
- 11:16billion to build a single inference factory?
- 11:19That massive capital expenditure seems entirely out of scale,
- 11:23with a 1.3% productivity gain. It does seem crazy on paper.
- 11:27It sounds like burning money to chase a trend that is barely
- 11:31moving the economic needle. Well, Huang's logic is that
- 11:34highly dense multibillion dollar factories are required to
- 11:37produce at the absolute lowest cost per Watt.
- 11:40Cost per Watt. Right, you need astronomical
- 11:43scale to make the energy consumption viable for the
- 11:45business. In a $50 billion budget, 20
- 11:48billion goes strictly to land acquisition, electricity,
- 11:52substations, and industrial cooling facilities.
- 11:55Wow, those are fixed costs regardless of the specific chips
- 11:59used inside the building. By maximizing the sheer
- 12:02throughput of the entire data center, the actual cost to
- 12:05produce one single token drops drastically.
- 12:08So the entire economic value of the IT sector becomes tethered
- 12:11directly to energy efficiency. Exactly.
- 12:13Tokens per Watt becomes the defining metric of corporate
- 12:16profitability. We have moved away from software
- 12:18algorithms and we are now dealing with a brutal physics
- 12:21problem of Power Distribution and heat management.
- 12:24And because cheap tokens require these massive physical $50
- 12:28billion factories and intense power grids, compute becomes a
- 12:31tangible physical choke point. Which means computing power is
- 12:35the ultimate lever for regulating and governing
- 12:38artificial intelligence globally.
- 12:40Yes, because algorithms and training data are intangible.
- 12:44They are easily shared across the Internet in a matter of
- 12:46seconds. But compute relies on a highly
- 12:49inelastic supply chain heavily dependent on EUV, lithography,
- 12:54and dedicated localized power infrastructure.
- 12:57Right? You can copy a data set
- 12:59instantly. You cannot copy a $50 billion
- 13:02power dense factory. Wait, EUV lithography?
- 13:05Just so everyone listening is on the same page, we are talking
- 13:08about the extreme ultraviolet lasers used to print the
- 13:11microscopic patterns onto the silicon chips, right?
- 13:14Exactly. These are some of the most
- 13:16complex, physical and rare their manufacturing machines on the
- 13:19planet. Compute is detectable, it's
- 13:21excludable, and it's completely quantifiable.
- 13:24You can literally see it, right? Governments can track large data
- 13:27centers from space using satellite imagery.
- 13:29They can subsidized specific research facilities or
- 13:32physically restrict international access to advanced
- 13:34chips. They know exactly where the
- 13:36power is flowing on the National Grid.
- 13:38So this physical reality opens up the potential for entirely
- 13:42new international control mechanisms.
- 13:44Like a Strategic Compute Reserve or even a shared international
- 13:48CERN for artificial intelligence.
- 13:51That's a fascinating concept. Yeah, nations can pool their
- 13:54financial resources to build this capital intensive
- 13:57infrastructure, ensuring they maintain strict control over the
- 14:00very physical substrate that the technology requires to function.
- 14:04We are entering an era where an engineer's value is directly
- 14:07measured by the volume of compute they can safely command.
- 14:11However, managing an autonomous workforce is far messier and
- 14:14slower than it feels, introducing severe security
- 14:17risks that require constant human oversight.
- 14:20If human engineers are no longer hired to do manual coding, how
- 14:23does someone ever get an entry level job to gain the experience
- 14:27needed to manage a fleet of 100 digital workers?
- 14:30If you're not subscribed yet, take a second and hit follow in
- 14:33whatever app you're using. It helps us keep making this.
- 14:35We appreciate you being here.