Latest / Elon Musk Podcast / Elon Musk fights California over First Amendment
Transcript
- 0:00Artificial intelligence developers are running out of
- 0:02high quality human data to train their systems, and it's forcing
- 0:06them to train new AI on text generated by older AI, which is
- 0:11a process that eventually causes the models to completely
- 0:13collapse and start spitting out gibberish about jackrabbits when
- 0:17you ask them to describe historical English architecture.
- 0:19Right the the jackrabbit hallucination it.
- 0:22It's just such a bizarre, striking visual to start with
- 0:25it. Really is because.
- 0:26You have these systems, right? They're engineered to be the
- 0:29most sophisticated heated synthesis of human knowledge
- 0:31ever assembled. They ingest entire libraries,
- 0:35thousands of years of recorded human thought, and they process
- 0:38it beautifully. But the moment they run out of
- 0:41our writing and they start consuming their own recycled
- 0:44data. So when an AI trains on the
- 0:45output of another AI, the internal logic just completely
- 0:49deteriorates. I mean, it's a concept engineers
- 0:51call synthetic data inbreeding, which sounds gross, but it's
- 0:54accurate. Yeah, over successive
- 0:56generations of training, that mathematical anchor to human
- 0:59reality just begins to drift by the 3rd or 4th generation of it
- 1:03basically eating its own tail. A model asked to discuss the
- 1:07structural elements of a 14th century cathedral will suddenly
- 1:10start generating paragraphs about the breeding habits of
- 1:13Jack rabbits. The whole system just eats
- 1:15itself and collapses into absolute nonsense.
- 1:17Exactly. And that bizarre reality sets up
- 1:20the core tension we're looking at today.
- 1:23You have massive technology companies actively hiding the
- 1:26exact ingredients of their models to protect these
- 1:30incredibly lucrative empires. Because high quality human data
- 1:34is suddenly the most scarce, valuable commodity on Earth.
- 1:37Right. And meanwhile, new laws are
- 1:39attempting to RIP the lid off that exact same black box to
- 1:42expose exactly whose data is being consumed.
- 1:45Which brings us to the real issue here.
- 1:47Can lawmakers actually force technology companies to expose
- 1:50the secret ingredients of their artificial intelligence without
- 1:53destroying the value of the models themselves?
- 1:56I mean, the sheer scale of the legal requirements now hitting
- 1:59these companies is staggering. New regulations require any
- 2:03developer releasing A generative AI system to the public to post
- 2:07a high level summary of their training data sets right on
- 2:10their website. Which seems huge.
- 2:12It completely shatters the old standard of simply releasing a
- 2:15product and telling you not to worry about how it got so smart.
- 2:18Right, but a high level summary sounds simple enough on the
- 2:21surface. Just give us a basic overview.
- 2:23But reading through the actual documentation required by these
- 2:27mandates, it is incredibly demanding.
- 2:29Yeah, the wording is definitely deceptive.
- 2:32The regulations dictate that companies must disclose the
- 2:35origins of the data and the specific sources or owners of
- 2:38those data sets. Which was a massive undertaking.
- 2:40And they have to provide the exact number of data points
- 2:42included in the sets. That alone is an astonishing
- 2:45requirement when you consider the architecture of these
- 2:47systems. We are not talking about
- 2:49counting books on a shelf here. We are talking about models that
- 2:52process billions, sometimes trillions, of individual tokens
- 2:56of information. They also have to disclose
- 2:59whether the data is copyrighted or in the public domain, whether
- 3:02it was purchased or licensed. Back up.
- 3:05Yeah. What exactly qualifies as a high
- 3:07level summary? Because listing out trillions of
- 3:10individual data points, categorizing them by copyright
- 3:13status and tracking down the original owners sounds like the
- 3:16exact opposite of a summary. Yeah.
- 3:17And just to ground that for you listening, a token isn't a whole
- 3:20word, right? It's more like a syllable or a
- 3:22fragment of a word. So trillions of tokens is an
- 3:26unimaginably granular breakdown of human language.
- 3:29Exactly. A token is essentially the
- 3:31smallest unit of meaning the machine can process.
- 3:34Sometimes it's a word, sometimes it's a prefix like on or pre,
- 3:38sometimes it is just a single character.
- 3:40Imagine being asked to provide a summary of a library, but the
- 3:43law requires you to classify every single syllable in every
- 3:47single book. And note where that syllable was
- 3:49printed. Right.
- 3:50And who holds the rights to the paper it is printed on and
- 3:53whether the ink used was synthetic?
- 3:55This mandate fundamentally limits the historic culture of
- 3:58technology development. How so?
- 4:01Well, for years the standard operating procedure was to write
- 4:04an automated script, scrape the entire public Internet as
- 4:08quickly as possible, feed it into a massive server farm, and
- 4:12just see what the neural network could learn.
- 4:15Nobody was meticulously cataloguing the input.
- 4:17They were just vacuuming it all up.
- 4:18Exactly. Now developers are forced to
- 4:21conduct exhaustive microscopic audits of all data collection
- 4:25practices retroactively. It changes how they categorize
- 4:29and store information before a single line of code is ever
- 4:32executed to build the model. If you have to publicly declare
- 4:35every single category of data, you can no longer afford to be
- 4:39messy with your ingestion pipelines.
- 4:41And if we look at the potential fallout of doing that
- 4:43accurately, combining these individual disclosures could
- 4:46theoretically allow competitors to reverse engineer the exact
- 4:49data recipes used by the industry leaders.
- 4:51Oh. Absolutely.
- 4:52Because if I know exactly what percentage of your training data
- 4:55came from specific types of copyrighted text, exactly how
- 4:59much synthetic data you mixed in to smooth out the edges, and
- 5:02exactly what proportion of conversational data you
- 5:05prioritize, I have a massive head start.
- 5:08You get the map without having to explore the wilderness
- 5:11yourself. Right.
- 5:12I can replicate your success without spending the billions of
- 5:15dollars you spent on trial and error.
- 5:17Which brings us to how the major players are actually responding
- 5:20to these regulatory mandates. Major companies, including
- 5:23Anthropic and Open AI are complying with the disclosure
- 5:26laws by using incredibly vague language.
- 5:29Of course they are. They are checking the required
- 5:31boxes by citing broad nebulous categories like publicly
- 5:35available information or non public data from third parties.
- 5:39They are actively avoiding naming a single specific data
- 5:42set or revealing the exact ratios of their mixtures.
- 5:46They're essentially saying yes, we use data and yes, it came
- 5:49from the Internet. It feels like compliance through
- 5:51malicious compliance. They are giving the regulators
- 5:55words without giving them any actual meaning.
- 5:58And then you have Elon Musk's XAI, which took a much more
- 6:02aggressive confrontational route.
- 6:04They attempted to secure a preliminary injunction to block
- 6:07these transparency laws entirely.
- 6:09Yeah, their argument in court was that the law forces them to
- 6:12surrender proprietary secrets, though a federal judge outright
- 6:15denied the request. Right.
- 6:17But from a strictly business and legal perspective, the logic
- 6:21behind attempting to block the law does make sense.
- 6:24Forcing companies to reveal their exact weights, parameters
- 6:27and training mixtures is essentially equivalent to an
- 6:29unconstitutional taking of private property and trade
- 6:32secrets. Do you think so?
- 6:34Well, think about it. The specific blend of data, the
- 6:36exact filtering mechanisms used to clean that data, the precise
- 6:40ratio of synthetic to human text.
- 6:42That is the intellectual property.
- 6:44That is the secret sauce that makes one company's model
- 6:47hallucinate less than another's. Forcing them to publish that
- 6:50recipe hands their hard earned competitive advantage over to
- 6:53anyone with a high speed Internet connection and a
- 6:55cluster of graphics processors. OK, I have to push back there.
- 6:58Go ahead. You cannot call it an
- 7:00unconstitutional taking of property when the entire system
- 7:04is quite literally built on the uncompensated property of
- 7:07millions of creators. You can't claim your secret
- 7:10recipe is a protected trade secret when the ingredients were
- 7:14scraped from personal blogs, news sites and digital
- 7:16portfolios without permission. That is the counter argument,
- 7:19yes. Corporate secrecy cannot
- 7:21function as an impenetrable shield, especially when state
- 7:25attorneys general are actively investigating the outputs of
- 7:27chat bots like Grok. When these systems are
- 7:30generating content that impacts elections, offering questionable
- 7:34medical advice, or mimicking the voices and likenesses of real
- 7:37people, that companies cannot just stand behind closed doors
- 7:40and say, trust us, our recipes a trade secret, right?
- 7:44The public and the lawmakers representing them have a right
- 7:47to know what is feeding the machine that is suddenly making
- 7:50decisions for them. And the companies hear that
- 7:53argument loud and clear. And that legal friction changes
- 7:56their entire compliance strategy.
- 7:58What are they doing instead? It opens up a totally new
- 8:01corporate approach, which is the of trust centers.
- 8:05Instead of listing raw data points and handing over the
- 8:07exact ingredients to appease the lawmakers, companies are
- 8:11publishing extensive, beautifully designed
- 8:13documentation focused entirely on their privacy preserving
- 8:16filters and their safety techniques.
- 8:19Redirection. Completely.
- 8:20They are building highly polished web portals dedicated
- 8:23to explaining how much they care about security, how many
- 8:27automated red teams they deploy, and how rigorously they test for
- 8:30bias. So it looks like transparency.
- 8:32It builds consumer confidence. It satisfies the surface level
- 8:36requirements of showing that they have governance structures
- 8:39in place while keeping the actual valuable raw data
- 8:42completely hidden from view. So they tell you exactly how
- 8:45strong the locks on the doors are.
- 8:47They show you the security cameras, they introduce you to
- 8:49the guards, but they never actually show you what is inside
- 8:51the vault. That is a perfect way to
- 8:53describe it. Let's talk about what happens
- 8:55when your personal diary entry or your private e-mail
- 8:59accidentally gets locked inside that vault with no way to get it
- 9:03out, because this leads directly to a massive collision with
- 9:07consumer privacy. Consumer privacy laws are being
- 9:10officially extended to cover artificial intelligence systems.
- 9:13Yes, that means personal data cannot be mined or regurgitated
- 9:17without accountability. And the massive challenge here
- 9:20is the very nature of the automated collection process.
- 9:24Personal information things like private e-mail addresses, cell
- 9:27phone numbers, or even highly sensitive health records
- 9:30discussed on obscure support forms a decade ago is frequently
- 9:34swept up as a purely incidental byproduct of massive Web
- 9:37crawling. They aren't looking for it
- 9:39specifically. Right.
- 9:40The developers building these systems are not actively
- 9:42searching for your specific personal e-mail address to teach
- 9:45their machine how to write, but when they deploy A crawler to
- 9:48vacuum up an entire domain to learn the structure of human
- 9:51conversation, your data comes along for the ride.
- 9:53Hold on, if the data is swept up accidentally, how do they remove
- 9:58it once the AI has already learned it?
- 10:00Because from everything in these technical briefs and AI does not
- 10:04store data like a standard computer or hard drive where you
- 10:07can just find a file, highlight it and hit delete.
- 10:10It does not work like a database at all.
- 10:12And that engineering reality changes the entire architecture
- 10:16of AI development because. You can't just run a search
- 10:18query. Exactly.
- 10:19In a normal database, if you request your data be deleted
- 10:22under privacy laws, a company runs a query, locates your row,
- 10:26and erases it. A neural network stores
- 10:29information in weights and connections across billions of
- 10:32parameters. If a user exercises their legal
- 10:35right to request their personal information be scrubbed,
- 10:37developers face the nearly impossible task of altering the
- 10:40foundational architecture of an already trained model.
- 10:43So how do you remove something that's baked in?
- 10:45Think of it like trying to remove a single drop of red dye
- 10:48from a fully baked cake. The dye isn't sitting in one
- 10:51spot, it has structurally altered the entire composition
- 10:55of the cake. This reality severely limits the
- 10:58indiscriminate vacuuming of the Internet that fueled the first
- 11:01wave of these tools. They have to be careful from the
- 11:03start now. And it opens up the absolute
- 11:05necessity for post training techniques like implementing
- 11:08specific guardrails, automated filters and a concept called
- 11:12machine unlearning to minimize personal information from ever
- 11:16appearing in models outputs, even if the model technically
- 11:20retains the mathematical memory of that data deep within its
- 11:23neural pathways. That is a terrifying thought.
- 11:26If your data gets vacuumed up, it is permanently baked into the
- 11:29model's brain. There is no actual delete
- 11:32button, just a muzzle put on the machine so it hopefully doesn't
- 11:34repeat what it knows about you. A muzzle is a good way to put
- 11:37it. And it is not just personal
- 11:39information getting caught in that indiscriminate web crawler.
- 11:43Tech giants deliberately transcribed copyrighted Internet
- 11:47videos and gathered protected text across the Web because
- 11:51negotiating licenses with individual publishers,
- 11:53independent artists and musicians would simply take too
- 11:56much time. It was a race.
- 11:58Internal communications from several companies revealed they
- 12:01knew exactly what they were doing, calculating that the risk
- 12:04of facing future lawsuits was entirely worth the immediate
- 12:08reward of building a smarter model faster than their
- 12:10competitors. And the proposed transparency
- 12:13measures take direct aim at that specific practice.
- 12:16Under new regulatory frameworks, developers would be required to
- 12:19respond to copyright owners with a complete, verified list of
- 12:23materials used. Which goes back to the indexing
- 12:25problem. Yes, and they would achieve this
- 12:27using an approximate content fingerprint to search for
- 12:30matches in their massive data sets.
- 12:32Essentially, a creator provides a digital signature of their
- 12:35work. Think of it like an audio
- 12:37identification app on your phone that can listen to a song in a
- 12:40noisy room and identify the underlying melody even if the
- 12:44instruments sound different. Although it makes sense.
- 12:46You create that same kind of signature for a textbook or a
- 12:49piece of code, and the AI company has to run that
- 12:52mathematical signature against their entire training library to
- 12:55see if there is a match. This proposed tracking mechanism
- 12:58completely changes the balance of power.
- 13:01Theoretically, it gives creators the exact ammunition they need
- 13:05to sue for copyright infringement or demand financial
- 13:08compensation. Yeah, if you can mathematically
- 13:11prove your fingerprint is in their system, you have a solid
- 13:14legal case. Yes, but there's a downside.
- 13:16Right. The massive downside is that it
- 13:18limits the speed of innovation. Verifying millions of individual
- 13:22fingerprints against massive data sets containing trillions
- 13:26of tokens is computationally heavy.
- 13:28It requires an enormous amount of processing power just to run
- 13:32the compliance checks. You are essentially asking a
- 13:34supercomputer to search for millions of needles in a
- 13:37haystack the size of a planet continuously every time a new
- 13:41creator registers a fingerprint. So the companies are being
- 13:44squeezed from both sides. They require infinite pristine
- 13:47data to make the model smarter and avoid the model collapse we
- 13:50talked about earlier. The Jack rabbits.
- 13:52Exactly. But every single piece of data
- 13:55they ingest now carries a legal tripwire that could result in
- 13:59massive financial penalties or, worse, forced architectural
- 14:03changes to the models themselves.
- 14:05If you've got a decent mic and laptop and some free time,
- 14:08Babble Audio is paying people to record speech data and annotate
- 14:11audio for AI training. No minimums, no fixed hours.
- 14:15You work when you want and get paid weekly via PayPal, Venmo,
- 14:18or bank transfer. They pay per recorded or
- 14:20annotated hour, plus bonus challenges for hitting weekly
- 14:23goals. And if you sign up through our
- 14:25link, you get priority processing and a $15 bonus.
- 14:28Links in the show notes. So getting back to it, beyond
- 14:30the massive copyright issues and the privacy nightmares, there is
- 14:33an entirely different category of regulation targeting the
- 14:36absolute biggest players in the room.
- 14:38Developers of massive frontier models, meaning those generating
- 14:41massive annual revenue and utilizing raising immense
- 14:44computing power for training, are now legally required to
- 14:47report critical safety incidents directly to state emergency
- 14:50services. Which is wild to think about for
- 14:52a software company. The internal governance
- 14:54requirements placed on these frontier developers are
- 14:57incredibly strict and completely unprecedented in the software
- 15:00industry. They must provide anonymous
- 15:02reporting hotline specifically for their engineering staff.
- 15:06They have to update those whistleblowers monthly on the
- 15:08status of any internal investigations, and they are
- 15:11required by law to report those findings directly to their board
- 15:15of directors. The legal framework completely
- 15:18bypasses the middle management layers that typically bury these
- 15:21types of concerns to keep product launches on schedule.
- 15:24Let me pause you right there 'cause I want to make sure I am
- 15:27wrapping my head around this. What kind of emergency are we
- 15:29talking about here? When I think of a chat bot, even
- 15:32a highly advanced one, I do not typically think of someone
- 15:35needing to dial emergency services because the software
- 15:38produce a bad output. No, you wouldn't normally, but
- 15:41this is specifically targeted at catastrophic risk to public
- 15:45health or physical safety. The regulators are looking at
- 15:48the trajectory of the technology.
- 15:50We are talking about models becoming capable of generating
- 15:54novel, highly lethal chemical weapons formulas that bypass
- 15:58known security filters. Yeah, we are talking about
- 16:03models tasked with coding that independently find and exploit 0
- 16:07day vulnerabilities in critical infrastructure grids like water
- 16:11treatment plants or power stations, or models given
- 16:14autonomous agent capabilities that engage in actions that
- 16:17could cause severe physical harm or massive financial disruption.
- 16:21So real world consequences, not just digital ones.
- 16:24Exactly. This mandate completely changes
- 16:27the accountability structure for corporate executives.
- 16:29It limits their ability to hide dangerous flaws or terrifying
- 16:33emergent capabilities in their technology to protect their
- 16:36public stock price. But.
- 16:37They have to report it. Right.
- 16:39It opens up an environment where AI employees have protected
- 16:42legally two mandated channels to sound the alarm if a model
- 16:46begins behaving dangerously during the training run, and it
- 16:49forces the company to engage with state and Federal Emergency
- 16:52services if those alarms are determined to be valid.
- 16:55That is a massive shift from the standard software development
- 16:58life cycle. Usually if software has a bug,
- 17:01you patch it. You don't call the authorities.
- 17:03But there is a massive catch to all of this.
- 17:06Force transparency and reporting.
- 17:08Academic researchers warned that simply forcing AI companies to
- 17:12dump massive amounts of data documentation on the public is a
- 17:15deeply flawed strategy. They compare it to the endless
- 17:18privacy policies or complex Federal Home Loan disclosures
- 17:23that consumers click agree to but never actually read.
- 17:26It's the fine print problem. Yes, it reminds me of those 50
- 17:29page terms of service updates we all blindly accept on our
- 17:32phones. Except now it is for a machine
- 17:35that might be writing our medical prescriptions or
- 17:37filtering our job applications. The psychological reality is
- 17:40that overwhelming the public with highly technical data
- 17:43points creates intense information overload.
- 17:46When a company complies with a mandate by publishing A500 page
- 17:49technical document detailing their data set categorization,
- 17:53their parameter counts, and their safety benchmarks, it
- 17:55actually reduces genuine transparency for the average
- 17:58person. Because nobody can read it.
- 17:59You hide the needle in a massive haystack of compliance jargon.
- 18:03If I hand you a textbook on fluid dynamics, I haven't
- 18:05actually helped you understand how your plumbing works.
- 18:07That's a great analogy. And this realization severely
- 18:10limits the effectiveness of raw data dumps as a tool for public
- 18:14accountability. It changes how policy makers
- 18:17must approach regulation entirely.
- 18:19Because direct to consumer doesn't work here.
- 18:21Right. Instead of focusing on direct to
- 18:23consumer disclosures that no one will ever read or understand,
- 18:27they are opening up a framework designed specifically for
- 18:30information intermediaries. These intermediaries are the
- 18:33investigative journalists, the academic researchers, and the
- 18:36independent nonprofit watchdogs. The translators, essentially.
- 18:40Exactly. They are the ones equipped with
- 18:42the technical knowledge to analyze the raw data, run the
- 18:45necessary adversarial tests against the models, and
- 18:48translate those dense, unreadable disclosures into
- 18:51actionable, understandable insights for the general public.
- 18:54And those intermediaries are going to need every advanced
- 18:57tool they can get their hands on, because the technical
- 18:59reality of catching these models breaking the rules is becoming
- 19:03incredibly difficult. Legal audits that attempt to
- 19:06catch AI companies stealing data by searching for exact word for
- 19:10word matches, which are known in computer science as substring
- 19:14searches, are becoming entirely obsolete.
- 19:17So the AI can steal your work, rewrite it just enough to pass a
- 19:21standard plagiarism check, and the law cannot prove it took
- 19:25your specific data. That is the exact problem the
- 19:27legal system is facing right now.
- 19:29Modern AI models do not operate like a search engine.
- 19:32They do not simply memorize and paste text from a database.
- 19:36They synthesize and regurgitate concepts by stitching together
- 19:39mathematical fragments. So they learned the idea, not
- 19:41just the words. The model maps language in a
- 19:44multi dimensional space. Words that have similar meanings
- 19:47are grouped closer together in that space.
- 19:49A model can perfectly replicate a highly specific copyrighted
- 19:53concept or a unique fictional narrative without ever using the
- 19:56exact original sequences of words.
- 19:59Because it understands the meaning.
- 20:00Right. The neural network understands
- 20:02the semantic meaning of the text, so it can express the
- 20:04exact same idea using an entirely different vocabulary
- 20:08and sentence structure. If your legal audit relies on
- 20:11finding identical strings of words looking for a copied
- 20:14paragraph, you will find absolutely nothing.
- 20:16The audit will come back clean even if the model was entirely
- 20:20trained on your specific protected book or article.
- 20:23So essentially trying to catch these models stealing data using
- 20:27old software tools cause like trying to nail Jelly to a wall,
- 20:31the law is trying to catch a ghost using a net designed for
- 20:35physical objects. It just phases right through.
- 20:37This technological reality severely limits the enforcement
- 20:40of both data transparency and copyright laws.
- 20:43It changes the technical requirements for future audits,
- 20:46opening up a massive demand for new advanced solutions like
- 20:49probabilistic testing. Which is fascinating.
- 20:52Yeah, that involves asking the model a highly specific question
- 20:54millions of times to statistically prove it knows
- 20:57something it shouldn't. It also opens up the demand for
- 21:00persistent data watermarks. Right, the invisible patterns.
- 21:03Invisible statistical patterns embedded into the original
- 21:06content itself that alter the Model S weights during training,
- 21:10creating a hidden signature that cannot be easily scrubbed or
- 21:13altered by the AI company. Which leaves us looking at an
- 21:16incredibly complex regulatory environment.
- 21:19The speed of the technology, the sheer scale of the mathematics
- 21:22involved, is actively outpacing the structural capacity of the
- 21:26legal system to measure it, let alone control it.
- 21:29Regulators are trying to force open the black box of artificial
- 21:32intelligence with mandatory disclosures and whistleblower
- 21:35protections. But the sheer complexity of how
- 21:38these models learn, and how they obscure their own origins is
- 21:42making true accountability a constantly moving target.
- 21:45What happens when the underlying technology becomes so advanced
- 21:49and the synthetic data loops become so dense that even the
- 21:52developers themselves can no longer trace the true origin of
- 21:55the machine's outputs? If you're not subscribed yet,
- 21:58take a second and hit follow on whatever app you're using.
- 22:00It helps us keep making this. We appreciate you being here.