Latest / Elon Musk Podcast / Latest Tesla Robotaxi news
Transcript
- 0:00Welcome to the debate. Today we're looking at a
- 0:03controversy that sits right at the bleeding edge of
- 0:07transportation technology. It's a dispute that really
- 0:11divides the engineering world right down the middle, and the
- 0:14outcome is going to determine how and frankly if our vehicles
- 0:19drive themselves in the next decade.
- 0:22We are talking about the war between vision only systems and
- 0:25sensor fusion, specifically regarding the use of Lidar.
- 0:29And this isn't just a theoretical argument anymore, is
- 0:31it? We're looking at recent and,
- 0:33well, quite alarming reports about the deployment of Tesla's
- 0:37robo taxi fleet. The headline data suggests a
- 0:39crash rate that's significantly higher than human drivers.
- 0:43Some reports are saying up to four times higher.
- 0:46It it just raises a very uncomfortable question.
- 0:49Can a system that relies only on cameras ever truly match the
- 0:52reliability of a system that uses active laser sensors?
- 0:56And that's really the core of it.
- 0:58Can a computer see the world well enough with just video
- 1:01feeds to navigate safely? Or does excluding depth sensing
- 1:04hardware like Lidar resent an insurmountable barrier?
- 1:08I'm the advocate. My position is that visual input
- 1:11is, well, theoretically sufficient because it mimics the
- 1:15biological model. It mimics us.
- 1:17I also suspect the current crash data is being heavily
- 1:20misinterpreted because of some serious reporting biases.
- 1:24And I'm the dissenter. My position is that abandoning
- 1:27Lidar is a dangerous cost cutting measure that ignores the
- 1:31kind of redundancy you absolutely need for safety
- 1:34critical systems. We're seeing higher crash rates
- 1:37not because of a reporting bias, but because of a fundamental
- 1:40hardware deficit. When you take away the sensor
- 1:43that tells you exactly how far away an object is, you are
- 1:46introducing a level of risk that just shouldn't be on public
- 1:48roads. OK, so let's get into the
- 1:50machinery of this. We really need to start with the
- 1:53first principles argument, because this is the hill the
- 1:56vision only proponents are willing to die on.
- 1:59There's a perspective shared by a contributor to our source
- 2:02material, Retroviridae 6, that's basically an existence proof.
- 2:07Ah, the humans do it argument. Exactly.
- 2:09It's the biological argument. You and I drove here today.
- 2:12We navigated traffic. We merged.
- 2:14We avoided pedestrians. We did all of that using two
- 2:17passive optical sensors, our eyes, and a biological neural
- 2:20network, our brain. We don't have lidar in our
- 2:23foreheads, we don't emit laser pulses to measure time of
- 2:25flight, we don't have radar. We rely entirely on optical flow
- 2:29and pattern recognition. Sure.
- 2:31So Retrovira day six's point is that if a biological neural net
- 2:35can drive a car using only passive optical sensors, then it
- 2:39is physically possible for a synthetic neural net to do the
- 2:42same thing. The physics allows it.
- 2:45Therefore, the argument that Lidar is required is just false.
- 2:49Lidar might be a shortcut, but it isn't a necessity.
- 2:53The photons entering the camera contain all the information you
- 2:56need to drive. I'm sorry but I just don't buy
- 2:58that. Let me tell you why that is a
- 3:00huge category error. Your conflating otential with
- 3:04execution. Just because humans can drive
- 3:06with their eyes doesn't mean robots should drive without
- 3:08laser precision. But why not if the goal is to
- 3:11replicate human capability? Theology is full of flaws we're
- 3:14trying to engineer out of the system.
- 3:17The promise of autonomy isn't to drive as well as a distracted
- 3:20ape, it's to drive perfectly. And to drive perfectly, you need
- 3:24data that the human eye simply cannot provide.
- 3:28But we're not talking about fatigue.
- 3:29We're talking about the sensory input required to build a model
- 3:32of the world. But you cannot compete with
- 3:34Lidar are using only visual cameras when it comes to what
- 3:37I'd call ground truth. As another observer, Wiggly Worm
- 3:41pointed out in the materials, Lidar provides absolute depth
- 3:44data. So maybe we should break that
- 3:46down a bit for anyone who isn't a robotics engineer.
- 3:49Right. A camera is a passive sensor.
- 3:52It takes in light and it creates a flat 2D image.
- 3:56To figure out how far away a car is, the software has to analyze
- 4:00the size of that car in the image, compare it to what it
- 4:03thinks a car looks like, and then infer the distance it's
- 4:07guessing. It's a very educated guess, but
- 4:09it is a guess. It's inference based on
- 4:11perspective and parallax, yeah. Lidar is active, It shoots out a
- 4:16laser pulse, it hits a car, and it measures exactly how long it
- 4:20takes for the light to bounce back.
- 4:21It's simple physics. Distance equals time multiplied
- 4:24by the speed of light. It doesn't guess, it knows.
- 4:27It says there is an object 12.4 meters away.
- 4:32Wiggly Worm's point is that when you remove that sensor, you're
- 4:35forcing the computer to hallucinate death.
- 4:37And when you look at the stats, Tesla's robo taxi is reportedly
- 4:41crashing at a rate 4 times higher than humans.
- 4:44That isn't just the learning curve, that is a failure of
- 4:47perception. I think that statistic, the four
- 4:49times higher crash rate, is doing a lot of heavy lifting in
- 4:52your argument, and I I want to contextualize it.
- 4:56We need to be very careful about comparing apples to oranges
- 4:59here. A crash is a crash, isn't it?
- 5:02Not necessarily. Another analyst, Eskrove 2 He
- 5:07noted that this specific headline is based on a really
- 5:10small sample size. 5 incidents in Austin in a single month.
- 5:15But if you look at the granularity of those incidents,
- 5:18the data includes really minor events like a tire touching a
- 5:21parking sign or bumping A curb while parking. 5 incidents in a
- 5:25month for a small fleet is still high.
- 5:28But think about human behavior. If I scrape my rim on a curb
- 5:31while I'm parallel parking, or if I tap a plastic Bullard at
- 5:35one mile per hour, do I call the police?
- 5:37Do I call my insurance company? No.
- 5:39It never enters the statistical record.
- 5:41It just vanishes. Right, it's unreported.
- 5:44Exactly, but for a robo taxi, every single sensor reading is
- 5:49logged. Every thump is a reported
- 5:51incident. The system self-reports
- 5:54everything. So you're comparing reported
- 5:57autonomous incidents where every scratch is scrutinized against
- 6:01reported human accidents, which are usually only the the one
- 6:04severe enough to require tow truck.
- 6:07Ask contributors UX Test and Jerkletos pointed out you're
- 6:11comparing A microscope to a telescope and then claiming the
- 6:13microscope sees more dirt I. Understand the reporting bias
- 6:17argument. It's valid to a point, but I
- 6:19also think it's a convenient way to wave away failure.
- 6:22You call it rubbing a curb. I call it a failure of object
- 6:25permanence. That feels like a bit of a
- 6:27stretch for a scratched rim. Is it?
- 6:30Contrast this with Waymo. We have user experiences from
- 6:33San Francisco contributors like Turbo Encapsulator and Luda lol
- 6:37who describe Waymo as flawless in the same complex urban
- 6:41environments where these vision based systems are struggling.
- 6:44Waymo uses Lidar. They have that spinning bucket
- 6:47on the roof. They aren't scraping rims.
- 6:49They aren't bumping signs. They're also driving in a
- 6:52fishbowl. You're driving in San Francisco.
- 6:54That's hardly officiable. It's a Geo fenced pre mapped
- 6:57environment that Waymo knows exactly where every curb is
- 7:01because it has a high definition map stored in its hard drive.
- 7:05It's not seeing the curb, it's remembering it.
- 7:07Tesla's vision approach is trying to do something much much
- 7:10harder. Drive anywhere on any road
- 7:13without a map, just like a human.
- 7:15Of course it's going to be clumsier in the beginning.
- 7:18It's learning general intelligence, not just
- 7:19memorizing a map. But that clumsiness has real
- 7:22world consequences. The reports of these vision only
- 7:25cars hitting stationary objects like parking signs?
- 7:29That indicates A fundamental flaw.
- 7:31If a vision system cannot calculate the distance to a
- 7:33concrete Bullard well enough to avoid hitting it, how can we
- 7:36possibly trust it to calculate the velocity of a child running
- 7:40into the street? Because the neural networks are
- 7:42weighted differently for those tasks, the system is likely
- 7:45hyper cautious around pedestrians, but has a higher
- 7:48tolerance for static objects to facilitate, say, parking.
- 7:52That is an assumption. Lidar solves the static object
- 7:56problem instantly. It doesn't need to infer or wait
- 7:59anything. It hits the Ballard with a laser
- 8:01and it knows it's there. The fact that these vision based
- 8:04cars are hitting stationary objects suggests that the
- 8:07software is hallucinating free space where there is solid
- 8:10matter. That is terrifying.
- 8:13It implies the car literally does not know the physical
- 8:15boundaries of its own environment.
- 8:17I will concede that static object detection is a hurdle
- 8:20right now, but identifying these edge cases is exactly how you
- 8:25train the network. Every time it hits a parking
- 8:28sign at 2 mph, it uploads that failure and the entire fleet
- 8:31learns not to do it again. And this brings us directly to
- 8:35the concept of systemic risk. We have to talk about the
- 8:38multiplier effect. Explain how you view that.
- 8:40This was articulated very well by the source contributor Becker
- 8:43Hollow. The argument is all about error
- 8:46scaling. When a human makes a mistake, it
- 8:48causes 1 accident. Human error is stochastic.
- 8:51It's random. You might get distracted by a
- 8:53text. I might drop my coffee.
- 8:55It's isolated. Sure, individual variants.
- 8:58But when a vision based software has a flaw in its programming,
- 9:01say a specific inability to distinguish a white truck
- 9:04against a bright sky, that error is replicated across every
- 9:07single device on the road. The centralized bug.
- 9:10Exactly. If the software misinterprets a
- 9:12specific shadow or glare, thousands of cars become
- 9:15dangerous simultaneously in the exact same way.
- 9:18You aren't dealing with one bad driver, you're dealing with a
- 9:21fleet of clones all sharing the same blind spot.
- 9:24That is a systemic risk profile that we have never, ever dealt
- 9:27with in automotive history. That's an interesting point,
- 9:30though I would frame it differently.
- 9:32That logic flips both ways. It's actually the strongest
- 9:35argument for autonomous systems. Yes, an error is distributed,
- 9:39but so is the solution. If they catch it in time.
- 9:42Think about it, when a human driver is bad at merging,
- 9:45they're usually bad at merging forever.
- 9:47You can't patch their brain, but if you solve the edge case in
- 9:51software, if you fix that white truck against the sky bug, you
- 9:54instantly fix every car on the road.
- 9:57You can upgrade the safety of the entire fleet overnight with
- 10:00an over the air update. The multiplier effect applies to
- 10:02safety even more than it applies to error.
- 10:05You're leveraging the collective learning of millions of miles.
- 10:08But. Until that fix arrives, the risk
- 10:10is distributed to the public without their consent.
- 10:12The public roads are becoming a beta testing environment.
- 10:15We're seeing a move fast and break things mentality applied
- 10:19to two ton metal projectiles. Beckerhollow's logic holds the
- 10:23error rate is a normal human error multiplied by the number
- 10:26of devices using the program. If the program is flawed, the
- 10:29carnage is scalable. I see why you think that, but
- 10:32let me give you a different perspective on the technical
- 10:34reliability piece. You keep going back to Lidar as
- 10:37this source of truth, but Lidar has its own failure modes.
- 10:41It does, but they are different from cameras.
- 10:44Lidar struggles with heavy rain. The laser pulses scatter off the
- 10:48water droplets. It struggles with fog.
- 10:50It can get confused by interference from other lidar
- 10:53units. It's not magic.
- 10:55And this is where the whole sensor fusion argument gets
- 10:57tricky. Go on.
- 10:58When you have a camera an A lidar, they will often disagree.
- 11:02The camera sees a plastic bag blowing across the road and
- 11:05thinks it's nothing. Lidar sees an object and says
- 11:10obstacle emergency brake. Now the computer has to decide
- 11:13which sensor to trust. This is the sensor fusion
- 11:16conflict. By removing Lidar, Tesla is
- 11:20arguing that you remove the noise.
- 11:22You force the neural net to resolve the visual data just
- 11:25like a human does, without getting confused by conflicting
- 11:28signals. That sounds like a very
- 11:30convenient engineering rationalization for saving
- 11:33money. Cameras are cheap, Lidar is
- 11:35expensive. It's definitely cheaper, but
- 11:38retroverted Z6 mentioned. We went from the horse and buggy
- 11:41to the moon in just a few decades.
- 11:44Assuming that computer vision can't bridge the gap just
- 11:46because it hasn't yet is premature.
- 11:48The issue isn't that the camera is blind, it's that the
- 11:51processing isn't yet sophisticated enough.
- 11:54But processing power is scaling exponentially.
- 11:57Software cannot conjure photons where there are none.
- 12:01That's the physics problem. But it can interpret context.
- 12:04Let's talk about those photons. Cameras are passive.
- 12:08They need light. What happens when you drive
- 12:10directly into the sunset? We've all done it.
- 12:12The visor goes down. You squint.
- 12:14You can barely see. Cameras get blinded by sun
- 12:16glare. They get obscured by mud.
- 12:18We have reports that Tesla is having to employ trailing chase
- 12:22cars with human safety monitors for their autonomous taxis.
- 12:26If the system is so theoretically sound, why does it
- 12:29need a human babysitter in a separate vehicle?
- 12:31Every developmental technology has safety protocols during
- 12:35testing, but. This is being sold as a future
- 12:37that is just around the corner. The reliance on cameras
- 12:41introduces A fragility that Lidar solves.
- 12:44Lidar cuts through sun glare. It works in total darkness.
- 12:48It is a second layer of truth. If the camera sees a shadow and
- 12:52thinks it's a hole in the road, the Lidar says no, the ground is
- 12:55flat. Removing that sensor removes a
- 12:58layer of survival. It's engineering hubris to
- 13:01believe you can derive 100% certainty from a sensor that is
- 13:04susceptible to optical illusions.
- 13:06I'm not convinced by that line of reasoning because it assumes
- 13:09we can't solve optical illusions with better AI.
- 13:12But let's pivot to the consequences of this, because
- 13:15the legal aspect is fascinating. It's a nightmare.
- 13:18This leads us to the inevitable question of accountability.
- 13:21When these systems do fail, whether it's a clumsy bump or a
- 13:25serious collision, who is responsible?
- 13:27This is the question posed by Shifty Mennonite in our source
- 13:30threads. Who is going to be held
- 13:33accountable when these things mow people down?
- 13:35It is a legal quagmire. Is it the driver, which in this
- 13:39case is the software? Is it the manufacturer or is it
- 13:42the limitations of the sensor suite itself?
- 13:44I think we need to distinguish between a software bug and a
- 13:48design choice. This is crucial.
- 13:51If a car crashes because of a line of bad code, that's one
- 13:54thing. But if a manufacturer knowingly
- 13:57removes a safety sensor like Lidar, a sensor that is industry
- 14:00standard for competitors like Waymo, and that removal leads to
- 14:04a crash because the camera couldn't estimate depth, that
- 14:07feels very distinct from a mere coding error.
- 14:09You're suggesting negligence. I'm saying it borders on it.
- 14:13There's this sentiment expressed by User Beneficial Soup 3699
- 14:17regarding blatant fraud. While that is, you know, strong
- 14:21language, the core sentiment is valid.
- 14:24If you claim a camera is sufficient and the physics
- 14:26suggest it isn't, and the data shows it crashing, at what point
- 14:30does adherence to a vision only philosophy become liability?
- 14:33But that assumes Lidar was prevented that specific crash.
- 14:36We don't know that. Like I said, LIDAR isn't a magic
- 14:40bullet. It's not magic, it's redundancy.
- 14:43In aviation, we don't fly with one altimeter, we have three.
- 14:47Why on earth should we drive with one type of eye?
- 14:49If the camera fails due to glare or a bug or mud, there's nothing
- 14:53to catch the car. It is a single point of failure
- 14:56system. But there's an economic argument
- 14:58here too. If you require Lidar, you make
- 15:01autonomous cars cost $100,000. They become toys for the rich.
- 15:05If you can solve it with vision, the hardware costs 500 $100.
- 15:09You can put it in every car. A vision based system that's 99%
- 15:13safe and available to everyone might save more total lives than
- 15:17a lighter system that's 99.9% safe but only 1000 people can
- 15:21afford it. That is a utilitarian calculus
- 15:24that works on a spreadsheet, but it doesn't work when you're the
- 15:27one crossing the street. The public Rd. should not be a
- 15:30testing ground for cost cutting measures, disguises innovation.
- 15:35The experiences of users in San Francisco and Austin show a
- 15:37clear divide. Waymo with Lidar is providing A
- 15:40flawless service while vision only systems are struggling with
- 15:44basic static objects. But again, Waymo is on rails.
- 15:47It's a local maximum. It's great for San Francisco,
- 15:50but it doesn't scale to the rest of the world.
- 15:53I'd rather have a safe local maximum than a dangerous global
- 15:56beta test. Until vision systems can match
- 15:58the redundancy and depth accuracy of Lidar, the safety
- 16:01consequences are real and they are statistically proven.
- 16:05We cannot verify the safety of a black box neural net without
- 16:08ground truth sensors. It ultimately comes down to that
- 16:11multiplier effect we talked about.
- 16:13It does. We are at a crossroads.
- 16:16We can take the safe, expensive route with Lidar, which might
- 16:19limit the scalability of the technology but provides that
- 16:22warm blanket of redundancy. Or we can push for the vision
- 16:26solution, which, if it works, multiply safety exponentially
- 16:31across the globe and solves general intelligence.
- 16:34But if it fails, it multiplies error.
- 16:36It multiplies the risk of a single software blind spot into
- 16:40a nationwide. Hazard, and that is the gamble.
- 16:44Is the current risk worth the future reward?
- 16:47I tend to believe that without taking that risk, we stagnate.
- 16:51We'd still be driving horses if we waited for the perfect car.
- 16:55And I would argue that safety is not a place for gambling when
- 16:59you're moving 2 tons of steel at 60 mph.
- 17:02Pretty good isn't good enough. You need absolute truth and
- 17:05cameras just don't provide that. A fundamental disagreement on
- 17:08the philosophy of engineering. Thank you for listening to the
- 17:12debate. We hope this exchange has
- 17:14illuminated the complexities behind the sensors.
- 17:17Drive safe everyone, and watch out for the robots.
- 17:20Goodbye.