Latest / Elon Musk Podcast / Tesla AI Day 2022 Optimus Robot reveal
Transcript
- 0:00The Black Friday savings at the Home Depot, you'll find top
- 0:02brand kitchen appliances with Innovative features that can do
- 0:05more to your holidays. Can be more ovens with built-in
- 0:08air, fryers for baking, the perfect cookies, dishwashers
- 0:11with smart Tech to clean everything from bakeware to
- 0:13festive mugs and high-capacity refrigerators to keep leftovers.
- 0:17Fresh chopped Black Friday savings and get up to 30% off,
- 0:21plus instantly save up, to 750 on select GE kitchen packages.
- 0:24At the Home, Depot a Dewar's, get more done offer.
- 0:27Valid m s through November 30th was only store online for
- 0:29details. If the Venture X card from
- 0:32Capital One gives you more of what you love.
- 0:34Like premium travel benefits and access to Taylor Swift tickets.
- 0:38I do love her earn five. Times miles on flights and 10
- 0:42times. Miles on hotels through Capital
- 0:44One travel. Enjoy your stay in sweet 13, 12
- 0:4613. That's Taylor's.
- 0:48Lucky number plus, get access to Taylor Swift.
- 0:50The air is tour presented by Capital One.
- 0:53Maybe I'll see you there. The Venture X card from Capital
- 0:56One, what's in your wallet terms of pliancy capitalone.com for
- 0:59details? The kissing run, Rashad JCPenney
- 1:05for thousands of deals solo. No coupons needed this weekend,
- 1:08save up to 50% on kitchen electrics from Brands like,
- 1:11Keurig and Cuisinart, say yes, please to Diamonds and
- 1:14gemstones. Now, 1999 each and bundle up the
- 1:16famine coats starting at 1499. We got your holiday.
- 1:24A valid on select items. 1118 311 20 excluded from coupons.
- 1:27Exclusions apply see storage AC p-- dot Comfort details.
- 1:34When you need Auto Parts, O Reilly Auto.com is just a few
- 1:37clicks away. We offer convenient options for
- 1:39you to get your parts quickly order online and pick up for
- 1:42free. At your local, O'Reilly Auto
- 1:44Parts store. Will even bring it out curbside
- 1:47or you can have your parts delivered right to your door
- 1:49with free shipping on most orders over. $35, visit O Reilly
- 1:52Auto.com The Robert can actually do a lot more than we just
- 2:04showed you. We just didn't want it to fall
- 2:06on its face, so will will show you some videos.
- 2:10Now of the robot, doing a bunch of other things.
- 2:15Yeah, which are less risky. Yeah, we're so close that screen
- 2:19guys. Yeah.
- 2:28Yeah, we wanted to show a little bit more what we've done over
- 2:30the past few months with a pad and just walking around and
- 2:33dancing on stage, just humble beginnings but you can see the
- 2:40autopilot neural networks running as is just retrained for
- 2:43the bud directly on that. On that new platform that's been
- 2:47watering can. Yeah, when you see a rendered
- 2:49view that's that's the robot. What's the best?
- 2:52The world of robots. He's so it's very clearly
- 2:55identifying objects that like, this is the object it should
- 2:58Pick up picking it up. Yeah.
- 3:05We use the same process as we did for the pie dough to connect
- 3:08data in train your on networks, that we didn't Deploy on your
- 3:11robot. That's an example.
- 3:13That illustrates the upper body, a little bit more, something
- 3:19that we like try to nail down in a few months over the next few
- 3:22months, I would say to Perfection.
- 3:26This is really an actual station in the Fremont Factory as well,
- 3:29but it's working at. Yep.
- 3:31So That's not the only thing we
- 3:43have to show today, right? Yeah, absolutely.
- 3:44So that was what you saw was what we call Bumble.
- 3:49See that's our sort of rough development robot using semi
- 3:54off-the-shelf actuators but we actually have gone a step
- 3:59further than that already. The teams that incredible job
- 4:03and we actually have an optimist bought with fully Tesla designed
- 4:07and built actuators battery pack control.
- 4:11System everything it wasn't quite ready to walk but I think
- 4:16it will work in a few weeks but we wanted to show you the robot.
- 4:21Something that's actually fairly close to.
- 4:24What will go into production and and show you all the things it
- 4:28can do. So let's bring it up.
- 4:32Do it. So hear you sing Optimus with?
- 5:24These are the with the three degrees of freedom that we
- 5:28expect to have in Optimus production unit 1, which is the
- 5:32ability to move. All the fingers independently
- 5:34move, the to have the thumb, have two degrees of freedom.
- 5:39So it has opposable thumbs and both left and right hand.
- 5:42So it's able to operate tools and do useful things.
- 5:46Our goal is to make a useful humanoid robot as quickly as
- 5:50possible. And we've also designed it using
- 5:54the same discipline that we use in designing the car, which is
- 5:58to say to design it for manufacturing, such that as
- 6:02possible to make The robot at in high volume at low cost with
- 6:07high reliability. So that's incredibly important.
- 6:10I mean, you've all seen very impressive, humanoid robot
- 6:14demonstrations, and that's great.
- 6:16But what are they missing? They're missing a brain that
- 6:20they don't have the intelligence to navigate the World by
- 6:24themselves. And they're also very expensive
- 6:27and made in low volume. Whereas this, this is often
- 6:32missed is society. And Extremely capable robot but
- 6:35made in very high volume. Probably ultimately, millions of
- 6:38units and it is expected to cost much less than a car.
- 6:44I just bring. So I would say probably less
- 6:47than 20 thousand dollars would be my guess.
- 6:56The potential for optimistic is I think appreciated by very few
- 7:01people. As usual has the demos and
- 7:08coming in hot. So yeah, the teams put put in
- 7:17and the team has put in an incredible amount of work.
- 7:20It's working days. You know, some days we see
- 7:23running the 3M oil that to get to the demonstration today.
- 7:28Super proud of what they've done is they've really done a great
- 7:31job. I just like, give a hand to the
- 7:32whole option was team. So you know that now there's
- 7:45still a lot of work to be done to refine Optimist and improve
- 7:51it obviously because there's just optimism version 1 and
- 7:55that's really why we're holding this event which is to convince
- 8:00some of the most talented people in the world like your guys to
- 8:04join Tesla and help make it a reality and bring it to fruition
- 8:09at. Such that it can help millions
- 8:13of people and the potential likes it is really boggles the
- 8:18mind because if say, like what is an economy and economy is
- 8:25sort of productive, entities times, the productivity capita
- 8:30times output productivity per capita at the point at which
- 8:33there is not a limitation on capita.
- 8:36The it's not clear what an economy.
- 8:38Even means that point it Me becomes quasi infinite.
- 8:43So What, you know, taken to fruition in the hopefully benign
- 8:49scenario the, this means a future of abundance, a future
- 8:57where there is no poverty where people, you can have whatever
- 9:03you want in terms of products and services.
- 9:09It really is a fundamental transformation of civilization
- 9:14as we know it. Obviously we want to make sure
- 9:18that transformation is a positive one and safe and but
- 9:24that's also why I think Tesla as an entity during this being a
- 9:29single class of stock publicly traded owned by the public is
- 9:34very important and should not be overlooked.
- 9:37I think this is essential because then if the public
- 9:41doesn't like, what Tesla's doing the public can buy shares and
- 9:44Tesla and vote. Differently.
- 9:47This is a big deal like it's very important that I can't just
- 9:52do what I want. You know, sometimes people think
- 9:55that that's not true. So You know, that it's very
- 10:02point that the corporate entity that has that makes this happen
- 10:07is something that the public can properly influence and so I
- 10:13think the Tesla structure is is ideal for that.
- 10:20Like I said that self-driving cars will certainly have a
- 10:25tremendous impact on the world. I think they will improve the
- 10:30productivity of Transport by at least a half order of magnitude,
- 10:35perhaps an order of magnitude perhaps more optimists.
- 10:40I think has Maybe a to order of magnitude potential Improvement
- 10:49in economic output. Like, it's not fair, it's not
- 10:54fair. What the limit we actually even
- 10:56is so, But we need to do is in the right way.
- 11:03We need to do it carefully and safely and ensure that the
- 11:07outcome is one, that is beneficial to civilization and
- 11:12one that Humanity wants can't this is also extremely important
- 11:17obviously. So And I hope you will consider
- 11:24joining Tesla to achieve those goals.
- 11:31It has a we really care about doing the right thing here or
- 11:34Spire to do the right thing and and really not pay the road to
- 11:39hell with good intentions. And I think the road is the road
- 11:41to hell is mostly paved with bad intentions, but every now and
- 11:43again, there's a good intention in there.
- 11:46So we wanted to do the right thing.
- 11:48So you know, consider joining us and helping make it happen with
- 11:52that. Let's, let's move on to the next
- 11:54phase. All right.
- 12:03So you seen a couple of robots today.
- 12:05Let's do a quick timeline recap. So last year we unveiled, the
- 12:09Tesla bought concept but a concept doesn't get us very far.
- 12:12We knew we needed a real development and integration
- 12:15platform to get real life learnings as quickly as
- 12:17possible. So that robot that came out and
- 12:20did the little routine for you guys.
- 12:21We had that within six months built working on software
- 12:25integration Hardware AIDS over the months since then.
- 12:29But in parallel, we've also been designing the Next Generation.
- 12:32This one over here. So this guy is rooted in the
- 12:37foundation of sort of the vehicle design process.
- 12:40You know, we're leveraging. All of those learnings that we
- 12:42already have. Obviously there's a lot that's
- 12:45changed since last year but there's a few things that are
- 12:47still the same. You'll notice we still have this
- 12:49really detailed focus on the true human form.
- 12:52We think that matters for a few reasons but it's fun we spend a
- 12:55lot of time thinking about how amazing the human body is.
- 12:59We have this incredible range of motion typically really amazing
- 13:03strength. Fun exercises.
- 13:05If you put your finger tip on the chair in front of you,
- 13:08you'll notice that there's a huge range of motion that you
- 13:12have in your shoulder and your elbow, for example, without
- 13:14moving your fingertip, you can move those joints all over the
- 13:17place. But the robot, you know, it's
- 13:20been a function is to do real useful work and it maybe doesn't
- 13:24necessarily need all of those degrees of freedom right away.
- 13:27So we've stripped it down to a minimum sort of 28 fundamental
- 13:30degrees of freedom. And then, of course, our hands
- 13:32in addition to that, Humans are also pretty efficient at some
- 13:37things and not so efficient in other times.
- 13:39So for example, we can eat a small amount of food to sustain
- 13:42ourselves for several hours, that's great, but when we're
- 13:46just kind of sitting around no offense, but we're kind of
- 13:49inefficient, we're just sort of burning energy.
- 13:52So on the robot platform where we're going to do is we're going
- 13:54to minimize that. Idle power consumption, drop it
- 13:56as low as possible. And that way we can just flip a
- 13:59switch and immediately the robot turns into something that does
- 14:02useful work. So let's talk about this latest
- 14:06generation in some detail, shall we?
- 14:09So on the screen here you'll see in Orange are actuators which
- 14:12we'll get to in a little bit. And in blue our electrical
- 14:15system. So now that we have our sort of
- 14:18human based research and we have our first development platform,
- 14:22we have both research and execution to draw from for this
- 14:25design. Again, we're using that vehicle
- 14:27design foundations. So we're taking it from concept,
- 14:30through design and Analysis and then build and validation.
- 14:36Along the way. We're going to optimize for
- 14:37things like cost and efficiency because those are critical
- 14:40metrics to take this product to scale eventually.
- 14:44How are we going to do that? Well, we're going to reduce our
- 14:46part count and our power consumption of every element
- 14:49possible. We're going to do things like
- 14:51reduce the sensing in the wiring at our extremities.
- 14:54You can imagine a lot of mass in your hands and feet is going to
- 14:57be quite difficult and power consumption of to move around.
- 15:01And we're going to centralize both our power distribution and
- 15:04our compute to the physical center of the platform.
- 15:08So in the middle of our tour, so actually it is the tour.
- 15:11So we have our battery pack. This is sized at 2.3 kilowatt
- 15:15hours which is perfect for about a full day's worth of work.
- 15:19What's really unique about this battery pack is it has all of
- 15:21the battery Electronics integrated into a single PC be
- 15:25within the pack. So that means everything from
- 15:27sensing to fusing charge management and power
- 15:32distribution is all on one, all in one place.
- 15:36We're also leveraging both our Vehicle Products and our Energy
- 15:40Products to roll all of those key features into this battery.
- 15:44So that's streamlined, manufacturing, really efficient
- 15:48and simple cooling, methods battery management and also
- 15:51safety. And of course we can leverage
- 15:54Tesla's, existing infrastructure and supply chain to make it.
- 15:59So going on to sort of our brain, it's not in the head but
- 16:02it's pretty close also in our tour so we have our Central
- 16:06Computer. So as you know, Tesla already
- 16:08ships, full self-driving, computers in every vehicle, we
- 16:11produce. We want to leverage both the
- 16:14autopilot hardware and the software for the humanoid
- 16:17platform, but because it's different in requirements and in
- 16:20form factor we're going to change a few things first.
- 16:23So we still are going to, it's going to do everything that a
- 16:26human brain does processing Vision data.
- 16:29You making split-second decisions based on multiple
- 16:32sensory inputs and also Communications.
- 16:35So to support Communications, it's equipped with wireless
- 16:38connectivity as well as audio support and then it also has
- 16:42Hardware level security features, which are important to
- 16:44protect both the robot and the people around the robot.
- 16:49So now that we have our sort of core, we're going to need some
- 16:53limbs on this guy and we'd love to show you a little bit about
- 16:56our actuators and are fully functional hands as well.
- 16:59But the first before we do that, I'd like to introduce Malcolm.
- 17:02Who's going to speak a little bit about our structural
- 17:04foundation for the robot. Thank you.
- 17:17Tessa have the capabilities finalized, highly complex
- 17:19systems. Don't get much more complex than
- 17:22a crash. You can see here a simulated
- 17:25crash and bottle three superimposed on top of the
- 17:27actual physical crash, it's actually incredible how accurate
- 17:31it is. Just to give you an idea of the
- 17:33complexity of this model. It includes every nut bolt and
- 17:36washer every spot Weld and it has 35 million degrees of
- 17:40freedom, quite amazing. And it's true to say that if we
- 17:44didn't have models like this, we wouldn't be able to make the
- 17:46safest Dazzle the world. So can we utilize our
- 17:50capabilities are methods from the automotive side to influence
- 17:54a robot? Well, we can make a bottle.
- 17:59And since we had crashed software use the same software
- 18:01here, we can make it fall down the purpose of.
- 18:04This is to make sure that if it falls down, ideally, it doesn't.
- 18:07But it's superficial damage. We don't want it to.
- 18:11For example, break its gearbox, that is arms.
- 18:13That's equivalent of a dislocated shoulder is a robot
- 18:17difficult and expensive to fix. So, we wanted to dust yourself
- 18:20off. Get on with the job.
- 18:21It's been different. We've also take the same bottle
- 18:27and we can drive the actuators, using the inputs from a
- 18:30previously solve bottle. Bring it to life.
- 18:34So, this is producing the Motions for the tasks.
- 18:37We want the robot to do, these tasks are picking up boxes.
- 18:40Turning squatting, walking upstairs, whatever the set of
- 18:44tasks are, we can place them bottle.
- 18:46This is showing just simple walking, we can create the
- 18:49stresses it all our components that helps us optimize the
- 18:52components These are not dancing robots.
- 18:58These are actually the modal Behavior, the first five modes
- 19:00of the robot. Typically when people make
- 19:03robots they make sure the first mode is up around the top single
- 19:07figures of towards 10 Hertz, who is a do this is to make the
- 19:11controls of walking easier for difficult to walk.
- 19:14If you can't guarantee where your footer, it wobbled around
- 19:17That's okay to make one robot. We want to make thousands, maybe
- 19:20Millions. We haven't got the luxury of
- 19:23making from carbon fibre and titanium.
- 19:25We want to make them on plastic, things are not quite as stiff.
- 19:29So we can't have these high targets of call them, dumb
- 19:32targets. We've got to make them work at
- 19:35lower targets. So is that is that going to
- 19:37work? Well, if you think about it,
- 19:39sorry about this but we just bags of soggy jelly and Bones
- 19:43thrown in. We're not high frequency.
- 19:46If I started my leg I don't vibrate at 10 Hertz.
- 19:50We people operate a lot of frequency so we know the robot
- 19:53actually. Can It just fits controls
- 19:55harder. So we take the information from
- 19:57this, the modal data and the stiffness of feeder into the
- 20:00control system that allows it to walk.
- 20:05Just changing tack slyly looking at the knee we could take some
- 20:09inspiration from biology and we can look to see what the
- 20:12mechanical advantages of the knee.
- 20:14Is it turns out? It actually represent quite
- 20:16similar for but a leak and that's quite nonlinear.
- 20:20That's not surprising really because if you think when you
- 20:22bend your leg down the talk on your knee is much more when it's
- 20:26bent, that is when it's straight.
- 20:28So you'd expect a nonlinear function and in fact that the
- 20:32biology is nonlinear this matched it quite accurately.
- 20:37So that's a reputation for by link is obviously don't
- 20:39physically for by link to the said the characteristics are
- 20:42similar but me Benny down that's not very scientific.
- 20:46Let's be a bit more scientific. We played all the tasks through
- 20:50the through this graph for this is showing picking things up
- 20:53walking squatting. The Tesla said we did on the
- 20:56stress and that's the talk at seen at the knee against the
- 21:02knee, bend on the horizontal axis, this is showing the
- 21:05requirement for the need to do all these tasks.
- 21:07And then put a curve through it. Surfing over the top of the
- 21:10piece and that's say this is what's required to make the
- 21:14robot. Do these tasks.
- 21:19So if we look at the four bar link, that's actually the green
- 21:22curve, and it's saying that the non-linearity of the 4 bar leak,
- 21:26is actually linearized. The characteristic of the force,
- 21:29what that really says, is that slowed, the force?
- 21:31That's the biggest, the actuator have the lowest possible Force,
- 21:34which is the most efficient we want to burn ended up.
- 21:36Slowly, what's the blue curve for the blue curve is actually,
- 21:40if we didn't have a four bar link, we just had an arm
- 21:42sticking out of my leg here, with an actuator on it.
- 21:46A simple tube our leak. That's the best you could do.
- 21:49Simple to bar link and it shows that that would create a much
- 21:52more force in the actuator we should not be efficient.
- 21:57So what they look like in practice?
- 21:59Well, It's just see, but it's very tightly packaged.
- 22:03In the knee you'll see a good chance found on the second.
- 22:05You'll see the full bar link there is operating on the
- 22:08actuator. This is determined the force and
- 22:11the displacement on the actuator and that pass you over to
- 22:14constitute us to tell you a lot more detail about how these
- 22:16actuators and they didn't design optimized.
- 22:19Thank you. Hey There, Michael.
- 22:28So I am, I would like to talk to you about the design process and
- 22:34actuator portfolio in our robot. So there are many similarities
- 22:39between a car and a robot. When it comes to powertrain
- 22:42design, they may be. The most important thing that
- 22:45matters here is energy, mass and cost.
- 22:48We are carrying over most of our designing experience from the
- 22:52car to the robot. So in the particular case, you
- 22:58see a car with to drive units and the drive units are used in
- 23:03order to accelerate the car zero to 60 MPH time or drive a City
- 23:09Drive side. While the robot that has 28
- 23:13actuators, it's not obvious. What are the tasks actuator
- 23:18level? So we have tasks that are higher
- 23:22level like walking or climbing stairs or Carrying a heavy
- 23:26object, which need to be translated into joint in to join
- 23:31specs. Therefore, we use our model,
- 23:34that generates the torque speed trajectories for our joints,
- 23:39which subsequently is going to be fed in our optimization model
- 23:44to run through the optimization process.
- 23:49This is one of the scenarios that the robot is capable of
- 23:52doing which is turning and walking.
- 23:55So when we have this torque speed trajectory, We Lay it over
- 23:59inefficiency map of an actuator. And we are able along the
- 24:04trajectory to generate the power consumption and the energy
- 24:09cumulative energy for the task versus time.
- 24:13So this allows us to define the system cost for the particular
- 24:17actuator and put a simple Point into the cloud.
- 24:20Then we do this for hundreds of thousands of actuators by
- 24:24solving. In our cluster and the red line,
- 24:26denotes, the part of round, which is the preferred area
- 24:30where we will look for optimal. So the X denotes the preferred
- 24:35actuator design, we have picked for this particular joint.
- 24:38So now we need to do this for every joint.
- 24:40We have 28 joints to optimize and we parse our Cloud.
- 24:45We parse our Cloud again for every joint spec.
- 24:49And the Red X is this time the note, the bespoke actuator
- 24:53designs for every joint. The problem here is that we have
- 24:57too many unique actuator designs and even if we take advantage of
- 25:01the Symmetry still, there are too many in order to make
- 25:04something must manufacturable. We need to be able to reduce the
- 25:08amount of unique actuator designs.
- 25:10Therefore we run something called commonality study with we
- 25:14parse our Cloud. Again looking this time for
- 25:18actuators that simultaneously Meet The Joint performance
- 25:21requirements for more than one joint the same time.
- 25:24So the resulting portfolio is six actuators and they show in a
- 25:28color map, the middle figure and the actuators can be also viewed
- 25:34in this slide. We have three rotary and three
- 25:38linear actuators all of which have a great output force or
- 25:43Torque per Mass. The rotary actuator in
- 25:48particular, has a mechanical clutch, integrated on the lock
- 25:51high-speed site, angular contact, ball bearing and on the
- 25:54high-speed side and on the low speed side across roller bearing
- 25:59into the gear. Train is a strain wave gear and
- 26:04there are three integrated sensors here and bespoke
- 26:09permanent magnet machine. The linear actuator.
- 26:19I'm sorry. The linear actuator has
- 26:22planetary rollers and an inverted planetary Screw As a
- 26:26gear train, which allows efficiency and compaction and
- 26:30durability. So in order to demonstrate the
- 26:34force capability of our linear actuators, we have set up an
- 26:39experiment in order to test it under its limits.
- 26:46And I will let you enjoy the video.
- 26:55So, our actuator is able to lift.
- 27:05Science, proves quality sleep, is vital to your mental
- 27:08emotional and physical health. The Sleep Number 360, smart bed,
- 27:12senses, your movements and automatically, adjust to help.
- 27:14Keep you both effortlessly comfortable, and it's
- 27:17temperature balancing. So you stay cool.
- 27:19So you're your best for yourself and those you care about most
- 27:22life-changing. Sleep only from Sleep Number.
- 27:24Don't miss our weekend, special save, 50% on the Sleep.
- 27:27Number 360 Limited Edition, smart bed plus special financing
- 27:29ends, Monday to learn more. Go to sleep number.com, special
- 27:32financing subject to credit approval.
- 27:33Minimum monthly payments quiet see store for Has the Black
- 27:36Friday savings at the Home Depot.
- 27:37You'll find top brand kitchen appliances with Innovative
- 27:40features that can do more to your holidays.
- 27:42Can be more ovens with built-in air, fryers for baking, the
- 27:44perfect cookies, dishwashers with smart Tech to clean
- 27:47everything from bakeware to festive mugs and high-capacity
- 27:51refrigerators to keep leftovers. Fresh chopped Black Friday
- 27:55savings and get up to 30% off, plus instantly save up, to 750
- 27:58on select GE kitchen packages, at the Home, Depot a Dewar's,
- 28:01get more done off of Alabama, second through November 30th us
- 28:04only store online. Details.
- 28:07A halftone. 9-foot concert grand, piano. and, This is a
- 28:19requirement, it's not something nice to have because our muscles
- 28:24can do the same when they are direct driven.
- 28:26When there are directly driven or quadricep muscles can do the
- 28:30same thing. It's just that they need is an
- 28:33up. Hearing linkage system that
- 28:35converts the force into velocity at the end effector of our heels
- 28:39for purposes of giving to the human body agility.
- 28:44So this is one of the main things that are amazing about
- 28:46the human body and I'm concluding my part at this point
- 28:50and I would like to welcome my colleague.
- 28:51Mike was going to talk to you about hand design.
- 28:54Thank you very much. Say something penis.
- 29:01We just saw how powerful a human and a humanoid actuator can be.
- 29:06However, humans are also incredibly dextrous Human hand
- 29:12has the ability to move at 300 degrees per second.
- 29:15There's tens of thousands of tactile sensors.
- 29:19It has the ability to grasp and manipulate almost every object
- 29:22in our daily lives. For our robotic hand design, we
- 29:26were inspired by biology. We have five fingers and
- 29:29opposable thumb. Our fingers are driven by
- 29:32metallic tendons that are both flexible and strong. because the
- 29:36ability to complete wide aperture power grasps while also
- 29:39being optimized for precision gripping of small, thin and
- 29:43delicate objects, So why a human like robotic hand?
- 29:48Well, the main reason is that our factories in the world
- 29:50around, us is designed to be ergonomic.
- 29:53So what that means is that it ensures that objects in our
- 29:55Factory are graspable. But it also ensures that new
- 29:58objects that we may have never seen before can be grasped by
- 30:01the human hand and by our robotic hand as well, the
- 30:05converse there is pretty interesting because it's saying
- 30:08that these objects are designed to our hand instead of having to
- 30:10make changes to our hand, to accompany a new object.
- 30:15Some basic stats about our hand is that a six actuators and 11
- 30:18degrees of freedom? It has an in hand controller,
- 30:21which drives the fingers and receive sensory feedback sensor,
- 30:25feedback is really important to learn a little bit more about
- 30:27the objects that were grasping and also for proprioception and
- 30:31that's the ability for us to recognize where hand is in
- 30:33space. One of the important aspects of
- 30:37our hand is that it's adaptive this adaptability is involved.
- 30:40Essentially as complex mechanisms that allow the hand
- 30:42to adapt, the objects being grasped.
- 30:46Another important part is that we have a non back drivable
- 30:48finger drive this clutching mechanism allows us to hold and
- 30:51transport objects without having to turn on the hand Motors.
- 30:55You just heard how we went about going, we went about designing
- 30:58the test about Hardware. Now I'll hand it off to Milan
- 31:01and our autonomy team to bring this robot to life.
- 31:12All right, so all those cool things we've shown earlier in
- 31:16the video we're posted possible just in a matter of a few
- 31:19months, thanks to the amazing work that we've done a little
- 31:22pilot over the past few years most of those components ported
- 31:26quite easily over to the but environment.
- 31:28If you think about it we're just moving from a robot on Wheels to
- 31:31our beloved on legs. So some of them components are
- 31:34pretty similar and some other require more heavy lifting The
- 31:39for example, our computer vision you are networks or poorly
- 31:43directly from autopilot to the Bots situation.
- 31:47It's exactly the same occupancy Network that will talk into a
- 31:50little bit more details later with the autopilot team, that is
- 31:53now running on the. But here in this video, the only
- 31:56thing that change really is the training data that we had to
- 31:58recollect. We also trying to find ways to
- 32:03improve those occupancy networks.
- 32:06Using work made on your Radiance fields to get really great
- 32:10volumetric, rendering of the butts environments, for example
- 32:13here. So Machinery, that the but might
- 32:15have to interact with. Another interesting problem to
- 32:21think about is in indoor environments, mostly with that
- 32:25sense of GPS signal. How do you get about to navigate
- 32:28to its destination? Say, for instance, to find its
- 32:31nearest charging station. So, we've been training, module
- 32:35networks to identify high frequency features, key points,
- 32:39within the buds, camera streams and track them across frame over
- 32:43time as the but navigates through its environment.
- 32:46And we're using those points to get a better estimate of the
- 32:50butts pose and trajectory within its environment as it's working.
- 32:57We also did quite some work on the simulation side.
- 32:59And this is literally the autopilot simulator to which
- 33:02we've integrated the robot Locomotion code.
- 33:05And this is a video of the motion control code running in
- 33:09your palate, similar simulator showing the evolution of the
- 33:12robots, work overtime. So, as you can see, we started
- 33:15quite slowly in April and start accelerating as you unlock more
- 33:18joints and deeper more Advanced Techniques, like arms balancing
- 33:22over the past few months. So Locomotion is specifically
- 33:26one component that's very different as we're moving from
- 33:29the car to the but environment and so I think it warrants a
- 33:32little bit more depth and I'd like my colleagues to start
- 33:35talking about this now Thank you Milan.
- 33:46Hi everyone. I'm Felix.
- 33:47I'm a robotics engineer on the project and I'm going to talk
- 33:50about walking walking seems easy, right?
- 33:54People do it every day, you don't even have to think about
- 33:57it but there are some aspects of walking which are challenging
- 34:01from an engineering perspective, for example, physical self
- 34:05awareness. That means having a good
- 34:08representation of yourself. What is the length of your
- 34:10limbs? What is the mass of your limbs?
- 34:13What is the size of your feet? That matters also having an
- 34:17energy-efficient gauge, you can imagine just different styles of
- 34:20walking and all of them are equally efficient.
- 34:25Most important, keep balance, don't fall and of course also,
- 34:29coordinate the motion of all of your limbs together.
- 34:33So now humans do all of this naturally but as Engineers or
- 34:37what is this? We have to think about these
- 34:39problems and if I'm going to show you how we address them in
- 34:42our Locomotion planning and control stack, so we start with
- 34:46Locomotion planning and our representation of the bond.
- 34:49That means a model of the robots kinematics, Dynamics, and a
- 34:52contact properties, and using our model and the desired path
- 34:57for the Box. Our Locomotion planner generates
- 35:00reference trajectory is for the entire system.
- 35:04This means feasible trajectories with respect to the assumptions
- 35:08of our model. The planet currently Works in
- 35:11three stages. It's starts planning footsteps
- 35:14and ends with the entire motion photosystem and let's dive a
- 35:18little bit deeper in how this works.
- 35:20So in this video we see footsteps being planned over a
- 35:23planning Horizon following the desired path and we start from
- 35:28this and add. Then for trajectories that
- 35:31connect these footsteps using too often heel strike just as
- 35:34the humans just as humans do and this gives us a larger stride
- 35:39and less You've been for high efficiency of the system.
- 35:43The last stage is then finding a center of mass trajectory, which
- 35:46gives us a feat dynamically feasible, motion of the entire
- 35:49system to keep balance. As we all know, plans are good,
- 35:54but we also have to realize them in reality.
- 35:57Let's say, how see how we can do this.
- 36:08Thank you, Felix. Hello everyone.
- 36:10My name is Anand. I'm going to talk to you about
- 36:12controls. So, let's take the motion plan
- 36:16that Felix just talked about and put it in the real world on a
- 36:20real robot. Let's see what happens.
- 36:25It takes a couple steps and falls down.
- 36:28Well, that's a little disappointing, but we are
- 36:31missing a few key pieces here, which will make it work.
- 36:36Now, as Felix mentioned, the motion planner is using an
- 36:40idealized version of itself and a version of reality around it.
- 36:45This is not exactly correct it. Also expresses its intention
- 36:50through trajectories and wrenches, wrenches, our forces
- 36:54and torques that it wants to exert on the World to locomote.
- 37:00Reality is way more complex than any similar model.
- 37:04Also, the robot is not simplified.
- 37:06It's got vibrations and modes, compliance sensor noise, and on,
- 37:11and on, and on. So what does that do to the real
- 37:15world when you put the body in the real world?
- 37:18Well, the unexpected forces caused on model Dynamics, which
- 37:22essentially the planet doesn't know about and that causes
- 37:24destabilization, especially For A system that is dynamically
- 37:29stable, like bipedal locomotion. So what can we do about it?
- 37:33Well, we measure reality, we use sensors and our understanding of
- 37:38the world to do state estimation and status to me here, you can
- 37:42see the attitude and pelvis pose, which is essentially the
- 37:45vestibular system in a human along with the center of mass
- 37:49trajectory being tracked when the robots walking in the office
- 37:52environment. Now we have all the pieces we
- 37:56need in order to close the loop. So we use our better bot model.
- 38:01We use the understanding of reality that we have gained
- 38:04through State estimation and we compare what we want, which is
- 38:08what we expect the reality expect, the reality is doing to
- 38:11us in order to add corrections to the behavior of the robot.
- 38:18Sure the robot certainly doesn't appreciate being poked but it
- 38:22doesn't admit Bill job of staying upright.
- 38:26The final Point here is a robot that walks is not enough.
- 38:31We need it to use its hands and arms to be useful.
- 38:35Let's talk about manipulation. Hi everyone.
- 38:49My name is Eric, robotics engineer on Shabbat and I want
- 38:53to talk about how we made the robot manipulator things in the
- 38:56real world. We wanted to manipulate objects
- 38:59while looking as natural as possible and also get there
- 39:03quickly. So, what we've done is, we've
- 39:06broken this process down into two steps.
- 39:08First is generating, a library of natural motion references or
- 39:12we could call them demonstrations and then we've
- 39:14adapted these motion references online, to the current
- 39:17real-world. Operation.
- 39:20So let's say we have a human demonstration of picking up an
- 39:23object. We can get a motion capture of
- 39:25that demonstration which is visualized right here as a bunch
- 39:29of key frames representing the locations, the hands, the elbows
- 39:32at or so we can map that to the robot using inverse kinematics.
- 39:37And if we collect a lot of leaves, now, we have a library
- 39:40that we can work with. But a single demonstration is
- 39:45not generalizable to the variation in the real world.
- 39:48For instance, this would only work for a box in a very
- 39:51particular local location. So we've also done is run these
- 39:56reference trajectories through a trajectory optimization program,
- 40:00which solves for where the hand should be, how the robot should
- 40:03balance during when it meets the adapt the motion to the real
- 40:08world. So for instance, if the box is
- 40:11In this location. Then our Optimizer will create
- 40:15this trajectory instead. Next Lawns going to talk about.
- 40:22What's next for The Optimist. Tesla lat.
- 40:24Thanks Thanks. Right.
- 40:33So hopefully by now you guys got a good idea of what we've been
- 40:35up to over the past few months. We start having something that's
- 40:39usable but it's far from being useful to still a long and
- 40:43exciting road ahead of us. I think the first thing within
- 40:46the next few weeks is to get Optimist at least at par with
- 40:50one both. See the other but prototype you
- 40:52saw earlier and probably Beyond. We also going to start focusing
- 40:57on the real, use case at one of our factories and really going
- 41:01to try to try to Nail this down and I run out all the elements
- 41:05needed to deploy. This product in the real world,
- 41:08I was mentioning earlier, you know, indoor navigation,
- 41:11graceful for management or even servicing.
- 41:14All components needed to scale this product up but I don't know
- 41:19about you. But after seeing what we shown
- 41:22tonight I'm pretty sure we can get this done within the next
- 41:24few months or years and make this project a reality and
- 41:28change the entire economy. So I would like to thank the
- 41:31entire Optimus team for their hard work over the past few
- 41:35months. I think it's pretty amazing.
- 41:36All of this was done in barely six or eight months.
- 41:39Thank you very much. The kids run, Richard JC Penney
- 41:51for thousands of deals solo. No coupons needed this weekend,
- 41:55save up to 50% on kitchen electrics from Brands, like,
- 41:57Keurig and Cuisinart, say yes, please to Diamonds and
- 42:00gemstones. Now, 1999 each and bundle up the
- 42:03famine coats starting at 1499. We got your holiday excluded
- 42:13from coupons. Exclusions.
- 42:15Apply see storage acb.com Tails the Venture X card from Capital.
- 42:19One gives you more of what you love like premium travel
- 42:22benefits and access to Taylor Swift tickets.
- 42:25I do love her or in five times miles on flights and 10 times.
- 42:29Miles on hotels through Capital One travel.
- 42:31Enjoy your stay in sweet 13, 12 13.
- 42:34That's Taylor's. Lucky number plus, get access to
- 42:37Taylor Swift. The era's tour presented by
- 42:39Capital One. Maybe I'll see you there.
- 42:41The Venture X card from Capital One.
- 42:43What's in your wallet terms applying see capitalone.com for
- 42:46Details. Hey, everyone.
- 42:58I am Ashok. I lead the autopilot team
- 43:00alongside Milan. God is coming so hard to top
- 43:04that Optimus section will try nonetheless.
- 43:09Anyway. Every Tesla that has been built
- 43:12over the last several years, you think has the hardware to make
- 43:16the car drive itself? We have been working on the
- 43:19software to add higher and higher levels of autonomy.
- 43:24This time around last year, we are roughly 2,000 cars driving.
- 43:28Our fsdb DOT software. Since then, we have
- 43:32significantly improved the soft press robustness and capability
- 43:35that we have. Now shifted to 160,000 customers
- 43:38as of today, This is not come for free.
- 43:49It came from the sweat and blood of the engineering team for the
- 43:52last one year. For example, we train 75,000
- 43:57neural network models. Just last one year, that's
- 44:00roughly model, every eight minutes that's coming out of the
- 44:04team and then we evaluate them on a large clusters and then we
- 44:07ship 280. One of those models that
- 44:10actually improve the performance of the car.
- 44:12And this pace of innovation is happening throughout the stack.
- 44:16The the planning software, the infrastructure of the tools even
- 44:19hiring everything is progressing to the next level.
- 44:25The fsdb, the software is quite capable of driving the car.
- 44:30You should be able to navigate from parking lot parking lot
- 44:32aling CDC driving stopping for traffic lights and stop signs
- 44:37negotiating with objects at intersections making turns and
- 44:40so on. All of this comes from the
- 44:46camera streams that go through neural networks that run on the
- 44:49car itself. It's not coming back to the
- 44:50server or anything. It runs on the car and produces
- 44:53all the outputs to form the world model around the car and
- 44:56the planning software drives the car based on that.
- 45:01Today, we'll go into a lot of the components that make up the
- 45:03system. The occupancy Network acts as
- 45:07the base geometry layer of the system.
- 45:10This is a multi-camera video. You don't Network That, from the
- 45:15images predicts, the full physical occupancy of the world
- 45:19around the robot. So anything that's physically
- 45:22present trees walls, buildings course balls, what have you it
- 45:27predicts, which physically present it?
- 45:28Predicts them along, with their future motion?
- 45:34On top of this base level of geometry, we have more semantic
- 45:38layers in order to navigate the roadways, we need the lens of
- 45:41course. But then the roadways have lots
- 45:44of different lanes and they connect in all kinds of ways.
- 45:47So it's actually really difficult problem for typical
- 45:50computer vision techniques to predict the set of planes and
- 45:52the connectivities. So we reach all the way into
- 45:55language Technologies. And then pull the
- 45:57state-of-the-art from other domains are not just computer
- 46:00vision to make this task possible.
- 46:04For vehicles, we need that full kinematic state to control for
- 46:07them. All of this directly comes from
- 46:11neural networks. Video streams raw video streams.
- 46:14Come into the Network's. Go through a lot of processing
- 46:17and then outputs the full kinematic state that positions
- 46:20velocities acceleration. Jerk all of that directly comes
- 46:24out a networks with minimal post-processing.
- 46:26That's really fascinating to me. Because how is this even
- 46:29possible? What world do we live in that?
- 46:31This magic is possible that this network's predicts fourth,
- 46:34derivatives of these positions, when people thought we couldn't
- 46:37even detect these objects, My opinion is that is not come for
- 46:42free, it recommend tons of data. So we had to bit sophisticated
- 46:47Auto leveling systems that Shone through raw sensor data, run on
- 46:51ton of offline compute on the server's.
- 46:53It can take a few hours run expensive, neural networks.
- 46:57This is the information into labels that frame.
- 47:00Our in-car neural networks On top of this.
- 47:04We also use our simulation system to synthetically create
- 47:07images and since it's a simulation, it's trivially have
- 47:10all the labels. All of this goes through a
- 47:15well-oiled data engine pipeline where we first trained, a
- 47:19baseline model with some data, ship it to the car, see what the
- 47:23failures are. And once we know the failures,
- 47:26we mine the fleet for the cases where it fails provide the
- 47:29correct labels and add the data to the training.
- 47:32Set this process, systematically fixes the issues.
- 47:36And we do this for every task that runs in the car.
- 47:40Yeah, and to train these new, massive new our networks.
- 47:42This year, we expanded our training infrastructure by
- 47:45roughly 40 to 50% so that it starts at about 14,000 gpus
- 47:50today across multiple training clusters in the United States.
- 47:55We also worked on Rai competitor, which now supports
- 47:58new operations needed by those neural networks and map them to
- 48:01the best of our underlying Hardware resources.
- 48:05And our inference engine today is capable of Distributing the
- 48:09execution of a single neuron Network.
- 48:11Across two independent system on chips.
- 48:14Essentially, two independent, computers interconnected within
- 48:17the simple, self-driving computer, and to make this
- 48:20possible, we have to keep a tight control on the end-to-end
- 48:23latency of this new system. So we deployed more advanced
- 48:26scheduling code across the 40 50 platform.
- 48:31All of this neural networks running in the car to get that
- 48:34produced the vector space, which is again, the model of the world
- 48:36around the robot of the car, then the planning system
- 48:39operates on top of this coming up with trajectories that avoid
- 48:42collisions or smooth, make progress towards the
- 48:45destination. Using a combination of model
- 48:47based optimization plus neural network that helps optimize it
- 48:51to be really fast. Today we are very excited to
- 48:56present progress on all of these areas.
- 48:59We have the engineering leads standing by to come in and
- 49:01explain these various blocks and these power not just the car,
- 49:05but the same components also run on The Optimist robot that Milan
- 49:08showed earlier. With that, I welcome people to
- 49:11start talking about the planning section.
- 49:20Hi, all I a real giant. Let's use this intersection
- 49:25ciliary to dive straight into how we do the planning and
- 49:28decision-making in autopilot. So, we are approaching this
- 49:32intersection from a side street and we have to yield to all the
- 49:35crossing vehicles. Right.
- 49:37Miss, as we are about to enter the intersection.
- 49:40The Pedestrian, on the other side of the intersection decides
- 49:43to cross the road without a crosswalk.
- 49:45Now, we need to yield to this pedestrian, yield to the
- 49:48vehicles from the right? And also understand the relation
- 49:51between The Pedestrian and the vehicle on the other side of the
- 49:54intersection. So, a lot of these entry object
- 49:58dependencies that we need to resolve in a quick glance.
- 50:03And humans are really good at this.
- 50:05We look at a scene, understand all the possible interactions
- 50:09evaluate the most promising ones, and generally end up.
- 50:12Choosing a reasonable one. so, let's look at a few of these
- 50:16interactions that autopilot system evaluated, We could have
- 50:20gone in front of this pedestrian with a very aggressive launch
- 50:23General lateral profile. Now obviously we are being a
- 50:25jerk to The Pedestrian and you would spook The Pedestrian and
- 50:28is cute but We could have moved forward slowly short for a gap
- 50:33between The Pedestrian, or L, the vehicle from the right.
- 50:36Again, we are being a jerk to the vehicle coming from the
- 50:38right, but you should not outright reject.
- 50:41This interaction in case this is only safe interaction available.
- 50:46Lastly, the interaction we ended up choosing stay slow initially,
- 50:51find reasonable Gap and then finish them anywhere.
- 50:53After all the agents past. Evaluation of all of these
- 50:58interactions is not trivial, especially when you care about
- 51:02modeling the higher order, derivatives for other agents.
- 51:06For example, what is The Logical jerk required by the vehicle
- 51:09coming from the right? When you assert in front of it,
- 51:13Relying purely on collision. Checks with modular predictions
- 51:16will only get you so far because you will miss out on a lot of
- 51:19valid interactions. This basically boils down to
- 51:22solving the multi-agent, joint trajectory planning problem over
- 51:26the trajectories of ego and all the other agents.
- 51:30Now how much ever you optimize? There's going to be a limit to
- 51:32how fast you can run this optimization problem, it will be
- 51:35close to close to order of 10 milliseconds even after a lot of
- 51:38incremental consummation Now, for a typical crowded
- 51:43unprotected, left say you have more than 20 objects, each
- 51:47object, having multiple different future modes, the
- 51:50number of relevant interaction combinations will blow up.
- 51:56Will the planner needs to make a decision, every 50 milliseconds.
- 51:59So how do we solve this in real time?
- 52:03we rely on a framework, what we call as interaction search,
- 52:05which is basically a paralyzed research over a bunch of
- 52:08maneuver trajectories The state space here corresponds to the
- 52:13kinematic state of ego, the kinematic state of other agents,
- 52:16the nominal future multiple multi model predictions and all
- 52:20the static entities in the scene.
- 52:24The action space is is where things get interesting.
- 52:28The user set of maneuver trajectory candidates to Branch
- 52:31over a bunch of interaction decisions and also incremental
- 52:35goals for a longer Horizon maneuver.
- 52:38Let's Walk Through This research very quickly, to get a sense of
- 52:41how it works. We start with a set of vision
- 52:44measurements namely Lanes occupancy, moving objects.
- 52:48These get represented as posture actions as well as latent
- 52:51features. We use this to create a set of
- 52:54gold candidates Lanes. Again, from the latest at work
- 52:58or unstructured regions which correspond to a probability
- 53:00mask, the right from Human demonstration.
- 53:05Once we have a bunch of these gold candidates, we create see
- 53:07trajectories using a combination of classical optimization
- 53:10approaches as well as our Network planner again, trained
- 53:13on data from the customer Fleet. Now, once we get a bunch of
- 53:17these three trajectories we use them to start, branching on the
- 53:21interactions. We find the most critical
- 53:24interaction in our case. This would be the interaction
- 53:27with respect to The Pedestrian, whether we are certain front of
- 53:30it or into it. Obviously, the option on the
- 53:32left is a high penalty option. It likely won't get prioritized.
- 53:37So, we Branch further onto the option on the right, and that's
- 53:40where we bring in more and more complex interactions building.
- 53:43This optimization problem incrementally, with more and
- 53:45more. Ants.
- 53:47And that research keeps flowing, dancing on more interactions,
- 53:50branching and more goals. Now, lot of tricks, here lie, in
- 53:54evaluation of each each of this node of the tree search.
- 53:59Inside each node initially, we started with creating
- 54:02trajectories using classical optimization approaches where
- 54:05the constraints, like I described would be added
- 54:07incrementally. And this would take close to 125
- 54:10milliseconds per action. Now even though this is fairly
- 54:14good. Number when you want to evaluate
- 54:16more than 100 percent fractions, this does not scale.
- 54:20So we ended up building light weight variable networks that
- 54:23you can run in the loop of the planner.
- 54:26These networks are trained on human demonstrations from the
- 54:28fleet as well as offline solvers with relaxed time limits.
- 54:34But this we were able to bring the rundown runtime down to
- 54:37close to 100 microseconds per action.
- 54:46Now, doing this alone is not enough because you still have
- 54:49this massive research that you need to go through and you need
- 54:52to efficiently through the search space.
- 54:55So you need to do a new scoring on each of these.
- 54:58Trajectories few of these are fairly standard.
- 55:00You do a bunch of collision checks, you do a bunch of
- 55:02comfort analysis, what is the jerk and accept required for a
- 55:05given maneuver. The customer Fleet data plays an
- 55:09important role here again. We run two sets of again, light
- 55:13weight variable networks both really augmenting each other.
- 55:17One of them train from interventions from the fsdb, the
- 55:19fleet which gives a score on How likely is a given manure to
- 55:23result in interventions over the next few seconds and second
- 55:27which is purely on human demonstrations.
- 55:29Human driven data, giving a score on how close is your given
- 55:32selected action to a human driven trajectory.
- 55:36This coating helps us to the search space t, branching
- 55:40further, all the interactions and focus the compute on the
- 55:42most promising outcomes. The cool part about this
- 55:48architecture is that it allows us to create a cool blend
- 55:51between our data driven approaches.
- 55:54There you don't have to rely on a lot of hand engineered costs
- 55:57but also grounded in reality with physics-based checks.
- 56:02Now a lot of what I described was with respect to the agents,
- 56:05we could observe in the scene but the same framework extends
- 56:08to objects behind occlusions We use the video feed from eight
- 56:14cameras to generate the 3D occupancy of the world.
- 56:18The blue mask here, corresponds to the visibility region, we
- 56:21call it. It basically gets blocked at the
- 56:25first occlusion. You see in the sea, we consume
- 56:27this visibility mask to generate what we call as ghost objects
- 56:30which you can see on the top left.
- 56:33Now, if you model this Pawn regionals and the state
- 56:35transitions of this ghost object correctly, If you tune your
- 56:40control response as a function of their existence likelihood,
- 56:43you can extract some really nice human like behaviors.
- 56:47Now, I'll pass it on to fill to describe more on how we generate
- 56:50is occupants in it works. Thank you.
- 56:59Hey guys, my name is Phil. I will share the details of the
- 57:02occupancy Network would be able to over the past year.
- 57:06This network is our solution to model, the physical world in 3D
- 57:09around our cars. And it is currently not shown in
- 57:13our customer facing visualization.
- 57:15And what you'll see here is the road Network output from our
- 57:19internal Dev tool, The octopus is that work takes video streams
- 57:25of all our other cameras input, produces a single unified,
- 57:30volumetric occupancy in Vector, space directly for every 3D
- 57:35location around all the car, it predicts the probability of that
- 57:39location being occupied or not, since it has video context, it
- 57:44is capable of predicting obstacles that are occluded
- 57:47instantaneously. For each location.
- 57:51It also produces a state of semantics such as curb car
- 57:56possession and road debris. As color-coded here.
- 58:04Occupancy flow is also predict for motion since the model is a
- 58:08generalized Network. It does not tell static and
- 58:11dynamic objects explicitly. It is able to produce and model
- 58:16the random motion such as a swerving trainer here.
- 58:21This network is currently running or Tesla's with FSD
- 58:25computers and it is incredibly efficient runs about every 10
- 58:29milliseconds with our new linear accelerator.
- 58:33So, how does this work? Let's take a look at the
- 58:35architecture. First, we Rectify each camera
- 58:39images with the camera calibration, and the images were
- 58:42shown here were given to the network is actually not a
- 58:45typical a bit RGB image. As you can see from the first
- 58:49image on top. We're giving the 12 Pedro photo
- 58:52count, imagery to the network. Since it has four bits, more
- 58:56information, it has 16 times better than a branch as well as
- 59:01reduced latency since we don't have to run ISP in the loop
- 59:04anymore. We use a set of red lights and
- 59:07by of FPS, as a backbone to extract images of space features
- 59:13next, we construct a set of 3D position query along with the
- 59:17empty space features as keys and values fit into a attention.
- 59:21Module, the output of the attention module is high
- 59:25dimensional space. Your features This special
- 59:28features are online temporarily using vehicle odometry to derive
- 59:34motion. Not this spatial temporal
- 59:39features goes through a set of deconvolution to produce the
- 59:42final occupancy, and occupancy, flow output.
- 59:45They're formed as fixed size voxel grid, which might not be
- 59:48precise enough for presenting on control in order to get a higher
- 59:53resolution. We also produced per voxel
- 59:55feature Maps, which will feed into MLP, with 3D spatial Point
- 1:00:00queries to get position and semantics at any arbitrary
- 1:00:04location. After knowing the model better,
- 1:00:09let's take a look at another example here.
- 1:00:11We have an articulated bus parked on right side row
- 1:00:14highlighted as an l shaped box up here.
- 1:00:17As we approach the bus start to move the blue, the front of the
- 1:00:21cart turns blue first indicating, the model predicts.
- 1:00:25The front of us has a nonzero occupancy flow.
- 1:00:30And as the bus keeps moving the entire bus turns blue and you
- 1:00:34can also see that the network predicts, the precise curvature
- 1:00:38of the bus. Well, this is a very complicated
- 1:00:42problem for traditional object detection Network.
- 1:00:45As you have, you see, whether I'm going to use one Cube or
- 1:00:48perhaps ask you to feed it the curvature, but for occupancy
- 1:00:51Network, since all we care about is the occupancy in the visible
- 1:00:55space and we will be able to model the curvature precisely.
- 1:01:01Besides the voxel grid, the occupancy Network, also produces
- 1:01:05a driver surface. The driver surface has both 3D
- 1:01:08geometry and semantics. They are very useful for
- 1:01:11control, especially on Healy and the curvy roads.
- 1:01:15The surface and the box of gray are not predicted independently.
- 1:01:19Instead the voxel grid actually aligns with the surface
- 1:01:23implicitly. Here, we are at a here Quest
- 1:01:28where you can see the 3D geometry of the surface bring
- 1:01:32predicted nicely. Printer can usually this
- 1:01:35information to decide the, perhaps we need to slow down low
- 1:01:38for the Hillcrest. And as you can also see the
- 1:01:41voxel, great alliance with the surface consistently, Besides
- 1:01:47the Box source and the surface were so very excited about the
- 1:01:50recent breakthrough, Enduro, Radiance field, or loaf.
- 1:01:55We're looking into both incorporate some of a lie
- 1:01:58smeller features into accessing Network training as well as
- 1:02:01using our Network output as the input state for Nerf.
- 1:02:07As a matter of fact, Ashok is very excited about this.
- 1:02:09This has been his personal weekend project for a while.
- 1:02:17This nurse because I think Academia is building out of
- 1:02:21these Foundation models for language using like tons of
- 1:02:24large data sets for language. Anything for vision, nerves are
- 1:02:28going to provide the foundation models for computer vision
- 1:02:31because they are grounded in geometry and geometry gives us a
- 1:02:35nice way to supervise this networks and freezes of the
- 1:02:37requirement Define. And on top when you need Auto
- 1:02:44Parts, O Reilly Auto.com is a few clicks away.
- 1:02:47We offer convenient options for you to get your parts quickly
- 1:02:50order online and pick up for free.
- 1:02:52At your local, O'Reilly Auto Parts store.
- 1:02:54Will even bring it out curbside or you can have your parts
- 1:02:57delivered right to your door with free shipping on most
- 1:02:59orders over. $35 visit O Reilly Auto.com and the supervision is
- 1:03:11essentially free because you just have to differential be
- 1:03:13rendered these images. So I think in the future this
- 1:03:17all coincidence work idea of our images come in and then the
- 1:03:20network produces consistent, volumetric representation of the
- 1:03:25scene that can then be differentiable rendered into any
- 1:03:28image. That was observed.
- 1:03:30I personally think is the future of computer vision and we do, we
- 1:03:34do some initial work on it right now.
- 1:03:36But I think in the future both Tesla and in the Academia we
- 1:03:40will see that these combination of One-Shot prediction of Full
- 1:03:45automatic occupancy will be the, that's my personal bet.
- 1:03:50Censorship. So here's an example, Ernie
- 1:03:53result of a 3D Reconstruction from our Fleet data.
- 1:03:57Instead of focusing on getting perfect RGB reprojection in
- 1:04:01image space, our primary goal here is to accurately represent
- 1:04:05the wording study space for driving and we want to do this
- 1:04:08for all our Fleet data over the world in all weather and
- 1:04:11lighting conditions. And obviously this is a very
- 1:04:14challenging problem and we're looking for you guys to help.
- 1:04:18Finally, the occupancy network is trained with large out
- 1:04:21labeled data set without any human in the loop.
- 1:04:25And with that, I'll pass to him to talk about what it takes to
- 1:04:28train this network. Thanks though.
- 1:04:36All right, everyone. Let's talk about some training
- 1:04:40infrastructure. So we've seen a couple of videos
- 1:04:43know, four or five, I think and Cara Moore and worry more about
- 1:04:48a lot more Clips on that. So we've been looking at the
- 1:04:52occupancy networks, just from Phil, just fills videos.
- 1:04:56It takes one point four billion frames the train that Network,
- 1:05:00what? You just saw.
- 1:05:01And if you have 100,000 gpus, it would take one hour.
- 1:05:05But if you have one GPU it, Would take 100,000 hours.
- 1:05:10So that is not a Humane time period that you can wait for
- 1:05:13your training job to run, right? We want to ship faster than
- 1:05:15that. So that means you're going to
- 1:05:17need to go parallel. So you need more compute for
- 1:05:20that. That means you're going to need
- 1:05:22a supercomputer. So this is why we've built
- 1:05:25in-house. Three supercomputers comprising
- 1:05:27of 14,000 gpus, where we use 10,000 gpus for training and run
- 1:05:324000 gpus for auto leveling. All these videos are stored in
- 1:05:3830 Theta. B of A distributed managed video
- 1:05:41cash. You shouldn't think of our data
- 1:05:44sets as fixed. Let's say, as you think of your
- 1:05:47image net or something, you know with like a million frames, you
- 1:05:50should think of it as a very fluid thing.
- 1:05:52So we've got a half a million of these videos flowing in and out
- 1:05:56of this cluster, these clusters every single day.
- 1:06:00And we track 400,000 of these kind of python video
- 1:06:04instantiations every And so that is that's a lot of calls.
- 1:06:09We are going to need a capture that in order to govern the
- 1:06:11retention policies of this distributed video cash.
- 1:06:15So, underlying, all of this is a huge amount of infra all of
- 1:06:18which we build and manage in house.
- 1:06:22So you cannot just by, you know, 14,000 views, and then a 30
- 1:06:26petabytes of Flash nvme and just put it together.
- 1:06:29And let's go train. It actually takes a lot of work
- 1:06:32and I'm going to go into a little bit of that.
- 1:06:35What you actually typically want to do is you want to take your
- 1:06:37accelerator so that could be the GPU or Dojo which we'll talk
- 1:06:41about later and because that's the most expensive component,
- 1:06:47that's where you want to put your bottleneck and so that
- 1:06:49means that every single part of your system is going to need to
- 1:06:53outperform this accelerator. And so that is really
- 1:06:56complicated. That means that your storage is
- 1:06:59going to need to have the size and the bandwidth to deliver all
- 1:07:02the data down into the nodes. These nodes need to The right
- 1:07:05amount of CPU and memory capabilities to feed into and
- 1:07:09your machine learning framework. This machine learning framework,
- 1:07:12then is to hand it off to your GPU and then you can start
- 1:07:15training, but then you need to do so across hundreds or
- 1:07:18thousands of GPU in a reliable way in lockstep, and in a way
- 1:07:23that's also fast. So you're also going to need
- 1:07:25interconnect, extremely complicated, we'll talk more of
- 1:07:28a dojo and second So first, I want to take you to some
- 1:07:33optimizations that we've done on our cluster so we're getting in
- 1:07:37a lot of videos and video is very much unlike let's say
- 1:07:41training on images or text, which I think is very well
- 1:07:44established. Video is quite literally a
- 1:07:46dimension, more complicated. And so that's why we needed to
- 1:07:51go end to end from the storage layer, down to the accelerating,
- 1:07:55optimize every single piece of that because we train on the
- 1:07:58photon count videos that come directly from our Fleet, we
- 1:08:02trained on those directly. We do not post processors those
- 1:08:05at all the way it's just done. Is we seek exactly to the
- 1:08:09frames? We select for our batch?
- 1:08:11We load those in including the frames that they depend on.
- 1:08:14So these are your iPhones or a key frames.
- 1:08:16We package those up movement, a shared memory, move them into a
- 1:08:19double bar from the GPU and then use Hardware.
- 1:08:22Decoder, that's only accelerated to actually decode the video.
- 1:08:27So, we do that on the GPU natively and it's all in a very
- 1:08:30nice python. Torch extension doing so unlock
- 1:08:34more than 30% training speed, increase for the occupancy,
- 1:08:37networks and Freda. Basically whole CPU to do any
- 1:08:41other thing. You cannot just do training with
- 1:08:46just videos. Of course, you need some kind of
- 1:08:47a ground truth and that is actually interesting problem as
- 1:08:51well. The objective for storing your
- 1:08:53ground. Truth is that you want to make
- 1:08:56sure you get to your ground truth that you need in the
- 1:08:58minimal amount of file system, operations and load in the
- 1:09:01minimal size of what you need in order to optimize for aggregate
- 1:09:05cross cluster, throughput. Because you should see a
- 1:09:08computer cluster as one big device, which has internally
- 1:09:12fixed constraints. And Households.
- 1:09:14So for this we rolled out a format that is native to us.
- 1:09:19That's called small. We use this for our ground
- 1:09:21truth, our feature cash, and any inference outputs.
- 1:09:24So a lot of dancers that are in there and so just the cartoon
- 1:09:27here, let's say, these are your, is your table that you want to
- 1:09:30store. Then that's how that would look
- 1:09:32out if you rolled out on disk. So, what you do is, you take
- 1:09:35anything you'd want to index on. So, for example, video
- 1:09:37timestamps, you put those all in the header.
- 1:09:40So that in your initial header, eat, you know, exactly where to
- 1:09:43go on this. Then, if you have any tensors,
- 1:09:46you're going to try to transpose the dimensions to put a
- 1:09:49different dimension last as he contiguous Dimension.
- 1:09:52And then also try different types of compression.
- 1:09:55Then you check out which one was most optimal and then store that
- 1:09:58one, this is actually a huge step.
- 1:10:00If you do feature caching unintelligible output from the
- 1:10:03machine Learning Network, rotate around the dimensions, a little
- 1:10:06bit, you can get up to 20 percent increase in efficiency
- 1:10:09of storage. Then when you store that we also
- 1:10:15ordered columns by size so that all your small columns and small
- 1:10:18values are together. So that when you see for a
- 1:10:21single value, you're likely to overlap with a read on more
- 1:10:24values, which we'll use later so that you don't need to do
- 1:10:27another file system, operation. So, I could go on and on, I just
- 1:10:32went on on touch, on two projects that we have
- 1:10:35internally, but this is actually part of a huge, continuous
- 1:10:38effort to optimize the compute that we have in-house.
- 1:10:42So accumulating in aggregating through all these optimizations,
- 1:10:45we now trainer occupancy networks, tries as fast just
- 1:10:48because it's twice as efficient. And now, if we add in bunch more
- 1:10:52compute and go parallel, we can now train this hours instead of
- 1:10:55days. And with that, I'd like to hand
- 1:10:58it off to the biggest user of compute John.
- 1:11:08Hi everybody. My name is John Emmons.
- 1:11:10I lead the autopilot, Vision team.
- 1:11:13I'm going to cover two topics with you today.
- 1:11:15The first is how we predict lanes, and the second is how we
- 1:11:18predict the future behavior of other agents on the road.
- 1:11:22In the early days of autopilot, we modeled the lane detection
- 1:11:25problem is an image space instance, segmentation task.
- 1:11:29Our network was super simple though.
- 1:11:31In fact, it was only capable of producing Lanes from a of a few
- 1:11:34different kinds of geometries specifically, it would segment
- 1:11:38the ego Lane, it could segment adjacent lanes, and then it has
- 1:11:41some special casing for forks and merges this simplistic
- 1:11:44modeling of the problem worked for highly structured roads.
- 1:11:46Like highways, But today we're trying to build a system that's
- 1:11:51capable of much more complex Maneuvers.
- 1:11:53Specifically we want to make left and right turns at
- 1:11:55intersections where the road topology can be quite a bit more
- 1:11:57complex and diverse. When we try to apply this
- 1:12:00simplistic modeling of the problem here, it just totally
- 1:12:03breaks down. Taking a step back for a moment.
- 1:12:07What we're trying to do here is to predict the sparse set of
- 1:12:09lame instances and their connectivity.
- 1:12:12And what we want to do is to have a neural network that
- 1:12:14basically predicts, this graph where the nodes are the lane
- 1:12:17segments in the edges and code the connectivity between these
- 1:12:20Lanes. So we have is our lane
- 1:12:24detection, neural network. It's made up of three
- 1:12:27components. In the first component, we have
- 1:12:30a set of convolutional layers attention layers, in other,
- 1:12:32neural network, layers that encode the video streams from
- 1:12:35our 8 cameras on the vehicle, in produce a rich visual
- 1:12:38representation. We then enhance this digital
- 1:12:42representation with a course roadmap level Road level map
- 1:12:47data, which we encode with a set of additional neural network
- 1:12:50layers that we call the lane guidance module.
- 1:12:53This map is not an HD map but it provides a lot of useful hints
- 1:12:56about the topology of lanes. Instead of intersections the
- 1:12:58lane counts and various roads in a set of other attributes that
- 1:13:01help us. The first two components here,
- 1:13:06produced a dense tensor, that sort of encodes the world.
- 1:13:09But what we really want to do is to convert this dense tensor
- 1:13:12into a smart set of lanes in their connectivities.
- 1:13:15We approach this problem like an image captioning task.
- 1:13:19Where the input is this dense tensor in the output text is
- 1:13:22predicted into a special language that we developed at
- 1:13:24Tesla for encoding Lanes in their connectivities.
- 1:13:27In this language of lanes, the words and tokens are the lane
- 1:13:30positions in 3D space in the ordering of the tokens
- 1:13:34introverted modifiers in the tokens encode, the connective
- 1:13:36relationships between these Lanes By modeling.
- 1:13:39The task is language problem. We can capitalize on recent
- 1:13:42autoregressive, architectures and techniques, from the
- 1:13:45language Community for handling the multiple daily of the
- 1:13:47problem. We're not just solving the
- 1:13:49computer vision problem at autopilot, we're also applying
- 1:13:52the state of the art and language modeling, the machine
- 1:13:53learning more generally I'm not going to dive into a little bit
- 1:13:57more detail, just language component.
- 1:14:01What I have depicted on the screen here is the satellite
- 1:14:03image which sort of represents the local area.
- 1:14:05Around the vehicle, the set of nodes and edges is what we refer
- 1:14:09to as the landgraf and it's ultimately what we want to come
- 1:14:11out of this neural network. We start with a blank slate.
- 1:14:17We're going to make our first prediction here at this Green
- 1:14:19Dot. This green dots position is
- 1:14:22encoded as an index into a coarse grid which discretize has
- 1:14:25the 3D World. Now, we don't predict this index
- 1:14:28directly because it would be too computationally expensive to do.
- 1:14:30So there's just too many grid points and printing a
- 1:14:33categorical distribution over. This has both implications that
- 1:14:36training time and test time. So instead what we do is we just
- 1:14:39recharge the world coarsely first.
- 1:14:41We predict a heat map over the possible locations and then we
- 1:14:44latch in the most probable location condition on this.
- 1:14:48We then refine the prediction and get the precise point.
- 1:14:53Now we know where the position of this token is, we don't know.
- 1:14:55Its type. In this case though, it's a
- 1:14:57beginning of a new Lane. So we refer to it as a start
- 1:15:00token. And because it's a star token,
- 1:15:03there's no additional attributes in our language.
- 1:15:07We then take the predictions from this first forward pass and
- 1:15:09we encode them using a learn traditional and wedding which
- 1:15:12produces a set of tensors that we combine together.
- 1:15:16Which is actually the first word in our language of lanes.
- 1:15:18We had this to the first position our sentence here.
- 1:15:22We then continue this process by preaching.
- 1:15:24The next link point in a similar fashion.
- 1:15:28Now this Lane point is not the beginning of a new Lane.
- 1:15:30It's actually a continuation of the previous Lane so it's a
- 1:15:34continuation token type. Now it's not enough just to know
- 1:15:38that this Lane is connected to the previously prepared to play
- 1:15:40in. We want to encode.
- 1:15:41It's precise geometry which we do by regressing, a set of
- 1:15:44spline coefficients. We then take this Lane, we
- 1:15:49encode it again and added as the next word in the sentence.
- 1:15:53We continue protecting these continuation, Lanes, until we
- 1:15:55get to the end of the prediction grid.
- 1:15:58We then move on to a different Lane segment, so you can see
- 1:16:01that signed up there. Now it's not top allegedly
- 1:16:03connected to that pink point. It's actually forking off of
- 1:16:06them that blue. Sorry that Greenpoint there.
- 1:16:10So it's got a fork type and Fork tokens actually point back to
- 1:16:15previous tokens from which their Fork originates.
- 1:16:19So you can see here, the fourth Point predictor is actually the
- 1:16:21index zero. So it's actually referencing
- 1:16:23back to tokens that has already created like you would in
- 1:16:25language. We continue this process over
- 1:16:29and over again. And so we've enumerated all of
- 1:16:30the tokens in the lane graph and in the network predicts, the end
- 1:16:34of sentence token. Yeah, I just wanted to note
- 1:16:37that. The reason we do this is not
- 1:16:40just because we want to build something complicated.
- 1:16:42It's almost feels like a turing-complete machine here.
- 1:16:44With neural networks though is that we tried simple approaches,
- 1:16:47for example, trying to just segment, the lanes along the
- 1:16:50road or something like that. But then the problem is, when
- 1:16:53this uncertainty say, you cannot see the road clearly and there
- 1:16:56could be two lanes or three lanes, and you can tell a simple
- 1:16:59segmentation based approach would just draw.
- 1:17:01Both of them is kind of a 2.5 Lane situation and the post
- 1:17:05processing, algorithm would collide History fail when the
- 1:17:07predictions are such again, the problems don't end there.
- 1:17:11I mean, you need to predict these connective conditions like
- 1:17:13these connective Lanes inside of intersections, which it's just
- 1:17:16not possible with the approach that ashok's mentioning, which
- 1:17:18is why we had to upgrade to this sort of event like or laps like
- 1:17:20this segmentation would just go Haywire but even if you try very
- 1:17:23hard to you know put them on separate layers it's just really
- 1:17:26hard problem. But languages offers a really
- 1:17:28nice framework for more getting a sample from the posterior, as
- 1:17:33opposed to, you know, trying to do all of this in post
- 1:17:35processing. I didn't exactly stop for this
- 1:17:38autopilot, right? John.
- 1:17:39This can be used for Optimus. Again, you know, I guess they
- 1:17:42wouldn't be called Lanes, but you can imagine, you know, sort
- 1:17:44of in this, you know, stage here that you might have sort of
- 1:17:47paths, that sort of, you know, encode the possible places that
- 1:17:50people could walk. Yeah.
- 1:17:52It's basically, if you're in a factory or a home setting, you
- 1:17:56can just ask the robot. Okay.
- 1:17:57Let's me, please route to the kitchen or please route to some
- 1:18:01location, The Factory, and then we predict the set of Pathways
- 1:18:04that would go through the aisles, like, the robot and say,
- 1:18:06okay, this is how you get to the kitchen.
- 1:18:08It just really gives us a nice framework to model these
- 1:18:10different paths. Let's simplify the navigation,
- 1:18:13problem for the downstream planner.
- 1:18:18All right, so ultimately what we get from this Lane detection
- 1:18:21network is a set of lanes in their connectivities which comes
- 1:18:24directly from the network. There's no additional step here
- 1:18:26for specifying these, you know, dense predictions into into
- 1:18:30dispersed ones. This is just the direct
- 1:18:31unfiltered output of the network.
- 1:18:36Okay, so I talked a little bit about Lanes, I'm going to
- 1:18:38briefly touch on how we model and predict the future paths in
- 1:18:42other semantics on objects. So I'm going to go really
- 1:18:45quickly through two examples. The video on the right here,
- 1:18:48we've got a car. That's actually running a red
- 1:18:50light and turning in front of us.
- 1:18:52What we do to handle situations. Like this is we predict a set of
- 1:18:55short time Horizon future trajectories on all objects.
- 1:18:58We can use these to anticipate the dangerous situation here and
- 1:19:02apply, whatever braking and steering actions required to
- 1:19:04avoid a collision. In the video on the right,
- 1:19:07there's two vehicles in front of us.
- 1:19:09The one of the left lane is parked apparently, it's being
- 1:19:12loaded unloaded, I don't know why the driver decided to park
- 1:19:14there, but the important thing is that our neural network
- 1:19:17predicted it, that it was stopped, which is the red color
- 1:19:20there. The vehicle in the other lane
- 1:19:22is, you notice also the stationary, but that one's
- 1:19:24obviously just waiting for that red light to turn green.
- 1:19:26So even though both objects are stationary and have zero
- 1:19:28velocity, it's the semantics that is really important here so
- 1:19:32that we don't get stuck behind that, awkwardly parked car.
- 1:19:37Predicting, all of these agent attributes, present, some
- 1:19:39practical problems, when trying to build a real-time system, we
- 1:19:42need to maximize the framerate of her optic section stack.
- 1:19:45So that autopilot can quickly react to the changing
- 1:19:47environment. Every millisecond, really
- 1:19:49matters here to minimize the inference latency.
- 1:19:52Our neural network is split into two phases.
- 1:19:55In the first phase we identified locations in 3D space where
- 1:19:59agents exist in the second stage.
- 1:20:01We then pull out tensors at those three locations.
- 1:20:04A pendant with additional data that's on the vehicle.
- 1:20:07And then we do the rest of the processing.
- 1:20:10This first vacation step allows the neural number to focus
- 1:20:12compute on the areas. That matter most, which gives us
- 1:20:15Superior performance for a fraction of the latency cost.
- 1:20:19So putting it all together, the autopilot business that creates
- 1:20:22more than just the geometry and kinematics the world.
- 1:20:24It also predicts a rich set of semantics which enable safe in
- 1:20:27human like driving. I'm not going to hand things off
- 1:20:30to streak. Will tell us how we run.
- 1:20:31All these cool neural networks on our FSD computer.
- 1:20:33Thank you. Hi everyone, I'm free today I'm
- 1:20:45going to get glimpse of what it takes to run this efficient
- 1:20:47networks in the car and how do we optimize for the inference
- 1:20:50latency? Today, I'm going to focus just
- 1:20:53on FS Elaine's Network that John just talked about.
- 1:20:59So when you started this track, we wanted to know if we can run
- 1:21:03this FS Elaine stadtwerke natively on the trip engine,
- 1:21:06which is our in-house neural network accelerator that we
- 1:21:09built in the officially computer.
- 1:21:11When we build this Hardware, we kept it simple and we made sure
- 1:21:16it can do one thing. Ridiculously fast dense dot
- 1:21:19products, but this architecture is auto regressive, and I
- 1:21:24traitor where it crunches through multiple attention,
- 1:21:27attention blocks in the Inner Loop.
- 1:21:29Producing sparse points directly at every step.
- 1:21:32So it's the challenge here was, how can we do this sparse Point,
- 1:21:36prediction, and sparse computation on a dense dot
- 1:21:38product engine. Let's see how we did this on the
- 1:21:41trip. So the network predicts the heat
- 1:21:46map of most probable spatial locations of the point.
- 1:21:50Now, we do a ARG Max and a one heart operation, which gives the
- 1:21:55one hard encoding of the index of the spatial location.
- 1:21:59Now we need to select the embedding associated with this
- 1:22:02index from an embedding table that is learned during training
- 1:22:07to do this on trip. We actually built a lookup table
- 1:22:10in s RAM and we engineer the dimensions of this embedding
- 1:22:14such that we could achieve all of this thing with just matrix
- 1:22:18multiplication. Not just that, we also wanted to
- 1:22:23store this embedding into a token cash so that we don't re
- 1:22:26compute this for every iteration rather, we use it for future
- 1:22:29Point prediction. Again, we pull some tricks here
- 1:22:32where we did all these operations just on the dot
- 1:22:35product engine. It's actually cool that our team
- 1:22:39found creative ways to map all these operations on the trip
- 1:22:42engine in ways that were not even imagined when this Hardware
- 1:22:46was designed. But that's not the only thing we
- 1:22:50had to do to make this work. We actually implemented a whole
- 1:22:54lot of operations and features to make this model compilable to
- 1:22:59improve the in Tate accuracy as well as to optimize performance.
- 1:23:03All of these things helped us run the 75 million parameter
- 1:23:07model just under 10 milliseconds of latency consuming just 8
- 1:23:11watts of power. But this is not the only
- 1:23:16architecture running in the car. There are so many other
- 1:23:18architectures modules and networks in running the car to
- 1:23:23give a sense of scale. There are about a billion
- 1:23:26parameters of all the networks combined producing around
- 1:23:29thousand your network single signals.
- 1:23:32So we need to make sure we optimize them jointly and such
- 1:23:37that we maximize the computer glaciation throughput and
- 1:23:40minimize the latency. So we built a compiler, just for
- 1:23:46neural networks that shares the structure to traditional
- 1:23:49compilers. As you can see, it takes the
- 1:23:52massive graph of urine Nets with 150 K nodes until scientific a
- 1:23:57connection takes this thing, partitions them.
- 1:24:00Into independent sub graphs and come Compares.
- 1:24:03Each of the sub graphs natively for the inference devices.
- 1:24:07Then we have a neural network Linker which shares the
- 1:24:10structure to traditional Linker. When we perform this link time
- 1:24:14optimization there, we solve an offline optimization problem for
- 1:24:20with computer memory and memory bandwidth constraints, so that
- 1:24:24it comes with an optimized schedule that gets executed in
- 1:24:26the car. On the run time.
- 1:24:29We designed a hybrid scheduling system, which basically does
- 1:24:33heterogeneous scheduling on one SOC and distributed scheduling
- 1:24:37across both the soc's to run these networks in a model
- 1:24:40parallel fashion, to get 100 tops of computerization.
- 1:24:45We need to optimize across all the layers of software, right?
- 1:24:49From turning the network, architecture the compiler, all
- 1:24:52the way to implementing a low latency high bandwidth RDMA link
- 1:24:56across both the soc's, And in fact going even deeper to
- 1:25:00understanding and optimizing the cache coherent and noncoherent
- 1:25:04data Paths of the accelerator in the soc.
- 1:25:07This is a lot of organization at every level in order to make
- 1:25:10sure we get the highest frame rate and as every millisecond
- 1:25:14counts here, And this is, this is just the, this is the
- 1:25:21visualization of the neural networks running in the car.
- 1:25:24This is a digital brain. Essentially, as you can see,
- 1:25:27these operators are nothing but just the matrix multiplication,
- 1:25:30convolution to name a few real operations running in the car.
- 1:25:35To train and train this network with the billion parameters, you
- 1:25:38need a lot of label data. So again is going to talk about
- 1:25:42how do we achieve this with the auto leveling pipeline Thank
- 1:25:54you, sure. Hi everyone.
- 1:25:56I'm Yang yanzhao and I'm leaving a symmetric fission and
- 1:25:59autopilot. So yeah, let's talk about Auto
- 1:26:04labeling. So we have several kinds of all
- 1:26:08the labeling Frameworks to support various types of
- 1:26:11networks. But today I'd like to focus on
- 1:26:14the awesome Lanes not here. So to successfully trained and
- 1:26:19generalize this network to everywhere.
- 1:26:22We think we went tens of millions of troops from probably
- 1:26:261 million, 1 million intersection or even more.
- 1:26:30So How to do that. So it is certainly achievable to
- 1:26:37Source sufficient amount of trips because we already have as
- 1:26:40team explained earlier, we already have like, five hundred
- 1:26:43thousand trips per day, cash rate, however, converting all
- 1:26:48those data into a training form is a very challenging technical
- 1:26:51problem. To solve this challenge, we've
- 1:26:56tried various ways of manual and auto labeling.
- 1:26:59So from the First Column to the second from the second to the
- 1:27:03third, each Advance provided as nearly hundred X Improvement in
- 1:27:07throughput, but still, we won an even better or labeling machine.
- 1:27:12That can provide providers providers, good quality,
- 1:27:16diversity, and scalability. To meet all these requirements.
- 1:27:23Despite the huge amount of engineering effort required
- 1:27:26here, we've developed a new auto labeling machine, powered by
- 1:27:30multitude of reconstruction. So these can replace five
- 1:27:35million hours of manual labeling with just 12 hours on cluster
- 1:27:39for labeling 10,000 troops. So how we solved, there are
- 1:27:44three big steps. The first step is high Precision
- 1:27:47trajectory and structure recovery by multi-camera fissure
- 1:27:50inertia odometry. So here all the features
- 1:27:54including ground, surface are inferred from videos by neural
- 1:27:56networks, then tracked and reconstructed in the vector
- 1:28:00space. So the typical drift weight of
- 1:28:04this trajectory in Coral is like 1.3 centimeters per meter and
- 1:28:080.45 million per meter, which is pretty decent considering its
- 1:28:13compact compared requirement. Then the recovery surface and
- 1:28:17roll details are also use as a strong guidance for the later.
- 1:28:20Manual, verification stop. This is also enabled in every
- 1:28:25FSD. Beagle.
- 1:28:26So we get pre-process trajectories and structures.
- 1:28:29Along with the trip data The second step is 42 reconstruction
- 1:28:37which is the peak and core piece of this machine.
- 1:28:40So the video shows how the previously shown trip is
- 1:28:43reconstructed and aligned with other trips.
- 1:28:46Basically other trees from different people, not the same
- 1:28:49vehicle. So this is done by multiple
- 1:28:51internships. Like coarse alignment pairwise
- 1:28:54matching join optimization. Then furthers surface
- 1:28:57refinement, in the end. The human analyst comes in and
- 1:29:02finalizes the label. So each heavy stuffs are already
- 1:29:06fully paralyzed on the cluster. So the entire process usually
- 1:29:11takes just a couple of hours The last tab is actually older
- 1:29:17labeling, the new trips. So, here we use the same multi
- 1:29:22trip, alignment engine, but only between pre-built reconstruction
- 1:29:26and each new trip. So, it's much much simpler than
- 1:29:30fully reconstructing all the clips or together.
- 1:29:33That's why it only takes 30 minutes per trip to order label,
- 1:29:37instead of manual several hours of manual labeling, And this is
- 1:29:43also the key of scalability of this machine.
- 1:29:49This machine easily scales, as long as we have available
- 1:29:53compute and trip data. So about 50 trips were newly
- 1:29:58Auto labeled from this scene, and some of them are shown here.
- 1:30:01So 53, from different vehicles. So this is how we capture and
- 1:30:08transform the space time, slices of the world into the network
- 1:30:11supervision. Yeah.
- 1:30:13One thing I'd like to note is that again just talked about how
- 1:30:16we oughta label our lanes but we have Auto laborers for almost
- 1:30:20every task that we do including our planner and many of these
- 1:30:23are fully automatic no humans involved, for example, for
- 1:30:25objects or the kinematics their shapes, their Futures,
- 1:30:29everything just comes from Auto labeling and the same is true
- 1:30:32for occupancy to. And we have really just build a
- 1:30:34machine or Round this. Yeah, so if you can go back, one
- 1:30:37slide. One more it says parallelized on
- 1:30:42cluster so so that sounds pretty straightforward but it really
- 1:30:47wasn't. Maybe it's fun to share our
- 1:30:49something like this comes about. So a while ago we didn't have
- 1:30:53any auto labeling at all and then someone makes a script, it
- 1:30:57starts to work, it starts working better, until you reach
- 1:31:00a volume. That's pretty high, and we
- 1:31:01clearly need a solution. And so there were two other
- 1:31:05engineers in our team who are like, you know, that's an
- 1:31:07interesting, you know, What we needed to do, was build a whole
- 1:31:10graph of essentially python functions that we need to run
- 1:31:14one after the other first, you pull the clip.
- 1:31:16They do some cleaning, then you do some Network inference, and
- 1:31:19another Network influence until you finally get this.
- 1:31:22But so you need to do that as a large-scale.
- 1:31:24So I tell them, we probably need to shoot for, you know, 100,000
- 1:31:28Clips per day or like hundred thousand items that seems good.
- 1:31:32And so, the engineer said, well, we can do a bit of postgres and
- 1:31:37bit of elbow grease. We can do.
- 1:31:39It? Meanwhile, we are bit later and
- 1:31:41we're doing 20 million of these functions every single day.
- 1:31:46Again, we pull in around half a million clips and on those we
- 1:31:49run it done a functions each of these in a streaming fashion and
- 1:31:52so that's kind of the back end. Infrared is also needed to not
- 1:31:55just run training, but also auto-leveling.
- 1:31:57Yeah, it really is like a factory that produces labels and
- 1:32:01it's like production lines. Yield quality inventory, like
- 1:32:04all of the same Concepts applied to this label.
- 1:32:07Factory that applies for you know, the factory for cars.
- 1:32:10That's right. Okay.
- 1:32:14Thanks Tim Allen Show. So, yeah, so concluding, this
- 1:32:18section, I'd like to share a few more challenging and interesting
- 1:32:21examples for Network, for sure. And even for humans, probably.
- 1:32:26So, from the top, there's like example, for like lack of Lights
- 1:32:30case or foggy night, or roundabout and occlusions pipe.
- 1:32:34Have your questions by parked cars and even rainy night with
- 1:32:37the raindrops on camera lenses, these are challenging.
- 1:32:41But once their original scenes are Fully reconstructed by other
- 1:32:44Clips. They all of them can be Auto
- 1:32:46labeled so that our cars can drive even better through these
- 1:32:50challenging scenarios. So, now let me pass the mic to
- 1:32:54David to learn more about how cities creating the new world on
- 1:32:56top of these labels. Thank you.
- 1:33:05Thank you again. My name is David and I'm going
- 1:33:07to talk about simulation. So simulation plays a critical
- 1:33:11role in providing data that is difficult to source and or hard
- 1:33:15to label. However 3D scenes are
- 1:33:18notoriously slow to produce. Take for example, the simulated
- 1:33:22seen playing behind me. A complex intersection from
- 1:33:26Market Street in San Francisco. It would take two weeks for
- 1:33:30artists to complete. And for us, that is painfully
- 1:33:33slow. However, I'm going to talk about
- 1:33:36using the agins automated ground truth labels along with some
- 1:33:39brand-new. Schooling that allows us to
- 1:33:41procedurally. Generate this seen in many like
- 1:33:43it and just five minutes. That's an amazing 1,000 times
- 1:33:47faster than before. So let's dive in to I was seeing
- 1:33:51like this is created. We start by piping the automated
- 1:33:55ground truth labels into our simulated World creator.
- 1:33:58Tooling inside the software, Houdini starting with Road
- 1:34:02boundary labels, we can generate a solid Road mesh and reads.
- 1:34:05Apologize it with the landgraf labels, this helps inform
- 1:34:09important Road details, like Crossroads slope and detailed
- 1:34:12material blending. Next, we can use the line data
- 1:34:16and sweep geometry across its surface and project it to the
- 1:34:19road creating Lane paint decals. Next using media and edges we
- 1:34:27can spawned Island geometry and populate it with randomize
- 1:34:30foliage. This drastically changes the
- 1:34:32visibility of the scene. Now the outside world can be
- 1:34:35generated through a series of randomized.
- 1:34:38Heuristics modular building generators, create visual
- 1:34:41obstructions while randomly placed objects like hydrants can
- 1:34:44change the color of the curbs, while trees can treat drop
- 1:34:47leaves below it, obscuring lines, or edges.
- 1:34:51Next, we can bring in map data to inform positions of things
- 1:34:55like traffic traffic lights or stop signs.
- 1:34:57We can trace along. It's normal to collect important
- 1:35:00information like number of lanes and even get accurate street
- 1:35:03names on the signs themselves. Next using land graph.
- 1:35:07We can determine Lane connectivity and spawned
- 1:35:10directional Road markings on the road and their accompanying road
- 1:35:13signs. And finally, with landgraf
- 1:35:16itself we can determine Lane adjacency and other useful
- 1:35:20metrics the spawn randomize traffic permutations inside our
- 1:35:23simulator. And again this is all automatic
- 1:35:26no artists in Loop and happens within minutes.
- 1:35:29And now the sets us up to do some pretty cool things.
- 1:35:34Since everything is based on data and heuristics, we can
- 1:35:36start to fuzz parameters to create visual variations of the
- 1:35:40single ground truth. It can be as subtle as object
- 1:35:43placement and random material swapping to more drastic
- 1:35:46changes. Like entirely new biomes are
- 1:35:48locations of environment like Urban Suburban or rural this
- 1:35:53allows us to create infinite, targeted permutations for
- 1:35:56specific ground truce that we need more ground Truth for and
- 1:36:01all this happens within a click of a button.
- 1:36:05We could even take this one step further by altering our ground
- 1:36:09truth itself, say John wants his Network to pay more attention.
- 1:36:12The directional Road markings to better detect an upcoming
- 1:36:15Capital left turn lane. We can start to procedurally
- 1:36:19alter our landgraf inside the simulator to help folks, to
- 1:36:22create entirely new flows through this intersection to
- 1:36:25help Focus. The Network's attention to the
- 1:36:28road markings to create more accurate predictions.
- 1:36:29And this is a great example of how this tooling allows us to
- 1:36:34create new data that could never be.
- 1:36:35Elected from The Real World. The true power of this tool is
- 1:36:41in its architecture and how we could run all tests in parallel
- 1:36:44to infinitely scale, So you saw the tile Creator tool and action
- 1:36:49converting the ground truth labels into their counterparts.
- 1:36:53Next, we can use our tile extractor tool to divide this
- 1:36:56data into geohash tiles about 150 meters square in size.
- 1:37:01We then save out that data into separate geometry and instance
- 1:37:05files, this gives us a clean source of data, that's easy to
- 1:37:08load and allows us to be rendering engine agnostic for
- 1:37:11the future. Then using a tile loader tool.
- 1:37:16We can summon any number of those cash tiles using a Geo
- 1:37:19hash ID. Currently we're doing about
- 1:37:22these five by five tiles or 3x3 usually centered around Fleet
- 1:37:25hotspots or interesting language off locations in the tile
- 1:37:29loader. Also converts these tile sets
- 1:37:32into you assets, for consumption, by the Unreal
- 1:37:35Engine gives you a finished project product from what you
- 1:37:38saw in the first slide. And this really sets us up for
- 1:37:42size and scale. As you can see, on the map
- 1:37:45behind us, we can easily generate most of San Francisco
- 1:37:49city streets. And this didn't take years or
- 1:37:52even months of work or rather two weeks by one person, and we
- 1:37:56can continue to manage and grow all this data using RPG Network
- 1:38:00inside of the tooling, this allows us to throw compute at it
- 1:38:04and regenerate all these tilesets overnight this ensures
- 1:38:08all environments are of consistent quality and features
- 1:38:12which is super important for training since new ontology Xin
- 1:38:15signals are constantly released And now if to come full circle
- 1:38:23because we generated all these tilesets from ground truth data,
- 1:38:26they contain all the weird intricacies from The Real World.
- 1:38:29We can combine that with the procedural Visual and traffic
- 1:38:32variety to create Limitless, targeted data for the network to
- 1:38:36learn from That concludes the same section.
- 1:38:39I'll pass it to Kate to talk about how we can use all this
- 1:38:42data to improve autopilot. Thank you.
- 1:38:54Thanks David. Hi everyone.
- 1:38:56My name is Kate Park and I'm here to talk about the data
- 1:38:59engine which is the process by which we improve our neural
- 1:39:02networks via data. We're going to show you how we
- 1:39:06deterministically solve interventions via data, and walk
- 1:39:09you through the life of this particular clip.
- 1:39:12In this scenario, autopilot is approaching a turn and in
- 1:39:16correctly, predicts, that Crossing vehicle as stopped for
- 1:39:19traffic and thus, the vehicle that we would slow down.
- 1:39:22For in reality, there's nobody in the car, it's just awkwardly
- 1:39:26parked. We built this, tooling to
- 1:39:28identify the mispredictions correct, the label and
- 1:39:32categorize this clip into an evaluation.
- 1:39:34Set this particular clip happens to be one of 126 that we've
- 1:39:39diagnosed as challenging parked cars at turns because of this
- 1:39:43infra, we can curate this evaluation set without any
- 1:39:47engineering resources custom to this particular challenge case.
- 1:39:52To actually solve that challenge case, requires mining thousands
- 1:39:56of examples like it and it's something Tesla can trivially,
- 1:39:59do we simply use our data sourcing infra request data, and
- 1:40:04use the tooling shown previously to correct.
- 1:40:06The labels by surgically targeting the mispredictions of
- 1:40:10the current model. We're only adding the most
- 1:40:13valuable examples to our training set.
- 1:40:16We surgically fix 13,900 clips and because those were examples
- 1:40:22where the current model struggles, we don't even need to
- 1:40:25change the model architecture, a simple weight update with this.
- 1:40:28New valuable data is enough to solve the challenge case.
- 1:40:32So you see, we no longer predict that Crossing vehicle as stopped
- 1:40:35as shown in Orange, but parked as shown in red, In Academia, we
- 1:40:41often see that people keep data constant but a Tesla, it's very
- 1:40:45much the opposite we see time and time and again, that data is
- 1:40:49one of the best if not the most deterministic lever to solving
- 1:40:52these interventions. We just showed you the data
- 1:40:56engine Loop for one challenge case.
- 1:40:58Namely, these parked cars at turns, but there are many
- 1:41:01challenge cases, even for one signal of vehicle Movement.
- 1:41:05We apply this data engine Loop to every single challenge case
- 1:41:08we've diagnosed, whether it's buses, curvy Road, stopped
- 1:41:11Vehicles, parking, lots, and we don't just add data once we do
- 1:41:16this again, and again to perfect.
- 1:41:17The semantics, in fact, this year, we updated our vehicle
- 1:41:21movement signal 5 times and with Weight update trained on the new
- 1:41:26data. We push our vehicle movement,
- 1:41:28accuracy up and up. This data engine framework
- 1:41:33applies to all our signals, whether they're 3D multicam
- 1:41:36video, whether the data is human labeled, Auto labeled or
- 1:41:40simulated, whether it's an offline model or an online model
- 1:41:43model and Tesla's able to do this at scale because of the
- 1:41:47fleet Advantage, the infra that are Eng team has built, and the
- 1:41:51labeling resources that feed our networks to train on all this
- 1:41:55data. We need a massive amount of
- 1:41:57compute, so I'll hand it off to Pete and Ganesh to talk about,
- 1:42:00Out, the dojo supercomputing platform.
- 1:42:03Thank you. A petty.
- 1:42:13Thanks everybody. Thanks for hanging in there.
- 1:42:14We're almost there. My name is Pete Benton.
- 1:42:18I run the custom silicon and low voltage teams at Tesla.
- 1:42:23And my name is Ganesh blanket, I had on the dodgy program.
- 1:42:32Thank you. I'm frequently asked.
- 1:42:35Why is the car company building a supercomputer for training and
- 1:42:39this question fundamentally misunderstands the nature of
- 1:42:44Tesla at its heart. Tesla is a hard core technology
- 1:42:48company. All across the company people
- 1:42:51are working hard in science and engineering to advance the
- 1:42:55fundamental understanding and methods that we have available
- 1:42:59to build cars, Energy, Solutions, robots, and anything
- 1:43:04else. So can we can do to improve The
- 1:43:06Human Condition around the world.
- 1:43:08It's a super exciting thing to be a part of and it's a
- 1:43:11privilege to run a very small piece of it in the semiconductor
- 1:43:15group. Tonight, we're going to talk a
- 1:43:16little bit about dojo and give you an update on what we've been
- 1:43:20able to do over the last year. But before we do that, I wanted
- 1:43:23to give a little bit of background on the initial design
- 1:43:26that we started a few years ago. When we got started the goal was
- 1:43:29to provide a substantial Improvement to the training
- 1:43:33latency for our autopilot team. Some of the largest neural
- 1:43:36networks, they train today run for over a month which inhibits
- 1:43:40their ability to rapidly explore Alternatives and evaluate them.
- 1:43:45So, you know, a 30 x speed up would be really nice if we could
- 1:43:49provide it at a cost competitive and energy competitive way to do
- 1:43:53that. We wanted to build a chip with a
- 1:43:56lot of arithmetic arithmetic. Units.
- 1:43:59That we could utilize that a very high efficiency.
- 1:44:01And we spent a lot of time studying whether we could do
- 1:44:04that using DRM various packaging ideas, all of which failed and
- 1:44:10in the end, even though it felt like an unnatural act, we
- 1:44:12decided to reject dram as the primary storage medium for this
- 1:44:16system. And instead focus on S Ram
- 1:44:19embedded. In the chip that's ramp
- 1:44:21provides. Unfortunately, a modest amount
- 1:44:24of capacity, but extremely high bandwidth and very low latency
- 1:44:27and that enables us to achieve High utilization with your
- 1:44:30arithmetic units. Those choices.
- 1:44:35That particular choice led to a whole bunch of other choices.
- 1:44:38For example, if you want to have a virtual memory, you need page
- 1:44:41tables. They take up a lot of space.
- 1:44:43We didn't have space, so no virtual memory.
- 1:44:46We also don't have interrupts. The accelerator is a bare-bones.
- 1:44:51Rob piece of Hardware that's presented to a compiler.
- 1:44:54And the compiler is responsible for scheduling everything that
- 1:44:57happens in a deterministic way. So there's no need or even with
- 1:45:00desire for interrupts in the system.
- 1:45:03We also chose to pursue Zu model, parallelism as a training
- 1:45:08methodology, which is not the typical situation most most
- 1:45:13machines today. Use data parallelism, which
- 1:45:15consumes additional memory capacity, which we obviously
- 1:45:18don't have. So all of those choices.
- 1:45:21Let us to build a machine. That is pretty radically
- 1:45:25different from what's available today.
- 1:45:28We also had a whole bunch of other goals, and one of the most
- 1:45:31important ones was no limits. So we wanted to build a compute
- 1:45:34fabric that would scale on in an unbounded way, for the most
- 1:45:37part. I mean, obviously, there's
- 1:45:39physical limits now and then, but, you know, pretty much.
- 1:45:43If your model was too big for the computer, you just had to go
- 1:45:45buy a bigger computer. That's what we were looking for
- 1:45:49today. The way package machines are
- 1:45:51package, there's a pretty fixed ratio of, for example, GPU CPUs
- 1:45:55and D Ram capacity and Work capacity and we really wanted to
- 1:45:59disaggregate all that. So that, as models involved, we
- 1:46:03could bury the ratios of those various elements and make the
- 1:46:07system more flexible to meet the needs of the autopilot team.
- 1:46:12Yeah. And it's so true.
- 1:46:13Be like No Limits philosophy was our guiding star all the way,
- 1:46:18all of our choices were centered around that.
- 1:46:22And, and to the point that we didn't want traditional data
- 1:46:25center infrastructure to limit, With our capacity to execute
- 1:46:30these programs at speed. So that's why we, that's why
- 1:46:35we're sorry about that. That's why we integrated
- 1:46:40vertically. Our data center entire data
- 1:46:44center by doing a vertical. Integration of the data center,
- 1:46:48we could extract new levels of efficiency.
- 1:46:51We could optimize power delivery Cooling and as well as system
- 1:46:56management across. Ross, the whole data center
- 1:46:59stack rather than doing Box by box and integrating that those
- 1:47:04boxes into Data Centers. And to do this, we also wanted
- 1:47:11to integrate early to figure out limits of scale for our software
- 1:47:16workloads. So we integrated Dojo
- 1:47:18environment into our autopilot software very early and we
- 1:47:22learned a lot of lessons. And today, Bill Chang will go
- 1:47:27over our hardware update as well as some of the challenges that
- 1:47:31we faced along the way and regime Korean will give you a
- 1:47:36glimpse. Of our compiler technology, as
- 1:47:39well as go over some of our cool results.
- 1:47:49Thanks Pete. Thanks, Ganesh.
- 1:47:52I'll start tonight with a high level vision of our system.
- 1:47:56That will that will help set the stage for the challenges and the
- 1:48:00problems were solving. And then also, how software will
- 1:48:04then leverage this for performance?
- 1:48:07Now, our vision for Dojo is to build a single unified
- 1:48:10accelerate, a very large, one software would see a seamless
- 1:48:15compute plane with globally, addressable, very fast memory
- 1:48:19and all connected together with uniform, high bandwidth and low
- 1:48:23latency. Now, to realize this we need to
- 1:48:29use density to achieve performance.
- 1:48:32Now, we leverage technology to get this density in order to
- 1:48:35break levels of hierarchy, all the way from the chip to the
- 1:48:39scale-out systems. Now silicon technology has has
- 1:48:44used, this has done this for decades.
- 1:48:46Chips, had followed Moore's law to for density and integration
- 1:48:51to get per performance scaling Now, a key step in realizing
- 1:48:56that Vision was our training tile.
- 1:48:59Not only can we integrate 25 dies at extremely high
- 1:49:03bandwidth, but we can scale that to any number of additional
- 1:49:07tiles by just connecting them together.
- 1:49:11Now last year, we showcased, our first functional training tile
- 1:49:16and at that time we already had workloads running on it.
- 1:49:21And since then, the team here has been working hard and
- 1:49:25diligently to deploy this at scale.
- 1:49:29Now we've made amazing progress and had a lot of Milestones
- 1:49:32along the way. And of course we've had a lot of
- 1:49:35unexpected challenges, but this is where our fail fast
- 1:49:39philosophy. Has allowed us to push our
- 1:49:41boundaries Now pushing density for performance presents, all
- 1:49:48new challenges, one area is power delivery.
- 1:49:53Here, we need to deliver the power to our compute Dy and this
- 1:49:57directly impacts our top-line compute performance.
- 1:50:01But we need to do this at unprecedented density.
- 1:50:04We need to be able to match. Our die pitch with a power
- 1:50:08density of almost 1 amp per mm Square.
- 1:50:12And because of the extreme integration, this needs to be a
- 1:50:15multi-tiered vertical power solution.
- 1:50:19And because there's a complex heterogeneous material stack up,
- 1:50:22we have to carefully manage the material transition, especially
- 1:50:26CTE. Now, why does the coefficient of
- 1:50:31thermal expansion matter in this case?
- 1:50:34CTE is a fundamental material property and if it's not
- 1:50:38carefully managed that Stack Up would literally rip itself
- 1:50:42apart. So, we started this effort by
- 1:50:47working with vendors to deliver to develop this power solution,
- 1:50:52but we realize that we actually had to develop this in house.
- 1:50:56Now, to balance schedule and risk, we built quick iterations
- 1:51:01to support, both our system bring up in software development
- 1:51:05and also to find the optimal design and stack up that would
- 1:51:08meet our final production goals. And in the end, we were able to
- 1:51:12reduce CTE over 50% and meet our performance by 3x over our
- 1:51:19initial version. Now, needless to say, finding
- 1:51:23this optimal material stack up while maximizing performance at
- 1:51:27density is extremely difficult. Now we did have unexpected
- 1:51:34challenges along the way, here is an example where we push the
- 1:51:38boundaries of integration that led to component failures.
- 1:51:43This started, when we scaled up to larger and longer workloads,
- 1:51:47and then intermittent intermittently a single site on
- 1:51:50a tile would fail. Now, they started out as
- 1:51:54recoverable failures, but as we pushed some much higher and
- 1:51:57higher power, these would become permanent values.
- 1:52:03Now, to understand this failure, you have to understand why and
- 1:52:07how we build our power modules solving density at every level
- 1:52:13is that is is the Cornerstone of actually achieving our system
- 1:52:16performance now, because our XY plane is used for high bandwidth
- 1:52:21communication, everything else must be stacked vertically.
- 1:52:26This means all other components other than our dye must be
- 1:52:30integrated into our power modules.
- 1:52:32Now that includes our clock and our power supplies and also our
- 1:52:36system controllers, Now, in this case, the Pharaohs were due to
- 1:52:41losing clock output from our oscillators.
- 1:52:45And after an extensive debug, we found that the root cause was
- 1:52:49due to vibrations on the module from piezoelectric effects.
- 1:52:53Are nearby capacitors. Now, singing caps are not a new
- 1:52:59phenomenon and in fact, very common in power design, but
- 1:53:03normally clock chips are placed in a very quiet area of the
- 1:53:06board and often not affected by power circuits.
- 1:53:10But because we needed to achieve this level of integration, these
- 1:53:14oscillators need to be placed in very close proximity.
- 1:53:18Now, due to our switching frequency and then the vibration
- 1:53:21resonance created, it caused out of plane vibration on our mems
- 1:53:26oscillator that caused it to crack.
- 1:53:31Now, the solution to this problem is a multi-prong
- 1:53:33approach. We can reduce the vibration by
- 1:53:36using soft terminal caps We can update our mems part with a
- 1:53:43lower Q factor for the out of plane Direction.
- 1:53:48And we can also update our switching frequency frequency to
- 1:53:52push the resonance further away from these sensitive bands.
- 1:53:58Now, addition to the density at the system level, we've been
- 1:54:03making a lot of progress at the infrastructure level.
- 1:54:07We knew that we had to read examine every aspect of the data
- 1:54:11center infrastructure in order to support our unprecedented
- 1:54:15power and cooling density. We brought in a fully custom
- 1:54:19design CD you to support dojos dense, cooling requirements.
- 1:54:24And the amazing part is we're able to do this at a fraction of
- 1:54:26the cost versus buying off the shelf and modifying it And since
- 1:54:32our Dojo cabinet, integrates enough, power and cooling to
- 1:54:35match an entire row of standard, it racks, we need to carefully
- 1:54:40design, our cabinet in an infrastructure together.
- 1:54:44And we've already gone through several iterations of this
- 1:54:47cabinet to optimize this. And earlier this year, we
- 1:54:51started load testing our power and cooling infrastructure.
- 1:54:54And we were able to push it over to megawatts before we tripped
- 1:54:57our substation, and got a call from the city.
- 1:55:04Now, last year we introduced only a couple of components of
- 1:55:07our system, the custom D1 die and the training tile, but we
- 1:55:11tease the exit pod. As our end goal will walk
- 1:55:15through the remaining parts of our system that are required to
- 1:55:18build out this exit pod. Now, the system tray is a key
- 1:55:23part of realizing our vision of a single accelerator, it enables
- 1:55:28us to seamless seamlessly connect tiles together.
- 1:55:31Not only within the cabinet, but between cabinets, we can connect
- 1:55:36these Tiles at very tight spacing across the entire
- 1:55:40accelerator. And this is how we achieve our
- 1:55:42uniform communication. This is a laminated busbar that
- 1:55:47allows us to integrate very high power mechanical, and thermal
- 1:55:51support and an extremely dense integration.
- 1:55:54It's 75. Mm in height and and support six
- 1:55:58Tiles. At 135 kg.
- 1:56:01This is the equivalent of three to four fully loaded, high
- 1:56:05performance racks. Next, we need to feed data to
- 1:56:11the training tiles. This is where we've developed
- 1:56:14the dojo interface processor. It provides our system with high
- 1:56:19bandwidth D Ram to Stage our training data and it provides
- 1:56:23full memory bandwidth to our training tiles.
- 1:56:26Using TTP, our custom protocol that we can use to communicate
- 1:56:30across our entire accelerator It also has high speed, ethernet.
- 1:56:35That helps us extend this custom protocol over standard ethernet
- 1:56:39and we provide native Hardware support for this with little to
- 1:56:43no software overhead. And lastly, we can connect
- 1:56:47connect to it through a standard Gen4 pcie interface. now we pair
- 1:56:5420 of these cards portray, and that gives us 640 gigabytes of
- 1:56:59high-bandwidth DRM And this provides our disaggregated
- 1:57:03memory layer for our training tiles.
- 1:57:06These cards are our high bandwidth in just path.
- 1:57:09Both through pcie and ethernet. They also provide a high rate x
- 1:57:14z connectivity path. That allows shortcuts across our
- 1:57:18large Dojo accelerator. Now we actually integrate the
- 1:57:24host directly underneath our system tray.
- 1:57:27These hosts provide our in just processing and connect to our
- 1:57:31interface processors through pcie.
- 1:57:34These host can provide Hardware video decoder support for
- 1:57:39video-based training. And our user applications land
- 1:57:43on these host that we so we can provide them with a standard x86
- 1:57:48Linux environment. Now, we can put two of these
- 1:57:54assemblies into one cabinet and pair it with redundant.
- 1:57:58Power supplies that do direct conversion of three-phase 480.
- 1:58:02Volt AC power 252 volt, DC power Now, by focusing on density, at
- 1:58:12every level, we can realize the vision of a single accelerator.
- 1:58:18Now, starting with the uniform nodes, on our custom D1 die, we
- 1:58:22can connect them together and are fully integrated training
- 1:58:25tile. And then finally seamlessly
- 1:58:29connecting them across cabinet boundaries to form our Dojo
- 1:58:33accelerator. And all together, we can house
- 1:58:37to full accelerators in our exit pod for a combined.
- 1:58:41One exaflop of ml compute. Now, all can all together, this
- 1:58:45amount of technology and integration has only ever been
- 1:58:49done a couple of times in the history of compute.
- 1:58:53Next, we'll see how software can leverage this to accelerate
- 1:58:56their performance. thanks Bill, my name is Rajiv and I'm going
- 1:59:08to talk some numbers SAR software stack begins with the
- 1:59:12pie torch extension that speaks to our commitment to one
- 1:59:15standard python Which models out of the box.
- 1:59:19We're going to talk more about our jit compiler and be in just
- 1:59:22pipeline that feeds the hardware with data.
- 1:59:25Abstractly performances tops times, utilization times
- 1:59:29accelerator occupancy. We've seen how the hardware
- 1:59:32provides Peak Performance, is the job of the compiler to
- 1:59:35extract utilization from the hardware while code is running
- 1:59:38on it. And it's the job of the interest
- 1:59:41pipeline to make sure that data can be fed at Ruppert high
- 1:59:44enough for the hardware to not ever starve.
- 1:59:48So let's talk about why communication bound models are
- 1:59:50difficult to scale. But before that, let's look at
- 1:59:53why resident 50 like models are easier to skill.
- 1:59:56You start off with a single accelerator, one of the forward
- 1:59:58and backward, pass has followed by the optimizer.
- 2:00:02Then to scale this up, you run multiple copies of this on
- 2:00:05multiple accelerators and while the gratings produced by the
- 2:00:08backward pass, do need to be reduced.
- 2:00:09And this introduces, some communication, this can be done
- 2:00:12Pipeline with the backward pass. This set of scales fairly well,
- 2:00:19almost linearly. For models with much larger
- 2:00:24activations, we run into a problem as soon as we want to
- 2:00:27run the forward. Pass the batch size that fits in
- 2:00:30a single accelerator is often smaller than the batch Norm
- 2:00:32surface. So, to get around this
- 2:00:34researchers, typically one the setup on multiple accelerators,
- 2:00:38in sync, bash, normal mode, this introduces latency bound,
- 2:00:41communication to the critical path of the forward pass, and we
- 2:00:44already have a communication bottleneck And while there are
- 2:00:48ways to get around this, they usually involve tedious.
- 2:00:50Manual work, best suited for a compiler and ultimately there's
- 2:00:55no skirting around. The fact that if your state does
- 2:00:57not fit in a single accelerator, you can be communication about
- 2:01:03and even with significant efforts from our ml Engineers,
- 2:01:06we see such models don't scale linearly the dojo system was
- 2:01:11built to make such models work at how utilization the high
- 2:01:15density and equation is was built to not only accelerate the
- 2:01:18compute bound portions of a model but also the latency bound
- 2:01:22portions, like a batch Norm or the bandwidth bound portions,
- 2:01:26like a gradient, all reduced or a parameter all gather.
- 2:01:31A slice of the dojo mesh can be carved out to run any model.
- 2:01:35The only thing you just need to do is to make the slice large
- 2:01:38enough to fit a bathroom surface for their particular model.
- 2:01:43After that, the partition presents itself as one large
- 2:01:46accelerator, being the users from having to worry about the
- 2:01:49internal details of execution. And as the job of the compiler
- 2:01:54to maintain this abstraction fine-grain synchronization
- 2:01:58Primitives in uniform. Low latency makes it easy to
- 2:02:01accelerate, all forms of parallelism across integration
- 2:02:04boundaries, tensors are usually stored, Chardon as RAM and
- 2:02:07replicated. Just in time for layers
- 2:02:10execution, we depend on the high Dojo being with to hide this
- 2:02:14replication time. Pensive replication and other
- 2:02:17day of transfers are overlapped with compute and the compiler
- 2:02:20can also recompute layers. One is profitable to do so, We
- 2:02:25expect most models to work out of the box is an example.
- 2:02:29We took the recently released stable diffusion model and got
- 2:02:32it running on dojo and minutes out of the box.
- 2:02:35The compiler was able to map it in a model.
- 2:02:37Pelham Manor on 25, Dojo dies. Here's some pictures of a cyber
- 2:02:42truck on Mars generated by stable, diffusion running on
- 2:02:45dojo looks Looks like it. Still has some ways to go before
- 2:02:56matching the Tesla Design Studio team.
- 2:03:00So we've talked about how communication bottlenecks can
- 2:03:02hamper scalability. Perhaps an acid test of a
- 2:03:05compiler, and the underlying Hardware is executing across
- 2:03:09diet Bachelor. Like mentioned before, this can
- 2:03:12be a Serial bottleneck. The communication phase of a
- 2:03:15bachelor begins with nodes Computing.
- 2:03:16Their local mean and standard deviations then coordinating to
- 2:03:20reduce these values, then broadcasting these values back
- 2:03:23and then they resume their work in parallel.
- 2:03:27So what would an ideal bathroom look like on 25 Dojo dies?
- 2:03:31Let's say the previous less activations are already split
- 2:03:34across the ice. We would expect the 350 nodes on
- 2:03:39each die to coordinate and produce die.
- 2:03:41Local mean and standard deviation values.
- 2:03:44Ideally, these would get rid further videos with the final
- 2:03:46value, ending somewhere in towards the middle of the tile.
- 2:03:51We would then hope to see a broadcast of this value.
- 2:03:53Radiating from the center. Let's see how the compiler,
- 2:03:57actually executes. A real bathroom operation across
- 2:04:0025. Does the communication trees
- 2:04:02were extracted from the compiler and the timing is from a real
- 2:04:06Hardware one. We're about to see eight
- 2:04:08thousand seven fifty nodes, on 25 dies, coordinating to reduce
- 2:04:12and then broadcast the bass drum mean and standard deviation
- 2:04:15values. Die, local reduction followed by
- 2:04:20global reduction towards the middle of the tie.
- 2:04:23Then the reduced value broadcast radiating from the middle
- 2:04:27accelerated by the Harbor's broadcast facility.
- 2:04:33This operation takes only five microseconds on 25 Dojo dice.
- 2:04:38The same operation takes 150 microseconds on 24 gpus.
- 2:04:42This is an orders of magnitude improvement over gpus and while
- 2:04:47we talked about and already saw operation in the context of a
- 2:04:49batch Norm, it's important to reiterate that the same
- 2:04:52advantages apply to all other communication Primitives and
- 2:04:56these Primitives are essential for large-scale training.
- 2:05:01So how about full model performance?
- 2:05:03So while we think that resonate, 50 is not a good representation
- 2:05:06of real world, Tesla workloads. It is a standard Benchmark so
- 2:05:10let's start there. We are already able to match
- 2:05:13the, a 100 die for dog. However, perhaps a hint of dojos
- 2:05:17capabilities is that were able to hit this number with just a
- 2:05:20batch of 8, / die, but Dojo was really built.
- 2:05:23A tackle larger complex models. So, when we set out to tackle
- 2:05:28real-world workloads, we looked at the usage patterns of her
- 2:05:31current GPU cluster. And to model, stood up the auto
- 2:05:34labeling networks across of offline models, that are used to
- 2:05:37generate ground truth, and the occupancy networks that you
- 2:05:40heard about. The are living that perks are
- 2:05:43large models that have higher asthmatic intensity.
- 2:05:46All the occupancy networks can be in just bound, we chose these
- 2:05:50models because together, they account for a large chunk of our
- 2:05:53current GPU cluster usage and they would challenge the system
- 2:05:56in different ways. So how do we do on these two
- 2:06:01networks? The results were about to see
- 2:06:03where measured on Multi die systems for both the GPU and
- 2:06:07Dojo but normalized important numbers.
- 2:06:10On our Auto labeling Network were already able to surpass the
- 2:06:14performance of an a 100 with our current Hardware running, on our
- 2:06:17older generation, prm's on our production Hardware, with our
- 2:06:21nerve, your arms. That translates to doubling the
- 2:06:23throughput of an a 100. And our model showed that with
- 2:06:27some key compiler optimizations we could get two more than 3x.
- 2:06:31The performance have been a 100, we see even bigger leaps on the
- 2:06:35occupancy Network. Almost 3x with our production
- 2:06:39Hardware, with room for more. So, what does that mean for
- 2:06:52Tesla with a current level of compiler performance, we could
- 2:06:55replace the ml computer of one, two, three, four, five and six
- 2:07:01GPU boxes with just a single Dojo tile.
- 2:07:13And this Dojo tile cost less than one of these GPU boxes.
- 2:07:20It is what it really means is that networks that took more
- 2:07:25than a month to train. Now take a less than a week.
- 2:07:31Last one, we measure things, it did not turn out so well at the
- 2:07:35pythons level we did not see our expected performance out of the
- 2:07:38gate and this timeline chart shows, our problem, the teeny
- 2:07:42tiny little green bars. That's the compiled code running
- 2:07:45on the accelerator. The row is mostly white space
- 2:07:48where the hardware is just waiting for data.
- 2:07:54With our dense. I'm old compute Dojo host
- 2:07:56effectively have 10x more ammo compute than the GPU host.
- 2:08:00The data loader is running on this one.
- 2:08:01Host simply couldn't keep up with all that ml hard work.
- 2:08:06So, to solve our data loader scalability issues, we knew we
- 2:08:09had to get over the limit of this single host.
- 2:08:12The Tesla transport protocol moves data seamlessly across,
- 2:08:15host tiles and ingest processors.
- 2:08:18So we extended the Tesla transport protocol to work over
- 2:08:21ethernet. We didn't build the dojo network
- 2:08:23interface card, that zdenek to leverage.
- 2:08:26TTP, over ethernet. This allows any host with a
- 2:08:29diner card, to be able to dma to.
- 2:08:31And from other TTP on points, So we started with the dojo mesh,
- 2:08:37then we added a tier of data loading host equipped with the
- 2:08:40Dina card. We connected These Hoes to the
- 2:08:45mesh via an ethernet switch every host in this data.
- 2:08:48Loading tier is capable of reaching all ttpm points in the
- 2:08:51dojo mesh via hardware-accelerated dma.
- 2:08:57After these optimizations went in, our occupancy run from 4% to
- 2:09:0297 percent. So the data loading sections
- 2:09:05have reduced They'll the data loading sections have reduced
- 2:09:13drastically. And the ml Hardware is kept
- 2:09:15busy. We actually expect this number
- 2:09:17to go to 100% pretty soon. After these changes went in, we
- 2:09:21saw the full expected speed up from the Peter Schuler and we
- 2:09:24were back in business. So we started with Hardware
- 2:09:29design that breaks through traditional integration
- 2:09:31boundaries in service of our vision of a single giant
- 2:09:34accelerator. We've seen how the compiler and
- 2:09:36in just layers build on top of that Hardware.
- 2:09:40So, after proving our performance on these complex,
- 2:09:42real world networks, we knew what our first large-scale
- 2:09:45deployment would Target our high rithmetic intensity
- 2:09:48auto-leveling Networks. Today, that occupies 4000 gpus
- 2:09:53over 72 GPU Rex with our dense computer in a high-performance.
- 2:09:58We expect to provide the same throughput with just for Dojo
- 2:10:02cabinets. And these for Dojo cabinets will
- 2:10:13be part of our first exit, but that we plan to build by quarter
- 2:10:16one of 2023 This will more than double Tesla's out of labeling
- 2:10:21capacity. The first exit part is part of a
- 2:10:30total of seven exit parts that we plan to build in Palo Alto
- 2:10:34right here across the wall. And we have a display cabinet
- 2:10:40from one of these extra parts for everyone to look at.
- 2:10:45Six tiles. Densely, packed on a tray. 54,
- 2:10:49petaflops of compute 640, GB of high, bandwidth memory, with
- 2:10:53power, and host defeated A lot of confusion.
- 2:11:02And we're building out new versions of all our cluster
- 2:11:05components, and constantly improving our software to hit
- 2:11:08new limits of skill. We believe that we can get
- 2:11:11another 10x improvement with our next Generation hard work.
- 2:11:16And to realize our ambitious goals, we need the best software
- 2:11:19and Hardware Engineers. So please come talk to us or
- 2:11:22visit Tesla.com a. I thank you.
- 2:11:39Let me know. All right, so we hopefully that
- 2:11:48was enough detail and now we can move to questions and guys like
- 2:11:58I think I got the team came out kind of came out on stage and we
- 2:12:04really wanted to show the depth and breadth of Tesla and
- 2:12:10artificial intelligence. The computer hardware robotics
- 2:12:14actuators and and try to really shift the perception of the
- 2:12:19company away from, you know, a lot of people think we're like,
- 2:12:24just a car company or we make cool cars, whatever, but they
- 2:12:28don't have most people have no idea that Tesla is arguably the
- 2:12:33leader in real-world, AI hardware and software and that
- 2:12:38we're building. What is arguably the first some
- 2:12:44of the most radical computer architecture since the cray-1
- 2:12:49supercomputer? And I think if you're interested
- 2:12:51in developing some some of the most advanced technology in the
- 2:12:55world that's going to really affect the world in a positive
- 2:12:58way tells us the place to be. So yeah, let's fire away with
- 2:13:03some questions. I think there's a mic at the
- 2:13:08front and a mic at the back. Or just throw Mike's of people
- 2:13:18jump off of the mic. Yeah.
- 2:13:21Hi, thank you very much. I was impressed here.
- 2:13:26Yeah, I was impressed very much by Optimist, but I wonder why
- 2:13:31they don't driven the hunt. Why did we choose attend?
- 2:13:34Do driven approach for the country because tendons are not
- 2:13:37very durable and why spring-loaded?
- 2:13:44This is pretty cool. Awesome.
- 2:13:46Yes, that's a great question. You know when it comes to any
- 2:13:49type of actuation scheme, there's trade-offs between, you
- 2:13:52know, whether or not it's a tender and system or some type
- 2:13:54of linkage based system. My close to your mouth.
- 2:13:57Let me closer. Jeremy, cool.
- 2:14:01So yeah, the main reason why we went for a tendon based system
- 2:14:05is that, you know, first we actually investigated some
- 2:14:07synthetic tendons but we found that metallic boating cables.
- 2:14:11Are, you know, a lot stronger. One of the advantages of these
- 2:14:15cables is that it's very good for part reduction.
- 2:14:19We do want to make a lot of these hands.
- 2:14:20So having a bunch of parts, a bunch of small linkages ends up
- 2:14:24being, you know, a problem when you're making a lot of
- 2:14:26something, one of the big reasons that, you know, Tendons
- 2:14:31are better than linkages in a sense is that you can be anti
- 2:14:34backlash. So anti backlash essentially,
- 2:14:37you know, allows you to not have any gaps or you know stuttering
- 2:14:41Motion in your fingers spring-loaded.
- 2:14:44Mainly what spring-loaded allows us to do is allows us to have
- 2:14:49active opening. So instead of having to have two
- 2:14:52actuators to drive the fingers closed and then open, we have
- 2:14:56the ability to, you know, have the tendon drive them close and
- 2:14:59then the springs passively. And this is something that's
- 2:15:02seen in our hands as well, right?
- 2:15:04We have the ability to actively flex and then we also have the
- 2:15:07ability to extend. Yeah, our goal with Optimus is
- 2:15:12to have a robot that is maximally useful as quickly as
- 2:15:16possible. So there's a lot of ways to
- 2:15:18solve the various problems of a humanoid robot and we're
- 2:15:23probably not barking up the right Tree on all the Technical
- 2:15:27Solutions. And I should say that we're
- 2:15:29we're open to evolving the Solutions that you see here,
- 2:15:32over time, we're not, they're not locked in stone but we do we
- 2:15:36have to pick something in and we want to pick it.
- 2:15:39Pick something that's going to allow us to produce the robot as
- 2:15:43quickly as possible. And have it likes it be useful
- 2:15:46as quickly as possible. We're trying to follow the goal
- 2:15:49of fastest path, to a useful robot that can be made at
- 2:15:53volume. And we're going to test the
- 2:15:55robot internally at Tesla in our Factory and to see he liked how
- 2:16:01useful is it? Because you have to have a
- 2:16:04you're going to close the loop on reality to confirm that the
- 2:16:06robot is, in fact, useful and yeah.
- 2:16:12So we're going to use it to build things and we're confident
- 2:16:17we can do that with the hand that we have currently designed.
- 2:16:19But there's I'm sure they'll be had version to version 3 and we
- 2:16:22may change the architecture of quite significantly over time.
- 2:16:30Hi. You're The Optimist.
- 2:16:33Robot is really impressive. That you did a great job bipedal
- 2:16:37robots are really difficult. But what I notice might be
- 2:16:41missing from your plan is to acknowledge the utility of the
- 2:16:47human spirit. And I'm wondering if Optimus
- 2:16:50will ever get a personality and be able to laugh at our jokes
- 2:16:53while they've well folds are closed. yeah, absolutely, I
- 2:16:58think we want to have Really fun versions of optimists.
- 2:17:04And so that Optimus can both do be your utilitarian and do
- 2:17:09tasks, but can also be kind of like a friend and a buddy and,
- 2:17:14and hang out with you. And I'm sure people will think
- 2:17:18of all sorts of creative uses for this robot and you know, the
- 2:17:25thing once you have the core intelligence and actuators
- 2:17:29figure it out, then you can Actually, you know put all sorts
- 2:17:34of costumes, I guess, unless on the robot, I mean you can make
- 2:17:39the Robert look, you can scan the robot in many different ways
- 2:17:47and I'm sure people will find very interesting ways to to,
- 2:17:52yeah, versions of optimist. Thanks for the great
- 2:17:59presentation. I wanted to know if there was an
- 2:18:01equivalent to interventions in Optimist, it seems like labeling
- 2:18:06through moments where humans disagree with what's going on is
- 2:18:08important. And in a humanoid robot that
- 2:18:12might be also desirable source of information.
- 2:18:19That's what I say. Yeah, I think we will have ways
- 2:18:26to remote operate. The robot and intervene.
- 2:18:28When it does something bad, especially when we are training
- 2:18:31the robot and bringing it up and hopefully we, you know, design
- 2:18:35it in a way that we can stop the robot from, if it's going to hit
- 2:18:38something, we can just like, hold it, and then we'll stop it.
- 2:18:40Won't like, you know, crush your hand or something and those are
- 2:18:43all intervention data. Yeah.
- 2:18:46And we can do not form of simulation systems to where we
- 2:18:48can check for collisions and supervise that there's a bad
- 2:18:51actions. Yeah, so after Miss we want
- 2:18:55overtime to for to be, you know, an Android kind of Android that
- 2:18:59you see in in in Sci-Fi movies like Star Trek the Next
- 2:19:03Generation like data, but obviously we could program the
- 2:19:06robot to be less robot like and more friendly and and you know,
- 2:19:11can obviously learn to emulate humans and feel very natural.
- 2:19:15All. So as AI in general improves, we
- 2:19:19can add that to the robot and it should be obviously able to do
- 2:19:26simple instructions or even into it what it is that you want.
- 2:19:31So you could give it a high level instruction and then it
- 2:19:34can break that down into a series of actions and and take
- 2:19:37those actions. Hi yeah, it's exciting to think
- 2:19:45that with The Optimist you will think that you can achieve
- 2:19:49orders of magnitude of improvement in economic output,
- 2:19:54that's really exciting. And when Tesla started the
- 2:19:57mission was to accelerate the Advent of renewable energy or
- 2:20:01sustainable transport. So with The Optimist, do you
- 2:20:05still see that mission being thus mission statement of Tesla?
- 2:20:09Or is it going to be updated with you know me In to
- 2:20:13accelerate the Advent of no infinite, abundance or it,
- 2:20:18Limitless Limitless, economy. Yeah, it mean it is not strictly
- 2:20:23speaking. Optimus is not strictly
- 2:20:25speaking. Directly in line with
- 2:20:31accelerating sustainable energy it you know to agree that it is
- 2:20:36more efficient at getting things done than a person.
- 2:20:39It is I guess help with if you know, sustainable energy.
- 2:20:43But I think the mission effectively does is somewhat
- 2:20:46broaden with the Advent of Optimus to, you know, I don't
- 2:20:51know making the future awesome. So you know I think you look at
- 2:20:55Optimist and I know about you but I I'm excited.
- 2:20:58To see what Optimus will become. And, you know, this is like, you
- 2:21:04know, if you could I mean you can tell like any given
- 2:21:07technology. Are you do you want to see what
- 2:21:11it's like in a year? Two years three years, four
- 2:21:14years, five years, ten. I'd say for sure.
- 2:21:17You definitely want to see what's happening with Optimus
- 2:21:20whereas you know a bunch of other Technologies or you know
- 2:21:23sort of plateaued but name names here.
- 2:21:28But you know so I think after this is gonna be incredible in
- 2:21:39like 5 years, 10 years, like mind-blowing.
- 2:21:41And I'm really interested to see that happen.
- 2:21:43I hope you are too. Thank you.
- 2:21:48I have a quick question here. Justin.
- 2:21:51And I was wondering, like, are you planning to extend like
- 2:21:55conversational capabilities for the robot and my second
- 2:21:59follow-up question to that is, what's like the end goal?
- 2:22:03What's the end goal with optimist?
- 2:22:07Yeah, Optimus were definitely of conversational capabilities.
- 2:22:10So You're be able to talk to it and have a conversation and it
- 2:22:16would feel quite natural. So from a technical standpoint
- 2:22:21I'm I do I think it's a keep evolving and I'm not sure where
- 2:22:29it ends up, but someplace interesting, for sure.
- 2:22:34You know, we always have to be careful about the, you know,
- 2:22:36don't go down the Terminator path.
- 2:22:39That's a I thought for my maybe we should start off with a video
- 2:22:43of like the Terminator. Starting off with the, you know,
- 2:22:46skull crushing, but that might be enough.
- 2:22:48You want to get too seriously? So, you know, we we do want
- 2:22:53Optimus to be safe. So we are designing in
- 2:22:57safeguards where you can locally, stop the robot and You
- 2:23:05know, with like basically a localized control ROM that you
- 2:23:08can't update over the Internet, which I think that's quite
- 2:23:11important essential frankly. So like a localized, stop
- 2:23:19button, remote, remote, control, something like that.
- 2:23:24That cannot be changed. But it's definitely going to be
- 2:23:32interesting. It won't be boring.
- 2:23:40Okay, yeah, I see you today. You have a very attractive
- 2:23:43product with told you and it's applications.
- 2:23:46So I'm wondering what's the future for Tokyo platform.
- 2:23:48We would like to provide to like a infrastructure infrastructure
- 2:23:52as a service, like a wa sou be like a sealed.
- 2:23:55Achieve like the Nvidia. So, basically what the future
- 2:23:58because of that, I say you use 70 meters, which is the
- 2:24:01development cost like a easily over 10 million u.s. dollars.
- 2:24:05How do you make the penis? Is like a penis wise?
- 2:24:09Yeah, I mean Dojo is a very big computer and actually will be
- 2:24:16use a lot of power and needs a lot of cooling so I think it's
- 2:24:19probably going to make more sense to have Dojo operate in
- 2:24:22like Amazon web services manner then to try to sell it to
- 2:24:26someone else. So the the most that would be
- 2:24:30the most efficient way to operate Dojo is just have it be
- 2:24:34a service that you can use that's available online.
- 2:24:39And that where you can train your models way faster and for
- 2:24:43less money. And as the world transitions to
- 2:24:48software 2.0, and that's on the bingo card.
- 2:24:55And someone I know, it has to know how to drink five, two
- 2:24:57killers. So, let's see.
- 2:25:03Software 2.0 will use a lot of neural, net training.
- 2:25:11So it kind of makes sense that over time.
- 2:25:16As there's more more neural, net stuff, people will want to use
- 2:25:21and the fastest lowest cost neural network training system.
- 2:25:25So I think is a lot of opportunity in that direction.
- 2:25:32Hi. My name is Alicia honey on.
- 2:25:35Thank you for this event is very inspirational.
- 2:25:39My question is I'm wondering what is your vision for Humanity
- 2:25:46robots? That understand our emotions and
- 2:25:51art and can contribute to our creativity?
- 2:25:58Well, I think there's this you're already seeing robots
- 2:26:01that At least. Are able to generate very
- 2:26:05interesting art with like, like Dolly and Delhi to and think
- 2:26:12we'll start seeing AI that can actually generate even movies
- 2:26:17that have a that have coherence like interesting movies and tell
- 2:26:20jokes. So it's quite remarkable how
- 2:26:23fast AI is advancing. At many companies besides Tesla.
- 2:26:32we're headed for a very interesting future and yeah, so
- 2:26:38You guys want to come out on that?
- 2:26:39Yeah, I guess The Optimist reward can come up with physical
- 2:26:43art, not just digital art, you can, you know, you can ask for
- 2:26:47some dance moves in text or voice and then you can produce
- 2:26:50those in the future. So, it's not affect physical
- 2:26:52heart, not just digital art. Oh yeah.
- 2:26:56Yeah. Computers can absolutely make
- 2:26:58physical art. Yeah.
- 2:26:59Arm sent an advance shortly. Soccer or whatever.
- 2:27:03It needs to get more agile but or time for sure.
- 2:27:09Thanks so much for the presentation for the Tesla
- 2:27:12autopilot. Slides I noticed that the models
- 2:27:15that you are using or heavily motivated by language models and
- 2:27:19I was wondering what the history of that was and how much of an
- 2:27:22improvement it gave. I thought that that was a really
- 2:27:24interesting. Curious choice to use language
- 2:27:26models for the lane transitioning.
- 2:27:29So there are sort of two aspects for why we transition to
- 2:27:31language modeling. So do a talk talk loud and
- 2:27:34close. Okay, okay, got it.
- 2:27:39Yeah, so the language models help us in two ways.
- 2:27:42The first way is that it lets us predict lanes that we couldn't
- 2:27:44have. Otherwise as I shook mentioned
- 2:27:46earlier basically when we predicted lanes and sort of a
- 2:27:49dense 3D fashion, you can only model certain kinds of lanes but
- 2:27:53we want to get those criss-crossing connections
- 2:27:55inside of intersections. It's just not possible to do
- 2:27:57that without making it a graph prediction.
- 2:27:59If you try to do this with dense segmentation is just doesn't
- 2:28:01work. Also the link prediction is a
- 2:28:05multi-modal problem, sometimes you just don't have sufficient
- 2:28:08visual information. To know precisely how things
- 2:28:10look, on the other side of the intersection.
- 2:28:12So you need a method that can generalize and produce coherent
- 2:28:16predictions. You don't want to be predicting
- 2:28:18two lanes in three lanes. At the same time, you want to
- 2:28:20commit to one in a generative model like these language models
- 2:28:23provides that All right. Oh hi.
- 2:28:30My name is Giovanni. Yeah.
- 2:28:33Thanks for the presentation. That's really nice.
- 2:28:37I have a question for FSD team. So for the neural networks, how
- 2:28:44do you test a hottie to unit test software units?
- 2:28:47That sounded like? Do you have like a punch or I
- 2:28:51don't know, maybe thousands or? Yes, cases.
- 2:28:56Where The neural network that after you train it, you have to
- 2:29:00pass it before you release it to as a product or it.
- 2:29:04Yeah, what's your software unit testing strategies for this?
- 2:29:08Yeah, glad you asked that's like a series of tests that we have
- 2:29:11defined starting from, you know, unit test for the software
- 2:29:14itself. But then for the neural network
- 2:29:15models, we have VIP sets defined where, you know, you can Define
- 2:29:20if you just have a large test set.
- 2:29:22That's not enough. What we find.
- 2:29:23We need like sophisticated VIP sets for different.
- 2:29:27Aw, and then we Q8 them and grow them over the time of the
- 2:29:30product. So, over the years, we have
- 2:29:33hundreds of thousands of examples where we have been
- 2:29:36failing in the past that we have curated.
- 2:29:38And so we for any new model we test against the enter history
- 2:29:42of these barriers and then keep adding to this test set on top
- 2:29:46of this. We have Shadow modes where we
- 2:29:48ship these models in silent to the car and we get data back on
- 2:29:51where they're failing or succeeding and there's extensive
- 2:29:55QA program. It's very hard to ship a
- 2:29:58regression. There's like nine levels of
- 2:30:00filters before it hits customers, but then we have
- 2:30:03really good infra to make this all efficient.
- 2:30:07I'm one of the curators, so I QA the car.
- 2:30:10Yeah, like, create a stir. Yeah, so I'm constantly in the
- 2:30:15car just being queuing, like, whatever.
- 2:30:17The latest alpha build is that doesn't totally crash finds a
- 2:30:22lot of bugs. Hi, great event.
- 2:30:27I have a question about foundational models for
- 2:30:31autonomous driving. We have all seen that big models
- 2:30:35that really can when your skill up with data and model
- 2:30:38parameter, right? From GT3 to Palm it, can
- 2:30:42actually. Now, do reasoning, do you see
- 2:30:44that is essential Skilling up foundational models with data
- 2:30:49and size? And then, at least you can get a
- 2:30:52teacher model, right? That potentially can All the
- 2:30:56problems and then you this tale to a student model is that how
- 2:31:00you see foundational models? Relevant 407.
- 2:31:04That's quite similar to our Auto labeling model so we don't just
- 2:31:07have models that run in the car. We train models that are
- 2:31:10entirely offline. There are extremely large that
- 2:31:13can't run in real time on the car.
- 2:31:15So we just run those offline on the server's.
- 2:31:17Producing. Really good labels that can then
- 2:31:21train the online networks. So that's one form distillation
- 2:31:24of these. Teacher student models, kind of
- 2:31:28the foundation models. We are building some really,
- 2:31:30really large data sets that you know, are multiple petabytes.
- 2:31:34And we are seeing that some of these tasks work really well
- 2:31:37when we have this large data sets like a Mattox.
- 2:31:39Like I mentioned we do in all the kinematics out of all the
- 2:31:42objects and up to the fourth derivative and people thought we
- 2:31:46couldn't do detection, IT cameras, detection depth,
- 2:31:48velocity acceleration. And imagine how precise this
- 2:31:52have to be for the signal higher order, derivatives to be
- 2:31:54accurate. And this all comes from these
- 2:31:57kind of large data sets and large models. so we're seeing
- 2:32:00the equivalent of foundation models in our own way for
- 2:32:03geometry and kinematics, and things like those You want to
- 2:32:08add anything job? Yeah, I'll keep it brief
- 2:32:11basically. Whenever we train on a larger
- 2:32:13data set, we see Bigfoot. Okay.
- 2:32:17Basically, whenever we train on a larger dataset, we see big
- 2:32:19improvements in our model performance, and basically,
- 2:32:22whenever we initialize our networks with, you know, some
- 2:32:24pre-training step from some other auxiliary task, we
- 2:32:27basically see improvements the self supervised door, supervised
- 2:32:30with large data sets, both helped a lot.
- 2:32:35Hi. So at the beginning Ilan said
- 2:32:38that Tesla was potentially interested in building
- 2:32:40artificial general intelligence systems given the potentially
- 2:32:44transformative impact of technology like that.
- 2:32:46It seems prudent to invest in technical AGI, safety expertise
- 2:32:51specifically. I know Tesla does a lot of
- 2:32:54technical narrow AI Safety Research I was curious if Tessa
- 2:32:58was intending to try to build expertise in technical
- 2:33:02artificial, general intelligence safety specific Lee.
- 2:33:06Well, if I mean, if we're starts looking like we're getting be
- 2:33:10making a significant contribution to artificial
- 2:33:13general intelligence and then we'll for sure.
- 2:33:15Invest in in safety, I'm a big believer in AI safety.
- 2:33:19I think there should be an AI sort of regulatory Authority at
- 2:33:24at the government level. Just as there is a regulatory
- 2:33:27Authority for anything that affects Public Safety.
- 2:33:30So we have regulatory Authority for aircraft and cars and sort
- 2:33:35of Food and Drugs. Yes.
- 2:33:36And because they affect Public Safety and AI also affects
- 2:33:40Public Safety. So I think and this is not
- 2:33:43really something that government I think understands yet, but I
- 2:33:46think I think there should be a referee that is ensuring or
- 2:33:50doing trying to ensure Public Safety for AGI.
- 2:33:57Anything of like, well what are the elements that are necessary
- 2:34:00to create a GI like the accessible data set is extremely
- 2:34:07important and if you've got a large number of cars and
- 2:34:12humanoid, robots processing, petabytes of video data and
- 2:34:20audio data from The Real World. Just like humans that that's
- 2:34:25that might be the biggest data set.
- 2:34:26Probably. The biggest data set because in
- 2:34:30addition to that you can obviously incrementally scan the
- 2:34:33internet. But what the internet can't
- 2:34:36quite do is have millions or hundreds of millions of cameras
- 2:34:40in the real world. And with like said with audio
- 2:34:44and and other senses as well. So I think we're probably will
- 2:34:49have the most amount of data and probably the most amount of
- 2:34:55training power. Therefore probably we will make
- 2:35:00a contribution to a GI. Hey, I noticed the semi was back
- 2:35:09there, but we haven't talked about it too much.
- 2:35:11I was just wondering for the semi truck.
- 2:35:13What are the changes? Your thinking, about?
- 2:35:15From a sensing perspective. Imagine there's very different
- 2:35:18requirements, obviously than just a car if and if you don't
- 2:35:22think that's true. Why is that true?
- 2:35:24No, I think basically you can drive a car.
- 2:35:28I mean think about what drives any vehicle.
- 2:35:30It's a biological neural net with with eyes were Essentially.
- 2:35:36So, if and really a, what is your primary sensors, are two
- 2:35:44cameras, on a slow gimbal, a very slow gimbal.
- 2:35:48That's, that's your head. So if, you know of biological
- 2:35:53neural net with with two cameras on a slow, gimbal can drive a
- 2:35:56semi truck then. If you've got like eight cameras
- 2:36:00with continuous 360-degree Vision, operating at a higher
- 2:36:04frame rate and much higher reaction, Run rate than I think
- 2:36:06it is obvious that you should be able to drive a semi or any of
- 2:36:09any vehicle much better than human.
- 2:36:14Hi, my name is Akshay. Thank you for the even assuming,
- 2:36:19you know, Optimus would be used for different use cases and
- 2:36:22would evolve at different piece. For these use cases, would it be
- 2:36:27possible to sort of develop and deploy different software and
- 2:36:31Hardware components independently and deploy them?
- 2:36:35You know in the in optimist so that the overall you know
- 2:36:40feature development is faster for optimism.
- 2:36:48Okay. All right.
- 2:36:49What we did not comprehend. Unfortunately, our neural net to
- 2:36:53not comprehend the question. So next question, Hi.
- 2:37:02I want to switch the gear to the autopilot.
- 2:37:04So when you guys plan to roll out the FST beta, two countries
- 2:37:09other than the US and Canada. And also my next question is
- 2:37:13what's the biggest bottleneck or the technology Co barrier using
- 2:37:16in the current or the power of the stack and how you envision
- 2:37:19to solve that to make the autopilot is considerably better
- 2:37:23than human in terms of web performance Matrix X safety
- 2:37:26assurance and the human component is I think you Also,
- 2:37:30mention 4V fsdb wherever you are, guys going to combine the
- 2:37:34highway and the city is a single stack and some architectural big
- 2:37:38Improvement, can you maybe explain a bit about on that?
- 2:37:41Thank you. Well, that's a whole bunch of
- 2:37:43questions, will we? We're hopeful to be able to, I
- 2:37:49think from a technical standpoint FSD beta should be,
- 2:37:53it should be possible to roll on SF is T beta worldwide by the
- 2:37:57end of this year. But we, you know, for a lot of
- 2:38:02countries we need regulatory approval.
- 2:38:05And so we are somewhat gated by the regulatory approval in other
- 2:38:08countries. But you know, but I think from a
- 2:38:14technical standpoint, it will be ready to go to a worldwide beta
- 2:38:19by the end of this year and there's quite a big Improvement
- 2:38:22that we're expecting to release next month.
- 2:38:25That will always be especially good at assessing the velocity
- 2:38:31of fast-moving, Crush traffic and a bunch of other things.
- 2:38:34So anyway, elaborate Yeah guess so there used to be a lot of
- 2:38:42differences between production autopilot and the pole,
- 2:38:45self-driving beta. But those differences have been
- 2:38:47getting smaller and smaller over time I think just a few months
- 2:38:50ago we now use the same vision. Only object in stack in both FSD
- 2:38:55and in the production autopilot on all vehicles there's still a
- 2:38:59few differences. The primary one being the way
- 2:39:01that we predict Lanes right now. So we upgraded the modeling of
- 2:39:04Lane so that I could handle these more complex.
- 2:39:06Geometries like I mentioned in the talk, In production on a pie
- 2:39:09that we still use a simpler lane model but we're extending our
- 2:39:13current Deputy beta models to work in all sort of Highway
- 2:39:17scenarios as well. Yeah.
- 2:39:19And the the version of FST better that I drive actually
- 2:39:22does have the integrated stack. So this uses the FST, stack both
- 2:39:27the city streets and Highway and works quite well for me.
- 2:39:32We need to validate it in all kinds of weather like heavy
- 2:39:35rains snow dust and just make sure it's working as better than
- 2:39:42the production stack in across a wide range of environments, but
- 2:39:48it was pretty close. ooh that I mean I think it's, I don't know,
- 2:39:52maybe Definitely be before the end of the year and maybe
- 2:39:57November. Yeah.
- 2:39:59In our personal drives, the FST stack on Highway drives, already
- 2:40:02way. Better than the production stack
- 2:40:03we have and we do expect to also include the parking lot stack as
- 2:40:08a part of the FSC stack before the end of this year.
- 2:40:11So that will basically bring us to you, sit in the car in the
- 2:40:15parking lot and drive till the end of the parking lot and The
- 2:40:17Parking Spot before the end of this year.
- 2:40:19And in terms of like the fundamental, the fundamental
- 2:40:23metric to optimize against is how many miles per hour in
- 2:40:27between in necessary intervention.
- 2:40:29So just massively improving the, how many miles the car can drive
- 2:40:36on in full autonomy before and intervention is required.
- 2:40:39That is safety critical. So, yeah, that's that's the
- 2:40:46fundamental metric that we're measuring every week and we're
- 2:40:50making radical improvements on that.
- 2:40:55Hi. Thank you.
- 2:40:57Thank you so much for the presentation, very inspiring.
- 2:41:00My name is Daisy. I actually have a non technical
- 2:41:03question for you. I'm curious, if you are back to
- 2:41:06your twenties, what are some of the things you wish you knew
- 2:41:10back then? What are some advice?
- 2:41:12You would give to your younger self?
- 2:41:25Well, I'm trying to figure out something useful.
- 2:41:28To say. Yeah.
- 2:41:32Yeah. Joint Tesla will do one thing.
- 2:41:38Yeah, I think just going to try to expose yourself to as many
- 2:41:43smart people as possible. And I read a lot of books.
- 2:41:52You know, I do that. Did do that though.
- 2:41:55So I think there's some Merit to just also like not being like
- 2:42:03necessarily too intense and and like enjoying the moment a bit
- 2:42:09more. I would say two twenty or twenty
- 2:42:11something me just a you know stop and smell.
- 2:42:15The roses occasionally would probably be a good idea, you
- 2:42:20know? It's like when we were
- 2:42:22developing the the Falcon 1 rocket And on the kwajalein
- 2:42:28atoll and we have this beautiful little island that we're
- 2:42:31developing the rocket on and not.
- 2:42:33Once did that during that entire time that I even have a drink on
- 2:42:36the beach? I'm like well I should have had
- 2:42:38a drink on the beach that would have been fine.
- 2:42:44Thank you very much. I think you have excited all of
- 2:42:47the robotics people with with Optimus.
- 2:42:50This feels very much like 10 years ago in driving but as
- 2:42:55driving has proved to be harder than it actually look 10 years
- 2:42:58ago. What do we know now that we
- 2:43:00didn't ten years ago? That would make for example, AGI
- 2:43:03on a humanoid come faster Well, I mean it it seems to me that
- 2:43:09HEI is advancing very quickly hardly a week goes by without
- 2:43:15some significant announcement. and yeah, I mean This point.
- 2:43:23Like AI seems to be able to win at almost any rule, based game.
- 2:43:29It's able to create extremely impressive art.
- 2:43:36Engage in conversations that are very sophisticated you know
- 2:43:43write essays and these these just keep improving and there's
- 2:43:50so much more so many more talented people working on AI.
- 2:43:55And the hardware is getting better.
- 2:43:56I think it's a, a eyes on a super like a strong exponential
- 2:44:01curve of improvements, independent of what we do at
- 2:44:05Tesla and obviously will benefit somewhat from that exponential
- 2:44:10curve of improvement. With a, i Excessive just also
- 2:44:16happens to be very good at actuators that Motors Motors
- 2:44:19gearboxes controllers, Power Electronics, batteries sensors.
- 2:44:25And you know, really like I say that, you know, the biggest
- 2:44:29difference between the robot on four wheels and the robot with
- 2:44:33arms and legs is is getting the actuators, right?
- 2:44:37Actually, it's an actuators and sensors problem.
- 2:44:41And obviously you know how you control those actuators and
- 2:44:44sensors but it's a yeah, actuators and sensors and how
- 2:44:49you control the actuators? It's I don't.
- 2:44:52We have to have like the ingredients necessary to create
- 2:44:54a compelling robot and we're doing it.
- 2:44:56So Hi Lon you are actually bringing the humanity to the
- 2:45:05next level literally Tesla and you are bringing the humanity,
- 2:45:08the next level. So you said Optimus Prime
- 2:45:12Optimus will be used in next Tesla Factory.
- 2:45:15My question is, will a new Tesla Factory will be fully run by
- 2:45:20Optimus program? and and when can general public
- 2:45:26order humanoid, Yeah, I think it'll, it'll we're going to
- 2:45:30start Optimist with very simple tests in the factory.
- 2:45:35You know, like maybe just like loading apart like you saw in
- 2:45:37the video loading apart. You know, carrying a pot from
- 2:45:42one place to another or loading apart into a one of our more
- 2:45:46conventional robot cells to, you know, that that world's body
- 2:45:52together. So we'll start you know just
- 2:45:55trying to how do we make it useful at all and then and then
- 2:45:58gradually expand the number of situations where it's useful.
- 2:46:03And I think that that whatever situation is where Optimus is
- 2:46:07useful will grow exponentially. Like really, really fast in
- 2:46:14terms of when people can order one.
- 2:46:17I don't know. I think it's not that far away.
- 2:46:20Well, I think you mean, when can people receive one?
- 2:46:25So, I don't know. I'm like, I'd say probably
- 2:46:29within three years, not more than five years went within
- 2:46:33three to five years. You could probably receive an
- 2:46:35optimist, I feel the best way to make the progress for a gi's to
- 2:46:44involve as many smart people across the water as possible,
- 2:46:47and given the size and resource of Tesla compared to robot
- 2:46:51companies and given the state of human research at the moment,
- 2:46:55wouldn't make sense for the kind of Tesla to sort of Open Source
- 2:46:59from of the simulation? How do our parts I think Tesla
- 2:47:03can still be the dominant platformer, where it can be
- 2:47:05something like an Android OS or like iOS stuff for the entire
- 2:47:10human research with Would that be something that rather than
- 2:47:13keeping The Optimist to just Tesla researchers or the factory
- 2:47:17itself can open it and let the whole world exploding, my
- 2:47:20research? I think we have to be careful
- 2:47:29about Optimus being potentially used in ways that are bad
- 2:47:34because that is one of the possible things to do.
- 2:47:37So, I think we're, you know, we're provide optimists where
- 2:47:45you can provide instructions to Optimist but where those
- 2:47:48instructions are, you know, governed by some laws of
- 2:47:52robotics that you cannot overcome.
- 2:47:58So, you know, not doing harm to others and Without I think
- 2:48:05probably have quite a few safety related things with Optimus.
- 2:48:09Yeah. So I wrote it will just take
- 2:48:12maybe a few more questions and then and then and then thank you
- 2:48:14all for coming. Questions one deep and one broad
- 2:48:21on the Deep for Optimus. What's the current and what's
- 2:48:24the ideal controller bandwidth? And then in the broader question
- 2:48:29there's this big advertisement for the depth and breadth of the
- 2:48:32company. What is it uniquely about Tesla
- 2:48:36that enables that? Anyone want to tackle the
- 2:48:40bandwidth question? So the technical bandwidth of
- 2:48:45the close to your mouth, okay? For the bandwidth question, you
- 2:48:49have to understand or figure out, what is the task that you
- 2:48:52wanted to do and what is the free if you took a frequency
- 2:48:56transform of that task? What is it that you want your
- 2:48:58limbs to do? And that's why you get your
- 2:49:00bandwidth from. It's not a number that you can
- 2:49:02specifically, just say, you need to understand your use case.
- 2:49:04And that's from, that's where the bandwidth comes from.
- 2:49:08What are the broad question? I don't.
- 2:49:23I researched on the band. The question I think we're
- 2:49:25probably will just end up increasing the bandwidth or
- 2:49:28rear, you know, which translates to the effect of dexterity and
- 2:49:34reaction time of the, of the robot, like, you get it safe
- 2:49:38safe. It's not one hurts and it's
- 2:49:42maybe you don't need to go all the way to 100 Hertz, but I do
- 2:49:46maybe 10, 20, 50, I don't know it.
- 2:49:48But over time, I think the bandwidth will increase quite a
- 2:49:52bit. Or translated to dexterity and
- 2:49:55latency. You'd want to minimize that over
- 2:49:59time. Yeah.
- 2:50:02Minimize latency maximize dexterity in terms of breadth
- 2:50:07and depth. I guess we're we've got, we're
- 2:50:11pretty big company at this point.
- 2:50:13So we've got a lot of different areas of expertise that we
- 2:50:15necessarily had to develop in order to make autonomous or in
- 2:50:18order to make electric cars and then order to make autonomous
- 2:50:21electric cars. So we've just Tesla is like a
- 2:50:26whole series of startups basically and so far, they've
- 2:50:32almost all been quite successful.
- 2:50:35So we must be doing something right.
- 2:50:37And I, you know, I consider one of my core responsibilities
- 2:50:42writing company is to have an environment where great
- 2:50:46Engineers can flourish and I think in a lot of companies on a
- 2:50:51maybe most companies, if somebody's a really talented
- 2:50:55driven engineer, the they're unable to actually their talents
- 2:51:01are suppressed affect a lot of companies and it's, you know,
- 2:51:06and some of the companies that the engineering Talent is
- 2:51:09suppressed in a way that is, maybe not obviously bad but
- 2:51:13where it's just so comfortable and you paid so much money and
- 2:51:17you, but your The output you actually have to produce is so
- 2:51:21low. That is like a Honey Trap, you
- 2:51:23know, so like there's a few Honey Trap.
- 2:51:26It's places in Silicon Valley, where they don't like
- 2:51:28necessarily don't seem like bad places for engineers but he have
- 2:51:32to say like a good engineer went in and what did they get out and
- 2:51:37the output of that? Engineering Talent is seems very
- 2:51:41low even though there seem to be enjoying themselves that's why I
- 2:51:47call it the the few Honey Trap companies in Silicon Valley,
- 2:51:50Tesla is not a Honey Trap. We're demanding and it's like
- 2:51:54going to get a lot of shit done and it's going to be really
- 2:51:57cool. And it's not going to be easy,
- 2:52:02but if you are a super talented engineer, your talents will be
- 2:52:10used. I think to, a greater degree
- 2:52:14than anywhere else. You know, SpaceX also that way.
- 2:52:19So Highland of I have two questions, so both to the
- 2:52:26autopilot team. So the thing is like I have been
- 2:52:28following your progress for the past few years.
- 2:52:30So today, you have made changes on like the lane detection, like
- 2:52:34you said that like previous, you're doing instance, somatic
- 2:52:36segmentation. Now you guys are built transfer
- 2:52:38models for like building the lanes.
- 2:52:41So what are another some other common challenges which you guys
- 2:52:44are facing right now, like which you are solving in future as a
- 2:52:47curious Engineers. So that like we as a researcher
- 2:52:50can work on those, Start working on those.
- 2:52:52And the second question is like, I'm really curious about the
- 2:52:54data engine. Like you guys have like stole a
- 2:52:58case, like why the car is stopped.
- 2:53:00So how are you finding cases, which is very much similar to
- 2:53:03that from the data, which you have like.
- 2:53:05So little bit more on the data engine would be great.
- 2:53:07So that's it for our star answer.
- 2:53:11The first question using occupancy that work as an
- 2:53:13example. So what you saw in the
- 2:53:17presentation did not exist a year ago.
- 2:53:19So we only spent one year About time, we're shipping, wouldn't
- 2:53:22Health occupancy Network and Q. Have a One Foundation model,
- 2:53:27actually to represent the entire physical world, around
- 2:53:31everywhere. And you always a condition is
- 2:53:34actually really, really challenging.
- 2:53:35So, only over the year ago, we're kind of like driving a 2d
- 2:53:39were if there's a war and it uses curb what kind of represent
- 2:53:43with the same static Edge, which is obviously, you know, not not
- 2:53:47ideal, right? There's a big difference between
- 2:53:49a curb and War when you drive you Make different choices,
- 2:53:52right? So after we realized that, what
- 2:53:54we go to 3D will have you basically rethink the entire
- 2:53:57problem and think about how we address that.
- 2:54:00So this will be like one example of a challenges, we have a web
- 2:54:04of conquer in the past year, Yeah, to answer the question
- 2:54:10about how we actually source examples of this tricky stopped
- 2:54:13cars. There's a few ways to go about
- 2:54:15this but two examples are one we can trigger for disagreements
- 2:54:19within our signals. So let's say that parked bit
- 2:54:22flickers between parked and driving will trigger that back.
- 2:54:26And the second is, we can leverage more of the Shadow mode
- 2:54:28logic. So if the customer ignores the
- 2:54:30car, but we think we should stop for it will get that data back
- 2:54:34to. So these are just different like
- 2:54:36various trigger logic that allows us to get those data.
- 2:54:39And pains back. Hi, thank you for the amazing
- 2:54:46presentation. Thanks so much.
- 2:54:48So there are a lot of companies that are focusing on the AGI
- 2:54:52problem, and one of the reasons why such a hard problem is
- 2:54:55because the problem itself is so hard to Define, several
- 2:54:58companies have several different definitions, they focus on
- 2:55:01different things. So, what is Tesla house, tester,
- 2:55:03defining, the AGI problem. And what are you focusing on
- 2:55:06specifically? Well, we're not actually
- 2:55:11specifically focused on a GI, I'm simply saying that AGI is so
- 2:55:15is seems likely to be an emergent property of what we're
- 2:55:19doing because we're creating all these autonomous cars, and
- 2:55:24autonomous humanoids that are actually with a truly gigantic
- 2:55:32data stream. That's coming in and being
- 2:55:34processed. It's by far the most amount of
- 2:55:38real-world data. And they do you can't get by
- 2:55:41just searching the internet because you have to be out there
- 2:55:43in the world and interacting with people and interacting with
- 2:55:46the roads. And, and just, you know, Earth
- 2:55:50is big place and reality is messy and complicated.
- 2:55:53So, so, I think it's sort of like, like you to just, it just
- 2:55:58seems likely to be an emergent property of if you've got, you
- 2:56:02know, tens or hundreds of millions of autonomous vehicles
- 2:56:04and, and maybe even a comparable number of humanoids, maybe more
- 2:56:08than that on the human heart. Run.
- 2:56:10Well, that's just the most amount of data and if that that
- 2:56:14videos being processed, it just seems likely that you know, the
- 2:56:19cars will will definitely get way better than human drivers
- 2:56:23and the humanoid robots will become increasingly
- 2:56:29indistinguishable from humans perhaps. and so then, like I
- 2:56:35said, you have a Emergent property of AGI.
- 2:56:45I believe you know, humans collectively are sort of a super
- 2:56:49intelligence as well especially as we improved the data rate
- 2:56:53between humans. Anything like that seem to be
- 2:56:55way back in the early days of the internet was like the
- 2:56:58internet was like Humanity, acquiring a nervous system where
- 2:57:03now all of a sudden any one element of humanity could know
- 2:57:07all of the knowledge of humans by connecting to the internet,
- 2:57:11almost all the knowledge on suddenly, here's part of it,
- 2:57:13whereas Ously, we would exchange information by osmosis it by you
- 2:57:18know, by would have liked in order to transfer data.
- 2:57:21So you would have to write a letter somewhere, we have to
- 2:57:23carry the letter by person to another person and then a whole
- 2:57:27bunch of things in between and then it was like yeah I mean
- 2:57:34insanely slow when you think about it and even if you were in
- 2:57:38the Library of Congress you still didn't have access to all
- 2:57:40the world's information and you certainly couldn't search it.
- 2:57:44And obviously very few people are in the Library of Congress.
- 2:57:48So I mean, one of the Great I sort of equality elements like
- 2:57:57the internet is has been the most, the biggest equalizer in
- 2:58:00history in terms of access to information and knowledge.
- 2:58:06And any student of History I think would agree with this
- 2:58:09because, you know, you go back a thousand years.
- 2:58:11There were very few books like like and books would be
- 2:58:14incredibly expensive but only a few people knew how to read and
- 2:58:17only if an even smaller number of people even had a book now.
- 2:58:22Now look at it like you can Access any book instantly?
- 2:58:25You can learn anything for basically, for free is pretty
- 2:58:29incredible. So, you know, I was asked
- 2:58:36recently, what period of History would?
- 2:58:40I prefer to be at the most? And my answer was right now.
- 2:58:46This is the most interesting time in history and I read a lot
- 2:58:49of history. So let's let's do our best to
- 2:58:53keep that going. Yeah.
- 2:58:57And to go back to one of the earlier questions.
- 2:58:58I would ask her like you can the the thing that's happened over
- 2:59:02time with respect to Tesla autopilot is that we've just the
- 2:59:09The neural Nets have gotten have gradually absorbed more, and
- 2:59:13more software. And in the limit, of course, you
- 2:59:16could simply take the videos as seen by the car and compare
- 2:59:22those to the the steering inputs from the steering wheel and
- 2:59:25pedals which very simple inputs. And it in principle, you could
- 2:59:30train with nothing in between because that's what humans are
- 2:59:34doing with the biological neural, net, you could train
- 2:59:37based on video and And the, what trains the video is the moving
- 2:59:43of the steering wheel and the pedals with no other software in
- 2:59:46between, in between, we're not there yet, but it's gradually
- 2:59:50going in that direction. All right, but we will last
- 2:59:56question. I think we got a crush on the
- 3:00:01front here. They're right there.
- 3:00:04I will do two questions. Fine.
- 3:00:08Thanks for such a great presentation.
- 3:00:10Well, the old question last with FSD being used by so many people
- 3:00:16do you think, what's the come, how do you evaluate the
- 3:00:18company's risk tolerance in terms of performance statistics
- 3:00:21and do you think there needs to be more transparency or
- 3:00:23regulation from third parties as to how what's good enough and
- 3:00:28defining, like Thresholds for performance across some many
- 3:00:33miles. Sure.
- 3:00:35We'll the, you know, the number one design requirement at Tesla
- 3:00:43is safety so that it goes across the board.
- 3:00:47So in terms of the mechanical, safety of the car, we have the
- 3:00:51lowest probability of injury of any cars ever tested by the
- 3:00:53government for just passive mechanical safety essentially,
- 3:00:58crash structure and airbags or whatnot.
- 3:01:03We have the best Best highest rating for active safety, as
- 3:01:07well. And is going to get to the point
- 3:01:12where your active safety is. So ridiculously good, it's like
- 3:01:16just absolutely better than human.
- 3:01:20And then with respect to autopilot, we do publish this.
- 3:01:24Broadly speaking the statistics on miles driven or with cars
- 3:01:30that have no autonomy or Tesla cars with no autonomy with Kind
- 3:01:34of Hardware one Hardware to Hardware 3 and then the ones
- 3:01:39that are in FST beta. And we see steady improvements
- 3:01:43all along the way. And, you know, sometimes there's
- 3:01:46this dichotomy of, you know, should you wait until the car is
- 3:01:52like on Earth, three times safer than a person before deploying
- 3:01:56any technology? But I think that's that's
- 3:01:58actually morally wrong at the point of which he believed that
- 3:02:02that adding autonomy. Deuces, injury and death.
- 3:02:09I think you have a moral obligation to deploy it, even
- 3:02:12though you're going to get sued and blamed by a lot of people,
- 3:02:16because the people whose lives you saved, don't know that their
- 3:02:19lives are saved. And the people, the people who
- 3:02:23do occasionally die or get injured, they definitely know or
- 3:02:26their state does that it was, you know, whatever.
- 3:02:29There's a problem with with, with autopilot.
- 3:02:32That's why you have to look at the The numbers, in terms of
- 3:02:35total miles driven, how many accidents occurred?
- 3:02:38How many accidents were serious? How many fatalities and we've
- 3:02:41got oh well over three million cars in the road.
- 3:02:43So this it's that's a lot of miles driven every day.
- 3:02:47It's not going to be perfect but what what matters it is that is
- 3:02:50that it is very clearly safer than not deploying it.
- 3:02:56Yeah. So last question, Thanks for the
- 3:03:06last question here. Okay.
- 3:03:14Hi. So I do not work on Hardware, so
- 3:03:19maybe the hardware dream and you guys can enlighten me, why is it
- 3:03:24required that there be symmetry in the design of optimist?
- 3:03:29Because humans, we have handedness, right?
- 3:03:33We are, we use some set of muscles more than others or
- 3:03:37time. There's wear and tear, right?
- 3:03:39So maybe you'll start. You see some joint failures or
- 3:03:42some actuator failures. More over time, I understand
- 3:03:47that this is extremely pre-staged.
- 3:03:49Also, we, as humans have based so much fantasy and fiction over
- 3:03:56super human capabilities, like all of us.
- 3:03:58Don't want to walk right over there.
- 3:04:00We want to extend our arms and like we have all these, you
- 3:04:04know, a lot of fantasy, Fantastical designs so
- 3:04:08considering everything else that is going on.
- 3:04:11Terms of batteries and intensity of compute.
- 3:04:15Maybe you can leverage all those aspects into coming up with
- 3:04:20something. Well, I don't know more
- 3:04:23interesting in terms of your the robot that you're building and
- 3:04:27I'm hoping you're able to explore those directions.
- 3:04:31Yeah. I mean I think it would be cool
- 3:04:32to have like, you know, make Inspector Gadget real.
- 3:04:35That would be pretty sweet. So yeah, I mean, right now we're
- 3:04:40just want to make Make basic humanoid what work well and our
- 3:04:44goal is fastest path to a useful humanoid robot.
- 3:04:48I think this is this will ground Us in reality literally and
- 3:04:53ensure that we are doing something useful.
- 3:04:57Like one of the hardest things to do is to be useful to
- 3:05:02actually and then and then to have high utility under the
- 3:05:05curve of like, how many people did you help taught?
- 3:05:08You know and how much help did you Provide to each person on
- 3:05:14average. And then how many people did you
- 3:05:15help the total utility like trying to actually ship useful
- 3:05:20product? That people like to a large
- 3:05:23number of people is so insanely hard.
- 3:05:26I it's boggles the mind, you know, it's like I said, like
- 3:05:29man, there's a hell of a difference between a company
- 3:05:31that has shipped product and one has not sure product.
- 3:05:34It's a game. This is night and day.
- 3:05:36And then even once you ship product, can you make the cost?
- 3:05:39The value of the output? Worth more than the cost of the
- 3:05:42input, which is again, insanely difficult, especially with
- 3:05:45Hardware, so. But I think over time, I think
- 3:05:49we cool to do creative things and have like eight arms and
- 3:05:53whatever and have different versions and maybe, you know,
- 3:05:58they'll be some Hardware like companies that are able to add
- 3:06:03things to an optimist like maybe we, you know, add a power port,
- 3:06:08or something like that or attached.
- 3:06:10Then you can add You know, add attachments to your Optimist,
- 3:06:13like you can add them to your phone, it could be.
- 3:06:16So a lot of cool things that could be done over time and it
- 3:06:18could be, maybe an ecosystem of small companies that or big
- 3:06:21companies that make add-ons for Optimus.
- 3:06:25So with that, I'd like to thank the team for their hard work.
- 3:06:31You guys are awesome and And I think and I thank you all for
- 3:06:40coming and for everyone online. Thanks for tuning in and I think
- 3:06:45this will be one of those great videos where you can.
- 3:06:47Like if you you can fast-forward to the bits that you find most
- 3:06:50interesting. But we try to give you a
- 3:06:53tremendous amount of detail literally.
- 3:06:56So you can look at the video at your leisure and you can focus
- 3:07:00on the parts that you find interesting and skip the other
- 3:07:02parts. So thank you all.
- 3:07:05And we'll do this to try to do this every year and we might do
- 3:07:08a part, a monthly podcast even so, but I think it'd be, you
- 3:07:12know, great to sort of bring you along for the ride and and like
- 3:07:17show you what cool things are happening.
- 3:07:20And yeah, thank you. All right, thanks.
- 3:07:33That gets run, Rashad JCPenney for thousands of deals solo.
- 3:07:36No coupons needed this weekend, save up to 50% on kitchen
- 3:07:40electrics from Brands like, Keurig and Cuisinart, say yes,
- 3:07:43please to Diamonds and gemstones.
- 3:07:44Now, 1999 each and bundle up the famine coats starting at 1499.
- 3:07:49We got your holiday. A valid on select items. 1118 3,
- 3:07:561120, excluded from coupons exclusions apply C Storage
- 3:07:59acb.com for details