Latest / Elon Musk Podcast / Tesla AI Day Featuring ELON MUSK
Transcript
- 12:26The show is brought to you by backblaze, I use back place to
- 12:29backup my podcast, my video files, all of my writing stuff,
- 12:33and all my photos. And you get unlimited computer
- 12:35backup, for Macs and PCs for just $7 a month.
- 12:39You can backup your own. Photos videos, drawings
- 12:42projects, all of your data and access your backed up data from
- 12:46anywhere in the world using the web app and you can access the
- 12:49files on your mobile to iOS Android apps all covered.
- 12:53And this is a cool part. This is my favorite part.
- 12:54You can restore it by mail, a hard drive will come to your
- 12:57house with all your data shipped to your door.
- 13:00I could come to your business to and you can restore return
- 13:03refund program so you can buy a hard drive restore, send the
- 13:06hard drive back within 30 days and get a full refund.
- 13:09So basically, they A ship you this hard drive and then you
- 13:12ship it back and you don't ever pay for it.
- 13:14Which is the perfect program for somebody who has huge files.
- 13:17You don't want to waste days and days, downloading terabytes and
- 13:21teraflops of data. And if you're worried about
- 13:23accidentally, deleting your files to bucks extra month, you
- 13:27can increase your retention history to one year, and I use
- 13:30it for all of my video. Files comes in super handy.
- 13:33So seven dollars plus two dollars nine dollars a month and
- 13:36you get everything backed up. He's a mind for up to a year and
- 13:40if you Use the URL back. Place.com slash Ilan, you get a
- 13:45fully featured 15-day. No credit card required.
- 13:49Free trial, check it out, play with it.
- 13:51Start protecting yourself from potential, bad times, back your
- 13:54stuff up. It's recommended by the New York
- 13:57Times ink macworld PC World. Life wire wired, Tom's, guide,
- 14:01925 mac and more and it's recently been listed on the
- 14:04NASDAQ stock exchange under be lze.
- 14:07So you know, they're legit back. Glaze is committed more than
- 14:10Ever to bring an easy and affordable data storage that you
- 14:13can trust don't be that person that forgot to back up their
- 14:16important files. We've got your back sign up for
- 14:19a free 15 day trial no credit card required go there sign up I
- 14:24you with it it's really powerful.
- 14:26It is really easy to use so go to back.
- 14:28Plays.com /l on backblaze.com Ilan, backblaze.com Ilan with
- 14:37progressives Name, Your Price tool, you can find options that
- 14:39fit your budget. Because giving you options is
- 14:42the right thing to do. Oh yeah.
- 14:44Like when I hold the door for someone sure it may be weird if
- 14:47I don't time it right in there, a little too far away.
- 14:49And now they're running and we're both down asking
- 14:52ourselves. Is it worth it to run instead of
- 14:55just, you know, letting them open their own door but still,
- 14:58it's the right thing to do. So, get options based on your
- 15:00needs. With progressives Name, Your
- 15:01Price tool Progressive Casualty Insurance Company of aliens and
- 15:03third party insurers. Comparison is not available in.
- 15:05All states are situations prices are days and how you buy.
- 21:48Armed. You made a mistake uncut to
- 40:59technological, told everyone everywhere, but I can tell you
- 41:03anything like this. Armed.
- 49:51Hello everyone. Sorry for the delay.
- 49:54Thanks for coming and sorry, I had some technical difficulties.
- 49:59Really AI for this. So what do you want to show
- 50:04today is that Tesla is much more than an electric car company,
- 50:08that we have deep AI activity in Hardware, on the inference level
- 50:15on the training level and obviously, I think arguably the
- 50:23leaders in real-world AI as it applies, the real world and
- 50:28those of you who have seen Seen the full self-driving beta can
- 50:32appreciate the rate at which the Tesla neural net is learning to
- 50:36drive. And this is a particular
- 50:40application of AI, but I think there's, there's they're more
- 50:43applications down the road that will make sense.
- 50:46And we'll talk about that later in the presentation.
- 50:49But yeah, we're basically want to encourage anyone who is
- 50:53interested in solving real-world, AI problems at
- 50:57either the hardware, or the software level.
- 50:59You join Tesla or considering Tesla?
- 51:01So let's see. We'll start off with Andre Yep.
- 51:23Okay. Hello great.
- 51:25Okay hi everyone and welcome. My name is Andre and I am.
- 51:30I lead the Byzantine here at Tesla autopilot and I'm
- 51:33incredibly excited to be here to kick off this section.
- 51:37Giving you a technical Deep dive into the autopilot stack and
- 51:41showing you all the under-the-hood components that
- 51:42go into making a car drive all by itself.
- 51:46So we're going to start off with the vision component here.
- 51:49Now in division component, what we're trying to do is we're
- 51:51trying to design a neural network that processes the raw
- 51:54information, which in our case is the eight cameras that are
- 51:57positioned around the vehicle and they send us images and we
- 52:00need to process that in real time into what we call the
- 52:03vector space and this is a three-dimensional representation
- 52:05of everything you need for driving.
- 52:07So this is the 3 Dimension positions, Appliance edges,
- 52:10curbs traffic signs, traffic lights cars, their positions,
- 52:14orientations depth, Steven so on.
- 52:20So here I am showing a video of actually hold on.
- 52:27Apologies. So here I am showing the video
- 52:30of the raw inputs that come into the stock and then neural
- 52:34processes that into the vector space and you are seeing parts
- 52:37of that Vector space rendered in the instrument cluster on the
- 52:40car. Now what I find fascinating
- 52:54about this is that we are effectively building a synthetic
- 52:56animal from the ground up. So the car can be thought of as
- 52:59an animal, it moves around its senses the environment and you
- 53:02know, act autonomously and intelligently.
- 53:05And we are building all the components from scratch in
- 53:08house. So we are building, of course,
- 53:09all the mechanical components of the body, the nervous system,
- 53:12which is all the electrical components and for our purposes,
- 53:14the brain of the autopilot and specifically for this section,
- 53:18the synthetic visual cortex. Now, the No visual cortex
- 53:21actually has quite intricate structure and a number of areas
- 53:25that organize the information flow of this brain.
- 53:28And so in particular, in our in your visual cortices, the
- 53:32information hits, the light hits the retina goes through the lgn
- 53:36all the way to the back of your visual cortex goes through areas
- 53:38V1, V2 V4. The it the Venture on the dorsal
- 53:41streams and the information is organized in a certain layout.
- 53:45And so when we are designing the visual cortex of the car, we
- 53:49also want to design the neural network architecture.
- 53:51Sure of how the information flows in the system.
- 53:55So the processing starts in the beginning, when light hits our
- 53:58artificial retina and we are going to process this
- 54:01information with neural networks.
- 54:02Now, I'm going to roughly organize this section or
- 54:05chronologically so starting off with some of the neural networks
- 54:08and what they look like, roughly four years ago, when I joined
- 54:10the team and how they have developed over time.
- 54:13So roughly four years ago, the car was mostly driving in a
- 54:17single Lane going forward on the highway and so it had to keep
- 54:21playing and it had to keep distance away from the car in
- 54:22front of us. And at that time, Processing was
- 54:26only on individual Image level. So a single image has to be
- 54:29analyzed by a neural net and make little pieces of the vector
- 54:31space process that into a little pieces of the vector space.
- 54:34So this processing took the following shape, we take a 1280
- 54:39by 960 input and this is 12 bit integers streaming in at roughly
- 54:4236 hurts. Now, we're going to process that
- 54:45with a neural network. So instantiate a feature
- 54:47extractor Backbone. In this case, we use residual
- 54:50neural networks. So, we have a stem and a number
- 54:52of residual blocks connected in series, Now, the specific class
- 54:55of resonance that we use are pregnant because we let this is
- 54:59like a very regnant offer a very nice design space for neural
- 55:02networks because they allow you to very nicely trade-off latency
- 55:05and accuracy. Now, these rig mats, give us as
- 55:10an output a number of features at different resolutions in
- 55:13different scales. So in particular on the very
- 55:15bottom of this feature here are key.
- 55:16We have very high resolution information with very low,
- 55:18channel counts and all the way at the top we have low spatial
- 55:23low-resolution spatially but Hi Channel counts.
- 55:26So on the bottom we have a lot of neurons that are really
- 55:28scrutinizing, the detail of the image and on the top, we have
- 55:30neurons that can see most of the image.
- 55:32And a lot of that context, have a lot of Baptists in context.
- 55:36We then like to process this with feature pyramid networks,
- 55:38in our case, we like to use by FPS and they get the more they
- 55:42get to multiple scales to talk to each other effectively and
- 55:44share a lot of information. So, for example, if you're in
- 55:47Iran all the way down and then Network and you're looking at a
- 55:49small patch and you're not sure this is a car, not, it
- 55:51definitely helps to know from the top layers that.
- 55:54Hey, you are actually in the vanishing point of this highway.
- 55:57And so that helps to disambiguate that this is
- 55:59probably a car. After a biased PN and a feature
- 56:02Fusion across scales, we then go into task-specific heads.
- 56:06So for example, if you are doing object detection, we have a one
- 56:09stage, you're like object detector here, where we
- 56:12initialize a raster, and there's a binary bit per position
- 56:15telling you whether or not there's a car there and then in
- 56:18addition to that, if there is here's a bunch of other
- 56:20attributes, you might be interested in.
- 56:21So the X Y width, height offset or any of the other attributes,
- 56:24like what type of a car is this? And so on.
- 56:26So this is for the detection by itself.
- 56:29Now very quickly, we discovered that we don't Just want to
- 56:31detect cars. We want to do a large number of
- 56:33tasks. So for example, we want to do
- 56:36traffic light, recognition and detection a link prediction and
- 56:38so on. So very quickly, we conversion
- 56:40this kind of architectural layout where there's a common
- 56:43shared backbone and then branches off into a number of
- 56:45heads. So, we call these therefore
- 56:48hydrants. And these are the heads of the
- 56:50Hydra. Now, this architectural layout
- 56:54has a number of benefits. So, number one, because of the
- 56:58feature sharing, we can amortize the forward pass inference in
- 57:01the car at test time. And so this is very efficient to
- 57:04run because if we had to have a backbone for every single task,
- 57:07that would be a lot of backbones in the car.
- 57:10Number two, this decouples all of the tasks.
- 57:12So we can individually work on everyone task in isolation.
- 57:15And for example we can we can approve any of the data sets or
- 57:18change some of the architecture of the head and so on and you
- 57:20are not impacting any of the other.
- 57:21Tasks. And so we don't have to
- 57:23revalidate all the other tasks which can be expensive and
- 57:26number three because there's this bottleneck here and
- 57:28features. What we do fairly often is that
- 57:31we actually cache these features to disk and when we are doing
- 57:34these fine-tuning workflows, we only find tune from from the
- 57:38cash features up and only finding the heads.
- 57:40So most often in terms of our training workflows, we will do
- 57:43an end-to-end training run once in a while where we train
- 57:46everything jointly, then we cash the features at the multiscale
- 57:51feature level. Oh, and then we fine-tune off of
- 57:53that for a while and then and to entrain once again and so on.
- 57:58So here's the kinds of predictions that we were
- 58:00obtaining I would say several years ago now from one of these
- 58:04Hydro Nets. So again we are processing,
- 58:06individual images. There we go.
- 58:09We are processing. Just individual image and we're
- 58:11making a large number of predictions about these images.
- 58:13So for example, here you can see predictions of the stop signs,
- 58:16the stop lines, the lines, the edges, the cars, the traffic
- 58:20lights, the curbs here, whether or not the car is parked all of
- 58:25the static objects like trash cans cones.
- 58:27And so on and everything here is coming out of the net here in
- 58:30this case out of the Hydra net. So that was all fine and great
- 58:33but as we work towards FSD, we quickly found that this is not
- 58:36enough so where this First started to break was when we
- 58:40started to work on Smart. Summon here I am showing some of
- 58:43the predictions of only the curb detection task and I'm showing
- 58:46it now for every one of the cameras so we'd like to wind our
- 58:49way around the parking lot to find the person who is summoning
- 58:51the car. Now the problem is that you
- 58:53can't just directly drive on image space predictions.
- 58:56You actually need to cast them out and form some kind of a
- 58:59vector space around you. So we attempted to do this using
- 59:02C++ and developed what we call the occupancy tracker at the
- 59:06time. So here we see that the curb
- 59:10deductions from the images are being stitched up, the cross
- 59:13Camera scenes, camera, boundaries and overtime.
- 59:16Now there are two prong two major problems.
- 59:18I would say with the setup. Number one, we very quickly
- 59:20discovered that tuning the occupancy tracker and all of its
- 59:22hyperparameters was extremely complicated.
- 59:24You don't want to do this explicitly by hand and C++ you
- 59:27want this to be inside a neural network and train that into and
- 59:30number two we very quickly discovered that the image space
- 59:33is not the correct output space. You don't want to make
- 59:36predictions in image space. You really want to make it
- 59:37directly. We in the vector space.
- 59:39So here's a way of illustrating the issue.
- 59:42So, here I'm showing on the first row, the predictions of
- 59:45our curbs and our lines in red and blue, and they look great in
- 59:50the image. But once you cast them out into
- 59:52the vector space things start to look really terrible and we are
- 59:55not going to be able to drive on this so you see how the
- 59:59predictions are quite bad in Vector space.
- 1:00:01And the reason for this fundamentally is because you
- 1:00:03need to have an extremely accurate depth per pixel in
- 1:00:06order to actually do this projection.
- 1:00:07And so you can imagine just how high of the bar.
- 1:00:10It is to predict that depth. So, Late in these tiny and every
- 1:00:14single Pixel of the image and also have there's any included
- 1:00:17area where you'd like to make predictions, you will not be
- 1:00:19able to because it's not an image space Concept in that
- 1:00:23case. So, we very quickly real, the
- 1:00:28other problems with this, by the way, is also for object
- 1:00:30detection. If you are only making
- 1:00:31predictions per camera, then sometimes you will encounter
- 1:00:35cases like this, where a single car actually spans five of the
- 1:00:38eight cameras. And so if you are making
- 1:00:41individual predictions, then no single camera since sees, all of
- 1:00:44the car. And so, obviously, you're not
- 1:00:46going to be able to do a very good job of predicting that
- 1:00:48whole car. And it's going to be incredibly
- 1:00:49difficult to fuse these measurements.
- 1:00:52So we have this intuition that what we'd like to do instead is
- 1:00:54we'd like to take all of the images and simultaneously feed
- 1:00:57them into a single neural, net and directly output in Vector
- 1:00:59space. Now, this is very easily, said,
- 1:01:02much more difficult to actually achieve but roughly we want to
- 1:01:06lay out a neural net in this way where we process every single
- 1:01:09image with a backbone. And then we want to somehow fuse
- 1:01:12them and we want to represent it.
- 1:01:15The the features from image space features to directly some
- 1:01:18kind of a vector space features and then go into the decoding of
- 1:01:21the head now. So there are two problems with
- 1:01:25this problem. Number one, how do you actually
- 1:01:28create the neural network components that do this
- 1:01:30transformation? And you have to make a
- 1:01:33differentiable? So that end-to-end training as
- 1:01:35possible. And number two, the if you want
- 1:01:39Vector space predictions from your neural, net, you Vector
- 1:01:41spaces based data sets. So just labeling the images and
- 1:01:44so on is not going to get you there.
- 1:01:45Any Vector space labels. We're going to talk a lot more
- 1:01:48about problem. Number two, later in the talk
- 1:01:50for now, I want to focus on the neural network architecture.
- 1:01:52So I'm going to Deep dive into problem, number one, so here's
- 1:01:57the rough problem, right? We're trying to have this bird's
- 1:01:59eye view prediction, instead of image space predictions.
- 1:02:02So for example, let's focus on the single Pixel in the output
- 1:02:04space in the yellow and this pixel is trying to decide.
- 1:02:06Am I part of a curb or not as an example and now we're The
- 1:02:11support for this kind of a prediction come from in the
- 1:02:13image space. Well, we know roughly how the
- 1:02:16cameras are positioned, and their extrinsic and intrinsic.
- 1:02:18So we can roughly project this point into the camera images
- 1:02:22and, you know, the evidence for the whether or not this is a
- 1:02:24curb meat. Come from somewhere here in the
- 1:02:25images. The problem is that this
- 1:02:27projection is really hard to actually, get correct because it
- 1:02:30is a function of the road surface.
- 1:02:31The road surface could be sloping upward sloping down or
- 1:02:33also, there could be other data dependent issues.
- 1:02:36For example, there can be inclusion, do to a car.
- 1:02:38So, there's a car occluding, this, this viewport, This part
- 1:02:42of the image, then actually, you may want to pay attention to a
- 1:02:44different part of the image, not the part where it projects.
- 1:02:47And so, because this is data dependent is really hard to have
- 1:02:49a fixed transformation for this component.
- 1:02:51So in order to solve this issue, we use a Transformer to
- 1:02:56represent this space and the Transformer, it uses
- 1:03:00multi-headed self attention and blocks of it.
- 1:03:02In this case, actually, we can get away with even a single
- 1:03:05block doing a lot of this work and effectively, what this does
- 1:03:10is you July's, a raster of the size of the output space that
- 1:03:13you would like and you tile it with positional and Coatings
- 1:03:16with Sines and cosines in the output space.
- 1:03:18And then these get encoded with an MLP into set of query vectors
- 1:03:22and then all of the images, and their features also emit their
- 1:03:24own keys and values and then the queries keys and values feed
- 1:03:28into the multi-headed, self attention and so effectively.
- 1:03:30What's happening? Is that every single image piece
- 1:03:32is broadcasting in its key. What it is, what is it a part
- 1:03:36of? So hey, I'm part of a pillar in
- 1:03:38roughly this location and I'm seeing this kind of stuff and
- 1:03:41that's It's in the key and then every query is something along
- 1:03:43the lines of hey I'm a pixel in the output space at this
- 1:03:45position and I'm looking for features of this type, then the
- 1:03:49keys and the queries interact multiplicatively, and then the
- 1:03:51value is get pooled accordingly. And so this represents the space
- 1:03:56and we found this to be very effective for this
- 1:03:58transformation. So if you do all of the
- 1:04:00engineering correctly, this again is very easily, said
- 1:04:03difficult to do. You do all of the engineering
- 1:04:05correctly? There's one more, the real
- 1:04:09problem actually. Before, I'm not sure what's up
- 1:04:13at the slides. So one more thing you have to be
- 1:04:15careful with some of the details here.
- 1:04:16When you are trying to get this to work, some particular, all of
- 1:04:19our cars are slightly cockeyed in a slightly different way.
- 1:04:23And so, if you're doing this transformation from image space
- 1:04:25to the output space, you really need to know what your camera
- 1:04:27calibration is. And you need to feed that in
- 1:04:29somehow into the neural net. And so you could definitely just
- 1:04:32like concatenate the cattle camera calibrations of all of
- 1:04:34the images and somehow feed them in with an MLP.
- 1:04:38But actually, we found that we can do much better by
- 1:04:40transforming all of the images into a synthetic virtual camera
- 1:04:43using a special rectification transform.
- 1:04:46So, this is what that would look like.
- 1:04:48We insert a new layer right above the image rectification
- 1:04:52layer. It's a function of camera
- 1:04:53calibration, and it translates. All of the images into a virtual
- 1:04:56common camera. So if you were to average up a
- 1:04:59lot of repeater images, for example, which face the back you
- 1:05:02would without doing this, you would get a kind of a blur but
- 1:05:05after doing the rectification transformation, you see that the
- 1:05:08the back mirror gets really crisp.
- 1:05:11So once you do this, this improves the performance quite a
- 1:05:14bit. So here are some of the results.
- 1:05:17So on the left we are seeing what we had before and on the
- 1:05:20right. We're now seeing significantly
- 1:05:22improve predictions coming, directly out of the neural net.
- 1:05:24This is a multi-camera network, predicting directly in Vector
- 1:05:27space. And you can see that it's
- 1:05:28basically night and day, you can actually drive on this and this
- 1:05:33took some time and some engineering and incredible work
- 1:05:36from the AI team to actually get this to work and deploy and make
- 1:05:39it efficient in the car. This also improved a lot of
- 1:05:44other object detection. So for example, here, in this
- 1:05:46video I'm showing single-camera predictions in Orange and
- 1:05:49multi-camera predictions in blue.
- 1:05:51And basically, if you if you can't predict these cars, if you
- 1:05:54are only seeing a tiny sliver of a car so your detections are not
- 1:05:57going to be very good and their positions are not going to be
- 1:05:59good but a multi-camera network does not have an issue.
- 1:06:02Here's another video from a more nominal sort of situation and we
- 1:06:05see that as these cars in this tight space cross camera
- 1:06:09boundaries, there's a lot of Jank that enters into the
- 1:06:11predictions and Really the whole setup just doesn't make sense
- 1:06:14especially for very large Vehicles like this one and we
- 1:06:16can see that the multi-camera networks struggle significantly
- 1:06:19less with these kinds of predictions.
- 1:06:22Okay, so at this point, we have multi-camera networks and
- 1:06:24they're giving predictions directly in Vector space, but we
- 1:06:26are still operating at every single instant in time
- 1:06:29completely independently. So very quickly, we discovered
- 1:06:32that there's a large number of predictions.
- 1:06:33We want to make that actually required the video context, and
- 1:06:36we need to somehow figure out how to feed this into the net.
- 1:06:39So, in particular, is this car parked or not?
- 1:06:41Is it moving? How fast is it moving?
- 1:06:43Is it still there? Even though it's temporarily
- 1:06:44occluded? Or for example, if I'm trying to
- 1:06:46predict the road geometry ahead, it's very helpful to know if the
- 1:06:50signs or the road. Markings that I saw 50 meters
- 1:06:52ago. So we try to develop, we try to
- 1:06:57insert video modules into our neural network architecture, and
- 1:06:59this is kind of one of the solutions that we converged on.
- 1:07:01So we have the multi scale features as we had them from
- 1:07:03before and what we are going to now, insert is a feature Q
- 1:07:06module, that is going to cash some of these features over time
- 1:07:10and then a video module that is going to diffuse this
- 1:07:12information to poorly and then we're going to continue into the
- 1:07:16heads that do the decoding. Now I'm going to go into both of
- 1:07:19these blocks, one by one, Also, in addition, notice here that we
- 1:07:23are also feeding in the kinematics.
- 1:07:24This is basically the velocity and acceleration that's telling
- 1:07:27us about how the car is moving. So, not only are not only are we
- 1:07:30going to keep track of what we're seeing from all the
- 1:07:32cameras, but also how the car has traveled.
- 1:07:36So here's the feature, cute and the rough layout of it.
- 1:07:38We are basically concatenating these features over time and the
- 1:07:42kinematics of how the car has moved and the positional
- 1:07:45encodings and that's been concatenated encoded and stored
- 1:07:48in the feature q, and that's going to be consumed by video
- 1:07:51module. Now there's a few details here
- 1:07:52again to get right? So in particular, with respect
- 1:07:55to the pop and push mechanisms, and when do you push and how?
- 1:07:58And especially when do you push basically?
- 1:08:01So here's a cartoon diagram illustrating, some of the
- 1:08:04challenges here. There's going to be a the ego
- 1:08:07cars coming from the bottom and coming up to this intersection
- 1:08:09here and then traffic is going to start Crossing in front of us
- 1:08:13and it's going to temporarily start including some of the cars
- 1:08:15ahead. And then we're going to be stuck
- 1:08:17at this intersection for a while and just waiting our turn.
- 1:08:19This is something that happens all the time and is a cartoon
- 1:08:21representation of some of the challenges here.
- 1:08:24So number 1, with respect to the feature q, and when we want to
- 1:08:27push into a queue, obviously, we'd like to have some kind of a
- 1:08:29time-based queue. Or, for example, we enter the
- 1:08:32features into the queue say every twenty, seven
- 1:08:34milliseconds, and If a car gets temporarily included, then the
- 1:08:38neural network. Now, has the power to be able to
- 1:08:40look and reference the memory in time and learn the association
- 1:08:44that. Hey, even though the same look
- 1:08:45secluded right now, there's a record of it in my previous
- 1:08:48features and I can I use this to still make a detection.
- 1:08:52So that's kind of like the more obvious one, but the one that we
- 1:08:53also discovered is necessary in our case is, for example,
- 1:08:57suppose you're trying to make predictions about the road
- 1:09:00surface, and the road geometry ahead, and you're trying to
- 1:09:03predict that I'm in the turning lane and the Laying next to us
- 1:09:06is going straight, then it's really necessary to know about
- 1:09:10the line markings and the signs and sometimes they occur long
- 1:09:13time ago. And so if you only have a
- 1:09:15time-based Q, you may forget the features while you're waiting at
- 1:09:19your red light. So in addition to a time-based
- 1:09:21key, we also have a space-based Q.
- 1:09:23So we push every time the car travels with certain fixed
- 1:09:26distance. So some of these details
- 1:09:28actually can matter quite a bit. And so, in this case, we have a
- 1:09:30time-based Q&A space Q 2 feet to cash our features and that
- 1:09:33continues into the video module. Now for the video module, we
- 1:09:37looked at a number of possibilities of how to fuse
- 1:09:39this information temporally. So we looked at three
- 1:09:42dimensional convolutions Transformers, axial Transformers
- 1:09:45in an effort to try to make them more, efficient recurrent
- 1:09:47neural, networks of a large number of flavors, but the one
- 1:09:50that we actually like quite a bit as well.
- 1:09:51And I want to spend some time on is a spatial recurrent neural
- 1:09:55network video module. And so, what we're doing here is
- 1:09:59because of the structure of the problem, we're driving on
- 1:10:01two-dimensional surfaces. We can actually organize the
- 1:10:03hidden State into a two-dimensional.
- 1:10:06And then as the cars driving around, we update only the parts
- 1:10:08that are near the car and where the car has visibility.
- 1:10:11So as the car is driving around, we are using the kinematics to
- 1:10:15integrate the position of the car and the hidden features
- 1:10:18grid. And we are only updating the RNN
- 1:10:22at the points, where we visit, where we have, that are nearby
- 1:10:24us sort of. So here's an example of what
- 1:10:28that looks like. Here what I'm going to show you
- 1:10:32is the car driving around and we're looking at the hidden
- 1:10:36state of this RNN. And these are different channels
- 1:10:42in the hidden state. So you can see that this is
- 1:10:44after optimization and training the neural net.
- 1:10:46You can see that some of the channels are keeping track of
- 1:10:48different aspects of the road. Like, for example, the centers
- 1:10:51of the road, the edges, the lines, the road surface and so
- 1:10:54on, here's another cool video of this.
- 1:10:57So this is looking at the mean of the first ten channels in the
- 1:11:00hidden state for Different traversals of different
- 1:11:04intersections and all I want you to see basically is that there's
- 1:11:07cool activity as the recurrent neural network is keeping track
- 1:11:10of what's happening at any point in time.
- 1:11:11And you can imagine that we've now given the power to the
- 1:11:13neural network to actually selectively read and write to
- 1:11:16this memory. So, for example, if there's a
- 1:11:18car right next to us and is including some parts of the
- 1:11:20road, then now the network has a has the ability to not write to
- 1:11:23those locations. But when the car goes away and
- 1:11:25we have a really good view, then the recurring your lot can say,
- 1:11:28okay, we have very clear visibility, we definitely want
- 1:11:30to write information. About what's in that part of
- 1:11:32space. Here's a few predictions that
- 1:11:35show what this looks like. So here we are making
- 1:11:40predictions about the road boundaries and red intersection
- 1:11:42areas and blue Road centers and so on.
- 1:11:45So we're only showing few of the predictions here, just to keep
- 1:11:48the visualization clean. And yeah, this is this done by
- 1:11:53the spatial RNN and this is only showing a single clip single
- 1:11:57traversal but you can imagine there could be multiple trips
- 1:12:00through here. And basically number of cars and
- 1:12:02number of Clips could be collaborating to build this map
- 1:12:05basically in effectively and HD map.
- 1:12:07Except it's not in the space of explicit items.
- 1:12:10It's in a space of features of a recurrent neural network which
- 1:12:13is kind of cool. I haven't seen that before.
- 1:12:17The video networks also improved our object detection quite a
- 1:12:20bit. So in this example, I want to
- 1:12:22show you a case where there are two cars over there and one car
- 1:12:26is going to drive by and include them briefly.
- 1:12:28So, look at what's happening with the single frame of the
- 1:12:29video predictions as the cars pass in front of us.
- 1:12:37Yeah. So that makes a lot of sense.
- 1:12:39So, a quick play through, what's happening when both of them are
- 1:12:43in view, the predictions are, roughly equivalent, and you are
- 1:12:46seeing multiple orange boxes because they're coming from
- 1:12:48different cameras when they are occluded.
- 1:12:53The single frame networks, drop the detection about the video
- 1:12:55module. Remembers it and we can persist
- 1:12:57the cars. And then when they are only
- 1:12:59partially occluded, the single frame that work is forced to
- 1:13:02make its best guess about what it's seeing, and it's forced to
- 1:13:05make a prediction and it makes a really terrible.
- 1:13:07Abel prediction but the video module knows that there's only a
- 1:13:09partial that you know it has the information and knows that this
- 1:13:13is not a very easily visible part right now and doesn't
- 1:13:17actually take that into account. We also saw significant
- 1:13:20improvements in our ability to estimate depth and of course
- 1:13:22especially velocity. So here I'm showing a clip from
- 1:13:25our remove, the radar, push, where we are seeing the radar
- 1:13:28depth and velocity in green. And we were trying to match or
- 1:13:32even surpass of course the signal just from video networks
- 1:13:35alone. And what you're seeing here is
- 1:13:37Is in Orange. We are seeing single frame
- 1:13:40performance and in blue, we are seeing again, video modules, and
- 1:13:44so, you see that the quality of depth is much higher and for
- 1:13:46velocity the orange signal. Of course, you can't get
- 1:13:49velocity out of a single frame Network.
- 1:13:50So we use, we just differentiate up to get that, but the video
- 1:13:53module actually is basically right on top of the radar
- 1:13:56signal. And so we found that this works
- 1:13:58extremely well for us. So here's putting everything
- 1:14:02together. This is what our architectural
- 1:14:04roughly looks like today. So we have raw images feeding on
- 1:14:08the bottom. They go through rectification
- 1:14:10layer to correct for camera, calibration and put everything
- 1:14:13into a common virtual camera. We pass them through reg Nets,
- 1:14:17residual networks to process them into a number of features
- 1:14:20at different scales. We fuse the multiscale
- 1:14:22information with by fpn this goes through a Transformer
- 1:14:25module to re represent it into the vector space in the output
- 1:14:28space. This feeds into a feature Q in
- 1:14:30time, or Space that gets processed by a video module like
- 1:14:33the spatial RNN and then continues into the branching
- 1:14:35structure of the Hydra net with Trunks and heads for all the
- 1:14:38different tasks and so that's the architecture roughly what it
- 1:14:42looks like today. And on the right, you are seeing
- 1:14:43some of its predictions for visualize both in a top-down
- 1:14:46Vector space and also in images. So, definitely this architecture
- 1:14:52has definitely complexified from just very simple image based.
- 1:14:55Single Network about three or four years ago, and continues to
- 1:14:57evolve is definitely quite impressive.
- 1:14:59Now, they're still up Panties for improvements that the team
- 1:15:01is actively working on. For example, you'll notice that
- 1:15:04our Fusion of time and space is fairly late in neural network
- 1:15:07terms so maybe we can actually do earlier Fusion of space or
- 1:15:10time and do for example cost volumes or Optical flow like
- 1:15:13networks on the bottom. Or for example, our outputs are
- 1:15:16dense rosters and it's actually pretty expensive to post
- 1:15:19process. Some of these dense rosters in
- 1:15:20the car and of course we are under very strict latency
- 1:15:23requirements. So this is not ideal.
- 1:15:24We actually are looking into what kinds of ways of predicting
- 1:15:27just a sparse structure of the road maybe like, you know, Point
- 1:15:30by point. Point or in some other fashion
- 1:15:31that is it doesn't require expensive post-processing.
- 1:15:36But this basically is how you achieve, very nice extra space
- 1:15:39and now I believe Kasich is going to talk about how we can
- 1:15:42run playing control tablet. Thank you, Andre.
- 1:15:54Hi everyone. My name is Ashok.
- 1:15:55I read the planning and controls auto-leveling and simulation
- 1:15:58teams. The like on to mention the
- 1:16:02vision networks take dense video data and then compress it down
- 1:16:05into a 3D Vector space. The role of the panel now is to
- 1:16:08consider this Vector space and get the car to the destination
- 1:16:11while maximizing the safety comfort and efficiency of the
- 1:16:14car. Even back in 2019 or plan of
- 1:16:17pretty capable driver. It was able to stay in the lens,
- 1:16:20make Lane changes as necessary and take exits of the highway,
- 1:16:23but CDC driving is much more complicated rally.
- 1:16:27There are such a lane lines. Make us do much more free from
- 1:16:31driving. The car has to respond to all of
- 1:16:34curtains and Crossing vehicles and pedestrians doing funny
- 1:16:38things. What is the key problem in
- 1:16:42planning? Merv on the action space is very
- 1:16:46non convex and number two, it is high-dimensional.
- 1:16:51What I mean by non convex is that can be multiple possible
- 1:16:56solutions, they can be independently good.
- 1:16:58But getting a globally, consistent solution is pretty
- 1:17:01tricky so they can be pockets of local Minima that the planet can
- 1:17:04start get stuck into. And secondly, the
- 1:17:07high-dimensional becomes pick as the car needs to plan for the
- 1:17:10next 10 to 15 seconds and introduce the position velocity
- 1:17:14and acceleration or the center window.
- 1:17:16This is a lot of parameters to produce at runtime.
- 1:17:20Discrete search methods are really great at solving non
- 1:17:23convex problems because they are discrete.
- 1:17:24They can they don't get stuck in local Minima well as continuous
- 1:17:27function, optimization can easily get stuck in local Minima
- 1:17:30and pretty sport solutions that are not great.
- 1:17:34On the other hand, for high dimensional problems.
- 1:17:36Discrete search sucks because of the discrete distribution, a
- 1:17:41great information. So literally as to go and
- 1:17:43explore. Each point know how good it is.
- 1:17:46Whereas countries, optimization is getting best methods to very
- 1:17:48quickly go to a good solution. Our solution to this little
- 1:17:53problem, is to break it down. Here are quickly.
- 1:17:55First, use a code search method to Crunch down the non
- 1:17:59convexity, and come up with a car next Corridor.
- 1:18:02And then use continuous optimization techniques to make
- 1:18:04the final smooth trajectory Let's see an example of how the
- 1:18:09search operates. So here we are trying to do a
- 1:18:13lane change. In this case, the car needs to
- 1:18:16do two back-to-back Lane changes to make the left turn up ahead.
- 1:18:21For this week are searches over different maneuvers.
- 1:18:27So in the first, the first one it's such as is Lane change,
- 1:18:30that's close by, but the car brakes pretty harshly, so it's
- 1:18:34pretty uncomfortable. The next maneuver tries does the
- 1:18:39lane change bit late so it speeds up goes by in the other
- 1:18:41course, in front of the other cars and find us a lane change.
- 1:18:44But now it risks missing the left turn We do thousands of
- 1:18:49searches in a very short time span because these are all
- 1:18:53physics based models. These features are very easy to
- 1:18:55simulate. And in the end we have a set of
- 1:18:58candidates and we finally choose one based on the optimality
- 1:19:00conditions of safety, comfort, and easily making the turn.
- 1:19:05So now the car has chosen this path and you can see that as the
- 1:19:08car executes its trajectory it pretty much matches.
- 1:19:11What we had planned, the cyan plot on the right side here.
- 1:19:14That one is the actual velocity of the car and the white line be
- 1:19:18underneath it is, was the plan. So we are able to plan for 10
- 1:19:22seconds here and able to match that when you see in hindsight.
- 1:19:26So this is a well-made plan. When driving the onset of their
- 1:19:31agents. It's important to not just plan
- 1:19:33for ourselves, but instead we have to plan for everyone
- 1:19:36jointly and optimized for the overall scenes traffic flow.
- 1:19:41Not to do this. What we do is we literally run
- 1:19:43the autopilot Planner on every single relevant object in the
- 1:19:45scene. Here's an example of why that's
- 1:19:48necessary. This is a narrow Corridor.
- 1:19:51I let you watch the video for a second.
- 1:20:07Yeah that was automated driving a narrow Corridor.
- 1:20:09Going around Park cars cones and poles here.
- 1:20:12This is 3D view of the same thing.
- 1:20:14The oncoming car is now and autopilot slows down a little
- 1:20:17bit but then realizes that we cannot yield to them because we
- 1:20:19don't have any space to our side but the other car can yield to
- 1:20:22us instead. So instead of just blindly
- 1:20:24breaking here out of a recent about that car has low enough
- 1:20:29Frosty that they can pull over and should yield to us because
- 1:20:31we can only look to them and assertively makes progress.
- 1:20:36A second oncoming car arrives. Now, this vehicle has higher
- 1:20:39velocity and like I said, earlier will literally run the
- 1:20:42autopilot planner for the other object.
- 1:20:44So, in this case, we're on the panel for them that objects plan
- 1:20:47now goes around their power, their sides Park cars.
- 1:20:50And then, after they pass the power cores, goes back to the
- 1:20:53right side of the road for them. Since we don't know what's in
- 1:20:56the mind of the driver, we actually have multiple possible
- 1:20:59futures for this car here. One future shown in red.
- 1:21:02The other one is shown is green, the green.
- 1:21:04One is a plan that eels to us but since this object velocity
- 1:21:07and acceleration are pretty high, we don't think that this
- 1:21:10person is going to yield to us and they actually want to go
- 1:21:12around the sport cars. So autopilot decides that okay I
- 1:21:15have space here, this person definitely going to come.
- 1:21:17So I'm going to pull over So, as autopilot spoiling our, we
- 1:21:22notice that that car has chosen to appeal to US, based on their
- 1:21:25your red and the acceleration not applied immediately changed
- 1:21:28his mind and continues to make progress.
- 1:21:31This is why we need to plan for everyone because otherwise we
- 1:21:33wouldn't know that this person is going to go on the other part
- 1:21:35course, and come back to their side.
- 1:21:38If you didn't do this auto power will be too timid and would not
- 1:21:40be a practical self driving car. So now we saw how the search and
- 1:21:46planning for other people set up connects Valley.
- 1:21:49Finally, we do a continuous optimization to produce the
- 1:21:52final trajectory that the panel needs to take here.
- 1:21:55The grave thing is the connects Corridor and we initialize a
- 1:21:59spline in heading and acceleration parameter is over
- 1:22:02the arc length of the plan. And you can see that the
- 1:22:05constellation continuously makes fine-grained changes to reduce
- 1:22:08all of its cost. Some of the costs, for example,
- 1:22:10or distance from obstacles, traversal time and Comfort for
- 1:22:15comfort. You can see that the lateral
- 1:22:17acceleration plots on the right. I have nice trapezoidal shapes,
- 1:22:20it's going to come for ya here. On the right side, the green
- 1:22:22plot is a nice trapezoidal shape, and if you recorded human
- 1:22:25trajectory, this is pretty much how it look like the letter joke
- 1:22:28is also minimized. So in summary, we do a search
- 1:22:32for both us and everyone else in the scene we set up a contact
- 1:22:35Corridor and then optimize for a smooth path together.
- 1:22:38This can be some really neat things, like shown above But
- 1:22:43driving looks a bit different in other places like, where I grew
- 1:22:46up from, it's very much more unstructured.
- 1:22:51Cars and pedestrians cutting each other harsh, braking
- 1:22:54honking. It's a crazy world.
- 1:22:58We can try to scale up these methods, but it's going to be
- 1:23:00really difficult to efficiently solve this at runtime.
- 1:23:03What we instead want to do is use learning based methods to
- 1:23:06efficiently solve them and I want to show why this is true.
- 1:23:10So we're going to go from this complicated problem to a much
- 1:23:12simpler toy product. Any problem, but still
- 1:23:14illustrates the core of the issue.
- 1:23:17Here, this is a parking lot. The ego cars in blue and H2 Park
- 1:23:21in the green parking spot here. So it needs to go around the
- 1:23:23curves, the power cores and the cones shown in Orange here.
- 1:23:28As the simple bass line, it's a star, the standard algorithm
- 1:23:31that uses a lot of space search and the heuristic here is a
- 1:23:35distance euclidean distance to the goal.
- 1:23:38So you can see that it directly shoots towards the goal, but
- 1:23:40very quickly gets trapped in a local Minima and it back tracks
- 1:23:43from there, and then such as a different path to try to go
- 1:23:46around this park or eventually it makes progress and gets to
- 1:23:50the goal, but ends up using 400,000 notes for making this.
- 1:23:55Obviously this is a terrible you re sick.
- 1:23:57We want to do better than this. So if error navigation router
- 1:24:02and has the code to for the navigation route, while being
- 1:24:04close to the goal, this is what happens.
- 1:24:08The navigation helps immediately but still when you enter
- 1:24:11encounters cones or other obstacles, it basically do the
- 1:24:16same thing as before backtracks and then searches are all New
- 1:24:18Path and the support search has no idea that these obstacles
- 1:24:22exist. It literally has to go there
- 1:24:24check if it's in Coalition and if it's in collusion, back up
- 1:24:28the navigation, arrows to help but still this took 22 thousand
- 1:24:31nodes. We can design more and more of
- 1:24:34these heuristics to help the search make go faster and
- 1:24:37faster, but it's really tedious and hard to design a globally.
- 1:24:41Optimal heuristic, even if he had a distance function from the
- 1:24:45cones that guided the search, this would not, this is not,
- 1:24:48this is only be effective for a single cone, but what we need is
- 1:24:50a global global value function. So instead, what we want to use
- 1:24:53is neural networks to give this heuristic for us.
- 1:24:56The division networks produces, a vector space, and we have cars
- 1:24:59moving around in it. It's basically looks like Atari
- 1:25:02game and its multiplayer version.
- 1:25:05So we can use techniques such as Mu 0, Alpha 0 etcetera, that was
- 1:25:08used to solve go and other Atari games to solve the same problem.
- 1:25:12So we're working on neural networks that can produce State
- 1:25:14and action distributions, that can then be plugged into Monte
- 1:25:17Carlo, tree search with various cost functions, some of the cost
- 1:25:20functions, can be explicit cost functions.
- 1:25:22Like, this is to collisions Comfort reversal, time, Etc.
- 1:25:25But they can also be Interventions from the actual
- 1:25:28panel driving events. We train such a network for the
- 1:25:32simple parking problem. So here again, same problem.
- 1:25:35Let's see how MCTS tree search does.
- 1:25:43Here you notice that the panel is basically able to in one
- 1:25:46shot. Make progress towards the goal
- 1:25:49to note is that this is not even humorous.
- 1:25:50Using a navigation eristic, just given the scene, the panel is
- 1:25:54able to go directly towards the goal or the other offshoots are
- 1:25:57seeing or possible options. It's not using any of them just
- 1:26:00using the option that directly takes it towards the goal.
- 1:26:03The reason is that the neural network is able to absorb the
- 1:26:05global context of the scene and then produce a very function
- 1:26:08that effectively guides it towards the global Minima as
- 1:26:10opposed to getting in certain any local Minima.
- 1:26:14So this one, it takes 288 nodes and several orders of magnitude
- 1:26:16less than what was done in the a star with the accordion distance
- 1:26:20realistic. So this is what a final
- 1:26:23architecture is going to look like the vision system is going
- 1:26:26to crush down the dance video, Raider into a vector space.
- 1:26:29It's going to be consumed by both a special planner and a new
- 1:26:31network. Planner in addition, to this,
- 1:26:33the network panel can also consume intermediate features of
- 1:26:35the network together, this producer rejected distribution
- 1:26:39and you can be optimized in to end both with explicit cost
- 1:26:42functions and human intervention.
- 1:26:44And other limitation data this then goes into explicit planning
- 1:26:47function. That does whatever is easy for
- 1:26:50that and produces the final state.
- 1:26:52And acceleration commands for the car.
- 1:26:56With that, we need to now explain how we train these
- 1:26:59networks and for training center works.
- 1:27:01We need large data sets and went on right to speak, briefly about
- 1:27:06menu labeling. Yes.
- 1:27:16So the data, the story of data sets is critical, of course, so
- 1:27:18far. We've talked only about neural
- 1:27:20networks, but neural networks only establish an upper bound on
- 1:27:23your performance. Many of these neural networks,
- 1:27:25they have hundreds of millions of parameters and these hundreds
- 1:27:28of millions of parameters, they have to be set correctly.
- 1:27:31If you have a bad setting of parameters is not going to work.
- 1:27:34So neural networks are just an upper bound.
- 1:27:35You also need massive data sets to actually train the correct
- 1:27:38algorithms inside them. Now, in particular, I mentioned,
- 1:27:41we want the data sets directly in the vector space.
- 1:27:43And so, the really the question becomes, how can you accumulate
- 1:27:46because our networks have hundreds millions of parameters?
- 1:27:48How do you accumulate millions and millions of vector space
- 1:27:51examples that are clean and diverse to actually train these
- 1:27:54neural networks effectively now. So there's a story of data sets
- 1:27:58and how they've evolved on the side of all, of the models in
- 1:28:02developments that we've achieved, Now in particular,
- 1:28:06when I joined roughly four years ago, we were working with a
- 1:28:08third party to obtain a lot of our data sets.
- 1:28:11Now unfortunately, we found very quickly that working with a
- 1:28:13third party to get data sets for something this critical, which
- 1:28:16is not going to cut it at the latency of working.
- 1:28:18With the third party was extremely high.
- 1:28:20And honestly, the quality was not amazing.
- 1:28:23And so in the spirit of full vertical integration at Tesla,
- 1:28:26we brought all of the labeling in-house and so over time, we've
- 1:28:30grown more than 1,000 person data, labeling org that is full
- 1:28:35of Professional labor leaders who are working very closely
- 1:28:37with the engineers. So actually, they're here in the
- 1:28:39US and co-located with the engineers here in Bay area, as
- 1:28:42well as we work, very closely with them.
- 1:28:44And we also build all the infrastructure for them from
- 1:28:47scratch ourselves. So we have a team, we are going
- 1:28:50to meet later today that develops and maintains, all of
- 1:28:53this infrastructure for data, labeling.
- 1:28:54And so here for example, I'm showing some of the screenshots
- 1:28:56of some of the latency throughput and quality
- 1:28:59statistics that we maintain about all of the labeling
- 1:29:01workflows and the individual people involved.
- 1:29:04And All the tasks and how the numbers of labels are growing
- 1:29:07over time. So we found this to be quite
- 1:29:11critical and we're very proud of this.
- 1:29:13Now, in the beginning, roughly three or four years ago, most of
- 1:29:16our labeling was in image space. And so you can imagine that this
- 1:29:20is taking quite some time to annotate an image like this.
- 1:29:23And this is what it looked like where we are sort of drawing,
- 1:29:25polygons and polylines on top of on top of these single
- 1:29:29individual images. As I mentioned, we need millions
- 1:29:31of vector space labels. So this is not going to cut it.
- 1:29:34So very he quickly we graduated to three dimensional or four
- 1:29:38dimensional labeling, where we are directly labeling in Vector
- 1:29:41space. Not an individual images.
- 1:29:43So here, what I'm showing Is a clip and you are seeing a very
- 1:29:47small reconstruction. You're about to see a lot more
- 1:29:50reconstructions soon, but it's very small reconstruction of the
- 1:29:52ground plane on which the car drove and a little bit of the
- 1:29:55point Cloud here, that was reconstructed.
- 1:29:57And what you're seeing here is that the labeler is changing,
- 1:30:01the labels directly in Vector space.
- 1:30:03And then we are re projecting those changes into camera
- 1:30:06images. So we're labeling directly in
- 1:30:08Vector space. And this gave us a massive
- 1:30:10increase in throughput for a lot of our labels because your label
- 1:30:12once in 3D and then you get to reproject But even this we
- 1:30:17realized was actually not going to cut it.
- 1:30:21So because people and computers have different pros and cons.
- 1:30:24So people are extremely good at things like semantics.
- 1:30:26But computers are very good at geometry reconstruction,
- 1:30:30triangulation tracking and so really for us, it's much more
- 1:30:33becoming a story of how do humans and computers collaborate
- 1:30:36to actually create these Vector space data sets.
- 1:30:38And so we're going to now talk about Auto labeling which is
- 1:30:41some of the infrastructure we've developed for labeling these
- 1:30:43clips at scale. Hi again.
- 1:30:53So, you know, we have lots of human laborers, the amount of
- 1:30:55training that are needed for training, the network
- 1:30:57significantly out a them. So we try to invest in a
- 1:31:00massive, auto-leveling pipeline. Here's an example of how we
- 1:31:04label a single clip. A clip is a entity that has
- 1:31:07dense sensor data like videos. I made a GPS automatically Etc.
- 1:31:12This can be 45 seconds to a minute long.
- 1:31:14These can be afforded by our own engineering course, or from
- 1:31:16customer cars. We collect this clips and then
- 1:31:19send them to service, where we run a lot of neural networks
- 1:31:23offline to produce intermediate results and segmentation massdep
- 1:31:27Point matching Etc. This and goes to a lot of
- 1:31:29Robotics and AI algorithms, pretty the final set of labels
- 1:31:32that can be used to train the networks.
- 1:31:36Now the first task we want to label is the road surface.
- 1:31:40Typically, we can use splines or measures to represent Road
- 1:31:42surface. But those are because of the
- 1:31:44territorial restrictions are not differentiable and not amenable
- 1:31:47to producing this. So, what we do instead is in the
- 1:31:49style of neural Radiance Fields work from last year, which is
- 1:31:52quite popular. So we use an implicit
- 1:31:54representation to represent the road surface.
- 1:31:57Here, we are querying X Y points on the ground and asking for a
- 1:32:00network to predict the height of the ground surface along with
- 1:32:04various semantics such as curves.
- 1:32:06And boundaries, Road surface travel space, Etc.
- 1:32:10So, even a single X Y. We get Z together.
- 1:32:13This make a 3D point and they can be represented in to all the
- 1:32:16camera views. So we make millions of search
- 1:32:19queries and get lots of points. These points are represented in
- 1:32:22to all the camera views in we are showing on the top right
- 1:32:26here. One, such camera image, with all
- 1:32:28these points. Three projected.
- 1:32:30Now, we can compare this point represented point with the image
- 1:32:35space prediction of the Imitations and joined the
- 1:32:38optimizing this or all the camera views was across space
- 1:32:41and time. For is the next term
- 1:32:42reconstruction. Here's an example of how the
- 1:32:46looks like. So here, this is an optimist
- 1:32:48Road surface that reproduction to the 8 cameras that the car
- 1:32:51has and across all of time. And you can see how it's
- 1:32:53consistent across both space and time.
- 1:32:59So a single car driving through some location.
- 1:33:01Can sweep out some patch around the trajectory using this
- 1:33:04technique. But we don't have to stop there.
- 1:33:08So here, we collect. Collect a different clips from
- 1:33:12the same location from different cars.
- 1:33:13Maybe. And each of them sweeps out some
- 1:33:16part of the road. Whole thing is, we can bring
- 1:33:19them all together into a single giant organization.
- 1:33:22So here, the 16 different trips are organized using aligned
- 1:33:26using various features such as Rogers Lane lines.
- 1:33:29All of them should agree with each other and also agree with
- 1:33:32all of their image space observations together.
- 1:33:34This is this produce an effective way to label the road
- 1:33:37surface, not just where the Cod Roe but also in other locations
- 1:33:40that it hasn't driven it. Again, the point of this is not
- 1:33:43to just build HD Maps or anything like that.
- 1:33:45It's only to label the clips through this intersections so we
- 1:33:48don't have to maintain them forever as long as the labels
- 1:33:51are consistent, with the videos that they were collected at.
- 1:33:55Optionally than humans can come on top of this and clean up any
- 1:33:57noise or add additional metadata to make it even richer.
- 1:34:03We don't have to stop at just the road surface.
- 1:34:05We can also orbit and reconstructed e, static
- 1:34:07obstacles here. This is reconstructed. 3D Point
- 1:34:11cloud from our cameras. The main Innovation here is the
- 1:34:16density of the point Cloud. Typically these points record
- 1:34:18texture to form associations from one frame to the next
- 1:34:22frame. But here we are able to produce
- 1:34:23these points even on textual or surfaces like the road surface
- 1:34:26or walls. And this is really useful to
- 1:34:28annotate arbitrary obstacles that we can see on the scene.
- 1:34:33World. For more cool advantage of doing
- 1:34:38all of this on server, on the servers offline, is that we have
- 1:34:41the benefit of hindsight. This is a super useful hack
- 1:34:44because say in the car, then the network needs to produce the
- 1:34:47velocity. He just has to use the
- 1:34:49historical information and guess what?
- 1:34:51The velocity is. But here we can look at both the
- 1:34:55history, but also the future and basically cheat and get the
- 1:34:58correct answer of the kinematics, like, velocity
- 1:35:01acceleration, Etc. One more Advantage is that we
- 1:35:04can have different tracks, but Can stitch them together, even
- 1:35:07through occlusions because we know the future, we are future
- 1:35:09tracks, we can match them and then associate them.
- 1:35:12So here, you can see the pedestrians on the other side of
- 1:35:14the road or persisted, even through multiple occlusions buy
- 1:35:17these cars, this is really important for the planner
- 1:35:20because the Pendleton, you know, if it's if it's awesome one it
- 1:35:23still needs to account for them even then they are occluded.
- 1:35:26So this is a massive advantage. Combining everything together.
- 1:35:32We can produce these amazing data sets that annotate all of
- 1:35:35the road texture or the static objects and out of the moving
- 1:35:39objects, even through occlusions producing, excellent kinematic
- 1:35:43labels, all you can see how the cards turn smoothly produce,
- 1:35:47really smooth labels or the predecessor.
- 1:35:49Consistently attract the power cores of this to 0 velocity so
- 1:35:53we can also know that they are part.
- 1:35:55So this is huge for us. This is one more example, of the
- 1:35:58same thing. You can see how everything is
- 1:36:01consistent. They want to produce a million
- 1:36:03such label clips and train or video multicam.
- 1:36:08Video networks with such large data set, and really crush this
- 1:36:11problem. We want to get the same view
- 1:36:13that is consistent. There are seeing here in the
- 1:36:15car. We started our first expression
- 1:36:19of this with the removed later project, we removed it in a very
- 1:36:22short time span I think within three months in the early days
- 1:36:26of the network, we noticed, for example, in low visibility
- 1:36:28conditions, the network can suffer understandably because
- 1:36:32obviously this truck just dumped a bunch of snow on as and it's
- 1:36:34really hard to see. But we should still remember
- 1:36:36that this car was in front of us.
- 1:36:39But on networks, early on did not do this because of the lack
- 1:36:42of data in such conditions. So what we did, we add the free
- 1:36:46to produce lots of similar clips and the feed responded it did.
- 1:36:51So it produces Play. Yeah, it produces lots of video
- 1:37:01clips where shits falling out of all other vehicles and we send
- 1:37:05this throttling pipeline that was able to label 10K Clips in
- 1:37:08within a week. This would have taken several
- 1:37:10months with humans labeling every single clip here.
- 1:37:14So we did this for two hundred different conditions and we were
- 1:37:18able to very quickly create large data sets and that's how
- 1:37:20we're able to remove this. So once we trained, the
- 1:37:23Network's with this data, you can see that it's totally
- 1:37:26working and keeps the memory that the subject was there and
- 1:37:31provides this. So finally, we wanted to
- 1:37:34actually get a cyber truck into data set for remove the radar.
- 1:37:38Can you all guess where we got this clip from?
- 1:37:41I'll give you a moment. Someone said it, yes, yes, it's
- 1:37:47render itself, simulation. It was hard for me to tell
- 1:37:50initially and I if I may, if I may say so myself, it looks
- 1:37:52pretty. It looks really pretty.
- 1:37:56So yeah, in addition to Auto leveling, we also invest heavily
- 1:37:59in using simulation for labeling our data.
- 1:38:03So this is the same scene as seen before, but from a
- 1:38:07different camera angle. So a few things that I wanted to
- 1:38:12point out, for example, the ground surface, it's not plain
- 1:38:15as fault, the lots of cars and cracks and door.
- 1:38:19Seems the some patch work done on top of it.
- 1:38:22Vehicles, move, realistically. The truck is articulated.
- 1:38:25Even ghosts of the curve and makes a right turn the other
- 1:38:28cars, behave smartly. They avoid collisions, go around
- 1:38:31curves, and also, smooth and actual date, smooth brake and
- 1:38:34accelerate smoothly. The car here with the logo on
- 1:38:39the top is Auto by actually Autobot is driving that car and
- 1:38:41it's making a left hand. And since this simulation, it
- 1:38:46starts from the vector space so it has perfect labels.
- 1:38:49Here we show a few of the labels that we produce, these are
- 1:38:52vehicle cuboids with kinematics, depth, surface normals
- 1:38:56segmentation. But Andre can name a new task
- 1:39:00that he wants next week and we can very quickly produce this
- 1:39:02because we already have the vector space, then we can write
- 1:39:05the code to produce this labels very very quickly.
- 1:39:09So when the simulation help, it helps number one.
- 1:39:12When read is difficult to source as large as our Fleet is, it can
- 1:39:16still be hard to get some crazy scenes like this couple and
- 1:39:19their dog running on the highway while there are other high-speed
- 1:39:22cards around. This is a pretty rare scene I'd
- 1:39:26say, but still can happen and autopilot still needs to handle
- 1:39:29it when it happens. When data is difficult to label,
- 1:39:33there are hundreds of pressing crossing the road.
- 1:39:35This could be a Manhattan. Downtown people crossing the
- 1:39:38road. It's going to take several hours
- 1:39:39for humans to label this clip and, you know, for automatic
- 1:39:41leveling algorithms. This is really hard to get the
- 1:39:43association bright and it can produce like bad velocities.
- 1:39:46When simulation, this is Trivial because you already have the
- 1:39:48objects, you served like spit out the cuboid in the velocities
- 1:39:52now. So finally, when we introduce
- 1:39:53closed-loop behaviour where the cars needs to be in a terminal
- 1:39:57situation or the letter depends on the actions.
- 1:40:00This is pretty much the only Way to get it, reliably.
- 1:40:04All this is great. What's needed to make this
- 1:40:06happen? Number one, accurate sensor,
- 1:40:11simulation. Again, the point of the
- 1:40:13simulation is not to Just Produce pretty pictures.
- 1:40:16It needs to produce what the camera.
- 1:40:18The car would see in other senses would see.
- 1:40:20So here we are stepping through different exposure settings of
- 1:40:23the real camera on the left side and the simulation on the right
- 1:40:25side. We're able to pretty much match
- 1:40:29what real cameras do. In order to do this, we had to
- 1:40:33Model A lot of the properties of the camera in our sensory,
- 1:40:36stimulation starting from sensor noise, motion blur Optical
- 1:40:40distortions, even Hitler Transmissions even like
- 1:40:45diffraction patterns of the windshield Etc.
- 1:40:48We don't use this, just for the autopilot software.
- 1:40:50We also use it to make hard decisions such as lens design.
- 1:40:53Camera design sensor placement. Even headlight, transmission
- 1:40:56properties, Second, we need to render the visuals in a
- 1:41:04realistic manner, you cannot have what in the game industry
- 1:41:07called jaggies. These are aliasing artifacts
- 1:41:10that are a dead giveaway that this is simulation.
- 1:41:12We don't want them. So we go through a lot of paints
- 1:41:15to produce. Nice spatio.
- 1:41:16Temporal spatial temporal anti-aliasing.
- 1:41:20We also are working on new rendering techniques to make
- 1:41:22this even more realistic. Yeah, in addition we also use
- 1:41:30Ray tracing to produce realistic lighting and Global
- 1:41:32illumination, okay? That's the last of the cop cars.
- 1:41:34I think We obviously cannot have really just four or five cars
- 1:41:40because it will easily over fit because it knows the sizes.
- 1:41:44So we need to have realistic assets, like the Moose on the
- 1:41:46road. Here we have thousands of Assets
- 1:41:49in our library and they can wear different shirts and actually
- 1:41:52can move realistically. So, this is really cool.
- 1:41:55We also have a lot of different locations map and created to
- 1:41:58create this Sim enrollments. We are actually two thousand
- 1:42:01miles of Road built and this is almost the length of the roadway
- 1:42:06from East Coast, the west coast of the United States, which I
- 1:42:08think is pretty cold. In addition, we have built
- 1:42:10efficient tooling to build several miles, more on a single
- 1:42:14day, on a single artist, But this is just the tip of the
- 1:42:19iceberg. Actually most of the data that
- 1:42:21we use to train is created procedurally, using algorithms
- 1:42:25as opposed to artists making these simulation scenarios.
- 1:42:29So these are all professionally created roads with lots of
- 1:42:32parameters, such as curvature, various varying trees cones,
- 1:42:35poles, cards with different velocities and the interaction
- 1:42:38produce an endless stream of data for the network, but a lot
- 1:42:41of this data can be boring because the network in order to
- 1:42:44get it correct. So, what we do is we use also a
- 1:42:46bass techniques to basically put up the network to see where it's
- 1:42:49failing at and create more data around the failure points of the
- 1:42:52network. So this is in closed loop.
- 1:42:54Trying to make the network performance.
- 1:42:56Be better. You don't stop there.
- 1:43:01Actually we want to create recreate any failures that the
- 1:43:04happens to the autopilot in simulation so that we can hold
- 1:43:06autopilot to the same bar from then on.
- 1:43:09So here on the left side, you are seeing a real clip there was
- 1:43:12cotton from a car. It then goes through or
- 1:43:15auto-leveling pipeline to produce a 3D reconstruction of
- 1:43:18the scene along with all the moving objects.
- 1:43:21With this combined, with the original visual information, we
- 1:43:24recreate the same scene synthetically and create a
- 1:43:27simulation scenario entirely out of it.
- 1:43:29So and then we replay autopilot on it out of bed, can do
- 1:43:32entirely new things and we can form new worlds, new outcomes
- 1:43:35from the original failure. This is amazing because we
- 1:43:38really don't want autopilot to fail and when it fails you want
- 1:43:41to capture it and keep it to that bar.
- 1:43:48Not just that. We can actually take the same
- 1:43:50approach that we said earlier and take it one step further.
- 1:43:53We can use new rendering techniques to make it look even
- 1:43:56more realistic. So, we take the original
- 1:43:59question video clip we create a synthetic simulation from it and
- 1:44:03then apply new rendering techniques on top of it.
- 1:44:05And it produces this which looks amazing in my opinion because
- 1:44:08this one is very realistic and looks almost like it was
- 1:44:11captured by the actual cameras. These are results from last
- 1:44:13night because it was cool and we wanted to present it, but yeah,
- 1:44:17it is Yeah. I'm very excited for what sin
- 1:44:19can achieve. This is not all bullshit because
- 1:44:22networks trained in the car already.
- 1:44:24Used simulation data, we used 300 million images with almost
- 1:44:28half a billion labels. And we want to crush down all
- 1:44:30the tasks that are going to come up for the next several months.
- 1:44:35With that. I invite Milan to see explain
- 1:44:37how we scale this operations and really build a label Factory and
- 1:44:41spit out. Millions of labels.
- 1:44:51I think I shook hey everyone I'm Ellen.
- 1:44:54I'm responsible for the integration of our networks in
- 1:44:57the car and for most of our neural network training and
- 1:44:59evaluation infrastructure. And so tonight I just like to
- 1:45:03start by giving you some perspective into the amount of
- 1:45:05compute that's needed to power. This type of data, generation
- 1:45:08Factory. And so in the specific context
- 1:45:11of the push, we went through as a team here, a few months ago to
- 1:45:14get rid of the dependency on the radar sensor, for the Pilot, We
- 1:45:18generated over 10 billion labels Across two and a half million
- 1:45:21clips. And so to do that, we had to
- 1:45:24scale a huge offline, you our networks, and our simulation
- 1:45:28engine across thousands of gpus and just a little bit shy of
- 1:45:3220,000 CPU cores. On top of that.
- 1:45:35We also included over 2,000 actual autopilot for
- 1:45:38self-driving computers in the loop with our simulation engine.
- 1:45:41And that's our smallest compute cluster.
- 1:45:48So I'd like to give you some idea of what it takes to take
- 1:45:51our neural networks and move them in the car.
- 1:45:55And so, the the two main constraints that we're working
- 1:45:58on, their here are mostly latency and frame rate, which
- 1:46:03are very important for safety but also to get proper estimates
- 1:46:07of acceleration and velocity of our surroundings.
- 1:46:11And so, the meat of the problem really is around Dai compiler,
- 1:46:15that we write an extent here within the group.
- 1:46:17That essentially Maps, the computer operations for my iPod,
- 1:46:20Touch model to a set of dedicated accelerated pieces of
- 1:46:25hardware and we do that while figuring out a schedule that's
- 1:46:29optimized for throat, put while working on their severe as from
- 1:46:33constraints. And so, by the way, we're not
- 1:46:36doing that, just on one engine but on across two engines on the
- 1:46:39autopilot computer. And the way we use those engines
- 1:46:41here at Tesla is such that at any given time.
- 1:46:44Only one of them will actually output control commands to the
- 1:46:46vehicle. While the other one is used as
- 1:46:48an extension of compute, but those roles are interchangeable.
- 1:46:52Both of the hardware and software level So, how do we
- 1:46:56tolerate quickly together as a group to this AI development
- 1:46:59Cycles? Well, first, we have been
- 1:47:02skating our capacity to evaluate our software in your network
- 1:47:05dramatically over the past few years.
- 1:47:07And today, we're running over a million evaluations per week on
- 1:47:11any code change. That the team is producing and
- 1:47:15those evaluations runs on over 3,000 actual food.
- 1:47:18So, driving computers that are hooked up together in a
- 1:47:20dedicated cluster. And so on top of this, we've
- 1:47:24been developing really cool debugging tools.
- 1:47:27And so, here is a video of one of our tools which is helping
- 1:47:30developers iterate through the development of neural networks
- 1:47:34and comparing life the outputs from different revisions of a
- 1:47:38senior Network model as reiterating life through a video
- 1:47:41clips. And so last but not least, we've
- 1:47:46been scaling, our neural network training, compute dramatically
- 1:47:49over the past few years and today we're barely shy of 10,000
- 1:47:53gpus which just to give you some sense in terms of number of GPU
- 1:47:58is more than the top five public in all supercomputers in the
- 1:48:00world but that's not enough. And so I'd like to invite Ganesh
- 1:48:05to talk about the next steps. Thank you, Milan.
- 1:48:21My name is Ganesh and I lead project dojo.
- 1:48:26It's an honor to present this project on behalf of the
- 1:48:29multidisciplinary Tesla team that is working on this project.
- 1:48:34As you saw from Milan. There's an insatiable demand for
- 1:48:40Speed as well as capacity for neural network training and Elan
- 1:48:44prefetch this. In a few years back, he asked us
- 1:48:47to design a super fast training computer and that's how we
- 1:48:50started project dojo. Our goal is to achieve best AI
- 1:48:56training performance and support.
- 1:48:57All these larger more complex models.
- 1:49:00That under a steam is dreaming of and B, power, efficient and
- 1:49:05cost effective at the same time. So, we thought about how to
- 1:49:10build this and we came up with a distributed computer
- 1:49:13architecture, after all, all the training computers are, there
- 1:49:18are distributed computers in one form or the other.
- 1:49:21They have compute elements in the Box out here, connected with
- 1:49:25some kind of network. In this case, it's a two
- 1:49:28dimensional network, but it could be any different network
- 1:49:31CPU, GPU accelerators. All of them have compute little
- 1:49:35memory and Network But one thing which is common Trend amongst
- 1:49:41this is it's easy to scale the compute.
- 1:49:45It's very difficult to scale up bandwidth and extremely
- 1:49:49difficult to reduce latency these.
- 1:49:51And you'll see how our design Point cater to that how our
- 1:49:55philosophy addressed. These aspects of traditional
- 1:49:59limits for dojo, we envisioned a large compute plane filled with
- 1:50:06very robust compute elements, backed with large pool of memory
- 1:50:10and interconnected with very high, bandwidth and low, latency
- 1:50:13Fabric. And in a 2d mesh format and onto
- 1:50:18this for extreme scale. Big neural networks will be
- 1:50:21partitioned and mapped to extract different, parallelism
- 1:50:26model, graph, data, parallelism, and then in neural.
- 1:50:30Allure of ours will exploit spatial and temporal locality
- 1:50:37such that it can reduce communication footprint to local
- 1:50:40zones and reduce Global Communication.
- 1:50:43And if we do that, our bandwidth utilization can keep scaling
- 1:50:48with the plain of compute that we desire out here.
- 1:50:54We wanted to attack this all the way, top to the bottom of the
- 1:50:59stack and remove any bottlenecks at any of these levels.
- 1:51:04And let's start this journey in an inside-out fashion, starting
- 1:51:07with the chip. As I described chips have
- 1:51:11compute elements are smallest entity or scale is called a
- 1:51:15training node and the choice of this node is very important to
- 1:51:20ensure seamless scaling. If you go too small, it will run
- 1:51:25fast but the overheads of synchronization will end,
- 1:51:29software will dominate if you pick it too big, it will have
- 1:51:34complexities in implementation in the real hardware and
- 1:51:38ultimately run into A bottle neck issues because we wanted to
- 1:51:42address. You want to address latency and
- 1:51:46bandwidth as our primary optimization Point?
- 1:51:49Let's see how he went about doing this.
- 1:51:52What we did was we picked the farthest distance is signal
- 1:51:56could Traverse in a very clock very high clock cycle.
- 1:51:59The in this case two gigahertz plus and we drew a box around
- 1:52:03it. This is the smallest latency
- 1:52:07that a signal can Traverse one cycle at a very high frequency.
- 1:52:11And then we filled up the box with wires to the This is the
- 1:52:15highest bandwidth, you can feed the box with and then we added
- 1:52:19machine learning computer underneath and then a large pool
- 1:52:22of s RAM and last, but not the least, a programmable core to
- 1:52:27control. And this gave us our high
- 1:52:31performance training node. What this is is a 64-bit
- 1:52:36superscalar CPU, optimized around Matrix multiply units and
- 1:52:40Vector Cindy. It supports, floating-point 32,
- 1:52:44Be float16 and a new format cfp, 8, configurable fe8, and it is
- 1:52:51backed by one. And a quarter, MB of fast, ECC
- 1:52:55protected, s RAM, and the low latency high band width fabric
- 1:52:59that we designed This might be our smallest entity of scale,
- 1:53:04but it packs a big punch, more than 1 teraflop of compute in
- 1:53:10our smallest entity of scale. So, let's look at the
- 1:53:13architecture of this. The computer Architects out
- 1:53:17here, may recognize this, this is a pretty capable architecture
- 1:53:21as soon as you see this. It is a superscalar in order CPU
- 1:53:26with four wide vector and to, I'd Vector to I'd for white
- 1:53:30scalar. And to, I'd Vector pipes, we
- 1:53:33call it in order. Although the vector and the
- 1:53:36scalar pipes can go out of order, but for the purists out
- 1:53:39there, we still call it in order and it also has four ways.
- 1:53:43Tithe reading this increases utilization because we could do
- 1:53:46compute and data transfers simultaneously, and our custom
- 1:53:50is a, which is the instruction set.
- 1:53:52Architecture is fully optimized for machine learning workloads.
- 1:53:56It has features like transpose gather linked reversals
- 1:54:01broadcast just to name a few And even in the Physical Realm, we
- 1:54:07made it extremely modular. Such that we could start
- 1:54:10abutting these training nodes in any direction and start forming
- 1:54:15the compute plane that we envisioned.
- 1:54:21When we click together 354 of these training nodes, we get our
- 1:54:25computer, a it's capable of delivering, 362, teraflops of
- 1:54:30machine learning compute. and of course, the high-bandwidth
- 1:54:34fabric that interconnects these And around this computer array,
- 1:54:39we surrounded it with high speed, low power Services, 576
- 1:54:45of them to to enable us to have extreme I/O bandwidth coming out
- 1:54:51of this chip. Just to give you a comparison
- 1:54:54point, this is more than two times, the bandwidth coming out
- 1:54:59of the state of the earth, networking switch chips, which
- 1:55:03are out there today and networks which tips are supposed to be
- 1:55:06the gold standard. That's for I/O bandwidth.
- 1:55:10If we put all of it together we get training, optimize chip, rd1
- 1:55:17chair. This trip is manufactured in
- 1:55:21seven, nanometer technology. It packs, 50 billion transistors
- 1:55:25in a miserly 6, 45 millimeter square, one thing, you notice, a
- 1:55:29hundred percent of the area out here is going towards machine
- 1:55:33learning training and bandwidth. There is no dark silicon, there
- 1:55:37is no Legacy support. This is a pure machine, learning
- 1:55:40machine. and, This is the D one chair in a flip chip BGA
- 1:55:53package. This was entirely designed by
- 1:55:57Tesla team internally, all the way from the architecture to TDs
- 1:56:04out and package. This trip is like a GPU level
- 1:56:11compute with a CPU. Level flexibility and twice the
- 1:56:15network. Chip level I/O bandwidth.
- 1:56:19If I were to plot the I/O bandwidth on the vertical scale
- 1:56:24versus teraflops of compute that is available in the
- 1:56:27state-of-the-art, machine learning chips, are there
- 1:56:30including some of the startups. You can easily see why our
- 1:56:34design Point, excels Beyond par. now, that we had this
- 1:56:39fundamental physical building block, how to design the system
- 1:56:45around it. Let's see since D1.
- 1:56:51Chips can seamlessly connect without any glue to each other.
- 1:56:56We just started putting them together.
- 1:56:59We just put 500,000 training notes together, to form our
- 1:57:10compute plane. This is thousand.
- 1:57:13Five hundred D1 chips, seamlessly connected to each
- 1:57:16other. And then we add Dojo interface
- 1:57:22process processors on each end. This is the host bridge to
- 1:57:28typical host in the data centers.
- 1:57:30It's connected with PCI Gen4. On one side, with a high
- 1:57:35bandwidth, fabric to our compute, plane the interface
- 1:57:38processors provide, not only the host bridge, but high-bandwidth
- 1:57:43dear, a shared memory for the compute plane in addition.
- 1:57:48The interface processors can also allow us to have a higher
- 1:57:52Radix network connection. In order to achieve this compute
- 1:57:58plane. We had to come up with a new way
- 1:58:01of integrating these chips together and this is what we
- 1:58:05call as a training tile. This is the unit of scale for
- 1:58:10our system. This is a groundbreaking
- 1:58:15integration of 25 known good D1 dies on to a fair fan-out, wafer
- 1:58:20process, tightly integrated, such that it preserves the
- 1:58:24bandwidth between them. The maximum bandwidth is
- 1:58:27preserved there and in addition we generated a connector a
- 1:58:33high-bandwidth high-density connector that preserves the
- 1:58:37bandwidth coming out of this training tile.
- 1:58:41And this style gives us nine petaflops of compute with a
- 1:58:46massive I/O bandwidth coming out of it.
- 1:58:51This perhaps is the biggest organic MCM in the chip industry
- 1:58:56multi-chip module. It was not easy to design this.
- 1:59:01They were no tools that existed. All the tools were croaking even
- 1:59:05our compute cluster couldn't handle it.
- 1:59:08We had to our Engineers came up with different ways of solving
- 1:59:12this. They created new methods to make
- 1:59:15this a reality now that we had our compute plane tile.
- 1:59:21With high bandwidth iOS. We had to feed it with power.
- 1:59:26And here we came up with a new way of feeding power vertically.
- 1:59:31We created a custom Voltage, regulator module, that could be
- 1:59:38reflowed, directly directly onto this fan out wafer.
- 1:59:42So what did we did out here is, we got chip package and we
- 1:59:47brought PCB level technology of Reflow on to this fan art way
- 1:59:51for technology. This is a lot of integration
- 1:59:55already out here, but we didn't stop here we integrated the
- 2:00:00entire electrical thermal and mechanical pieces out.
- 2:00:04Out here. To form our training tile fully
- 2:00:09integrated interfacing with a 52 volt DC input.
- 2:00:17It's unprecedented. This is an amazing piece of
- 2:00:20engineering. Our compute plane is completely
- 2:00:25orthogonal to power supply and cooling that makes
- 2:00:31high-bandwidth compute planes possible.
- 2:00:41What it is is a nine petaflop training tile.
- 2:00:45This becomes our unit of scale for our system.
- 2:00:50And this. Is real.
- 2:01:08I can't believe I'm holding 9 petaflops out here.
- 2:01:21And in fact, last week we got our first functional training
- 2:01:26tile. And on a limited limited
- 2:01:30cooling, benchtop setup. We got some networks running.
- 2:01:35And I was told, Andre doesn't believe that we could run
- 2:01:38networks till we could run one of his Creations.
- 2:01:42Andre, this is minji pt.2 running under Joe.
- 2:01:47Do you believe it? Next up.
- 2:01:56How to form a compute cluster out of it.
- 2:01:59By now, you must have realized our modularity story is pretty
- 2:02:04strong. You just put together some
- 2:02:07tiles. We just tied together tiles.
- 2:02:11A 2 by 3. Tile in a tray, makes our
- 2:02:15training Matrix and two trays in a cabinet gave hundred petaflops
- 2:02:21of compute. Did we stop here?
- 2:02:26No. We just integrated seamlessly,
- 2:02:32we broke the cabinet was we integrated the style seamlessly,
- 2:02:36all the way through preserving the bandwidth.
- 2:02:39There is no bandwidth divot out here.
- 2:02:41There's no bandwidth Cliffs all the ties are seamlessly
- 2:02:45connected with the same bandwidth and with this We have
- 2:02:54a neck support. This is one exaflop of compute.
- 2:03:02In 10 cabinets. It's more than a million
- 2:03:07training nodes that you saw. We paid meticulous attention to
- 2:03:10that training node in there. Are 1 million nodes out here
- 2:03:15with uniform bandwidth. Not just the hardware software
- 2:03:23aspects are so important to ensure scaling.
- 2:03:28And not every job requires a huge cluster so we planned for
- 2:03:32it, right from the get-go. Are compute.
- 2:03:37Plane can be, subdivided can be partitioned into units called
- 2:03:43Dojo processing unit. ADP you consists of one or more
- 2:03:50D1 chips. It also has our interface
- 2:03:54processor and one or more hosts. And this can be scaled up or
- 2:04:00down as per the needs of any algorithm any network running on
- 2:04:05it. What does the user have to do?
- 2:04:09They have to change their scripts minimally?
- 2:04:14And this is because of our strong compiler Suite.
- 2:04:19It takes care of fine, grained, parallelism and mapping the
- 2:04:22problems of mapping the neural networks very efficiently on to
- 2:04:27our compute plane. Our compiler is uses multiple
- 2:04:33techniques to extract parallelism.
- 2:04:36It can transform the Network's to achieve, not only fine
- 2:04:40grained, parallelism using data model graph.
- 2:04:44Parallelism techniques. It also can do optimizations to
- 2:04:50reduce memory footprints. One thing because of our high
- 2:04:56bandwidth nature of the fabric is enabled out.
- 2:04:59Here is model parallelism could not have been extended to the
- 2:05:03same level. As what we can, it was limited
- 2:05:06to chip boundaries now we can because of our high bandwidth,
- 2:05:10we can extend it to training tiles and Beyond.
- 2:05:13Thus large networks can be efficiently.
- 2:05:17Matt here at low batch sizes and extract utilization, a new
- 2:05:21levels of performance. In addition, our compiler is
- 2:05:26capable of handling. High-level Dynamic control
- 2:05:29flows, like Loops if-then-else Etc.
- 2:05:34And our compiler engine is just part of our entire software
- 2:05:38suite. The stack consists of Extension
- 2:05:44to Pirate Arch that ensures the same user level interfaces.
- 2:05:48That ml scientists are used to and our compiler, generates code
- 2:05:52on the Fly such that it could be reused for subsequent execution.
- 2:05:58It has a llvm back-end that generates the binary for the
- 2:06:01hardware and this ensures, we can create optimized code for
- 2:06:07the hardware without relying on. Even single line of handwritten
- 2:06:11colonel, A driver stack takes care of the multi host multi
- 2:06:18partitioning that you saw few slides back.
- 2:06:22And then we also have profilers and debuggers in our software
- 2:06:29stack. So with all this, we integrated
- 2:06:34In a vertical fashion, we broke the traditional barriers to
- 2:06:39scaling and that's how we got modularity up and down the stack
- 2:06:44to add to new levels of performance to sum it all.
- 2:06:49This is what it will be. It will be a fastest AI training
- 2:06:54computer for X the performance. At the same cost 1.3 X better
- 2:07:01performance per watt that is energy, saving and Phi.
- 2:07:04A smaller footprint, this will be Dojo computer.
- 2:07:19And we are not done. We are assembling our first
- 2:07:23cabinets pretty soon and we have a whole Next Generation plan.
- 2:07:28Already, we are thinking about 10x more with different aspects
- 2:07:33that we can do all the way from Silicon to the system.
- 2:07:37Again, we will have this journey again your recruiting heavily
- 2:07:41for all of these areas. Thank you very much.
- 2:07:55And next up, alone will update us on what's beyond our vehicle
- 2:08:00Feet Fleet for AI. All right.
- 2:09:32Thank you. Unlike John unlike Dojo
- 2:09:43obviously that was not real, so don't use real.
- 2:09:47The Tesla bought will be real, but basically, if you think
- 2:09:52about what we're doing right now with the cars, Tesla is arguably
- 2:09:55the world's biggest robotics company because our cars are
- 2:09:58like, said Sammy sentient robots on Wheels and with the full self
- 2:10:05driving. And computer, essentially, the
- 2:10:06inference engine on the car, which we will keep evolving,
- 2:10:09obviously and dojo, and all the neural Nets.
- 2:10:15Recognizing the world understanding, how to navigate
- 2:10:17through the world. It kind of makes sense to put
- 2:10:20that onto a humanoid form and also quite good at senses and
- 2:10:25batteries and Actuators. So we think we'll probably have
- 2:10:33a prototype sometime next year. That is basic, looks like this
- 2:10:39and it's intended to be friendly, of course, and
- 2:10:49navigate to a world, built for humans and eliminates it
- 2:10:54dangerous, repetitive and boring tasks.
- 2:10:57We're setting it. Such that it is Add a mechanical
- 2:11:01level at a physical level. You can run away from it and
- 2:11:08most likely overpower it. So hopefully that doesn't ever
- 2:11:13happen but you never know. So it's a, it'll be a, you know,
- 2:11:21a light, light light, five miles an hour.
- 2:11:24You get run pass and I'll be fine. so yeah, it's a right
- 2:11:34around five foot eight has sort of a screen where the head is
- 2:11:40for useful information but as otherwise basically the
- 2:11:44autopilot system in it so it's a great cameras, got eight cameras
- 2:11:47and yeah What's the driving computer and making use of all
- 2:11:58of the same tools that were using the car so many things I
- 2:12:03think that are really hard about having a useful humanoid robot
- 2:12:07is cannot navigate through the world without being explicitly
- 2:12:10trained. I mean without explicit like
- 2:12:14line-by-line instructions. Can you can you talk to it and
- 2:12:18say please pick up that bolt and attach it to the car with that
- 2:12:26ranch and it should be able to do that.
- 2:12:30It should be able to please, you know, please go to the store and
- 2:12:33get me the following groceries that kind of thing.
- 2:12:37So Yeah, I think we can do that. And yeah this I think will be
- 2:12:48quite quite profound because if you say it like what is the
- 2:12:51economy it is at the foundation of his labor.
- 2:12:55So what happens when there is, you know, no shortage of Labor
- 2:13:04is why I think long-term that they will need to be Universal
- 2:13:06basic income. All right now because this
- 2:13:12Robert doesn't work. So we didn't need a minute.
- 2:13:17So yeah. But I think it's essentially in
- 2:13:21the future of physical work will be a choice if you want to do
- 2:13:24it, you can but you won't need to do it.
- 2:13:26And yeah, I think obviously has profound implications for the
- 2:13:30economy because given that the economy at its foundational
- 2:13:34level is labor. I mean capital is Capital
- 2:13:37Equipment is just distilled labor.
- 2:13:39Then, is there any actual limit to the economy?
- 2:13:43Maybe not. So yeah.
- 2:13:51Join our team and help hold this.
- 2:13:53All right, so I think will will have everyone come back on the
- 2:13:56stage and you guys can ask questions if you'd like.
- 2:13:59Yeah. Maybe try to get the camera
- 2:14:48angle from there, but not from the side.
- 2:14:57All right, cool. So we probably turn the lights
- 2:15:00back on and Yeah, right. So we're happy to answer any
- 2:15:08questions. You have about anything on the
- 2:15:11software, Hardware side where things are going and yeah, fire
- 2:15:15away. We have because the lights are
- 2:15:18like, interrogation lights. So, we actually cannot see other
- 2:15:21ago. Great.
- 2:15:24All right, cool. I can just okay, there we go.
- 2:15:40First, I mean, thanks to all the presenters, that was just super
- 2:15:42cool to see everything. I'm just curious at a high level
- 2:15:45and this is kind of a question for, really, anyone who wants to
- 2:15:47take it to? What extent are you interested
- 2:15:50in publishing or open sourcing anything that you do for the
- 2:15:55future? Well, I mean it is a
- 2:16:11fundamentally extremely expensive to create the system
- 2:16:15so somehow that has to be paid for, I'm not sure how to pay for
- 2:16:20it if it's fully open sourced. Yeah, most people want to work
- 2:16:26for free. But, but I should say that this
- 2:16:34is if other car companies want to license it and use it in
- 2:16:38their cars, that would be cool. This is not intended to be just
- 2:16:41limited to Tesla cars. It's for the dojo supercomputer.
- 2:16:49So, did you solve the compiler problem of scaling, to these
- 2:16:52many nodes or is, or if it is solved, is it only applicable to
- 2:16:58do Joe because I'm doing research in deep learning
- 2:17:03accelerators and getting the correct scale ability or the
- 2:17:07distribution. Even in one ship is extremely
- 2:17:11difficult from the research projects perspective.
- 2:17:14So I was just curious Excuse me. Mike for You want to have?
- 2:17:22We solved the problem not yet? Are we confident we will solve
- 2:17:25the problem? Yes, we have a demonstrated
- 2:17:28networks on Prototype Hardware. Now, we have models performance
- 2:17:31model showing the scaling. The difficulty is, as you said,
- 2:17:35how do we keep the localities? If we can do enough model,
- 2:17:38parallel enough data, parallel, to keep most of the things
- 2:17:42local, we just keep scaling. I have to fit the parameters in
- 2:17:45our working, set in rsrm that we have and we flow through the
- 2:17:48pipe. There's plenty of opportunities.
- 2:17:54Sorry, as we get further scale for further process or nodes
- 2:17:57have more local memory memory. Trade-offs with bandwidth, we
- 2:18:00can do more things, but as we see it now, the applications
- 2:18:04that Tesla halves we see a clear path.
- 2:18:09And our mortality story means we can have different ratios
- 2:18:13different aspects created out of it.
- 2:18:16I mean, this is something that we chose for our applications
- 2:18:19internally. Sure.
- 2:18:27Portion of it given that training is such a soft scaling
- 2:18:30application, even though you have all these compute and have
- 2:18:34a high bandwidth, hi van with interconnect.
- 2:18:39It could not give you that performance because you are
- 2:18:42doing computations on limited memory are different locations.
- 2:18:46So, I was the that's very curious to me when you said it's
- 2:18:49solved because I just jumped onto the opportunity and would
- 2:18:52love to know more given that. How much you can open source?
- 2:18:56Yeah. Yeah, I guess proofs in the
- 2:19:00pudding. So we're we should have Dojo
- 2:19:04operational next year and I will, obviously use it for
- 2:19:10training video training. It's I mean, I'm not like this
- 2:19:12is about like the, the primary application initially is I we've
- 2:19:17got vast amounts of video and how we train vast amounts of
- 2:19:20video as efficiently as possible.
- 2:19:23And Also shorten the amount of time like, if you're trying to
- 2:19:29train train to a task, like just in general, Innovation is how
- 2:19:35many iterations and what is the average progress between each
- 2:19:38iteration. And so if if you can reduce the
- 2:19:41time between iterations, the rate of improvement is much
- 2:19:45better. So, you know, takes like
- 2:19:48sometimes couple days for models train versus a couple hours.
- 2:19:52That's that's a big deal. But the the acid test here and
- 2:19:58you know what of told that Dojo team is like it's successful.
- 2:20:03If the software team wants to turn off the GPU cluster but if
- 2:20:09they want to keep the GPU cluster on it's not successful.
- 2:20:14So Hi. Bert.
- 2:20:21Over here. Love the presentation.
- 2:20:23Thank you for getting us out here.
- 2:20:25Love to everything. Especially the simulation part
- 2:20:27of the presentation. I was wondering look very, very,
- 2:20:31very realistic. Are there any plans to maybe
- 2:20:33expand simulation to other parts of the company in any way?
- 2:20:38Mike too high. I mean glow.
- 2:20:41I manage the autopilot simulation team.
- 2:20:44So as we go down the path to full self-driving, we're going
- 2:20:47to simulate more and more of the vehicle currently.
- 2:20:50We're simulating vehicle Dynamics, we're going to be Ms.
- 2:20:54We're going to need the MCU. We're going to every single part
- 2:20:55of the vehicle integrated, and then actually makes the
- 2:20:58autopilot simulator really useful for places outside of
- 2:21:01autopilot. So I want to expand our we want
- 2:21:03to expand eventually to being a universal simulation platform.
- 2:21:07But I think before that we're going to be spending a lot of
- 2:21:09Optimus support and then a little bit further down the line
- 2:21:12whether we have some rough ideas and potentially how to get the
- 2:21:16simulation infrastructure. And some of the cool things
- 2:21:18we've built into the hands of people outside of the company.
- 2:21:22Optimus is the code name for the Tesla bot.
- 2:21:25Oops, Optimus up Ryan. Yeah hi this is a Dijon Ian.
- 2:21:40Thank you for the great presentation and putting all of
- 2:21:43these cool things together. Yeah for a while, I have been
- 2:21:47thinking that the car is already a robot.
- 2:21:50So why not a human? It's robot.
- 2:21:53And I'm so happy that today you mentioned that you're going to
- 2:21:58build such thing especially I think that this can give
- 2:22:02opportunity for raise of putting multi-modal.
- 2:22:06Tea together for instance, we know that in the example that
- 2:22:11you are showed that there was a dog and with some passengers or
- 2:22:16running together the language and symbolic processing can
- 2:22:23really help for visualizing that.
- 2:22:26So I was wondering if I could hear a little more about this
- 2:22:33type of putting modalities together, Eluding language and
- 2:22:37vision because I have been working with for instance, mini
- 2:22:41GPT the and Andre put out there. And yeah, I didn't hear much
- 2:22:47about other modalities that's going into the car, or at least
- 2:22:52in this simulation. Is there any comment that you
- 2:22:57could tell us? Well, driving is fundamentally
- 2:23:02basically almost entirely Vision neural Nets.
- 2:23:06Basically, it's running on a biological Vision.
- 2:23:09Neural, net. And we're doing here is a
- 2:23:12silicon camera, neural net. So, there are there's some
- 2:23:17amount of audio, you know, you want to hear if there's like an
- 2:23:21emergency vehicles or, you know, I guess converse with the people
- 2:23:28in the car. You know, somebody's yelling
- 2:23:32something at the at you at the car that car.
- 2:23:35Nice understand. Well, that is so old things that
- 2:23:41are necessary for to be fully autonomous.
- 2:23:45Yeah, thank you. All right, thank you for all the
- 2:23:54great work that you've shown. My question is for the team
- 2:23:58because the data that were shown was seems to be predominantly
- 2:24:01from the United States that the the FST computer is being
- 2:24:04trained on. But as it is being as it gets
- 2:24:07rode out to different countries which have their own Road
- 2:24:09systems and challenges that come with it.
- 2:24:12How do you think that it's going to scale like like I'm assuming
- 2:24:16like Roundup is not a very viable solution so how does the
- 2:24:20transfer to different countries? Well, there we actually do
- 2:24:24trained on using data from probably, like, 50 different
- 2:24:29countries, but we have to pick it, you know, as we're trying to
- 2:24:36Advance full self-driving, we need to pick one country.
- 2:24:38And since we're located here, we pick the US.
- 2:24:42And there are a lot of questions like why not even Canada like
- 2:24:44well because the roads are a little different in Canada
- 2:24:46different enough and so we're trying to solve a hard problem.
- 2:24:51You want to Say like okay what's the let's not add additional
- 2:24:57complexity right now. Let's just solve it for the US
- 2:25:00and then will extrapolate to the rest of the world.
- 2:25:02But we do use video from all around the world.
- 2:25:04Yeah. I think a lot of a lot of what
- 2:25:06we are building, as very country of Gnostic fundamentally all the
- 2:25:09computer, vision components, and so on, don't care too much about
- 2:25:11the country specific. Sort of features everything, you
- 2:25:15know, different countries have roads and they have curbs and
- 2:25:18they have cars and everything were building is fairly general
- 2:25:20for that and the The prime directive is don't crash, right?
- 2:25:25And that's true for every country.
- 2:25:26Yes. So sleep prime directive.
- 2:25:30And even right now the car is pretty good at not crashing.
- 2:25:35And so just basically whatever it is, don't hit it.
- 2:25:40Even if it's a UFO that crashed landed on the highway and to
- 2:25:45learn hit, it should not need to recognize it in order to not hit
- 2:25:49it. That's very important.
- 2:26:01Haitian and I wanted to ask that when you do the photo magic
- 2:26:06process, multi-view geometry, how much of an error GOC is at
- 2:26:10like 1? Mm, 1 cm.
- 2:26:13So I'm just if it's not confidential, All right
- 2:26:16question. What is the, what's the?
- 2:26:18What's the difference between the synthetic?
- 2:26:21Sure. What is the difference between a
- 2:26:24synthetically created geometry to the actual geometry?
- 2:26:29Yeah, it's usually within a couple centimeters three or four
- 2:26:32centimeters. That's the standard deviation.
- 2:26:38We're different kind of modalities to bring down that
- 2:26:41error. We primarily try to find
- 2:26:45scalable ways to label in some occasions.
- 2:26:47We use other senses to inch help Benchmark, but we primarily use
- 2:26:51cameras for this system. Okay.
- 2:26:54Thanks. Yeah, I think we want to aim for
- 2:26:57the car to be positioned accurately to the social
- 2:27:01centimeter level you know something on that order.
- 2:27:05I was a it'll depend on Distance by closed I think is a much more
- 2:27:08accurate than farther away things because and they would
- 2:27:10matter less because the car does not make decisions much farther
- 2:27:13away, and as it comes close, it will become more and more
- 2:27:15accurate. Exactly.
- 2:27:20A lot of questions. Thanks everybody.
- 2:27:23My question has to do with sort of AI and Manufacturing.
- 2:27:26It's been a while since we've heard about the alien.
- 2:27:28Dreadnought concept is the humanoid that's behind you guys.
- 2:27:32Is that kind of brought out of the production house, I'm line
- 2:27:34and saying that humans are underrated in that process.
- 2:27:40Well, sometimes like, you know, something that I say is taken to
- 2:27:44too much of an extreme there. There are parts of the tell the
- 2:27:50system that are almost completely automated.
- 2:27:52And then there are some parts that are almost completely
- 2:27:54manual. And if you were to walk through
- 2:27:58the whole production system, you would see a very wide range from
- 2:28:02yeah, like sip fully automatic to almost completely manual but
- 2:28:07the vast It's most of it is, is already automated.
- 2:28:14So, and then with the, some of the design architecture changes,
- 2:28:19like going to large, aluminum high-pressure diecast
- 2:28:24components. We can take the entire rear,
- 2:28:27third of the car and cast it as a single piece.
- 2:28:29And now we're going to do that. The front third of the cars, a
- 2:28:33single piece. So the body line drops by like
- 2:28:3860 to 70 percent in size, but yeah, the, the robot is not is
- 2:28:45not prompted by specifically by manufacturing needs.
- 2:28:50It's just that, we're just obviously making the pieces that
- 2:28:56are needed for a useful, humanoid robot.
- 2:29:00So I guess we probably should make it and if we don't someone
- 2:29:03else with willing so I guess we should make it.
- 2:29:08And make sure it's safe. I should say like also
- 2:29:15manufacturing volume manufacturing is extremely
- 2:29:17difficult and underrated and we've gotten pretty good at
- 2:29:20that. It's also important for that
- 2:29:23humanoid robot. Like how do you make that you
- 2:29:25wanted robot not be super expensive and Hi.
- 2:29:32Thank you for the present presentation and my question
- 2:29:36will be about scaling of the jaw and in particular, how do you
- 2:29:42scale the compute nodes in terms of thermal thermals and power
- 2:29:48delivery? Because there is only so much
- 2:29:51heat that you can dispense and only so much power that you can
- 2:29:56bring to like cross track. And how do you point to scale it
- 2:30:02and how they point to scale it in multiple data centers?
- 2:30:08Sure. Hi, I'm Bill.
- 2:30:16I wanted to do Joe Engineers. The so from a thermal standpoint
- 2:30:21and power standpoint, we've designed it very modular.
- 2:30:25So what you saw in the compute tile that will that will cool
- 2:30:28the entire tile. So what we once we hook it up to
- 2:30:32it, is liquid cooled on both the top and the bottom side.
- 2:30:36It doesn't need anything else. And so when we talk about
- 2:30:40clicking these together, once we click it, To power.
- 2:30:44And we once we click it to cooling, it will be fully
- 2:30:48powered and fully cooled. And all of that is less than a
- 2:30:51cubic foot. Yet, and so Tessa has a lot of
- 2:30:55expertise in power electronics, and in cooling.
- 2:30:59So, we took the Power Electronics expertise, from the
- 2:31:03vehicle power train and the sort of advanced cooling that we
- 2:31:06developed for the power electronics and for the vehicle
- 2:31:10and applied that to the supercomputer because as you
- 2:31:13point out, getting heat out is extremely important, just really
- 2:31:18heat Limited. So yes, it was funny that the at
- 2:31:24the computer level its operating at less than a bolt, which is a
- 2:31:29very low voltage is a lot of amps.
- 2:31:32So therefore a lot of heat I squared.
- 2:31:34R is a robot really bites you on the ass.
- 2:31:40Hi. My questions also really a
- 2:31:42question of scaling, so it seems like a natural consequence of
- 2:31:46using, you know, significantly faster, training Hardware is
- 2:31:49that you'd be either training models over a lot more data or
- 2:31:52you'd be training a lot more complex models, which would be
- 2:31:55potentially significantly more expensive to run at inference
- 2:31:58time on the cars. I guess I was wondering like if
- 2:32:01there was a plan to like, also apply do JoJo as something that
- 2:32:07you'd be using like, on this, all driving cars.
- 2:32:11And if so, like, you know, do you foresee additional
- 2:32:14challenges there? I can.
- 2:32:18So as you could see, like Andres models are not just for cars.
- 2:32:22Like there are two labeling models.
- 2:32:24There are other models that are like Beyond Car application, but
- 2:32:29they feed into the car stack. So, so does, your will be used
- 2:32:33for all of those too. Not just the car inference, part
- 2:32:38of the training. Yeah, I mean if the teachers
- 2:32:43first application will be consuming video data for
- 2:32:46training for that would then be run in the adverse inference
- 2:32:49engine on the car, but and that I think is an important test to
- 2:32:54see if it actually is good or is actually better than GPU cluster
- 2:32:58or not. So, but then beyond that, it's
- 2:33:02basically a general journalized neural net training computer,
- 2:33:06but it's very much optimized to be a neural net.
- 2:33:09So You know, CPUs gpus. There are there weren't they're
- 2:33:15not made to be, they're not they're not designed
- 2:33:18specifically for training neural.
- 2:33:20Nets been able to make GPS especially very efficient for
- 2:33:26portraying neural Nets but that's not that was never their
- 2:33:28design intent. So it's basically gpus are so
- 2:33:32essentially running it neural net training in emulation mode.
- 2:33:35So with with Dojo were saying like, Okay let's just let's just
- 2:33:39a sec. The whole thing was just a
- 2:33:41statement, have this thing that's both for one purpose and
- 2:33:44that is Is neural net training and just generally any system
- 2:33:48that is designed for a specific purpose will be better than one
- 2:33:51that is designed for a general purpose.
- 2:33:55Hey, I have a question here. So you describe two separate
- 2:34:01systems one. Observation, therefore planner
- 2:34:03and control this Dojo love you. Train networks that cross that
- 2:34:07boundary and second thing is if you are able to paint such
- 2:34:11networks, would you have the on-board computer capability in
- 2:34:13FST system to be able to run that when in the under your
- 2:34:16tight latency constraints? Thanks.
- 2:34:21Yeah, I think we should be able to train panel networks on dojo
- 2:34:24or any GPS. It's really invariant to the
- 2:34:27platform and I think if anything, once you make this
- 2:34:30entire thing and to n it be more efficient than decoding or of
- 2:34:33this intermediate States, so we should be able to run faster.
- 2:34:35If you make the entire thing into a new node networks, we can
- 2:34:38avoid a lot of decoding of the intermediate States and only
- 2:34:41decode essential things required for driving the car.
- 2:34:44Yep. Certainly.
- 2:34:45And to understand, is the guiding principle behind a lot
- 2:34:48of the network developments and overtime in the stack.
- 2:34:50Back neural networks have taken on more and more functionality.
- 2:34:53And so we want everything to be trained and to end because we
- 2:34:56see that that works best, but we are building it incrementally.
- 2:34:59So right now, the interface there is a vector space and we
- 2:35:01are consuming it in the planner but nothing really fundamentally
- 2:35:04prevents you from actually taking features and eventually
- 2:35:07fine-tuning end to end. So I think that's definitely
- 2:35:09where this is headed. And the discovery really is like
- 2:35:12what are the right architectures that we need to place in network
- 2:35:15blocks to make it amenable to the task.
- 2:35:17So, like on a describe, we can place special RNs.
- 2:35:20Help with the perception problem and now it's just neural
- 2:35:23network. So, similarly for planning, we
- 2:35:24need to bake in search optimization into the planning,
- 2:35:27enter the network architecture, and once we do that, you should
- 2:35:29be able to do planning very quickly.
- 2:35:31Similar to C++, algorithms. Okay.
- 2:35:41Okay I think I had a question very similar to what he was
- 2:35:45asking about. It seems like a lot of neural
- 2:35:48Nets around computer vision and kind of traditional planning you
- 2:35:51had model predictive control and solving convex optimization
- 2:35:54problems very quickly. And I'd wonder if there's a
- 2:35:58computer architecture that's more suited for convex
- 2:36:01optimization or the model predictive control Solutions
- 2:36:05very quickly. Yeah, 100% vegan bacon, like I
- 2:36:07said earlier, we want to bake in these architectures that do say
- 2:36:10model Critter Control, but it just like replay some of the
- 2:36:12blocks with neural networks or if we know the physics of it.
- 2:36:15They can also use fixed base models, part of the neural
- 2:36:17networks for pass itself. So they are going to go towards
- 2:36:20a hybrid system where we will have nutrient of blocks into
- 2:36:27place together with physics-based blocks and Mornin
- 2:36:30folks later. So it will be a hybrid stack.
- 2:36:32And what we know to do. Well, we place was explicitly
- 2:36:34and what the networks are, graded will use the network.
- 2:36:36To optimize this. So brilliant, when stack with
- 2:36:39this architecture baked in, When I do think that so long as
- 2:36:44you've got like surround video neural, nets for understanding
- 2:36:50what's going on and can convert those around video into Vector
- 2:36:57space. Then you basically have a video
- 2:36:59game and if you know if you it's like if you're in Grand Theft
- 2:37:04Auto, whatever you can, you can make the cars drive around and
- 2:37:06pedestrians walk around without crashing.
- 2:37:08So you can do You don't have to have a neural net for control
- 2:37:14and planning but it's probably ultimately better, so but I
- 2:37:21think you probably get to. In fact, I'm sure you can get
- 2:37:24too much safer than human with control and planning primarily
- 2:37:28in C++ with perception Vision in neural net.
- 2:37:37Hi, my question is, we've seen other companies for example, use
- 2:37:43reinforcement learning and machine learning to optimize
- 2:37:46power consumption and data centers and all kinds of other
- 2:37:48internal processes. My question is, are, is Tesla
- 2:37:50using machine learning within its manufacturing design or
- 2:37:55other engineering processes. I discourage use of machine
- 2:38:05learning because it's really difficult unless you basically,
- 2:38:10unless you have to use machine learning, don't do it.
- 2:38:15It's usually a red flag. When somebody says we want us to
- 2:38:18use machine learning to solve the stack.
- 2:38:19I'm like, that sounds like bullshit.
- 2:38:22So, 99.9 percent time you do not need it.
- 2:38:28So, Yeah. So it's kind of like a you reach
- 2:38:35for machine learning when you when you need to know.
- 2:38:39But it's I've not found it to be a convenient.
- 2:38:41Easy thing to do is to super hard things to do.
- 2:38:45That may change. If you got a humanoid robot that
- 2:38:46can you no understand normal instructions, but, yeah,
- 2:38:54generally minimize use of machine learning in the rectory
- 2:39:02Hi based on your videos from the simulator.
- 2:39:05It looked like a combination of graphical and neural approaches.
- 2:39:09I'm curious what the set of underlying techniques that are
- 2:39:12used for your simulator and specifically for neural
- 2:39:16rendering if you can share. Yeah.
- 2:39:20So we're doing at the bottom of the stack, it's just traditional
- 2:39:24game techniques. Just rasterization real-time,
- 2:39:29you know, very similar to what you'd see in like GTA on top of
- 2:39:33that, we're doing real time, Ray tracing.
- 2:39:35And then those results were really hot off the press.
- 2:39:39I mean, we had that little asterisk at the bottom that was
- 2:39:41from last night, we're going into the neural rendering space.
- 2:39:44We're trying out a bunch of different things.
- 2:39:46We want to get to the point where the neural run.
- 2:39:48During is the the Tyrion the top that pushes it to the point
- 2:39:51where the models will never be able to over fit on our
- 2:39:54simulator. Currently we're doing things
- 2:39:58similar to photorealism enhancement.
- 2:40:00There's a paper, a recent paper photo enhancing photorealism
- 2:40:03enhancement but we can do a lot more than what they could do in
- 2:40:06that paper because we have way more labeled data way more
- 2:40:09compute and also much. We have a lot more control over
- 2:40:13environments and we also have a lot of people who can help us
- 2:40:16make this run at real time. But we're going to try whatever
- 2:40:19we can do to get to the point where we can train everything,
- 2:40:23just with the simulator if we had to, but we will never have
- 2:40:27to, because we have so much real-world data.
- 2:40:29That no one else has. It's just to fill in the little
- 2:40:32gaps in the real world. Yeah, I mean, as the simulator,
- 2:40:37is very helpful. When there's like these rare
- 2:40:40cases, like like, you know, like collision avoidance, right
- 2:40:44before an accident. And then, ironically, the better
- 2:40:48our car Has become at avoiding accidents.
- 2:40:50The fewer accidents they are so they're not training.
- 2:40:52Set is small, so then we have to make them crash in the
- 2:40:54simulation. So it's like okay minimize
- 2:40:58potential injury to pedestrians and people in the car at you
- 2:41:04have five meters, you're traveling at you know, 20 meters
- 2:41:08per second. What actions would minimize
- 2:41:12prior probability of injury? We can run that in some cars
- 2:41:23driving down the wrong side of the highway that kind of thing,
- 2:41:26happens occasionally but not that often.
- 2:41:30For your humanoid context. I'm wondering if you've decided
- 2:41:33on what use cases, you're going to start with and what the grand
- 2:41:36challenges are in that context to make this viable.
- 2:41:42Well, thank for the human foot, the Tesla bot after this.
- 2:41:49It's basically going to start with just dealing with work.
- 2:41:53That is boring, repetitive and dangerous.
- 2:41:58Basically, what is the work that people would least like to do Hi
- 2:42:11so quick. Question about your simulations,
- 2:42:14obviously, they're not perfect right now.
- 2:42:16So are you using any sort of domain adaptation techniques to
- 2:42:20basically, bridge the gap between your simulated data, and
- 2:42:23your actual real-world data? Because I imagine it's kind of
- 2:42:26dangerous to just deploy models which are solely trained on
- 2:42:30simulated data. So maybe some sort of explicit
- 2:42:32domain adaptation, or something. Is that going on anywhere in
- 2:42:36your pipeline? So currently, we're producing
- 2:42:40the videos straight out of the simulator, the the full clips of
- 2:42:43kind of Maddox and everything and then we're just immediately
- 2:42:46training on them but it's not the entire data set.
- 2:42:49It's just a small targeted segment and we only
- 2:42:51re-evaluating based on real-world video.
- 2:42:54We're paying a lot of attention to make sure we don't ever fit
- 2:42:57in a, we have to start doing fancier things we will but
- 2:43:00currently it's we're not having an issue with it, over fitting
- 2:43:03on the simulator. We will as we scale up the data
- 2:43:06and that's what we're hoping to use.
- 2:43:07Use neural rendering to bridge that Gap to push that even
- 2:43:10further out. We've already done things we're
- 2:43:12we're using like the same network is in the car, but
- 2:43:14retrain it to detect some versus real to drive art decisions.
- 2:43:18And that's actually helped prevent some of these things as
- 2:43:21well. Yeah, just emphasize the
- 2:43:25overwhelmingly. The data set is the real video
- 2:43:28from the Cars on the actual roads lovings weirder or has
- 2:43:32more Corner cases than reality. It's gets really strange out
- 2:43:36there but but then if we find see a few examples of something
- 2:43:41very odd and this one, very some force, a very odd pictures.
- 2:43:46We've seen then, in order to train it effectively we want to
- 2:43:51create simulations. Tape thousand simulations that
- 2:43:55are that are variants of that quirky thing that we saw the
- 2:43:59foot to filling the quit of some important gaps and make the
- 2:44:03system better. And really all of this is about
- 2:44:05overtime, just reducing the probability of a crash or an
- 2:44:10injury and it's good that the march of nines like, how do you
- 2:44:14get to 99.999999% safe, you know, and it.
- 2:44:22Yeah, each night isn't it? ER, of magnitude difficulty
- 2:44:25increase. Hey, thanks so much for the
- 2:44:30presentation. I was curious about the Tesla
- 2:44:33bot. Specifically, I'm wondering if
- 2:44:36there are any specific applications that you think the
- 2:44:39humanoid form factor lends itself to and secondary Because
- 2:44:46of its human form factor is emotion or companionship at all.
- 2:44:50Thought about on the product roadmap at all.
- 2:45:00We certainly hope this does not feature in a dystopian sci-fi
- 2:45:04movie. But you know, like really at
- 2:45:11this point we're saying like maybe this robot can just we're
- 2:45:15trying to we're trying to as little as possible.
- 2:45:17Can it do boring dangerous repetitive jobs that people
- 2:45:22don't want to do? And You know, once you can have
- 2:45:27it do that and maybe could do other things too.
- 2:45:29But that's the that's the thing that we really great to have.
- 2:45:34so, It could be a buddy to, if my, by want to have, to be your
- 2:45:39friend, and whatever for the people think of some very
- 2:45:45creative uses. So, so, Firstly, thanks for the,
- 2:45:56the really incredible presentation.
- 2:45:58My questions on the AI side. So one thing we've been seeing
- 2:46:02is that with some of these language modeling a eyes, we've
- 2:46:05seen that scaling has just had incredible impacts in their
- 2:46:08capabilities and what they're able to do.
- 2:46:11So I was wondering, whether you're seeing similar kinds of
- 2:46:13effects of scaling in your neural networks in your
- 2:46:16applications. Absolutely a bigger Network.
- 2:46:20Typically we see it performs better provided you have the
- 2:46:22data to also train it with. This is also what we see for
- 2:46:25ourselves, definitely in the car, we have some latency
- 2:46:28consideration to be mindful of and so there we have to get
- 2:46:31creative to actually deploy much much, larger networks.
- 2:46:34But as we mentioned, we don't only train your own networks for
- 2:46:37what goes in the car. We have these are labeling
- 2:46:40pipelines that can utilize models of arbitrary size.
- 2:46:42So in fact, we trade the number of models that are not
- 2:46:44Deployable. There are significantly larger
- 2:46:46and work much better because we want 100% likely We want much
- 2:46:49higher accuracy for the auto labeling and so we've done a lot
- 2:46:52of that. And there we definitely see this
- 2:46:53trend. Yeah.
- 2:46:55That the order of labeling is an extremely important part of this
- 2:47:00whole situation without the or labeling.
- 2:47:04I think we would not be able to solve the self-driving problem.
- 2:47:07It's Kind of a Funny form of distillation where you're using
- 2:47:09these very massive models plus the structure of the problem to
- 2:47:13do this reconstruction and then you distill that into neural
- 2:47:15networks that you deploy to the car and we basically have a lot
- 2:47:18of neural networks and a lot of tasks that are never intended to
- 2:47:20go into the car. And also as time goes on that
- 2:47:25you get new friends information. So you really want to make sure
- 2:47:27your computer has disappeared across all the information as
- 2:47:30opposed to just taking a single frame and hugging on it for say
- 2:47:33200 milliseconds, you actually will have new frames coming in.
- 2:47:35So you want to use all of the information and not just use
- 2:47:38that one frame. I think it was one of the things
- 2:47:42we're seeing is that the cause predictive ability is, is quite
- 2:47:46as eerily good. It's really getting better than
- 2:47:50human in terms of predicting like you said like what predict
- 2:47:54what this road will look like when it's out of sight, like
- 2:47:58it's around the bend and predicts the road with very high
- 2:48:01accuracy and protect pedestrians or cyclists.
- 2:48:07We're behind. You know, where it's just these
- 2:48:09little corner of the bicycle and the little bit Through the
- 2:48:12Windows of the bus, is its ability to predict things is
- 2:48:16going to be much better than humans.
- 2:48:18Like, really wait, Way, Beyond right here.
- 2:48:21We see this often where we have something that is not visible
- 2:48:23but the neural network is making up stuff.
- 2:48:25That actually is very sensible. Sometimes it's eerily good and
- 2:48:28you have to like you're wondering this isn't a training
- 2:48:30set and actually actually in the limit you can imagine the neural
- 2:48:34net has enough parameters to potentially remember Earth.
- 2:48:36So in the limit, it could actually give you the correct
- 2:48:39answer. Sorry.
- 2:48:39It's kind of like an HD map back baked into the weights of the
- 2:48:42neural nut. Yeah.
- 2:48:50I have a question about the design of the Tesla bought
- 2:48:53specifically in order. How is it important?
- 2:48:56Is it to maintain that humanoid form to build hands with five
- 2:49:01fingers that also respects the weight limits could be quite
- 2:49:05challenging. You might have to use cable
- 2:49:06driven and then that also causes all kinds of issues.
- 2:49:12I mean this is just going to be but version one, I mean we'll
- 2:49:15see. So the it's it needs to be able
- 2:49:20to do things that people do. And there be a generalized.
- 2:49:26Get a humanoid robot. I mean, you could make
- 2:49:32potentially have give it like, you know, two fingers and a
- 2:49:35thumb or something like that, you know, for now, we'll give it
- 2:49:39five fingers and, and see see if that works out, okay?
- 2:49:43Probably will. It doesn't need to be like, you
- 2:49:45know, have like an incredible grip strength, but it needs to
- 2:49:49be able to work with tools. So, I carry a bag that kind of
- 2:49:54thing. Alright, thanks a lot for the
- 2:49:58presentation. So an old professor of mine told
- 2:50:02me that the thing he disliked a lot about his Tesla was that the
- 2:50:05autopilot ux didn't really Inspire much confidence in the
- 2:50:08system. Especially one like objects are
- 2:50:09spinning. Classifications are flickering.
- 2:50:12I was wondering, like, even if you have a good self-driving
- 2:50:16system, how are you working on convincing Tesla owners other
- 2:50:20Road users or other Road users and just the general public that
- 2:50:24your system is safe and reliable.
- 2:50:27Well, I think that's that's the cars while back cars you
- 2:50:31suspend, they don't they don't spend any more not in.
- 2:50:33If you're seen the FSD beta videos, they are pretty solid
- 2:50:39and they will be getting more solid.
- 2:50:41Yeah, I think I had more and more data and train these multi
- 2:50:44camera and works. Like these are pretty recent
- 2:50:45artistic few months old and I still improving.
- 2:50:47It's not done product and we know Minds, we can clearly see
- 2:50:51how this just going to be like perfect.
- 2:50:53Perfect way to space because why not all the information is there
- 2:50:56in? Videos, it should produce a
- 2:50:57given lots of data and good architectures and this is just
- 2:51:02intermediate point in the timeline.
- 2:51:05It's pretty, it's clearly headed to way better than even without
- 2:51:08question. Return height here.
- 2:51:17I was wondering if you could talk a little bit about the
- 2:51:19short to medium term economics of the, but I guess I understand
- 2:51:24the long-term vision of replacing physical labor, but I
- 2:51:28also think repetitive dangerous and boring tasks tend to not be
- 2:51:33so highly compensated. So I just don't see how to
- 2:51:36reproduce, you know, start with a Supercar and then break into
- 2:51:40like the lower end of the market.
- 2:51:42How do you do that for our robot humanoid?
- 2:51:44No. Well, I guess you'll just have
- 2:51:47to see. Hello.
- 2:51:58Hi. I was curious to know how the
- 2:52:01car guy. Prioritises, occupant safety
- 2:52:04versus pedestrian safety. And what thought process goes
- 2:52:07into? Like deciding how to bake this
- 2:52:09into the AI? Well we the thing to appreciate
- 2:52:18is that from the computer standpoint, everything is moving
- 2:52:21slowly. So think, you know, to a human
- 2:52:26things are moving fast to the computer, they are not moving
- 2:52:28fast. So I think this is in reality
- 2:52:31somewhat of a false dichotomy not that it will never happen,
- 2:52:34but it will be very rare. You know if you think it was
- 2:52:38like you know going the other direction like rendering you
- 2:52:43know with full Ray tracing neural net enhanced graphics on
- 2:52:48something like cyberpunk or in any you know, Advanced video
- 2:52:52game you know doing 60 frames a second perfectly rendered.
- 2:52:58Like how long would it take a person to the even render one
- 2:53:00frame? And without any mistakes can't
- 2:53:05be done. I mean, we're take like a month
- 2:53:08just to just render 11 frame out of 60 in a second in a video
- 2:53:14game. It's computers are fast and
- 2:53:19humans are slow. I mean, for example, on the
- 2:53:27rocket side, the you cannot steer the rocket to orbit.
- 2:53:33We actually hooked up a joystick to see if anyone could steal the
- 2:53:36rocket orbit, but you need to react at roughly 67 hurts.
- 2:53:44People can't do it. Not even a that's pretty low,
- 2:53:48you know, we're talking while like even for like 30 Hertz type
- 2:53:53of thing. Hi with the over here with
- 2:54:01Hardware three. There's been lots of speculation
- 2:54:03that with larger Nets. It's hitting the limits, of what
- 2:54:06it can provide, how much Headroom has the extended
- 2:54:09compute modes provided and what point would hardware for be
- 2:54:12required if at all Well, I'm confident that Hardware three or
- 2:54:19four self-driving. Computer one will be able to
- 2:54:23achieve full surviving at a safety level much greater than a
- 2:54:27human probably. I don't know at least two or
- 2:54:29three hundred percent better than a human.
- 2:54:32Then obviously there will be a future hardware for or
- 2:54:35self-driving computer to which will probably introduced with
- 2:54:38the Cyber truck. So maybe in about a year or so
- 2:54:44that is Probably well they'll be about four times more capable
- 2:54:49roughly but it's really going to be like can we take it from say
- 2:54:56for argument's sake 300% if than a person to 1000% save Earth
- 2:55:00you're just like there are people on the road who with with
- 2:55:04varying driving abilities but we still let people drive it.
- 2:55:08You don't have to be the world's best driver to be on the road.
- 2:55:12So, as we see so yeah. All right.
- 2:55:23So are you worried at all, since you don't have any depth sensors
- 2:55:26on the car, that people might try like adversarial attacks,
- 2:55:30like print it out photos or something to try to trick the
- 2:55:34RGB neural network. Yeah.
- 2:55:38Like or the pull some like Wiley Coyote stuff and like paint the
- 2:55:42tunnel on the on the wall. It's like off.
- 2:55:49We haven't really seen much of that, I mean.
- 2:55:54For sure. Like like right now, if you Most
- 2:55:58likely if you had like a t-shirt with us that there's t-shirt,
- 2:56:01would like a stop sign on it, which I actually have a t-shirt
- 2:56:03with a stop sign on it and and then you like flash the car, it
- 2:56:09will, it will stop. I proved that.
- 2:56:16But we can obviously as we see these adversarial tax then we
- 2:56:20can retrain the car's too. You know, notice that will it's
- 2:56:25actually a person wearing a t-shirt.
- 2:56:27The stop sign on. It's probably not real.
- 2:56:29Stop sign. Hi, my question is about the
- 2:56:38prediction and the planning I'm curious.
- 2:56:41How do you incorporate uncertainty into your planning
- 2:56:45algorithms? Do you just basically assume you
- 2:56:49know you mentioned that you run the autopilot for all of the
- 2:56:52other cars on the road. Do you assume that they're all
- 2:56:55going to follow those rules or you accounting for the
- 2:56:57possibility that well they might be bad drivers for example.
- 2:57:02If you do, I'll confirm our commodity Futures.
- 2:57:04It's not that we just choose one.
- 2:57:06We account for this person can actually do many things and we
- 2:57:10use that actual physics and kinematics to make sure that
- 2:57:13they're not doing a thing that would interfere with us before
- 2:57:16we act. So, if there's any uncertainty
- 2:57:18we are conservative and then or deal to them, of course, there's
- 2:57:21a limit to this because if you have to consider then it's
- 2:57:24probably not practical. So, at some point we had to
- 2:57:26assert and we even then we make sure that the other person can
- 2:57:30heal to us and accent. Sibley.
- 2:57:36I should say like like before we introduce something into the
- 2:57:40fleet, we will run it in Shadow mode and so and we'll see what
- 2:57:46would this neural net for example have done in this
- 2:57:50particular situation because and then effectively the drivers are
- 2:57:55training training the net. So if the neural net would have
- 2:58:00controlled and you know, and say beard right?
- 2:58:02But the person actually You went left.
- 2:58:04It's like oh there was a difference.
- 2:58:06Why was there that difference? And secondly, obviously all the
- 2:58:11human drivers are essentially training, the neural net as to
- 2:58:15what is the correct course of action for producing?
- 2:58:18It doesn't then ended up in a crash, you know, doesn't count
- 2:58:21in that case. Yeah, and secondly we have
- 2:58:22various estimates of uncertainty, like Flickr.
- 2:58:25And when we observe this, we actually save, you're not able
- 2:58:29to see something. We actually slow down the car to
- 2:58:31be again, safe and get more information before acting if you
- 2:58:35don't want to be Brazen, and it's going to something that we
- 2:58:36don't know about. We only go into places, where we
- 2:58:38know about About. Yeah.
- 2:58:42Yeah. It should be like the
- 2:58:44aspirationally that the car should be the less.
- 2:58:46It knows the slit, you know, the slower.
- 2:58:48It goes. Yeah.
- 2:58:49That's just not true at some point but now.
- 2:58:52Yeah, yeah. We've yeah, should be speed
- 2:58:55proportionate to confidence. Thanks for the presentation.
- 2:59:05So I am curious, appreciate the fact that the FSD is improving,
- 2:59:11but if you have the ability to improve one component along the,
- 2:59:15I stack the presented today, whether it is simulation data
- 2:59:19collections planning and control Etc.
- 2:59:21Which one in your opinion is going to have the biggest impact
- 2:59:24for the performance of the full self-driving.
- 2:59:27Justin. It's really the area under the
- 2:59:34curve of this like multiple points and if you increase or
- 2:59:36improve anything measuring for the area, I mean the short term,
- 2:59:41it's arguably. We need all of the Nets to be
- 2:59:45surround video and so, we still have some Legacy.
- 2:59:48This is a very short term. Obviously, we're fixing it fast.
- 2:59:51But there's, there's still some Nets that are not using surround
- 2:59:54video and ideally, they're all use around video.
- 3:00:00Yeah. Very yeah, I think a lot of
- 3:00:02puzzle pieces are there for Success.
- 3:00:03We just need more strong people. To also just help us make it
- 3:00:07work. Yeah, I didn't actually box.
- 3:00:08The bet is the actual bottom. Like I would say, I'm really one
- 3:00:10of the reasons that we are putting on this event exactly
- 3:00:13what well said, Andre that there's just a tremendous amount
- 3:00:16of work to do to make, make it work.
- 3:00:19So that's why we need talent people to join and solve the
- 3:00:24problem. Thank you for the great
- 3:00:32presentation, most of my questions answered, but one
- 3:00:35thing is, when imagine that now you have a large amount of data
- 3:00:42even unnecessary. How do you consider that?
- 3:00:45Like, there's a for getting problem in neural networks.
- 3:00:48Like, how are you considering those aspects?
- 3:00:52And also another one, are you considering an online learning
- 3:00:56or continuous learning so that maybe Each driver can have their
- 3:01:01version of self-driving. I think, I think I know the
- 3:01:08literature that you're referring to.
- 3:01:09That's not some of the problems that we've seen, and we haven't
- 3:01:11done too much. Continuous learning.
- 3:01:13We train the system. Once we find in the few times,
- 3:01:16that sort of goes into the car, we need something stable that we
- 3:01:18can evaluate extensively, and then we think that that's good
- 3:01:21and that goes into cars. So we don't do too much learning
- 3:01:24on spot or continuous learning and don't face the for getting
- 3:01:26problem. But there will be settings that
- 3:01:29you can say. Like if you do you want are used
- 3:01:31typically a conservative driver or do you want to drive fast or
- 3:01:34slow? You know, it's like I'm late for
- 3:01:36my I'm late for the airport. Could you go faster than you
- 3:01:39know basically the kind of instructions you give to you a
- 3:01:41driver it's like a million late for the flight please hurry or
- 3:01:45take it easy or whatever your style is A few more questions
- 3:01:56here, so then we'll call it a day.
- 3:01:58Alright. So as our models have become
- 3:02:00more and more capable, and I guess you're deploying these
- 3:02:04models into the real world. One thing I guess that's
- 3:02:06possible is for AI to become more, I guess, Miss aligned with
- 3:02:10what humans desire. So I guess is that something
- 3:02:13that you guys are worried about, is you guys deploy more and more
- 3:02:15robots, or do you guys like will solve that problem when we got?
- 3:02:19There. Yeah, I think that we should be
- 3:02:23worried about a. I know like what we're trying to
- 3:02:28do here is say a narrow AI threatened, pretty narrow, like
- 3:02:32just make the car drive better than a human and then have the
- 3:02:37humanoid robots be able to do basic stuff, you know.
- 3:02:43So at the point at which you sort of talk, get to superhuman
- 3:02:48intelligence, I don't know. All bets are off. but, you know,
- 3:02:55that's you know, That'll that'll probably happen but but what
- 3:02:59we're trying to do here at Tesla is make useful AI that people
- 3:03:04love and is unequivocally good. That's our, you know, try to aim
- 3:03:10for that. Tell me one question.
- 3:03:17My question is about the camera sensor.
- 3:03:19In the beginning of the talk, you had mentioned about building
- 3:03:22a synthetic animal and if you think about it, as a camera is a
- 3:03:26very poor approximation of a human eye, and the human eye
- 3:03:29does lot more than take a sequence of frames have.
- 3:03:33You looked into like The Relic? These days are like cameras,
- 3:03:36like, even cameras, have you look into them?
- 3:03:38Or are you looking into a more flexible camera, design, or
- 3:03:41building your own camera for example?
- 3:03:46Well, with hardware for we will have a next-generation camera
- 3:03:50but I have to say that the current cameras we have not
- 3:03:53reached the limit of the current cameras.
- 3:03:55So and I'm confident we can achieve for self-driving with
- 3:04:01much higher safety than humans with the current cameras and
- 3:04:04current computer hardware. But you know, very good to be
- 3:04:111000% better rather than 300 better.
- 3:04:14So we will see continued Evolution on all levels in
- 3:04:19pursuit of that goal. And I think in the future people
- 3:04:22will look back and say, wow, I can't believe we have to drive
- 3:04:25these cars ourselves. You know, it's it like
- 3:04:28self-driving cars will just being just a normal like
- 3:04:31self-driving elevators, you know, elevators used to have
- 3:04:34elevator operators and there's someone there would like, you
- 3:04:37know, big big relay switch operating.
- 3:04:40Later. And then every now and then that
- 3:04:42get tired or you know some make a mistake and sure somebody to
- 3:04:46have. So so now we you know, we made
- 3:04:50elevators automatic and you just go and you press the button and
- 3:04:53you can be in 100 story skyscraper and don't really
- 3:04:57worry about it. Just go ahead and press a button
- 3:04:58and the elevator takes you where you want to go.
- 3:05:02But it used to be that all elevators were operated manually
- 3:05:04manually, it'll be the same thing.
- 3:05:05Like four cars, all cars will be automatic and then A tent
- 3:05:11electric obviously. So there will be some gasoline
- 3:05:15cars and some manual cars just like there are still some
- 3:05:17horses. So all right, well thanks
- 3:05:22everyone for coming and I hope you enjoyed the presentation and
- 3:05:26thanks for the great questions. All right?