Latest / Elon Musk Podcast / Elon Musk's Xai grok 4 announcement presentation
Transcript
- 0:01Hey everybody. Welcome back to the Elon Musk
- 0:04Podcast. This is a show where we discuss
- 0:07the critical crossroads, the Shape, SpaceX, Tesla X, The
- 0:11Boring Company, and Neurolink. I'm your host, Will Walden.
- 0:16All right, welcome to the Grok 4 release here.
- 0:20This is the smartest AI in the world, and we're going to show
- 0:23you exactly how and why. And it really is remarkable to
- 0:30see the advancement of artificial intelligence, how
- 0:33quickly it is evolving, I sometimes think, compare it to
- 0:42the growth of a human and how fast a human learns and gains
- 0:49conscious awareness and understanding.
- 0:51And AI is advancing just vastly faster than any human.
- 1:00I mean, we're going to take you through a bunch of benchmarks
- 1:04that that Grok 4 is able to achieve incredible numbers on.
- 1:10But it's, it's, it's actually worth noting that like Grok 4,
- 1:15if, if given, like the SAT would get perfect Sats every time,
- 1:20even if it's never seen the, the questions before.
- 1:24And if even going beyond that to say like graduate student exams,
- 1:29like the GRE, it will get near perfect results in, in every
- 1:36discipline of, of education. So from the humanities to like
- 1:41languages, math, physics, engineering, pick anything.
- 1:46And we're talking about questions that it's never seen
- 1:49before. These are not on, not on the
- 1:50Internet. And it's Grok 4 is smarter than
- 1:56almost all graduate students in all disciplines simultaneously.
- 2:02Like it's actually just important to appreciate the like
- 2:06that's really something. And the, the reasoning
- 2:13capabilities of Grok are incredible.
- 2:16So there's some people out there who who think AI can't reason
- 2:19and look, it can reason at superhuman levels.
- 2:24So yeah. And frankly, it only gets better
- 2:28from here. So we'll, we'll take you through
- 2:32the Grok 4 release and, yeah, cheer you back the pace of pace
- 2:40of progress here. Like, I guess the first part is
- 2:44like in terms of the training, we're going from Grok 2 to Grok
- 2:493 to Grok 4, we've essentially increased the training by an
- 2:53order of magnitude in each case. So it's a, you know, 100 times
- 2:58more training than than Grok 2. And and that that's only going
- 3:03to increase. So it's, yeah, frankly, I mean,
- 3:09I don't know, in some ways a little terrifying, but the
- 3:14growth of intelligence here is, is remarkable.
- 3:16Yes. It's important to realize there
- 3:18are two types of training compute. 1 is the pre training
- 3:21compute that's from grad 2 to grad 3.
- 3:24But for from grad 3 to grad 4, we're actually putting a lot of
- 3:28compute in reasoning in IR. Yeah, and just like you said,
- 3:34this is literally the fastest moving field and Grad 2 is like
- 3:37the high school student by today's standard.
- 3:39If you look back in the last 12 months, Grad 2 was only a
- 3:43concept. We didn't even have graph 212
- 3:46months ago. And then by training gratitude,
- 3:49that was the first time we scale up like the pre training.
- 3:51We realized that if you actually do the data oblation really
- 3:54carefully and the infra and also the algorithm, we can actually
- 3:59push the pre training quite a lot by amount of 10X to make the
- 4:03model the best pre trained based model.
- 4:06And that's why we build classes, the world's supercomputer with
- 4:09100,000 H 100 and then with the best pre trained model.
- 4:14And we realized if you can collect these verifiable outcome
- 4:17reward, you can actually train this model to start thinking
- 4:20from the first principle, start to reason, correct it's own
- 4:22mistakes. And that's where the graphical
- 4:23reasoning comes from. And today we ask the question,
- 4:27what happens if you take expansion of clauses with all
- 4:30200,000 GPUs, put all these into oil, 10X more compute than any
- 4:37of the models out there on reinforcement learning,
- 4:40unprecedented scale, What's going to happen?
- 4:42So this is a story of Grok 4 and you know Tony.
- 4:47Share some insight with the audience.
- 4:50Yeah. So, yeah, let's just talk about
- 4:53how smart Grok 4 is. So I guess we can start
- 4:57discussing this benchmark on humanities last exam.
- 5:00And this this benchmark is a very, very challenging
- 5:03benchmark. Every single problem is curated
- 5:06by subject matter experts. It's in total 2500 problems and
- 5:13it consists of many different subjects, mathematics, natural
- 5:16sciences, engineering and also all the humanity subjects.
- 5:21So essentially when when it was first released actually like
- 5:25earlier this year, most of the models out there can only get
- 5:30single digit accuracy on this benchmark.
- 5:34Yeah. So we can we can look at some of
- 5:36those examples, you know. So there there is this
- 5:40mathematical problem which is about natural transformations in
- 5:44category theory. And there's this organic
- 5:47chemistry problem that talks about electrical cyclic
- 5:51reactions. And also there's this linguistic
- 5:54problem that tries to ask you about this distinguishing
- 5:57between closed and open syllabus from a Hebrew source text.
- 6:03So you can see also it's a very wide range of problems and every
- 6:08single problem is PhD or even Advanced Research level
- 6:12problems. Yeah.
- 6:14I mean, these, there are no humans that can actually answer
- 6:17these can get a good score. I mean, if you actually say like
- 6:20any given human, what, like what's the best that any human
- 6:25could score? I mean, I'd say maybe 5%
- 6:29optimistically. Yeah.
- 6:32So this, this is much harder than than what any any human can
- 6:36do. It's it's incredibly difficult.
- 6:38And you can see from the types of questions like you might be
- 6:41incredible in linguistics or mathematics or chemistry or
- 6:44physics or any one of a number of subjects, but you're not
- 6:46going to be at a post grad level in everything.
- 6:51And Grok 4 is a post grad level in everything.
- 6:55Like it's it's just some of these things are just worth
- 6:57repeating. Like Grok 4 is postgraduate.
- 7:01Like PhD level in everything better than pH, but like most
- 7:05PHD's would fail so it's better. That said, I mean at least with
- 7:10respect to academic questions. It I want to just emphasize this
- 7:13point with respect to academic questions.
- 7:16Group 4 is better than PhD level in every subject, no exceptions.
- 7:23Now, this doesn't mean that it's, you know, times.
- 7:26It may lack common sense and it has not yet invented new
- 7:31technologies or discovered new physics.
- 7:35But that is just a matter of time.
- 7:37If it, I think it may discover new technologies as soon as
- 7:43later this year. And I would be shocked if it is
- 7:47not done so next year. And so I would expect Croc to,
- 7:51yeah, literally discover new technologies that are actually
- 7:53useful no later than next year and maybe end of this year.
- 7:57And it might discover new physics next year.
- 8:00And within two years, I'd say almost certainly like.
- 8:04So just let that sink in. Yeah, so.
- 8:15Yeah, how? OK, So I guess we can talk about
- 8:19the the what, what's behind the scene of Graph 4?
- 8:21As Jimmy mentioned, we actually saw in a lot of compute into
- 8:26this training, you know, when it started, it's only also a single
- 8:31digit. Sorry, the previous slide,
- 8:33sorry. Yeah, it's only a single digit
- 8:37number. But as you start putting in more
- 8:40and more training compute, it started to gradually become
- 8:44smarter and smarter and eventually solved 1/4 of the HLA
- 8:49problems. And this is without any tools.
- 8:53The next thing we did was to adding a tools capabilities to
- 8:57the model. And unlike GRAS three, I think
- 9:01you know, GRAS 3 actually is able to use cruel as well.
- 9:03But here we actually make it more native in the sense that.
- 9:07We put the. Tools into training.
- 9:10But graph 3 was only relying on generalization.
- 9:12Here we actually put the tools into training and it turns out
- 9:16this significantly improves the models capability of using those
- 9:19tools. Yeah, I remember we had like
- 9:22deep search back in the days. Yeah.
- 9:24So how is this different? Yeah, yeah, yeah.
- 9:26Exactly. So Deep Search was exactly the
- 9:29Graph 3 reasoning model but without any specific training.
- 9:33But we only asked it to use those tools.
- 9:36So compared to this, it was much weaker in terms of its tool
- 9:40capabilities and reliable and unreliable, I guess, yes.
- 9:44And to be clear, like these are still, I'd say fairly, this is
- 9:47still fairly primitive tool use. If you compare it to say the
- 9:51tools that are used at Tesla or SpaceX where you're using, you
- 9:55know, finite element analysis and computational fluid dynamics
- 10:00and, and you're you're able to run or say like Tesla does, like
- 10:05crash simulations, where the simulations are so close to
- 10:08reality that if the test doesn't match the simulation, you assume
- 10:12that the test article is wrong. That's how good the simulations
- 10:15are. So Crock is not currently using
- 10:17any of the tools that the really powerful tools that a company
- 10:21would use, but but that is something that we will provide
- 10:24it with later this year. So we'll have the tools that
- 10:28that a company has and and have very accurate physics simulator.
- 10:34Ultimately, the, the thing that'll make the biggest
- 10:37difference is being able to interact with the real world via
- 10:40humanoid robots. So you combine sort of grok
- 10:42with, with Optimus and it can actually interact with the real
- 10:45world and figure out if, if it's high, if it has, if it's
- 10:49formulate and hypothesis, and then confirm if that hypothesis
- 10:54is, is true or not. So we're really, you know, I
- 11:00think about like where we are today.
- 11:01We're at the beginning of an immense intelligence explosion.
- 11:06We're in, we're in the intelligence Big Bang right now
- 11:12and the we're at the most interesting time to be alive of
- 11:16any time in history. Yeah, now that's it.
- 11:21We need to make sure that the AI is a good AI, good croc.
- 11:29And the thing that I think is most important for AI safety, at
- 11:32least my biological neural net tells me the most important
- 11:36thing for AI is to be maximally truth seeking.
- 11:40So this is this is a very fundamental, but you can think
- 11:45of AI as this super genius child that ultimately will outsmart
- 11:50you. But you can still and you can
- 11:52install the right values and encourage it to be sort of, you
- 11:58know, truthful, I don't know, honorable, you know, good, good
- 12:06things like the values you want to instill in in a child that
- 12:13that that would grow, ultimately grow up to be incredibly
- 12:17powerful. Yeah.
- 12:24So, yeah. So this is really, I'd say we're
- 12:27saying, we say tools. These are still primitive tools,
- 12:30not the kind of tools that that serious commercial companies
- 12:35use, but we will provide it with those tools and it, I think it
- 12:39will be able to solve with those tools real world technology
- 12:42problems. In fact, I'm certain of it.
- 12:44It's just a question of how long it takes.
- 12:47Yes, yes, exactly. So is it just compute all you
- 12:51need, Tony, sorry, is it just compute all you need at this
- 12:55point? Well, you need compute plus plus
- 12:59the right tools and and then ultimately to be able to
- 13:04interact with the physical world, yes.
- 13:07And then I mean we'll effectively have an economy that
- 13:12is well, ultimately an economy that is thousands of times
- 13:18bigger than our current economy or maybe millions of times.
- 13:22I mean, if you if you think of civilization as percentage
- 13:26completion of the Khadashev scale where Karashev one is
- 13:30using all the energy output of a planet and Karashev 2 is using
- 13:35all the energy output of a sun and three is all the energy
- 13:38output of a Galaxy. We're we're only in my opinion
- 13:42probably like close closer to 1% of Karashev one then we are up
- 13:48to 10%. So like maybe a pointer 11 or 2%
- 13:54of Karashev 1. So we we will get to most of the
- 14:02weight like 8090% Kaddashev 1 and then hopefully if
- 14:05civilization doesn't self annihilate and then Kadashev 2.
- 14:11Like it's the, the actual notion of, of a human economy assuming
- 14:15civilization continues to progress will seem very quaint
- 14:19in in retrospect, it will, it will seem like sort of Cavemen
- 14:25throwing sticks into a fire level of economy compared to
- 14:30what the future will hold. I mean, it's very exciting.
- 14:34I mean, I, I've been at times kind of worried about like,
- 14:38well, you know, is this, this seems like it's somewhat
- 14:45unnerving to have intelligence created that is far greater than
- 14:49our own. And will this be bad or good for
- 14:54humanity? It's like, I, I, I think it'll
- 14:59be good. Most likely it'll be good.
- 15:03Yeah. Yeah.
- 15:07But if someone reconciled myself to the fact that even if I, if,
- 15:12even if it wasn't going to be good, I'd at least like to be
- 15:15alive to see it happen. So you know so.
- 15:22Actually, one, yeah, yeah. I think what 1 technical problem
- 15:28that we still need to solve besides just compute is how do
- 15:31we unblock the data, data bottleneck.
- 15:35Because when we try to scale up the RL, in this case, we did
- 15:41invent a lot of new techniques, innovations to allow us to
- 15:45figure out how to find a lot of a lot of challenging RL problems
- 15:50to work on. It's not just a problem itself
- 15:52needs to be challenging, but also it needs to be you.
- 15:55You also need to have like a reliable signal to tell the
- 15:59model you did it wrong, you did it right.
- 16:01This is the sort of the principle of reinforcement
- 16:03learning. And as the models get smarter
- 16:06and smarter, the number of cool problem or challenging problems
- 16:10will be lesser. Unless, yeah.
- 16:12So it's going to be a new type of challenge that we need to
- 16:16surpass besides just compute. Yeah, yeah.
- 16:19We actually are running out of of actual test questions to ask.
- 16:23So there's like even ridiculously questions that are
- 16:26ridiculously hard, if not essentially impossible for
- 16:29humans that are written down questions are becoming swiftly
- 16:35becoming trivial for for AI. So then there's, but you know
- 16:42what? The one thing that is an
- 16:44excellent judge of things is reality.
- 16:46So because if physics is the law, ultimately everything else
- 16:50is a recommendation. You can't break physics.
- 16:53So the ultimate test, I think for whether an AI is the
- 16:58ultimate reasoning test is reality.
- 17:01So you invent a new technology like say, improve the design of
- 17:04a car or a rocket or create a new medication that and, and,
- 17:11and does it work? Yeah.
- 17:14Does does the rocket get to orbit?
- 17:15Does the does the car drive? Does the medicine work?
- 17:19Whatever the case may be, reality is the ultimate judge
- 17:22here. So it's going to be a
- 17:25reinforcement learning closing loop around reality.
- 17:32We asked the question how do we even go further?
- 17:34So actually we are thinking about now with single agent we
- 17:40are able to solve 40% of the problem.
- 17:42What if we have multiple agents running at the same time?
- 17:46So this is what's called test and compute.
- 17:49And as we scale up the test and compute, actually we are able to
- 17:53solve almost more than 50% of the text only subset of the HIV
- 17:58problems. So it's a remarkable
- 18:01achievement. I think you know this.
- 18:03Is this is this is insanely difficult.
- 18:05These are it's it's what we're saying is like a majority of the
- 18:09of the of the text based of humanities, you know, scarily
- 18:14named humanities last exam Grafoe can solve and you can you
- 18:18can try it out for yourself. And the, with the Grafoe heavy,
- 18:22what, what it does is it spawns multiple agents in parallel.
- 18:25And all of those agents do, do work independently.
- 18:29And then they compare their work and they, they decide which one
- 18:34like, it's like a study group. And it's not as simple as
- 18:38majority vote because often only one of the agents actually
- 18:42figures out the trick or figures out the solution.
- 18:46And, and, and, but once they share the, the trick or, or
- 18:50figure out what, what the real nature of the problem is, they
- 18:53share that solution with the other agents and then they
- 18:56compare. They essentially compare notes
- 18:59and then, and then yield, yield an answer.
- 19:02So that's, that's the, the heavy part of Grok 4 is, is where we,
- 19:06you scale up the test time, compute by roughly an order of
- 19:09magnitude, have multiple agents tackle the task and then they
- 19:15compare their work and they, they put forward what they think
- 19:19is the best result. Yeah.
- 19:22So we will introduce graph 4 and graph 4 Heavy so you can click
- 19:26the next slide. Yeah, yes.
- 19:29So yeah. So basically Graph 4 is a single
- 19:33agent version and Graph 4 heavy is the multi agent version.
- 19:38So let's take a look how they actually do on those exam
- 19:43problems and also some real real life problems.
- 19:45Yeah, So we're going to start out here and we're actually
- 19:48going to look at one of those HLE problems.
- 19:50This is actually one of the easier math ones.
- 19:53I don't really understand it very well.
- 19:55I'm not that smart, but I can launch this job here and we can
- 19:58actually see how it's going to go through and start to think
- 20:01about this problem. While we're doing that, I also
- 20:04want to show a little bit more about like what this model can
- 20:06do and launch a Grok 4 Heavy as well so everyone knows
- 20:11Polymarket. It's extremely interesting.
- 20:13It's the, you know, seeker of truth.
- 20:15It aligns with what reality is most of the time.
- 20:18And with Grok, what we're actually looking at is being
- 20:21able to see how we can try to take these markets and see if we
- 20:25can predict the the future as well.
- 20:27So as we're letting this run, we'll see how Grok for Heavy
- 20:31goes about predicting the, you know, the World Series odds for
- 20:35like the current teams and the MLB.
- 20:38And while we're waiting for these to process, we're going to
- 20:40pass it over to Eric and he's going to show you an example of
- 20:43his. Yeah.
- 20:48So I guess one of the coolest things about Grok 4 is its
- 20:53ability to understand the world and to solve hard problems by
- 20:57leveraging tools like Tony discuss.
- 21:00And I think one kind of cool example of this, we asked it to
- 21:03generate a visualization of two black holes colliding.
- 21:09And of course, you know, it took some, there are some liberties.
- 21:13It's in my case, actually pretty clear in its thinking trace
- 21:16about what these liberties are. For example, in order for it to
- 21:19actually be visible, you need to really exaggerate the the scale
- 21:23of the, you know, the the the waves.
- 21:29And yeah, so here's like, you know, this kind of inaction, it
- 21:35exaggerates the scale in like multiple ways.
- 21:38It drops off a bit less in terms of amplitude it over distance.
- 21:43And but yeah, we can kind of see the basic effects that, you
- 21:49know, are actually like, you know, correct.
- 21:51It starts with the in spiral, it merges and then you have the
- 21:57ring down. And like, this is basically
- 22:01largely correct. Yeah, module some of the
- 22:06simplifications that need to do, you know, it's actually quite
- 22:10explicit about this. You know, it uses like post,
- 22:13post Newtonian approximations instead of actually like
- 22:16computing. The general relativistic effects
- 22:19are like near the center of the black hole, which is, you know,
- 22:22incorrect and, you know, will lead to, you know, some
- 22:26incorrect results. But the overall, you know,
- 22:28visualization is yeah, is basically there and you can
- 22:33actually look at the kinds of resources that are references.
- 22:36So here it it actually, you know, it obviously is a search.
- 22:41It gathers results from a bunch of links, but also reads through
- 22:45a undergraduate text in analytical analytic
- 22:50gravitational wave models. It's, yeah, it reasons quite a
- 22:57bit about the actual constants that it should use for a
- 23:02realistic simulation. It references, I guess, existing
- 23:06real world data. And yeah, it, yeah, it's a, it's
- 23:12a pretty good model, yeah. But like actually going forward,
- 23:17we can, we can, we can plug, we can give it the same model that
- 23:20physicists use so it can run the the same level of compute that
- 23:26so leading physics researchers are using and and give you a
- 23:30physics accurate backhoe simulation.
- 23:32Exactly. Just right now is running in
- 23:35your browser, so. Yeah, this is just running in
- 23:37your browser exactly. Pretty simple.
- 23:39So swapping back real quick here, we can actually take a
- 23:42look at the math problem is finished.
- 23:44The model was able to let's look at its thinking trace here so
- 23:48you can see how it went through the problem.
- 23:51I'll be honest with you guys, I really don't quite fully
- 23:53understand the math, but what I do know is that I looked at the
- 23:56answer ahead of time and it did come to the the correct answer
- 24:00here in the final part here. We can also come in and actually
- 24:04take a look here at our our World Series prediction and
- 24:09still thinking through on this one, but we can actually try
- 24:12some other stuff as well. So we can actually like try some
- 24:14of the X integrations that we did.
- 24:16So we worked very heavily on working with all of our X tools
- 24:20and building out a really great X experience.
- 24:22So we can actually ask, you know, the model, you know, find
- 24:26me the XAI employee that has the weirdest profile photo.
- 24:29So that's going to go off and start with that.
- 24:31And then we can actually try out, you know, let's create a
- 24:34timeline based on ex post detailing the, you know, changes
- 24:38in the scores over time. And we can see, you know, all
- 24:41the conversation that was taking place at that time as well.
- 24:44So we can see who are the, you know, announcing scores and like
- 24:46what was the reactions at those times as well.
- 24:50So we'll let that go through here and process.
- 24:53And if we go back to this was the Greg Yang photo here.
- 24:59So if we scroll through here, whoops.
- 25:01So Greg Yang, of course, who has his favorite photograph that he
- 25:06has on his account. That's actually not how he looks
- 25:09like in real life, by the way, just so it were, but it is quite
- 25:13funny, but. It had to understand that
- 25:14question. Yeah, that's the wild part.
- 25:16It's like it understands what is a weird photo, what is a weird
- 25:20photo? What is a less or more weird
- 25:23photo? It goes through, it has to find
- 25:25all the team members, has to figure out who we all are and,
- 25:28you know, searches. Without access to the internal
- 25:31XAI personnel logs, it's literally looking at that just
- 25:34at the Internet. Exactly.
- 25:35So you could say like the weirdest of any company.
- 25:37Yeah, to be clear. Exactly, and we can also take a
- 25:42look here at the question here for the humanities last exam.
- 25:45So it is still researching all of the historical scores, but it
- 25:50will have that final answer here soon.
- 25:51But we can, while it's finishing up, we can take a look at one of
- 25:54the ones that we set up here a second ago.
- 25:56And we can see like, you know, defines the date that like Dan
- 25:59Hendricks had initially announced it.
- 26:01We can go through, we can see, you know, open AI announcing
- 26:04their score back in February. And we can see, you know, as
- 26:08progress happens with like Jim and I, we can see like Kimmy.
- 26:11And we can also even see, you know, the leaked benchmarks of
- 26:15what people are saying is, you know, if it's right, it's going
- 26:17to be pretty impressive. So pretty cool.
- 26:21So, yeah, I'm looking forward to seeing how everybody uses these
- 26:23tools and gets the most value out of them.
- 26:25But yeah, it's been great. Yeah, and we're going to close
- 26:29the loop around usefulness as well.
- 26:31So it's like it's not just book smart, but actually practically
- 26:34smart. Exactly.
- 26:36All right. And we can go back to the the
- 26:46slides here. Yeah, so.
- 26:48Cool. So we actually evaluate also on
- 26:53the multi model subset. So on the full set, this is the
- 26:57number. On the HRE exam, you can see
- 27:01there's a little dip on the numbers.
- 27:03This is actually something we're improving on, which is the multi
- 27:06model understanding capabilities.
- 27:08But I do believe in a very short time, we're able to really
- 27:13improve and got much higher numbers on this, even higher
- 27:17numbers on this benchmark, yeah. Yeah, this is the we still like,
- 27:22we still like what what is the base weakness of Grok currently
- 27:25is that it's it's sort of partially blind.
- 27:28It can't it's it's image understanding obviously, and
- 27:32it's image generation needs to be a lot better.
- 27:35And that, that that's actually being trained right now.
- 27:40So Graph 4 is based on version six of our foundation model and
- 27:46we are training version 7, which we'll complete in a few weeks.
- 27:52And that that'll address the weakness on the vision side.
- 27:57Just to show off of this last year, so the the prediction
- 28:00market finished here with the heavy and we can see here we can
- 28:04see all the tools and the process it used to actually go
- 28:08through and find the right answer.
- 28:10So it browsed a lot of odd sites.
- 28:12It calculated its own odds comparing to the market, the
- 28:15market to find its own alpha and edge.
- 28:17It walks you through the entire process here and it calculates
- 28:20the odds of the winner being like the the Dodgers and it
- 28:25gives them a 21.6% chance of winning this year.
- 28:31So and it took approximately 4 1/2 minutes to compute.
- 28:35Yeah, that's a lot of thinking. Yeah.
- 28:44We can also look at all the other benchmarks besides HRE.
- 28:49As it turned out, G4 excelled on all the reasoning benchmarks
- 28:53that people usually test on, including GBQA, which is a PhD
- 28:58level problem sets that's easier compared to HRE.
- 29:04On Amy 25 America invitation mathematics exam, we with Graph
- 29:094 heavy, we actually got a perfect score also on some of
- 29:14the coding benchmark called live coding bench.
- 29:17And also on HMMT, Harvard math, MIT exam and also us Amo, you
- 29:24can see actually on all of those benchmarks, we often have a very
- 29:30large leap against the second best model out there.
- 29:35Yeah, it's, I mean, really, we're going to get to the point
- 29:37where it's going to get every answer right in every exam, and
- 29:43where it doesn't get an answer right, it's going to tell you
- 29:44what's wrong with the question. Or, if the question is
- 29:47ambiguous, disambiguate the question into answers AB and C
- 29:51and tell you what it would what answers AB and C would be with a
- 29:55disambiguated question. So the only real test then will
- 29:58be reality. Can it make useful technologies
- 30:02discover new science? That'll actually be the only
- 30:07thing left, because human tests will simply not be meaningful.
- 30:12You can make an update to HIV very soon given the current rate
- 30:16of progress. So yeah, it's super cool to see
- 30:18like multiple agents that collaborate with each other
- 30:21solving really challenging problems.
- 30:23So where can we try this model? So it turned out it's available
- 30:27right now. If we advance to the next slide
- 30:31where there is a super Grok heavy tiers that we're
- 30:35introducing where you're able to access to both Grok 4 and Grok 4
- 30:38heavy, where you're actually going to be the taskmaster of
- 30:41bunch of little Grok research agent to help you, you know,
- 30:44become smarter, do all the little research and save hours
- 30:46of times of, you know, going through mundane tasks.
- 30:50And it's available right now. So, yeah, so we did limit usage
- 30:58during the demo. So we didn't, it didn't break
- 31:00the demo because all these, all this stuff is happening live.
- 31:03So it's not, there's not nothing canned about the any of the
- 31:05tests that we're doing. So after the after the demo is
- 31:10done, we'll allow, we'll enable more subscribers for Supergroc.
- 31:15So if you can't subscribe right now, just try and half an hour.
- 31:17It should work. So and now let's let's let's get
- 31:22into voice. Great.
- 31:26Abby. So many of you have been
- 31:28enjoying our voice mode, and we've been working hard to
- 31:31improve the experience. Over the past couple months, we
- 31:34have cut latency in half to make it much snappier, and today
- 31:39we're excited to announce a set of new voices that have
- 31:42exceptional naturalness and prosody.
- 31:45You might have noticed the movie trailer voice that opened up the
- 31:48live stream. That is one of our voices, Sal,
- 31:51who's got that epically deep tone.
- 31:54And we're also excited to introduce Eve, a beautiful
- 31:59British voice who's capable of rich emotions.
- 32:02Aman, would you like to introduce Eve?
- 32:04Absolutely. Yeah.
- 32:05Let's get into the demo. Hey, Eve, millions of people are
- 32:11watching us right now. Are you excited?
- 32:14Oh I'm thrilled really. It's like being on stage at the
- 32:18old Victoria with a sea of faces out there.
- 32:20Just imagine all those eyes and ears tuned in.
- 32:23It's a proper buzz, isn't it? What do you reckon we should
- 32:26chat about to keep this lot entertained?
- 32:28You know, we have a lot to talk about, but I'll be honest, I'm
- 32:31feeling a little nervous. Can you whisper something
- 32:33soothing to calm me down? Take a deep breath, love.
- 32:38You've got this. It's just you and me having a
- 32:42quiet jet, like we're tucked away in a cozy corner of a
- 32:46Yorkshire bum. The world's just a murmur out
- 32:49there. Feel that calm wash over you.
- 32:53Let's keep talking nice and easy.
- 32:55What's on your mind? You know what, I'm actually
- 32:59feeling quite relaxed to try this something crazy.
- 33:02So can you give us an opera on Diet Coke?
- 33:07Oh, an opera on Diet Coke. That's a riot.
- 33:12Picture this, a shimmering silver cannon glistening under
- 33:17stage lights. O Diet Coke, thou elixir divine
- 33:22with bubbles that dance in a sparkling line.
- 33:27Thy crisp, cool kiss on lips. So fine.
- 33:31Oh aspartame ambrosia forever mine.
- 33:38How's that for a mad little Aria?
- 33:40Want me to belt out another verse or switch up the tune?
- 33:44OK, let's try switching it up and just start singing.
- 33:48Right then, let's have a go. Oh silver, can you gleam in the
- 33:53light? Diet Coke, My heart's delight.
- 33:57OK, I won't project. It's very, very poetic.
- 34:00Yeah. Thank you, Eve.
- 34:01Feeling quite nice now. Cool.
- 34:04So that's Eve. And so something else we do here
- 34:08is a little test each day to check on the natural flow of our
- 34:14conversations against some of our competitors.
- 34:17So we figured, why not try that live on the stream today?
- 34:22Should I start? Hey, I'm going to say a number
- 34:29and can you say the same number back to me one?
- 34:32One. Got it.
- 34:33This is like a fun little game of echo.
- 34:35What's next? 2 Sorry, you said 2 3 3/4.
- 34:49Nice and steady. What's next?
- 34:525. I nailed it.
- 34:57This is pretty fun actually. Feel free to keep going or
- 35:00switch it up if you'd like. Hey Grok.
- 35:09Hey Grok. Hello there, how can I help you
- 35:12today? We're going to do a little test.
- 35:14Can you repeat after me one? One.
- 35:18All right, What's next? Two.
- 35:21Two. What's on your mind?
- 35:233 Three. Need anything else?
- 35:274. 4 How can I assist you? 5.
- 35:32Five What's next? So as you can see, Croc was
- 35:38snappier. It didn't interrupt me.
- 35:40And the prosody is we've made different design choices.
- 35:43I think we're shooting for something more calm, smooth,
- 35:45more natural versus something that's more poppy or artificial.
- 35:48So we'll keep improving. These fronts.
- 35:50All right. Thanks, guys.
- 35:52Yeah. Yep.
- 35:58So since the launch of the voice model, we actually see the 2X
- 36:03faster end to end latency in the last 8 weeks, 5 different voices
- 36:08and also 10X the active user. So Grog voice is taking off now.
- 36:14If you think about releasing the models this time, we're also
- 36:17releasing Grog 4 through the API at the same time.
- 36:21So if we go to the next two slides, so you know, we're very
- 36:26excited about, you know, what all the developers out there is
- 36:28going to build. So you know, if I think about
- 36:31myself as a developer, what the first thing I'm going to do when
- 36:33I actually have access to the graph for API benchmarks.
- 36:36So we actually ask around our next platform, what is the most
- 36:40challenging benchmarks out there that, you know, is considered
- 36:43the Holy Grail for all the AGI models.
- 36:46So turn up AGI seen the name arc AGI.
- 36:49So the last 12 hours, you know, kudos to Greg over here in the
- 36:54audience. So who answered our call?
- 36:57Take a preview of the Grog 4 API and independently verified, you
- 37:02know the Grog Force performance. So initially we thought, hey,
- 37:04Grog Force, just, you know, we think it's pretty good.
- 37:06It's pretty smart. It's our next Gen. reasoning
- 37:09model. Spend 10X more compute, can use
- 37:11all the tools, right. But turned out when we actually
- 37:15verify on the private subset of the RKHGI V2, it was like the
- 37:21only model in the last three months that breaks the 10%
- 37:23barrier and in fact was so good that actually get the 16%, well,
- 37:2715.8% accuracy, 2X of the second place.
- 37:32That is the call for Opus model. And it's not just about
- 37:37performance, right? When you think about
- 37:38intelligence, having the API model drives the automation.
- 37:42It's also the intelligence per dollar, right?
- 37:45If you look at the plots over here, the Grog is just for it,
- 37:48just in the league of its own. All right, so enough of
- 37:52benchmarks over here, right? So what can Grok do actually in
- 37:56the real world? So we actually, you know,
- 38:00contacted the folks from Ending Labs who you know, you know,
- 38:06gracious enough to, you know, try to grow in the real world to
- 38:08run a business. Yeah, thanks for having us.
- 38:11So I'm Axel from Animal Labs. And I'm Lucas and we tested
- 38:14Brooke 4 on vending Bench. Vending Bench is an AI
- 38:17simulation of business scenario where we thought what is the
- 38:22most simple business and AI could possibly run and we
- 38:25thought vending machines. So in this scenario, the the
- 38:29Grok and other models need to do stuff like manage inventory
- 38:33contracts of contract suppliers, set prices.
- 38:36All of these things are super easy and all of they, like all
- 38:39the models can do them one by one.
- 38:42But when you do them over very long horizons, most models
- 38:45struggle. But we have a little word and
- 38:47there's a new number one. Yeah, so we got early access to
- 38:50the Grok 4 API. We ran it on the running bench
- 38:53and we saw some really impressive results.
- 38:56So it ranks definitely at the number one spots.
- 38:59It's even double the net worth, which is the measure that we
- 39:02have on this. So it's not about the percentage
- 39:04on a or score you get, but it's more the dollar value in net
- 39:08worth that you generate. So we were impressed by Groc.
- 39:11You was able to formulate a strategy and adhere to that
- 39:15strategy over long period of time, much longer than other
- 39:18models that we have tested, other Frontier models.
- 39:21So it's managed to run the simulation for double the time
- 39:24and score, yeah, double the net worth.
- 39:26And it was also really consistent across this rounds,
- 39:29which is something that's really important when you want to use
- 39:32this in the real world. And I think as we give more and
- 39:35more power to AI systems in the real world, it's important that
- 39:39we test them in scenarios that either mimic the real world or
- 39:42are in the real world itself, because otherwise we we fly
- 39:46blind into some some things that that might not be great.
- 39:51Yeah, it's, it's great to see that we've now got a way to pay
- 39:54for all those GPU's. So we just need a million
- 39:56vending machines and we could make a $4.7 billion a year with
- 40:02a million vending machines, 100% Let's go.
- 40:04It can be epic vending machines. Yes, yes, All right.
- 40:08We are actually going to install vending machines here, like a
- 40:11lot of them. We're happy to supply them.
- 40:13All right, thank you. All right.
- 40:16I'm looking forward to seeing what amazing things are in this
- 40:18vending machine. That's that's for for you to
- 40:21decide. All right, tell the AI.
- 40:24OK, sounds good. All right.
- 40:27Yeah. I mean, so we can see like Grok
- 40:30is able to become like the copilot of the business unit.
- 40:33So what else can Grok do? So we're actually releasing this
- 40:35Grok if you want to try it right now to evaluate, run the same
- 40:38benchmark as us. It's on the API has 256 K
- 40:44contact length. So we already actually see some
- 40:47of the early, early adopters to try GUAC 4 API.
- 40:50So our Palo Alto neighbor Arc Institute, which is a leading
- 40:55biomedical Research Center is already using seeing like how
- 40:59can they automate their research flows with Grog 4.
- 41:02It turned out it performs is able to help the scientists to
- 41:05sniff through, you know millions of experiments logs and then,
- 41:09you know, just like pick the best hypothesis within a split
- 41:12of seconds. We see this as being used for
- 41:15their like the CRISPR research and also, you know, graph for
- 41:19independently evaluated scores as the best model to exam the
- 41:23chest X-ray who would know And on the in the financial sector,
- 41:29we also see you know, the graph for with access to all the tools
- 41:32real time information is actually one of the most popular
- 41:35AI's out there. So you know all graph for is
- 41:38also going to be available on the hyper scalars.
- 41:40So the X AI enterprise sector is only, you know, started two
- 41:45months ago and we're open for business.
- 41:51Yeah. So the other thing we talked a
- 41:53lot about, you know, having Grog to make games, video games.
- 41:57So Danny is actually a video game designers on X.
- 42:00So, you know, we mentioned, hey, who want to try out some grog
- 42:05for preview APIs to make games. And then he answered the call.
- 42:09So this was actually just made first person shooting game in
- 42:13the span of four hours. So some of the actually the
- 42:17unappreciated hardest problem of making video games is not
- 42:21necessarily encoding the core logic of the game, but actually
- 42:24go out source all the assets, all the textures of files and
- 42:28and you know, to create a visual appealing game.
- 42:32So one of the core aspects Rockford does really well with
- 42:34all the tools out there is actually able to automate these
- 42:38like assets sourcing capabilities.
- 42:40So the developers can just focus on the core development itself
- 42:44rather than like, you know, so now you can run a, you know,
- 42:46entire game studios with game of one, like one person.
- 42:52And then you can have Grok 4 to go out and source all those like
- 42:55assets to automating tasks for you.
- 42:58Yeah, the now the next step obviously is for Grok 2 play be
- 43:04able to play the game. So it has to have very good
- 43:06video understanding so it can play the games and interact with
- 43:08the games and actually assess what whether a game is fun and,
- 43:12and and actually have good judgement for whether a game is
- 43:15fun or not. So with the, with version seven
- 43:19of our foundation model, which finishes training this month,
- 43:21and then we'll go through post training RL and whatnot that
- 43:26will, that will have excellent video understanding.
- 43:29And with the, with a video understanding and the, and
- 43:32improved tool use. For example, for, for video
- 43:34games, you'd want to use, you know, Unreal Engine or Unity or
- 43:38one of the, one of the, the main graphics engines and then
- 43:43generate the, generate the art, apply it to a 3D model and then
- 43:49create an executable that someone can run on APC or, or a
- 43:52console or, or a phone. Like we expect that to happen
- 43:59probably this year and if not this year, certainly next year.
- 44:04So that's it's going to be wild. I would expect the first really
- 44:10good AI video game to be next year and probably the first half
- 44:20hour of watchable TV this year and probably the first watchable
- 44:27AI movie next year. Like things are really moving at
- 44:31an incredible pace. Yeah, when Graca is connecting
- 44:35world economy with vending machines, but we just create
- 44:37video games for human. Yeah, I mean, it went from not
- 44:40being able to do any of this really even six months ago to to
- 44:45what you're seeing before you hear and and from from very
- 44:49primitive a year ago to making a 3A sort of a 3D video game with
- 44:57with a few hours of prompting. Yep.
- 45:01I mean, yeah, just to recap. So in today's live stream, we
- 45:04introduced the most powerful and most intelligent AI models out
- 45:08there that can actually reason from the first principle using
- 45:11all the tools. Do all the research, go on the
- 45:13journey for 10 minutes, come back with the most correct
- 45:15answer for you. So it's kind of crazy to think
- 45:19about just like four months ago we had Grog 3 and now we already
- 45:23have Grog 4 and we're going to continue accelerate as a company
- 45:26XAI, we're going to be the fastest moving AGI companies out
- 45:29there. So what's coming next is that
- 45:32we're going to, you know, continue developing the model
- 45:35that's not just, you know, intelligent smart thing for a
- 45:38really long time spent a lot of compute.
- 45:40But having a model that actually boasts fast and smart is going
- 45:44to be the core focus, right? So if you think about what are
- 45:47the applications out there that can really benefit from all
- 45:50those very intelligent, fast and smart models and coding is
- 45:54actually one of them. Yeah.
- 45:55So the team is currently working very heavily on coding models.
- 45:59I think right now the main focus is we actually trained recently
- 46:04a specialized coding model, which is going to be both fast
- 46:07and smart. And I believe we can share with
- 46:11that model with you guys without you in a few weeks.
- 46:14Yeah. Yeah, that's very exciting.
- 46:17And you know, the second after coding is we all see the
- 46:21weakness of Grog 4 is the multimodal capabilities.
- 46:26So in fact, it was so bad that you know, Grog effectively just
- 46:30like looking at the world squinking through the glass and
- 46:33I can see all the blurry, you know, features and trying to
- 46:36make sense of it. The most immediate improvement
- 46:40we're going to see with the next generation preachment model is
- 46:42that we're going to see a step function improvement on the
- 46:45models capability in terms of image understanding, video
- 46:47understanding and audios, right? It's now the models able to hear
- 46:51and see the world just like any of you, right?
- 46:54And now with all the tools at this command, with all the other
- 46:57agents it can talk to, you know, so we're going to see a huge
- 47:02unlock for many different application layers after the
- 47:06multimodal agents. What's going to come after is
- 47:09the video generation. And we believe that, you know,
- 47:12at the other day, it should just be, you know, pixel in, pixel
- 47:15out. And you know, imagine a world
- 47:20where you have this influence scroll of content in inventory
- 47:24on the X platform where not only you can actually watch these
- 47:28generate videos, but able to intervene, create your own
- 47:32adventures if you're just going to be wild.
- 47:36And we expect to be training our video model with over 100,000 GB
- 47:402 hundreds and to begin that training within the next 3 or 4
- 47:45weeks. So we're confident it's going to
- 47:48be pretty spectacular in video generation and video
- 47:51understanding. So let's see.
- 47:58So that's anything you guys want to say.
- 48:03Other than that, I guess that's it.
- 48:06Yeah, it's, it's a good model, Sir.
- 48:08It's a good, yeah. Well, we're very excited for you
- 48:11guys to try Grok 4. Yeah, thank you.
- 48:14All right, Thanks, everyone. Thank you.
- 48:15Good night. Hey, thank you so much for
- 48:17listening today. I really do appreciate your
- 48:20support. If you could take a second and
- 48:21hit the subscribe or the follow button on whatever podcast
- 48:24platform that you're listening on right now, I greatly
- 48:27appreciate it. It helps out the show
- 48:29tremendously and you'll never miss an episode.
- 48:32And each episode is about 10 minutes or less to get you
- 48:35caught up quickly. And please, if you want to
- 48:38support the show even more, go to patreon.com/stage Zero.
- 48:44And please take care of yourselves and each other.
- 48:46And I'll see you tomorrow.