Latest / Elon Musk Podcast / AI Update - New Free Video Generator, Bytedance, Google and More! - Sponsor : Fish.Audio
Transcript
- 0:01Hey everyone, welcome back to the Elon Musk Podcast.
- 0:05I'm thrilled to share some exciting news with you over the
- 0:08next two weeks. We're evolving.
- 0:10We'll be broadening our focus to cover all the tech Titans
- 0:14shaping our world. And with that, our show will
- 0:16become Stage 0. You'll still get the latest
- 0:20insights on Elon Musk plus so much more.
- 0:23So stay tuned for our official relaunch at Stage 0 coming soon.
- 0:28I'd like to thank the sponsor of this show, Phish dot Audio, a
- 0:33great service. So if you're a podcaster, a
- 0:36voice over artist, or a content creator like myself and you're
- 0:39looking for hyper realistic AI generated voices, you need to
- 0:44check out Phish Audio. Ranked among the top 2 most
- 0:48realistic voice generators in 13 languages, Phish Audio lets you
- 0:52create high quality voice overs with just a few clicks.
- 0:57Now I'm going to be honest with you, sometimes I make mistakes
- 0:59during the show and I need to replace some words because I
- 1:02misspoke or something was in the background, there was some noise
- 1:06or something. I just can't get rid of it.
- 1:07So sometimes I use AI generated me in order to fill in those
- 1:11gaps. And I tried Phish Audio for this
- 1:14and it worked so good, so much better than the others that I've
- 1:19tried. And it's open source friendly.
- 1:21It's super fast, it's even faster than 11 labs if you're
- 1:24familiar with that. And best of all, it's
- 1:26affordable. It has the lowest prices on the
- 1:29market. So whether you need multilingual
- 1:32voice cloning, real time speech to speech conversion, or
- 1:36professional quality narration, Fish Audio has you completely
- 1:42covered with whatever you need done.
- 1:43Try it today for free at Fish dot Audio and see why it's the
- 1:49absolute best AI vice tool out there today.
- 1:53Go check them out at Fish dot Audio.
- 1:56That's Fish dot Audio fish dot audio.
- 2:00Head to Fish dot Audio and start creating with AI powered voices
- 2:05right now. Welcome to this week's AI
- 2:08update. So how many times can artificial
- 2:11intelligence reinvent itself in a single week?
- 2:14And what happens when it starts to reshape the very definition
- 2:18of creativity, design, and human likeness?
- 2:21This week in AI has introduced more than just flashy demos or
- 2:25iterative upgrades. It has opened a new chapter in
- 2:28how we visualize the future of interaction, media and autonomy.
- 2:33From image generation that understands multiple references
- 2:36to hyper realistic avatars that lip sync in multiple languages,
- 2:40to humanoid robots executing front flips and punches, all
- 2:44within a few days, the question is no longer what AI can do, but
- 2:50what it will leave for us to do. Bytedance, Google, Meta,
- 2:54Alibaba, Vivago AI and a constellation of open source
- 2:58contributors have each launched new products, updates or
- 3:01jaw-dropping showcases, many of which feel like they came from
- 3:05different futures. This isn't about incremental
- 3:08progress, it's about a cascade of systems each pushing
- 3:11boundaries once thought unscalable.
- 3:14But what's buried beneath the spectacle is the unsettling
- 3:17realization. How fast are we letting machines
- 3:21to what once required human hands, eyes, and voices, or
- 3:24evened our own emotions? So let's start with Bytedance's
- 3:29latest creation, Uno. This new AI model redefines what
- 3:33it means to generate an image. Uno responds to text prompts,
- 3:39but it also fuses multiple image references into a single
- 3:42coherent output, something that previous models often struggle
- 3:46with. Most impressively, it retains
- 3:49character fidelity and object accuracy, regardless of how
- 3:53surreal or stylistically diverse the prompt might be.
- 3:57Now, say you upload a logo and a photo of AT shirt.
- 4:00Uno doesn't just overlay it, it visualizes the logo printed on
- 4:04the shirt with convincing fabric, shadows and appropriate
- 4:08scale for the logo. Or consider uploading a doll and
- 4:11a plush toy. Uno will place them together in
- 4:15a new realistic scene. It even let's users generate A
- 4:19stylized portrait of a person using a single input photo.
- 4:23Anime, Pixar, or even video game style.
- 4:27Uno handles all of that stuff while maintaining identity
- 4:30trades. Now for creators, Uno is a
- 4:33silent assassin. If you're running a fashion
- 4:36brand, Uno is basically your model and your whole photo shoot
- 4:39team. Upload 2 clothing items and
- 4:41prompt the model to generate a person wearing them.
- 4:45Set in a cityscape or flower field of sunset.
- 4:48You can even try different models, body types, or urban
- 4:50settings. And for influencers, Uno becomes
- 4:53a face swapping studio. Give it one video of yourself
- 4:56and it'll generate your likeness in any scene, doing anything,
- 5:00wearing anything. This kind of tool doesn't just
- 5:03help with photorealistic image creation, it enables entirely
- 5:07new workflows for people who don't have a studio full of
- 5:11people. Now, what used to require hours
- 5:14and possibly days in Photoshop and multiple production rounds
- 5:17is now available with just a few clicks and a smart prompt.
- 5:23Now, Uno's value isn't just in flexibility, it's in fidelity.
- 5:27In benchmark comparisons with tools like Omni IP Adapter and
- 5:31OEM control, Uno consistently produce more accurate
- 5:35generations. For instance, when fed a
- 5:37reference image of a uniquely designed clock, Uno was the only
- 5:41tool able to correctly recreate it on green grass surrounded by
- 5:45sunflowers with even the yellow 3 intact.
- 5:49The same held true for altering the color of toys or preserving
- 5:53intricate shapes. Accuracy at the level means less
- 5:56trial and error and more trust in AI generated visuals.
- 6:00That's a quiet revolution for designers, developers and
- 6:04content creators who work under very tight deadlines.
- 6:08Now Byte Dance has made Uno freely available through a
- 6:11Hugging Face demo with an optional GitHub repository for
- 6:15local installs. You'll need a machine with at
- 6:17least 16 gigs of RAM and some familiarity with launching
- 6:21Python scripts. Quantized model variants are
- 6:24also available for those with limited GPU resources and this
- 6:28dual offering browser based ease and local control makes Uno
- 6:31accessible without compromising its power.
- 6:34Gives professionals and hobbyists alike the ability to
- 6:37integrate next Gen. visual tooling into their own
- 6:39workflows. Meanwhile, a whole different
- 6:43team has pushed video generation to the next level.
- 6:46Their latest tool, called One Minute Video Generation with
- 6:49Test Time Training, allows AI to produce full length, minute long
- 6:53videos that maintain consistent characters, environments, and
- 6:58art styles. Instead of relying on a single
- 7:01long prompt, you just breakdown a video into storyboard style
- 7:04scenes. Each scene contains a short
- 7:06description, like Jerry the brown mouse holds cheese or Tom
- 7:10the cat snatches the cheese and the model stitches these
- 7:13together to form a coherent, evolving story.
- 7:17Yes, the generated videos have rough edges, the lip sync is
- 7:20glitchy, backgrounds jitter, and text in scenes is often
- 7:24illegible. But the consistency of
- 7:26character, motion and design throughout a full minute is a
- 7:29major leap forward. Prior tools could barely sustain
- 7:33coherence over 3 to 5 seconds. This was made possible by
- 7:37layering test time training modules over Cog Video's base
- 7:40model. These modules act like memory
- 7:43units, absorbing and maintaining the visual state across scene
- 7:46transitions. It points to a near future where
- 7:50AI can generate a 10 minute short film with nothing but a
- 7:54script and a few references. Now then there's Hollow Point,
- 7:59another new AI system focused on 3D modeling.
- 8:02The tool doesn't just segment object, it reconstructs what's
- 8:06hidden from view. If you upload a ring, the model
- 8:09identifies the band in the settings, then completes the
- 8:12invisible side of each. That completion isn't guesswork.
- 8:16It uses a two stage system. First it segments visible parts.
- 8:20Next, it sends each piece through a 3D diffusion model to
- 8:25fill in occluded areas. The end result is a fully
- 8:29editable mesh with complete logical subparts.
- 8:33For 3D artists, this means they can now expand or customize
- 8:36individual parts of a model, resize the diamond, change a rim
- 8:40texture, or separate a chair's legs from its seat, all without
- 8:45needing to manually remodel or make speculative guesses.
- 8:49This kind of automation saves hours of work in production
- 8:53design, animation, and game development.
- 8:55Like many other tools this week, Hollow Part is available on
- 8:58Hugging Face for interactive testing and through GitHub for
- 9:01full installation. Open source models just got
- 9:04better than commercial ones. One major development this week
- 9:08came from a lesser known but quickly rising company called
- 9:10Vivigo AI. It launched Hydram, a new open
- 9:14source image generation model that has overtaken the
- 9:17competition, and that's at least in benchmark ratings.
- 9:21Unlike other high performing systems that remain closed or
- 9:24partially commercialized, High Dream is fully uncensored and
- 9:28openly available. According to Artificial Analysis
- 9:32Independent Model Ranking Organization, High Dream now
- 9:35holds the 3rd place position among all text to image models.
- 9:39It is the highest ranked open source model on the list,
- 9:42surpassing not only stable to Fusion, but also Flux 1D.
- 9:47Flux Pro, which ranks slightly higher, remains closed source,
- 9:51making Hydream the best available option for developers
- 9:54and researchers working without licensing restrictions.
- 9:58Early tests show that Hydream is more than a benchmark leader,
- 10:02and it performs well under practical creative conditions.
- 10:06Text prompts that previously produced vague or distorted
- 10:09visuals now return precise, detailed, and highly stylized
- 10:12results. This has opened up new use cases
- 10:15for illustrators, concept designers, and product marketers
- 10:18who want control over the visual fidelity of AI generated art.
- 10:22Importantly, I Dream allows full creative expression without the
- 10:26usual guardrails or sensors. This is a rare stance in the AI
- 10:30space, where most companies apply aggressive filtering or
- 10:34content restrictions. That raises ethical questions
- 10:37for many users, particularly in animation satire and adult
- 10:41design solves a problem they've long encountered with other
- 10:44models. Now Omni SVG is also rising it's
- 10:50text to vector generation. Another is stand out Omni SVG,
- 10:54an AI tool designed to generate scalable vector graphics, or
- 10:58SVGS, from text prompts or images.
- 11:00Unlike pixel based images, SVGS are resolution independent means
- 11:05that designers can scale these visuals from handheld to
- 11:10billboard sizes without losing fidelity, and developers can
- 11:14drop them into apps without worrying about file size or
- 11:17weight on the SVG stands out for its ability to produce complex,
- 11:21detailed vector illustrations that actually match user
- 11:24prompts. Previous models SVG generation
- 11:27often failed to capture layout or character details, but this
- 11:31one completely nails it. Examples include cartoon
- 11:34characters with mushroom hats, stylized buttons with web UI
- 11:37concepts, and even photo based SVG conversations.
- 11:41The retain line fidelity, so vectors still matter.
- 11:45This matters because the web still runs on vectors.
- 11:48Icons, logos, app UI elements, and infographics all rely on
- 11:53SVGS. And for developers who need
- 11:56sharp design across multiple screen sizes, or for companies
- 12:00localizing assets into dozens of languages, tools like Omni SVG
- 12:04reduce the time between idea and development.
- 12:07And what's more, the model also accept raster images and
- 12:11converts them into precise SVG renderings.
- 12:14This brings it into direct competition with some commercial
- 12:17vectorization tools and according to testing it
- 12:20outperforms most of them in detail.
- 12:22Preservation and line geometry code and data sets for Omni SVG
- 12:27are expected to be released soon.
- 12:30Meanwhile, Alibaba has debuted a new version of the talking head
- 12:34generator dubbed Omni Talker. The tool takes a video of a
- 12:38person talking just a few seconds and allows users to
- 12:41change what that person says using a custom transcript.
- 12:46And the results is a realistic looking video where the person
- 12:49says anything you want them to say.
- 12:51The output is not only visually smooth, but supports different
- 12:55languages, maintains mouth movements In Sync with the
- 12:58audio, and preserves facial expressions.
- 13:02This means you can take an input of Jackie Chan speaking in
- 13:05Mandarin and generate a video of him delivering a speech in
- 13:08English, accent intact, expressions unbroken.
- 13:13Of course, it's not perfect. The output can become uncanny if
- 13:16the person moves their head rapidly or if there are long
- 13:19monologues without breaks. Some cases, the AI renders the
- 13:22face smoothly or loses subtle emotional transitions if the lip
- 13:26sync and frame continuity far surpass most tools currently
- 13:31available. And the more fascinating part is
- 13:34emotional modulation. If the reference video shows a
- 13:38sad expression, the generated face will appear somber even if
- 13:42the transcript is upbeat. A happy reference video injects
- 13:46cheerful expressions into whatever script is applied.
- 13:49This emotional anchoring makes the system feel less robotic and
- 13:53more like a tool for actors or presenters who want to alter
- 13:56delivery without rerecording. Now most avatar generators top
- 14:02out at about 10 to 15 seconds. Omni Talker handles multi minute
- 14:06videos, maintaining voice sync and facial movements throughout.
- 14:11In one test, a 2 minute political speech was generated
- 14:14using just a headshot and a transcript, no voice acting
- 14:18needed. You can even interact with the
- 14:20avatars live in real time, giving them queries and getting
- 14:24real time spoken answers. This makes the system useful for
- 14:27educators, sales, customer service or political messaging.
- 14:32Any scenario where one person must deliver dynamic scripts at
- 14:36scale without reshooting or editing the video, removes AI
- 14:40from behind the scenes to center stage.
- 14:42And another related tool launched this week, though not
- 14:46from Alibaba, can take a static photo and animate the face using
- 14:50either an audio clip or a full reference video of facial
- 14:54movement feeded a Steve Jobs speech and a photo of Einstein.
- 14:58It'll make Einstein deliver Jobs words, lip syncing each syllable
- 15:01with Uri accuracy. And what's remarkable here isn't
- 15:04just the lip sync, it's the ability to mimic non verbal
- 15:07cues, blinking, nodding, pausing or even glancing away.
- 15:12The AI tracks motion and emotional shifts, mapping them
- 15:15to any face, real or fictional. Whether it's Anne Hathaway or a
- 15:19digital clone, the face mimics both speech and sentiment, and
- 15:24compared to rivals like X portrait or live portrait, this
- 15:27tool performs better on both fidelity and emotional nuance.
- 15:32The head doesn't stay locked, the eyes shift naturally, and
- 15:35when combined with expressive audio, it creates A convincing
- 15:38illustration of live speech. The GitHub repository is listed,
- 15:44but the code hasn't been released yet.
- 15:46It's dual mode functionality, supporting both speech only
- 15:49input and full motion transfer, means that users can choose
- 15:53between simple lip sync generation or full body
- 15:56animation. The results are strong enough to
- 15:59enter entertainment production, especially for animated
- 16:02interviewers, explainer videos or social content.
- 16:07Now, robots. Aside of software, the most
- 16:11viral moments this week came from humanoid robots.
- 16:14Unitri's robot has performed flips and complex dances before,
- 16:17but Engine AI's robot raises eyebrows with boxing humans
- 16:22reacting to punches and landing counter blows.
- 16:25The demo, though, was captured during a live stream with a well
- 16:28known streamer visiting the company's facility in China.
- 16:31Critics have long claimed these humanoid demos were faked CGI
- 16:35stitched into real environments, but the live stream confirmed
- 16:39otherwise. The robot fell, got back up,
- 16:42aimed some punches and adjusted its stance all autonomously with
- 16:46AI. Not pre recorded, not
- 16:48choreographed. Unlike dancing routines or
- 16:50rehearsed acrobatics, boxing requires real time adaptation.
- 16:54The robot had to interpret opponent movements, adjust its
- 16:58balance, and select appropriate actions, all within
- 17:00milliseconds. Like human.
- 17:03It's about machines competing with humans.
- 17:06Physically a human versus a robot.
- 17:09We've seen this movie before, and the humans don't fare too
- 17:12well in it. Now what sets this apart is the
- 17:15autonomy system wasn't remote controlled.
- 17:17It executed decisions based on internal algorithms tracking
- 17:21visual input and calculating body position.
- 17:24If it fell, it knew how to stand up.
- 17:26If it gets too close, it backed off.
- 17:29For the first time, the idea of reacting humanoid robotics feels
- 17:33actually real. Then there's Kawasaki.
- 17:37They introduced a concept that sounds absurd until you watch
- 17:41it, an AI powered mechanical horse.
- 17:46Unlike traditional scooters or two Wheelers, this machine walks
- 17:49on 4 robotic legs. Its body mimics the movement of
- 17:53an actual horse and it's designed to be powered by
- 17:56hydrogen, releasing only water vapor.
- 17:59At this stage the prototype is non functional.
- 18:02The riding demo is CGI, and the model shown at Osaka's Expo is a
- 18:07static mock up. But the idea is absolutely
- 18:09serious. Kawasaki envisions a transport
- 18:12method that's clean, adaptive, and capable of navigating uneven
- 18:16terrain, but early feedback has been skeptical.
- 18:19Most people agree that wheels still outperform legs when it
- 18:22comes to energy efficiency and also stability.
- 18:26Now, all of these tools and models and demos are very
- 18:31flashy. They're tangible shifts in how
- 18:34humans can work, communicate, and create.
- 18:36If you're a designer, AI will now generate not only your
- 18:39visuals, but your style. If you're a performer, AI will
- 18:42now speak with your face and your voice.
- 18:44If you're a robotics engineer, the next leap may come not from
- 18:48code, but from carbon fiber and real time planning algorithms.
- 18:53The acceleration isn't evenly distributed, but the
- 18:57opportunities and risks are everywhere.
- 18:59These tools aren't just for labs or corporations anymore.
- 19:02They're for us. They're available on GitHub.
- 19:05Hugging Face or even browser based apps anyone can build with
- 19:10them. Hey, thank you so much for
- 19:14listening today. I really do appreciate your
- 19:16support. If you could take a second and
- 19:18hit this subscribe or the follow button on whatever podcast
- 19:21platform that you're listening on right now, I greatly
- 19:24appreciate it. It helps out the show
- 19:26tremendously and you'll never miss an episode.
- 19:28And each episode is about 10 minutes or less to get you
- 19:32caught up quickly. And please, if you want to
- 19:35support the show even more, go to Atreoncom Stage Zero.
- 19:41And please take care of yourselves and each other, and
- 19:43I'll see you tomorrow.