AI Search SxiRANj0xLs

RIP Claude Fable, open-source AI unleashed, full body avatars, new Google models, new TTS: AI NEWS

Transcript: Done Yayin: 2026-06-13 20:29 YouTube
RIP Claude Fable, open-source AI unleashed, full body avatars, new Google models, new TTS: AI NEWS
Kanala don
Job gecmisi
Job Durum Deneme Worker Istek Baslama Bitis
FetchYouTubeTranscript #72 Done 1 learning-prod-worker-1 2026-06-18 23:34:34 2026-06-18 23:34:58 2026-06-18 23:35:12

Ozet

Bu video icin henuz ozet yok.

Transcript

Video metni
en markdown 2026-06-18 23:35 youtube-transcript-api:generated
Indir
AI never sleeps and this week has been
absolutely insane. Anthropic releases
their best model, Claude Fable 5, but
then it also gets killed in less than a
week. This is now the best open- source
AI for animating characters with a
reference video. Google drops an
incredibly fast open- source model, plus
a free realtime translator. Kimmy just
released their latest model, and it's an
absolute beast. And also my favorite
open source lab ZAI quietly drops their
latest model GLM 5.2. Miniax also
released their best open source model
and an incredibly efficient
architecture. This AI can create 4D full
body avatars. We have a ton of really
cool 3D model generators, plus another
top Texas speech generator and a lot
more. So, let's jump right in. First up,
we have probably the best open- source
AI for transferring motion from one
video onto another video. It's called
Scale 2. And first of all, here are some
examples. As you can see, this can even
transfer the motion of more than one
character. So on the top is the original
video. We can change the input frame
instead to these two characters fighting
in a forest. And here's our result. Note
that it matches the movement of the
original video very well. Even though
this is quite a high action and tricky
scene, this can also do non-human
animals. So, for example, at the top is
the original video. We can change the
woman into a flamingo, and it's still
able to preserve roughly the same
movement. This even works if the input
or output character has some weird
aspect ratios or dimensions. For
example, if your output is a chubby
creature like this with abnormal
proportions, it's still able to detect,
you know, the skeleton of this character
and animate it appropriately. Here's
another example where we can easily
replace one character with another
character in a video. Not only does this
work with realistic videos, but also
anime and other artistic styles. And if
you compare this with other open-source
motion animators like Wan Animate or the
first version of Scale, note that this
new Scale 2 is just a lot better. I'd
say the quality is even on par with the
closed source Cling 3. Here's another
example for your reference. Note that
Scale 2 is even able to capture the
camera movement of the original driving
video, whereas Cling was not able to do
so. And then here's a comparison with
multiple characters. And as you can see,
scale 2 is by far the most performant.
It's able to consistently maintain the
details of both characters. The awesome
thing is this is out already. So at the
top of the page, if you click on this
code button and you scroll down a bit
here, it contains all the instructions
on how to download and run this locally
on your computer. Note that if I click
on the hugging face repo, the total size
of all the models is 81 GB. So, we'll
probably need to wait for some quantized
versions or GGUFS to be able to run it
on most consumer devices. By the way,
interesting thing is if you look at the
affiliations here, this is also by ZAI,
which is the company behind GLM, one of
the top open source models out there.
Interesting how they're also
contributing to AI videos as well. If
you're interested in reading further,
I'll link to this main page in the
description below. Also this week we
have a pretty interesting AI system
called actionable world representation.
This is basically a way to build a
moving digital twin of real world
objects. And the reason it's interesting
is that it doesn't just predict what an
object looks like. It tries to model how
that object can actually change, bend,
or move. So this takes real world 3D
data like point clouds or even just a
depth video and it outputs a
controllable 3D model of that object as
it moves. So instead of just generating
a physical 3D object, it also generates
the object's motion. And what makes this
impressive is the range of objects it
can handle. The demos show articulated
objects like a hand or an entire human
moving to things like deformable
earphones or even the Unitere dog robot.
And this matters because if we want
robots or AI agents to act in the real
world, we need more than hard objects in
simulations. We also need accurate
representations of how these objects
move. The awesome thing is they've
released this already. So if you click
on this code button and you scroll down
a bit here, it contains all the
instructions on how to run this on your
computer. If you're interested in
reading further, I'll link to this main
page in the description below. Also this
week we have a new AI called Oscar and
this is a world model for robots. In
simple terms, it tries to predict or
simulate what would happen when you tell
the robot to take a certain action. For
example, clearing the dining table or
inserting a capsule in a coffee machine
or lifting a pot by the handle,
inserting a plug, etc. The impressive
part is this is designed to work across
different robot bodies, not just one
specific arm or one specific setup.
Instead of relying on a robot's exact
appearance, Oscar uses a 2D skeleton
style motion as the control signal, as
you can see here. This makes it easier
to transfer across different robot arms
or different models. And the model can
focus on motion structure instead of
memorizing what each particular robot
looks like. Here you can see that it's
able to generate videos doing a variety
of tasks like manipulating different
objects to making lunch. And you know,
this AI is really valuable because right
now we don't have enough real world
video data to train humanoid robots.
Well, so what we could do is just train
them in a virtual environment or use AI
to create a ton of synthetic training
videos like this to train robots. If you
compare Oscar with other robot video
simulators, you can see that Oscar is a
lot more similar to the ground truth
compared to the other models. Here's
another example for your reference. Now,
at the top of the page, they have
released the code to this. So, if you
scroll down here, it contains all the
instructions on how to run this on your
computer. Note that here they do
recommend an Nvidia GPU with over 24 GB
of VRAM. Plus, the awesome thing is this
is under the Apache 2 license, which has
very minimal restrictions. If you're
interested in reading further, I'll link
to this main page in the description
below. Also, this week, Google releases
Gemini 3.5 Live Translate. This is their
new real-time translation model.
Basically, you can talk in one language
and it would speak out a translated
version using your same voice almost in
real time. Here are some examples.
[music] Good morning.
>> I'm not sure where I am. Let me look
around first. I think I'm at the main
gate, but I don't see your car. I am
further down near the
I understand. I'm at the other pickup
point. Let me walk a bit further to find
your car.
>> It's great to see you both. Cassie,
how's the weather in Shanghai?
>> It's nice to meet you. The weather here
is really nice. I spent the whole day
with my family in the park.
>> That sounds wonderful. Anna, you
mentioned last week that you were
celebrating your birthday. How was
[clears throat] it? Oh, they were
fantastic.
>> That was fantastic. I had dinner at my
favorite restaurant with some friends.
If you visit Sweden, you must try this
restaurant.
>> Okay, Cassie, can you update us on the
website roll out in Asia?
Okay, that's great. How is the Europe?
Note that here it can automatically
detect over 70 languages. And it also
tries to preserve the speaker's
intonation, pacing, and pitch. And it
also tries to keep the conversation
natural. Unlike turnbyturn translation
systems that wait until someone finishes
talking, Gemini 3.5 L translate
generates continuously while staying
only a few seconds behind the speaker.
So that means fewer awkward pauses and
smoother back and forth. The awesome
thing is this feature is already
available via their API as well as
Google's AI studio which you can try for
free. So simply click on this link and
you can try it here. And apparently this
is available for everyone via Google
Translate on Android and iOS.
>> Welcome to our tour of the beautiful
city of Cartahana. We will walk through
its cobblestone streets to admire its
impressive architecture. And we will end
the day enjoying a beautiful sunset.
>> So now we can all be lazy and not need
to even learn a new language. We can
just pull up live translate when we are
traveling in a foreign country and it
will help us translate whatever we say
on the fly. Super useful tool. If you're
interested in reading further, I'll link
to this main page in the description
below. Also this week, Google releases
diffusion Gemma. This is their
open-source model which is incredibly
fast. Now, this is a bit different from
regular language models which you may be
used to. You see, most language models
write text from left to right, one token
or word at a time. But this diffusion
Gemma model generates text more like an
image generator by using a diffusion
model. Instead of generating text one
word at a time from left to right, this
actually drafts blocks of text in
parallel and then refineses them over
multiple passes. So rather than just
predicting the next word and then the
next word, this works on a whole block
of text at once. And here Google says
that this can reach up to four times
faster text generation compared to
regular autoprogressive language models.
And you know the really cool thing is
you can even get this to do pretty
complex stuff like solving sudoku or
generating 3D models as you can see from
this example. Now historically these
diffusion language models have performed
pretty badly in terms of generating text
reasoning and world understanding.
They're not as good as normal language
models. But if you look at these
benchmarks, then this new diffusion
Gemma in light blue actually performs
pretty close to Gemma 4, which is a
regular auto reggressive language model
of the same size across these knowledge
benchmarks like MMLU or GPQA, which is
graduate level scientific knowledge as
well as competitive math and coding. The
nice thing is this model is already
released and this is under a very
permissive Apache 2 license which has
very minimal restrictions. You can even
use this for commercial purposes. Now,
on this hugging face repo, note that the
total size of the model and some
additional files is around 52 GB. So,
you'll still need a high-end GPU to run
this. And that's because this is 26
billion parameters. If you're interested
in reading further, I'll link to this
main page in the description below. Next
up, this AI is quite interesting. It's
called streamforce and this is a video
generator where you can control the
motion by applying forces directly in
the video. So here are some examples of
this in action. For example, you can
apply force this way and it'll basically
generate a video of the clothes being
blown in that direction. Or here are a
few other examples. Instead of just
using a prompt to dictate how the
objects move, you can now just push
objects in a video using a force signal.
And you can also change that force over
time. Think of it like giving an AI
video model a physics joystick. Now,
this supports local and global forces.
So here, if we use a global force, it's
like applying wind that affects the
whole scene. Or if we use a local force,
then it only applies that force on one
part of the scene. For example, like
moving just one chair or just moving
this apple. Now the important part of
this is it's also causal and streaming
which means it can react in real time as
you change the four signals. The page
says that streamforce runs up to 16.6
frames pers on just a single CPU. Now at
the top of the page it does say the code
is coming soon. So hopefully they will
stick their word and release this. For
now if you're interested in reading
further I'll link to this main page in
the description below. Also this week we
have a new very important benchmark
called agents last exam. You see this is
basically a stress test for AI models
doing real professional work not just
answering random knowledge questions or
solving tiny coding tasks. The idea is
simple. If AI agents are supposed to
help with actual jobs, then we need
benchmarks that look more like actual
jobs. So instead of asking an agent just
one isolated question, this benchmark
gives it much longer and practical
workflows with clear outcomes. So here
you can see it covers tasks across 55
subindustries including animation,
neuroscience, 3D modeling, architecture,
manufacturing, game development,
engineering, and more. So, for example,
agents might have to complete animation
and VFX work in After Effects or create
3D models or set up scenes in Unreal
Engine or run some analyses in medical
software or do architectural modeling
using another app. Basically, it's
trying to measure whether agents can
survive real workflows where the task is
multi-step. It's specific to a certain
domain and it's actually used in real
work. And if you look at the
leaderboard, then you can see that GPT
5.5 Codeex is actually the best. It even
beats Claude Fable 5. Here there's also
a footnote. Most of the time Claude
Fable 5 will just outright gate you. And
even if it does decide to answer you, it
might silently and secretly give you a
dumber answer by default. Even rerunning
it can't correct that. And here, as you
can see, GPT 5.5 using codeex, which is
what I use personally for most of my
daily tasks, does perform the best in
terms of these realworld agentic tasks.
followed by GPT 5.5 and OpenClaw and
then surprisingly the new composer model
by cursor also does very well. So a
pretty interesting benchmark for you to
put on your radar. At the top they have
released a GitHub button which contains
instructions on how you can run this
yourself. If you're interested in
reading further I'll link to this main
page in the description below. Also this
week we have a new system called Arbor
which is really useful for autonomous
research. The easiest way to think about
it is this. as an AI agent doing
research. This makes it so that it
doesn't just try one idea and then fail
and then forget what happened. Instead,
it builds a living research tree. It
takes the research objective and breaks
it down into hypotheses. It runs
experiments. It saves the evidence and
then it uses those results to decide
what to do next. Basically, it turns
research from a messy sequence of
disconnected attempts to a more
cumulative process where you can refine
and iterate further. The key idea is
called hypothesis tree refinement. So we
have this overarching coordinator which
manages the overall strategy and then we
also have these executors that test
individual ideas in isolated instances.
These are like the temporary research
assistants that implement just one
experiment. They run it and then they
report back and then each node in the
tree basically keeps track of each
hypothesis and experiment. And this
matters because AI agents often struggle
with continuity. They can edit code and
run short tests and do one specific
task. But they often lose sight of the
big overall picture. But here with
Arbor, this is a system where the agents
can keep track of what it has learned
and compare competing directions and
iterate further. Now the really cool
thing is if you look at these benchmarks
comparing arbor with standard harnesses
like cloud code or codecs it performs
even better across all these different
categories like optimizer design or
architecture design agentic coding math
reasoning etc. And some of these gains
are pretty significant and that's why
this arbor system is so interesting.
This is a system for AI models to do
more structured and persistent work. The
awesome thing is they've released this
already. So, if you click on this GitHub
repo and you scroll down a bit here, it
contains all the instructions on how to
get started on your computer. Also, this
is under the Apache 2 license, which is
very permissive. You can even use this
for commercial purposes. If you're
interested in reading further, I'll link
to this main page in the description
below. If you've ever had a creative
idea that sounded exciting at first, but
then turned into a mess of tabs and
different tools, then you should
definitely check out Luma Agents by Luma
AI, the sponsor of this video. Luma
Agents is basically an AI workspace
designed to help you bring creative
projects to life from start to finish.
But the key difference is that it's not
just another AI generator. It's built
around agents that can collaborate with
you across the entire creative process.
So instead of jumping between separate
tools for brainstorming, planning,
design, video generation, and execution,
you can work inside one connected canvas
where the agents understand what you're
trying to make and keep the whole
project in context. And that context is
a big deal. Luma's technology is built
around understanding the physical world,
including motion, space, camera
movement, and how things behave
visually. So, when you're working on
something creative, the agents are not
just responding to isolated prompts.
They can help reason about the actual
world your project is trying to
represent. You can start with a rough
concept for a video, a brand idea, or a
visual scene. From there, the agents
help you develop the idea, organize the
next steps, and coordinate different
parts of the workflow behind the scenes.
In other words, you stay in the role of
creative director. You bring the taste,
vision, and final judgment. The agents
help you structure the chaos and handle
more of the execution layer so you don't
get stuck managing 10 different tools
just to make one thing. And that's what
makes it feel really useful for
creators. You can test more ideas,
iterate more freely, and keep the whole
project moving without losing track of
the original vision. Whether you're
making videos, designing visuals, or
building campaigns, Luma Agents is the
perfect tool to use. Simply scan the QR
code, or use the link in the description
below to get started. Also this week,
one of the top Chinese labs out there,
Kimmy, released their best and latest
model, Kimmy K2.7 Code. Now, Kimmy has
always been at the top of the
leaderboards in terms of open- source
models, and their K2.6 6 model was
already very good, but this week they
made an even more performant model,
which is super exciting. So, as you can
see from all these different benchmarks,
not only is Kimmy K 2.7 better than
K2.6, but it's also edging very close to
the top closed models out there,
including GPT 5.5 and Opus 4.8. Super
impressive for an open- source model.
Also, if you look at this chart, the
X-axis is the number of tokens used. So,
the lower the better. It means it's more
efficient. And then Y-axis is
performance. So higher is better.
Ideally, you want to move towards the
upper left corner. And that's indeed
what you get with Kimmy K2.7, which is
the blue dots. Here are some additional
specs for your reference. So as a
state-of-the-art open source model, this
is quite huge at a trillion parameters.
This is a mixture of experts models. So
when you use it, only 32 billion
parameters are active. So this is very
efficient. Now, as before, they've
already open- sourced Kim K 2.7. So, if
you go on HuggingFace and you look at
the model at a trillion parameters, this
is quite huge at almost 600 GB in size.
So, you'll need a pretty massive setup
in order to run this locally. Here, they
say that K2.7 is designed to do even
better in terms of reasoning efficiency.
So, it overthinks less of the time.
Plus, it has improved instruction
following and it's more optimized for
long horizon coding tasks. On this page,
if you scroll down a bit, it contains
all the instructions on how to download
and run this locally on your computer.
Alternatively, if you don't have a beast
of a setup at home to run this, you can
also subscribe to a Kimmy Code to run
this. If you're interested in reading
further, I'll link to this main page in
the description below. Now, a few weeks
ago, one of my favorite labs, Miniax,
announced their latest and best model,
Miniax M3, and they said it will be open
weights. And this is one of the best
open- source models you can use. If you
look at this chart from Artificial
Analysis, then you can see that Miniax
M3 is indeed the leading open-source
model out there, even ahead of Kimmy and
Mimo and Deepseek V4. Well, this week
they stuck to their word and released
the model, which is fantastic. By the
way, this is an incredibly performant
but relatively small model. So this is
only like 427 billion parameters.
Whereas Kimmy or Deepseek are a trillion
parameters. So this is less than half
the size. Now this is a mixture of
experts models. So when you use it only
23 billion parameters are active. Plus
this has a massive 1 million token
context window. Meaning you can fit a
ton of information into your prompt at
once. Now at over 400 billion
parameters, this model is still pretty
huge. So the full model is over 850 GB
in size. If you want to run this
locally, again, you'll need some insane
hardware to do so. They also released a
smaller FXMP8 version if you have the
right hardware to run this. And this is
much smaller at 444 GB in size. Now, the
reason why they were able to jam-pack
all this intelligence in just 400
billion parameters, whereas the other
top open models are over a trillion
parameters, is largely due to this
mechanism called the Miniax sparse
attention. This allows the model to
understand huge amounts of context
faster and cheaper. The awesome thing is
they've also released the code and
technical report on this. So in really
simple terms, this takes part of the
language model that normally tries to
look at everything in the context window
and instead teaches it to quickly pick
the most useful blocks of information
before doing the expensive attention
step. The idea is to add a lightweight
indexing branch that stores chunks of
memory, chooses the top chunks for each
attention group, and then lets the main
attention branch focus on only the
selected blocks. You can think of it as
like giving the model a smart table of
contents before it reads a massive book.
By the way, Miniax isn't the only one to
release such a mechanism. Deepseek also
has their own technique called Deepseek
sparse attention. And then Kimmy also
released their own way to handle a huge
context window efficiently. In fact, if
you're interested in learning a bit
about both these architectures, I'll
link to this video and this video which
go over these details in simple terms.
Now, Miniax also released the technical
details to their own sparse attention
mechanism. It's great that these top
Chinese labs are actually publishing
this information for the community to
learn from, whereas the top closed labs
haven't really contributed anything.
Anyways, if you are interested, I will
link to this page on miniax sparse
attention as well as this page on the
miniax m3 model in the description
below. Also, this week we have another
frontier open-source model. This is
called Nex and2 and this is built around
the idea that reasoning should be used
for action including coding and long
horizon tasks. Note that this is trained
from Quen 3.5 as the base model. Instead
of creating search, coding, and tool use
as totally separate behaviors, NexN2
tries to use one consistent thinking
pattern across all of them. The useful
part is this adaptive reasoning feature.
The model can decide when to think
harder and when not to instead of
forcing expensive reasoning on every
single task. And if you look at these
benchmarks, this performs incredibly
well. This even beats leading open
models like Deepseek V4 and GLM 5.1
across various agentic and coding
benchmarks. In terms of deep suite, it
even scores 33.6 significantly beating
the other open models. Now note that
this family consists of two different
models. There's a pro variant which is
397 billion parameters. This is a
mixture of experts models. So think of
it as like a team of AIs working
together. And when you use it, only 17
billion parameters are active. So this
is fairly efficient. And then they also
have a mini model. This is a lot smaller
at 35 billion parameters and only 3
billion parameters are active when you
use it. The awesome thing is this is out
already. So at the top here, if you
click on this GitHub link, it takes you
to this page which contains all the
instructions on how to install this.
Note that for the Pro model, because it
is quite huge at almost 400 billion
parameters, the total size of the model
is 794 GB whereas the mini model is a
lot more accessible. This is only 70 GB
in size. If you're interested in reading
further, I'll link to this main page in
the description below. Also, this week
we have yet another Frontier texttospech
model which is actually very performant.
So, this is quite a small 2 billion
parameter model. You can just input a
few seconds of someone's voice and this
AI can basically get that voice to speak
out this transcript. So, it has like
zeroot voice cloning. Let's listen to a
few examples. The fact we were able to
complete the construction work on
schedule is a testament to everyone's
hard work. How could you possibly
believe such an obvious lie? Daniel
questioned incredulously, his eyebrows
raised in disbelief. The story doesn't
even make logical sense if you think
about it for more than 5 seconds.
Unfortunately, we haven't made any real
progress. Uh, well, Lisa, Trick
stuttered, nervously extending the ring
he bought for God knows how much.
Will will you marry me?
>> As you can hear, it sounds almost
exactly like the original reference
voice. Plus, it can even do subtle stuff
like stuttering. Or here's another
example where this can even do
whispering.
[snorts]
Yeah.
[snorts]
[gasps]
[sighs]
This can also do multiple languages,
even if the input reference voice
doesn't speak that language. For
example, let me play you the input voice
first and then let's get it to speak out
all these different languages,
>> which is what has brought us to this
point in the first place.
Fore.
Now, if you look at this chart, the
x-axis is the error rate. So, the lower
the better and the y-axis is the
similarity compared with the reference
voice. So, the higher the better.
Ideally, you want to be in this upper
left corner. And that's exactly where.
TTS is even out competing other recent
TTS models like CTDTS, Vox CPM, Quen 3
TTS, and Index TTS 2. All of which I
featured on my channel before. The
awesome thing is they've released this
already and it's under the Apache 2
license which has very minimal
restrictions. You can even use this for
commercial purposes. If you click on
this GitHub repo here, it contains all
the instructions on how to download and
run this on your computer. Now, at 2
billion parameters, this is fairly tiny.
So, the base model is only around 5 GB
in size. You should be able to fit this
on most consumer GPUs. If you're
interested in reading further, I'll link
to this main page in the description
below. Also, this week, Anthropic
releases their flagship model, Claude
Fable 5, which they claim is Mythos
quality. Now, I already covered a ton of
personal tests and specs and benchmarks
in my previous video, so see this if you
haven't already. I'll put it in the
description below. But a ton of other
crazy things happened over the course of
the week which I didn't cover in that
video. So let's talk about this a bit
more. I already mentioned that Fable is
severely nerfed. It appears to be one of
the most locked down frontier models
released to the public, especially
around AI research, model training,
cyber security, and biology. In many
cases, when you ask it questions related
to these fields, it will simply refuse.
But here's the biggest red flag and the
most controversial part of its design,
especially if you're asking it to do
things in AI research, machine learning,
or model training. Hidden deep in its
system card paper, which is over 300
pages long, is this paragraph, which
states the following. If you ask it to
do things in regards to like AI
research, training models, machine
learning, it secretly gives you a dumber
answer. It doesn't outright refuse you.
Instead, the model could intentionally
give you a significantly weaker or
incomplete answer or steer you somewhere
else, essentially sabotaging you. For
developers, especially in AI or machine
learning, this is a serious trust
problem. Now, of course, many users,
including myself, were not happy about
this. And in the middle of this week,
they did retract the sabotaging
mechanism, but honestly, the trust has
already been broken. They say that it's
not going to sabotage you and instead
it's going to explicitly refuse your
request. But we don't really know if
they're still secretly getting the cloud
models to weaken your answers behind the
scenes. And that's the main problem I
have with Anthropic. Even though they do
have very capable models, they
constantly gatekeep and rugpole their
users. It's really challenging to keep
using them. Now, that wasn't the only
update about Fable this week. So, last
night, Anthropic announced this. They
say that the US government issued a
directive requiring it to suspend all
access to Fable 5 and Mythos 5 for
foreign nationals, including foreign
national anthropic employees. Because of
that, Enthropic says it must disable
Fable 5 and Mythos 5 for all customers
to stay compliant, including all
Americans, by the way, while other
anthropic models are not affected. So,
uh, yeah, I guess we all don't even have
access to Fable now. What a crazy turn
of events. Now, this opens up a ton of
different unanswered questions. If the
US can issue this ban for Anthropic,
will they potentially also do it for
OpenAI or Google once they release
better and better models? And if so, is
US still the best place for AI labs if
the government can just suddenly require
you to suspend your product? And here
they're removing access for everyone.
But let's say they can offer it to US
citizens. How would that even be
enforced? Would they need to do like
KYC? Do they need to scan your ID or
scan your iris or something? There are
just a lot of unanswered questions.
Anyway, this whole Fable release took
quite a sharp turn at the end there. Let
me know in the comments what you think
of all of this. Now, shortly after the
Fable lockdown, one of my favorite open
labs, ZAI, was quick to jump in and
announce that GLM 5.2 2 is now available
on all GLM coding plans and the model
will be officially open sourced next
week which is fantastic. This is the
open source model I'm most looking
forward to. Now currently there's no
like system card or technical report or
benchmarks released for GLM 5.2. It's
all coming next week, but if you want a
head start and access GLM 5.2, I'll link
to this page where you can subscribe to
the GLM coding plan. Also, this week we
have a new AI called world tracing. This
is a new way to turn a single image or
even a short video into a layered 3D
model. So, this takes an image and
outputs not just the visible depth, but
multiple layers of hidden geometry
behind what the camera can actually see.
Think of it as like giving every pixel a
little stack of 3D points. The first
point is the surface which you can see,
but the later points are the parts
behind it. Things like the back of an
object or the wall behind furniture,
etc. What makes this useful is that once
you have this layered geometry, you can
do much more than just make a nice 3D
preview. It can lift the layered point
cloud into a textured mesh without extra
retraining. It can also support 3D edits
like adding or removing objects or even
guide video generation with 3D geometry
references. So, a really powerful and
flexible tool. The awesome thing is
they've released everything already. So
if you click on this code button and you
scroll down a bit here, it contains all
the instructions on how to install and
run this on your computer. Note that
there are three different models. One
model is for generating a static object.
Another one is for generating a static
scene. Another one is for generating a
moving 3D object. The model is quite
tiny. It's only 6.2 GB in size. So you
should be able to fit this on most
consumer GPUs. If you're interested in
reading further, I'll link to this main
page in the description below. Also,
this week we have a new AI called Flex
4D human. And this can turn simple human
videos into full body 4D
reconstructions. 4D here means 3D with
the added dimension of time. So
basically a moving 3D human model which
you can view from different angles. Now
this can take either a single video. So,
for example, here is one video
reference, and as you can see, it's able
to create a 3D model of this character
moving, which you can view at different
angles. And this is much more performant
than other methods. Now, instead of just
one video, you can also input multiple
videos. So, here's an example of two
different videos, and this gives it more
data and makes its generation even more
accurate. In addition to two references,
here's an example with four references,
which gives it even more data and makes
it even more consistent. What's
interesting is that it doesn't rely on
explicit skeletons or depth maps or
surface normals or any traditional
methods. This just uses AI to
reconstruct everything from the raw
camera footage. And because it's now
created a 3D model, you can then plug
this 3D model into any scene or edit it
further like this. If you scroll up to
the top of the page, they have released
a code button. And here it says the code
release is in preparation. So stay tuned
for that. If you're interested in
reading further, I'll link to this main
page in the description below. Also,
this week we have a new AI called video
MDM. This is about training AI to
generate 3D human motion without needing
any 3D mocap data. So, here what they
did is they took just ordinary videos in
2D, but they extracted accurate body
poses and used that to train a model
which can then generate coherent 3D
human motion. Here are some examples for
your reference. Just from a text prompt,
it can generate the full 3D motion of a
human performing the action, such as a
person waving with his right hand or a
person walking backwards in a straight
line. It understands a ton of different
actions like deadlifts or push-ups or
barbell dead rows. At the top of the
page, they have released the code to
this. So, if you click on this button
and you scroll down a bit here, it
contains all the instructions on how to
download and run this locally on your
computer. Note that this is under the
MIT license, which is fairly permissive.
If you're interested in reading further,
I'll link to this main page in the
description below. Also, this week, this
AI is pretty cool. It's called Surflow,
and this is a system for turning
multiple images of a scene into one
clean 3D model. It fuses all the input
views into one shared global 3D state.
It even works if the images are not
aligned or even if the camera positions
are not known. It keeps one compact
global understanding of the scene and
that understanding simply gets better
and better as you give it more views. So
instead of carrying around a huge amount
of repeated image information, Surflow
uses those extra photos to fill in
missing parts of the 3D model more
cleanly. Now at the top of the page,
they have released a code button to this
here. It says, "We will make our code
and data available as soon as possible."
So stay tuned for that. If you're
interested in reading further, I'll link
to this main page in the description
below. Also, this week we have Moverse.
This is an AI that turns a normal image
into a 360° panorama world, which you
can move around and interact with in
real time. So, here are some examples of
this in action. All you need is just one
image, and it would transform it into
pretty consistent 3D world. The
impressive part is that this works real
time at 8 frames per second on just a
single RTX490.
This isn't some really expensive
enterprise GPU. The important trick is
that Mover separates world construction
from video rendering. So, first it fills
in the missing surroundings with a 360°
panorama. So, the model is not only
working with a limited small part of the
image. Instead, it expands the original
image into a huge wide-angle panorama
like this. Afterwards, it turns that
panorama into a 3D Gausian model, which
acts like a reusable spatial memory of
the scene. Finally, the video renderer
takes views from that reference along
with a user controlled camera path, and
turns it into a coherent video. You can
see here it works with a ton of
different environments, including a
cartoon or 3D environment like this, or
even an oil painting like this, or even
Gibli style. Now, the quality isn't
great. This is quite blurry and low
resolution, but this is rendering eight
frames per second in real time on just a
consumer GPU. So, this is incredibly
efficient. At the top of the page, they
have released a code button. And here it
says the code and models are currently
under corporate compliance and review.
It'll likely take about a month to
review and then afterwards, they will
release all materials. If you're
interested in reading further, I'll link
to this main page in the description
below. Also this week we have a new
fully open-source image model from
Princeton called I1. This is fairly
performant. So here are some examples of
its generations. This can do
photorealistic images as well as
different art styles and anime. And this
is also great at text rendering. So here
are some examples of its generations
with a ton of text. Now that being said,
this is not the most performant model
out there and they were quite honest
about this. So this is behind Z image
and Quinn image likely behind ideoggram
and Ernie image as well which were
released recently. However, the biggest
value from this is that they open
sourced everything. Here it says that
leading open models just released the
models but not the training data or the
full recipe. So we don't really know how
this works. If we want to train an image
model from scratch, we don't really get
a lot of insights. But here the awesome
thing is they've fully open sourced
this. They released the models as well
as the training and inference code and
the data processing pipeline for this.
So if you click on this code button, it
takes you to this page and down here it
contains all the information on how you
can run this yourself as well as the
links to the data set and the data
pipelines. Note that this model is quite
tiny at 3 billion parameters and this is
around 12 GB in size. So it should be
able to fit on most consumer devices.
So, if you're looking to just generate
images, this isn't the best model out
there. I would not recommend this. But
if you want to study how an image model
can be trained from scratch, if you want
to study how to correctly format the
data set and the data pipelines to train
this, this is a very valuable resource.
If you're interested in reading further,
I'll link to this main page in the
description below. Also, this week, this
AI is pretty cool. It's called Anchor
World and this is basically a
firstperson world simulator where the
model is controlled by real human
motion. So here are some examples of how
it works. You would input some human
motion moving through a scene in 3D like
this. And the output is a firstperson
video of this person moving around the
scene. Think of it like instead of only
typing a prompt or pressing keys in a
world simulator, you guide it with a
person moving through the scene instead.
Now, in addition to inputting just the
human action in 3D, you also input some
anchor views like this. This is
basically an image for how that part of
the scene should look like. Then, after
plugging in the human motion, it's able
to render the scene from a firstperson
perspective of the human. Here are some
additional examples for your reference.
Note that this isn't just getting the
person to walk around. This simulator
can also get the person to interact with
objects such as operating the sink or
holding this knife. So, a pretty
interesting idea. This could be
potentially used for creating egocentric
videos for training humanoid robots to
do various tasks. Now, at the top of the
page, it does say that the code is under
review. So, it looks like they are
planning to release the models to this
in the future, which is great. For now,
if you're interested in reading further,
I'll link to this main page in the
description below. Also, this week, Meta
surprisingly released a really fast 3D
generator called Mesh Flow. So, this can
either take a prompt or a point cloud or
a reference image and generate an actual
3D mesh with vertices and edges, not
just a cloud of points or some soft
shape. So here are some generated meshes
for your reference. And the key here is
speed. A lot of generation methods
create meshes step by step. Almost like
writing a sentence a token at a time.
Well, mesh flow instead compresses that
mesh into latent space using something
called meshvae. And it's able to sample
from that space much faster. In fact,
here they claim that it's 18 times
faster than other generation methods.
The nice thing is they've released the
code to this already. So, if you click
on this code button and you scroll down
a bit here, it contains all the
instructions on how to run this on your
computer. If you're interested in
reading further, I'll link to this main
page in the description below. Also,
this week, we have this new AI system
called Millie Vid. This tackles one of
the hardest problems in AI video, which
is making long consistent videos. You
see, current video models can make some
impressive short clips that are like 5
to 20 seconds long, but once you ask
them to continue for a longer time, they
often lose consistency or forget what
different parts of the scene look like,
or it just gets a lot more errorprone.
Well, Milly Vid aims to solve this
problem. As you can see, even for really
long videos, it's able to maintain
consistency much better than other
methods. Now how they solve this is they
use something called a hierarchical
autoenccoder. Basically each frame is
represented at multiple levels of
detail. The coarse level captures the
big stuff like the layout and semantics
while the finer levels capture texture
and appearance. And then the model
generates video in a coarse to fine
rollout. And apparently this lets it
remember the scene and structure much
better and for much longer. Now at the
top they do have a code button. And here
it says that the paper will be released
on June 12th. At the time of this
recording, it's not out yet, but by the
time this video is published, the code
should already be out. If you're
interested in reading further, I'll link
to this main page in the description
below. And that sums up all the
highlights in AI this week. Let me know
in the comments what you think of all of
this. Which piece of news was your
favorite? And which tool are you most
looking forward to trying out? As
always, I will be on the lookout for the
top AI news and tools to share with you.
So, if you enjoyed this video, remember
to like, share, subscribe, and stay
tuned for more content. Also, there's
just so much happening in the world of
AI every week, I can't possibly cover
everything on my YouTube channel. So, to
really stay up to date with all that's
going on in AI, be sure to subscribe to
my free weekly newsletter. The link to
that will be in the description below.
Thanks for watching and I'll see you in
the next one.