AI Search kkLlzQqa7MY read

New robot waifus, GLM 5.2 craze, AI spas, new world models, new science agents: AI NEWS

Transcript: Done Yayin: 2026-06-20 20:27 YouTube
New robot waifus, GLM 5.2 craze, AI spas, new world models, new science agents: AI NEWS
Kanala don
Job gecmisi
Job Durum Deneme Worker Istek Baslama Bitis
SummarizeYouTubeTranscript #268 Done 1 learning-prod-worker-1 2026-06-22 01:38:04 2026-06-22 01:40:48 2026-06-22 01:40:59
FetchYouTubeTranscript #258 Done 1 learning-prod-worker-1 2026-06-22 00:27:31 2026-06-22 00:31:36 2026-06-22 00:32:23

Ozet

openai/gpt-4.1-mini-2025-04-14 - 2026-06-22 01:40
Indir

Ozet

Bu hafta yapay zeka dünyasında çok sayıda önemli gelişme yaşandı. Yeni tam vücut robot waifu Moya tanıtıldı; yüz ifadeleri ve bazı ev işleri yapabiliyor. GLM 5.2 adlı açık kaynak dil modeli sektörde büyük yankı uyandırdı; yüksek performansına rağmen uygun fiyatıyla dikkat çekiyor ve kullanıcılar için çeşitli boyutlarda sıkıştırılmış versiyonları çıktı. DreamXWorld adlı yeni dünya modeli, Unreal Engine ve gerçek dünya videoları ile eğitilerek uzun süreli, tutarlı ve etkileşimli sanal ortamlar yaratabiliyor. Ayrıca, Permavid adlı video düzenleme yapay zekası, görünüm ve 3D yapıyı ayrı hafızalarda tutarak video düzenlemelerinde tutarlılığı artırıyor.

Alibaba'nın Logos modeli, proteinlerden kimyasal reaksiyonlara kadar çoklu bilimsel alanları tek bir dilde anlayabilen açık kaynaklı bir model olarak öne çıktı. OpenAI ise Codex ile ekran kaydı üzerinden otomasyon yapmayı sağlayan "record and replay" özelliğini sundu. Sony'nin profesyonel oyunculara karşı üstün performans gösteren masa tenisi robotu ve Peking Üniversitesi'nin iki ayaklı humanoid robotu da dikkat çekti. Midjourney ise tamamen farklı bir alana yönelerek, su altında ultrasonik tarama yapabilen ve vücut sağlığını hızlıca analiz eden bir spa konsepti geliştirmeye başladı.

Ana Fikirler

  • GLM 5.2, açık kaynak en iyi dil modellerinden biri ve düşük hata oranıyla öne çıkıyor.
  • DreamXWorld, uzun süreli ve tutarlı sanal dünya yaratabilen yeni bir açık kaynak dünya modeli.
  • Permavid, video düzenlemede görünüm ve yapıyı ayrı tutarak tutarlılığı artırıyor.
  • Alibaba Logos, çoklu bilimsel alanları tek bir modelde birleştiren açık kaynak AI.
  • OpenAI Codex "record and replay" ile kullanıcı hareketlerini öğrenip otomasyon oluşturuyor.
  • Sony'nin masa tenisi robotu, gerçek zamanlı spin algılayarak profesyonel oyuncuları yenebiliyor.
  • Midjourney, tıbbi görüntüleme için su altı ultrasonik tarama yapan spa konsepti geliştiriyor.
  • Yeni tam vücut robot waifu Moya, temel yüz ifadeleri ve ev işleri yapabiliyor.
  • LTX Trainer 2, video modeli LTX için esnek eğitim ve ince ayar platformu.
  • Stil transferi için Telestyle V2, farklı içerik ve stil kombinasyonlarında başarılı sonuçlar veriyor.

Uygulanabilir Notlar

  • GLM 5.2’nin sıkıştırılmış versiyonları, yüksek performanslı modelleri tüketici cihazlarında çalıştırmak için ideal.
  • DreamXWorld ve Permavid gibi modeller, oyun ve video prodüksiyonlarında yeni yaratıcı araçlar sunabilir.
  • Alibaba Logos modeli, biyoteknoloji ve kimya alanlarında araştırma ve tasarım süreçlerini hızlandırabilir.
  • OpenAI Codex’in "record and replay" özelliği, karmaşık iş akışlarının otomasyonunda kullanıcı eğitimi gerektiren durumlarda kullanılabilir.
  • Midjourney Medical spa konsepti, tıbbi görüntüleme alanında yeni bir kullanıcı deneyimi yaratabilir; ancak regülasyonlar kritik.
  • Robotik alanında Sony ve Peking Üniversitesi’nin gelişmeleri, insana benzer hareket kabiliyeti ve hassasiyet gerektiren uygulamalarda ilerleme sağlıyor.
  • Stil transferi ve görüntü düzenleme araçları, yaratıcı içerik üreticileri için yeni seçenekler sunuyor.

Anahtar Kavramlar

  • GLM 5.2: Açık kaynak yüksek performanslı dil modeli.
  • DreamXWorld: Tutarlı ve etkileşimli sanal dünya modeli.
  • Permavid: Video düzenlemede tutarlılığı artıran AI.
  • Logos (Alibaba): Çoklu bilimsel alanları kapsayan birleşik AI modeli.
  • Codex "record and replay": Ekran kaydından otomasyon öğrenme özelliği.
  • Masa tenisi robotu (Sony): Spin algılayabilen profesyonel robot.
  • Midjourney Medical Spa: Su altı ultrasonik vücut tarama sistemi.
  • LTX Trainer 2: Video modeli için eğitim ve ince ayar aracı.
  • Telestyle V2: Gelişmiş stil transfer AI aracı.
  • Universal Manipulation Exoskeleton: Robot hareketlerini insan hareketleriyle kontrol eden dış iskelet.

Transcript

Video metni
en markdown 2026-06-22 00:32 youtube-transcript-api:generated
Indir
AI never sleeps and this week has been
absolutely insane. Your full body robot
waifu is coming soon. The new number one
open source model GLM 5.2 takes the
industry by storm and now you can even
run this locally on high-end consumer
devices. We have a new open source image
model which can also edit images just
like Nano banana. We also have a new
open video editor that's super
consistent. Midjourney, the image
generation company, now turns into a spa
hotel. Like seriously, we have an
incredibly powerful and open source
world generator. Alibaba releases an
open unified model for science. It can
understand things like proteins, small
molecules, chemical reactions and more.
We have some ridiculous humanoid robot
demos and a lot more. So let's jump
right in. First up, we have a new world
model called DreamXWorld. With just a
prompt or reference images, this lets
you create a world that can be explored,
controlled and changed with actions or
event prompts. Basically, it takes in
instructions like movements or camera
control or events and outputs this
environment which keeps updating and
evolving over time. Notice that this is
super flexible. This works with a ton of
different subjects and environments. For
example, you can even add this drone
here or ride a motorcycle like this or
drive a car or do a variety of different
stuff. Now, what's interesting is that
it's trained on a mix of Unreal Engine
and gameplay data as well as real world
videos. So it learns both realistic
movement and more game-like interaction.
It can generate really long videos over
hundreds of frames with persistent
memory. So even if you look away and you
look back, the scene remains consistent
the whole time. Especially if you
compare this with other similar models
like World or HY World Play, then as you
can see DreamXWorld is even more
consistent. The awesome thing is they
released this already. So if you scroll
up to the top and you click on this code
button and you scroll down a bit here it
contains all the instructions on how to
download and run this locally on your
computer. Note that this is based off of
12.2. There's a smaller 5 billion
parameter version, which is what they
released now, but they're also planning
to release a larger and better quality
14B version in the future. Currently,
this 5B version is fairly tiny at like
21 GB in size, so you should be able to
fit this on both consumer devices. If
you're interested in reading further,
I'll link to this main page in the
description below. Also this week we
have a new AI called Permavid. This is
focused on one of the biggest problems
in AI video editing, which is
consistency. Right now when you edit a
generated video, for example, adding an
object or replacing an object or
changing the style of the video, the
model often forgets the change later on.
So, it doesn't really have nice
persistent memory and its edits are not
really consistent. But, Permavid tries
to fix this by giving the video a better
memory system. Think of it like two
separate memory banks. One remembers
what things look like and the other
remembers the 3D structure behind the
scene. So, if you make a global edit
like changing the whole style, it can
keep the geometry stable while carrying
the new look forward. And if you make a
local edit like changing an object, it
can remember that object while keeping
the rest of the scene untouched. In
other words, it separates appearance
from structure, which makes long video
edits much more stable over time. Now,
at the top of the page there's a code
button, so if you click on this and you
scroll down a bit, it contains all the
instructions on how to download and run
this locally on your computer. Note that
this is 29 GB in size and this uses 12.1
and Vace as the base video generator.
So, you'll need a high-end GPU to be
able to fit everything. Plus they also
released the data set to this, which is
like over 400 GB. If you're interested
in reading further, I'll link to this
main page in the description below. Also
this week, we have a new AI from Kling
called Omni Director. This is a system
that allows you to clone the camera
motion of a video onto another video. So
here are some examples of this in
action. This is really important because
instead of just describing the camera
movement with a text prompt like pan
left or zoom in, instead you can now
give it a reference video and it learns
the camera behavior from that video.
Then it uses that to animate a new
source image. So basically the input is
a reference video with camera motion
plus a new image and it would generate a
new video from that new image based on
the camera movement. What's really cool
is that it supports much more than just
simple pans and tilts. It can also do
quite dramatic aerial moves and diving
motions, multi-shot sequences and
coherent transitions. For example here,
notice that the reference video on the
left is made of two different shots.
It's able to apply the same cut on the
reference video. Here's another example
where for the reference video, we have a
third-person shot of the biker and then
a first-person view of the biker. Well,
it's able to apply the same thing to
this eagle flying in the sky. And it's
also able to separate this clip into
multiple shots just like the reference
video. Or here's another example where
we can apply the exact camera movements
from an ad onto a completely different
video. The really cool thing is that
this can even handle special techniques
like dolly zooms or bullet time effects
or lens distortions. So here's a bullet
time effect and then here's a dolly zoom
effect and here's some really wide fish
eye distortion. It's able to understand
this and apply it to the new video. Now
at the top, they have released a code
button but currently they haven't
released any code or models yet.
Hopefully they will open source this.
For now, if you're interested in reading
further, I'll link to this main page in
the description below. Also this week we
have a very performant open source image
generator and editor called Boo Goo
image. What a crazy name by the way. So
this is an image model just like Nano
Banana or GPT image. Not only can this
generate images using a text prompt, but
this can also take in reference images
and edit images. So here are some
examples. You can see it's able to
generate really photorealistic images
like this and it's also great at
generating text and infographics and
posters. So you can feed it a ton of
instructions and text in your prompt and
it's able to handle and generate all of
these different components. Here are
some additional examples for your
reference. And it's also pretty good at
handling anatomy. It has a really good
world understanding of like different
company logos or celebrities or even
existing interfaces like Instagram or
TikTok. So for example, you could get it
to generate an Instagram profile
screenshot like this. Pretty impressive.
Here if you look at their leaderboard
called Boo Goo Arena, then here they
claim that it does outperform existing
open source image generators like Z
image or Kwen or Hi Dream. Same with
image editing. So here they claim that
it's even better than other image
editors like Kwen image edit or Long Cat
image edit. Although they didn't include
Flux Client on this list. Now for my
initial test, I actually did not find it
as good as Z image for photorealistic
stuff and it's not as good as Flux
Client in terms of photo editing. Plus
this does tend to generate slower than
Flux. So that's why I'm on the fence of
doing a full tutorial on it. I don't
think the quality is as good as the best
open image models out there right now.
If you scroll up to the top of the page,
they have released a GitHub to this. So
if you click on this button and you
scroll down here, it contains all the
instructions on how to download and run
this locally on your computer. Plus
support for ComfyUI is already here. The
really awesome thing about this is it's
under the Apache 2 license, which has
very minimal restrictions. You can even
use this for commercial purposes.
Whereas, the best open image generator
out there, Audiogram 4, has a quite
restrictive non-commercial license. Same
with the best open image editor out
there, Flux Cline, it's also, I believe,
under their own Flux license. So, this
new Bubu model is much more permissive.
Note that there are two different types
of models. There's a base model, which
is just like text-to-image, and then
there's an edit model, which is for
editing images. Now, the full base model
is 20 GB in size. The FP8 version is a
bit smaller at 10 GB. They also have a
turbo version, which allows you to
generate images a lot faster at only
four steps, but at the sacrifice of some
quality. And then, in terms of image
edit, they don't have a turbo version.
They only have the full model, which is
also 20 GB in size, and a compressed FP8
version, which is roughly 10 GB. You
should be able to fit this on like
medium-to-high-end GPUs. If you're
interested in reading further, I'll link
to this main page in the description
below. Also this week, we have something
called Universal Manipulation
Exoskeleton built by Alibaba's Ant
Group. As you can see from this video,
this is a wearable robot control system
for teaching robots how to move and
handle real-world physical tasks. So,
the human would wear this upper body
exoskeleton, and this allows it to
control a robot by cloning their exact
movements. The system records both the
arm motion and also the force or torque
feedback. So, it doesn't just capture
where or how the human moves, but also
what the human feels when touching,
pushing, pulling, or manipulating
objects. And this is important because,
especially with household robots, they
need to deal with contact. They need to
understand when something is stuck or
heavy or blocked or hidden from view.
So, these demos are very hands-on. So,
with this, you can first demonstrate how
to do a ton of different household
tasks, and then afterwards, it will be
able to learn how to do these
autonomously. So, here's the training
process, and then afterwards, the robot
can just open the fridge and take out a
can of Coke by itself. Currently, they
haven't released anything yet, but here
it says the code is coming soon, so
hopefully they will open source this.
For now, if you're interested in reading
further, I'll link to this main page in
the description below. Also this week,
the Tongyi Lab from Alibaba releases a
really useful AI model for science. It's
called Logos, which stands for language
of generative objects, and here's some
background on this idea. You see,
science has many different domains or of
languages. For example, we have
proteins, small molecules, materials,
antibodies, reaction systems, and each
domain has their own format. For
example, a protein is built completely
differently compared to a small
molecule. Well, Logos is basically an AI
model that tries to understand all of
this in just one shared grammar. Think
of it as like a unified model for
understanding all these different
domains. Now, in simple terms, it breaks
data from all these different subjects
down into tokens, and then they trained
a model on that. Much like how a large
language models turn words into tokens
and then learn how to understand,
predict, and generate natural language.
So, without going into that much detail,
it's essentially the same thing, but
instead of training it on English, this
model learns how to predict and generate
across multiple scientific domains. This
means the same framework can be used for
things like designing proteins or
antibodies, predicting binding sites, or
generating materials, and so on and so
forth. Now, they released a family of
three different models. The largest is 8
billion parameters, whereas the smallest
is 1 billion. But as you can see, if you
compare this with other similar
competitors, this model is a lot more
performant across all these different
benchmarks, including ligand design and
material generation and protein editing
and antibody design. And the awesome
thing is they released this. So, on this
page it contains all the instructions on
how to download and run this locally on
your computer. And this is under the
Apache 2 license, which has very minimal
restrictions. You can even use this for
commercial purposes. If you look at the
largest 8 billion parameter model, this
is fairly tiny at only 16 GB in size.
So, you should be able to run this on
like most consumer devices. If you're
interested in reading further, I'll link
to this main page in the description
below. Also this week we have LTX
Trainer 2, which is a training and
fine-tuning package for the video model
LTX. If you're not familiar with LTX,
their latest 2.3 model is like the
leading open-source video model with
audio natively built in. If you're
interested in learning more, see this
video for a full tutorial. Well, this
week they released something called LTX
2 Trainer, which allows you to train and
fine-tune this LTX model. For example,
if you want a model to consistently
generate a certain character or object
or VFX, you can plug your data through
this to train a LoRA from it. Now, this
is a super flexible platform. It
supports various workflows like
extension, video inpainting and
outpainting, as well as text to audio,
audio extension or inpainting, video to
video transformations. So, if you want
to get LTX to create a really specific
thing, whether it's a character or
camera effect, or a certain style of
transformation, this is the official way
to fine-tune a model on that. Now, on
this page it contains a quick start
guide plus the data set preparation
instructions and everything else you
need to know to run this. If you're
interested in reading further, I'll link
to this main page in the description
below. Also this week OpenAI releases a
feature which I think is really useful.
It's called record and replay, and just
like the name implies, here's how it
works. You can just take a screen
recording of you doing a certain task
and then feed that video into Codex.
Codex watches and understands what you
did and then turns that process into a
reusable skill. So, here's a demo of how
this works.
>> This time, I'm going to have Codex watch
me so it can learn how we do it.
As I go, I pull in the title and
description, add the thumbnail and
English captions, and save the video as
private.
When I'm done, Codex reviews the
recording and turns what it learned into
a skill.
It remembers where our metadata lives,
how the upload package is organized, and
how we add captions, save, and verify
each upload.
Now, I'm going to open a fresh thread,
attach the next video package, and ask
Codex to handle it.
And now Codex handles this next one for
me.
It matches the package to the right row
in the spreadsheet, fills in the
metadata, adds the thumbnail and English
captions, uploads the video as private,
and then verifies [music] everything was
saved correctly.
>> You can see how powerful this is,
especially for workflows where it's
really hard to just describe with text.
Instead, you can just take a screen
recording of you doing the entire
workflow. This could be like filing an
expense or booking something, publishing
a video, whatever your workflow might
be, you can just feed that recording of
yourself doing it through Codex and it
can create a skill which can be used
again. Now, it doesn't work all of the
time, so the catch is that it works best
when the workflow is stable and the
success criteria is clear. Currently,
they say that this feature is only
available on Mac OS with computer use
enabled and it's currently not available
in the EU or related areas. But you
know, this is a really useful feature.
This makes automation feel less like
randomly prompting an AI agent and
hoping it gets it right to more like
training an AI assistant by just
demonstrating it yourself. Hopefully,
they will also release this for the
Windows Codex app in the future. If
you're interested in reading further,
I'll link to this main page in the
description below. If you want to create
unlimited AI videos, definitely check
out Higgsfield, the sponsor of this
video. They just launched Higgsfield
Seed Dance Unlimited, which gives you
unlimited access to Seed Dance 2.0 fast,
one of the best AI video models in the
world. And this is a pretty big deal
because everywhere else, AI video is
usually capped by credits or generation
quotas. But on Higgsfield, you can run
Seed Dance with no limits. If you grab
any eligible plan between June 20th to
27th, you unlock unlimited Seed Dance
2.0 fast through to July 17th. The
earlier you lock in, the more
generations you get. Seed Dance is
especially impressive because it can
generate cinematic videos with realistic
motion, strong physics, and even native
audio that's synced to the visuals. So
instead of generating a silent clip and
adding sound later, the video and audio
are created together. You can use it for
short films, action scenes, product
videos, music videos, and even social
media. I'm especially impressed by the
multi-shot storytelling. You can
generate scenes with consistent
characters across different cuts while
also controlling the camera movement,
motion style, and overall story
direction. And the workflow is super
creator-friendly. Inside Higgsfield, you
can combine text, images, videos, and
audio as references. You can upload
images, video, or audio and then use
them to guide the composition, camera
language, motion, and even sound. Try
Higgsfield's Seat Dance Unlimited today
using the link in the description below.
It's for a limited time only.
In robotics news this week we have an
incredibly impressive demo from this
autonomous table tennis robot called
Ace. Now we've had plenty of these table
tennis robots before, but previous ones
were pretty bad. This one actually
played against a professional human
player and it absolutely dominated the
game. So this robot is by Sony and as
you can see the impressive thing is if
you've ever played table tennis mildly
seriously you'll know it's not just like
hitting the ball. It's actually entirely
dictated by spin. So like slicing the
ball, there's top spin, back spin, side
spin. The ball doesn't just bounce in a
straight line. So for a robot to
actually successfully play against a pro
player and return shots from a pro
player it needs to have real-time spin
detection. It needs to have a vision
system that calculates the ball's
rotations and axis of spin in like less
than milliseconds just by watching it.
And then it also needs to know how to
slice the ball and do these different
types of spins itself when it hits the
ball. This is much harder to train
compared with just getting a paddle to
hit the ball at any angle. Now this
robot doesn't have legs so this is
mounted on a heavy duty motorized rail,
but it has incredible high speed latency
and actuations. It has to like instantly
move its entire mass from left to right,
also adjust its arms wrist angle and
strike the ball within a fraction of a
second. Table tennis is incredibly fast
and it's even faster when you play with
a professional human player so it needs
to be able to react to incoming balls
almost instantly. And you can see how it
plays is also really impressive. It like
kind of uses an adaptive strategy. It's
not just mindlessly or passively trying
to save balls or just hit the ball back.
It actively changes the ball's placement
or spin to force the human player into
making mistakes. It's like actively
trying to beat the human player. So, a
very impressive demo. This is like by
far the best robot we have right now in
terms of playing table tennis. Now,
speaking of table tennis, we also have
another demo, this time from the AGI bot
A3, and this is able to also
autonomously play table tennis against a
human. Now, this is not as impressive as
the Sony demo from before. However, this
is just a humanoid robot designed to do
a ton of other stuff like manipulating
objects, walking up and down stairs,
moving things, etc. Plus, this looks
more like a human with two legs. So, not
only does it have to maintain balance of
its body, but also react to and hit the
ball in a matter of seconds. In fact,
they say this is powered by something
called the Spike Ping Pong algorithm by
the Peking University, and this allows
its vision response to be 10 times
faster and enables millimeter level
precision for continuous rallies,
trajectory tracking, and whole body
planning. Very impressive how it's
actually able to do this while
maintaining balance on two legs. Also
this week, we have a new full body waifu
demo. So, this company called Droid Up
teases their full body humanoid robot
called Moya. You can see that this
robot's face is moderately realistic.
She can blink and tilt her head, and
even demonstrate some facial
expressions. And you can see how from
this demo, this robot is intended for
companionship. It's also able to do some
chores. So, it's like picking up this
bottle of orange juice and walking back
to the desk. It's even able to pour a
glass for the woman, or this could also
be used for elderly care, as you can see
over here. Now, I would say her face
isn't as realistic as some of the other
waifu humanoids I featured before. The
face of this Moya one seems to be kind
of rigid and hard. Something just looks
off about this face. Maybe it's the
shape of the facial features as well.
Something just doesn't really look
human. But, if you're looking for a
full-body robot waifu, here's another
option to put on your radar. Also this
week, OpenAI showed a near-autonomous AI
chemist that actually improved a real
medicinal chemistry reaction. Not just
on paper, but in an actual lab
experiment. So, the basic idea is this.
They connected GPT to this Maria system,
which is like an AI chemistry platform
hooked into a high-throughput lab. And
they gave it a pretty open-ended goal.
Find a way to improve an important class
of reactions. That's pretty much it. The
model came up with research ideas,
helped design the experiments, analyzed
the results, and even suggested
follow-up experiments while the human
chemists stayed in the loop for steering
corrections, and actually doing the lab
operations. Now, the reaction it focused
on is called the Chan-Lam reaction. This
is quite technical, but basically it's a
way to connect molecules by forming
carbon-nitrogen bonds. And this matters
because these kinds of bonds show up all
over medicinal chemistry. And the
specific challenge here was with the
sulfonamides, which are like useful
drug-like chemical groups. Now, in the
past, this reaction has not worked very
well with them. So, GPT suggested using
this oxidant called TEMPO as an
additive. And then after they tested
this at scale and found that indeed this
TEMPO additive significantly improved
the yield of that reaction. Here are
some further graphs showing how this new
TEMPO additive identified by GPT was
able to significantly outperform all
other oxidants. So, that's pretty cool.
This is an AI system that can help move
through the scientific loop. It can
read, propose, test, analyze, and
refine. In chemistry, this is usually
really slow and expensive with a bunch
of dead ends. But, with the help of AI,
we can actually automate a ton of
different steps in this workflow. And
here they showed that it actually helped
make a genuine discovery in medicinal
chemistry. At the top here, they
released a full technical paper on this
with even more details. If you're
interested in reading further, I'll link
to this main page in the description
below. Also this week, by far the best
open model was released. It's called GLM
5.2. Now, this is from my favorite lab,
ZAI. I've mentioned this countless times
on my channel before. I have high
expectations of them, but even this
release caught me by surprise. You see,
the previous model was just 5.1, so I
thought 5.2 would just be a tiny
upgrade, but it's actually much bigger
than I thought. I already did a full
review video on it this week with a ton
of impressive demos, so see this if you
want to learn more. I'm not going to
repeat anything from that video, but
after that was published, we have a ton
of new data on its performance on
various benchmarks and leaderboards. So,
let's go in some more detail. First of
all, GLM 5.2 is finally published on
Artificial Analysis. So, if you look at
their intelligence index, then you can
see it's among the best of the best
models out there, only behind the best
GPT and the best Claude. And it's by far
the best open model out there. I mean,
the gap between this and the next best
model, MiniMax M3, is huge. And if
you're wondering about the new Kimikaze
2.7 code that was also recently
released, that's all the way back here.
Now, what's even more impressive about
this is the price. So, even though it's
close to the level of the best GPT and
the best Opus, GLM 5.2 is way cheaper,
like half the cost of GPT 5.5 and five
times less than Claude Opus 4.8. So, if
you're looking for the most
cost-efficient option, which is just as
performant, GLM 5.2 is definitely the
best option. A few other things to note
here, if you look at this Artificial
Analysis omniscient's hallucination
rate, you can see that by far GLM 5.2
has the lowest hallucination rate
compared to the other frontier models.
If you look at Opus 4.8, it's all the
way over here. And also note that Claude
Fable is over here. So, GLM 5.2
hallucinates like 50% less than Claude
Fable. And then GPT 5.5 is all the way
over here. So, this hallucinates like
three to four times more than GLM 5.2.
Notice that this 28% doesn't mean it
hallucinates 28% of the time on average.
It's just 28% of this benchmark, which
has really tricky questions aimed to get
the model to hallucinate. So, if you're
working in a field that requires
factually accurate information like law
or medicine, GLM 5.2 is actually the
best frontier model you can use right
now. Now, some doubters say that GLM
might be bench maxed, but here's a
completely new benchmark that didn't
even exist before GLM was released. So,
artificial analysis released this new
benchmark called a a briefcase. And this
is their new benchmark for testing
models on long horizon knowledge work
tasks. So, workflows or really complex
projects that need to be run for like
many weeks and it's linked to many tasks
and thousands of input source files.
These were built by industry experts,
and after testing all these frontier
models, you can see that GLM 5.2 is
actually ranked number three. Even
beating the best GPT and only slightly
behind the best Opus. You can't bench
max for a benchmark that didn't even
exist before. Now, it's important to
refer to other leaderboards to get a
sense of how good a model is because
each leaderboard can be very different.
So, if you look at this other
leaderboard by LM Arena, you can see
that GLM 5.2 is all the way down here in
10th place, even below GPT 5.4 and the
older Opus models. And it seems like it
scored so low because it's not as
steerable as the other frontier models.
And if you look at live bench by Abacus
AI, note that GLM 5.2 is all the way
down here, even below the older Opus
models and below Gemini 3.1 Pro. And if
you look at the details of this, it
seems to do really well in terms of
agentic coding. In fact, this is the
model that scored the highest, but it's
not so good in terms of reasoning or
instruction following. Now, this model
is completely open weights under the MIT
license, which is super permissive, and
the full model is like 1.5 terabytes in
size. So, it's not going to be possible
for most consumers to run the full
model. But, behold the power of open
source. Because the model is released
already, people can fine-tune it and
mess with it however they want. So, in
just the span of like 2 days, we already
have a ton of different GLM 5.2 variants
and fine-tunes from the community. For
example, Unsloth just released very
compressed GGUF versions of GLM 5.2. So,
they released many different models of
different compression and sizes, so you
can select the one that fits your
hardware. For example, the smallest
1-bit version is just 223 GB, which is
pretty insane compression from like 1.5
TB. This can potentially be run on just
like two or three RTX 6000s or just one
or two DGX Sporks or even just a Mac
Studio. Though, I would not recommend
Apple for running local AI because it's
way slower. If you can, always get
Nvidia. So, that's the 1-bit and then
2-bit is also pretty small and
accessible and only 245 GB. Now, with
compression, of course, you're going to
get some degradation in quality. So,
here it says that the 1-bit model gets
around 76% accuracy compared to the full
model, but it's like 86% smaller. And
then the 2-bit version gets 82% accuracy
while being 84% smaller. So, it's still
pretty usable. It's not like these
compressed models are significantly
dumber. So, that's the power of open
source. Because the models are released,
the community can build on top of them
and create their own fine-tunes or more
compressed versions or even uncensored
versions of GLM. And in the span of just
one or two days, we already have a model
that can fit on high-end consumer
devices. You can potentially run GLM 5.2
at home. You don't have to sell your
house or kidney to run this. If you're
interested in reading further, I'll link
to this main page in the description
below. Also, this week we have a new
style transfer AI called Telestyle V2.
This is a really easy way for you to
apply the style of a reference image
onto a new image. For example, let's say
this is my input image and I want to
convert it into this art style. Well, I
can plug this through Telestyle V2 and
here's what I get. Or here's another
example. Let's say this is the input
image. I want to convert it into this
painting style. Well, here's the result.
Or let's say this is my input image. I
want to convert it into this watercolor
chibi style. After plugging it through
Telestyle V2, it gives me this. Here are
some additional examples for your
reference. Now, the problem with older
style transfer tools is that they only
work well in a very narrow setup. For
example, the content image has to be
realistic and the style image has to be
artistic. But sometimes, if you mix
those two up, then it doesn't really
transfer the style well. Well, Telestyle
V2 is able to handle all different
combinations. For example, you can first
take this input image and convert it
into these different styles. But then
you can also take these different styles
and then plug it through Telestyle V2
again to convert it into other styles.
And it's able to handle this very well.
The awesome thing is they've released
the code to this already. So, at the top
of the page, if you click on this code
button and you scroll down a bit, here
it contains all the instructions on how
to download and run this locally on your
computer. Here they say that they tested
it on an H100 with 80 GB, but you don't
need that much. This is based off of
Quant Image Edit, which requires like
just a mid-to-high-end consumer GPU to
run. If you're interested in reading
further, I'll link to this main page in
the description below. Also this week,
the image generation company Midjourney
is now pivoting to building spas. But
all jokes aside, they just announced
something completely different from
their image generation platform. So they
announced Midjourney Medical. It's not
an image generator, but instead this is
a kind of body scanner where you immerse
yourself in water. So it's more like
going to a spa instead of seeing a
doctor. Instead of waiting for an MRI or
booking appointments and only scanning
your body when something feels wrong,
Midjourney wants to create a system
where you can casually scan your body in
about just 60 seconds and track changes
over time. This prototype apparently can
build a much richer picture of your
health. Basically, they're trying to
make full-body imaging faster, cheaper,
and more widespread. The scanner itself
actually sounds pretty wild. So you step
into this pool of warm golden light and
slowly descend into water, and then it's
surrounded by these underwater sensors,
which use ultrasound. So it's kind of
like how dolphins communicate using
echolocation. So these sensors send
sound waves through your body from many
different angles and they listen to how
these waves bounce and change. And then
they use that data to reconstruct what's
inside you. Think of it as like taking
thousands and thousands of sound-based
snapshots of your body from every
direction and then using all of that
data to recreate a 3D map of your body.
Midjourney says the goal is for the
entire scan to take no more than 60
seconds. You go into the water, you come
out, and you're done. The technical side
is pretty crazy. So the scanner uses a
ring made of around half a million tiny
elements, each about the size of a grain
of sand. Then each one sends out
ultrasound waves and record the
returning ripples millions of times per
second producing terabytes of data every
second. That's pretty crazy. So, this
isn't just a medical device problem.
This is also a massive computing
problem. The hard part is taking all of
this wave pattern data and turning them
into useful images. Apparently, they say
it's able to do so because these waves
behave differently when they pass
through water, skin, fat, bone, muscle,
organs. And so, the system can use these
differences to reconstruct a detailed
internal map of the body. They say it's
similar to today's MRI scanning
technology but at nearly 100 times the
speed. The interesting part is how they
want people to experience it. They're
not building a scanner and putting it in
a clinic. Instead, they're building the
Midjourney Spa. This is first planned
for San Francisco in 2027 with hot tubs,
saunas, cold plunges, and scanning pools
built into the experience. So, you go to
the spa, you enjoy the space, and then
you also get a body scan there. The road
map is pretty ambitious. So, over the
next year Midjourney says they'll refine
the algorithms and hardware, run
research trials, and build better
prototypes, and prepare the first
research spa. They also point out that
regulation is a huge bottleneck here
because these medical features usually
require FDA approval. So, they're
starting with detailed body composition
maps first while submitting test results
for expanded capabilities over time. And
then in 2028 they want to scale to more
cities and move to an even better
scanner with custom silicon where they
expect much better image quality and
scan speed. The core idea is indeed
fascinating but this is a completely
different pivot compared to what they're
known for, which is image generation.
So, I'm not sure if they have the talent
or expertise to pull this off. Let me
know in the comments what you think of
this. If you're interested in reading
further, I'll link to this main page in
the description below. And that sums up
all the highlights in AI this week. Let
me know in the comments what you think
of all of this. Which piece of news was
your favorite and which tool are you
most looking forward to trying out? As
always, I will be on the lookout for the
top AI news and tools to share with you.
So, if you enjoyed this video, remember
to like, share, subscribe, and stay
tuned for more content. Also, there's
just so much happening in the world of
AI every week. I can't possibly cover
everything on my YouTube channel. So, to
really stay up-to-date with all that's
going on in AI, be sure to subscribe to
my free weekly newsletter. The link to
that will be in the description below.
Thanks for watching, and I'll see you in
the next one.