New robot waifus, GLM 5.2 craze, AI spas, new world models, new science agents: AI NEWS
Job gecmisi
| Job | Durum | Deneme | Worker | Istek | Baslama | Bitis |
|---|---|---|---|---|---|---|
| SummarizeYouTubeTranscript #268 | Done | 1 | learning-prod-worker-1 | 2026-06-22 01:38:04 | 2026-06-22 01:40:48 | 2026-06-22 01:40:59 |
| FetchYouTubeTranscript #258 | Done | 1 | learning-prod-worker-1 | 2026-06-22 00:27:31 | 2026-06-22 00:31:36 | 2026-06-22 00:32:23 |
Ozet
Ozet
Bu hafta yapay zeka dünyasında çok sayıda önemli gelişme yaşandı. Yeni tam vücut robot waifu Moya tanıtıldı; yüz ifadeleri ve bazı ev işleri yapabiliyor. GLM 5.2 adlı açık kaynak dil modeli sektörde büyük yankı uyandırdı; yüksek performansına rağmen uygun fiyatıyla dikkat çekiyor ve kullanıcılar için çeşitli boyutlarda sıkıştırılmış versiyonları çıktı. DreamXWorld adlı yeni dünya modeli, Unreal Engine ve gerçek dünya videoları ile eğitilerek uzun süreli, tutarlı ve etkileşimli sanal ortamlar yaratabiliyor. Ayrıca, Permavid adlı video düzenleme yapay zekası, görünüm ve 3D yapıyı ayrı hafızalarda tutarak video düzenlemelerinde tutarlılığı artırıyor.
Alibaba'nın Logos modeli, proteinlerden kimyasal reaksiyonlara kadar çoklu bilimsel alanları tek bir dilde anlayabilen açık kaynaklı bir model olarak öne çıktı. OpenAI ise Codex ile ekran kaydı üzerinden otomasyon yapmayı sağlayan "record and replay" özelliğini sundu. Sony'nin profesyonel oyunculara karşı üstün performans gösteren masa tenisi robotu ve Peking Üniversitesi'nin iki ayaklı humanoid robotu da dikkat çekti. Midjourney ise tamamen farklı bir alana yönelerek, su altında ultrasonik tarama yapabilen ve vücut sağlığını hızlıca analiz eden bir spa konsepti geliştirmeye başladı.
Ana Fikirler
- GLM 5.2, açık kaynak en iyi dil modellerinden biri ve düşük hata oranıyla öne çıkıyor.
- DreamXWorld, uzun süreli ve tutarlı sanal dünya yaratabilen yeni bir açık kaynak dünya modeli.
- Permavid, video düzenlemede görünüm ve yapıyı ayrı tutarak tutarlılığı artırıyor.
- Alibaba Logos, çoklu bilimsel alanları tek bir modelde birleştiren açık kaynak AI.
- OpenAI Codex "record and replay" ile kullanıcı hareketlerini öğrenip otomasyon oluşturuyor.
- Sony'nin masa tenisi robotu, gerçek zamanlı spin algılayarak profesyonel oyuncuları yenebiliyor.
- Midjourney, tıbbi görüntüleme için su altı ultrasonik tarama yapan spa konsepti geliştiriyor.
- Yeni tam vücut robot waifu Moya, temel yüz ifadeleri ve ev işleri yapabiliyor.
- LTX Trainer 2, video modeli LTX için esnek eğitim ve ince ayar platformu.
- Stil transferi için Telestyle V2, farklı içerik ve stil kombinasyonlarında başarılı sonuçlar veriyor.
Uygulanabilir Notlar
- GLM 5.2’nin sıkıştırılmış versiyonları, yüksek performanslı modelleri tüketici cihazlarında çalıştırmak için ideal.
- DreamXWorld ve Permavid gibi modeller, oyun ve video prodüksiyonlarında yeni yaratıcı araçlar sunabilir.
- Alibaba Logos modeli, biyoteknoloji ve kimya alanlarında araştırma ve tasarım süreçlerini hızlandırabilir.
- OpenAI Codex’in "record and replay" özelliği, karmaşık iş akışlarının otomasyonunda kullanıcı eğitimi gerektiren durumlarda kullanılabilir.
- Midjourney Medical spa konsepti, tıbbi görüntüleme alanında yeni bir kullanıcı deneyimi yaratabilir; ancak regülasyonlar kritik.
- Robotik alanında Sony ve Peking Üniversitesi’nin gelişmeleri, insana benzer hareket kabiliyeti ve hassasiyet gerektiren uygulamalarda ilerleme sağlıyor.
- Stil transferi ve görüntü düzenleme araçları, yaratıcı içerik üreticileri için yeni seçenekler sunuyor.
Anahtar Kavramlar
- GLM 5.2: Açık kaynak yüksek performanslı dil modeli.
- DreamXWorld: Tutarlı ve etkileşimli sanal dünya modeli.
- Permavid: Video düzenlemede tutarlılığı artıran AI.
- Logos (Alibaba): Çoklu bilimsel alanları kapsayan birleşik AI modeli.
- Codex "record and replay": Ekran kaydından otomasyon öğrenme özelliği.
- Masa tenisi robotu (Sony): Spin algılayabilen profesyonel robot.
- Midjourney Medical Spa: Su altı ultrasonik vücut tarama sistemi.
- LTX Trainer 2: Video modeli için eğitim ve ince ayar aracı.
- Telestyle V2: Gelişmiş stil transfer AI aracı.
- Universal Manipulation Exoskeleton: Robot hareketlerini insan hareketleriyle kontrol eden dış iskelet.
Transcript
AI never sleeps and this week has been absolutely insane. Your full body robot waifu is coming soon. The new number one open source model GLM 5.2 takes the industry by storm and now you can even run this locally on high-end consumer devices. We have a new open source image model which can also edit images just like Nano banana. We also have a new open video editor that's super consistent. Midjourney, the image generation company, now turns into a spa hotel. Like seriously, we have an incredibly powerful and open source world generator. Alibaba releases an open unified model for science. It can understand things like proteins, small molecules, chemical reactions and more. We have some ridiculous humanoid robot demos and a lot more. So let's jump right in. First up, we have a new world model called DreamXWorld. With just a prompt or reference images, this lets you create a world that can be explored, controlled and changed with actions or event prompts. Basically, it takes in instructions like movements or camera control or events and outputs this environment which keeps updating and evolving over time. Notice that this is super flexible. This works with a ton of different subjects and environments. For example, you can even add this drone here or ride a motorcycle like this or drive a car or do a variety of different stuff. Now, what's interesting is that it's trained on a mix of Unreal Engine and gameplay data as well as real world videos. So it learns both realistic movement and more game-like interaction. It can generate really long videos over hundreds of frames with persistent memory. So even if you look away and you look back, the scene remains consistent the whole time. Especially if you compare this with other similar models like World or HY World Play, then as you can see DreamXWorld is even more consistent. The awesome thing is they released this already. So if you scroll up to the top and you click on this code button and you scroll down a bit here it contains all the instructions on how to download and run this locally on your computer. Note that this is based off of 12.2. There's a smaller 5 billion parameter version, which is what they released now, but they're also planning to release a larger and better quality 14B version in the future. Currently, this 5B version is fairly tiny at like 21 GB in size, so you should be able to fit this on both consumer devices. If you're interested in reading further, I'll link to this main page in the description below. Also this week we have a new AI called Permavid. This is focused on one of the biggest problems in AI video editing, which is consistency. Right now when you edit a generated video, for example, adding an object or replacing an object or changing the style of the video, the model often forgets the change later on. So, it doesn't really have nice persistent memory and its edits are not really consistent. But, Permavid tries to fix this by giving the video a better memory system. Think of it like two separate memory banks. One remembers what things look like and the other remembers the 3D structure behind the scene. So, if you make a global edit like changing the whole style, it can keep the geometry stable while carrying the new look forward. And if you make a local edit like changing an object, it can remember that object while keeping the rest of the scene untouched. In other words, it separates appearance from structure, which makes long video edits much more stable over time. Now, at the top of the page there's a code button, so if you click on this and you scroll down a bit, it contains all the instructions on how to download and run this locally on your computer. Note that this is 29 GB in size and this uses 12.1 and Vace as the base video generator. So, you'll need a high-end GPU to be able to fit everything. Plus they also released the data set to this, which is like over 400 GB. If you're interested in reading further, I'll link to this main page in the description below. Also this week, we have a new AI from Kling called Omni Director. This is a system that allows you to clone the camera motion of a video onto another video. So here are some examples of this in action. This is really important because instead of just describing the camera movement with a text prompt like pan left or zoom in, instead you can now give it a reference video and it learns the camera behavior from that video. Then it uses that to animate a new source image. So basically the input is a reference video with camera motion plus a new image and it would generate a new video from that new image based on the camera movement. What's really cool is that it supports much more than just simple pans and tilts. It can also do quite dramatic aerial moves and diving motions, multi-shot sequences and coherent transitions. For example here, notice that the reference video on the left is made of two different shots. It's able to apply the same cut on the reference video. Here's another example where for the reference video, we have a third-person shot of the biker and then a first-person view of the biker. Well, it's able to apply the same thing to this eagle flying in the sky. And it's also able to separate this clip into multiple shots just like the reference video. Or here's another example where we can apply the exact camera movements from an ad onto a completely different video. The really cool thing is that this can even handle special techniques like dolly zooms or bullet time effects or lens distortions. So here's a bullet time effect and then here's a dolly zoom effect and here's some really wide fish eye distortion. It's able to understand this and apply it to the new video. Now at the top, they have released a code button but currently they haven't released any code or models yet. Hopefully they will open source this. For now, if you're interested in reading further, I'll link to this main page in the description below. Also this week we have a very performant open source image generator and editor called Boo Goo image. What a crazy name by the way. So this is an image model just like Nano Banana or GPT image. Not only can this generate images using a text prompt, but this can also take in reference images and edit images. So here are some examples. You can see it's able to generate really photorealistic images like this and it's also great at generating text and infographics and posters. So you can feed it a ton of instructions and text in your prompt and it's able to handle and generate all of these different components. Here are some additional examples for your reference. And it's also pretty good at handling anatomy. It has a really good world understanding of like different company logos or celebrities or even existing interfaces like Instagram or TikTok. So for example, you could get it to generate an Instagram profile screenshot like this. Pretty impressive. Here if you look at their leaderboard called Boo Goo Arena, then here they claim that it does outperform existing open source image generators like Z image or Kwen or Hi Dream. Same with image editing. So here they claim that it's even better than other image editors like Kwen image edit or Long Cat image edit. Although they didn't include Flux Client on this list. Now for my initial test, I actually did not find it as good as Z image for photorealistic stuff and it's not as good as Flux Client in terms of photo editing. Plus this does tend to generate slower than Flux. So that's why I'm on the fence of doing a full tutorial on it. I don't think the quality is as good as the best open image models out there right now. If you scroll up to the top of the page, they have released a GitHub to this. So if you click on this button and you scroll down here, it contains all the instructions on how to download and run this locally on your computer. Plus support for ComfyUI is already here. The really awesome thing about this is it's under the Apache 2 license, which has very minimal restrictions. You can even use this for commercial purposes. Whereas, the best open image generator out there, Audiogram 4, has a quite restrictive non-commercial license. Same with the best open image editor out there, Flux Cline, it's also, I believe, under their own Flux license. So, this new Bubu model is much more permissive. Note that there are two different types of models. There's a base model, which is just like text-to-image, and then there's an edit model, which is for editing images. Now, the full base model is 20 GB in size. The FP8 version is a bit smaller at 10 GB. They also have a turbo version, which allows you to generate images a lot faster at only four steps, but at the sacrifice of some quality. And then, in terms of image edit, they don't have a turbo version. They only have the full model, which is also 20 GB in size, and a compressed FP8 version, which is roughly 10 GB. You should be able to fit this on like medium-to-high-end GPUs. If you're interested in reading further, I'll link to this main page in the description below. Also this week, we have something called Universal Manipulation Exoskeleton built by Alibaba's Ant Group. As you can see from this video, this is a wearable robot control system for teaching robots how to move and handle real-world physical tasks. So, the human would wear this upper body exoskeleton, and this allows it to control a robot by cloning their exact movements. The system records both the arm motion and also the force or torque feedback. So, it doesn't just capture where or how the human moves, but also what the human feels when touching, pushing, pulling, or manipulating objects. And this is important because, especially with household robots, they need to deal with contact. They need to understand when something is stuck or heavy or blocked or hidden from view. So, these demos are very hands-on. So, with this, you can first demonstrate how to do a ton of different household tasks, and then afterwards, it will be able to learn how to do these autonomously. So, here's the training process, and then afterwards, the robot can just open the fridge and take out a can of Coke by itself. Currently, they haven't released anything yet, but here it says the code is coming soon, so hopefully they will open source this. For now, if you're interested in reading further, I'll link to this main page in the description below. Also this week, the Tongyi Lab from Alibaba releases a really useful AI model for science. It's called Logos, which stands for language of generative objects, and here's some background on this idea. You see, science has many different domains or of languages. For example, we have proteins, small molecules, materials, antibodies, reaction systems, and each domain has their own format. For example, a protein is built completely differently compared to a small molecule. Well, Logos is basically an AI model that tries to understand all of this in just one shared grammar. Think of it as like a unified model for understanding all these different domains. Now, in simple terms, it breaks data from all these different subjects down into tokens, and then they trained a model on that. Much like how a large language models turn words into tokens and then learn how to understand, predict, and generate natural language. So, without going into that much detail, it's essentially the same thing, but instead of training it on English, this model learns how to predict and generate across multiple scientific domains. This means the same framework can be used for things like designing proteins or antibodies, predicting binding sites, or generating materials, and so on and so forth. Now, they released a family of three different models. The largest is 8 billion parameters, whereas the smallest is 1 billion. But as you can see, if you compare this with other similar competitors, this model is a lot more performant across all these different benchmarks, including ligand design and material generation and protein editing and antibody design. And the awesome thing is they released this. So, on this page it contains all the instructions on how to download and run this locally on your computer. And this is under the Apache 2 license, which has very minimal restrictions. You can even use this for commercial purposes. If you look at the largest 8 billion parameter model, this is fairly tiny at only 16 GB in size. So, you should be able to run this on like most consumer devices. If you're interested in reading further, I'll link to this main page in the description below. Also this week we have LTX Trainer 2, which is a training and fine-tuning package for the video model LTX. If you're not familiar with LTX, their latest 2.3 model is like the leading open-source video model with audio natively built in. If you're interested in learning more, see this video for a full tutorial. Well, this week they released something called LTX 2 Trainer, which allows you to train and fine-tune this LTX model. For example, if you want a model to consistently generate a certain character or object or VFX, you can plug your data through this to train a LoRA from it. Now, this is a super flexible platform. It supports various workflows like extension, video inpainting and outpainting, as well as text to audio, audio extension or inpainting, video to video transformations. So, if you want to get LTX to create a really specific thing, whether it's a character or camera effect, or a certain style of transformation, this is the official way to fine-tune a model on that. Now, on this page it contains a quick start guide plus the data set preparation instructions and everything else you need to know to run this. If you're interested in reading further, I'll link to this main page in the description below. Also this week OpenAI releases a feature which I think is really useful. It's called record and replay, and just like the name implies, here's how it works. You can just take a screen recording of you doing a certain task and then feed that video into Codex. Codex watches and understands what you did and then turns that process into a reusable skill. So, here's a demo of how this works. >> This time, I'm going to have Codex watch me so it can learn how we do it. As I go, I pull in the title and description, add the thumbnail and English captions, and save the video as private. When I'm done, Codex reviews the recording and turns what it learned into a skill. It remembers where our metadata lives, how the upload package is organized, and how we add captions, save, and verify each upload. Now, I'm going to open a fresh thread, attach the next video package, and ask Codex to handle it. And now Codex handles this next one for me. It matches the package to the right row in the spreadsheet, fills in the metadata, adds the thumbnail and English captions, uploads the video as private, and then verifies [music] everything was saved correctly. >> You can see how powerful this is, especially for workflows where it's really hard to just describe with text. Instead, you can just take a screen recording of you doing the entire workflow. This could be like filing an expense or booking something, publishing a video, whatever your workflow might be, you can just feed that recording of yourself doing it through Codex and it can create a skill which can be used again. Now, it doesn't work all of the time, so the catch is that it works best when the workflow is stable and the success criteria is clear. Currently, they say that this feature is only available on Mac OS with computer use enabled and it's currently not available in the EU or related areas. But you know, this is a really useful feature. This makes automation feel less like randomly prompting an AI agent and hoping it gets it right to more like training an AI assistant by just demonstrating it yourself. Hopefully, they will also release this for the Windows Codex app in the future. If you're interested in reading further, I'll link to this main page in the description below. If you want to create unlimited AI videos, definitely check out Higgsfield, the sponsor of this video. They just launched Higgsfield Seed Dance Unlimited, which gives you unlimited access to Seed Dance 2.0 fast, one of the best AI video models in the world. And this is a pretty big deal because everywhere else, AI video is usually capped by credits or generation quotas. But on Higgsfield, you can run Seed Dance with no limits. If you grab any eligible plan between June 20th to 27th, you unlock unlimited Seed Dance 2.0 fast through to July 17th. The earlier you lock in, the more generations you get. Seed Dance is especially impressive because it can generate cinematic videos with realistic motion, strong physics, and even native audio that's synced to the visuals. So instead of generating a silent clip and adding sound later, the video and audio are created together. You can use it for short films, action scenes, product videos, music videos, and even social media. I'm especially impressed by the multi-shot storytelling. You can generate scenes with consistent characters across different cuts while also controlling the camera movement, motion style, and overall story direction. And the workflow is super creator-friendly. Inside Higgsfield, you can combine text, images, videos, and audio as references. You can upload images, video, or audio and then use them to guide the composition, camera language, motion, and even sound. Try Higgsfield's Seat Dance Unlimited today using the link in the description below. It's for a limited time only. In robotics news this week we have an incredibly impressive demo from this autonomous table tennis robot called Ace. Now we've had plenty of these table tennis robots before, but previous ones were pretty bad. This one actually played against a professional human player and it absolutely dominated the game. So this robot is by Sony and as you can see the impressive thing is if you've ever played table tennis mildly seriously you'll know it's not just like hitting the ball. It's actually entirely dictated by spin. So like slicing the ball, there's top spin, back spin, side spin. The ball doesn't just bounce in a straight line. So for a robot to actually successfully play against a pro player and return shots from a pro player it needs to have real-time spin detection. It needs to have a vision system that calculates the ball's rotations and axis of spin in like less than milliseconds just by watching it. And then it also needs to know how to slice the ball and do these different types of spins itself when it hits the ball. This is much harder to train compared with just getting a paddle to hit the ball at any angle. Now this robot doesn't have legs so this is mounted on a heavy duty motorized rail, but it has incredible high speed latency and actuations. It has to like instantly move its entire mass from left to right, also adjust its arms wrist angle and strike the ball within a fraction of a second. Table tennis is incredibly fast and it's even faster when you play with a professional human player so it needs to be able to react to incoming balls almost instantly. And you can see how it plays is also really impressive. It like kind of uses an adaptive strategy. It's not just mindlessly or passively trying to save balls or just hit the ball back. It actively changes the ball's placement or spin to force the human player into making mistakes. It's like actively trying to beat the human player. So, a very impressive demo. This is like by far the best robot we have right now in terms of playing table tennis. Now, speaking of table tennis, we also have another demo, this time from the AGI bot A3, and this is able to also autonomously play table tennis against a human. Now, this is not as impressive as the Sony demo from before. However, this is just a humanoid robot designed to do a ton of other stuff like manipulating objects, walking up and down stairs, moving things, etc. Plus, this looks more like a human with two legs. So, not only does it have to maintain balance of its body, but also react to and hit the ball in a matter of seconds. In fact, they say this is powered by something called the Spike Ping Pong algorithm by the Peking University, and this allows its vision response to be 10 times faster and enables millimeter level precision for continuous rallies, trajectory tracking, and whole body planning. Very impressive how it's actually able to do this while maintaining balance on two legs. Also this week, we have a new full body waifu demo. So, this company called Droid Up teases their full body humanoid robot called Moya. You can see that this robot's face is moderately realistic. She can blink and tilt her head, and even demonstrate some facial expressions. And you can see how from this demo, this robot is intended for companionship. It's also able to do some chores. So, it's like picking up this bottle of orange juice and walking back to the desk. It's even able to pour a glass for the woman, or this could also be used for elderly care, as you can see over here. Now, I would say her face isn't as realistic as some of the other waifu humanoids I featured before. The face of this Moya one seems to be kind of rigid and hard. Something just looks off about this face. Maybe it's the shape of the facial features as well. Something just doesn't really look human. But, if you're looking for a full-body robot waifu, here's another option to put on your radar. Also this week, OpenAI showed a near-autonomous AI chemist that actually improved a real medicinal chemistry reaction. Not just on paper, but in an actual lab experiment. So, the basic idea is this. They connected GPT to this Maria system, which is like an AI chemistry platform hooked into a high-throughput lab. And they gave it a pretty open-ended goal. Find a way to improve an important class of reactions. That's pretty much it. The model came up with research ideas, helped design the experiments, analyzed the results, and even suggested follow-up experiments while the human chemists stayed in the loop for steering corrections, and actually doing the lab operations. Now, the reaction it focused on is called the Chan-Lam reaction. This is quite technical, but basically it's a way to connect molecules by forming carbon-nitrogen bonds. And this matters because these kinds of bonds show up all over medicinal chemistry. And the specific challenge here was with the sulfonamides, which are like useful drug-like chemical groups. Now, in the past, this reaction has not worked very well with them. So, GPT suggested using this oxidant called TEMPO as an additive. And then after they tested this at scale and found that indeed this TEMPO additive significantly improved the yield of that reaction. Here are some further graphs showing how this new TEMPO additive identified by GPT was able to significantly outperform all other oxidants. So, that's pretty cool. This is an AI system that can help move through the scientific loop. It can read, propose, test, analyze, and refine. In chemistry, this is usually really slow and expensive with a bunch of dead ends. But, with the help of AI, we can actually automate a ton of different steps in this workflow. And here they showed that it actually helped make a genuine discovery in medicinal chemistry. At the top here, they released a full technical paper on this with even more details. If you're interested in reading further, I'll link to this main page in the description below. Also this week, by far the best open model was released. It's called GLM 5.2. Now, this is from my favorite lab, ZAI. I've mentioned this countless times on my channel before. I have high expectations of them, but even this release caught me by surprise. You see, the previous model was just 5.1, so I thought 5.2 would just be a tiny upgrade, but it's actually much bigger than I thought. I already did a full review video on it this week with a ton of impressive demos, so see this if you want to learn more. I'm not going to repeat anything from that video, but after that was published, we have a ton of new data on its performance on various benchmarks and leaderboards. So, let's go in some more detail. First of all, GLM 5.2 is finally published on Artificial Analysis. So, if you look at their intelligence index, then you can see it's among the best of the best models out there, only behind the best GPT and the best Claude. And it's by far the best open model out there. I mean, the gap between this and the next best model, MiniMax M3, is huge. And if you're wondering about the new Kimikaze 2.7 code that was also recently released, that's all the way back here. Now, what's even more impressive about this is the price. So, even though it's close to the level of the best GPT and the best Opus, GLM 5.2 is way cheaper, like half the cost of GPT 5.5 and five times less than Claude Opus 4.8. So, if you're looking for the most cost-efficient option, which is just as performant, GLM 5.2 is definitely the best option. A few other things to note here, if you look at this Artificial Analysis omniscient's hallucination rate, you can see that by far GLM 5.2 has the lowest hallucination rate compared to the other frontier models. If you look at Opus 4.8, it's all the way over here. And also note that Claude Fable is over here. So, GLM 5.2 hallucinates like 50% less than Claude Fable. And then GPT 5.5 is all the way over here. So, this hallucinates like three to four times more than GLM 5.2. Notice that this 28% doesn't mean it hallucinates 28% of the time on average. It's just 28% of this benchmark, which has really tricky questions aimed to get the model to hallucinate. So, if you're working in a field that requires factually accurate information like law or medicine, GLM 5.2 is actually the best frontier model you can use right now. Now, some doubters say that GLM might be bench maxed, but here's a completely new benchmark that didn't even exist before GLM was released. So, artificial analysis released this new benchmark called a a briefcase. And this is their new benchmark for testing models on long horizon knowledge work tasks. So, workflows or really complex projects that need to be run for like many weeks and it's linked to many tasks and thousands of input source files. These were built by industry experts, and after testing all these frontier models, you can see that GLM 5.2 is actually ranked number three. Even beating the best GPT and only slightly behind the best Opus. You can't bench max for a benchmark that didn't even exist before. Now, it's important to refer to other leaderboards to get a sense of how good a model is because each leaderboard can be very different. So, if you look at this other leaderboard by LM Arena, you can see that GLM 5.2 is all the way down here in 10th place, even below GPT 5.4 and the older Opus models. And it seems like it scored so low because it's not as steerable as the other frontier models. And if you look at live bench by Abacus AI, note that GLM 5.2 is all the way down here, even below the older Opus models and below Gemini 3.1 Pro. And if you look at the details of this, it seems to do really well in terms of agentic coding. In fact, this is the model that scored the highest, but it's not so good in terms of reasoning or instruction following. Now, this model is completely open weights under the MIT license, which is super permissive, and the full model is like 1.5 terabytes in size. So, it's not going to be possible for most consumers to run the full model. But, behold the power of open source. Because the model is released already, people can fine-tune it and mess with it however they want. So, in just the span of like 2 days, we already have a ton of different GLM 5.2 variants and fine-tunes from the community. For example, Unsloth just released very compressed GGUF versions of GLM 5.2. So, they released many different models of different compression and sizes, so you can select the one that fits your hardware. For example, the smallest 1-bit version is just 223 GB, which is pretty insane compression from like 1.5 TB. This can potentially be run on just like two or three RTX 6000s or just one or two DGX Sporks or even just a Mac Studio. Though, I would not recommend Apple for running local AI because it's way slower. If you can, always get Nvidia. So, that's the 1-bit and then 2-bit is also pretty small and accessible and only 245 GB. Now, with compression, of course, you're going to get some degradation in quality. So, here it says that the 1-bit model gets around 76% accuracy compared to the full model, but it's like 86% smaller. And then the 2-bit version gets 82% accuracy while being 84% smaller. So, it's still pretty usable. It's not like these compressed models are significantly dumber. So, that's the power of open source. Because the models are released, the community can build on top of them and create their own fine-tunes or more compressed versions or even uncensored versions of GLM. And in the span of just one or two days, we already have a model that can fit on high-end consumer devices. You can potentially run GLM 5.2 at home. You don't have to sell your house or kidney to run this. If you're interested in reading further, I'll link to this main page in the description below. Also, this week we have a new style transfer AI called Telestyle V2. This is a really easy way for you to apply the style of a reference image onto a new image. For example, let's say this is my input image and I want to convert it into this art style. Well, I can plug this through Telestyle V2 and here's what I get. Or here's another example. Let's say this is the input image. I want to convert it into this painting style. Well, here's the result. Or let's say this is my input image. I want to convert it into this watercolor chibi style. After plugging it through Telestyle V2, it gives me this. Here are some additional examples for your reference. Now, the problem with older style transfer tools is that they only work well in a very narrow setup. For example, the content image has to be realistic and the style image has to be artistic. But sometimes, if you mix those two up, then it doesn't really transfer the style well. Well, Telestyle V2 is able to handle all different combinations. For example, you can first take this input image and convert it into these different styles. But then you can also take these different styles and then plug it through Telestyle V2 again to convert it into other styles. And it's able to handle this very well. The awesome thing is they've released the code to this already. So, at the top of the page, if you click on this code button and you scroll down a bit, here it contains all the instructions on how to download and run this locally on your computer. Here they say that they tested it on an H100 with 80 GB, but you don't need that much. This is based off of Quant Image Edit, which requires like just a mid-to-high-end consumer GPU to run. If you're interested in reading further, I'll link to this main page in the description below. Also this week, the image generation company Midjourney is now pivoting to building spas. But all jokes aside, they just announced something completely different from their image generation platform. So they announced Midjourney Medical. It's not an image generator, but instead this is a kind of body scanner where you immerse yourself in water. So it's more like going to a spa instead of seeing a doctor. Instead of waiting for an MRI or booking appointments and only scanning your body when something feels wrong, Midjourney wants to create a system where you can casually scan your body in about just 60 seconds and track changes over time. This prototype apparently can build a much richer picture of your health. Basically, they're trying to make full-body imaging faster, cheaper, and more widespread. The scanner itself actually sounds pretty wild. So you step into this pool of warm golden light and slowly descend into water, and then it's surrounded by these underwater sensors, which use ultrasound. So it's kind of like how dolphins communicate using echolocation. So these sensors send sound waves through your body from many different angles and they listen to how these waves bounce and change. And then they use that data to reconstruct what's inside you. Think of it as like taking thousands and thousands of sound-based snapshots of your body from every direction and then using all of that data to recreate a 3D map of your body. Midjourney says the goal is for the entire scan to take no more than 60 seconds. You go into the water, you come out, and you're done. The technical side is pretty crazy. So the scanner uses a ring made of around half a million tiny elements, each about the size of a grain of sand. Then each one sends out ultrasound waves and record the returning ripples millions of times per second producing terabytes of data every second. That's pretty crazy. So, this isn't just a medical device problem. This is also a massive computing problem. The hard part is taking all of this wave pattern data and turning them into useful images. Apparently, they say it's able to do so because these waves behave differently when they pass through water, skin, fat, bone, muscle, organs. And so, the system can use these differences to reconstruct a detailed internal map of the body. They say it's similar to today's MRI scanning technology but at nearly 100 times the speed. The interesting part is how they want people to experience it. They're not building a scanner and putting it in a clinic. Instead, they're building the Midjourney Spa. This is first planned for San Francisco in 2027 with hot tubs, saunas, cold plunges, and scanning pools built into the experience. So, you go to the spa, you enjoy the space, and then you also get a body scan there. The road map is pretty ambitious. So, over the next year Midjourney says they'll refine the algorithms and hardware, run research trials, and build better prototypes, and prepare the first research spa. They also point out that regulation is a huge bottleneck here because these medical features usually require FDA approval. So, they're starting with detailed body composition maps first while submitting test results for expanded capabilities over time. And then in 2028 they want to scale to more cities and move to an even better scanner with custom silicon where they expect much better image quality and scan speed. The core idea is indeed fascinating but this is a completely different pivot compared to what they're known for, which is image generation. So, I'm not sure if they have the talent or expertise to pull this off. Let me know in the comments what you think of this. If you're interested in reading further, I'll link to this main page in the description below. And that sums up all the highlights in AI this week. Let me know in the comments what you think of all of this. Which piece of news was your favorite and which tool are you most looking forward to trying out? As always, I will be on the lookout for the top AI news and tools to share with you. So, if you enjoyed this video, remember to like, share, subscribe, and stay tuned for more content. Also, there's just so much happening in the world of AI every week. I can't possibly cover everything on my YouTube channel. So, to really stay up-to-date with all that's going on in AI, be sure to subscribe to my free weekly newsletter. The link to that will be in the description below. Thanks for watching, and I'll see you in the next one.