AI Search
6d__WOpZswY
New #1 open-source AI model is here!
Job gecmisi
| Job | Durum | Deneme | Worker | Istek | Baslama | Bitis |
|---|---|---|---|---|---|---|
| FetchYouTubeTranscript #71 | Done | 1 | learning-prod-worker-1 | 2026-06-18 23:34:33 | 2026-06-18 23:34:36 | 2026-06-18 23:34:48 |
Ozet
Bu video icin henuz ozet yok.
Transcript
Video metni
We have a new number one open-source model and the gap is insane. So, my favorite lab ZAI just released GLM 5.2 and get this, it's by far the best open model out there. Plus, it even beats the best GPT and the best Gemini across multiple benchmarks. That's how ridiculously capable this model is. In this video, we're going to go over all the incredible things that it can do. Plus, we're going to go over its specs, performance, and benchmarks against other competitors, as well as how and where to use it. Let's jump right in. Now, there are various places where you can use this new GLM 5.2. You can try it for free using their online chat interface, which I'll link to in the description below. It's just chat.z.ai. And then at the top here is where you should be able to select the latest GLM 5.2. However, for all these top Frontier models, they can already do simple stuff like drafting emails. summarizing things, even helping you write an essay. This is way too easy. In fact, if you use GLM 5.2 in the chat interface, it's not really doing it justice. To really unleash its full potential, you should try to use these frontier models with a gentic frameworks or harnesses like OpenClaw or Hermes or Claude Code or ZAI also has their own agentic framework called Zcode which you can download for free. It supports Mac, Windows, and Linux. And after you download that, it should look something like this. It's pretty much a Codeex clone actually. So, very similar to Codeex. And what I like about it is that you can set up multiple projects with each project being in a folder and it can work on multiple files within a project like this. All right, for the first test, let's make it super hard already. I'm going to get it to build a fully interactive 3D digital twin of Earth that has the following features. Allow users to zoom seamlessly from outer space down to individual city streets. When I hover over a city, the country's outline should be highlighted and there should be a pop-up displaying stats like area, population, GDP, show a realistic planet Earth with toggles for atmospheric cloud cover, flight traffic, day and night, and then night mode to show city streets. Make sure it loads efficiently on a regular web browser. I'm going to use GLM 5.2 and then set the thinking mode to max. Let's press run. All right, so it gave me an initial prototype, but there were some minor things that were not working. For example, the cloud cover and the country borders didn't work. So, I just prompted it further. And then it worked for another 3 minutes and 47 seconds. And then also the flight traffic looked pretty ugly. So, I prompted it to animate the flights with path streaks. It worked for another 3 minutes. And that's pretty much it. So, afterwards, here is our planet Earth. First of all, it includes various toggles. So, let's go over these. For satellite imagery, it's basically just the image of Earth. So, let's turn this back on. for cloud cover. It tries to overlay some cloud cover layer on the earth like this, but it doesn't really work well. So, this is still broken. Let me turn that back off. And then you can see that country borders work. And then if I hover over each country, it does show me the stats like area, population, and GDP. And then next, let's turn on flight traffic. And this looks beautiful. Very nice. So, let's turn off flight traffic. And then day and night terminator looks like this. So half of the earth is daytime and the other half is nighttime. So that works. Next, here are night city lights. So that also works. Very nice. Let me turn that off. Next, let's try zooming into individual cities. So let's try New York. And indeed, it automatically flies me to New York. And it zooms all the way into city street view. You can even zoom all the way in to see people in cars. And then next, let's try Tokyo. So, it's going to zoom out and then zoom into Tokyo. And here's what we get. So, you know, most features just work. Now, it did require some additional prompting. So, it took like around 15 minutes for it to complete the entire thing. Whereas the best of the best model out there, Cloud Fable 5, which you can't even use right now, that was able to get it in pretty much just one prompt. That being said, for an open source model, this is still extremely impressive. Next, let's try something even trickier and probably more relevant to everyday workflows. I'm going to get it to work with a ton of tools at once to create a promo video. So, the prompt is from this product page, which I just put the wise landing page, which looks like this. Let me just scroll down a bit so you can see what this looks like. Create a promo video about this product, but also you need to include a voice over using Gemini TTS. And then you also need to use this hyperframes framework to create the animation. And I just simply linked to the GitHub repo of hyperframes. This is an open-source package that allows you to create animations using code. So, here's what the GitHub repo looks like. It needs to go to this web page and figure out how to install this itself. I didn't give it any explicit instructions on how to do so. You could also feed it another tool like Remotion if you prefer. And then it should be 16 to9 around a minute long. Use this track as the background audio. It's just some random background music. Reduce the volume so the voiceover can be heard clearly. And then here's an example of how to use Gemini TTS. So what I did was I just went to AI Studio from Google and then clicked on speech and music and then Gemini flash DTS and then I just clicked on the add voice over. I selected a voice down here. So let's go with this voice. >> What's a skill you'd like to develop? >> And then afterwards I just put enter transcript here. And then at the top I clicked on get code. I copied the code and then just pasted it into ZI. And then I also gave it my API key at the bottom here. And that's pretty much it. And afterwards, it gave me the final deliverable. That's it. I didn't need to prompt it any further. So, here's the video. Sending money across borders shouldn't cost you a fortune. Meet Wise, the international account built for a [music] world without borders. Hold and convert over 40 currencies at the real mid-market exchange rate. No hidden markups, no surprise fees, just money that [music] moves the way it should. Send to family, pay suppliers, or get paid like a local with your own account details for up to nine currencies for up to nine currencies. Spend anywhere with the [music] Wise debit card and keep every cent you save. From the everyday to the enterprise, over 16 million people and businesses already trust Wise with their [music] money. One account, 40 plus currencies, zero nonsense. Join them at wise.com [music] and start sending money the smarter way. yours for less. So, you can see how capable this model is. The thing I like about GLM is it requires very minimal handholding. It just works right out of the box most of the time and it has very few errors. By the way, let me also hover over the stats. So, you can see the tokens used and the cash hit rate for just this one prompt. So, it used roughly 100k tokens and it took around 20 minutes to complete. All right. Next, let's test how good it is at generating 3D models. So, I'm going to input this. Create a beautiful 3D animated model of a V8 engine. Include a slider to show a smooth transition from the fully assembled engine into an exploded view which shows the inner workings such as pistons, spark plugs, valves, connecting rods, and central crankshaft in motion. Use a single HTML file. All right, so it worked for 13 minutes and in just one prompt, here's what we got. Let's open this index.html file. And here is our V8 engine. Very nice. Let's explode this. And as you can see, the explosion works. It's very nicely animated. We can even like increase the engine speed to make this even faster or decrease it. Really cool. So, it's very capable at creating 3D models. Let me also show you the tokens used. So, this was really quick and lightweight. It only used like 28,000 tokens plus the average cash hit rate is 96%. Or here's an even trickier example which Claude Fable could not even get correct. Create a beautiful 3D animated model of a traditional mechanical watch, including inner workings such as the dials, hands, hour markers, etc. Show the watch mechanism in motion with the gears turning, balance wheel oscillating, etc., etc. Let's see if it can handle this. All right, so it proceeded to code this up. Now, creating a watch is a lot trickier. So, it did hit an error. I just pasted the error back into here and then it fixed the error. And then I just wanted to make it even more appealing. So I wrote explode it further and allow me to turn on or off layers. Make it even more visually impressive. And here is our result. Let's explode this. So you can see the individual components separated. Everything looks decent and the clock is moving. The minute, hour, and second hands are all correct. Plus the time scale also looks correct. Now, let's turn on or off certain layers. So, we have the sapphire crystal. Let's turn that off. Hands. So, that works. Let's try the dials and markers. That works. Let me get rid of this bulky metal case over here. So, you can see the inner components. All right. So, the components look like this. However, if I zoom into the gears, here's where we start to see some errors. So, you can see that I guess some of these dials are moving. The others are not really moving. Plus, they're not really aligned. So, this is an incredibly challenging prompt. And note that even Claude Fable 5 was not able to get this correct. It was not able to get the dials to align completely. Nevertheless, still a very impressive generation. I did not expect it to get it 100% correct. Note that here it used around 80,000 tokens with an average cash hit rate of around 94%. Now, you don't have to use GLM 5.2 within Zcode. You can also use another harness like Cloud Code or Open Claw or Hermes. So, let me also show you a quick example of that. A lot of people don't know this, but you can actually use Claude code or even codeex with models other than Claude or GPT. I'll link to this page in the description below in case you want to link this to Claude code. So, what you need to do is first of all sign up for a plan and then use the API key and then install this coding helper. So, I'm just going to copy this line and then on my computer I'm going to open command prompt and then paste this in. And then afterwards, it gives you a ton of options. For example, for API key, you'll need to input an API key if this is your first time using it. So, let me paste this in here and then press enter. Then afterwards, you would select the agentic platform you want to configure. So, yes, let's go with cloud code. And then you just need to select the first option to update cloud code to use GLM instead. So, let's press enter. And that's pretty much it. So, you can either start cloud code or I would recommend exiting the terminal and starting from scratch. Now there's one more step you need to do which is if you click on how to switch models down here you also need to explicitly state like which GLM model to use to replace haiku sonnet or opus and you would enter this in your claude settings.json file which is usually located in your user folder. So for example for me it's located over here. So let me open this in vs code or you can also open this in notepad or whatever. And then we just need to copy and paste in these four lines into this environment field. So, let me just paste it over here and fix the formatting and then press Ctrl S to save this. That's pretty much it. So, afterwards, if I open up a terminal and then use cloud code, you can see that now I'm using GLM 5.2 as the model. Or if I type in /model, you can see that these GLM models have now replaced the default cloud models. So, that's how you can use GLM in cloud code. If you prefer cloud code instead of, you know, Zcode or some other agentic framework, let's try a quick example with cloud code. So, I'm going to write develop a ray tracing simulation featuring one sphere, one cube, and one pyramid. It should be a blue sky with a checkered ground. Add adjustable parameters like position, reflectivity, roughness, transparency, and other material properties for each shape. The key here is do not use 3JS or other external libraries. It needs to code this up completely from scratch. So, this tests its understanding of physics and lighting. AI agents can help you do a ton of work. But did you know with the right tools, you can even use them for content creation? That's where Higfield, the sponsor of this video, comes in. With Higsfield MCP, you can turn Claude, Codeex, OpenClaw, Hermes, or other agent systems into a full creative production studio. You give it one prompt and it can help turn that idea into images, videos, ads, marketing concepts, and full creative workflows. What makes this so powerful is that great AI content usually comes down to great prompting. And Claude or other agents are extremely good at understanding what you mean, refining your idea, and directing the right creative steps. So instead of manually trying to write the perfect prompt, Claude can act like your creative director, while Higsfield handles the actual generation. For example, with Higsfield MCP, you can connect Higsfield directly inside Claude. That means Claude can plan the concept, write the creative brief, and generate videos through Higsfield and even place the final assets into your working directory. No switching tabs, no creating prompts back and forth, and no separate creative handoff. You can use it for things like making launch videos, creating UGC style ads, rebuilding viral formats in your own brand, or even building full marketing campaigns from one thread. And if you want everything in one place, Higsfield supercomputer gives you a full AI creative team inside a single chat. You can choose Claude or another agent as the brain and then let it work with Higsfield's creative models like Seed Dance 2.0, 0 Cinema Studio and Marketing Studio to go from idea to finished deliverable. Whether you're making ads, launch videos, or social media content, Higsfield is a gamecher that will supercharge your production workflow. Try today using the link in the description below. All right, so afterwards, it proceeds to code this up and it even spun it up in a local server and then took a screenshot to verify that everything works. Now, this doesn't have vision capabilities by default. So, it's just using another tool to verify that the interface is working. For example, that the sky is blue and there's a checkered ground. And afterwards, that was it. That took around 20 minutes. But let's pull this up. And indeed, we have a very beautiful ray tracing simulation of these three shapes built from scratch without 3JS or any external libraries. Really impressive. In fact, this loads even better than what I got from Claude Fable, which had some very weird latency issues. Next, let's play around with all these settings. So, the position of the sphere works. The radius also works. And then, let's play with the reflectivity. So, decreasing this would make it non-reflective. Increasing this would make it super reflective, as you can see here. Let's also play with the roughness. So, the roughness slider works. Let's drag this all the way to zero. And you can clearly see the other shapes reflected in the sphere. Actually, let me set the color to white. And then let's increase the transparency while also decreasing the reflectivity to make this fully transparent. And then I can play with the IU settings like this. So everything just works. Very cool. Let's also play with emissions. That just turns it brighter like a light source. Interesting. Let's also play with the cube settings. The position sliders work. The size also works. Let's also change the color. Let's make this super reflective and zero roughness. And then moving on to the pyramid. Let's also adjust the settings of this so the pyramid also works. Let's change the color to red. And then let's make this incredibly opaque. So I'm going to drag the reflectivity to zero, roughness all the way to one, and then transparency to zero. And then let me drag this closer to the cube. So if I view it like over here, you can see all the physical properties of these three shapes are accurate. So the sphere is completely transparent. So I should not see any reflections in the sphere. The pyramid is completely opaque. The cube is completely reflective. So I should see the reflection of the sphere and the pyramid within the cube. So everything is correct. Really cool. There's also additional settings here like adjusting the sun's setting. Very cool. and the sun's elevation as well as the intensity and the ambience. I can even adjust the light color. So, let's set this to something like red. And indeed, it turns the light to red. And then I can even change the sky color. So, let's change this to something like green. And that works. I can change the checkered background. So, let's set this to something like red. I can adjust the checkered scale as well as the reflectivity of the ground and even the camera distance and camera height. So, in just one prompt, zero shot, it was able to code this ray tracing simulation completely from scratch without any external libraries. And I actually like this a lot better than what I got from Cloud Fable 5. Again, the thing I like about GLM 5.2 or even the previous GLM models is that everything just tends to work right out of the box. You have to do very minimal handholding. It contains very few errors, and if it does, you can often just fix it with one or two follow-up prompts. All right, next, let's test how good it is at music composition. So, first I'm going to get it to make a music interface. I'm going to get it to create a JW with these instruments. Piano, synth, pluck, strings, drums, and bass. For each instrument, there should be a piano roll interface where I can drag and drop notes on the timeline. Each track should have pan, volume, and other standard settings. Add play, pause, and other settings, etc. All right, so it was able to code this up no problem. I didn't even need to prompt it further. The previous model could also do this, so it's nothing impressive. What I really wanted to test is how good it is at composing music. So next I put by default show a powerful expressive 32 bar song rich in complexity that would win a Grammy. Include effects, automation, proper panning for each track. Make sure everything is mastered well. So it worked for 8 minutes and 23 seconds. It composed the song with the full arc including intro, verse, preorus, chorus, and then afterwards there were some alignment issues. So, I wrote some notes on the piano roll are above the view. Add vertical scroll so I can scroll higher. And that's pretty much it. So, it worked for an additional 34 seconds. So, let's open this up and play it. All right. Here's our 32 bar Grammy awardwinning song. [music] [music] Heat. Heat. N. [music] Heat. Heat. [music] Heat. Heat. [music] [music] Heat. Heat. N. [music] >> [music] [music] [music] >> So, it's not bad. It definitely won't win a Grammy. It's still lacking in complexity. It did add some cool effects like reverb and delay. However, it didn't really auto pan the tracks like I prompted it to. It also did not add automation across the tracks. So, quite basic, but nevertheless, still very impressive. This actually sounds very similar to the generation I got from Claude Fable 5. In fact, if you want to hear the generation from Fable, see this video. All right. Next, let's also see if it can create some complex mathematical animations using a package called Manom. So my prompt is create a beautiful manom animation where for your circles draw a butterfly use rotating circles connected by vector arms with the end points tracing the butterflyy's wings body and antenna etc etc. Save the output as MP4. The challenge here is I don't even have manom installed on here. So it needs to also figure out how to install the package man first in order to create the animation. All right. So first it needs to set up the environment and download all the required packages like manom, numpy, sci etc. These are all packages that are required to make the video and then afterwards it proceeded to create the animation. Plus it also used its built-in image analysis tool to verify that the frames actually consist of a butterfly. It also was able to self-verify and catch some bugs along the way. Now I didn't like the look of the butterfly too much, so I wrote make the butterfly more detailed with sharper top wings. Make sure everything is accurate. Also, for animations, make sure the circles and butterfly outline are visible. Make the circle outlines brighter because they were too dull before. So, it worked for another 22 minutes. Keep in mind, a portion of this time is spent like just rendering the video. Here is what we got. So, indeed, this looks like a butterfly kind of. And here, it's using 12 circles. So, it's a very basic butterfly shape. Here, it's adding even more circles to make it even more intricate. And then here, it's adding even more details. And finally, here even more details. And you know, this animation does look correct. So overall, very impressive. All right, those were some of my tests using Zcode. Next, let's also test it on some regular stuff using just the online chat interface. You can try this out for free. And at the top here is where you can select the latest GLM 5.2. Now GLM does not have vision capabilities unlike their competitor Miniax M3, which is multimodal. So, it can still use external tools to analyze images, but it doesn't have like native image understanding. I'm going to try it on my classic test. So, I'm going to feed it this image. And there's a frog hidden somewhere in this image. I don't expect it to get it correct because it doesn't even have vision capabilities, but let's see how it does. So, I'm going to upload this frog image over here and then ask it to find the frog hidden in this photo. I'm going to set it to max and deep think. All right. Now, it doesn't have vision capabilities. So, it's actually using an external vision language model to analyze this image and then attempt to find the frog. And after like 10 minutes or so, here is the answer I got, which unfortunately is wrong. So, the frog is not here. All right. Next, let's also test it on its deep research capabilities. So, here the prompt is describe the molecular drivers of this type of leukemia. Discuss all these different things. include relevant tables and visualizations. I'm going to set it to deep think and max and also I'm going to enable advanced search for more in-depth research. Let's press generate. All right, here is what I got. So, here's the intro paragraph. Here are the different sections. It gives me a nice table. And then we have constitutive kaise activity and downstream ancogenic signaling. It even gives me a nice flowchart. And then here, number two, evolution of targeted therapy. It was able to code up this really nice timeline plus a very thorough table of all these different drugs including their features, toxicities, etc. The next section is key trial endpoints at long-term follow-up. The thing I like about GLM is it's no BS. It's very short and to the point. And then afterwards, here are some resistance mechanisms. It's even able to go up this really nice web chart. And then here is a table on comparative resistance profile by mechanism class. It's able to go super in-depth while being very concise. And then next section we have survival outcomes. It has no problem coding up very nice tables and charts. And finally, we have clinical implications. So definitely a very performant model for deep research. All right. So those are some of my personal demos and where and how to use it. Again, there are various places you can try for free on their online chat platform or you can also use their harness called Zcode. So you can work on multiple projects and multiple files on your computer. Or you can also connect this to other frameworks like Claude Code or OpenClaw or Hermes. Next, let's go over the specs of this. First of all, I love that this new GLM 5.2 now supports a 1 million token context window. This is like how much stuff you can put into your prompt at once. And a million tokens is like over 700,000 words or a small to medium-sized codebase. So this can jam-pack a ton of information. What I also love about this is this has an MIT open-source license which is very permissive. If you look at these benchmarks here is where its performance is pretty damn insane. So across these long autonomous coding task benchmarks like Frontier Suite or Post Train Bench, you can see that GLM 5.2 2 not only beats Open AI's best model GPT 5.5 as well as Google's best model Gemini 3.1 Pro by a huge margin and it's edging very close to Claude's best model Opus 4.8. That's pretty crazy considering that this is an open model. Here are some other benchmarks for your reference. So SWEBench Pro again this even beats the best version of GPT and Gemini and is edging very close to the best Claude. Same with Terminal Bench. And then here's the new kid in town called Deep Suite. They claim this is a more accurate measure of software engineering capabilities. And as you can see, GLM 5.2 scores incredibly high at 46.2. In fact, if you look at this official leaderboard by Deep Suite, this is by far the highest scoring open model out there. It's not even close. This benchmark called humanity's last exam is also pretty noteworthy. So, this tests an AI model's knowledge on some really obscure scientific knowledge. And again, it's even more knowledgeable than the best GBT and the best Gemini. That's how crazy this model is. The awesome thing is they actually revealed how they built this. So, they implemented some insane architecture tweaks to make this work. First of all, this includes a design called index share, which basically reuses the same indexer across its sparse attention layers. And this effectively reduces the compute by 2.9 times. This is a bit technical, but they also have this improved MTP layer which increases its decoding length by up to 20% which also helps with generation efficiency. Now, at the time of this recording, some of the main independent leaderboards and benchmarks have still not added GLM 5.2 yet, but here are some that have. If you look at this leaderboard called Frontier, you can see that GPT 5.2 is even Opus 4.8 level across most of their leaderboards. It even beats GPT 5.5, which is insanely impressive. It's also great at front end and design. Here you can see that GLM 5.2 even beats Opus 4.8 in front-end coding. Isn't that crazy? Note that Claude Fable 5, which is number one, is currently banned, so no one can even use it. Here's another leaderboard from Design Arena. And get this, GLM 5.2 even beats Claude Fable 5, as well as the rest of the other closed models in terms of front end and design. Here's another fun benchmark called Runescape Bench, which tests how good an AI model is at playing Runescape. And here you can see it's by far the best open model out there. The rest of the open models are all the way down here. The awesome thing is they've released the models already. They are fully open weights under the MIT license. So if you click on this hugging face link, note that it's actually relatively small compared to other Frontier models which are over a trillion parameters like Deepseek V4. So this one is only 753 billion parameters. That being said, this is still pretty damn huge at 1.51 terabytes in size. So, I mean, good luck running this at home unless you have a freaking data center in your basement. So, it's not really meant for consumer devices. But here's the value of opensource AI models. If you've been following my channel, you should know that the other closed labs like Anthropic constantly gatekeep their models. They intentionally make it dumber for certain use cases. Plus also the US government just outright banned the latest clawed model so people can't even use it. But with open models like GLM or DeepSeek or Quinn, you don't need to be dependent on these other closed source labs. The importance of open source is that it brings the power of intelligence back to the people. You could potentially host this yourself. This gives you ultimate sovereignty. Another value of open source is privacy. So for closed source models, you're essentially sending data to their servers. If your messages contain some sensitive or confidential information, they potentially have access to it. Whereas, if you host an open- source model locally and on prem, your data can stay where you control it. This is very important for things like legal documents, medical data, or government records, financial records, etc. And of course, because the model is completely open weights, people can fine-tune it or make it better. You know, the community builds on top of this. That's why I'm such a big fan of open source. That's what makes open- source so much more attractive and valuable compared to these other closed AI labs which never share anything. Anyway, that sums up my review of GLM 5.2. ZAI has always been my favorite AI lab, and once again, they were released an incredibly performant model that even beats the best GPT and the best Gemini on some benchmarks, and it's completely open. Let me know in the comments what you think of this. As always, I will be on the lookout for the top AI news and tools to share with you. So, if you enjoyed this video, remember to like, share, subscribe, and stay tuned for more content. Also, there's just so much happening in the world of AI every week. I can't possibly cover everything on my YouTube channel. So, to really stay uptodate with all that's going on in AI, be sure to subscribe to my free weekly newsletter. The link to that will be in the description below. Thanks for watching and I'll see you in the next one.