Anthropic Just Dropped Fable 5 And It’s Terrifying
Job gecmisi
| Job | Durum | Deneme | Worker | Istek | Baslama | Bitis |
|---|---|---|---|---|---|---|
| SummarizeYouTubeTranscript #88 | Done | 1 | learning-prod-worker-1 | 2026-06-19 02:35:51 | 2026-06-19 02:37:22 | 2026-06-19 02:37:33 |
| FetchYouTubeTranscript #83 | Done | 1 | learning-prod-worker-1 | 2026-06-19 02:32:55 | 2026-06-19 02:34:30 | 2026-06-19 02:34:50 |
Ozet
Ozet
Anthropic, yapay zeka alanında çığır açan Claude Fable 5 modelini yayınladı. Model, siber güvenlik, biyoloji ve kimya gibi yüksek riskli alanlarda çok güçlü yeteneklere sahip olduğu için, Anthropic bu yeteneklerin kötüye kullanılmasını önlemek amacıyla özel güvenlik mekanizmaları geliştirdi. Bu mekanizmalar, riskli sorular algılandığında modelin daha az güçlü bir versiyonu olan Claude Opus 4.8'e geçiş yapmasını sağlıyor. Böylece, modelin potansiyel olarak zararlı kullanımları engellenmeye çalışılıyor.
Fable 5, yazılım mühendisliği, bilgi işleme, görsel görevler ve bilimsel araştırmalarda üstün performans gösteriyor. Örneğin, Stripe şirketi, Fable 5'in aylar sürecek kod tabanı güncellemelerini günler içinde tamamladığını belirtti. Ancak bu güç, kötü niyetli kullanımlar için de risk oluşturuyor; model siber saldırılar için keşif ve saldırı süreçlerini kolaylaştırabiliyor. Anthropic, bu nedenle kapsamlı testler ve dış denetimler yaparak güvenlik önlemlerini sıkılaştırıyor. Ayrıca, biyoloji ve kimya alanlarında da modelin potansiyel kötüye kullanım riskleri nedeniyle ekstra dikkat gösteriliyor.
Ana Fikirler
- Claude Fable 5, yüksek riskli alanlarda çok güçlü ve potansiyel olarak tehlikeli bir yapay zeka modeli.
- Model, riskli sorular algılandığında daha az güçlü Claude Opus 4.8 versiyonuna geçiş yapıyor.
- Fable 5, yazılım mühendisliği, finansal analiz, görsel tanıma ve bilimsel araştırmalarda üstün performans sergiliyor.
- Siber güvenlik alanında model, siber saldırıların keşfi ve yürütülmesinde kullanılabilir, bu da ciddi riskler doğuruyor.
- Anthropic, güvenlik için ayrı AI sınıflandırıcıları ve kapsamlı dış denetimler uyguluyor.
- Biyoloji ve kimya alanlarında modelin virüs tasarımı gibi kötüye kullanım potansiyeli var.
- Distillation (model yeteneklerinin izinsiz çıkarılması) tehdidine karşı önlemler alınıyor.
- Mythos 5, Fable 5’in tam yetenekli versiyonu ve sadece güvenilir kurumlara sunuluyor.
- Modelin kullanımı için yeni veri saklama politikaları getirildi; veriler saldırı tespiti için 30 gün tutuluyor.
- Anthropic ve diğer büyük AI şirketleri halka açılma hazırlığında, AI güvenliği ve düzenlemesi gündemde.
Uygulanabilir Notlar
- Yüksek kapasiteli AI modelleri kullanılırken güvenlik mekanizmalarının ve sınırlamaların önemi vurgulanmalı.
- Siber güvenlik ve biyoloji gibi kritik alanlarda AI kullanımında etik ve yasal çerçeveler oluşturulmalı.
- AI modellerinin kötüye kullanımını önlemek için sürekli dış denetim ve kırmızı takım testleri yapılmalı.
- Kurumlar, güçlü AI modellerini kullanmadan önce risk değerlendirmesi yapmalı ve gerekli önlemleri almalı.
- Veri saklama politikaları, AI güvenliği için standart hale getirilebilir; bu konuda şeffaflık sağlanmalı.
- AI yeteneklerinin izinsiz çoğaltılması (distillation) tehdidine karşı teknik ve politik önlemler geliştirilmeli.
- AI alanındaki gelişmeler yakından takip edilmeli ve yeni çıkan modellerin kapasiteleri iyi anlaşılmalı.
Anahtar Kavramlar
- Claude Fable 5: Anthropic’in güçlü ve çok yetenekli AI modeli.
- Claude Opus 4.8: Fable 5’in güvenlik nedeniyle devreye giren daha az güçlü versiyonu.
- Siber güvenlik riskleri: AI’nin siber saldırı keşfi ve yürütme kapasitesi.
- Distillation: AI model yeteneklerinin izinsiz çıkarılması ve çoğaltılması.
- Agentic hacking: AI’nin otomatik ve çok aşamalı siber saldırı gerçekleştirmesi.
- Red teaming: Güvenlik açıklarını bulmak için yapılan dış denetim ve saldırı simülasyonları.
- Veri saklama politikası: Kullanıcı verilerinin güvenlik amacıyla belirli süre tutulması.
- Recursive self-improvement: AI’nin kendi kendini geliştirme yeteneği.
- Mythos 5: Fable 5’in tam yetenekli ve kısıtlamasız versiyonu, sadece seçilmiş kurumlara açık.
Transcript
All right. So, what Anthropic just did today is kind of unprecedented. They just released Claude Fable 5, and this thing is so powerful that they're actually scared to give it to you at full strength. So, basically, they built an AI model that's so capable at hacking, at biology research, at finding vulnerabilities in code, that they had to create an entirely separate system just to stop it from answering certain questions. They're literally censoring their own AI, not because of content policies or whatever, but because they genuinely believe this thing could be weaponized. And the crazy part is, they're still releasing it to the public anyway, just with a safety net that kicks in when things get too dangerous. But let me back up a second and give you some context here. Back in April, Anthropic released something called Claude Mythos preview, and they were super cautious about it. They only gave it to a handful of partners because they were concerned about cybersecurity risks. We're talking about major organizations that manage critical infrastructure, [music] the kind of stuff where if something goes wrong, it's not just a minor inconvenience. Then last week, [music] they expanded access to hundreds of organizations across 15 countries. But again, very controlled, very selective. Now though, they're bringing a version of that technology to everyone through their Claude API and enterprise plans. But here's where it gets interesting. They're calling it Fable 5 instead of Mythos [music] because it's not quite the same thing. It's the same underlying model, yeah, but with these safety mechanisms built in that fundamentally change how it operates. The name actually comes from the Latin word fabula, which means that which is told, and it's connected to the Greek word mythos. So, there's this whole linguistic connection between [music] the two, but the safeguards are what really separate them. So, what exactly are these safeguards? Well, when Fable 5 detects that you're asking about high-risk [music] stuff like cybersecurity, biology, chemistry, or something called [music] distillation, it just straight-up refuses to answer with its full capabilities. [music] Instead, it falls back to Claude Opus 4.8, which is still a really capable model, but nowhere near as powerful as Fable. The thing is, Anthropic says this only happens in less than 5% of sessions based on their early data. So, at least 95% of the time you're getting the full Fable experience. But still, that 5% is there for a reason. Now, let's talk about what makes Fable 5 so powerful that they felt the need to do [music] all this. According to Anthropic, this thing is state-of-the-art on basically every benchmark they tested. Software engineering, [music] knowledge work, vision tasks, scientific research, you name it, Fable 5 excels at it. And the longer and more complex the task, [music] the bigger the gap between Fable and their other models. Stripe, the payment company, said that Fable 5 compressed months of engineering work into just days. They had this massive 50 million-line Ruby codebase, and Fable did a codebase-wide migration in a single day that would have taken an entire team over 2 months to do manually. That's absolutely insane when you think about it. For knowledge work, Fable scored the highest of any model on something called Hebbius Finance Benchmark, which tests senior-level reasoning. >> [music] >> IMC, a trading firm, said it aced their trading analysis evaluations across the [music] board, factual lookup, conceptual reasoning, root cause analysis, expected value analysis, all of it. And in third-party testing, an analytics company called Hex, said Fable was the first model to hit 90% on their core analytics benchmark of complex, long-running analytical tasks. >> [music] >> They specifically noted that on the hardest questions, it shows strong judgment and attention to nuance. The vision capabilities are particularly impressive. >> [music] >> Fable can extract precise numbers from detailed scientific figures and even rebuild a web app source code from screenshots alone. They showed this demo where Fable played Pokémon Fire Red from start to finish [music] using only raw game screenshots. No maps, no navigation aids, nothing. Earlier Claude models needed this whole complex helper system to play Pokémon, but Fable [music] just did it with vision alone. They also had it build a simulation of the solar system where it derived the planet's orbital motion from physics first principles and use that to predict solar eclipses. It's genuinely pretty wild stuff. But, here's where things get a bit concerning and why Anthropic is being so careful. >> [music] >> Mythos class models have apparently reached a threshold where they present significant risks. In cybersecurity specifically, these models excel at discovering and exploiting software vulnerabilities. They can make cyber attacks substantially easier and cheaper to commit. And it's not just about finding exploits. They can perform multiple different parts of a cyber attack including reconnaissance, discovery, lateral movement, the whole nine yards. That's what they call agentic hacking [music] and it's genuinely dangerous in the wrong hands. So, Anthropic built these classifiers, which are basically separate AI systems that detect potential misuse including jailbreak attempts and prevent the main model from responding. When the classifiers catch something related to cybersecurity, biology, and chemistry, or distillation, boom, you get routed to Opus 4.8 instead. And they're pretty upfront about it. They tell you when this happens. They tested these safeguards pretty extensively, too. They ran an external bug bounty that produced no universal jailbreaks in over 1,000 hours of testing. They worked [music] with external red teaming organizations which also failed to find universal jailbreaks, though the UK AIS apparently made some progress toward one within a brief initial testing window. Now, Anthropic [music] admits it's probably impossible to completely prevent universal jailbreaks. But their goal is to make any remaining jailbreaks sufficiently slow and costly that they can detect [music] and prevent them before they're used at scale. The biology and chemistry stuff is interesting because Anthropic is being extra cautious here. They tested Mythos 5's ability to complete a challenging step in designing adeno-associated viruses, which are used in gene therapy. And here's the thing. These models outperformed specialized protein language models using their biological reasoning alone, even though they weren't explicitly trained for this task. That's both promising for therapeutic development and kind of scary because the same capability could enable the design of dangerous viruses in the wrong hands. There's also this whole distillation thing, which is basically when someone tries to extract [music] Claude's capabilities to train competing models. Apparently, Anthropic has identified large-scale attempts to do this from authoritarian countries. If distillation of Fable 5's abilities were to succeed, it could lead to the proliferation of near-frontier AI capabilities without the appropriate safeguards, which is obviously problematic. Now, for the people who actually need the full unrestricted power of these models, Anthropic is launching Claude Mythos 5. It's the exact same model as Fable 5, but with the safeguards lifted in some areas. Initially, this is going to organizations that were already approved for Mythos preview through something called Project Glasswing, which is being done in collaboration with the US government. Anthropic says Mythos 5 has the strongest cybersecurity capabilities of any model in the world. Soon, they're planning to expand access through a broader trusted access program for both cybersecurity organizations and biology researchers. The pricing for both models is $10 per million input tokens and $50 per million output tokens, which is double the price of Opus 4.8. That's actually less than half the price of Claude Mythos preview, though. And honestly, the price alone might serve as a deterrent for widespread use, but some companies think it's worth it. Rakuten, a shopping rewards platform, said that at the highest effort level, Fable reflects on and validates its own work, and that's what makes highly autonomous operations possible for them. The extra thinking pays for itself. There's also this new data retention policy that's pretty significant. For Fable 5, Mythos 5, and future models with similar or higher capability levels, Anthropic is requiring 30-day [music] retention for all traffic, even for enterprises that previously had zero retention agreements. They say they won't use this data for training and will only use it to defend against complex and novel attacks, >> [music] >> including new jail breaks, and to identify and reduce false positives. But this could set an industry precedent where access to increasingly powerful models comes with mandatory data retention policies framed as a safety measure. The timing of all this is pretty interesting, too. Anthropic recently announced they've confidentially filed a draft S1, which basically means they're preparing for an IPO. OpenAI announced the same thing on Monday, and SpaceX, [music] which includes Elon Musk's xAI, is set to go public on Friday. So, all the major AI companies are making these big moves at the same time. Anthropic's initial announcement of Mythos back in April apparently spooked financial markets and governments worldwide, raising concerns that AI models had advanced to the point where they could uncover major vulnerabilities in software and cybersecurity. President Trump even signed an executive [music] order that allows AI companies to voluntarily give the federal government access [music] to frontier AI models up to 30 days before their release, though it explicitly prohibits the government from imposing mandatory review. This whole release also comes after Anthropic urged major global AI labs to establish what they called a coordinated brake pedal on frontier AI development. They warned that systems are advancing so rapidly that they may soon achieve recursive self-improvement, which is basically when AI can autonomously improve itself [music] without human intervention. That's a pretty concerning prospect if you think about it. For subscription plans, Anthropic is rolling things out in stages. From today through June 22, Fable 5 is included in Pro, Max, Team, and seat-based enterprise plans at no extra cost. But on June 23, they're pulling Fable 5 from those plans, and using it after that will require usage credits. They say they plan to restore it as a standard subscription feature as soon as possible when capacity allows. So, yeah. This is a pretty massive development in the AI space. We're seeing these companies grapple with models that are becoming so powerful that they genuinely pose risks if used maliciously, and they're trying to find ways to release them safely while still making the technology available. Whether these safeguards will hold up long term remains to be seen, but it's definitely a fascinating time to be watching this space. Also, if you want more content around science, space, and advanced tech, we've launched a separate channel for that. Links in the description. Go check it out. So, what do you think? Is Anthropic being responsible here, or >> [music] >> are we entering a stage where AI models are getting too powerful to release normally? Let me know in the comments. Subscribe for more AI and tech updates. Like the video if you found it useful, and thanks for watching. I'll catch you in the next one.