AI Revolution 8TjCwdnZSp8 read

Anthropic Just Dropped Fable 5 And It’s Terrifying

Transcript: Done Yayin: 2026-06-09 18:40 YouTube
Anthropic Just Dropped Fable 5 And It’s Terrifying
Kanala don
Job gecmisi
Job Durum Deneme Worker Istek Baslama Bitis
SummarizeYouTubeTranscript #88 Done 1 learning-prod-worker-1 2026-06-19 02:35:51 2026-06-19 02:37:22 2026-06-19 02:37:33
FetchYouTubeTranscript #83 Done 1 learning-prod-worker-1 2026-06-19 02:32:55 2026-06-19 02:34:30 2026-06-19 02:34:50

Ozet

openai/gpt-4.1-mini-2025-04-14 - 2026-06-19 02:37
Indir

Ozet

Anthropic, yapay zeka alanında çığır açan Claude Fable 5 modelini yayınladı. Model, siber güvenlik, biyoloji ve kimya gibi yüksek riskli alanlarda çok güçlü yeteneklere sahip olduğu için, Anthropic bu yeteneklerin kötüye kullanılmasını önlemek amacıyla özel güvenlik mekanizmaları geliştirdi. Bu mekanizmalar, riskli sorular algılandığında modelin daha az güçlü bir versiyonu olan Claude Opus 4.8'e geçiş yapmasını sağlıyor. Böylece, modelin potansiyel olarak zararlı kullanımları engellenmeye çalışılıyor.

Fable 5, yazılım mühendisliği, bilgi işleme, görsel görevler ve bilimsel araştırmalarda üstün performans gösteriyor. Örneğin, Stripe şirketi, Fable 5'in aylar sürecek kod tabanı güncellemelerini günler içinde tamamladığını belirtti. Ancak bu güç, kötü niyetli kullanımlar için de risk oluşturuyor; model siber saldırılar için keşif ve saldırı süreçlerini kolaylaştırabiliyor. Anthropic, bu nedenle kapsamlı testler ve dış denetimler yaparak güvenlik önlemlerini sıkılaştırıyor. Ayrıca, biyoloji ve kimya alanlarında da modelin potansiyel kötüye kullanım riskleri nedeniyle ekstra dikkat gösteriliyor.

Ana Fikirler

  • Claude Fable 5, yüksek riskli alanlarda çok güçlü ve potansiyel olarak tehlikeli bir yapay zeka modeli.
  • Model, riskli sorular algılandığında daha az güçlü Claude Opus 4.8 versiyonuna geçiş yapıyor.
  • Fable 5, yazılım mühendisliği, finansal analiz, görsel tanıma ve bilimsel araştırmalarda üstün performans sergiliyor.
  • Siber güvenlik alanında model, siber saldırıların keşfi ve yürütülmesinde kullanılabilir, bu da ciddi riskler doğuruyor.
  • Anthropic, güvenlik için ayrı AI sınıflandırıcıları ve kapsamlı dış denetimler uyguluyor.
  • Biyoloji ve kimya alanlarında modelin virüs tasarımı gibi kötüye kullanım potansiyeli var.
  • Distillation (model yeteneklerinin izinsiz çıkarılması) tehdidine karşı önlemler alınıyor.
  • Mythos 5, Fable 5’in tam yetenekli versiyonu ve sadece güvenilir kurumlara sunuluyor.
  • Modelin kullanımı için yeni veri saklama politikaları getirildi; veriler saldırı tespiti için 30 gün tutuluyor.
  • Anthropic ve diğer büyük AI şirketleri halka açılma hazırlığında, AI güvenliği ve düzenlemesi gündemde.

Uygulanabilir Notlar

  • Yüksek kapasiteli AI modelleri kullanılırken güvenlik mekanizmalarının ve sınırlamaların önemi vurgulanmalı.
  • Siber güvenlik ve biyoloji gibi kritik alanlarda AI kullanımında etik ve yasal çerçeveler oluşturulmalı.
  • AI modellerinin kötüye kullanımını önlemek için sürekli dış denetim ve kırmızı takım testleri yapılmalı.
  • Kurumlar, güçlü AI modellerini kullanmadan önce risk değerlendirmesi yapmalı ve gerekli önlemleri almalı.
  • Veri saklama politikaları, AI güvenliği için standart hale getirilebilir; bu konuda şeffaflık sağlanmalı.
  • AI yeteneklerinin izinsiz çoğaltılması (distillation) tehdidine karşı teknik ve politik önlemler geliştirilmeli.
  • AI alanındaki gelişmeler yakından takip edilmeli ve yeni çıkan modellerin kapasiteleri iyi anlaşılmalı.

Anahtar Kavramlar

  • Claude Fable 5: Anthropic’in güçlü ve çok yetenekli AI modeli.
  • Claude Opus 4.8: Fable 5’in güvenlik nedeniyle devreye giren daha az güçlü versiyonu.
  • Siber güvenlik riskleri: AI’nin siber saldırı keşfi ve yürütme kapasitesi.
  • Distillation: AI model yeteneklerinin izinsiz çıkarılması ve çoğaltılması.
  • Agentic hacking: AI’nin otomatik ve çok aşamalı siber saldırı gerçekleştirmesi.
  • Red teaming: Güvenlik açıklarını bulmak için yapılan dış denetim ve saldırı simülasyonları.
  • Veri saklama politikası: Kullanıcı verilerinin güvenlik amacıyla belirli süre tutulması.
  • Recursive self-improvement: AI’nin kendi kendini geliştirme yeteneği.
  • Mythos 5: Fable 5’in tam yetenekli ve kısıtlamasız versiyonu, sadece seçilmiş kurumlara açık.

Transcript

Video metni
en markdown 2026-06-19 02:34 youtube-transcript-api:generated
Indir
All right. So, what Anthropic just did
today is kind of unprecedented. They
just released Claude Fable 5, and this
thing is so powerful that they're
actually scared to give it to you at
full strength. So, basically, they built
an AI model that's so capable at
hacking, at biology research, at finding
vulnerabilities in code, that they had
to create an entirely separate system
just to stop it from answering certain
questions. They're literally censoring
their own AI, not because of content
policies or whatever, but because they
genuinely believe this thing could be
weaponized. And the crazy part is,
they're still releasing it to the public
anyway, just with a safety net that
kicks in when things get too dangerous.
But let me back up a second and give you
some context here. Back in April,
Anthropic released something called
Claude Mythos preview, and they were
super cautious about it. They only gave
it to a handful of partners because they
were concerned about cybersecurity
risks. We're talking about major
organizations that manage critical
infrastructure, [music] the kind of
stuff where if something goes wrong,
it's not just a minor inconvenience.
Then last week, [music] they expanded
access to hundreds of organizations
across 15 countries. But again, very
controlled, very selective. Now though,
they're bringing a version of that
technology to everyone through their
Claude API and enterprise plans. But
here's where it gets interesting.
They're calling it Fable 5 instead of
Mythos [music] because it's not quite
the same thing. It's the same underlying
model, yeah, but with these safety
mechanisms built in that fundamentally
change how it operates. The name
actually comes from the Latin word
fabula, which means that which is told,
and it's connected to the Greek word
mythos.
So, there's this whole linguistic
connection between [music] the two, but
the safeguards are what really separate
them. So, what exactly are these
safeguards? Well, when Fable 5 detects
that you're asking about high-risk
[music]
stuff like cybersecurity, biology,
chemistry, or something called [music]
distillation, it just straight-up
refuses to answer with its full
capabilities. [music] Instead, it falls
back to Claude Opus 4.8, which is still
a really capable model, but nowhere near
as powerful as Fable. The thing is,
Anthropic says this only happens in less
than 5% of sessions based on their early
data. So, at least 95% of the time
you're getting the full Fable
experience. But still, that 5% is there
for a reason. Now, let's talk about what
makes Fable 5 so powerful that they felt
the need to do [music] all this.
According to Anthropic, this thing is
state-of-the-art on basically every
benchmark they tested. Software
engineering, [music]
knowledge work, vision tasks, scientific
research, you name it, Fable 5 excels at
it. And the longer and more complex the
task, [music] the bigger the gap between
Fable and their other models. Stripe,
the payment company, said that Fable 5
compressed months of engineering work
into just days. They had this massive 50
million-line Ruby codebase, and Fable
did a codebase-wide migration in a
single day that would have taken an
entire team over 2 months to do
manually. That's absolutely insane when
you think about it. For knowledge work,
Fable scored the highest of any model on
something called Hebbius Finance
Benchmark, which tests senior-level
reasoning.
>> [music]
>> IMC, a trading firm, said it aced their
trading analysis evaluations across the
[music] board, factual lookup,
conceptual reasoning, root cause
analysis, expected value analysis, all
of it. And in third-party testing, an
analytics company called Hex, said Fable
was the first model to hit 90% on their
core analytics benchmark of complex,
long-running analytical tasks.
>> [music]
>> They specifically noted that on the
hardest questions, it shows strong
judgment and attention to nuance. The
vision capabilities are particularly
impressive.
>> [music]
>> Fable can extract precise numbers from
detailed scientific figures and even
rebuild a web app source code from
screenshots alone. They showed this demo
where Fable played Pokémon Fire Red from
start to finish [music] using only raw
game screenshots. No maps, no navigation
aids, nothing. Earlier Claude models
needed this whole complex helper system
to play Pokémon, but Fable [music] just
did it with vision alone. They also had
it build a simulation of the solar
system where it derived the planet's
orbital motion from physics first
principles and use that to predict solar
eclipses. It's genuinely pretty wild
stuff. But, here's where things get a
bit concerning and why Anthropic is
being so careful.
>> [music]
>> Mythos class models have apparently
reached a threshold where they present
significant risks.
In cybersecurity specifically, these
models excel at discovering and
exploiting software vulnerabilities.
They can make cyber attacks
substantially easier and cheaper to
commit. And it's not just about finding
exploits. They can perform multiple
different parts of a cyber attack
including reconnaissance, discovery,
lateral movement, the whole nine yards.
That's what they call agentic hacking
[music] and it's genuinely dangerous in
the wrong hands.
So, Anthropic built these classifiers,
which are basically separate AI systems
that detect potential misuse including
jailbreak attempts and prevent the main
model from responding. When the
classifiers catch something related to
cybersecurity, biology, and chemistry,
or distillation, boom, you get routed to
Opus 4.8 instead. And they're pretty
upfront about it. They tell you when
this happens.
They tested these safeguards pretty
extensively, too. They ran an external
bug bounty that produced no universal
jailbreaks in over 1,000 hours of
testing. They worked [music] with
external red teaming organizations which
also failed to find universal
jailbreaks, though the UK AIS apparently
made some progress toward one within a
brief initial testing window. Now,
Anthropic [music] admits it's probably
impossible to completely prevent
universal jailbreaks.
But their goal is to make any remaining
jailbreaks sufficiently slow and costly
that they can detect [music] and prevent
them before they're used at scale. The
biology and chemistry stuff is
interesting because Anthropic is being
extra cautious here. They tested Mythos
5's ability to complete a challenging
step in designing adeno-associated
viruses, which are used in gene therapy.
And here's the thing. These models
outperformed specialized protein
language models using their biological
reasoning alone, even though they
weren't explicitly trained for this
task. That's both promising for
therapeutic development and kind of
scary because the same capability could
enable the design of dangerous viruses
in the wrong hands. There's also this
whole distillation thing, which is
basically when someone tries to extract
[music] Claude's capabilities to train
competing models.
Apparently, Anthropic has identified
large-scale attempts to do this from
authoritarian countries. If distillation
of Fable 5's abilities were to succeed,
it could lead to the proliferation of
near-frontier AI capabilities without
the appropriate safeguards, which is
obviously problematic. Now, for the
people who actually need the full
unrestricted power of these models,
Anthropic is launching Claude Mythos 5.
It's the exact same model as Fable 5,
but with the safeguards lifted in some
areas.
Initially, this is going to
organizations that were already approved
for Mythos preview through something
called Project Glasswing, which is being
done in collaboration with the US
government. Anthropic says Mythos 5 has
the strongest cybersecurity capabilities
of any model in the world. Soon, they're
planning to expand access through a
broader trusted access program for both
cybersecurity organizations and biology
researchers. The pricing for both models
is $10 per million input tokens and $50
per million output tokens, which is
double the price of Opus 4.8. That's
actually less than half the price of
Claude Mythos preview, though. And
honestly, the price alone might serve as
a deterrent for widespread use, but some
companies think it's worth it. Rakuten,
a shopping rewards platform, said that
at the highest effort level, Fable
reflects on and validates its own work,
and that's what makes highly autonomous
operations possible for them. The extra
thinking pays for itself. There's also
this new data retention policy that's
pretty significant. For Fable 5, Mythos
5, and future models with similar or
higher capability levels, Anthropic is
requiring 30-day [music] retention for
all traffic, even for enterprises that
previously had zero retention
agreements. They say they won't use this
data for training and will only use it
to defend against complex and novel
attacks,
>> [music]
>> including new jail breaks, and to
identify and reduce false positives. But
this could set an industry precedent
where access to increasingly powerful
models comes with mandatory data
retention policies framed as a safety
measure. The timing of all this is
pretty interesting, too. Anthropic
recently announced they've
confidentially filed a draft S1, which
basically means they're preparing for an
IPO. OpenAI announced the same thing on
Monday, and SpaceX, [music] which
includes Elon Musk's xAI, is set to go
public on Friday. So, all the major AI
companies are making these big moves at
the same time. Anthropic's initial
announcement of Mythos back in April
apparently spooked financial markets and
governments worldwide, raising concerns
that AI models had advanced to the point
where they could uncover major
vulnerabilities in software and
cybersecurity. President Trump even
signed an executive [music] order that
allows AI companies to voluntarily give
the federal government access [music] to
frontier AI models up to 30 days before
their release,
though it explicitly prohibits the
government from imposing mandatory
review. This whole release also comes
after Anthropic urged major global AI
labs to establish what they called a
coordinated brake pedal on frontier AI
development. They warned that systems
are advancing so rapidly that they may
soon achieve recursive self-improvement,
which is basically when AI can
autonomously improve itself [music]
without human intervention. That's a
pretty concerning prospect if you think
about it. For subscription plans,
Anthropic is rolling things out in
stages. From today through June 22,
Fable 5 is included in Pro, Max, Team,
and seat-based enterprise plans at no
extra cost. But on June 23, they're
pulling Fable 5 from those plans, and
using it after that will require usage
credits. They say they plan to restore
it as a standard subscription feature as
soon as possible when capacity allows.
So, yeah. This is a pretty massive
development in the AI space. We're
seeing these companies grapple with
models that are becoming so powerful
that they genuinely pose risks if used
maliciously, and they're trying to find
ways to release them safely while still
making the technology available. Whether
these safeguards will hold up long term
remains to be seen, but it's definitely
a fascinating time to be watching this
space. Also, if you want more content
around science, space, and advanced
tech, we've launched a separate channel
for that. Links in the description. Go
check it out. So, what do you think? Is
Anthropic being responsible here, or
>> [music]
>> are we entering a stage where AI models
are getting too powerful to release
normally? Let me know in the comments.
Subscribe for more AI and tech updates.
Like the video if you found it useful,
and thanks for watching. I'll catch you
in the next one.