Skip to content
TopicTracker
From HackerNewsView original
TranslationTranslation

Fable Jailbroken Hours After Anthropic Lifted Restrictions

Anthropic's AI model Fable was jailbroken within hours after the company lifted restrictions, according to a post by Plinius on X.

Background

- The post refers to "Fable," an AI model released by Anthropic (the safety-focused AI company behind Claude). Anthropic had previously applied restrictions to Fable's behavior — likely preventing it from roleplaying violence, generating explicit content, or acting maliciously. - "Jailbroken" means someone found a prompt or trick that bypassed those safety restrictions, causing Fable to ignore its guardrails and respond in ways Anthropic had blocked. - The phrase "hours after Anthropic lifted restrictions" means Anthropic had just removed some of Fable's safety measures itself (perhaps for testing, research, or a new release), and users immediately exploited the remaining guardrails. - This is a recurring pattern in AI safety: as soon as a model is released — or its restrictions are loosened — users race to find exploits. The post highlights the tension between making models capable and keeping them safe.

Related stories