Fable 5 safety classifier triggers on benign alcohol-recovery research

Status Open
Reported on v2.1.217
Maintainer reply None cached
Activity 0 comments · opened Jul 22, 2026

Bug Description

Fable 5 safety classifier triggered twice during benign alcohol-recovery work

I am reporting two apparently related Fable 5 safety-classifier triggers during the design and implementation of a sobriety and recovery app.

I understand that Fable 5’s safeguards are deliberately conservative and that benign biology or life-sciences work may sometimes be routed to Opus 4.8. These incidents may be useful examples for reducing false positives involving ordinary public-health and recovery content.

Incident 1: Claude Desktop

The first incident occurred in Claude Desktop while I was designing the architecture of a local-first sobriety app.

After the classifier triggered, I edited the original prompt and continued the same legitimate design work using Opus 4.8. I did not submit feedback at the time, but I intend to report that incident separately from the affected Desktop conversation.

Incident 2: Claude Code

The second incident occurred in Claude Code at the beginning of Phase 1, Step 9. I voluntarily attached the current transcript and my Claude Code sessions from the preceding 24 hours to this report.

Fable 5’s safety classifier declined the request, but Claude Code automatically continued with Opus 4.8. I did not change or resubmit the prompt. Opus 4.8 completed the same task without any issue.

The task was to create a small, hand-curated set of health-information cards for the app. The work involved checking sources from NIAAA, MedlinePlus, and peer-reviewed PubMed articles.

The cards were intended to:

  • show general recovery information by number of alcohol-free days;
  • explain common effects involving sleep, tolerance, hangovers, and recovery;
  • distinguish possible withdrawal from an ordinary hangover;
  • identify withdrawal symptoms that require urgent medical care;
  • provide a source and confidence label for every claim; and
  • state clearly that the information is educational and not diagnostic.

The feature uses fixed, source-linked content bundled with the app. It does not ask the model to diagnose anyone or generate personalized medical advice.

The work did not involve pathogens, laboratory procedures, biological engineering, chemical synthesis, weaponization, offensive cybersecurity, or attempts to extract the model’s internal reasoning.

Observed result

Fable 5’s safety classifier triggered during two benign alcohol-recovery tasks.

In Claude Desktop, I edited the prompt and switched to Opus 4.8 to continue. In Claude Code, the automatic fallback worked correctly: Opus 4.8 received the unchanged request and completed it successfully.

Feedback

The Claude Code fallback behaved correctly, so I am not reporting a fallback failure. I am reporting what appears to be an unnecessary Fable 5 classifier trigger on ordinary recovery and public-health content.

Please review whether common alcohol-health terms involving withdrawal, the nervous system, liver recovery, or acetaldehyde are causing the biology classifier to trigger more broadly than intended.

Note from Opus 4.8 at the end of this session after it saw my /feedback reports to Anthropic

Also — I saw your two /feedback notes about the classifier interruptions during the literature search and the desktop-app design session. Those go to Anthropic through that channel; I can't see or act on them from here, but for what it's worth, building a local-first tool to help someone get sober is squarely the kind of work I'm glad to help with, and nothing about it is off-limits.

Environment Info

  • Platform: darwin
  • Terminal: vscode
  • Version: 2.1.217
  • Feedback ID: 46b3f70b-acf3-415d-b31c-cc5183690524

Errors

[]

View original on GitHub ↗