concept · concept/responsible-scaling-policy
Responsible Scaling Policy
Also called Anthropic RSP, RSP, AI Safety Level, AI Safety Levels, ASL
Anthropic describes its Responsible Scaling Policy (RSP) as a voluntary framework for managing catastrophic risks from advanced AI. The current version recorded here is As accessed 2026-09-23, Anthropic's policy page listed version 3.4 as current, effective 2026-07-08.source, accessed 2026-09-23. Its basic mechanism is conditional: when an evaluation reaches a specified capability threshold, stronger safeguards become a requirement for training, storing, or deploying the model. The policy draws a line between capability evidence and the protections needed to address the risk.
The original ASL ladder
The September 2023 announcement introduced AI Safety Levels (ASLs), loosely modeled on biosafety levels. Its short definitions were capability gates paired with progressively stricter safety and security expectations:
- ASL-1: Anthropic's 2023 high-level summary put ASL-1 at systems with no meaningful catastrophic risk, such as an AI that only plays chess.source, accessed 2026-09-23
- ASL-2: Anthropic's 2023 summary described ASL-2 as early dangerous capabilities that remain unreliable or no more useful than information from a search engine.source, accessed 2026-09-23
- ASL-3: Anthropic's 2023 summary placed ASL-3 at capabilities that substantially increase catastrophic-misuse risk over non-AI baselines or show low-level autonomy; the level calls for stronger deployment and security safeguards.source, accessed 2026-09-23
- ASL-4 and above: Anthropic's 2023 public summary left ASL-4 and higher undefined, anticipating a qualitative escalation in catastrophic-misuse potential and autonomy rather than publishing a complete ASL-4 safeguard checklist.source, accessed 2026-09-23
RSP version 2.0 then made the ladder operational: Anthropic's October 15, 2024 RSP update described ASL-2 as its baseline safety and security standard and said capability thresholds would trigger stronger safeguards.source, accessed 2026-09-23 The 2024 update maps a capability threshold to the safeguard standard needed once that threshold is reached.
In the original announcement, ASL-1 was summarized as “no meaningful catastrophic risk,” ASL-2 as “early signs of dangerous capabilities,” and ASL-4 and above as “not yet defined.” Those are Anthropic's 2023 shorthand, not a current rating scale to apply to every model. The announcement described ASL-3 as a step up in both deployment protection and weight security. When Anthropic activated ASL-3 protections for Claude Opus 4 in May 2025, it called that deployment a “precautionary and provisional action” and said it had not yet determined that Opus 4 had definitively crossed the capability threshold. The activation post also said Claude Sonnet 4 did not require the ASL-3 Standard. A level label could therefore describe safeguards chosen under uncertainty, not just a confirmed capability classification.
What ASL means in the current policy
RSP version 2.0 separated the model's capabilities from the safeguards applied to it: capability thresholds specify when protections should increase, and required safeguards specify what must be in place. The current policy says earlier editions used ASL for lists of controls and model categories; it still uses ASL for existing mitigation levels, but says future mitigations are better described by the risk argument and the threat actors they must address. The policy explains that future capability thresholds do not naturally fall into discrete levels. The current policy PDF records this change in Appendix B. RSP v3.4, effective 2026-07-08, says ASL still describes present levels of mitigations for existing models; for future capabilities it favors risk arguments and threat actors over fixed lists of controls grouped into ASL-N levels.source, accessed 2026-09-23 This is why the 2023 four-tier shorthand is useful context but incomplete as a description of Anthropic's current policy.
Model labels in the latest public reporting
Anthropic's August 2026 Risk Report covers activity through July 15. It says model capabilities and safeguards vary across multiple dimensions and that it is “no longer using” ASL-2 and ASL-3 to label both. Instead, its production model table gives Level 1, 2, or 3 for the robustness of biological blocking classifiers. Anthropic's August 2026 Risk Report, covering models and actions through 2026-07-15, says its previous use of ASL-2 and ASL-3 for both model capability and safeguards is no longer used because those vary across multiple dimensions.source, accessed 2026-09-23 In the Risk Report with coverage through 2026-07-15, biological classifier robustness was Level 3 for Claude Mythos 5 and Claude Fable 5, Claude Mythos Preview, Claude Opus 4.8, Claude Opus 4.7, and Claude Sonnet 5; Level 2, with some product surfaces at Level 1, for Claude Opus 4.6, Opus 4.5, Sonnet 4.6, and Sonnet 4.5; Level 1 for Opus 4.1 and Opus 4; and no blocking classifiers (N/A) for Sonnet 4, Haiku 4.5, and other legacy models. These are classifier-robustness labels, not ASL designations.source, accessed 2026-09-23 Those classifier levels measure resistance to jailbreaks; they are not ASL tiers and do not summarize a model's overall risk.
The report predates the later releases recorded below. Their launch pages describe safeguard packages without assigning an overall ASL number: Anthropic's September 2026 launch page says Mythos 5.1 falls short of the next RSP risk tier and uses the safeguards applied to Mythos 5; Fable 5.1 is the same underlying model with different safeguards. The page does not assign either model an ASL number.source, accessed 2026-09-23 Anthropic's 2026-09-22 Opus 5.5 release page says the model is available on all platforms and has safeguards similar to Fable 5.1; it does not assign Opus 5.5 an ASL number.source, accessed 2026-09-23 For these releases, reporting the named safeguards and access conditions is more precise than inferring an ASL level.
Facts
- latest published version
- As accessed 2026-09-23, Anthropic's policy page listed version 3.4 as current, effective 2026-07-08.source, accessed 2026-09-23
- asl 1 gate
- Anthropic's 2023 high-level summary put ASL-1 at systems with no meaningful catastrophic risk, such as an AI that only plays chess.source, accessed 2026-09-23
- asl 2 gate
- Anthropic's 2023 summary described ASL-2 as early dangerous capabilities that remain unreliable or no more useful than information from a search engine.source, accessed 2026-09-23
- asl2 baseline standard
- Anthropic's October 15, 2024 RSP update described ASL-2 as its baseline safety and security standard and said capability thresholds would trigger stronger safeguards.source, accessed 2026-09-23
- asl 3 gate
- Anthropic's 2023 summary placed ASL-3 at capabilities that substantially increase catastrophic-misuse risk over non-AI baselines or show low-level autonomy; the level calls for stronger deployment and security safeguards.source, accessed 2026-09-23
- asl 4 gate
- Anthropic's 2023 public summary left ASL-4 and higher undefined, anticipating a qualitative escalation in catastrophic-misuse potential and autonomy rather than publishing a complete ASL-4 safeguard checklist.source, accessed 2026-09-23
- current asl use
- RSP v3.4, effective 2026-07-08, says ASL still describes present levels of mitigations for existing models; for future capabilities it favors risk arguments and threat actors over fixed lists of controls grouped into ASL-N levels.source, accessed 2026-09-23
- latest model asl reporting
- Anthropic's August 2026 Risk Report, covering models and actions through 2026-07-15, says its previous use of ASL-2 and ASL-3 for both model capability and safeguards is no longer used because those vary across multiple dimensions.source, accessed 2026-09-23
- classifier robustness snapshot
- In the Risk Report with coverage through 2026-07-15, biological classifier robustness was Level 3 for Claude Mythos 5 and Claude Fable 5, Claude Mythos Preview, Claude Opus 4.8, Claude Opus 4.7, and Claude Sonnet 5; Level 2, with some product surfaces at Level 1, for Claude Opus 4.6, Opus 4.5, Sonnet 4.6, and Sonnet 4.5; Level 1 for Opus 4.1 and Opus 4; and no blocking classifiers (N/A) for Sonnet 4, Haiku 4.5, and other legacy models. These are classifier-robustness labels, not ASL designations.source, accessed 2026-09-23
- mythos fable 5 1 public status
- Anthropic's September 2026 launch page says Mythos 5.1 falls short of the next RSP risk tier and uses the safeguards applied to Mythos 5; Fable 5.1 is the same underlying model with different safeguards. The page does not assign either model an ASL number.source, accessed 2026-09-23
- opus 5 5 public status
- Anthropic's 2026-09-22 Opus 5.5 release page says the model is available on all platforms and has safeguards similar to Fable 5.1; it does not assign Opus 5.5 an ASL number.source, accessed 2026-09-23
Timeline
- Anthropic publishes its August 2026 Risk Report, with a model coverage date of July 15source
- RSP version 3.4 takes effectsource
- RSP version 3.0 separates Anthropic's company plans from industry-wide recommendationssource
- RSP version 2.0 adds capability thresholds and required safeguardssource
- Anthropic publishes RSP version 1.0 and the first ASL summarysource