Claude Fable 5.1 and Claude Mythos 5.1
热度趋势
趋势数据积累中
百分比基于当前可用热度信号,而非评论数或独立用户人数。
官方发布涉及Claude 模型访问、订阅权益规则,适合跟踪产品开放节奏和用户影响。
- 判定依据
- 热度约为该来源近期上榜条目中位水平的 7.7 倍
- 指标对比
- 636 vs 中位 82.5(20 条基线样本)
- 检出时间
- 09/01 19:00
Anthropic 推出了 Claude Fable 5.1 和 Claude Mythos 5.1,被誉为全球最先进的编码和知识工作模型。这些模型展现了研究能力,其中 Mythos 5.1 在 Terminal-Bench 4.0 和 CursorBench 3.2.0 上的自主编码性能有所提升。尽管 Mythos 5.1 的能力超越了 Mythos 5,但评估表明其在化学和生物风险方面仍未达到下一个风险等级。因此,Mythos 5.1 将沿用 Mythos 5 的安全防护措施进行部署,限制其在研究生物学方面的使用权限。
We’re introducing Claude Fable 5.1 and Claude Mythos 5.1. They’re the world’s most advanced models for coding and knowledge work—and their research capabilities offer an early glimpse of how AI models will contribute to scientific progress.
Claude Fable 5.1 and Claude Mythos 5.1 are the same model, but with different levels of safeguards. Fable 5.1 is generally available, while Mythos 5.1 is available only through our trusted access programs; its safeguards are specifically designed to support work in cybersecurity and the life sciences.
Alongside its increased capabilities, Fable 5.1 takes important steps towards addressing the feedback we’ve received from customers on price, data retention, and safeguards.
Price. Fable 5.1 will cost an estimated 25% less than Fable 5 for typical workloads, wherever usage is billed by token. This is because we’re reducing our pricing on cache reads (where the model reads inputs that have already been processed and stored). For highly agentic work, the savings will often be much larger—up to approximately 45%.
Data retention.Our new system of Enterprise Frontier Safeguards (EFS) gives customers complete privacy (the same as a zero data retention policy) while still being state-of-the-art at preventing adversarial use. EFS works by storing data in cloud infrastructure controlled entirely by the customer, not Anthropic. It will be made available to enterprise customers in phases, beginning later this fall. Until EFS is available, eligible customers will be able to use Fable 5.1 with zero data retention.
Safeguards.We’ve improved our safeguards to reduce false positives (where the system flags benign content). In cybersecurity, our newest safeguards block 60% fewer false positives than before. In part, this is because Fable 5.1 can now be used to discover software vulnerabilities—though not to develop exploits for them. In biology, we’ve established an access program, developed in partnership with the US government, to enable access to Claude Mythos 5.1’s advanced biology capabilities. We expect to open enrollment for scientists soon.
A new performance frontier
Claude Fable 5.1 sets a new standard for coding, knowledge work, and long-running problem-solving tasks. The charts below show that Fable 5.1 is capable of much higher performance than its predecessor, Fable 5. And when set to Low or Medium effort, Fable 5.1 achieves results similar to or better than Fable 5’s at a much lower cost. (Note that Fable 5.1 defaults to High effort in Claude Code, and to Medium in Claude Cowork and on Claude.ai.)
Terminal-Bench-Science 0.1 Accuracy vs Cost - **Fable 5.1**
- Fable 5
0 10 20 30 40 50 60 Score (%) 10 15 20 30 40 50 Mean cost per task (USD, log scale) low med high xhigh max low med high xhigh max
Terminal-Bench-Science 0.1: The standard error is ±3.5–4.5 pts per model. The public leaderboard (3 trials/task, Claude Code harness) reports Claude Opus 5 at 30.0% and Claude Fable 5 at 21.4%; our setup reproduces them at 29.0% and 24.7%, respectively, both within noise.
Fable 5.1 avoids shortcuts that result in poorer-quality work, and it’s smart enough to fix the root causes of software issues. For example, in testing by the investment firm Millennium, Fable 5.1 found the cause of a rare crash in its internal systems that none of its engineers (or any other model) had been able to explain after several years of trying.
Here, you can see how Fable 5.1 compares across various benchmarks:
Fable 5.1 Fable 5 Opus 5 GPT-5.6 Sol Agentic scientific research Terminal-Bench-Science 0.1 [1] 52.6% 24.7% 29.0% 22.4% Agentic coding Terminal-Bench 4.0 55.8% 60.9% (Mythos 5.1) 42.0% 52.3% 37.3% Knowledge work GDPval-AA v2 1853 1723 1824 1711 Computer use OSWorld 2.0 [2] 77.9% partial 72.9% partial 75.4% partial partial Computer use OSWorld 2.0 41.7% strict 36.1% strict 39.6% strict strict Multidisciplinary reasoning Humanity's Last Exam 60.9% no tools 57.8% no tools 56.6% no tools no tools 65.0% with tools 63.8% with tools 63.6% with tools with tools Business workflows AutomationBench 31.4% 17.1% 26.9% 19.6% Agentic coding CursorBench 3.2.0 73.4% 70.5% 70.0% 67.2%
Fable 5.1 was evaluated with its production safeguards enabled. On tasks where these safeguards intervened, Fable 5.1 and Fable 5 scored a zero on OSWorld 2.0, and Fable 5 scored a zero on AutomationBench. In all other interventions from our safeguards, cybersecurity tasks were completed by Claude Opus 4.8, and biology tasks were completed by Claude Opus 5. This likely reduces the performance of Fable 5.1 and Fable 5 on these benchmarks.
Our early-access partners noticed these performance upgrades, and also picked up on more qualitative improvements in the model’s outputs. Here’s what they told us:
Quote
“In internal benchmarks, Claude Fable 5.1 solves more of our coding problems than Fable 5 or Opus 5, and achieves state of the art on trading intuition. While prior models became hard to follow the longer they worked, Fable 5.1 remains readable over long, multi-step tasks.”
Company Jane Street Capital
Author Craig Falls, Head of Quantitative Research
01 /
22
Scientific research
We’re introducing Claude Fable 5.1 and Claude Mythos 5.1. They’re the world’s most advanced models for coding and knowledge work—and their research capabilities offer an early glimpse of how AI models will contribute to scientific progress. Claude Fable 5.1 and Claude Mythos 5.1 are the same model, but with different levels of safeguards. Fable 5.1 is generally available, while Mythos 5.1 is available only through our trusted access programs; its safeguards are specifically designed to support work in cybersecurity and the life sciences. Alongside its increased capabilities, Fable 5.1 takes important steps towards addressing the feedback we’ve received from customers on price, data retention, and safeguards. **Price.** Fable 5.1 will cost an estimated 25% less than Fable 5 for typical workloads, wherever usage is billed by token. This is because we’re reducing our pricing on cache reads (where the model reads inputs that have already been processed and stored). For highly agentic work, the savings will often be much larger—up to approximately 45%. **Data retention.**Our new system of Enterprise Frontier Safeguards (EFS) gives customers complete privacy (the same as a zero data retention policy) while still being state-of-the-art at preventing adversarial use. EFS works by storing data in cloud infrastructure controlled entirely by the customer, not Anthropic. It will be made available to enterprise customers in phases, beginning later this fall. Until EFS is available, eligible customers will be able to use Fable 5.1 with zero data retention. **Safeguards.**We’ve improved our safeguards to reduce false positives (where the system flags benign content). In cybersecurity, our newest safeguards block 60% fewer false positives than before. In part, this is because Fable 5.1 can now be used to discover software vulnerabilities—though not to develop exploits for them. In biology, we’ve established an access program, developed in partnership with the US government, to enable access to Claude Mythos 5.1’s advanced biology capabilities. We expect to open enrollment for scientists soon. **A new performance frontier** Claude Fable 5.1 sets a new standard for coding, knowledge work, and long-running problem-solving tasks. The charts below show that Fable 5.1 is capable of much higher performance than its predecessor, Fable 5. And when set to Low or Medium effort, Fable 5.1 achieves results similar to or better than Fable 5’s at a much lower cost. (Note that Fable 5.1 defaults to High effort in Claude Code, and to Medium in Claude Cowork and on Claude.ai.) Terminal-Bench-Science 0.1 Accuracy vs Cost - **Fable 5.1** - **Fable 5** 0 10 20 30 40 50 60 Score (%) 10 15 20 30 40 50 Mean cost per task (USD, log scale) low med high xhigh max low med high xhigh max Terminal-Bench-Science 0.1: The standard error is ±3.5–4.5 pts per model. The public leaderboard (3 trials/task, Claude Code harness) reports Claude Opus 5 at 30.0% and Claude Fable 5 at 21.4%; our setup reproduces them at 29.0% and 24.7%, respectively, both within noise. Fable 5.1 avoids shortcuts that result in poorer-quality work, and it’s smart enough to fix the root causes of software issues. For example, in testing by the investment firm Millennium, Fable 5.1 found the cause of a rare crash in its internal systems that none of its engineers (or any other model) had been able to explain after several years of trying. Here, you can see how Fable 5.1 compares across various benchmarks: Fable 5.1 Fable 5 Opus 5 GPT-5.6 Sol Agentic scientific research Terminal-Bench-Science 0.1 [1] 52.6% 24.7% 29.0% 22.4% Agentic coding Terminal-Bench 4.0 55.8% 60.9% (Mythos 5.1) 42.0% 52.3% 37.3% Knowledge work GDPval-AA v2 1853 1723 1824 1711 Computer use OSWorld 2.0 [2] 77.9% partial 72.9% partial 75.4% partial partial Computer use OSWorld 2.0 41.7% strict 36.1% strict 39.6% strict strict Multidisciplinary reasoning Humanity's Last Exam 60.9% no tools 57.8% no tools 56.6% no tools no tools 65.0% with tools 63.8% with tools 63.6% with tools with tools Business workflows AutomationBench 31.4% 17.1% 26.9% 19.6% Agentic coding CursorBench 3.2.0 73.4% 70.5% 70.0% 67.2% Fable 5.1 was evaluated with its production safeguards enabled. On tasks where these safeguards intervened, Fable 5.1 and Fable 5 scored a zero on OSWorld 2.0, and Fable 5 scored a zero on AutomationBench. In all other interventions from our safeguards, cybersecurity tasks were completed by Claude Opus 4.8, and biology tasks were completed by Claude Opus 5. This likely reduces the performance of Fable 5.1 and Fable 5 on these benchmarks. Our early-access partners noticed these performance upgrades, and also picked up on more qualitative improvements in the model’s outputs. Here’s what they told us: Quote “In internal benchmarks, Claude Fable 5.1 solves more of our coding problems than Fable 5 or Opus 5, and achieves state of the art on trading intuition. While prior models became hard to follow the longer they worked, Fable 5.1 remains readable over long, multi-step tasks.” Company Jane Street Capital Author Craig Falls, Head of Quantitative Research 01 / 22 **Scientific research** We tested the scientific research capabilities of Claude Fable 5.1 and Claude Mythos 5.1 across a wide range of domains. What we found—which includes the early examples we share below—adds to the evidence that AI models will soon make important contributions to scientific discovery. **Molecular design.**Many modern medicines work by binding to targets within the body to block, activate, or deliver something to them. High-affinity binders are necessary for drugs to work at lower doses; designing one is the first step in the development process for many common drug modalities. To see how well Claude Mythos 5.1 could do at this task, we gave the model access to open-source protein design and folding tools and sent its designs to two external organizations for experimental validation. Mythos 5.1 proved able to design very high-affinity binders. On three targets, [3] its binding affinities were 10 times higher than the best designs submitted to Adaptyv Bio’s protein design competitions. Its hit rate (that is, the number of designs that were viable binders) was the strongest we’ve measured to date: it reached nearly 50% across 12 targets. (Hit rates of 10–15% are typical in protein design today.) **Computational analysis and modeling**. Claude Fable 5.1 trained a neural network to create a new, high-resolution elevation map of a third of the planet Venus. Its work was based on radar images taken by NASA’s Magellan mission more than 30 years ago and a map that already existed for one-fifth of the planet. Claude’s new map now reveals details down to two to three kilometers, rather than 10 to 20, and shows heights up to 25% more accurately than before. We’re releasing this map under a Creative Commons license in advance of upcoming NASA VERITAS and ESA EnVision missions, in hopes that it might help them determine which geologic features to target for future observation. **Computational biology.** In computational biology, it’s common to run task-specific machine learning models on GPUs. The speed of these models is therefore a bottleneck to research progress. Mythos 5.1 provided one solution to this problem: by writing custom GPU kernels and caching their intermediate results, it sped up seven open-source deep learning models by up to 2.5 times (with identical outputs). The benefits of such speed-ups accumulate quickly. In any given experiment, biologists might run these models thousands of times (for example, testing every possible mutation near every human gene). On analyses like these, the optimized models cut estimated GPU costs by 30–60%. This kind of optimization would normally take a team of performance engineers weeks, and is often unaffordable for academic labs. Mythos 5.1 was able to do it in just days, using the publicly available source code alone. We plan to open-source these optimizations soon. Inference speedup 0 1 2 3 Speedup on an NVIDIA H100 (×) Original implementation ChromBPNet (6M) 2.1-kb DNA sequence Flashzoi (200M) 524-kb DNA sequence Enformer (250M) 196-kb DNA sequence Profluent-E1 (600M) 1,024-amino-acid protein ProGen2 (6.4B) 512-amino-acid protein Evo 2 (7B) 8-kb DNA sequence Evo 2 (40B) 8-kb DNA sequence 1.6× 1.8× 1.4× 1.6× 2.5× 1.6× 1.4× Inference speedup for seven open-source protein and genomics models on an NVIDIA H100 As our models’ scientific capabilities improve, our investment in scientific progress is also growing. Last week, we previewed the Model Hardware Standard, which allows Claude to directly and safely operate laboratory equipment. We’ve also recently expanded our support for scientists through our AI for Science program, which provides free credits to researchers working on high-impact scientific projects, and we are offering steeply discounted usage through our new Claude Team plan for scientists. **Safety, security, and alignment** AI models’ agentic capabilities have become much more powerful over the past two years. But as we’ve documented, greater autonomy comes with new risks. Work on safety, security, and alignment needs to advance at the same pace as AI capabilities. Yesterday, we published a report describing how we are improving our own alignment and security efforts Prior to releasing Claude Fable 5.1 and Claude Mythos 5.1, we (and, in some cases, external researchers) subjected the models to extensive testing for risks across many areas. We describe these efforts in full in our System Card; below is a brief summary. **Chemical and biological risks**. We tested the extent to which Claude Mythos 5.1 could help create chemical or biological weapons. This involved expert red-teaming, automated evaluations, and a tabletop exercise that paired PhD-level biologists with AI experts, testing whether the models could match human specialists’ performance. Mythos 5.1’s capabilities are greater than those of Mythos 5. However, our evaluations indicate that it still falls short of the next risk tier defined in our Responsible Scaling Policy. We are therefore deploying Mythos 5.1 with the same safeguards that we applied to Mythos 5, which restrict access to research biology capabilities.