Sakana Fugu Deep Dive: Why Multi-Agent Orchestration Is Hot in the Fable 5 Era
A benchmark-backed look at Sakana Fugu, Fable 5, model routing, pricing, access limits, and why orchestration is becoming the real AI product layer.

원문 링크: WordPress 원문
EN · English KO · 한국어
AI Notes · model orchestration · English
Sakana Fugu Deep Dive
Fugu is not hot because it is simply another large model. It is hot because it packages model routing, multi-agent coordination, benchmarks, pricing, and access trade-offs behind one product surface.
Fugu is best read as an orchestration layer: one API, a model pool, routing, verification, and a final answer.
Start here: Fugu is a productized coordination layer, not just a model name
The most useful way to read Sakana Fugu is not as “a new Japanese LLM.” Fugu behaves like one model from the API user’s point of view, but the product story is about selecting models by task, coordinating multiple agents when needed, and adding verification into the workflow.
That is why the news matters. The AI market is moving from “which single model is strongest?” toward “who chooses, combines, and verifies models for the job?” Fugu becomes especially interesting in the same news cycle as Fable 5: Fable 5 represents a powerful single frontier model with safety and fallback layers, while Fugu represents a model-pool orchestration product.
Why Fugu became a hot issue
-
It reduces single-vendor dependence. Sakana frames Fugu as frontier-level performance without locking the user into one provider.
-
It targets long workflows. Code review, automated research, security evaluation, and similar multi-step tasks are exactly where orchestration matters.
-
The benchmark claims are aggressive. Sakana’s official table includes GPQA-D 95.5, LiveCodeBench 92.9, Terminal-Bench 2.1 80.2, and Fugu Ultra 82.1.
-
The product shape is different. The user sees one API and one pricing model, while Sakana handles model-pool coordination behind the scenes.
The benchmark numbers: Fugu is strong, but not universally dominant
Sakana’s official table compares Fugu and Fugu Ultra with several frontier models. The important reading is not “Fugu wins everything.” It does not. The table is more interesting because it separates the areas where orchestration looks strong from areas where a top single model still leads.
Benchmark Fugu official numbers Comparison point How to read it
Terminal-Bench 2.1 Fugu Ultra 82.1 / Fugu 80.2 Fable 5 80.4 Fugu Ultra leads; regular Fugu and Fable 5 are very close in Sakana’s table.
LiveCodeBench Fugu Ultra 93.2 / Fugu 92.9 Fable 5 89.8 Fugu is presented as stronger on this live coding benchmark.
GPQA-D Fugu Ultra 95.5 / Fugu 95.5 Mythos Preview 94.6 Fugu is slightly ahead in the official table.
CharXiv Reasoning Fugu Ultra 86.6 / Fugu 85.1 Mythos Preview 86.1 Fugu Ultra is high; regular Fugu is below Mythos Preview.
SWE-Bench Pro Fugu Ultra 73.7 / Fugu 59.0 Fable 5 80.0 This is the clearest counterpoint: Fable 5 is stronger here.
Humanity’s Last Exam (text) Fugu Ultra 50.0 / Fugu 48.5 Fable 5 53.3 Fable 5 is higher on this hard general reasoning text benchmark.
CTI-REALM Fugu Ultra 69.4 / Fugu 67.5 Opus 4.8 69.6 / Mythos Preview 68.5 Opus 4.8 is slightly higher; the field is close.
Source: Sakana AI Fugu release / technical report benchmark chart. Some baseline scores may be provider-reported rather than independently re-run under one lab setup, so this should be read as company-published comparison evidence, not a final independent ranking.

Official Sakana Fugu benchmark comparison chart from Sakana AI / SakanaAI fugu repository.
Compared with Fable 5: the real difference is product philosophy
The obvious comparison is Fable 5. Anthropic’s Fable 5 / Mythos 5 materials present strong long-horizon work scores such as SWE-Bench Pro 80.3%, Terminal-Bench 2.1 88.0%, ExploitBench 78.0%, and HealthBench Professional 66.0%. But Anthropic’s tables also include safety and fallback caveats, especially around cyber and biology-related evaluation.
Fugu should therefore not be judged only by asking whether it beats Fable 5 on every benchmark. It does not. On SWE-Bench Pro and Humanity’s Last Exam, Fable 5 is stronger in the cited tables. But Fugu is not selling a single model. It is selling the ability to coordinate multiple models behind one interface. That makes the comparison less about a leaderboard and more about product architecture.

Fugu and Fable 5 represent two different product strategies: orchestration layer versus single frontier model plus guardrails and fallback.
Category Sakana Fugu Claude Fable 5
Basic shape A model pool coordinated behind one API A strong single Claude-family model exposed directly
Strength Model combination, long workflows, vendor diversification, task-level routing Strong long-horizon work performance, especially in software benchmarks
Overlapping numbers Terminal-Bench 2.1: Fugu Ultra 82.1 / Fugu 80.2; LiveCodeBench: Fugu 92.9 Terminal-Bench 2.1: 88.0 in Anthropic’s table; SWE-Bench Pro: 80.3
Caveat Sakana does not disclose the underlying model used for each query; EU/EEA availability is limited; Fugu Ultra uses a fixed full pool Safety classifiers, fallback, and access conditions are part of the product experience
Best reading A product experiment in multi-agent orchestration A case study in strong single-model capability plus control layers
Direct comparison is strongest only on overlapping benchmark names and public company-reported numbers. Because evaluation setup, access conditions, and fallback behavior differ, the safer reading is scenario-based rather than absolute ranking.
Pricing and operating numbers matter too
Fugu is also interesting because its business model tries not to make multi-agent usage feel like a simple sum of all model costs. Sakana says regular Fugu charges the standard rate for the active underlying model when only one agent is active. When multiple agents are active, it does not simply stack every model charge; it uses a single rate based on the highest-tier model involved.
Product number Detail
Fugu Ultra fixed token plan Input $5 / 1M tokens, output $30 / 1M tokens, cached input $0.50 / 1M tokens
Long-context surcharge For context > 272K: input $10, output $45, cached input $1.00 / 1M tokens
Subscription Standard $20/month, Pro $100/month, Max $200/month
Update cadence Sakana says new frontier models are targeted for Fugu integration after about two weeks of training and evaluation
Beta signal The launch post mentions feedback from close to 500 early users
Availability Not available in the EU/EEA until Sakana completes GDPR and regional regulatory work
Source: Sakana Fugu product page, release FAQ, and pricing section. Pricing and availability can change.
The biggest advantage is also the biggest risk: routing is powerful but opaque
The appeal is obvious: the user does not need to choose a model, verifier, or fallback path by hand. But Sakana also says the exact underlying model used for each query and the orchestration policy are proprietary. That creates an enterprise question: convenience improves, but auditability, data policy, and compliance review become more important.
Sakana says regular Fugu lets users opt out of certain providers or models, while Fugu Ultra uses the full agent pool and keeps that pool fixed. In other words, Fugu separates “maximum automatic performance” from “more organizational control.”
Where Fugu fits
Fugu is most attractive when the task does not end in one answer. Code review, paper reproduction, security assessment, mechanical design, Japanese handwriting analysis, one-shot chess, and financial time-series prediction all fit the pattern: multi-step work where routing, checking, and retrying may matter more than a single raw model score.
Bottom line
Fable 5 is a case study in how a strong single model is opened with guardrails. Sakana Fugu is a case study in what happens when many strong models need an operating layer. That is why Fugu is not only a benchmark story; it is an AI workflow operating-system story.
References
-
Sakana AI — Sakana Fugu release
-
Sakana AI — Sakana Fugu product page
-
SakanaAI/fugu GitHub repository
-
arXiv — Sakana Fugu Technical Report
-
Anthropic — Claude Fable 5 and Claude Mythos 5
-
Anthropic — Redeploying Fable 5
-
Claude Platform Docs — Introducing Claude Fable 5 and Claude Mythos 5
다음 액션
실전 운영/리서치 사례를 주간으로 받아보려면 블로그를 북마크하고, 필요한 주제는 문의로 남겨주세요.

