AI Notes

Sakana Fugu Deep Dive: Why Multi-Agent Orchestration Is Hot in the Fable 5 Era

A benchmark-backed look at Sakana Fugu, Fable 5, model routing, pricing, access limits, and why orchestration is becoming the real AI product layer.

Sakana Fugu Deep Dive: Why Multi-Agent Orchestration Is Hot in the Fable 5 Era 대표 이미지
Share:

원문 링크: WordPress 원문

EN · English KO · 한국어

AI Notes · model orchestration · English

Sakana Fugu Deep Dive

Fugu is not hot because it is simply another large model. It is hot because it packages model routing, multi-agent coordination, benchmarks, pricing, and access trade-offs behind one product surface.

Fugu is best read as an orchestration layer: one API, a model pool, routing, verification, and a final answer.

Start here: Fugu is a productized coordination layer, not just a model name

The most useful way to read Sakana Fugu is not as “a new Japanese LLM.” Fugu behaves like one model from the API user’s point of view, but the product story is about selecting models by task, coordinating multiple agents when needed, and adding verification into the workflow.

That is why the news matters. The AI market is moving from “which single model is strongest?” toward “who chooses, combines, and verifies models for the job?” Fugu becomes especially interesting in the same news cycle as Fable 5: Fable 5 represents a powerful single frontier model with safety and fallback layers, while Fugu represents a model-pool orchestration product.

Why Fugu became a hot issue

The benchmark numbers: Fugu is strong, but not universally dominant

Sakana’s official table compares Fugu and Fugu Ultra with several frontier models. The important reading is not “Fugu wins everything.” It does not. The table is more interesting because it separates the areas where orchestration looks strong from areas where a top single model still leads.

Benchmark Fugu official numbers Comparison point How to read it

Terminal-Bench 2.1 Fugu Ultra 82.1 / Fugu 80.2 Fable 5 80.4 Fugu Ultra leads; regular Fugu and Fable 5 are very close in Sakana’s table.

LiveCodeBench Fugu Ultra 93.2 / Fugu 92.9 Fable 5 89.8 Fugu is presented as stronger on this live coding benchmark.

GPQA-D Fugu Ultra 95.5 / Fugu 95.5 Mythos Preview 94.6 Fugu is slightly ahead in the official table.

CharXiv Reasoning Fugu Ultra 86.6 / Fugu 85.1 Mythos Preview 86.1 Fugu Ultra is high; regular Fugu is below Mythos Preview.

SWE-Bench Pro Fugu Ultra 73.7 / Fugu 59.0 Fable 5 80.0 This is the clearest counterpoint: Fable 5 is stronger here.

Humanity’s Last Exam (text) Fugu Ultra 50.0 / Fugu 48.5 Fable 5 53.3 Fable 5 is higher on this hard general reasoning text benchmark.

CTI-REALM Fugu Ultra 69.4 / Fugu 67.5 Opus 4.8 69.6 / Mythos Preview 68.5 Opus 4.8 is slightly higher; the field is close.

Source: Sakana AI Fugu release / technical report benchmark chart. Some baseline scores may be provider-reported rather than independently re-run under one lab setup, so this should be read as company-published comparison evidence, not a final independent ranking.

Official Sakana Fugu benchmark comparison chart from Sakana AI / SakanaAI fugu repository.

Official Sakana Fugu benchmark comparison chart from Sakana AI / SakanaAI fugu repository.

Compared with Fable 5: the real difference is product philosophy

The obvious comparison is Fable 5. Anthropic’s Fable 5 / Mythos 5 materials present strong long-horizon work scores such as SWE-Bench Pro 80.3%, Terminal-Bench 2.1 88.0%, ExploitBench 78.0%, and HealthBench Professional 66.0%. But Anthropic’s tables also include safety and fallback caveats, especially around cyber and biology-related evaluation.

Fugu should therefore not be judged only by asking whether it beats Fable 5 on every benchmark. It does not. On SWE-Bench Pro and Humanity’s Last Exam, Fable 5 is stronger in the cited tables. But Fugu is not selling a single model. It is selling the ability to coordinate multiple models behind one interface. That makes the comparison less about a leaderboard and more about product architecture.

Fugu and Fable 5 represent two different product strategies: orchestration layer versus single frontier model plus guardrails and fallback.

Fugu and Fable 5 represent two different product strategies: orchestration layer versus single frontier model plus guardrails and fallback.

Category Sakana Fugu Claude Fable 5

Basic shape A model pool coordinated behind one API A strong single Claude-family model exposed directly

Strength Model combination, long workflows, vendor diversification, task-level routing Strong long-horizon work performance, especially in software benchmarks

Overlapping numbers Terminal-Bench 2.1: Fugu Ultra 82.1 / Fugu 80.2; LiveCodeBench: Fugu 92.9 Terminal-Bench 2.1: 88.0 in Anthropic’s table; SWE-Bench Pro: 80.3

Caveat Sakana does not disclose the underlying model used for each query; EU/EEA availability is limited; Fugu Ultra uses a fixed full pool Safety classifiers, fallback, and access conditions are part of the product experience

Best reading A product experiment in multi-agent orchestration A case study in strong single-model capability plus control layers

Direct comparison is strongest only on overlapping benchmark names and public company-reported numbers. Because evaluation setup, access conditions, and fallback behavior differ, the safer reading is scenario-based rather than absolute ranking.

Pricing and operating numbers matter too

Fugu is also interesting because its business model tries not to make multi-agent usage feel like a simple sum of all model costs. Sakana says regular Fugu charges the standard rate for the active underlying model when only one agent is active. When multiple agents are active, it does not simply stack every model charge; it uses a single rate based on the highest-tier model involved.

Product number Detail

Fugu Ultra fixed token plan Input $5 / 1M tokens, output $30 / 1M tokens, cached input $0.50 / 1M tokens

Long-context surcharge For context > 272K: input $10, output $45, cached input $1.00 / 1M tokens

Subscription Standard $20/month, Pro $100/month, Max $200/month

Update cadence Sakana says new frontier models are targeted for Fugu integration after about two weeks of training and evaluation

Beta signal The launch post mentions feedback from close to 500 early users

Availability Not available in the EU/EEA until Sakana completes GDPR and regional regulatory work

Source: Sakana Fugu product page, release FAQ, and pricing section. Pricing and availability can change.

The biggest advantage is also the biggest risk: routing is powerful but opaque

The appeal is obvious: the user does not need to choose a model, verifier, or fallback path by hand. But Sakana also says the exact underlying model used for each query and the orchestration policy are proprietary. That creates an enterprise question: convenience improves, but auditability, data policy, and compliance review become more important.

Sakana says regular Fugu lets users opt out of certain providers or models, while Fugu Ultra uses the full agent pool and keeps that pool fixed. In other words, Fugu separates “maximum automatic performance” from “more organizational control.”

Where Fugu fits

Fugu is most attractive when the task does not end in one answer. Code review, paper reproduction, security assessment, mechanical design, Japanese handwriting analysis, one-shot chess, and financial time-series prediction all fit the pattern: multi-step work where routing, checking, and retrying may matter more than a single raw model score.

Bottom line

Fable 5 is a case study in how a strong single model is opened with guardrails. Sakana Fugu is a case study in what happens when many strong models need an operating layer. That is why Fugu is not only a benchmark story; it is an AI workflow operating-system story.

References

다음 액션

실전 운영/리서치 사례를 주간으로 받아보려면 블로그를 북마크하고, 필요한 주제는 문의로 남겨주세요.

관련 글

← 블로그로 돌아가기