News Weekly
LV 10 XP
0% read
S47model-release
#47 Issue #1Confirmed

Sakana AI releases Fugu Max and Fugu Ultra v2.0 efficiency-model updates

On Friday 2026-09-11 (2:15 AM JST), Sakana AI announced and shipped two new members of its Fugu orchestration family, both "available today via our standard OpenAI-compatible API": Fugu Max (v1.0) — the cost-performance variant. Orchestrates "our largest pool of open and specialized models to date," explicitly including NVIDIA's Nemotron open-model family (via the Sakana–NVIDIA collaboration announced on NVIDIA/open-model work). Priced at $2 per 1M input tokens and $6 per 1M output tokens (cached input $0.25; web_search/web_fetch tools $0.007/call; flat rate regardless of context); 1M-token context window; 128K max output. Fugu Ultra (v2.0) — the peak-performance flagship. Priced at $5 input / $30 output / $0.50 cached input per 1M tokens, rising to $10/$45/$1.00 above 272K context; 1M-token context; 128K max output. Training cutoff 2026-08-28. Sakana stresses that Fable 5, Fable 5.1 and GPT-6-Astra are NOT in Fugu Ultra v2's model pool — the system reaches its benchmark results by orchestrating a swappable pool of open and specialized models instead of relying on individual closed frontier models.

An empty conductor's rostrum faces a semicircle of interchangeable module sockets while a mechanical arm reshuffles modules, leaving one rail socket unfitted.
How do you want to read this?

Tailored emphasis while keeping the full article available.

Best for you · Builder

⌘ Jump to architecture, developer details, and the hands-on route.

At a glance

The essential information in 30 seconds

What happened

On Friday 2026-09-11 (2:15 AM JST), Sakana AI announced and shipped two new members of its Fugu orchestration family, both "available today via our standard OpenAI-compatible API":

  1. Fugu Max (v1.0) — the cost-performance variant. Orchestrates "our largest pool of open and specialized models to date," explicitly including NVIDIA's Nemotron open-model family (via the Sakana–NVIDIA collaboration announced on NVIDIA/open-model work). Priced at $2 per 1M input tokens and $6 per 1M output tokens (cached input $0.25; web_search/web_fetch tools $0.007/call; flat rate regardless of context); 1M-token context window; 128K max output.
  2. Fugu Ultra (v2.0) — the peak-performance flagship. Priced at $5 input / $30 output / $0.50 cached input per 1M tokens, rising to $10/$45/$1.00 above 272K context; 1M-token context; 128K max output. Training cutoff 2026-08-28. Sakana stresses that Fable 5, Fable 5.1 and GPT-6-Astra are NOT in Fugu Ultra v2's model pool — the system reaches its benchmark results by orchestrating a swappable pool of open and specialized models instead of relying on individual closed frontier models.

Headline claims (COMPANY CLAIM unless noted): Fugu Max achieves the best overall score on six benchmarks (Terminal Bench 2.1, GPQAD, AA-LCR, GDP.pdf, AutomationBench, SWEFish — the last being Sakana's internal benchmark), expands the cost-performance Pareto frontier on seven of ten benchmarks, and its output pricing is 40-60% lower than Sonnet 5, GPT 5.6 Terra and Kimi K3. Fugu Ultra v2 achieves best-or-joint-best on five of eight benchmarks (GDP.pdf, Chartography, SWEFish, DeepSWE, Toolathon) and top-2 on seven of eight; Chartography 48.3 vs Opus 5's 27.3 and Fable 5's 29.5; DeepSWE 74.3, outperforming models that cost three-to-five times more per token (BigGo's independent recap puts DeepSWE 74.3 vs GPT-6 Astra 74.1 and Claude Fable 5.1 67.4).

The release page frames the two as "the same core orchestration architecture optimized for two distinct missions" — Fugu Max = best output at lowest cost; Fugu Ultra v2 = highest capability on complex multi-step tasks. "Both models are available today… upgrading to Max or Ultra v2 requires a single-line parameter change. No migration. New architecture, same API." Distribution: direct console (console.sakana.ai) plus OpenRouter, Vercel AI Gateway, opencode (models.dev), Creao and Merge; subscription tiers $20/$100/$200 per month also include Fugu Max and Ultra.

An in-window follow-up: on Sep 17, 2026, Sakana Chat (the consumer product) rolled Fugu Max out as its default model alongside a new cross-conversation memory feature — the largest live consumer exposure of the new orchestrator within the same window.

Why it matters
  • The pricing war moved to "orchestration economics." Fugu Max's $2/$6 output pricing (40-60% below Sonnet 5 / GPT 5.6 Terra / Kimi K3 per Sakana; the per-token prices were independently confirmed via OpenRouter's live catalog during this research) is a different kind of price cut: instead of a single model getting cheaper, a routing layer gets cheaper per task. Forkast's same-day analysis ("The Orchestration Arbitrage") argues value migrates from model providers to orchestrators — if an orchestrator can deliver frontier-level output by routing to cheaper specialists, single-model providers lose premium-margin pricing power.
  • Counter-narrative to monolithic scaling: Sakana's core argument — "the most capable AI will never come from a single monolithic model" — lands the same week the industry is debating pacing (S15 Amodei, S22 von der Leyen, S36 Trump): orchestration is offered as a scaling path of its own, competing with the "bigger model" axis without needing a trillion-parameter flagship.
  • Export-control / AI-sovereignty positioning realized: born from the June 2026 Anthropic Fable/Mythos export-control disruption, Fugu is explicitly marketed as "the resilient, vendor-agnostic infrastructure required for true AI sovereignty." The Sep 11 release demonstrates the thesis with benchmarks that exclude the export-controlled frontier models entirely — a Japan-based answer to U.S. closed-frontier dependence.
  • Open-model ecosystem leverage: Fugu Max's bet is that the Pareto frontier of the future will be "built out of many open, specialized models working in concert," with NVIDIA Nemotron folded in via the July 2026 partnership — giving NVIDIA a distribution story for its open models and Sakana a moat-able orchestration layer.
  • Japanese AI leadership signal: a Tokyo lab shipping an enterprise-grade orchestration product with a fully bilingual release, consumer rollout (Sakana Chat), and a clear sovereignty pitch to Japanese enterprises/government is a significant marker for Japan's AI industrial policy this quarter.
Evidence

CONFIRMED

17 sources · 75 min read
Story identity
  • What: Sakana AI released Fugu Max (v1.0) and Fugu Ultra (v2.0) on September 11, 2026 — the next generation of its "Sakana Fugu" multi-agent orchestration system, sold as a single OpenAI-compatible API. Fugu Max is the cost-performance variant ($2/$6 per 1M input/output tokens) that orchestrates Sakana's largest pool of open and specialized models to date (including NVIDIA Nemotron); Fugu Ultra v2 is the peak-performance flagship ($5/$30 per 1M tokens) that pushes benchmark scores higher without relying on Anthropic's Fable 5/5.1 or OpenAI's GPT-6-Astra in its agent pool. Both are live immediately; existing Fugu customers upgrade with a single-line parameter change.
  • Event-date verification for the orchestrator (S45-style re-verification performed): The discovery event date is 2026-09-11 and it is CONFIRMED against primary sources — no re-anchoring required, no mismatch to flag.
    • Sakana AI's release page "Introducing Fugu Max and Fugu Ultra v2: Orchestrating the Pareto Frontier" is explicitly dated September 11, 2026 (https://sakana.ai/fugu-max-release/).
    • Sakana AI's official account posted the announcement at 2:15 AM · Sep 11, 2026 (JST) and follow-up "release notes" at 8:37 PM · Sep 11, 2026 (https://x.com/SakanaAILabs/status/2098233826816205275 — announcement; /status/2098511247134392562 — release notes).
    • Independent outlets uniformly date the event Sep 11: Forkast (2026-09-11 2:45 PM UTC, "On September 11, 2026, Sakana AI launched Fugu Max v1.0 and Fugu Ultra v2.0"), DataNorth (published 11 September 2026), BigGo Finance (announced September 11), GIGAZINE (Sep 14 report on the Sep 11 release), impress 窓の杜 (Sep 15, "Sakana AIは9月11日…「Fugu Max」と「Fugu Ultra v2」のリリースを発表した").
    • Noise (does not change the date): one aggregator (theresanaiforthat.com) and MarkTechPost's URL slug say "September 10" — clear US-timezone artifacts of a release announced at 02:15 JST on Sep 11 (= 10:15 AM PDT, Sep 10). The primary source is definitive: Sep 11, 2026, which is inside the 2026-09-10..2026-09-17 window.
  • Correction to the discovery characterisation (evidence discipline): the discovery record describes this as the "bio-inspired structured LLM family emphasizing Japanese-language and efficiency performance." The bio-inspired/efficiency framing is fair (Sakana's evolved-coordinator lineage — TRINITY/Conductor, ICLR 2026), but the "Japanese-language emphasis" is a mischaracterization of this release: Fugu Max/Ultra v2's pitch is cost-performance Pareto-frontier orchestration and vendor-resilient collective intelligence, not Japanese-language modeling. Japanese-language specialization at Sakana lives in the separate Namazu model (Kimi K2.6-based, per the Sakana Chat/model catalog). The Japanese connection here is real but different: Sakana is a Tokyo lab, the announcement is fully bilingual (EN/JA), and one flagship demo (kana-letter reading-order recovery) showcases classical-Japanese document analysis. This nuance is reflected in §6/§13.
  • Labels used: FACT (release date, availability, pricing, model-pool facts as published, benchmark scores as published, OpenRouter listing/prices verified directly), COMPANY CLAIM ("expands the Pareto efficiency frontier", "40-60% lower output pricing than Sonnet 5 / GPT 5.6 Terra / Kimi K3", "best overall score on six benchmarks", "surpasses GPT-6 Astra" framing, "AI sovereignty" positioning), INDEPENDENT EVIDENCE (OpenRouter catalog check executed during this research — models live, prices match; independently written recaps: Forkast, GIGAZINE, 窓の杜, BigGo, DataNorth; June-launch independent coverage: VentureBeat, TechCrunch), INTERPRETATION (orchestration-arbitrage value-capture shift), PREDICTION (pool-update cadence, price responses, EU availability).
✓

What happened?

On Friday 2026-09-11 (2:15 AM JST), Sakana AI announced and shipped two new members of its Fugu orchestration family, both "available today via our standard OpenAI-compatible API":

  1. Fugu Max (v1.0) — the cost-performance variant. Orchestrates "our largest pool of open and specialized models to date," explicitly including NVIDIA's Nemotron open-model family (via the Sakana–NVIDIA collaboration announced on NVIDIA/open-model work). Priced at $2 per 1M input tokens and $6 per 1M output tokens (cached input $0.25; web_search/web_fetch tools $0.007/call; flat rate regardless of context); 1M-token context window; 128K max output.
  2. Fugu Ultra (v2.0) — the peak-performance flagship. Priced at $5 input / $30 output / $0.50 cached input per 1M tokens, rising to $10/$45/$1.00 above 272K context; 1M-token context; 128K max output. Training cutoff 2026-08-28. Sakana stresses that Fable 5, Fable 5.1 and GPT-6-Astra are NOT in Fugu Ultra v2's model pool — the system reaches its benchmark results by orchestrating a swappable pool of open and specialized models instead of relying on individual closed frontier models.

Headline claims (COMPANY CLAIM unless noted): Fugu Max achieves the best overall score on six benchmarks (Terminal Bench 2.1, GPQAD, AA-LCR, GDP.pdf, AutomationBench, SWEFish — the last being Sakana's internal benchmark), expands the cost-performance Pareto frontier on seven of ten benchmarks, and its output pricing is 40-60% lower than Sonnet 5, GPT 5.6 Terra and Kimi K3. Fugu Ultra v2 achieves best-or-joint-best on five of eight benchmarks (GDP.pdf, Chartography, SWEFish, DeepSWE, Toolathon) and top-2 on seven of eight; Chartography 48.3 vs Opus 5's 27.3 and Fable 5's 29.5; DeepSWE 74.3, outperforming models that cost three-to-five times more per token (BigGo's independent recap puts DeepSWE 74.3 vs GPT-6 Astra 74.1 and Claude Fable 5.1 67.4).

The release page frames the two as "the same core orchestration architecture optimized for two distinct missions" — Fugu Max = best output at lowest cost; Fugu Ultra v2 = highest capability on complex multi-step tasks. "Both models are available today… upgrading to Max or Ultra v2 requires a single-line parameter change. No migration. New architecture, same API." Distribution: direct console (console.sakana.ai) plus OpenRouter, Vercel AI Gateway, opencode (models.dev), Creao and Merge; subscription tiers $20/$100/$200 per month also include Fugu Max and Ultra.

An in-window follow-up: on Sep 17, 2026, Sakana Chat (the consumer product) rolled Fugu Max out as its default model alongside a new cross-conversation memory feature — the largest live consumer exposure of the new orchestrator within the same window.

Δ

What changed?

  • Before: Sakana Fugu existed as Fugu (balanced latency/quality) and Fugu Ultra (v1.1-class, performance-optimized) plus Fugu Cyber, launched GA June 22, 2026 with an April beta; the July v1.1 release added Fugu-Cyber and a Claude Code interface; the NVIDIA/Nemotron partnership was announced July 16, 2026. There was a single "performance vs balanced" axis, and Fugu Ultra carried no versioned "Max"-style cost tier. Pricing war positioning was "frontier capability without export-control risk" (vs Anthropic's Fable/Mythos export-control revocation of June 12, 2026).
  • Change (Sep 11): Sakana split the product line onto two explicit axes of the cost-performance plane: Fugu Max pushes the cost Pareto frontier outward (biggest pool, cheapest output class among frontier-adjacent products), and Fugu Ultra v2 pushes the performance frontier upward (explicitly without Fable 5/5.1 or GPT-6-Astra in the pool). Pricing became public and aggressive ($2/$6 for Max; $5/$30 for Ultra v2), the Nemotron pool integration landed, and "orchestration" was productized as a standardized, API-compatible category.
  • After: As of the window's end, Sakana Fugu is a four-model lineup (Fugu, Fugu Ultra v2, Fugu Max, Fugu Cyber) behind one OpenAI-compatible API; Sakana Chat defaults to Fugu Max; independent commentary (Forkast, DataNorth) has begun analyzing "orchestration arbitrage" — value migrating from model providers to orchestrators; and the pricing war framing has moved from per-token model prices to per-task orchestration economics, in the same week DeepSeek's V4.1-Flash (S46) reset open-weight efficiency economics.
↔

Before → Change → After

PhaseState
BeforeSakana Fugu GA (Jun 22, 2026): Fugu + Fugu Ultra v1; April 2026 beta (fugu-mini/fugu-ultra); Jul 2026 v1.1 adds Fugu-Cyber + Claude Code interface; Jul 16, 2026 NVIDIA partnership announced (Nemotron to enter pool); positioning = "frontier capability without the risk of export controls" after Anthropic's Fable 5/Mythos 5 access revocation (Jun 12, 2026); single API, single-product-line pricing ($5/$30-class Ultra, per VentureBeat's June snapshot).
ChangeSep 11, 2026 launch of Fugu Max v1.0 ($2/$6/$0.25; 1M ctx; 128K max out; flat-rate; largest-ever model pool incl. NVIDIA Nemotron; Pareto-frontier cost claims) and Fugu Ultra v2.0 ($5/$30/$0.50; $10/$45/$1.00 >272K; cutoff 2026-08-28; Fable 5/5.1 and GPT-6-Astra explicitly excluded from the pool); single-line-parameter upgrade; OpenRouter/Vercel/opencode/Creao/Merge distribution; subscription tiers include both.
AfterFugu family = Fugu / Fugu Ultra v2 / Fugu Max / Fugu Cyber behind one OpenAI-compatible API; Sakana Chat defaults to Fugu Max (Sep 17, in-window); "orchestration economics" becomes an independent analytical frame (Forkast "Orchestration Arbitrage"; DataNorth per-task-cost caution); competitive pressure on single-model API pricing; EU/EEA still excluded (GDPR compliance work).
⚙

How it works

⌘ For Builder
  • Orchestration model, not a monolithic LLM: Fugu is itself a (comparatively small) language model trained to call other LLMs — "Fugu is itself a language model trained to call various LLMs in an agent pool, including instances of itself recursively." It decides when to answer directly and when to assemble a team of experts, handling model selection, delegation, verification and synthesis internally. The user hits one OpenAI-compatible endpoint.
  • Research foundation: the learned-coordination approach builds on two ICLR 2026 papers — TRINITY ("An Evolved LLM Coordinator": a lightweight evolved coordinator assigning Thinker/Worker/Verifier roles; arXiv:2512.04695) and the Conductor (RL-trained natural-language coordination; arXiv:2512.04388) — plus the Fugu Technical Report (arXiv:2606.21228; submitted Jun 19, 2026, revised Jun 23, 2026), which describes large-scale fine-tuning, evolutionary algorithms and reinforcement learning behind the production system.
  • Model-pool design: the pool is "swappable by design" — a resilience property: if a provider restricts access, Fugu routes around the disruption. Fugu Ultra v2's headline point is that it excludes the very frontier models it benchmarks against (Fable 5, Fable 5.1, GPT-6-Astra are not in the pool; training cutoff 2026-08-28), relying on open-weights and specialized models, including NVIDIA Nemotron (coding/tool-calling/instruction-following strengths) in Fugu Max's larger pool.
  • Pool-configuration policy (from the product FAQ): Fugu Ultra's pool is fixed (it needs the full pool for peak performance); Fugu Max's pool is fixed (it must jointly optimize cost and performance); only base Fugu supports opting specific providers/models out for compliance reasons. Enterprise customers can get customized pool configurations via sales.
  • Pricing mechanics: token-plan rates are blended per request — "we never stack model fees; you are charged a single rate based on the top tier model involved" — so orchestrating four models does not mean four bills. Tools (web_search, web_fetch) are metered separately at $0.007/call under Fugu Max. Token usage and cost are reported per request.
  • What stays hidden: the FAQ states plainly (Q9) that the specific models selected and the coordination used are proprietary and not exposed by design — users cannot see which underlying models handled a given query. This is the central trade-off of the product (see Risks).
  • Timeline of the "Fugu journey" (from the release page): April 2026 beta → June 2026 GA + Fugu Ultra v1 → July 2026 Fugu-Cyber & Claude Code interface → August 2026 Sakana Chat + NVIDIA Nemotron integration → September 2026 (today) Fugu Max + Fugu Ultra v2. (Nuance: the partnership page itself is dated July 16, 2026, while the release page's timeline puts "NVIDIA Partnership" under August; minor internal inconsistency, immaterial to the event.)
!

Why it matters

  • The pricing war moved to "orchestration economics." Fugu Max's $2/$6 output pricing (40-60% below Sonnet 5 / GPT 5.6 Terra / Kimi K3 per Sakana; the per-token prices were independently confirmed via OpenRouter's live catalog during this research) is a different kind of price cut: instead of a single model getting cheaper, a routing layer gets cheaper per task. Forkast's same-day analysis ("The Orchestration Arbitrage") argues value migrates from model providers to orchestrators — if an orchestrator can deliver frontier-level output by routing to cheaper specialists, single-model providers lose premium-margin pricing power.
  • Counter-narrative to monolithic scaling: Sakana's core argument — "the most capable AI will never come from a single monolithic model" — lands the same week the industry is debating pacing (S15 Amodei, S22 von der Leyen, S36 Trump): orchestration is offered as a scaling path of its own, competing with the "bigger model" axis without needing a trillion-parameter flagship.
  • Export-control / AI-sovereignty positioning realized: born from the June 2026 Anthropic Fable/Mythos export-control disruption, Fugu is explicitly marketed as "the resilient, vendor-agnostic infrastructure required for true AI sovereignty." The Sep 11 release demonstrates the thesis with benchmarks that exclude the export-controlled frontier models entirely — a Japan-based answer to U.S. closed-frontier dependence.
  • Open-model ecosystem leverage: Fugu Max's bet is that the Pareto frontier of the future will be "built out of many open, specialized models working in concert," with NVIDIA Nemotron folded in via the July 2026 partnership — giving NVIDIA a distribution story for its open models and Sakana a moat-able orchestration layer.
  • Japanese AI leadership signal: a Tokyo lab shipping an enterprise-grade orchestration product with a fully bilingual release, consumer rollout (Sakana Chat), and a clear sovereignty pitch to Japanese enterprises/government is a significant marker for Japan's AI industrial policy this quarter.
✦

What became possible?

  • One-line migration to a cheaper/faster or stronger tier: existing Fugu API users change a single parameter to move to Fugu Max or Fugu Ultra v2 — no SDK migration, same OpenAI-compatible surface.
  • Frontier-adjacent agent workloads at $6/1M output: Fugu Max's price point makes sustained agent loops (which burn output tokens across many sub-calls) materially cheaper, within a pool that includes open-weights workhorses rather than premium closed frontier APIs.
  • Benchmark-competitive autonomous execution without the export-controlled frontier: Fugu Ultra v2 claims SWEFish/DeepSWE/Chartography-class results while excluding Fable 5/5.1 and GPT-6-Astra from its pool — so buyers in export-constrained or sovereignty-conscious environments can access frontier-adjacent capability.
  • Vendor-agnostic orchestration as a buyable product: instead of wiring multi-agent frameworks themselves, teams can buy "multi-agent system as a model" with per-request cost visibility, subscription or token billing, and integrations through OpenRouter/Vercel/opencode/Merge.
  • Consumer-scale demonstration: Sakana Chat defaulting to Fugu Max (Sep 17) puts the orchestration bet in front of everyday users, with memory features — evidence of orchestration running interactive product loads.
◎

Implications

⌘ For Builder

Technical

  • Orchestration as a first-class product category: Fugu Max/Ultra v2 treat the routing layer as the product, benchmarked on agentic workloads (Terminal Bench 2.1, AutomationBench, DeepSWE, Toolathon, SWEFish) rather than classic academic evals — consistent with the industry's agent-benchmark pivot this week (see MLPerf Inference v6.1's agentic workloads, S18).
  • Benchmark-rigor caveats (COMPANY CLAIM, not independently re-run): SWEFish is an internal benchmark "reflecting Sakana AI's own coding challenges" — not reproducible by third parties; baseline scores are provider-reported (Sakana says so in the June release); Industry-wide, Sakana's table-methodology flags "all scores other than Fugu's are reported by the model providers." No independent re-run of the Chartography/DeepSWE v2 numbers existed in the window beyond press recaps of Sakana's charts.
  • Pool opacity by design: routing and coordination are proprietary and invisible (FAQ Q9). This trades auditability for performance — a privacy/compliance surface that enterprises must evaluate, and a reproducibility problem for independent evaluation.
  • No weights, no self-host: unlike open-weight releases (cf. DeepSeek V4.1-Flash, S46, same week), Fugu Max/Ultra v2 are API-only products; the value is in the live pool, not in a downloadable artifact. Architecture details are public via the technical report; runtime behavior is not.
  • Efficiency claims are pool-level, not model-level: "40-60% lower cost" compares Sakana's blended orchestrator rates against single-model list prices; per-task cost depends on routing behavior and token burn, which DataNorth explicitly cautions must be measured per finished task, not per token (orchestration tokens are billed as ordinary tokens; a task that quietly consults four models can erase the per-token advantage).
  • The "excluded frontier" claim is technically load-bearing: Fugu Ultra v2's notability depends on achieving 74.3 DeepSWE without Fable 5/5.1/GPT-6-Astra — an anti-lock-in demonstration; if the scores hold up under independent re-runs, it materially strengthens the orchestration scaling thesis.

Developer

  • Try-before-you-buy is trivial: OpenAI-compatible endpoint; available on OpenRouter (verified live during this research: sakana/fugu-max and sakana/fugu-ultra-v2) so developers can test with existing tooling and minimal credit top-up; also Vercel AI Gateway, models.dev (opencode), Creao, Merge.
  • Integration cost is near zero; behavior cost is not: migration is one parameter, but per-request latency/quality/cost differ from any single model — re-tune timeouts, max tokens (128K cap), and budget logic; Ultra's >272K context surcharge doubles rates, so long-context workers should expect $10/$45 at extreme lengths.
  • Pool controls matter for compliance: only base Fugu allows agent opt-outs; Max and Ultra pools are fixed. Teams with strict data-flow requirements must either use Fugu, buy enterprise custom configs, or skip the product (and EU/EEA-based developers cannot use it at all — no EU/EEA service while GDPR work proceeds).
  • Cost accounting: measure per task, not per token (DataNorth's advice): track tokens-per-finished-task, since orchestration spends are opaque inside the request; Sakana reports usage/cost per request, so build per-request cost telemetry into the harness.
  • Tool metering note: Fugu Max bills web_search/web_fetch at $0.007/call — an extra line item agent-heavy apps need to budget (relevant to the week's larger theme that tool-using agents multiply API spend).

Enterprise

  • A vendor-resilience option, explicitly marketed: positioned for buyers burned by API revocations and export controls — enterprises and governments wanting "AI sovereignty" can route around single-vendor dependence; the pool exclusion of Fable/GPT-6-Astra is a feature for export-sensitive buyers.
  • Compliance boundary — EU/EEA exclusion: not available in the EU/EEA during GDPR compliance work; watch for EU AI Act GPAI interaction (S34) — an orchestration layer built on third-party models creates a tangled provider-attribution chain for regulated entities.
  • Auditability gap: enterprises cannot see which models processed their data (FAQ Q9) — a real blocker for regulated sectors (finance, health, legal) and for anyone with data-residency obligations; opt-out pools exist only on base Fugu.
  • Managed-agent economics: at $2/$6 flat-rate with per-request cost reporting and subscription tiers, Fugu Max is positioned as a predictable-budget managed-agent API — attractive for deploying agent fleets without running orchestration frameworks in-house.
  • Japanese public/enterprise angle: the same week's Japanese coverage (窓の杜) emphasizes "国産" (domestic) orchestration; with Sakana's broader SCSK/Sumitomo collaboration and Japan MoD AI research contract (both announced around this period), Fugu Max/Ultra v2 give Japanese enterprises a domestic orchestration alternative to U.S. closed APIs — an enterprise-procurement signal worth tracking.

Strategic

  • Value migrates to the routing layer (INDEPENDENT EVIDENCE/INTERPRETATION): Forkast's "Orchestration Arbitrage" thesis — model-agnostic orchestrators capture value previously held by model providers, challenging both the "compute landlord" and "model-as-a-product" theses — is a structural claim about where AI value accrues, contested by the model labs.
  • Japan stakes a position in the orchestration layer: a major Japanese lab productizing learned orchestration with NVIDIA's open-model stack partners Japan's AI industry strategy with the anti-lock-in narrative; expect government-procurement and sovereignty-policy follow-ons (ties to the week's "AI sovereignty" and export-control thread, cf. S40, S63).
  • Same-week efficiency double-punch: Fugu Max ($2/$6 orchestrator) + DeepSeek V4.1-Flash (cheap open-weight efficiency model, S46) + reported frontier price cuts (Forkast's "three labs cut frontier prices in 72 hours") make the week's pricing signal unmistakable: efficiency and orchestration economics, not scaling alone, are the competitive frontier.
  • NVIDIA's open-model play: Nemotron's presence in Fugu Max shows NVIDIA distributing its open models through third-party orchestration — strengthening the open-model ecosystem against closed frontier APIs, in the same week NVIDIA pushed the "agents program themselves" enterprise narrative at Dreamforce (S38).
  • Closed labs may respond by tightening ecosystems (Forkast's "what to watch"): if orchestration becomes the primary developer interface, monolithic providers risk commoditization — expect either competitive orchestration layers or API-ecosystem hardening (and note the same-week agent-platform offensive from OpenAI/Microsoft/Salesforce, S29/S30/S26).
⚠

Risks & limitations

Risks
  • Black-box routing / compliance risk (highest): the pool is invisible by design; a request's data may flow through undisclosed third-party models; "you cannot audit which model read your data, and Sakana can change both without telling you" (DataNorth) — a serious concern for regulated data, IP, and residency-sensitive workloads, and a reputational single point of failure if a pool model misbehaves.
  • Hidden-cost risk: orchestration tokens are billed as ordinary tokens; multi-model consultations can erase the per-token price advantage (DataNorth's caution) — enterprises that buy $2/$6 expecting a monolith's bill may see per-task costs approach frontier levels.
  • Single-vendor concentration re-created: Fugu removes dependence on model vendors but creates dependence on Sakana as the orchestrator — a new single point of failure (availability, pricing, policy, company fate).
  • Self-reported benchmarks: SWEFish is internal; baseline scores are provider-published; no independent reproduction of the v2 headline numbers existed in-window — "beats GPT-6 Astra" should be treated as a claim.
  • Agent-safety surface: orchestrated autonomous agents capable of long-horizon, tool-using work (research, code, security per Fugu Cyber) amplify the week's agent-containment concerns (S01, S02, S14); darwinian learned coordination is harder to audit than a single model's behavior.
  • EU/EEA unavailability and geo-policy: no EU/EEA service during GDPR compliance; US/China export-control dynamics continue to shape which pools stay available to whom — a product explicitly built on geopolitical turbulence is itself exposed to it.
  • Fixed pools limit control: Max and Ultra pools cannot be trimmed for compliance; customers needing "no model X" must downgrade to Fugu or negotiate enterprise terms.
Limitations
  • No open weights: Fugu Max/Ultra v2 are API-only; no self-hosting, no local runs, no artifact-level INSPECT of a model file (the lab for this story therefore verified availability and pricing via the public OpenRouter catalog instead — see labs/S47.md).
  • Benchmark provenance: the Pareto-frontier/performance tables are Sakana-published (COMPANY CLAIM); GIGAZINE, BigGo and others restated Sakana's numbers; no third-party re-run of Chartography 48.3 / DeepSWE 74.3 existed in-window.
  • Opaque internals: routing topology, pool contents and orchestration prompts are undisclosed — "IP by design" — which limits technical evaluation to black-box probing.
  • Discovery-characterisation correction: "Japanese-language emphasis" was not the story of this release (see §1); Japanese-language modeling remains Namazu's lane at Sakana.
  • Price-page lags: DataNorth noted Sakana's pricing page still showed v1.1-era rates at publication time; the product page now carries v2.0 rates — third-party price snapshots should be re-checked against the console.
  • In-window evidence bounds: only six days of post-launch evidence existed in-window; Sakana Chat's Fugu Max rollout (Sep 17) is inside the window but user reaction postdates it.
?

Open questions

  1. How large is Fugu Max's pool, exactly? "Largest pool to date, including NVIDIA Nemotron" — no count, no roster is published; pool size and composition are core product facts that remain undisclosed.
  2. Do the headline benchmark numbers survive independent re-runs? Chartography 48.3 / DeepSWE 74.3 / Terminal Bench 2.1 leadership need third-party execution (the same standard applied to DeepSeek's V4.1-Flash tables this week).
  3. When does the pool update next? Sakana's FAQ says ~two weeks to train/evaluate a new Fugu version after a major model release — with GPT-6 Astra and Fable 5.1 now mainstream, will they ever enter the pool, and does the "excluded frontier" claim degrade as those models age?
  4. Does per-task orchestration spend actually beat single-model spend in production? DataNorth's "$1,100-volume" arithmetic and per-task-cost advice await real-world telemetry.
  5. When does EU/EEA availability arrive under GDPR work, and how will EU AI Act GPAI attribution treat the hidden pool?
  6. What will closed labs do in response — build their own orchestration layers, tighten API ecosystems, or keep cutting single-model prices (the Forkast "what to watch")?
  7. How much margin does Sakana hold on $2/$6 blended rates when underlying open/specialized models carry their own token costs, and does the blend survive a pool-model price change?
↗

What happens next?

  • In-window (Sep 17, 2026): Sakana Chat defaults to Fugu Max with a new memory feature — first large consumer-scale rollout of the new orchestrator.
  • Weeks ahead: per the FAQ's ~2-week pool-update cadence, a next Fugu iteration could land quickly as new models stabilize; first independent re-runs of the v2 benchmark claims; third-party per-task cost analyses (Forkast/DataNorth follow-ups); possible pricing responses from OpenAI/Anthropic/Google on single-model tiers.
  • Quarter ahead: EU/EEA availability decision under GDPR compliance work; enterprise custom-pool deals; Japan public-sector adoption signals (MoD/SCSK/Sumitomo cluster already visible); whether major closed labs launch competitive orchestration layers or harden API ecosystems.
★

Editorial takeaway

Fugu Max and Fugu Ultra v2 are not primarily a benchmark story — they are the first clean productization of "the routing layer is the product," with pricing ($2/$6, $5/$30) that reframes the AI cost conversation from per-token to per-finished-task. The technical claim that earns attention is Fugu Ultra v2 delivering frontier-adjacent results without Fable 5, Fable 5.1 or GPT-6-Astra in its pool: orchestration as an anti-lock-in scaling path, born from the June export-control shock and aimed squarely at "AI sovereignty." The discipline note for readers: every headline number is Sakana-published, the pool internals are deliberately invisible, and real-world per-task economics — not list prices — will decide whether the orchestration-arbitrage thesis holds. In a week dominated by calls to pace the frontier, Sakana's answer (with DeepSeek's V4.1-Flash alongside it) was to make frontier-class agent capability cheaper, swappable and less dependent on any single lab — a forceful market-level counterpoint to the pause narrative.

Evidence-status summary: CONFIRMED (event, date, availability, pricing — primary sources + OpenRouter catalog check); COMPANY CLAIM (Pareto-frontier performance, benchmark leadership, "40-60% cheaper," "surpasses GPT-6 Astra," sovereignty framing); INDEPENDENT EVIDENCE (OpenRouter live listing/prices verified during this research; Forkast, GIGAZINE, 窓の杜, BigGo, DataNorth recaps; June-launch VentureBeat/TechCrunch background); INTERPRETATION/PREDICTION as labeled inline.

Two smooth quality curves rise together, the cheaper frontier stepping outward then bending sharply upward past a drawn step into an expensive regime.
⌘

Lab: NO-LAB

⌘ For Builder
≡

Research sources

Primary Sources (8)
Primary
Sakana Fugu console — Models pageOfficial console listing of Fugu / Fugu Ultra / Fugu Cyber (v1.0-era benchmark table incl. Terminal Bench 2.1 82.1 for Fugu-Ultra, SWE Bench Pro 73.7; Fugu Cyber CyberGym 86.9 / CTI-REALM 72.1); confirms console access as the direct distribution surface. — Primary / FACT (console availability, published v1.0 benchmark table).Date: accessed 2026-09-19 during research
Visit source ↗
Primary
Sakana Fugu Technical Report — arXivPeer-archived technical report (Yujin Tang et al., 14 authors): Fugu family described as "orchestrator models" that "are themselves language models trained to understand user queries and dynamically devise agentic scaffolds"; training paradigm (large-scale fine-tuning, evolutionary algorithms, reinforcement learning); SOTA claims on SWE-Bench Pro, Terminal Bench, LiveCodeBench, GPQA-Diamond, HLE, CharXiv Reasoning; Fugu (latency) vs Fugu-Ultra (quality) split. — Primary / FACT (architecture and training description), COMPANY CLAIM (SOTA results as reported).Date: submitted 2026-06-19, revised 2026-06-23 (v2)
Visit source ↗
Primary
Sakana AI announcement post on X — @SakanaAILabsSecond primary confirmation of event date; announcement framing ("next evolution of Sakana Fugu's multi-agent orchestration system; Fugu Max expands the Pareto efficiency frontier… orchestrating our largest pool of open-weights and specialized models to date, including NVIDIA Nemotron family"); DeepSWE claim ("outperforms models that cost three to five times more per token… without Fable 5, Fable 5.1, or GPT-6-Astra in its agent pool"); resilience positioning ("vendor lock-in, API revocations, sudden service cutoffs"). — Primary / FACT (date, availability), COMPANY CLAIM (performance claims).Date: 2026-09-11 (2:15 AM)
Visit source ↗
Primary
Sakana AI Teams With NVIDIA to Advance Open Model Innovation from Japan — partnership announcementNVIDIA Nemotron integration into Sakana Fugu's agent pool (coding, tool-calling, instruction-following strengths); David Ha and Kari Briski quotes; "open models become more useful when orchestrated together" thesis that Fugu Max's pool claim rests on. — Primary / FACT (partnership, Nemotron-in-pool intent).Date: 2026-07-16 (partnership page; note the Sep 11 release page's timeline loosely labels this "August")
Visit source ↗
Primary
Sakana Fugu beta announcement — April 2026Lineage of the product (fugu-mini/fugu-ultra beta variants); foundation of Fugu on ICLR 2026 papers TRINITY and Conductor; GPQAD/LCBv6/SWEPro beta table; OpenAI-compatible-API design intent. — Primary / FACT (product lineage, research grounding).Date: 2026-04-24
Visit source ↗
Primary
Sakana Fugu: One Model to Command Them All — June GA announcementBackground on Fugu architecture (Fugu is itself a language model trained to call LLMs in an agent pool, including itself recursively); Fugu vs Fugu Ultra positioning; export-control context (Anthropic Fable/Mythos access revocation); benchmark methodology note ("all scores other than Fugu's are reported by the model providers"); early-user testimony (code review, security assessment, patent landscape); links to Trinity/Conductor ICLR 2026 papers and technical report. — Primary / FACT (architecture narrative), COMPANY CLAIM (GA benchmark claims), INTERPRETATION (anti-vendor-lock-in rationale).Date: 2026-06-22 (GA + Fugu Ultra v1 launch)
Visit source ↗
Primary
Sakana Fugu — Multi-Agent System as a Model (official product page)Four-model lineup (Fugu, Fugu Ultra, Fugu Max, Fugu Cyber) behind one OpenAI-compatible API; Fugu Ultra (v2.0) pricing $5/$30 per 1M in/out, $0.50 cached, $10/$45/$1.00 above 272K context; Fugu Max $2/$6/$0.25 flat-rate regardless of context, web_search/web_fetch billed at $0.007 per call (billed per call; added as separate line item); subscription tiers $20/$100/$200/month (each includes Fugu, Fugu Ultra and Fugu Max); FAQ material: Fugu Ultra and Fugu Max pools are FIXED, only base Fugu supports agent opt-outs; routing/coordination proprietary and not exposed (Q9); token-plan never stacks model fees (single top-tier rate); ~2-week pool-update cadence after a new frontier model release; Fugu not available in EU/EEA while GDPR compliance work proceeds; per-request cost reporting; enterprise custom-pool services via sales. — Primary / FACT (pricing, pool policy, availability constraints, architecture positioning).Date: accessed 2026-09-19 during research (product page reflecting v2.0 lineup)
Visit source ↗
Primary
Introducing Fugu Max and Fugu Ultra v2: Orchestrating the Pareto Frontier — Sakana AI official release pageTHE event-date anchor; full official announcement: Fugu Max (largest pool to date incl. NVIDIA Nemotron family; $2/$6 per 1M input/output tokens; best overall score on six benchmarks — Terminal Bench 2.1, GPQAD, AA-LCR, GDP.pdf, AutomationBench, SWEFish; Pareto-frontier expansion on 7 of 10 benchmarks; 40-60% lower output pricing than Sonnet 5, GPT 5.6 Terra, Kimi K3) and Fugu Ultra v2 (best-or-joint-best on five of eight benchmarks — GDP.pdf, Chartography, SWEFish, DeepSWE, Toolathon; top-2 on seven of eight; Chartography 48.3 vs Opus 5 27.3 and Fable 5 29.5; DeepSWE 74.3; training cutoff 2026-08-28; Fable 5 / Fable 5.1 / GPT-6-Astra NOT in agent pool); both live immediately via OpenAI-compatible API; single-line parameter change to upgrade; Fugu journey timeline (Apr beta → Jun GA → Jul Fugu-Cyber/Claude Code → Aug Sakana Chat + NVIDIA → Sep today); "AI sovereignty" framing; links to product page, console and technical report. — Primary / FACT (event, date, availability, published specs/prices), COMPANY CLAIM (performance, cost-frontier and benchmark claims).Date: 2026-09-11 (page dated "September 11, 2026")
Visit source ↗
Independent Sources (8)
Independent
No Claude Fable 5? No problem: Sakana achieves frontier performance with new Fugu multi-model, auto synthesis system — VentureBeatIndependent background on the June GA: routing logic hidden by design ("the specific models Fugu selects and how it coordinates them are proprietary"); Fugu vs Fugu Ultra tiers; June pricing snapshot table (Sakana Fugu Ultra $5/$30/$35 per 1M) used to establish the "before" pricing in §4; Anthropic June 12 revocation context; David Ha positioning. — Independent / CONFIRMED (background, historical pricing), INTERPRETATION (vendor-lock-in hedge).Date: 2026-06-22
Visit source ↗
Independent
Asian AI startups launch Mythos-like models as Anthropic's export ban drags on — TechCrunchIndependent background on the June launch context: Fugu named after the Japanese word for blowfish; "delivering frontier capability without the risk of export controls" positioning; David Ha on orchestration as the next frontier; Fugu as hedge strategy in the Anthropic export-control environment. — Independent / CONFIRMED (background context), INTERPRETATION (geo-strategic framing).Date: 2026-06-27
Visit source ↗
Independent
BigGo Finance — Sakana AI Launches "Fugu Ultra v2," Surpassing GPT-6 Astra; Low-Cost Version Also UnveiledIndependent recap with specific cross-model numbers used in this story: DeepSWE 74.3 (Fugu Ultra v2) vs GPT-6 Astra 74.1 vs Claude Fable 5.1 67.4; Chartography 48.3 vs Claude Opus 5 27.3; Fugu evolution timeline (Apr beta → Jun GA/Ultra v1 → Jul Cyber/Claude Code → Aug Sakana Chat/Nemotron → Sep 11 Max/Ultra v2); both models share the same orchestration foundation differing only in optimization targets; single-setting-parameter migration. — Independent / CONFIRMED (event, timeline), COMPANY CLAIM (numeric comparisons restated from Sakana).Date: 2026-09-14 (published 2026-09-14T05:05Z)
Visit source ↗
Independent
国産の「Sakana Fugu」に「Fable 5.1」「GPT-6 Astra」超えの強化モデルが登場 — impress 窓の杜 (Mado no Mori)Japanese-market independent coverage: "Sakana AIは9月11日…「Fugu Max」と「Fugu Ultra v2」のリリースを発表" (Sept 11 phrasing); positioning of the two variants (cost-efficiency vs peak-performance), visual-reasoning and real-world-software-engineering benchmarks; fixed-structure avoidance of single-vendor lock-in. — Independent / CONFIRMED (event date in Japanese coverage), COMPANY CLAIM (benchmark claims restated).Date: 2026-09-15
Visit source ↗
Independent
Sakana AI has released 'Fugu Ultra v2,' which surpasses the GPT-6 Astra… — GIGAZINE (EN)Independent Japanese-tech-press recap (EN edition): Fugu Ultra v2 outperformed GPT-6 Astra and Claude Fable 5.1 on several benchmarks per Sakana's graphs; Fugu Max wider model variety for cost-effective processing; Terminal Bench 2.1 cost-vs-score chart discussion; note that Fable 5/5.1 and GPT-6 Astra are not in Ultra v2's pool. — Independent / CONFIRMED (release event, dates), COMPANY CLAIM (benchmark numbers restated).Date: 2026-09-14
Visit source ↗
Independent
The Orchestration Arbitrage: How Sakana's Fugu Max Rewrites the Pricing War — ForkastIndependent analysis anchoring the "orchestration economics" thesis: "On September 11, 2026, Sakana AI launched Fugu Max v1.0 and Fugu Ultra v2.0"; 40-60% undercut framing; Fugu Ultra v2 pricing ($5/$30/$0.50; $10/$45/$1.00 >272K); TRINITY/Conductor technical grounding; strategic reading that value migrates from model providers to orchestrators and challenges the "compute landlord" thesis; "what to watch" (closed labs commoditizing orchestration or tightening ecosystems); integrations (OpenRouter, Vercel, opencode, Creao, Merge); Fugu Cyber numbers (86.9% CyberGym, 72.1% CTI-REALM). — Independent / INDEPENDENT EVIDENCE (event date, pricing), INTERPRETATION (value-capture shift).Date: 2026-09-11 (2:45 PM UTC)
Visit source ↗
Independent
Sakana AI launches Fugu Max and Fugu Ultra v2: $2 per 1M tokens — DataNorthIndependent same-day recap; pricing table (Fugu Max $2/$6, cached $0.25; Fugu Ultra v2 $5/$30, cached not published for v2 then; 1M context; 128K max output; long-context surcharge above 272K); the skeptic's framing used in this story's Risks/§9: invisible pool ("you cannot see the pool, you cannot audit which model read your data, and Sakana can change both without telling you"), orchestration tokens billed as ordinary tokens, advice to compare cost per finished task rather than per token. — Independent / CONFIRMED (pricing details), INTERPRETATION (cost-auditability critique).Date: 2026-09-11
Visit source ↗
Independent
OpenRouter public model catalog API — /api/v1/models (verified during research)INDEPENDENT VERIFICATION LAB (see labs/S47.md): among 447 listed models, `sakana/fugu-max` (prompt $0.000002 = $2/1M, completion $0.000006 = $6/1M, context 1,000,000) and `sakana/fugu-ultra-v2` (prompt $0.000005 = $5/1M, completion $0.00003 = $30/1M, context 1,000,000) are both live and priced exactly at Sakana's published list rates; also present: `sakana/fugu-ultra` (v1-class) and `sakana/sakana-namazu`. Confirms "available today via OpenAI-compatible API" is objectively true via a third-party provider by Sep 19, 2026. — Independent / INDEPENDENT EVIDENCE (availability + pricing verified directly, outside Sakana's own materials).Date: queried 2026-09-19 (during this research, via curl)
Visit source ↗
Secondary Sources (1)
Secondary
Sakana Chat rolls out Fugu Max model and memory feature — AI WeeklyIn-window follow-up event: Sakana Chat now defaults to Fugu Max (orchestrator routing each prompt to a chosen open model), plus a cross-conversation memory feature; Namazu and Fugu Max both selectable; Sakana Chat history (Namazu alpha Mar 2026; Fugu, code execution and attachments Aug 2026). — Secondary / CONFIRMED (in-window consumer rollout, dated).Date: 2026-09-17 (published 00:12 UTC)
Visit source ↗