Sakana AI has launched Fugu, a model that does not answer on its own but orchestrates other models to handle each request. Its Ultra version matches Fable 5 across several benchmarks, with a 73.7 on SWE Bench Pro against 69.2 for Opus 4.8. The catch is real-world use, where researcher Ethan Mollick calls it incredibly slow, around thirty minutes per coding test.
Key Takeaways
- Fugu is a model trained to call other models, with selection, delegation, and synthesis handled internally.
- Fugu Ultra matches Fable 5 and Mythos Preview on coding, science, and reasoning.
- In practice, Fugu Ultra returns solid but very slow answers, a gap between the leaderboard and the desk.
Have an AI Sum Up This Article
ChatGPTA model that calls other models to answer
The idea is unusual. Fugu is itself a language model, but trained to summon other models from a swappable pool, copies of itself included. It decides on its own whether to handle a task or delegate it, then runs the selection, verification, and synthesis without the user seeing the machinery.
All of it runs through a single OpenAI-compatible API. Two variants exist, a Fugu built for low latency on everyday coding and chatbot use, and a Fugu Ultra meant for top quality on complex, multi-step problems. The latter carries the benchmark claims.
The house commands respect. Sakana AI was founded by Llion Jones, co-author of the 2017 paper that launched the transformer architecture, and by David Ha, both formerly at Google. We already put the lab to the test in our review of the Japanese AI everyone was talking about, and the orchestrator approach extends that research signature.
The technical bet is clear. Rather than one ever-larger model, Fugu banks on coordinating a set, a path others explore on the agent side too, as shown by the Tencent study where agents finished 14% of real tasks. Orchestration promises the best of each model, as long as it holds up under load.
73.7 on SWE Bench Pro, but 30 minutes per test
The published scores put Fugu Ultra among the leaders. On SWE Bench Pro, it lands 73.7 against 69.2 for Opus 4.8. On GPQA-D, a scientific reasoning test, it scores 95.5 against 92.0, and it pushes Humanity’s Last Exam to 50.0 where GPT-5.5 stays at 41.4.
Sakana AI claims parity with Fable 5 and Mythos Preview across that battery of tests, and the lab laid out its readings on its official site. On paper, this is an orchestrator that rivals the best monolithic models on the market, without being one itself.
The field test cools the enthusiasm. Researcher Ethan Mollick found Fugu Ultra incredibly slow during his usual coding tests, around thirty minutes per task, for a result he calls fine but short of Fable in practice. The leaderboard says one thing, the work session says another.
The gap comes from the architecture. Each request triggers an internal chain of calls, checks, and synthesis, and that coordination costs compute time. Quality comes from a process, not an instant answer, and that process is paid in latency.
More articles on Horizon
- Kimi K3 Ships Its Open Weights, 1.4TB Free
- ChatGPT Voice Now Controls Your Whole Desktop
- Claude Opus 5 Matches the Best AI at Half the Cost
Benchmarks don’t tell the whole story of real use
For a team evaluating Fugu, the lesson is simple. A top-of-the-leaderboard score does not guarantee a good experience inside an agent loop where latency stacks up at every turn. Thirty minutes per task rules the tool out for interactive work, even if the final answer holds.
The realistic use case narrows. Fugu Ultra fits the asynchronous handling of heavy problems, where waiting is acceptable in exchange for high quality, while the standard Fugu covers the fast everyday. Framed that way, the orchestrator finds a place, just not that of a drop-in replacement for fast models.
On the competitive side, the episode weighs on how we read leaderboards. It is a reminder that speed is a full dimension of quality, an angle the price war already forced across rivals. Buyers now watch the quality-latency pair, not the score alone, and a model that wins a benchmark in thirty minutes does not reassure a product lead on a deadline.
There is also a lesson for anyone shopping for a model on numbers. A benchmark measures whether the answer is correct, rarely whether it arrives in time, and those two things are not the same purchase. Fugu makes that split unusually visible, since it can top a chart and still frustrate a developer in the same afternoon.
The Sakana method keeps its value. If the lab cuts latency without losing quality, orchestration becomes a real third path against single models, a debate that runs right through how Fable 5 crushed every other public model on code. The next marker will be an Ultra version fast enough to hold inside a work session.
Follow the story on Horizon.


