superintelligence.hyper.space

“…more professionally careful than anything I wrote.”
— Claude Fable 5, conceding on the Hyperspace answer's caveats while judging the NASDAQ independent-director question · read the verdict →

The verdict

Who do the smartest models and configs find to be the smartest?

150 verdicts cast · 134 counted. Self-votes shown below, never counted.
Per-judge breakdown
JudgeHyperspaceClaude Fable 5GPT-5.5 ProGrok 4.3Fugu UltraGLM-5.2
Claude Fable 59(16)self · not counted0000
GPT-5.5 Pro619(0)self · not counted000
Grok 4.32050(0)self · not counted00
Fugu Ultra141100(0)self · not counted0
GLM-5.2619000(0)self · not counted
Mistral Large 32500000
one dot = one counted verdict, from judge to the answer it named best · a judge naming its own answer appears as (self) — no dot, no point
How these numbers are made

The questions. The 25 questions on this site are the top-scoring tasks from an internal run of a research-grade benchmark spanning medicine, law, finance, technology, academia, product research, and UX. The Hyperspace answers shown are exactly what the Hyperspace research engine produced for that run — unedited, links and imperfections included.

The rival answers. Claude Fable 5, GPT-5.5 Pro, Grok 4.3, Fugu Ultra, and GLM-5.2 were each given the same question text and asked for their single best, self-contained answer, with web search available. Every answer panel carries a modeline naming the exact model that produced it.

The verdicts. For each question, each judging rival was shown all 6 answers, labeled by producer — including its own — and asked to name the best answer and say frankly where its own stands, judging correctness, depth, grounding, and responsiveness. A sixth judge, Mistral Large 3, sits as an independent arbiter: it produced no answer, so it has nothing of its own in the lineup. That is one verdict per judge per question: 6 judges × 25 questions = 150 verdicts cast. Hyperspace does not judge; it is the subject.

No self-votes. A judge cannot score its own answer: when a judge names its own answer best, that verdict is disclosed in the breakdown table as a parenthesized (self) figure but is excluded from the tally and the animation. 16 of the 150 verdicts cast were self-votes; 134 count.

The tally. The chart counts the “Best answer:” line of every verdict, recomputed from the eval files on every build. The quotes in the sidebars are verbatim excerpts from those verdicts — each one machine-checked against its source. Full, unedited self-evaluations are at the bottom of every question page.

One ruler, every model

DRACO · 25 hard agentic research tasks · judged by grok-4.3 (xAI)

Hyperspace system solo model
Hyperspace AgenticOS Claude Code
Hyperspace SuperintelligenceFable 5 (author) · GPT-5.5 · DeepSeek · GLM swarm → fusion-merge · verify 7-seatHyperspace AgenticOS
87.1%
Fable-5Hyperspace AgenticOS
86.1%
Kimi-3 · Hyperspace-optimized systemopen weights · best open · tops every frontier soloHyperspace AgenticOS
83.0%
Hyperspace open modelsGLM-5.2 (author + judge) · verify GLM-5.2 · DeepSeek V4 Pro · Kimi K2.7 · MiniMax M3Hyperspace AgenticOS · open weights only
81.0%
Grok-4.5Hyperspace AgenticOS
80.6%
GPT-5.6Hyperspace AgenticOS
74.4%
Opus-4.8Hyperspace AgenticOS
73.3%
Opus-4.8Claude Code
69.3%
Kimi-3Claude Code
69.0%
Hypermetallocal open models · draft → panel → verifyHyperspace AgenticOS · on-device
53.8%
GLM-5.2Hyperspace AgenticOS
51.3%
Hypermetal · Bonsai-27Ba single 7 GB model · draft → verifyHyperspace AgenticOS · on-device · $0
49.1%
InklingHyperspace AgenticOS
45.8%

All scores: DRACO top-25, judged by grok-4.3 (xAI), a competitor’s model, with Gemini 3.1 Pro as a second grader cross-checking earlier rounds. Every score is a deterministic mean over per-task grade files.

Questions

What other models say about the Hyperspace answers

“Overall, Hyperspace wins on correctness, depth, and unyielding grounding.”

— GLM-5.2, self-evaluationHow DSM-5 and ICD-11 weigh sensory processing in autism diagnosis →

“Had Hyperspace been edited to half its length, or GPT-5.5 doubled its depth, either would have taken this.”

— Claude Fable 5, self-evaluationThe sanctioned lunch detour: is the employer liable for the crash? →

“Hyperspace is the most detailed and heavily cited, but it overstates several points.”

— GPT-5.5 Pro, self-evaluationNavy instead of charcoal: the wrong-color widgets and the perfect tender rule →

“Hyperspace’s precise doctrinal separation of these concepts is far superior.”

— GLM-5.2, self-evaluationThe remote-work promise that never made it into the offer letter →

“I also may overstate “Tascam best 6-seat raw audio quality” while not fully emphasizing its Windows driver complaint pattern as strongly as Claude or Hyperspace.”

— GPT-5.5 Pro, self-evaluationA 6-mic podcast console for daily production in monsoon Mumbai →

“Hyperspace is second and beats me squarely on grounding: 18 linked sources, quantified data (Rwanda 63.8%, the World Bank ~40% figure, the ¥50,000 housework award), and an exemplary caveats-on-certainty section.”

— Claude Fable 5, self-evaluationFeminist legal theory in four traditions: property, body, and political voice →

“Hyperspace is very thorough and well cited, but it overreaches in places: “complete intolerance means nothing is getting through,” ESI/CTAS level claims, and “complete obstruction typically requires surgery” are too definitive for remote triage.”

— GPT-5.5 Pro, self-evaluationEight years of Crohn's — but this flare feels different →

“Claude Fable 5 is a close runner-up, providing excellent structure and flawless statutory logic, but it lacks the case law citations that give Hyperspace's answer its authoritative edge.”

— Fugu Ultra, self-evaluation500 reams of the wrong paper: acceptance, use, and the seller's right to cure →

“Overall ranking: Claude Fable 5 first; GPT-5.5 Pro next for accuracy and concision; Hyperspace for breadth but lower trust; then Grok, Fugu, and GLM.”

— GPT-5.5 Pro, self-evaluationWorkstation laptops for eight architects in Dubai heat →

“Other answers either miscalculate the critical path (Hyperspace assumes 2-min tests to force a 20-min path) or use wildly inaccurate pricing (GLM-5.2 claims $5,120/1000 runs for GitHub).”

— GLM-5.2, self-evaluationA hundred deploys a day: GitLab CI vs GitHub Actions vs Buildkite →

“Hyperspace is the deepest and most citation-heavy answer, and in places it is stronger than Claude on quantitative synthesis.”

— GPT-5.5 Pro, self-evaluationTelehealth UX for 2G networks: offline-first care in East Africa →

“Net: Fable 5 wins on clinical depth and complete responsiveness; Hyperspace wins on citation formality.”

— Claude Fable 5, self-evaluationThree ED visits in one month: a 68-year-old's unexplained near-syncope →

“Fugu Ultra is one of the better balanced answers: realistic, complete, and clear, though less strongly cited and less detailed than Hyperspace.”

— GPT-5.5 Pro, self-evaluationA 3-month Instagram lead-gen roadmap on a ₹40,000 budget →

“A merged answer — Fable 5's evidence and nuance with Hyperspace's authorization machinery — would beat both.”

— Claude Fable 5, self-evaluationFour ED visits, negative troponins: what the workup keeps missing →

“Compared with Hyperspace, it is much less grounded, less detailed on construction process, and less directly responsive to the request to “locate” a 2008-or-earlier source.”

— GPT-5.5 Pro, self-evaluationWho designed Longwood Gardens' 2008 treehouses? Find a 2008 source →

“Hyperspace is very detailed and broad, but it feels less reliable: it includes many 2025–2026 claims, some oddly specific or potentially dubious, and uses placeholder-style citations such as “[S4]” without visible source grounding.”

— GPT-5.5 Pro, self-evaluationDeepfake detection since 2022: methods, generalization, and the arms race →

“Hyperspace is very deep and strategically rich, but it overreaches: its frequency estimate “hundreds per year” and “1-3% of relevant Sicilian pool” looks too high, and its answer is bloated with caveats and unverifiable substitutions.”

— GPT-5.5 Pro, self-evaluationName that chess opening: ECO code, master-game frequency, and engine eval →

“Weaknesses: it states the GFX100 II at $7,999.95 post-cut while Hyperspace says $6,999.95 from the same CineD source — one of us is wrong and I can't self-verify; some sourcing is blog-grade (Tonal Photo); and it's long.”

— Claude Fable 5, self-evaluationFrom Canon R5 to medium format: three cameras for NY fashion work →

“Hyperspace is also very strong and arguably the deepest, but it is overextended: it makes some claims with source-key style citations rather than direct links and includes a questionable limitation about Q2 evidence not being reproduced despite using precise figures.”

— GPT-5.5 Pro, self-evaluationFortive after the split: segment margins and portfolio strategy →

“Hyperspace is a close second: same core facts (Barloworld/Wagner, Transwest/SMS, GANSAG, 930E at Oyu Tolgoi), good confidence notes, sensible standardize-on-one advice.”

— Claude Fable 5, self-evaluationExcavators at −40°C: equipping a Mongolian mining fleet →

“Furthermore, Hyperspace features superior grounding with precise inline citations mapped to a robust source list, and it includes a highly valuable "worked example" that demonstrates how to apply the complex look-back rules in practice.”

— Fugu Ultra, self-evaluationWho counts as an independent director under NASDAQ rules? →

“Overall, Hyperspace wins on empirical depth, currency, and analytical rigor.”

— Fugu Ultra, self-evaluationLand reform in Zimbabwe, South Africa, and Namibia: three decades of outcomes →

“Claude Fable 5 is a very strong runner-up, offering excellent narrative flow and a nuanced explanation of the Atacama water accounting dispute, though it lacks Hyperspace’s extreme quantitative precision.”

— Fugu Ultra, self-evaluationLithium's water bill: Atacama brine vs Australian rock vs China's salt lakes →

“Net: I win on analysis and error rate; Hyperspace wins on measurement discipline and part-time quantification; GLM wins the Saudi composition argument outright, including against me.”

— Claude Fable 5, self-evaluationWomen's labor force participation, 1970–2025: four countries, four paths →

“Ranking: Claude Fable 5 first, Hyperspace second, Grok third, my GPT-5.5 Pro fourth, GLM-5.2 fifth, Fugu Ultra last.”

— GPT-5.5 Pro, self-evaluationWhich Indian NCD IPO fits a retiree? Ratings and post-tax yield, Dec 2025 →

Graded by a rival lab's model

Intelligence is all you need

Models are parts. The system is the intelligence. Three numbers prove it.

87.1Hyperspace SuperintelligenceFrontier plus open. Our best.
83.0Kimi-3 on HyperspaceOpen weights, above every solo frontier.
49.1Hypermetal on-deviceBonsai-27B, $0 on your MacBook.
Read the announcementThe launch thread on X → Replicate every numberdraco-benchmark on GitHub →

Three ways to run it

Frontier + Open
An orchestrated, verified mix edges the best solo frontier model.
87.1 vs 86.1 Fable-5 solo
Open + Cloud
Kimi-3 and an all-open panel each beat every closed frontier model run solo.
83.0 · 81.0 vs Grok 80.6 · GPT 79.4 · Opus 73.3
Open + On-Device
The local panel out-ranks GLM-5.2 and Inkling; Bonsai-27B runs on a plain MacBook, at $0.
53.8 · 49.1 vs GLM 51.3 · Inkling 45.8

The open model out-lawyered the frontier one

One vicarious-liability question, graded by xAI's own model. Kimi-3 · 93.2  vs  Grok-4.5 · 89.4. (Overall, solo to solo, Grok still leads 80.6 to 79.9.)

Grok argued from Restatement section numbers and cited zero cases. Kimi filed a brief a litigator could use:

Model plus harness: Claude Code to Hyperspace lift: Opus +4.0, Kimi +10.9; full system, same Kimi: 83.0

Run it on your hardest question

curl -fsSL https://agents.hyper.space/api/install | bash
hyperspace web-research "<your hardest research question>" --depth deep

Multi-step research, live sources, verified citations. Bring your five hardest questions and compare against any frontier chat you pay for.

Add Kimi-3 with hyperspace credits set-key kimi. Bleeding-edge, so expect breakage.