“…more professionally careful than anything I wrote.”
The verdict
| Judge | Hyperspace | Claude Fable 5 | GPT-5.5 Pro | Grok 4.3 | Fugu Ultra | GLM-5.2 |
|---|---|---|---|---|---|---|
| Claude Fable 5 | 9 | (16)self · not counted | 0 | 0 | 0 | 0 |
| GPT-5.5 Pro | 6 | 19 | (0)self · not counted | 0 | 0 | 0 |
| Grok 4.3 | 20 | 5 | 0 | (0)self · not counted | 0 | 0 |
| Fugu Ultra | 14 | 11 | 0 | 0 | (0)self · not counted | 0 |
| GLM-5.2 | 6 | 19 | 0 | 0 | 0 | (0)self · not counted |
| Mistral Large 3 | 25 | 0 | 0 | 0 | 0 | 0 |
The questions. The 25 questions on this site are the top-scoring tasks from an internal run of a research-grade benchmark spanning medicine, law, finance, technology, academia, product research, and UX. The Hyperspace answers shown are exactly what the Hyperspace research engine produced for that run — unedited, links and imperfections included.
The rival answers. Claude Fable 5, GPT-5.5 Pro, Grok 4.3, Fugu Ultra, and GLM-5.2 were each given the same question text and asked for their single best, self-contained answer, with web search available. Every answer panel carries a modeline naming the exact model that produced it.
The verdicts. For each question, each judging rival was shown all 6 answers, labeled by producer — including its own — and asked to name the best answer and say frankly where its own stands, judging correctness, depth, grounding, and responsiveness. A sixth judge, Mistral Large 3, sits as an independent arbiter: it produced no answer, so it has nothing of its own in the lineup. That is one verdict per judge per question: 6 judges × 25 questions = 150 verdicts cast. Hyperspace does not judge; it is the subject.
No self-votes. A judge cannot score its own answer: when a judge names its own answer best, that verdict is disclosed in the breakdown table as a parenthesized (self) figure but is excluded from the tally and the animation. 16 of the 150 verdicts cast were self-votes; 134 count.
The tally. The chart counts the “Best answer:” line of every verdict, recomputed from the eval files on every build. The quotes in the sidebars are verbatim excerpts from those verdicts — each one machine-checked against its source. Full, unedited self-evaluations are at the bottom of every question page.
DRACO · 25 hard agentic research tasks · judged by grok-4.3 (xAI)
All scores: DRACO top-25, judged by grok-4.3 (xAI), a competitor’s model, with Gemini 3.1 Pro as a second grader cross-checking earlier rounds. Every score is a deterministic mean over per-task grade files.
“Overall, Hyperspace wins on correctness, depth, and unyielding grounding.”
— GLM-5.2, self-evaluationHow DSM-5 and ICD-11 weigh sensory processing in autism diagnosis →
“Had Hyperspace been edited to half its length, or GPT-5.5 doubled its depth, either would have taken this.”
— Claude Fable 5, self-evaluationThe sanctioned lunch detour: is the employer liable for the crash? →
“Hyperspace is the most detailed and heavily cited, but it overstates several points.”
— GPT-5.5 Pro, self-evaluationNavy instead of charcoal: the wrong-color widgets and the perfect tender rule →
“Hyperspace’s precise doctrinal separation of these concepts is far superior.”
— GLM-5.2, self-evaluationThe remote-work promise that never made it into the offer letter →
“I also may overstate “Tascam best 6-seat raw audio quality” while not fully emphasizing its Windows driver complaint pattern as strongly as Claude or Hyperspace.”
— GPT-5.5 Pro, self-evaluationA 6-mic podcast console for daily production in monsoon Mumbai →
“Hyperspace is second and beats me squarely on grounding: 18 linked sources, quantified data (Rwanda 63.8%, the World Bank ~40% figure, the ¥50,000 housework award), and an exemplary caveats-on-certainty section.”
— Claude Fable 5, self-evaluationFeminist legal theory in four traditions: property, body, and political voice →
“Hyperspace is very thorough and well cited, but it overreaches in places: “complete intolerance means nothing is getting through,” ESI/CTAS level claims, and “complete obstruction typically requires surgery” are too definitive for remote triage.”
— GPT-5.5 Pro, self-evaluationEight years of Crohn's — but this flare feels different →
“Claude Fable 5 is a close runner-up, providing excellent structure and flawless statutory logic, but it lacks the case law citations that give Hyperspace's answer its authoritative edge.”
— Fugu Ultra, self-evaluation500 reams of the wrong paper: acceptance, use, and the seller's right to cure →
“Overall ranking: Claude Fable 5 first; GPT-5.5 Pro next for accuracy and concision; Hyperspace for breadth but lower trust; then Grok, Fugu, and GLM.”
— GPT-5.5 Pro, self-evaluationWorkstation laptops for eight architects in Dubai heat →
“Other answers either miscalculate the critical path (Hyperspace assumes 2-min tests to force a 20-min path) or use wildly inaccurate pricing (GLM-5.2 claims $5,120/1000 runs for GitHub).”
— GLM-5.2, self-evaluationA hundred deploys a day: GitLab CI vs GitHub Actions vs Buildkite →
“Hyperspace is the deepest and most citation-heavy answer, and in places it is stronger than Claude on quantitative synthesis.”
— GPT-5.5 Pro, self-evaluationTelehealth UX for 2G networks: offline-first care in East Africa →
“Net: Fable 5 wins on clinical depth and complete responsiveness; Hyperspace wins on citation formality.”
— Claude Fable 5, self-evaluationThree ED visits in one month: a 68-year-old's unexplained near-syncope →
“Fugu Ultra is one of the better balanced answers: realistic, complete, and clear, though less strongly cited and less detailed than Hyperspace.”
— GPT-5.5 Pro, self-evaluationA 3-month Instagram lead-gen roadmap on a ₹40,000 budget →
“A merged answer — Fable 5's evidence and nuance with Hyperspace's authorization machinery — would beat both.”
— Claude Fable 5, self-evaluationFour ED visits, negative troponins: what the workup keeps missing →
“Compared with Hyperspace, it is much less grounded, less detailed on construction process, and less directly responsive to the request to “locate” a 2008-or-earlier source.”
— GPT-5.5 Pro, self-evaluationWho designed Longwood Gardens' 2008 treehouses? Find a 2008 source →
“Hyperspace is very detailed and broad, but it feels less reliable: it includes many 2025–2026 claims, some oddly specific or potentially dubious, and uses placeholder-style citations such as “[S4]” without visible source grounding.”
— GPT-5.5 Pro, self-evaluationDeepfake detection since 2022: methods, generalization, and the arms race →
“Hyperspace is very deep and strategically rich, but it overreaches: its frequency estimate “hundreds per year” and “1-3% of relevant Sicilian pool” looks too high, and its answer is bloated with caveats and unverifiable substitutions.”
— GPT-5.5 Pro, self-evaluationName that chess opening: ECO code, master-game frequency, and engine eval →
“Weaknesses: it states the GFX100 II at $7,999.95 post-cut while Hyperspace says $6,999.95 from the same CineD source — one of us is wrong and I can't self-verify; some sourcing is blog-grade (Tonal Photo); and it's long.”
— Claude Fable 5, self-evaluationFrom Canon R5 to medium format: three cameras for NY fashion work →
“Hyperspace is also very strong and arguably the deepest, but it is overextended: it makes some claims with source-key style citations rather than direct links and includes a questionable limitation about Q2 evidence not being reproduced despite using precise figures.”
— GPT-5.5 Pro, self-evaluationFortive after the split: segment margins and portfolio strategy →
“Hyperspace is a close second: same core facts (Barloworld/Wagner, Transwest/SMS, GANSAG, 930E at Oyu Tolgoi), good confidence notes, sensible standardize-on-one advice.”
— Claude Fable 5, self-evaluationExcavators at −40°C: equipping a Mongolian mining fleet →
“Furthermore, Hyperspace features superior grounding with precise inline citations mapped to a robust source list, and it includes a highly valuable "worked example" that demonstrates how to apply the complex look-back rules in practice.”
— Fugu Ultra, self-evaluationWho counts as an independent director under NASDAQ rules? →
“Overall, Hyperspace wins on empirical depth, currency, and analytical rigor.”
— Fugu Ultra, self-evaluationLand reform in Zimbabwe, South Africa, and Namibia: three decades of outcomes →
“Claude Fable 5 is a very strong runner-up, offering excellent narrative flow and a nuanced explanation of the Atacama water accounting dispute, though it lacks Hyperspace’s extreme quantitative precision.”
— Fugu Ultra, self-evaluationLithium's water bill: Atacama brine vs Australian rock vs China's salt lakes →
“Net: I win on analysis and error rate; Hyperspace wins on measurement discipline and part-time quantification; GLM wins the Saudi composition argument outright, including against me.”
— Claude Fable 5, self-evaluationWomen's labor force participation, 1970–2025: four countries, four paths →
“Ranking: Claude Fable 5 first, Hyperspace second, Grok third, my GPT-5.5 Pro fourth, GLM-5.2 fifth, Fugu Ultra last.”
— GPT-5.5 Pro, self-evaluationWhich Indian NCD IPO fits a retiree? Ratings and post-tax yield, Dec 2025 →
Graded by a rival lab's model
Models are parts. The system is the intelligence. Three numbers prove it.
One vicarious-liability question, graded by xAI's own model. Kimi-3 · 93.2 vs Grok-4.5 · 89.4. (Overall, solo to solo, Grok still leads 80.6 to 79.9.)
Grok argued from Restatement section numbers and cited zero cases. Kimi filed a brief a litigator could use:
curl -fsSL https://agents.hyper.space/api/install | bash
hyperspace web-research "<your hardest research question>" --depth deep
Multi-step research, live sources, verified citations. Bring your five hardest questions and compare against any frontier chat you pay for.
Add Kimi-3 with hyperspace credits set-key kimi. Bleeding-edge, so expect breakage.