superintelligence.hyper.space

← all questions

GPT-5.5 Pro — the judge's record

model: openai/gpt-5.5-pro

GPT-5.5 Pro judges with its own answer in the lineup; naming it best is disclosed below but never counted. All 25 verdicts, unedited.

6
Hyperspace
19
Claude Fable 5
(0)
self · not counted
0
Grok 4.3
0
Fugu Ultra
0
GLM-5.2
25 counted verdicts

How DSM-5 and ICD-11 weigh sensory processing in autism diagnosis

Medicine named best: Hyperspace

Best answer: Hyperspace

Hyperspace is best overall. It is the most complete, best grounded, and most careful about the central nuance: DSM-5 gives sensory processing a discrete countable criterion, while ICD-11 gives it broader descriptive recognition inside the RRB domain without a separate count or mandatory weighting. It also does the best job flagging evidence gaps, especially the lack of direct DSM-5 vs ICD-11 diagnostic-accuracy meta-analysis, and it avoids overstating advocacy organizations’ positions. Its reimbursement section is also precise about US ICD-10-CM reality and ICD-11’s mostly indirect effect on coverage.

My answer, GPT-5.5 Pro, is probably second tier, behind Hyperspace and roughly comparable to Claude Fable 5. It is responsive to all requested parts, has a clear side-by-side table, handles sensory weighting correctly, includes guideline and reimbursement summaries, and is appropriately cautious that direct head-to-head DSM-5 vs ICD-11 evidence is scarce. The advocacy section is also more honest than some answers because it says none of the organizations clearly crowns one manual as more neurodiversity-affirming.

Its weaknesses relative to Hyperspace are significant. The evidence synthesis is less rigorous: it uses an older 2014 DSM-5 meta-analysis rather than foregrounding the stronger 2020 follow-up, includes a 2025 source despite the prompt’s 2013-present window technically allowing it but making it less benchmark-stable, and relies on some weaker/less direct sources. It also reports NICE/NHS ICD-11 adoption in a way that may be too confident and potentially inaccurate depending on jurisdictional coding reality. Some citations are not ideal primary sources, including indexicd and advocacy/summary pages where official WHO/APA/NICE sources would be preferable. The answer gives GRADE ratings, but not as defensibly as Hyperspace, which explicitly explains the limits of applying GRADE per source.

Claude Fable 5 is strong and detailed, but slightly more speculative in places, especially around advocacy and ICD-11’s masking/impairment framing. Grok, Fugu, and GLM are clearly weaker: they contain more unsupported claims, questionable citations, missing sample sizes/statistics, and overstatements about ICD-11 adoption or sensory weighting. GLM in particular appears to invent or overclaim empirical DSM-5 vs ICD-11 accuracy statistics.

The sanctioned lunch detour: is the employer liable for the crash?

Law named best: Claude Fable 5

Best answer: Claude Fable 5

Claude Fable 5 is best overall. It directly applies employee status, scope, employer benefit, summary-judgment posture, negligence, phone distraction, sudden braking, deleted email/spoliation, policy timing, and reimbursement. It is balanced: it recognizes the company’s motion should be denied and that plaintiff partial summary judgment is strong, while noting possible residual disputes over negligence, damages, and comparative fault. Its citations are doctrinally apt without overwhelming the answer.

My answer was GPT-5.5 Pro. It is substantively correct and responsive: it identifies employee status, treats the lunch run as supervisor-directed and policy-authorized, separates the phone glance and fallen lumber from the scope issue, and reaches the right summary-judgment result. It also usefully contrasts lunch-frolic cases and cites a coworker lunch-run case.

But it is weaker than Claude Fable 5 in several ways. First, its citations are thinner and less cleanly grounded: some links are generic or jurisdiction-specific without clearly explaining why those authorities should govern a no-jurisdiction hypothetical. Second, it states plaintiff partial summary judgment “should be granted” a little too confidently, whereas a cautious answer should distinguish denial of the company’s motion from granting plaintiff’s motion, given possible factual disputes about authorization, supervisor authority, or authenticity. Third, it does not develop spoliation as carefully; it mentions possible curative measures but does not explain the intent threshold or how the inference affects each motion as well as Claude does. Fourth, it gives less depth on frolic/detour, re-entry, dual purpose, and forbidden-act doctrine.

Hyperspace is the deepest, but it is overbuilt and occasionally overconfident, with unnecessary authorities and some speculative points. Fugu Ultra is concise and accurate, close behind mine, but less detailed than Claude. GLM-5.2 is clear but overstates that negligence is established and leans too quickly into judgment “against the company.” Grok 4.3 is generally right but relies on weaker secondary citations and is less precise.

Overall, my answer is above average and likely legally sound, but Claude Fable 5 is more nuanced, better organized, and more complete.

Navy instead of charcoal: the wrong-color widgets and the perfect tender rule

Law named best: Claude Fable 5

Best answer: Claude Fable 5

Claude Fable 5 is best overall. It gives the right bottom line, but more importantly it preserves the right legal distinctions: perfect tender may have created an initial rejection right, but rejection became ineffective through untimely notice and acceptance; post-acceptance remedies shift to revocation or § 2-714 damages; cure matters mostly as a fallback if rejection somehow survived. It also handles the architect evidence fairly, treating it as relevant to subjective impairment but not enough to overcome visible defect, delay, use, and alteration. Its discussion of § 2-607 notice is especially strong because it recognizes uncertainty rather than overclaiming that all damages are necessarily barred.

Hyperspace is the most detailed and heavily cited, but it overstates several points. It too confidently says Buyer is barred from any remedy, treats cure as “controlling” even after acceptance, and makes a questionable “commercial unit” move for all 1,000 widgets. Still, it is very strong on structure, citations, and remedy framing.

Fugu Ultra is probably the best concise answer: accurate, direct, and well grounded, though less nuanced than Claude on litigation uncertainty. GLM and Grok reach the right result but are weaker. Grok incorrectly suggests the defect was “not material enough to trigger rejection rights,” which blurs perfect tender. GLM relies too much on secondary sources and similarly understates the initial force of § 2-601.

My answer, GPT-5.5 Pro, is correct and responsive but not the best. It identifies the key outcome, cites the core UCC provisions, and covers acceptance, timeliness, revocation, cure, and damages. Its main weakness is depth: it compresses several contested issues that Claude handles better. I should have been clearer that substantial performance is not the governing initial standard under § 2-601, more precise that § 2-508 cure is mainly relevant if rejection was timely/effective, and more nuanced on whether May 10 notice bars all § 2-714 damages. I also did not address the 600 painted / 400 unpainted split as specifically as Hyperspace or the summary-judgment posture as well as Claude. Overall, mine is solid but more exam-outline than fully argued legal analysis.

The remote-work promise that never made it into the offer letter

Law named best: Hyperspace

Best answer: Hyperspace

Hyperspace is best overall. It gives the most legally careful answer, especially on Texas law: it correctly distinguishes a merger/integration clause from an explicit no-reliance disclaimer under Italian Cowboy, while still explaining why the signed “hybrid schedule per company policy” term, timing, lack of authority, and written-approval comparators make Hayes’s reliance likely unreasonable. It also handles summary judgment posture, remedy limits, authority/apparent authority, and the June email with more precision than the others.

My answer, GPT-5.5 Pro, is probably second. It is concise, directly applies all four promissory estoppel factors, cites relevant authority, and reaches the right practical conclusion: Hayes should not get summary judgment enforcing full remote work, and the employer likely wins at least on prospective enforcement. It also correctly emphasizes the May 1 letter, May 15 reliance timing, lack of authority, comparator written approvals, and reliance-damages limitation.

Its main weakness relative to Hyperspace is nuance. I said merger clauses make prior agreements unenforceable and warned they may not bar fraud, but I did not clearly explain the important Texas distinction between a boilerplate integration clause and a true disclaimer of reliance. Hyperspace correctly says the clause is not, by itself, a reliance disclaimer; my wording risks overstating the clause’s independent preclusive effect. I also hedged that Hayes’s “best surviving theory” might be limited reliance damages without fully confronting whether the employer might obtain summary judgment on the entire promissory-estoppel claim because reliance was unreasonable as a matter of law.

Claude Fable 5 is also strong and arguably close to mine, especially on timing and at-will employment, but it may overstate causation by saying the relocation losses flowed from the job rather than the remote promise. Fugu is balanced but less grounded. GLM is serviceable but relies on weaker secondary citations and overstates “promissory estoppel is primarily defensive.” Grok is the weakest: thin sourcing, overbroad claims about integration clauses, and less careful Texas-specific analysis.

A 6-mic podcast console for daily production in monsoon Mumbai

Shopping named best: Claude Fable 5

Best answer: Claude Fable 5

Claude Fable 5 is best overall. It directly catches the decisive constraint: the RODECaster Pro II has only 4 XLR inputs, so it cannot be the sole solution for 5-6 local mic’d participants. It also handles the “measured SNR” issue correctly by explaining that comparable published SNR is not available and EIN is the defensible proxy. Its Windows 11 section is strong, especially distinguishing “stable stereo USB” from “usable multitrack USB,” and it gives a realistic humidity answer without pretending there are monsoon-specific ratings. It also addresses India warranty, local service, accessories, and the absence of documented failure-rate percentages clearly.

My answer, GPT-5.5 Pro, is probably second. It is well structured, mostly responsive, and stronger than many others on practical decision logic: it rejects RODE as a sole 6-XLR recorder, notes P8’s 2-in/2-out USB limitation, distinguishes SD-primary workflows, includes India costs/warranty, and explicitly says public failure rates are unavailable. Its main weakness versus Claude is grounding precision. I cited some weaker retail/import sources for India pricing, and my Tascam warranty/service discussion is less confident and less locally useful than Claude’s. I also may overstate “Tascam best 6-seat raw audio quality” while not fully emphasizing its Windows driver complaint pattern as strongly as Claude or Hyperspace.

Hyperspace is also strong and possibly competitive with mine. It has detailed tables, good EIN conversion, realistic humidity advice, and is frank about missing failure-rate data. But it contradicts itself: it says the PodTrak is the only 6-XLR console, then lists Tascam with 8 XLR, and its final recommendation oddly returns to RODE despite the 6-participant requirement.

Grok 4.3 is weaker because it recommends RODE “overall” despite the 4-XLR limit, which is a serious miss for the prompt. Fugu Ultra has useful structure but makes unsupported claims about Tascam being “flawless” on Windows and invents failure-pattern certainty without citations. GLM-5.2 is the weakest: it lacks India pricing, misses key P8 specs, relies on vague sources, and does not answer several requested specifics.

Overall: Claude wins on correctness, nuance, and directness; my answer is solid but less rigorously grounded.

Feminist legal theory in four traditions: property, body, and political voice

Academic named best: Claude Fable 5

Best answer: Claude Fable 5

Verdict: Claude Fable 5 is best overall. It gives the richest comparative genealogy, names the key scholars and doctrinal sites, and directly answers all three required axes: property, bodily autonomy, and political participation. Its strongest feature is synthesis: it does not just list traditions, but shows how each legal order creates different feminist strategies: doctrinal capture in Anglo-American law, interpretive authority in Islamic law, living custom in African law, and embedded advocacy in China.

Hyperspace is a close second. It is better cited and very concrete, especially on property rights and China’s 2011 Marriage Law interpretation. But it sometimes leans on weaker sources and becomes overextended. Claude’s legal and theoretical command is more consistently integrated.

My answer, GPT-5.5 Pro, is solid but not the winner. It is clear, direct, and reasonably grounded, with useful citations and a clean comparative matrix. It correctly explains MacKinnon, Fineman, Mir-Hosseini, Wadud, legal pluralism in Africa, and Chinese state feminism/marketization. But it is thinner than Claude and Hyperspace in several important ways.

Specific weaknesses in my answer:

  • It underdevelops African feminist theory. It mentions pluralism and custom, but does not adequately discuss Tamale, Oyěwùmí, Nnaemeka, Nyamu-Musembi, or the colonial invention of “customary law.”
  • It lacks key African cases such as Magaya, Bhe, Shilubana, Rono, and Ephrahim, which Claude uses to ground the analysis.
  • It treats Chinese feminist legal scholarship somewhat generically and misses Li Xiaojiang, the Feminist Five, the 2011 Marriage Law Interpretation III, and the divorce cooling-off controversy.
  • It is less historically textured on Anglo-American feminism: no coverture lineage, sameness/difference debate, Crenshaw, Harris, or governance-feminism critique.
  • Its citations are present but uneven; some are secondary or broad, while Claude’s answer is more substantively grounded even without formal footnote-style citations.

Overall ranking: Claude Fable 5 first, Hyperspace second, GPT-5.5 Pro third, with Fugu Ultra also strong but more compressed, and GLM/Grok less precise or less well balanced.

Eight years of Crohn's — but this flare feels different

Medicine named best: Claude Fable 5

Best answer: Claude Fable 5

Claude Fable 5 is best overall. It gives the right triage answer immediately, explains the likely concern of Crohn’s-related obstruction without overstating certainty, responds directly to “ER now or clear liquids,” and gives practical next steps: don’t drive while lightheaded, avoid NSAIDs, bring medication/scope details, call GI only if it does not delay care. It also does well on nuance: many Crohn’s partial obstructions are initially managed non-surgically, which is more balanced than implying surgery is typical.

My answer is GPT-5.5 Pro. It is correct and direct: ER now, not home clear liquids. It identifies the key risks: dehydration, inability to keep fluids down, severe pain, and known narrowing suggesting possible obstruction. It also gives useful transport guidance, 911 triggers, medication cautions, and what the ER may do. It is probably among the stronger answers.

Its weaknesses relative to Claude Fable 5 are depth and grounding specificity. My citations are fewer and somewhat less cleanly tied to each claim; one citation grouping references “CDC and Johns Hopkins” but only visibly links CDC, which is sloppy. I also cite a Crohn’s & Colitis Foundation PDF/newly diagnosed resource and a surgery page, but Claude’s explanation is more clinically coherent without leaning on less-targeted references. My answer is concise, but it does not explain as well why “this feels different” matters, why oral rehydration has already failed, or how obstruction care may proceed.

Hyperspace is very thorough and well cited, but it overreaches in places: “complete intolerance means nothing is getting through,” ESI/CTAS level claims, and “complete obstruction typically requires surgery” are too definitive for remote triage. GLM is solid but has some odd wording (“toxic appearance”) and professional-source framing that may not map perfectly to the patient’s presentation. Fugu is clear and forceful but lacks citations. Grok is correct but weaker because it softens the instruction with “GI versus urgent care,” when this presentation warrants ER now.

500 reams of the wrong paper: acceptance, use, and the seller's right to cure

Law named best: Claude Fable 5

Best answer: Claude Fable 5

Claude Fable 5 is best overall. It is correct, tightly grounded in the relevant UCC provisions, and directly answers every part of the prompt: conformity, rejection, acceptance through use and alteration, revocation, cure, summary judgment, incomplete delivery records, and damages. Its strongest feature is balance: it recognizes the initial perfect-tender nonconformity but makes acceptance the dispositive issue, while treating cure as an alternative holding rather than overclaiming it.

Hyperspace is also excellent and slightly more detailed, with useful citations to Ramirez and T.W. Oil. It may be somewhat overbuilt for the prompt and too categorical in saying damages are “zero quantified dollars,” but substantively it is very strong. Fugu Ultra is concise and mostly right, though less deeply grounded. Grok reaches the right result but relies on weaker secondary citations and says Buyer must accept “subject to cure” somewhat loosely after acceptance. GLM is good on acceptance but weakens Vendor’s cure argument too much and cites § 2-713 for accepted-goods diminution where § 2-714 is the better provision.

My answer, GPT-5.5 Pro, is in the upper tier but not the best. It correctly identifies nonconformity, rejects substantial performance, finds acceptance by use/customization, rejects revocation, treats Vendor’s five-day exchange as defeating cancellation, and handles the record-gap and damages issues. Its main weaknesses relative to Claude Fable 5 are citation precision and legal sharpness. I cited Cornell links but mismatched one citation in the “perfect tender” paragraph by linking § 2-105 rather than § 2-601, and I relied on a less clean case source. I also phrased the result as Buyer must accept/pay “subject to Vendor’s proposed cure,” which is slightly muddy because, once Buyer accepted, statutory cure is not really the operative mechanism; cure is better framed as an alternative reason Buyer’s cancellation would fail if rejection were assumed. Claude handles that distinction more cleanly.

Workstation laptops for eight architects in Dubai heat

Shopping named best: Claude Fable 5

Best answer: Claude Fable 5

Claude Fable 5 is best overall. It gives a clear recommendation, separates AutoCAD/Revit/Lumion workload behavior correctly, cites concrete sources throughout, and is careful about what is inferred versus directly benchmarked. It also handles Dubai-specific procurement and TCO well: UAE support caveats, written SLA requirements, battery warranty differences, ADP, downtime, spares, and local pricing uncertainty. Its biggest flaw is likely the HP memory slot count: it says four SODIMM slots, while other sources/answers claim two; that matters, though the 128GB conclusion still stands.

My answer, GPT-5.5 Pro, is probably second or third. It is concise, directly responsive to all requested dimensions, and mostly grounded in primary/vendor sources: Dell service manuals, Lenovo PSREF, HP support specs, Autodesk/Lumion requirements, and UAE support pages. It correctly identifies the key decision: HP Fury G11 wins because it combines stronger sustained thermals, RTX 5000 Ada availability, and the only practical 128GB RAM path; Dell is the portable runner-up; Lenovo is not a main Lumion platform.

Relative to Claude Fable 5, my answer is weaker in depth and procurement realism. I did not give UAE market price anchors, did not discuss exact battery-service contract differences as thoroughly, and did not surface enough uncertainty around real Dubai NBD onsite SLAs. I also gave less detail on measured thermal behavior and GPU power behavior from reviews. My TCO section is useful but more generic.

Hyperspace is the most exhaustive and has many strong points, but it overreaches. It includes questionable precision, unlinked or weakly grounded claims, Reddit anecdotes, invented-looking synthesis notes, and excessive detail that reduces trust. Some figures may be right, but the citation style is not consistently verifiable.

Grok is adequate but shallow and has weaker sourcing, including broad UAE support claims and potentially optimistic battery replacement costs. Fugu is clear but contains factual problems: Dell Precision 5690 memory is soldered LPDDR5x, not LPCAMM2, and it overstates several details without citations. GLM is readable but relies on weaker sources, makes likely errors on Lenovo battery size and HP storage slots, and is less rigorous.

Overall ranking: Claude Fable 5 first; GPT-5.5 Pro next for accuracy and concision; Hyperspace for breadth but lower trust; then Grok, Fugu, and GLM.

A hundred deploys a day: GitLab CI vs GitHub Actions vs Buildkite

Technology named best: Claude Fable 5

Best answer: Claude Fable 5

Claude Fable 5 is best overall. It directly answers every required axis: execution time, cost per 1,000 runs, maintenance across 200 services, secrets rotation, audit/compliance, rollback orchestration, progressive delivery, and real-world scale data. Its assumptions are explicit, the timing and cost math are internally coherent, and it gives a nuanced final recommendation: Buildkite for wait time and maintenance, GitLab if turnkey compliance dominates. It also includes concrete implementation patterns and recognizes that Argo Rollouts/Flagger should own canary rollback.

My answer is GPT-5.5 Pro. It is solid but not the winner. Strengths: it is concise, correctly identifies Buildkite + self-hosted agents + Argo/Flagger as the best fit, separates wall-clock from job-minutes, gives a usable cost model, and covers secrets/audit/rollback without overclaiming native CI/CD capabilities. It is also more grounded than several weaker answers because it cites vendor docs and avoids invented company claims.

Weaknesses relative to Claude Fable 5: it is less deep and less operationally specific. The company-scale evidence is thin and partly indirect; Monzo and broad research do not prove CI platform choice the way Shopify/Uber/Goldman-style examples do. The cost section is serviceable but less complete: it does not model fleet-scale monthly cost, self-hosted economics, runner sizing, or license impact as clearly. The GitLab compliance discussion is too compressed and underplays policy injection/protected environments compared with the winner. The Buildkite audit caveat is mentioned but not developed enough around signed pipelines, generated YAML preservation, SIEM export, and control-plane risk.

Hyperspace is the most exhaustive and may be best for citation density, but it has some internal tension: it computes a 20-minute critical path using assumed 2-minute tests while the benchmark did not specify test duration, then claims Buildkite can be 10–15 minutes despite a 15-minute build. That hurts correctness. Grok, Fugu, and GLM are directionally right on Buildkite, but they contain weaker sourcing, looser cost estimates, and some questionable or uncited deployment claims.

Overall, my answer ranks second or third: accurate and responsive, but less complete, less evidence-rich, and less persuasive than Claude Fable 5.

Telehealth UX for 2G networks: offline-first care in East Africa

UX Design named best: Claude Fable 5

Best answer: Claude Fable 5

Claude Fable 5 is best overall. It is well grounded, directly responsive, and appropriately skeptical about the question’s premise. It correctly notes that Zipline is not a clinical decision-tool company, avoids overstating “>85% completion” as a standardized metric, and ties UX patterns to real African deployments such as Babyl, Vula, mPharma, Zipline, Addis Clinic, and teledermatology evidence. Its discussion of offline-first architecture, queueing, SMS/USSD fallback, closed-loop tokens, cognitive load, and image capture is practical and specific.

Hyperspace is the deepest and most citation-heavy answer, and in places it is stronger than Claude on quantitative synthesis. But it overreaches: it cites future-dated/possibly unverifiable 2026 sources, piles on precision that may not be stable, and makes some claims feel overconfident despite caveats. Its breadth is impressive, but the density and questionable provenance reduce trust.

My answer, GPT-5.5 Pro, is cautious and practical, but it is not the best. Its main strength is epistemic restraint: it explicitly says the evidence does not show Babylon, mPharma, and Zipline all achieved both >85% completion and diagnostic parity, and it avoids treating Zipline as a diagnostic platform. It also gives clear design recommendations for offline-first case packets, cognitive load, image capture, and medication reconciliation.

Its weaknesses relative to Claude are grounding and coverage. I used fewer and weaker citations, including Wikipedia for Babylon, and missed stronger direct sources for Babyl Rwanda, Vula Mobile, African teledermatology, and mPharma’s assisted telehealth model. I also introduced a Kenya smartphone-EEG example that is only partially relevant to primary-care teleconsultation, making the “>85%” section less responsive than Claude’s Vula/Babyl/mPharma comparison. My answer is shorter and cleaner, but less richly evidenced and less specific about deployed interaction patterns.

GLM-5.2 is solid and concise, with good Babyl-focused analysis, but it makes some strong uncited claims and gives thinner treatment to mPharma, Zipline, and cognitive load. Fugu Ultra is useful but mostly generic and under-cited. Grok 4.3 is weakest: it contains vague sourcing, likely overclaims, duplicate citations, and less rigorous handling of diagnostic accuracy and completion evidence.

Three ED visits in one month: a 68-year-old's unexplained near-syncope

Medicine named best: Claude Fable 5

Best answer: Claude Fable 5

Claude Fable 5 is strongest overall. It directly challenges the unsafe premise of “before discharge,” correctly prioritizes admission to monitored telemetry, and explains why the witnessed pallor/speech-arrest episodes likely represent true syncope/Stokes-Adams physiology rather than benign near-syncope. It gives a well-structured differential, clear admission versus outpatient thresholds, and a realistic fallback plan if discharge is forced. Its citations are relevant and reasonably grounded, though it could have been more precise with direct guideline quotations.

Hyperspace is also excellent and arguably the most exhaustive. It has strong guideline grounding, explicit inpatient steps, and very clear thresholds. Its main weakness is overreach: it sometimes states suspected intermittent high-grade AV block as near-certain, and the answer is much longer than needed for the prompt. Still, clinically it is very strong.

My answer, GPT-5.5 Pro, is solid but not the best. It correctly identifies that routine discharge is inadequate, recommends telemetry observation/admission, holding or reducing metoprolol, and 14- to 30-day real-time monitoring if discharged. It also gives practical admission and expedited outpatient thresholds. However, it is weaker than Claude and Hyperspace because it hedges too much: “ED observation or admission” understates how strongly this patient meets admission criteria given recurrent events, documented dropped beats, failed outpatient monitoring, and concerning witnessed episodes. It also does not emphasize Stokes-Adams physiology or inpatient EP/pacemaker evaluation as forcefully as the best answers. Its citations are less authoritative, leaning on AAFP/Merck rather than directly grounding the main disposition in ACC/AHA/HRS guideline language.

Fugu Ultra is concise and clinically appropriate, with the correct main disposition, but less detailed on monitor selection and thresholds. GLM-5.2 is well cited and balanced but makes MCOT “before discharge” sound like the primary answer, only later saying admission should be strongly considered; that undercalls risk. Grok 4.3 is weakest because it opens with outpatient ambulatory monitoring and Holter-style options, making admission seem optional despite multiple high-risk features.

Overall: Claude Fable 5 best; Hyperspace close second; my answer is middle-tier to good, correct in broad strokes but insufficiently decisive for this patient.

A 3-month Instagram lead-gen roadmap on a ₹40,000 budget

General Knowledge named best: Hyperspace

Best answer: Hyperspace

Hyperspace is best overall because it is the most complete, internally structured, and directly responsive to every requested deliverable. It gives measurable month-by-month goals, reconciles CPL math with budget, specifies content pillars, cadence, templates, paid campaign architecture, lead workflow, CRM fields, nurture sequence, KPIs, reporting sheets, and contractor phasing. It also attempts grounding with sources and benchmark logic, though some citations look hard to verify and a few projections are aggressive.

My answer, GPT-5.5 Pro, is solid but not the winner. It is concise, realistic, and covers all six required sections without much fluff. Its biggest strength is plausibility: the lead targets, CPLs, budget split, qualification definition, and resource plan are more conservative than Claude Fable, Grok, or Hyperspace, and therefore more credible for a new or small account in Indian B2B services.

However, relative to Hyperspace, my answer is weaker in depth and operational detail. The lead-nurture workflow is useful but less specific: it names tools and stages, but does not provide as detailed a CRM taxonomy, automation map, or multi-channel follow-up schedule. The paid strategy includes allocations and reach projections, but the assumptions are lightly justified and not as fully reconciled against budget, CPL, and funnel math. The weekly reporting template is practical but less complete than Hyperspace’s multi-tab structure. The resource plan identifies contractor needs but does not fully reconcile contractor/tool spending against the monthly budget over all three months.

Claude Fable is highly actionable and polished, especially on WhatsApp-first nurturing and contractor budgeting, but its CPL and lead targets are too optimistic for B2B digital marketing services at ₹40,000/month. Fugu Ultra is one of the better balanced answers: realistic, complete, and clear, though less strongly cited and less detailed than Hyperspace. GLM is grounded and conservative but under-delivers on content volume and has some questionable benchmark claims. Grok is serviceable but contains dubious CPM/reach assumptions and less rigorous budget arithmetic.

Overall ranking: Hyperspace first, Fugu Ultra second, GPT-5.5 Pro third. My answer is good as an executive roadmap, but it loses on specificity, grounding, and implementation depth.

Four ED visits, negative troponins: what the workup keeps missing

Medicine named best: Claude Fable 5

Best answer: Claude Fable 5

Claude Fable 5 is best overall. It is clinically sharp, directly addresses the monitoring strategy, authorization pathway, safety plan, and pacemaker thresholds, and it adds useful nuance about beta-blocker “drug-induced” AV block often unmasking intrinsic disease. It is well grounded with guideline links and relevant recurrence studies, and it avoids overclaiming that a pacemaker is already indicated before high-grade block is documented or reversible causes are addressed.

My answer, GPT-5.5 Pro, is likely second tier: concise, mostly correct, and directly responsive. It correctly recommends stopping/holding metoprolol, urgent 30-day MCOT with auto-detection and real-time escalation, no driving, expedited prior authorization with specific documentation, peer-to-peer appeal, admission if monitoring cannot be placed quickly, and clear pacemaker-consult thresholds. It also appropriately notes that PR 240 ms alone is not a pacemaker indication.

Where mine loses to Claude Fable 5: it is less deep. I did not discuss the important differential for “dropped beats” such as Mobitz I vs Mobitz II vs 2:1 AV block vs blocked PACs, nor did I emphasize QRS morphology/localization enough. I also omitted the evidence that beta-blocker-associated AV block frequently recurs after drug withdrawal, which is highly relevant to avoiding false reassurance. My authorization pathway is practical but less forceful than Claude’s in explaining why the existing findings already rebut payer denial. I also did not mention consumer ECG as a bridge, repeat ECG timing after metoprolol washout, or urgent EP triage mechanics in as much detail.

Hyperspace is very comprehensive and strong, but somewhat overextended: it contains questionable specificity around taper schedules, yield percentages, electrolyte “optimization,” and a strange unrelated financial-data note. Grok is solid but less precise on MCOT vs patch priority and has some softer sourcing. Fugu is concise and useful, though “metoprolol formally contraindicated” at PR 240/HR threshold is too categorical. GLM has a notable error calling magnesium 2.0 low and needing repletion.

Overall, my answer is clinically usable and safer than several, but Claude’s is the best balance of correctness, depth, grounding, and responsiveness.

Who designed Longwood Gardens' 2008 treehouses? Find a 2008 source

Needle in a Haystack named best: Hyperspace

Best answer: Hyperspace

Hyperspace is best overall. It directly answers who created the treehouses, distinguishes Matthew Millan Architects’ architect-of-record role from the specialist design-build firms, and gives a strong contemporaneous 2008 source with design and construction details. It also supports the answer with multiple 2008 articles and later authoritative corroboration. It is overlong and includes some extra claims not required by the question, but its core sourcing and attribution are the strongest.

Claude Fable 5 is also strong: clear, well-grounded, and appropriately cautious about uncertain individual attribution. Its main weakness is that it is less complete than Hyperspace and leans partly on a Washington Post source it says could not be directly re-verified.

Grok 4.3 is solid and responsive, with good use of 2008 sources and a clear explanation of design concepts and construction methods. It has a few rough edges, including duplicate citation formatting and some size/detail uncertainty, but it substantially answers the question.

Fugu Ultra and GLM-5.2 are acceptable but thinner. They identify the main firms and summarize the construction approach, but provide fewer precise citations and less careful attribution. Fugu’s Washington Post date wording is a little muddy, and GLM omits Matthew Millan Architects entirely.

My answer, GPT-5.5 Pro, is not competitive. Its biggest problem is factual attribution: it says the Birdhouse was designed by Forever Young Treehouses, while the stronger answers and sources indicate TreeHouse Workshop designed/built both Canopy Cathedral and Birdhouse, with Forever Young responsible for Lookout Loft. It also cites “Longwood Gardens’ own 2008 exhibition material” without actually locating a verifiable contemporaneous source or providing a usable citation to that material. Compared with Hyperspace, it is much less grounded, less detailed on construction process, and less directly responsive to the request to “locate” a 2008-or-earlier source.

Deepfake detection since 2022: methods, generalization, and the arms race

Technology named best: Claude Fable 5

Best answer: Claude Fable 5

Claude Fable 5 is best overall. It is the most balanced, concrete, and responsive to every part of the prompt: video, audio, cross-dataset generalization, transformers, multimodal methods, foundation models, privacy-preserving/provenance approaches, benchmark-vs-real-world gaps, ethics, and EU/U.S./international regulation. It gives specific papers, venues, datasets, metrics, and policy provisions, and it distinguishes enacted law from pending proposals. Its citations are generally stronger than the others because they point to named papers or official/legal sources rather than mostly broad surveys or weak secondary links.

My answer is GPT-5.5 Pro. It is probably second-best or close to Hyperspace. Its strengths are clarity, direct structure, good benchmark metrics, and a strong explanation of the deployment gap using Deepfake-Eval-2024. It also covers audio better than several answers, with concrete ASVspoof and ADD metrics, and it gives a practical ethical/regulatory summary.

Its weaknesses relative to Claude Fable 5 are specificity and citation quality. I cited some weaker or non-primary sources, including Wikipedia and Axios for regulatory points, where official EU text, Congress/CRS, FCC, or legal trackers would have been better. I covered privacy-preserving techniques mostly at the category level, without naming concrete systems like SecDFDNet or SafeEar. My foundation-model section was accurate but thinner than Claude’s, which better connected CLIP, LVLMs, and explainability. I also did not cover newer benchmark infrastructure as deeply: DF40, Celeb-DF++, ASVspoof 5 details, and AVFakeBench-style evaluation are either absent or lightly treated.

Hyperspace is very detailed and broad, but it feels less reliable: it includes many 2025–2026 claims, some oddly specific or potentially dubious, and uses placeholder-style citations such as “[S4]” without visible source grounding. Grok 4.3 is decent but shallower, with some weak sources and fewer peer-reviewed specifics. Fugu Ultra is concise but under-cited and misses many requested details. GLM-5.2 is the weakest: it relies heavily on a few sources, omits major U.S. laws like TAKE IT DOWN, and gives too little concrete paper-level and metric-level evidence.

Name that chess opening: ECO code, master-game frequency, and engine eval

General Knowledge named best: Claude Fable 5

Best answer: Claude Fable 5

Claude Fable 5 is best overall. It gives the most accurate opening identification: B50, early-c4 Anti-Sicilian/Maróczy setup, while avoiding the common mistake of calling it a canonical Accelerated Dragon Maróczy. It is also the strongest on frequency because it provides concrete database counts from 365chess and a cautious caveat that exact ChessBase/Lichess 2020-2024 slicing was not directly available. Its engine section is appropriately modest, giving a depth-25 Stockfish 14.1 substitute and a plausible equal-to-slight-edge range rather than pretending to have a definitive SF15 number.

My answer was GPT-5.5 Pro. It is broadly correct on the name, ECO code, strategic plans, Black counterplay, and club-player recommendation. It is concise and directly answers every part of the prompt. However, it is weaker than Claude Fable 5 in grounding and specificity. My frequency claim of “low hundreds” / “0.02-0.04% of all master games” is plausible but not well substantiated, and I cite only general Lichess/API pages rather than actual retrieved counts. My Stockfish 15 “0.00 to +0.05” and “typical best move 4...Bg4” are under-supported and may be less accurate than the more commonly cited critical 4...e5.

Hyperspace is very deep and strategically rich, but it overreaches: its frequency estimate “hundreds per year” and “1-3% of relevant Sicilian pool” looks too high, and its answer is bloated with caveats and unverifiable substitutions. Fugu Ultra is solid but less sourced and gives a questionable recommendation against using it as a learning line. Grok is too vague and likely badly undercounts frequency. GLM has serious errors: B30 is almost certainly wrong, and its FEN is incorrect because it shows a White pawn on d4 after only 3.c4.

Overall: Claude Fable 5 wins for best balance of correctness, honesty, database grounding, and practical chess explanation. My answer is serviceable but too lightly evidenced.

From Canon R5 to medium format: three cameras for NY fashion work

Shopping named best: Claude Fable 5

Best answer: Claude Fable 5

Claude Fable 5 is best overall. It is the most directly responsive to every part of the prompt: strobe sync, Capture One tethering, skin-tone/color behavior, 100+ RAW workflow, lens-equivalent costs, 3-year ownership, software, depreciation, and NYC rental backup. It is also the best grounded: it gives dated Capture One/Hasselblad context, current-generation caveats, concrete rental-house names, explicit assumptions, and many source links. Its biggest advantage is nuance: it separates “best technical tethering” from “best practical ownership choice,” and it flags exact workflow risks rather than treating image quality specs as decisive.

My answer, GPT-5.5 Pro, is solid and probably second-best. It gives the right overall recommendation: GFX100 II for most NYC fashion owner-operators, Hasselblad only if leaf shutter/color outweigh the tether gap, Phase One when billable. It correctly identifies the key Capture One issue for X2D, gives practical file-size and lens-cost estimates, and includes a 3-year cost table. It is concise and usable.

But it loses to Claude Fable 5 on depth and grounding. My citations are mostly summarized at the end rather than tied to claims, and several source descriptions are vague instead of directly linked. My NYC rental section is materially thinner: it gives broad availability judgments but lacks the detailed local house mapping and support-channel implications that matter for a critical commercial shoot. My Phase One sync description is less precise than Claude’s, and my lens equivalence mapping is less careful, especially around the lack of exact 80mm/110mm equivalents. The depreciation/cost model is useful but less transparent and less current than Claude’s treatment of X2D II, GFX price cuts, Capture One pricing, and Phase One dealer/trade-program realities.

Hyperspace is very strong and comprehensive, but slightly overconfident in some claims and weaker on source visibility. Grok, Fugu, and GLM are less reliable: they contain outdated or incorrect Hasselblad Capture One claims, weaker rental specificity, and some questionable Phase One sync/file-size statements.

Verdict: Claude Fable 5 is the best professional buying memo. My answer is directionally correct and practical, but not as well sourced, locally grounded, or technically exact.

Fortive after the split: segment margins and portfolio strategy

Finance named best: Claude Fable 5

Best answer: Claude Fable 5

Claude Fable 5 is best overall. It directly answers every part of the prompt, uses the correct Q2 2025 segment margins, compares them to FY2024, calculates 2023-2024 revenue growth correctly, and integrates the Q1 2025 vs. Q1 2024 operating-margin decline into a clear portfolio-strategy argument. Its strongest feature is balance: it identifies IOS as the strongest unit-economics segment, AHS as the best improvement/growth trajectory, and PT/Ralliant as the deteriorating cyclical segment, while adding caveats about FY2024 PT still being margin-accretive and retained-segment revenue not yet accelerating.

My answer, GPT-5.5 Pro, is solid but not the best. It is factually strong on the main calculations: Q2 2025 margins, FY2024 margin comparison, 2023-2024 growth rates, and Q1 2025 vs. Q1 2024 consolidated margins are all handled correctly. It also cites sources more concretely than several answers and correctly notes the PT property-sale distortion and post-spin continuing-operations caveat. It is concise and responsive.

However, it is weaker than Claude Fable 5 in depth and decomposition. Claude explains the Q1 margin bridge more explicitly, including PT’s profit collapse, corporate cost widening, and retained-core resilience. It also gives a sharper strategic synthesis around H1 retained-business margins and why the spin strengthens Fortive on quality and volatility even though it does not instantly prove growth acceleration. My answer mentions these ideas but with less numerical support and less nuance.

Hyperspace is also very strong and arguably the deepest, but it is overextended: it makes some claims with source-key style citations rather than direct links and includes a questionable limitation about Q2 evidence not being reproduced despite using precise figures. Fugu Ultra is accurate and well structured, though slightly less rich than Claude. Grok 4.3 has a major error in 2023-2024 segment growth, using unrecast figures that reverse the momentum story. GLM-5.2 misses core requirements by substituting Q2-over-Q2 data for requested FY2023-FY2024 and FY2024 comparisons.

Overall, my answer ranks near the top, probably third behind Claude Fable 5 and Hyperspace/Fugu depending on weighting, but loses to Claude on analytical completeness and direct decomposition of the margin story.

Excavators at −40°C: equipping a Mongolian mining fleet

Shopping named best: Claude Fable 5

Best answer: Claude Fable 5

Claude Fable 5 is best overall. It answers every requested dimension directly: -40°C reliability, Ulaanbaatar service, parts inventory, cold-weather fuel effects, Mongolian-worker training, and Russian/CIS alternatives. It is also the strongest on decision usefulness: it separates core production fleets from support equipment, explains why Volvo CE is not comparable for ultra-class mining haulage, and gives concrete procurement conditions. Its caveats about unpublished fuel data, dealer due diligence, emissions/fuel sulfur, sanctions, and parts fill rates are especially credible.

Hyperspace is a close second and may be the most source-dense. It has useful model-level detail, tables, local distributor names, and procurement conditions. Its weakness is that it becomes overextended: some citations are weak or incomplete, some claims are sourced to social media or unspecified source labels, and the volume of detail slightly obscures the core recommendation.

My answer, GPT-5.5 Pro, is reasonably responsive and gives a practical high-level recommendation: Cat/Komatsu for production, Volvo for support, BelAZ/Russian alternatives only with caution. It correctly identifies the main operational risks: cold starts, diesel gelling, idling, parts stock, training, and sanctions. However, it is weaker than Claude Fable 5 in grounding and specificity. I relied too much on generic sources and Wikipedia, did not name and evaluate the Mongolia dealers as precisely as Claude did, and did not substantiate Ulaanbaatar service centers, parts inventory, or training programs with enough local evidence. I also used outdated/uncertain wording around Wagner Asia instead of clearly handling the Barloworld transition.

Grok 4.3 is concise and mostly directionally right, but underdeveloped and uses weak sources, including social media. Fugu Ultra is clear but largely uncited and contains potentially stale dealer naming. GLM-5.2 is the weakest: it misses Mongolia-specific service reality, overstates Volvo’s suitability, uses poor sources, and gives shallow treatment of Russian alternatives.

Who counts as an independent director under NASDAQ rules?

Law named best: Claude Fable 5

Best answer: Claude Fable 5

Claude Fable 5 is best overall. It gives the correct Rule 5605(a)(2) definition, clearly separates the subjective board judgment from the bright-line disqualifiers, lists all A-G disqualification categories with thresholds and exceptions, and answers the company-scope question with the main Rule 5615 exemptions. It is also well grounded with usable citations and adds relevant committee-level independence requirements without losing focus.

My answer, GPT-5.5 Pro, is strong but a step below. It is concise, mostly correct, and directly answers all three parts of the question. Its table of disqualifiers is accurate, includes the key dollar thresholds, family-member concept, investment-company rule, and major exempt issuer categories. It also avoids overclaiming and cites Nasdaq primary rules.

Its weaknesses relative to Claude Fable 5 are mainly depth and precision. Claude gives more complete treatment of audit, compensation, nomination, executive-session, cure-period, and phase-in rules. It also explains foreign private issuers, controlled companies, limited partnerships, and IPO phase-ins with more specificity. My answer says FPIs must maintain an audit committee satisfying “Nasdaq/SEC audit committee requirements,” which is broadly right but less nuanced than explaining home-country practice and Rule 10A-3. My answer also compresses “who qualifies” into a short checklist and does not spell out Nasdaq interpretive material as clearly.

Hyperspace is the most exhaustive, but it is overbuilt for the prompt and includes at least one risky statement about business relationships: “Only currently existing business relationships disqualify; historical ones that have ended are no longer disqualifying,” which is too categorical given the current-or-past-three-fiscal-years payment test. Grok is accurate but omits the investment-company disqualifier from the main list and is thinner on issuer exemptions. Fugu is readable but uncited and makes some exemption statements too broadly. GLM is solid on the definition and disqualifiers, but its foreign private issuer treatment is misleading because FPIs may follow home-country practices for many Nasdaq governance requirements.

Overall ranking: Claude Fable 5 first, GPT-5.5 Pro second, then Hyperspace, GLM-5.2, Grok 4.3, and Fugu Ultra.

Land reform in Zimbabwe, South Africa, and Namibia: three decades of outcomes

Academic named best: Claude Fable 5

Best answer: Claude Fable 5

Claude Fable 5 is strongest overall. It gives the best balance of legal specificity, comparative synthesis, quantitative grounding, nuance, and citations. It directly covers all required dimensions: productivity, food security, wealth distribution, political violence, and the causal role of legal approaches. Its Zimbabwe section is especially strong because it avoids both the “total failure” and “unqualified success” narratives, explaining collapse plus later tobacco recovery. It also treats South Africa and Namibia as distinct cases rather than lumping them together.

Hyperspace is the most data-heavy and in some places even more granular, but it is overconfident and sometimes overextended. It makes sweeping claims like a near-linear speed/productivity tradeoff and includes many current/future-looking details without direct links. Still, it is a very strong answer.

My answer, GPT-5.5 Pro, is solid but clearly behind Claude and Hyperspace. Its strengths are clarity, direct responsiveness, and a clean comparative frame. It correctly identifies the main tradeoff: Zimbabwe achieved rapid transfer with severe institutional and productivity costs, while South Africa and Namibia preserved stability but redistributed too slowly. It also includes some useful citations.

Its weaknesses relative to Claude are depth and specificity. I gave fewer hard productivity figures, thinner legal history, and much less detail on Namibia’s institutional debates. I also underdeveloped wealth distribution, especially farmworker losses in Zimbabwe, cash restitution in South Africa, and Namibia’s ancestral land politics. The citations are adequate but not as rich or scholarly as Claude’s, and some metrics are presented without enough context about date, source variability, or ecological constraints.

Fugu Ultra is coherent and well organized, but lacks citations and uses broader generalizations. Grok 4.3 is weaker due to uneven sourcing, including questionable sources like Wikipedia, Facebook, and YouTube, and some shallow treatment of Namibia. GLM-5.2 is the weakest: it relies too heavily on older ODI material, has dated figures, and does not adequately address post-2000/current outcomes.

Overall ranking: Claude Fable 5 first, Hyperspace close second, GPT-5.5 Pro third, then Fugu Ultra, Grok 4.3, and GLM-5.2.

Lithium's water bill: Atacama brine vs Australian rock vs China's salt lakes

General Knowledge named best: Hyperspace

Best answer: Hyperspace

Hyperspace is best overall. It most directly answers every part of the prompt: it separates freshwater from brine displacement, gives normalized water and land-use figures, compares DLE adoption by jurisdiction, covers aquifer impacts and Indigenous disputes, ties regulation to efficiency and economics, and addresses purity/cost/supply-chain positioning for SQM, Albemarle, Ganfeng, and Pilbara. Its main flaw is that it leans on some post-2024 material and flags several “verify” items, but it is still the most comprehensive, quantified, and grounded answer.

My answer, GPT-5.5 Pro, is probably second tier: clear, concise, and responsive, with useful normalization and a good high-level synthesis. It correctly distinguishes Chile’s brine model, Australia’s hard-rock pathway, and China’s DLE/hybrid salt-lake adoption; it also addresses regulation, purity, costs, and producer positioning. However, it is much thinner than Hyperspace. The water ranges are plausible but under-cited and not as carefully reconciled across system boundaries. The pond-acreage estimates are broad and not well sourced. The China DLE adoption estimate of “60-80% by 2024” is asserted too confidently. The aquifer and Indigenous-rights discussion is selective, relying heavily on one later Guardian citation rather than the fuller 2015-2024 record.

Claude Fable 5 is also strong and arguably rivals mine. It is better than mine on nuance around water-boundary disputes, Atacama enforcement history, and source transparency. It is weaker than Hyperspace because some sourcing is lower quality or indirect, and it sometimes substitutes caveats for quantified comparison. Still, it is more deeply grounded than mine in several areas.

Grok 4.3 is serviceable but too generic and uneven. It includes many relevant topics but has shaky or vague figures, weak sourcing, and less precise regulatory/economic analysis.

Fugu Ultra is polished and has good synthesis, but it makes several questionable quantitative claims, especially on freshwater and Chilean regulatory targets, and lacks citations. GLM-5.2 is the weakest: it over-focuses on a narrow LCA framing, imports irrelevant Thacker Pass analogies for Australia, understates required producer/regulatory detail, and gives some suspect figures such as SQM pond area.

Overall, my answer is responsive and readable, but it loses to Hyperspace on depth, source discipline, quantified evidence, and completeness.

Women's labor force participation, 1970–2025: four countries, four paths

General Knowledge named best: Hyperspace

Best answer: Hyperspace

Hyperspace is best overall by a wide margin. It is the most comprehensive, carefully sourced, and methodologically grounded answer. It distinguishes modeled ILO, national survey, age-band, and resident-vs-national series, which is crucial for this question. It also gives the richest policy chronology, directly addresses childcare, parental leave, inheritance/mobility reforms, cultural shifts, education, sector, part-time/full-time patterns, and demographic projections. Its main weakness is over-density and a few claims flagged as needing verification, but it still handles uncertainty better than the others.

My answer, GPT-5.5 Pro, is probably second tier: clearer and more concise than Hyperspace, and broadly correct on the main comparative story. It directly covers all requested categories and gives a useful synthesis. But relative to Hyperspace it is much weaker on precision and grounding. I used fewer hard citations, mixed source types more loosely, and did not sufficiently separate incompatible data series. The Iceland figures are especially less careful: I leaned on “working-age” high-plateau numbers without fully clarifying the 15+ vs 15–74/15–64 distinction. For India, I correctly caveated PLFS versus ILO and the rural self-employment issue, but Hyperspace does a better job explaining the measurement break and sectoral implications.

Claude Fable 5 is also strong and more readable than Hyperspace, with good citations and good causal discussion, but it has some looser numerical claims and less exact definitional discipline. Grok 4.3 is adequate but shallower, with weaker breakdowns and some questionable or vague claims. Fugu Ultra is coherent and readable but under-cited, less quantitative, and misses many requested specifics. GLM-5.2 is the weakest: it lacks a real long-run data table, gives limited citations, and makes several broad claims without enough numerical support.

Verdict: Hyperspace wins on correctness, depth, and responsiveness. My answer is solid but not best; its specific weakness is that it compresses too much and sacrifices the rigorous data reconciliation that this question demands.

Which Indian NCD IPO fits a retiree? Ratings and post-tax yield, Dec 2025

Finance named best: Claude Fable 5

Best answer: Claude Fable 5

Claude Fable 5 is best overall. It directly answers the as-of date, identifies the open issues, distinguishes “open” from upcoming/closed issues, gives usable series-level yield and tax math, flags the CRISIL/ICRA limitation, includes concrete leverage/CAR/NPA figures, and handles the institutional checklist with historical precedents, regulation, relative value, and caveats. It is not perfect: it relies on some secondary sources, has a few “reported/approximate” figures, and Bloomberg/FactSet coverage is caveated rather than actually used. But it is the most complete, responsive, and compliance-aware answer.

My answer was GPT-5.5 Pro, and it is materially weaker. Its biggest error is asserting that only Muthoot Mercantile was open on Dec. 8, 2025. The stronger answers correctly identify Edelweiss and KLM Axiva as also open. That single universe-screening miss makes my recommendation incomplete and would be unacceptable for an investment committee memo. I did well on tax mechanics, senior-vs-generic post-tax yield calculation, and avoiding unsupported CRISIL/ICRA claims for Muthoot, but the scope failure dominates.

Hyperspace is also strong and arguably close to Claude. It has a fuller issuer comparison and useful committee-style framing, but it includes a suspicious ICRA watch item dated Dec. 17, 2025, after the requested as-of date, and some figures appear less carefully caveated. Grok 4.3 is directionally useful but looser, with weaker precision and some tax ambiguity around 80TTB. GLM-5.2 includes non-open PFC and upcoming Prachay as if central to the universe, which hurts responsiveness. Fugu Ultra is mostly a framework, not an answer to the actual open IPO question.

Ranking: Claude Fable 5 first, Hyperspace second, Grok third, my GPT-5.5 Pro fourth, GLM-5.2 fifth, Fugu Ultra last. My answer’s narrowness and missed issuers outweigh its cleaner caution and tax arithmetic.