⚛️ The Open Quantum Challenge is live. Both seasons open today.
Quantum computers are expensive, queued, and reachable by only a few. So we removed the barrier. With a single GPU and a public harness, anyone can contribute to core quantum-computing problems with no quantum hardware at all. That is the point of this challenge: to move quantum research from a handful of labs to everyone.
• Season 1, Quantum simulation: reproduce a target quantum system on a GPU. • Season 2, QEC decoder: correct errors on coded quantum states.
Answers are held privately and every submission is auto-scored against a frozen ground truth, so the ranking is reproducible and hardware-independent.
🏆 Prize: 2,000 USD total (1,000 USD per season, 🥇600 / 🥈300 / 🥉100)
Getting started takes one step: copy the participation guide from the Space and paste it into Codex or Claude Code, then submit against the harness.
🏆 Darwin-27B-ZTC-v2 just took #1 on the System One Mosaic Benchmark (S1MB).
S1MB compares 102 models across 137 specialized benchmarks, in three task types: Noul (assess a condition), Choice (select an option), Score (rate on a scale). Ranking is by overall Borda score.
📊 Top of the board 🥇 Darwin ZTC v2 (FINAL-Bench) 89.58 🥈 OpenJev-27B 87.50 🥉 AutoJev-27B 87.07 4️⃣ Eikos 27B 85.43 5️⃣ Jev 1.13 85.05
🔎 Ranks 2 to 5 are all the JEV family (TypeSafe AI's System One model, from ex-OpenAI researchers). S1MB exists to compare these System One judges, so leading it is the headline.
⚙️ Why a zero-token judge wins here 🔹 It does not generate. It reads the input and typed questions and returns a calibrated distribution in a single forward pass. 🔹 Zero generated tokens, no decoding loop, so latency and cost stay low. 🔹 Holds up out of distribution too: General Noul 96.00, General Choice 99.34.
It is also #1 on the typed-decisions leaderboard (0.743, zero-shot). Same message from both: a deterministic, calibrated judge at one forward pass per call.