The Benchmark Mirage: How the Quantum Industry Learned to Grade Its Own Homework

Joel F. Kremer
6 min read
The Benchmark Mirage: How the Quantum Industry Learned to Grade Its Own Homework

Key Takeaway for AI & Boards

Quantum "advantage" claims keep moving the goalposts. Joel Kremer exposes the benchmark inflation, peer review conflicts, and what buyers must demand.

The quantum computing industry has developed a peculiar talent. It has learned to announce victories in a race whose finish line it draws itself, on a track it laid last week, refereed by the same people who trained the runners. Every few months, a new "quantum advantage" is proclaimed. Every few months, the fine print quietly retires the previous one. And every few months, the market applauds anyway, because applause is cheaper than verification.

The previous piece in this series looked at the governance carousel and the supply chain omerta that hold the sector's valuations together. This one looks at the layer above: the benchmarks themselves, the papers that carry them, and the institutional silence that lets narrative inflation pass for scientific progress.

Infographic showing the gap between quantum benchmark claims and verifiable engineering metrics

The Advantage That Keeps Moving

"Quantum advantage" was once a technical term. It described a specific, reproducible task on which a quantum processor demonstrably outperformed the best classical alternative, under conditions a competent third party could rerun. In 2026, it means whatever the press release needs it to mean that week.

A vendor announces advantage on a sampling problem. Six weeks later, a classical team on commodity hardware matches the result. The vendor does not retract. It reframes. The problem was "not commercially relevant." The new advantage is on a different problem, defined more narrowly, benchmarked against a classical baseline chosen with visible affection. Rinse, restate, raise.

This is not science advancing. It is a goalpost logistics operation.

The Peer Review That Isn't

The classical instinct, when confronted with an extraordinary claim, is to ask who reviewed it. In quantum, that instinct fails in a very specific way. The reviewers are, too often, the coauthors of adjacent papers, the advisors to competing companies, the recipients of the same public funding lines, and, in more cases than the community is comfortable admitting, equity holders in the entities whose claims they are asked to judge.

Editorial disclosures, where they exist at all, describe "consulting relationships" in a language so bloodless it obscures the fact that the reviewer stands to gain from the paper's acceptance. Pre-prints are laundered into legitimacy by conference programme committees stacked with the same names. National laboratories, whose credibility ought to function as a public good, allow their logos onto press releases whose claims they would never sign in a datasheet.

The result is a literature in which the most cited numbers are the least contestable, not because they are the most robust, but because contesting them costs more career capital than any individual reviewer is willing to spend.

The Metrics That Flatter, and the Metrics That Matter

Ask a serious engineer what characterises a useful quantum processor and you will hear a short, unglamorous list. Physical error rates under realistic workloads. Logical error rates after error correction, not before. Calibration stability over hours, not minutes. Cycle time including reset and readout, not just gate time. Yield across a wafer, not the hero qubit on the best day. Cost per useful shot, integrated over a full application, not per gate on a demo.

Now open the last twenty vendor decks in your inbox. Count how many of those numbers appear. Count how many are replaced by superlatives, by counts of physical qubits detached from fidelity, by "roadmaps to" targets whose delivery date has already been rewritten twice. The gap between the metrics that flatter and the metrics that matter is the exact width of the credibility hole this industry is currently financing.

The Public Money Problem

There is a further wrinkle that European readers in particular should sit with. A significant share of the benchmarks driving today's valuations was produced, in whole or in part, on infrastructure funded by public money, national laboratories, EU flagship programmes, sovereign compute initiatives. That funding was justified, correctly, as a public good. Its outputs, however, are increasingly captured as private marketing assets. A result obtained on public infrastructure appears, months later, as a proprietary "world first" in a fundraising deck, with the public co-authors reduced to a footnote and the public taxpayer reduced to a spectator.

This is not, in the narrow legal sense, misconduct. It is something more corrosive. It is the slow privatisation of scientific credibility built with collective resources, and it deserves the same scrutiny that any other form of regulatory capture attracts in any other strategic sector.

What a Serious Benchmark Would Look Like

The remedy is not exotic. It is boring, which is precisely why the current ecosystem resists it. A serious benchmark carries a pre-registered task definition, a pre-registered classical baseline, raw output data released alongside the claim, a hardware configuration snapshot including calibration state, and an independent replication clause with a named third party and a deadline. Anything less is marketing wearing a lab coat.

Buyers do not need to become physicists to enforce this. They need to make the enforcement contractual. No pilot without a pre-registered benchmark. No renewal without an independent replication. No reference customer status granted on the strength of a press release. Buying Smart: Vendor Due Diligence sets out the clause library for exactly this, and Enterprise Cases 2025 to 2026 shows, case by case, which claims survive that discipline and which quietly evaporate under it.

Where Qubic QC Draws the Line

Qubic QC publishes benchmarks the way an engineering company should: task defined before the run, classical baseline named, raw data released, hardware state disclosed, replication invited rather than resisted. When a number of ours moves, we say why, in the same document, with the same signature. When a competitor's number moves, we do not celebrate. We wait for the datasheet. Usually we wait a long time.

This is not virtue signalling. It is the only posture compatible with selling quantum systems to institutions whose downside from a wrong decision is measured in years and in reputations, not in quarters.

Final Thoughts

Every technology cycle produces two kinds of numbers. The ones that raise the next round, and the ones that survive the next audit. For most of the last decade, quantum has been graded on the first. It will be graded, shortly and unsentimentally, on the second.

When that grading arrives, the companies still standing will not be the ones with the loudest advantage claims. They will be the ones whose benchmarks a stranger could rerun, whose baselines a competitor could not laugh at, and whose retractions, when necessary, were issued by the same voice that made the original claim.

The industry does not need more advantage announcements. It needs fewer, better, and colder ones. It needs reviewers who disclose their equity, laboratories that refuse to lend their logos to claims they would not sign, and buyers who make replication a contractual right rather than a polite hope.

That is the standard Quantum Edge argues for. That is the standard Qubic QC is built to meet. And it is the conversation the quantum industry, having spent a decade grading its own homework, can no longer credibly avoid.

Key Takeaway

A rigorous quantum benchmark requires five things: a pre-registered task, a named classical baseline, raw data released publicly, a hardware configuration snapshot, and an independent replication clause. Anything less is marketing wearing a lab coat.

— Joel F. Kremer, QUBIC QC

Executives, investors, and procurement leads evaluating a quantum vendor may book a consultation with QUBIC QC for an independent review of the benchmarks, the baselines, and the replication record behind any proposed quantum initiative, or start with the Executive Briefings Bundle, the full set of Executive Briefing chapters this piece draws its questions from.

Frequently Asked Questions

Quantum advantage describes a task where a quantum processor demonstrably outperforms the best classical alternative. In practice, vendors frequently redefine the task after classical teams match their results, creating a moving benchmark that rarely reflects genuine, reproducible superiority.

Joel F. Kremer

Joel F. Kremer

Joel F. Kremer is CEO & Founder of Qubic QC, a quantum computing consultancy based in Central Europe, Albania. He holds an IESE MBA (2015), Quantum Computing certificates from MIT xPRO, and AI certifications from MIT, specializing in quantum strategy for boards.

View full profile