roboticsQUANTUM ASSURANCE
/
Get a quote
← All research

Computational research note · Executed experiment

Calibration Before Quantum Claims

A reproducible control experiment for scientific classifiers.

Abstract

A classifier can rank events well while reporting unreliable probabilities. This note constructs a synthetic binary classification problem with deliberately overconfident scores, fits one temperature on a separate validation sample and evaluates an untouched test sample. Ranking and threshold accuracy remain unchanged, while probability-based metrics improve. The result motivates a calibration-aware comparison protocol for classical, quantum-inspired and quantum classifiers.

1 / A ranking is not a probability model

Scientific workflows may use a score both to select events and to estimate probabilities. These are different tasks. ROC AUC assesses ordering across classes; it does not determine whether a reported probability of 0.9 corresponds to a 90 percent positive-class frequency in comparable cases.

The proposed evaluation principle is to compare representations using a common account of discrimination, probability quality, uncertainty and cost. A quantum feature map should face the same held-out protocol as a classical model. A change in representation should not silently change what counts as a useful scientific result.

2 / A controlled synthetic construction

Two independent standard-normal variables define a logit z = 1.2x₀ − 0.9x₁ + 0.7x₀x₁. Labels are Bernoulli draws with probability sigmoid(z). The raw classifier uses 2.5z, deliberately preserving the ranking while exaggerating confidence. This is a constructed scoring rule, not a trained neural network or a particle-detector model.

Using seed 20261012, the program generates 4,000 validation examples followed by 20,000 independent test examples. It minimizes validation negative log-likelihood over 901 positive temperatures from 0.5 to 5.0. The selected temperature is fixed before test evaluation. Positive temperature scaling divides the raw logit by T and is a published calibration method.

Scientific context: [1]

3 / Executed test-set results

The validation sample selects T = 2.635. The table below is computed only on the separate 20,000-example test sample. Lower is better for negative log-likelihood, Brier score and binary ECE.

MetricRaw scoresCalibrated scores

ROC AUC

0.8126850.812685

Accuracy at p = 0.5

0.7424500.742450

Negative log-likelihood

0.6583480.522470

Brier score

0.1946140.174826

Binary ECE · 15 bins

0.1293260.013254
Reliability curves for raw and temperature-scaled synthetic probabilities.
Figure 2. Fifteen equal-width probability bins. Each point pairs a bin’s mean predicted probability with its observed positive-class fraction. The diagonal is the calibration reference; no confidence bands are plotted.

4 / What the control experiment establishes

Dividing logits by a positive temperature is monotonic, so the score ranking and ROC AUC stay the same. The decision boundary at probability 0.5 also stays the same. Negative log-likelihood and Brier score can nevertheless change substantially because they evaluate the probabilities themselves. The observed contrast is the purpose of this control experiment.

Binary expected calibration error is reported using fifteen equal-width probability bins, weighting the absolute gap between each bin’s mean prediction and positive-class fraction by its test-sample share. ECE depends on binning and finite sample size, so the table also reports unbinned proper scoring rules. The reliability plot is a descriptive point estimate from one test sample.

5 / A comparison protocol for scientific quantum AI

CERN’s public QML research context includes comparisons with classical methods in precision, accuracy, performance and power consumption. The independent proposal here adds explicit probability calibration to that comparison, particularly when a classifier’s output enters a probabilistic inference workflow.

A next study would compare a classical baseline and a small quantum classifier on a shared task, keeping data splits and evaluation criteria fixed. It would measure AUC, proper scoring rules, calibration, sampling error, encoding cost and behavior under prespecified distribution shifts. Physical constraints should be checked where the chosen scientific task supplies them. The present experiment tests the evaluation logic; no quantum classifier was executed.

Scientific context: [2]

6 / Reproduce and inspect

The public script specifies the data-generating equation, split sizes, temperature grid, seed and metric definitions. It produces the full JSON record, CSV metrics and reliability figure. Keeping the validation and test samples separate makes the result useful as an inspectable calibration control for a larger research workflow.

python -m pip install numpy==2.3.5 matplotlib==3.10.8
python reproduce_research.py --output-dir results

References

  1. [1] Guo, Pleiss, Sun & Weinberger (2017) — On Calibration of Modern Neural Networks ↗
  2. [2] CERN QTI — Applications of quantum machine learning ↗

Author

Maurizio Viviani

Independent research · Robotics

Develop the next experiment.

Discuss scientific benchmarking, hybrid computing or a reproducible evaluation study.

Discuss a research collaboration