roboticsQUANTUM ASSURANCE
/
Get a quote
← All research

Research proposal · Evaluation protocol

An Evidence Contract for Hybrid Scientific Workflows

A proposed evaluation protocol across classical and quantum execution.

Abstract

Hybrid scientific computing joins data preparation, classical processing, quantum execution and interpretation. This proposal defines a compact evidence record that keeps the scientific objective, artifact identity, execution conditions, uncertainty and complete resource accounting together. Its research hypothesis is that a common result contract can make heterogeneous executions comparable without requiring identical numerical outputs or identical hardware.

1 / Compare scientific results, not just jobs

A completed job is an execution event. A useful scientific result is an answer to a defined question with an interpretable uncertainty and a record of how that answer was obtained. In a mixed CPU, GPU and quantum workflow, the distance between these two objects can grow: preprocessing, compilation, queues, sampling and post-processing each change the final account.

The proposed contract makes the scientific objective and acceptance conditions explicit before execution. A task should state what observable or predictive quantity it returns, which constraints must hold, what precision is required and which reference calculation will be used. The contract is an evaluation interface, not a replacement for the scientific model.

2 / The minimum evidence record

The record should identify the input data or synthetic generator, model definition, parameters, random streams, code version and numerical environment. It should also identify the returned artifact, metric definitions, uncertainty estimate, stage timings and scientific acceptance decision. These fields let another researcher distinguish a changed result from a changed experiment.

An artifact hash establishes whether the inspected bytes match a recorded artifact. It does not establish that the model is correct or the computation scientifically valid. That judgment requires numerical checks, suitable references and the relevant domain constraints. Keeping identity and scientific validation as separate fields avoids an ambiguous claim of verification.

task: scientific objective + acceptance criteria
inputs: artifact identity + generator + parameters
execution: code version + environment + random streams
result: artifact identity + metric definitions
uncertainty: estimator + assumptions + precision
resources: stage timings + timing boundary
validation: reference comparison + constraint checks

3 / Account for the complete workflow

Resource accounting should report preparation, transfer, queue, execution and analysis components under an explicit timing boundary. Parallel tasks may overlap, so critical-path elapsed time should be reported separately from summed task durations. A small circuit runtime can coexist with an expensive full workflow.

For stochastic outputs, the evidence record should connect resource consumption to estimator quality. The executed QAOA study supplies one simple example: a shot count belongs beside the resulting standard error. For classifiers, the calibration study shows why an ordering metric and a probability-quality metric answer different questions. A single score should not silently collapse these distinctions.

4 / A falsifiable evaluation protocol

The first evaluation would use one public or synthetic workload and compare serial execution, batched classical execution and an available quantum simulator. Each route would produce the same evidence schema. Prespecified changes to seeds, inputs and numerical precision would test whether the record explains expected variation and exposes incompatible runs.

The hypothesis would fail if reviewers could not recover the stated metrics, if materially different inputs appeared comparable, or if the recorded timing boundaries concealed the principal cost. Successful evaluation would require recoverable metrics, correctly detected mismatches and measured record-keeping overhead. A multi-platform or CERN deployment has not yet been performed; this page defines the experiment to be run.

5 / Scientific computing context

CERN QTI’s public program connects hybrid algorithms, orchestration and distributed infrastructure. ROOT’s RDataFrame documentation describes parallel data analysis and distributed execution. These provide a relevant public setting for an independent study of result comparability across execution routes.

A focused collaboration would implement one adapter and one workload, publish the schema and benchmark scripts, and report both scientific consistency and operational overhead. The intended contribution is practical experimental infrastructure: results that can be inspected, compared and extended by another team.

Scientific context: [1] [2] [3]

References

  1. [1] CERN QTI — Hybrid computing infrastructures, algorithms and applications ↗
  2. [2] CERN QTI — Distributed quantum computing infrastructure ↗
  3. [3] ROOT — RDataFrame: data analysis and parallel execution ↗

Author

Maurizio Viviani

Independent research · Robotics

Develop the next experiment.

Discuss scientific benchmarking, hybrid computing or a reproducible evaluation study.

Discuss a research collaboration