Oracle Services

Acceptance Tests for Commissioned Research, Including Negative Results

2026-09-10

Research acceptance tests should assess whether the agreed investigation was completed to its specified standard, not whether it confirmed the client's preferred hypothesis. Define the question, inputs, analyses, deliverables and review procedure before execution. A supported negative finding can satisfy those criteria. A broken pipeline, omitted control or unsupported conclusion cannot be excused by describing uncertainty as part of science. Commercial terms need their own written agreement.

Separate the research outcome from delivery quality

A client commissions a comparison because the answer is not known. If acceptance requires the new method to outperform the baseline, the supplier has an incentive to keep changing the evaluation until it does. The risk is that the project rewards a favourable chart rather than a trustworthy investigation. If acceptance is intentionally performance-contingent, freeze the evaluation and state separately what a correctly executed but unsuccessful result means for payment and scope.

The opposite extreme is equally unhelpful: a supplier should not be able to deliver an unreadable notebook and claim that any uncertainty counts as success. Agree two separate questions. Was the promised analysis carried out competently and transparently? What did that analysis establish? The first governs technical delivery review. The second informs the client's next scientific decision.

Write tests a reviewer can actually perform

Use observable criteria. Avoid wording such as provide groundbreaking insight or demonstrate world-class intelligence. A reviewer cannot reliably inspect those phrases. The following example concerns an existing public forecasting dataset and is a technical scope illustration, not legal contract language.

DeliverableAcceptance checkNot required for acceptance
Dataset auditVersion, exclusions and timing documentedData proves the preferred thesis
Baseline comparisonAgreed split and metric reproducedCandidate beats baseline
Robustness reviewNamed sensitivity checks reportedEvery result remains positive
Decision memoConclusions trace to outputs and limitsRecommendation is always proceed

Define the difference between a defect and a disagreement

If the report uses the wrong dataset version, that is a defect against the agreed specification. If the analysis is correctly performed but the client favours a different scientific interpretation, that is an evidential disagreement to resolve explicitly. If the client now wants a new dataset or a causal claim instead of a predictive comparison, that is likely a scope change.

ACM's artifact framework distinguishes functioning research artifacts from independently validated results. This supports separating different kinds of review rather than treating all approval as one event. For a commission, name the reviewer, the files they inspect and how a documented defect is corrected. Do not assume that an attractive PDF proves the underlying computation is complete.

Sources: ACM: Artifact Review and Badging.

Protect the original evaluation from moving targets

Freeze the central metric and comparison before inspecting the decisive results. Additional analyses may be useful, but label them as later exploration. A client should be able to tell whether the original question succeeded, failed or remained unresolved, even if a new framing looks promising.

The Center for Open Science explains that pre-analysis planning helps separate planned from exploratory analyses and that existing-data work requires transparency about prior exposure to outcomes. In a private project, a dated internal protocol can preserve that distinction without pretending the study was publicly preregistered. Keep the initial version and record amendments rather than silently replacing it.

Sources: Center for Open Science: Preregistration.

Make milestones buy inspectable progress

For a deeper commission, one milestone might deliver a data-readiness decision, another a reproducible comparison, and another a challenge memo. The appropriate sequence depends on the question. Payment timing, ownership and any dispute process must be negotiated separately; an AI score should not be presented as inherently impartial or a substitute for those terms.

A Direction Preview is narrower. It offers one exploratory direction, a bounded preliminary evidence check, a written recommendation and a further-research proposal when justified. It should not be evaluated as if it included a finished scientific solution. Matching the acceptance tests to the actual purchase helps the buyer recognise useful early work without overpaying for ambiguous promises.

A good final review can end with complete and stop. That is not a contradiction. It means the commissioned work was delivered and the evidence does not justify continuing along that route.

Questions this raises

Must I pay for a negative finding?

That depends on the agreement. Technically, a negative finding can be a complete deliverable when it results from the agreed analysis and meets the agreed quality criteria.

Can an AI judge replace delivery criteria?

No. Any assessment tool still needs clear criteria, evidence access, known limitations and an agreed review process. Model agreement does not make an ambiguous scope precise.

Sources and their limits

Prepared with AI assistance. The linked sources support the specified technical points; they do not validate applied psionics as a whole or guarantee a result for a client.

Read the editorial and evidence standard.

Continue reading

Explore Scientific Oracle consultingfor a scoped review of an existing-data research decision. Start with a non-confidential outline of the question, available evidence and the decision it needs to inform.