Oracle Services
Acceptance Tests for Commissioned Research, Including Negative Results
2026-09-10
Research acceptance tests should assess whether the agreed investigation was completed to its specified standard, not whether it confirmed the client's preferred hypothesis. Define the question, inputs, analyses, deliverables and review procedure before execution. A supported negative finding can satisfy those criteria. A broken pipeline, omitted control or unsupported conclusion cannot be excused by describing uncertainty as part of science. Commercial terms need their own written agreement.
Separate the research outcome from delivery quality
A client commissions a comparison because the answer is not known. If acceptance requires the new method to outperform the baseline, the supplier has an incentive to keep changing the evaluation until it does. The risk is that the project rewards a favourable chart rather than a trustworthy investigation. If acceptance is intentionally performance-contingent, freeze the evaluation and state separately what a correctly executed but unsuccessful result means for payment and scope.
The opposite extreme is equally unhelpful: a supplier should not be able to deliver an unreadable notebook and claim that any uncertainty counts as success. Agree two separate questions. Was the promised analysis carried out competently and transparently? What did that analysis establish? The first governs technical delivery review. The second informs the client's next scientific decision.
Write tests a reviewer can actually perform
Use observable criteria. Avoid wording such as provide groundbreaking insight or demonstrate world-class intelligence. A reviewer cannot reliably inspect those phrases. The following example concerns an existing public forecasting dataset and is a technical scope illustration, not legal contract language.
| Deliverable | Acceptance check | Not required for acceptance |
|---|---|---|
| Dataset audit | Version, exclusions and timing documented | Data proves the preferred thesis |
| Baseline comparison | Agreed split and metric reproduced | Candidate beats baseline |
| Robustness review | Named sensitivity checks reported | Every result remains positive |
| Decision memo | Conclusions trace to outputs and limits | Recommendation is always proceed |
Define the difference between a defect and a disagreement
If the report uses the wrong dataset version, that is a defect against the agreed specification. If the analysis is correctly performed but the client favours a different scientific interpretation, that is an evidential disagreement to resolve explicitly. If the client now wants a new dataset or a causal claim instead of a predictive comparison, that is likely a scope change.
ACM's artifact framework distinguishes functioning research artifacts from independently validated results. This supports separating different kinds of review rather than treating all approval as one event. For a commission, name the reviewer, the files they inspect and how a documented defect is corrected. Do not assume that an attractive PDF proves the underlying computation is complete.
Sources: ACM: Artifact Review and Badging.
Protect the original evaluation from moving targets
Freeze the central metric and comparison before inspecting the decisive results. Additional analyses may be useful, but label them as later exploration. A client should be able to tell whether the original question succeeded, failed or remained unresolved, even if a new framing looks promising.
The Center for Open Science explains that pre-analysis planning helps separate planned from exploratory analyses and that existing-data work requires transparency about prior exposure to outcomes. In a private project, a dated internal protocol can preserve that distinction without pretending the study was publicly preregistered. Keep the initial version and record amendments rather than silently replacing it.
Sources: Center for Open Science: Preregistration.
Make milestones buy inspectable progress
For a deeper commission, one milestone might deliver a data-readiness decision, another a reproducible comparison, and another a challenge memo. The appropriate sequence depends on the question. Payment timing, ownership and any dispute process must be negotiated separately; an AI score should not be presented as inherently impartial or a substitute for those terms.
A Direction Preview is narrower. It offers one exploratory direction, a bounded preliminary evidence check, a written recommendation and a further-research proposal when justified. It should not be evaluated as if it included a finished scientific solution. Matching the acceptance tests to the actual purchase helps the buyer recognise useful early work without overpaying for ambiguous promises.
A good final review can end with complete and stop. That is not a contradiction. It means the commissioned work was delivered and the evidence does not justify continuing along that route.
Questions this raises
Must I pay for a negative finding?
That depends on the agreement. Technically, a negative finding can be a complete deliverable when it results from the agreed analysis and meets the agreed quality criteria.
Can an AI judge replace delivery criteria?
No. Any assessment tool still needs clear criteria, evidence access, known limitations and an agreed review process. Model agreement does not make an ambiguous scope precise.
Sources and their limits
- ACM: Artifact Review and Badging. Supports separating artifact functionality from independent result validation.
- Center for Open Science: Preregistration. Supports explicit prior plans and disclosure of existing-data exposure; does not supply contractual terms.
Prepared with AI assistance. The linked sources support the specified technical points; they do not validate applied psionics as a whole or guarantee a result for a client.
Read the editorial and evidence standard.
Continue reading
- When a bounded fixed-price review fits
- Preview and commissioned research are different scopes
- Explore the topic library
Explore Scientific Oracle consultingfor a scoped review of an existing-data research decision. Start with a non-confidential outline of the question, available evidence and the decision it needs to inform.