Scientific Discovery

Rank Hypotheses Before and After AI to See What the Model Actually Added

2026-09-10

Save the candidate list and ranking before AI review, then record exactly which candidates the model adds, removes or reorders and why. Evaluate the resulting choices against a fixed external criterion and a fair AI-only or ordinary baseline when feasible. A better-written explanation is not automatically a better research decision, and a changed ranking is not automatically an improvement.

The final answer hides the contribution history

A human proposes a mechanism, a model supplies a polished explanation, and the report presents the result as a combined insight. That may describe collaboration accurately, but it does not show which step improved the scientific choice or whether either component helped at all.

Andrei Ursachi's exploratory workflow combines trained intuition with frontier AI. To examine that combination, preserve a pre-AI snapshot rather than reconstructing it from memory after the model has offered alternatives. The record should distinguish generating an idea from defending it.

Use a stable candidate identifier. If a model rewrites a vague candidate into a materially different mechanism, assign a new version instead of treating the change as cosmetic. Otherwise the original intuition receives credit for specificity introduced later.

A hypothetical ranking of four computational routes

Consider an existing anomaly-detection dataset. The human ranking is A, B, C, D. AI review proposes that A may reflect duplicated records, promotes C and adds E. The post-review ranking becomes C, E, B, A, D.

This record supports a concrete attribution statement: the human proposed the initial set; the model identified a possible artifact and introduced a new candidate. It does not yet show that C or E is scientifically better. That requires the agreed data comparison.

Save the model version, the supplied context and the reason for each change. An AI receiving extra literature or outcome information has an information advantage, not necessarily a superior reasoning method. Compare like with like when interpreting the contribution.

SnapshotRankingWhat is preserved
Before AIA, B, C, DHuman candidate ordering
AI reviewPromote C; add E; challenge ASpecific additions and objections
After reviewC, E, B, A, DDecision before evaluation
After dataPendingNo improvement claimed in advance

Define improvement before measuring it

The outcome might be the quality of the top-ranked candidate, the number of useful candidates found under a fixed budget or the time required to reject an artifact. These are different objectives. Choose the one that corresponds to the buyer's decision.

If candidates receive explicit probabilities for a defined outcome, proper scoring rules can evaluate those probabilities against results. Gneiting and Raftery provide the relevant statistical foundation. Do not convert a model's confident prose into a numerical probability without an agreed interpretation.

A ranking comparison can use a simpler criterion, such as whether the first funded route passes a specified held-out check. Report the full sequence as well, so a method that needs many attempts is not made to resemble one that succeeds immediately.

Sources: Gneiting and Raftery (2007), Strictly Proper Scoring Rules, Prediction, and Estimation.

Do not let the evaluation become another prompt revision

Once the final comparison is opened, changes to prompts, candidate definitions and preprocessing are exploratory. A new model version may be genuinely better, but testing it repeatedly on the same revealed answers does not provide fresh evidence.

Preprocessing is part of this boundary. The scikit-learn documentation warns against fitting transformations on test data. In a human-AI review, summaries derived from the evaluation set can leak information just as effectively as an improperly fitted scaler.

Where feasible, preserve an AI-only route with the same allowed information and budget. If the combined route performs similarly, that is useful attribution evidence. If no fair comparison is possible, report the observed workflow rather than claiming a uniquely intuitive advantage.

Sources: scikit-learn, Common pitfalls and recommended practices.

Ask which contribution would justify the next investment

A buyer may care more about eliminating a bad route early than about proving which participant deserves intellectual credit. The before-and-after record can serve both purposes by showing how the recommendation changed and whether that change improved the decision.

For a first consulting exchange, describe the current candidate list, what AI tools already contribute and the existing evidence available for comparison. A bounded review can identify the missing challenge and propose a fair way to evaluate the revised shortlist.

The practical takeaway is to preserve provenance before polishing the story. A combined workflow becomes more credible when it can show what each component added, what the data rejected and what remains unproven.

Questions this raises

Does a different ranking mean AI improved the result?

No. It establishes that AI changed the decision. Improvement requires an external outcome criterion chosen before evaluation.

Must every project include an AI-only baseline?

Not always. But without a fair comparison, claims about the unique value of intuition or the combined workflow should remain limited.

Sources and their limits

Prepared with AI assistance. The linked sources support the specified technical points; they do not validate applied psionics as a whole or guarantee a result for a client.

Read the editorial and evidence standard.

Continue reading

Explore Scientific Oracle consultingfor a scoped review of an existing-data research decision. Start with a non-confidential outline of the question, available evidence and the decision it needs to inform.