AI for Science

Three AI Critics Can Still Make the Same Scientific Mistake

2026-09-10

Agreement among several AI critics is useful only to the extent that their errors are not simply shared. Different brand names do not establish statistical independence. Models may receive the same incomplete brief, retrieve the same mistaken source or favour the same persuasive wording. Use diverse tests, separate initial reviews and external evidence checks. Report model consensus as a review signal, never as proof of a scientific claim or automatic impartiality.

Count independent checks, not model logos

A research proposal is sent to three models. All approve it. The result feels stronger than one approval, but the models may all have seen the same summary, accepted the same assumption and never examined the underlying data. The review has multiplied readers without necessarily multiplying evidence.

This is not an argument against multiple models. It is an argument for assigning them genuinely different jobs. A source checker, a numerical reproducer and an alternative-explanation reviewer may expose different weaknesses. Three general requests to assess whether an idea is good can produce a reassuring consensus whose basis remains opaque.

See the independence assumption in the arithmetic

In a deliberately simplified illustration, suppose each of three critics has a 10% chance of making a particular error. If the errors were independent, the chance that all three make it would be 0.1 × 0.1 × 0.1 = 0.1%. If their errors were perfectly correlated, the chance all three make it would instead remain 10%.

These are invented assumptions, not measured error rates for any model or service. Real dependence is more complicated and varies with the task. The example shows why multiplying confidence from model count alone is unjustified. The essential empirical question is how often the committee fails together on the types of claims that matter to your project.

Audit the judge as well as the answer

Research on LLM judges has identified position, verbosity and self-enhancement biases. These findings concern particular evaluations, not a universal error rate, but they show why a judge's approval cannot be treated as a neutral measurement by default. A longer explanation or its position in the prompt may influence the assessment.

Before relying on a review procedure, use a small set of known examples with clear defects and correct alternatives. Swap answer order, remove author labels and vary irrelevant presentation details. If the verdict changes while the evidence does not, the review protocol needs work. Keep this calibration separate from the scientific result the client ultimately cares about.

Sources: Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

Design critics around different failure modes

Consider a hypothetical claim that a model predicts future demand better than a seasonal baseline. One critic checks whether the cited literature says what the report claims. A second inspects the timing of features for leakage. A third reruns the metric from saved predictions. Their disagreement can now identify a concrete weakness rather than a difference in rhetorical preference.

A study of multi-agent judging examined several biases, including position and bandwagon effects. That motivates testing the committee structure itself instead of assuming discussion removes bias. Start with independent written reviews, then reconcile them against sources and outputs. Do not let the most persuasive agent rewrite the evidence presented to all the others.

Sources: Judging with Many Minds: Do More Perspectives Mean Less Prejudice?.

Give consensus the right place in the memo

The report can state which reviewers agreed, which tests they performed and what remained unresolved. It should not say that science has been validated because multiple models approved it. A numerical reproduction, a source audit and an expert interpretation are different kinds of support and should remain distinguishable.

The accountable researcher owns the final recommendation. If a critic identifies a fatal data issue, one well-supported objection outweighs several unsupported approvals. If all critics agree but none checked a key assumption, that assumption remains open. This is how frontier AI can make an unconventional hypothesis easier to challenge without turning a committee of language models into a substitute for evidence.

Questions this raises

Are different providers independent judges?

They may differ in useful ways, but provider identity alone does not demonstrate independent errors, especially when all receive the same flawed evidence.

Should a majority vote decide whether a hypothesis is true?

No. Votes can organise review. Scientific conclusions must follow the relevant evidence, including decisive objections raised by a minority reviewer.

Sources and their limits

Prepared with AI assistance. The linked sources support the specified technical points; they do not validate applied psionics as a whole or guarantee a result for a client.

Read the editorial and evidence standard.

Continue reading

Explore Scientific Oracle consultingfor a scoped review of an existing-data research decision. Start with a non-confidential outline of the question, available evidence and the decision it needs to inform.