AI for Science

A Four-Field Citation Audit Worksheet for AI Literature Summaries

2026-09-10

For a numerical literature claim, record four things: the draft sentence, the source's evaluation conditions, the exact reported contrast, and a verdict with corrected wording. Extract the metric, reference value, evaluation unit and time horizon before repeating a percentage. This worksheet catches a common failure: an accurate number becomes an inaccurate scientific claim when its denominator or scope disappears.

Start where ordinary source verification finishes

Suppose an AI summary says a forecasting method is twenty percent more accurate across climates. Its citation resolves to the correct paper. The remaining question is not whether the publication exists, but whether that sentence preserves what was measured. A number can survive copying while its meaning changes.

For checking publication identity and general source support, use Source Verification for AI-Assisted Scientific Research, linked below. This companion worksheet focuses on numerical scope inflation. NIST identifies confabulated reasoning and citations among generative AI risks; the worked example here is an original audit exercise, not a result reported by NIST.

Sources: NIST: Generative AI Profile.

Locate the row, metric and evaluation window

First pin the publication version and the table or figure containing the result. Crossref provides scholarly metadata lookup, including DOI records. That can help identify the work; it cannot tell you whether its table supports the proposed sentence.

Imagine a fictional paper evaluating 120 twenty-four-hour-ahead forecasts from one regional archive over two months. Its table reports mean absolute error of 10 units for a baseline and 8 units for Method A. These are invented values for the worksheet. No confidence interval or cross-region evaluation is supplied in this example.

Write down the unit of evaluation: forecasts, not necessarily independent days, sites or participants. Also distinguish the forecast horizon from the evaluation period. Twenty-four hours ahead and two months of evaluation answer different questions. Neither field should disappear into the vague phrase long-term performance.

Sources: Crossref: Metadata Retrieval.

Fill four fields before approving the sentence

The card preserves enough context to correct the claim without reproducing an entire paper. Attach the source identifier and table location to the card itself. The four fields below illustrate a completed audit for the fictional forecasting result.

  • State the metric instead of replacing every improvement with accuracy.
  • Preserve the comparator and the reference value used for a relative percentage.
  • Mark missing uncertainty or transfer evidence rather than supplying it from intuition.
FieldIllustrative entry
1. Draft sentenceMethod A is 20% more accurate across climates
2. Evaluation conditions120 forecasts; one region; two months; 24-hour horizon
3. Reported contrastMean absolute error: baseline 10 units, Method A 8 units; relative reduction 2/10 = 20%
4. Verdict and rewriteScope inflated. Method A had 20% lower mean absolute error than the stated baseline in this regional evaluation

Recalculate the percentage and challenge the transfer

The illustrative reduction is (10 minus 8) divided by 10, which equals 20 percent. It is not a twenty-percentage-point increase in classification accuracy. It is also not a twenty-percent saving in operating costs. Those would require different measurements and a separate evidential bridge.

Next check whether the result applies to the setting in the proposed recommendation. A regional retrospective comparison does not itself establish performance across climates. Even the corrected sentence should identify whether the reported evaluation was held out and what information the models could use at prediction time, when the source makes that available.

If the paper does not report the necessary denominator, aggregation method or timepoint, mark that field unresolved. Do not reverse-engineer a convenient value from rounded prose and present it as the authors' analysis. The audit should make uncertainty visible, not remove it by arithmetic.

Deliver corrected claims, not a longer bibliography

For a buyer, package the cards for the few quantitative claims that change the decision. Preserve the original sentence alongside its replacement so the difference is visible. Separate supported numerical comparisons from extrapolations requiring another dataset or analysis. This makes the literature review useful for scoping subsequent computational work.

AI can draft the cards and locate candidate passages, while a reviewer checks the exact table and arithmetic. If the cited item is a review, follow consequential numbers to the original result. Multiple summaries repeating one table remain one underlying observation.

The outcome may be that a method still looks promising, but for a narrower reason. That is not a failed review. It is a cleaner research question: whether the improvement survives the client's relevant horizon, comparator and permitted data, rather than whether an inflated headline sounds convincing.

Questions this raises

What if the paper never reports the denominator?

Mark it unresolved and avoid repeating a percentage that cannot be interpreted reliably. A missing source field is a limitation to disclose, not a value for AI to guess.

Can a correct percentage still mislead?

Yes. It may describe a different metric, comparator, horizon or setting from the buyer's question. Preserve those conditions even when the arithmetic is correct.

Sources and their limits

Prepared with AI assistance. The linked sources support the specified technical points; they do not validate applied psionics as a whole or guarantee a result for a client.

Read the editorial and evidence standard.

Continue reading

Explore Scientific Oracle consultingfor a scoped review of an existing-data research decision. Start with a non-confidential outline of the question, available evidence and the decision it needs to inform.