AI for Science
A Four-Field Citation Audit Worksheet for AI Literature Summaries
2026-09-10
For a numerical literature claim, record four things: the draft sentence, the source's evaluation conditions, the exact reported contrast, and a verdict with corrected wording. Extract the metric, reference value, evaluation unit and time horizon before repeating a percentage. This worksheet catches a common failure: an accurate number becomes an inaccurate scientific claim when its denominator or scope disappears.
Start where ordinary source verification finishes
Suppose an AI summary says a forecasting method is twenty percent more accurate across climates. Its citation resolves to the correct paper. The remaining question is not whether the publication exists, but whether that sentence preserves what was measured. A number can survive copying while its meaning changes.
For checking publication identity and general source support, use Source Verification for AI-Assisted Scientific Research, linked below. This companion worksheet focuses on numerical scope inflation. NIST identifies confabulated reasoning and citations among generative AI risks; the worked example here is an original audit exercise, not a result reported by NIST.
Sources: NIST: Generative AI Profile.
Locate the row, metric and evaluation window
First pin the publication version and the table or figure containing the result. Crossref provides scholarly metadata lookup, including DOI records. That can help identify the work; it cannot tell you whether its table supports the proposed sentence.
Imagine a fictional paper evaluating 120 twenty-four-hour-ahead forecasts from one regional archive over two months. Its table reports mean absolute error of 10 units for a baseline and 8 units for Method A. These are invented values for the worksheet. No confidence interval or cross-region evaluation is supplied in this example.
Write down the unit of evaluation: forecasts, not necessarily independent days, sites or participants. Also distinguish the forecast horizon from the evaluation period. Twenty-four hours ahead and two months of evaluation answer different questions. Neither field should disappear into the vague phrase long-term performance.
Sources: Crossref: Metadata Retrieval.
Fill four fields before approving the sentence
The card preserves enough context to correct the claim without reproducing an entire paper. Attach the source identifier and table location to the card itself. The four fields below illustrate a completed audit for the fictional forecasting result.
- State the metric instead of replacing every improvement with accuracy.
- Preserve the comparator and the reference value used for a relative percentage.
- Mark missing uncertainty or transfer evidence rather than supplying it from intuition.
| Field | Illustrative entry |
|---|---|
| 1. Draft sentence | Method A is 20% more accurate across climates |
| 2. Evaluation conditions | 120 forecasts; one region; two months; 24-hour horizon |
| 3. Reported contrast | Mean absolute error: baseline 10 units, Method A 8 units; relative reduction 2/10 = 20% |
| 4. Verdict and rewrite | Scope inflated. Method A had 20% lower mean absolute error than the stated baseline in this regional evaluation |
Recalculate the percentage and challenge the transfer
The illustrative reduction is (10 minus 8) divided by 10, which equals 20 percent. It is not a twenty-percentage-point increase in classification accuracy. It is also not a twenty-percent saving in operating costs. Those would require different measurements and a separate evidential bridge.
Next check whether the result applies to the setting in the proposed recommendation. A regional retrospective comparison does not itself establish performance across climates. Even the corrected sentence should identify whether the reported evaluation was held out and what information the models could use at prediction time, when the source makes that available.
If the paper does not report the necessary denominator, aggregation method or timepoint, mark that field unresolved. Do not reverse-engineer a convenient value from rounded prose and present it as the authors' analysis. The audit should make uncertainty visible, not remove it by arithmetic.
Deliver corrected claims, not a longer bibliography
For a buyer, package the cards for the few quantitative claims that change the decision. Preserve the original sentence alongside its replacement so the difference is visible. Separate supported numerical comparisons from extrapolations requiring another dataset or analysis. This makes the literature review useful for scoping subsequent computational work.
AI can draft the cards and locate candidate passages, while a reviewer checks the exact table and arithmetic. If the cited item is a review, follow consequential numbers to the original result. Multiple summaries repeating one table remain one underlying observation.
The outcome may be that a method still looks promising, but for a narrower reason. That is not a failed review. It is a cleaner research question: whether the improvement survives the client's relevant horizon, comparator and permitted data, rather than whether an inflated headline sounds convincing.
Questions this raises
What if the paper never reports the denominator?
Mark it unresolved and avoid repeating a percentage that cannot be interpreted reliably. A missing source field is a limitation to disclose, not a value for AI to guess.
Can a correct percentage still mislead?
Yes. It may describe a different metric, comparator, horizon or setting from the buyer's question. Preserve those conditions even when the arithmetic is correct.
Sources and their limits
- NIST: Generative AI Profile. Supports the risk of confabulated reasoning and citations.
- Crossref: Metadata Retrieval. Supports DOI and scholarly metadata lookup, not verification of a paper's substantive claims.
Prepared with AI assistance. The linked sources support the specified technical points; they do not validate applied psionics as a whole or guarantee a result for a client.
Read the editorial and evidence standard.
Continue reading
- Start with the source-verification primer
- Place the worksheet inside a literature-review workflow
- Explore the topic library
Explore Scientific Oracle consultingfor a scoped review of an existing-data research decision. Start with a non-confidential outline of the question, available evidence and the decision it needs to inform.