Research
The Missing Denominator: Count Every Intuitive Research Attempt
2026-09-10
An intuition success rate is uninterpretable without a denominator that includes all eligible attempts. Define an attempt before evaluation, preserve revisions and separate untested ideas from tested outcomes. Report exclusions and missing results explicitly. A memorable successful hypothesis may be useful, but it cannot establish the performance of a method whose unsuccessful search history is absent.
A highlight reel answers the wrong question
A portfolio of three striking hypotheses can show what a researcher finds interesting. It cannot by itself tell a buyer how often the process yields a useful direction. That requires knowing how many questions were considered, how many candidates were generated and how many were actually evaluated.
For Andrei Ursachi's exploratory applied-psionics framework, the relevant commercial question is whether intuition helps allocate research effort. A complete search record can address that question more honestly than a collection of anecdotes, even when some anecdotes are genuinely impressive.
Decide whether the unit is a client question, a candidate hypothesis or a computational test. Each denominator measures something different. Switching between them after the results are known can make the same workflow appear dramatically stronger or weaker.
One question is not always one attempt
Imagine a hypothetical review containing ten scientific questions. For each question, the analyst proposes four candidates, creating forty candidates in total. Only twelve receive a computational test, and three pass their specified comparisons. These illustrative counts yield several distinct summaries.
Three out of twelve describes tested-candidate performance. Three out of forty describes the fraction of all generated candidates reaching that outcome. Neither alone states whether three different client questions benefited, because all three successes might concern the same question.
Keep a hierarchy of question ID, candidate ID, revision ID and test ID. This lets the summary answer the actual buyer question without erasing the work that preceded the final memo.
| Unit | Illustrative count | What it tells you |
|---|---|---|
| Questions | 10 | Number of decisions presented |
| Candidates | 40 | Breadth of proposed search |
| Tested candidates | 12 | What entered evaluation |
| Candidates passing | 3 | Narrow test outcomes, not universal discoveries |
Handle untested ideas without inventing outcomes
An untested candidate is not automatically a scientific failure. The necessary data may be unavailable, the test may be too costly or another candidate may have higher information value. Record those reasons instead of converting every omission into either a success or a failure.
At the same time, a process that frequently produces untestable ideas has a practical limitation. Report testability yield separately from performance on the tested subset. This prevents excellent results on a tiny, selected subset from concealing poor usefulness across the original briefs.
Simmons and colleagues demonstrate how flexible analysis and reporting can inflate false-positive findings. The safeguard proposed here is operational: specify the counting rules and publish the full status distribution before choosing which individual examples to feature.
Sources: Simmons, Nelson and Simonsohn (2011), False-Positive Psychology.
Compare workflows at the same budget
Suppose one workflow generates hundreds of candidates and another generates five. Comparing only the best candidate from each rewards search volume without accounting for its cost. Use a fixed question set and comparable time or compute budgets when judging which approach helps more.
Selection can also leak information. If candidates are chosen because their evaluation outcomes were already visible, a later test summary is no longer an independent evaluation. Kapoor and Narayanan's analysis of leakage explains why information boundaries matter for scientific prediction claims.
Record whether outcomes influenced admission to the tested set. If they did, label the comparison exploratory and identify what would be needed for a cleaner assessment. A complete denominator helps reveal this issue, but does not automatically repair it.
Ask for a decision-level report
A buyer should be able to ask: for how many eligible questions did the workflow produce a testable direction, how often did that direction survive the agreed check, and what did the review cost? These are different questions and deserve different numbers.
If you have an existing archive of briefs and outcomes, a bounded consulting discussion can assess whether those quantities are recoverable. Start with aggregate counts and non-confidential descriptions. There is no need to disclose proprietary scientific content before agreeing appropriate handling.
The practical takeaway is to retain the entire path from brief to recommendation. Fast progress is not the number of attractive ideas generated. It is the quality of decisions improved relative to the total effort required.
Questions this raises
Should every raw thought count as a prediction?
No. Define when an impression becomes an entered candidate. The threshold should be fixed before results, and raw notes should remain available for provenance.
Can selected case studies still be published?
Yes, if clearly labeled as selected examples rather than evidence of an overall success rate.
Sources and their limits
- Simmons, Nelson and Simonsohn (2011), False-Positive Psychology. Undisclosed flexibility in analysis and reporting can inflate false-positive findings.
- Kapoor and Narayanan (2023), Leakage and the reproducibility crisis in machine-learning-based science. Information leakage can produce overly optimistic scientific prediction results.
Prepared with AI assistance. The linked sources support the specified technical points; they do not validate applied psionics as a whole or guarantee a result for a client.
Read the editorial and evidence standard.
Continue reading
- Preserve creativity without losing evaluation discipline
- Choose which candidates deserve a test
- Explore the topic library
Explore Scientific Oracle consultingfor a scoped review of an existing-data research decision. Start with a non-confidential outline of the question, available evidence and the decision it needs to inform.