Skip to content

Benchmark limitations

Research cut-off date: 2026-07-22. Status: Phase 2, not yet written. This file states the scope of the work, the arguments it owns, and the research required to complete it. It contains no findings, because none have been sourced. See the changelog for phase status.

Scope

The cross-cutting analysis of why benchmark scores fail to transfer: contamination, saturation and ceiling effects, prompt and harness sensitivity, sampling policy inflation, scaffold attribution, construct validity, and selective reporting. Each limitation is stated with its detection method and its consequence for the comparison rules applied throughout this repository.

Required structure for each benchmark

Every benchmark answers all nine questions, in this order. The ninth is a judgement and is argued rather than asserted.

  1. What it measures
  2. Task format
  3. Scoring method
  4. Known limitations
  5. Contamination risk
  6. Tool permissions
  7. Sampling policy
  8. Human baseline comparability
  9. Suitability for procurement

Research checklist

  • Locate and register the construction paper or methodology page for every benchmark listed above, before any result from it is recorded.
  • Answer all nine required questions for each benchmark, citing the construction source for the first three.
  • For each limitation, cite at least one Grade A source demonstrating it rather than asserting it as common knowledge.
  • Provide the detection method for contamination that this repository actually applies, and state its weakness where training data is undisclosed.
  • Quantify, from a source, the score variation attributable to harness and prompt formatting alone.
  • Confirm this file develops the arguments that chapter 09 summarises, without either file restating the other.
  • Record every extracted result in data/benchmarks.csv with harness, sampling policy, tool permissions, evidence grade, and date.
  • Run python scripts/validate_tables.py --check-ranges --check-comparability.

Completion criteria

This file is complete when every listed benchmark answers all nine questions from a registered source, when no result appears here that is absent from data/benchmarks.csv, and when the quality-control checklist in the research methodology passes.