A board is asked to approve deployment of an AI system into a consequential workflow: a customer-facing decision, a financial risk assessment, a step that touches patient information. The vendor's pitch deck includes a capability score. An internal pilot report includes a success rate. Somewhere in the discussion, someone cites a benchmark result showing the system outperforms a competitor. An evaluation number can be treated as though it were a portable property of the system, something fixed that will hold regardless of how the system ends up being used. Whether that inference travels has to be established, not assumed.

A related article on this site asked whether a human reviewer actually catches what an AI system gets wrong. That is a question about a control: does the reviewer add something real, or simply approve what was already there. This article asks a different question. Whatever review process sits downstream of a capability number, the number describing what a system can do deserves scrutiny as evidence in its own right.

An evaluation result is an estimate produced under specific conditions. The part worth pausing on is this: it does not have to be wrong to mislead. A number can be completely accurate and still be the wrong evidence for the decision in front of a board. Two recent studies, one from the UK's AI Security Institute (AISI), one from the US National Institute of Standards and Technology (NIST), show this from two different directions.

How compute budget changes measured capability

AISI's Science of Evaluations team tested frontier AI agents across a range of compute budgets rather than scoring each model once at a fixed cutoff. The finding: agent capability is not a single number but a curve that rises as the compute budget grows, and if that curve is still rising when an evaluation stops, the reported score is a lower bound, not a ceiling. For one recent frontier model, AISI's estimated 80 percent time horizon, the duration of task the model can complete with an 80 percent success rate, rose from around 40 minutes at a 2.5-million-token budget to around 4 hours at a 50-million-token budget. The model and task set were unchanged; the evaluation compute budget was not.

The consequence appears at the level of trend estimates as well, not just individual scores. AISI's own previous estimate of how fast frontier cyber capability has been advancing, a doubling time of 4.7 months, was measured at a 2.5-million-token budget. Measured at 50 million tokens instead, the fitted trend comes out roughly 60 percent steeper. The apparent rate of frontier progress therefore also depends partly on the compute budget used to estimate capability.

This does not mean more compute always reveals more capability, and that limitation is itself decision-relevant rather than a caveat to set aside. On HealthBench, AISI found every model plateaued within its usual budget; extra compute helped most where a system could check its own work, such as code or cybersecurity tasks, and helped little where feedback was weak or absent. Whether compute budget is material therefore depends on the task setting; it cannot be treated as an incidental evaluation detail in one setting simply because it proved immaterial in another.

Benchmark accuracy and generalized accuracy

NIST's contribution comes from a different angle. Its report on statistical modeling in AI benchmarking formally separates two things a single accuracy number can mean: benchmark accuracy, performance on the specific set of questions in that benchmark, and generalized accuracy, performance across the broader population of similar questions the benchmark is meant to represent. These are not the same quantity, and treating them as interchangeable can mislead. In an illustrative analysis of 22 frontier language models across three benchmarks, NIST found that some model pairs had significantly different benchmark accuracy while showing no significant difference in generalized accuracy. One such comparison, on GPQA-Diamond, was Llama 3 (40B) versus Phi 4. Two systems that appeared statistically different on the fixed benchmark questions were not shown to differ significantly in generalized accuracy over the broader population of similar questions.

Statistical significance is not the same thing as business materiality, and NIST's framework only speaks to the former. Even where a comparison is statistically supported, management still has to decide whether the size of that difference matters for the decision at hand.

From an evaluation result to a decision

A vendor's outperformance claim, an internal pilot's success rate, and a benchmark citation in a proposal do not all inherit exactly the same vulnerability. AISI's finding concerns agent evaluations and the resource budget allowed during testing. NIST's concerns the statistical relationship between a fixed set of test questions and the broader population those questions are meant to stand in for. A board's task is not to demand identical documentation for every number it is shown. It is to understand which conditions and assumptions are material to the inference being drawn from that number for the decision at hand.

Concretely: was the resource budget appropriate to the capability claim being made and to the way the system will actually be used, and would a materially different budget change the conclusion? And is the apparent advantage specific to the exact questions in this benchmark, or does the evidence support expecting it across the broader class of tasks the organization actually cares about?

What management should be able to explain

Management, not the board itself, should be able to answer four things about any evaluation number put forward as evidence, two of them AISI's territory and two of them NIST's.

From AISI

Compute budget and resource conditions

  1. What compute or resource budget produced this number, and had performance plateaued at that budget, or was the curve still rising when the evaluation stopped?
  2. Was that budget appropriate to how the system will actually be resourced once deployed, such that a materially different budget would not change the conclusion?

From NIST

What the number estimates, and how confidently

  1. What exactly is this number estimating: the specific test items used, or a broader population of similar tasks?
  2. How much uncertainty surrounds any comparison being drawn from it, and is the apparent difference actually supported for the quantity that matters?

A board does not need to interrogate an estimand or a confidence interval. It needs assurance that management can justify the connection between the evaluation and the decision.

Both AISI and NIST arrive at a version of the same conclusion from different directions. Change a material compute budget, and measured capability can change. Change what quantity is being estimated or how uncertainty is treated, and the comparison the evidence supports can change. Neither finding says evaluators were careless. Both show that what an evaluation result supports depends on its measurement target and on the conditions under which it was produced. This sits inside what is increasingly called AI assurance, but the evaluation number itself is where that assurance either holds or does not. A result does not have to be wrong to be inadequate evidence for a decision. The job of governance, in this narrower sense, is not to reproduce the evaluation. It is to ensure that the organization can justify why the evaluation supports the decision being made from it.