Evals ML
LLM Evaluation & Benchmark Engineering

Pass@k Calculator and Eval Cost Estimator

Enter how many samples you draw per item (n) and how many of them pass (c) to get the unbiased pass@k estimate for that item. Then size the whole run: total generations, API cost at the rates you enter, and the margin of error your item count allows.

1 to 500. Draw at least as many samples as the largest k you report.

Must be no larger than n.

How many of that item's n samples passed its tests. Must be no larger than n.

pass@k for this item
30.0%
The benchmark score is the mean of this over all items
Total generations
1,640
items x n
Total evaluation cost
$8.61
At the rates entered above
95% margin of error (worst case)
± 7.7 pts
From the item count

pass@k uses the unbiased estimator from Chen et al. (2021) in its numerically stable product form, 1 - (1 - k/(n-c+1)) x ... x (1 - k/n), which is 1 whenever n - c < k. It is the estimate for one item: a benchmark's pass@k is the mean of this value over every item, so compute it from your per-item pass counts. Cost re-sends the prompt tokens with every sample at the input rate and charges the completion tokens at the output rate; prompt caching, where a provider offers it, cuts the input half. The preset rates are list prices for the named models on 3 October 2026, so check the provider's current pricing page before budgeting. The margin of error is the 95% normal-approximation interval at the worst-case pass rate of 0.5, an upper bound on the interval around a score from a set of this size. It describes the sampling error of the eval set only, not grader error.

How to read these numbers

  • pass@k counts an item as solved if any of k attempts passes. Estimating it from n > k samples, rather than drawing exactly k, is what makes the estimate unbiased and lowers its variance.
  • Generations scale with n, not with the dataset. HumanEval at the Codex paper's n = 200 is 32,800 model calls, not 164.
  • Cost is usually driven by output tokens, because output rates run several times input rates on most price lists. Long few-shot prompts are the exception, so set the prompt field to your real prompt length.
  • Margin of error is what decides whether a comparison is real. On a 164-item set the 95% interval is roughly ±8 points, which is wider than most of the gaps quoted between models.