
GUI-Primitives: Spatial Reasoning Failures in Vision Models
GUI-Primitives tests spatial reasoning in vision-language models for GUI grounding, finding weak localization…
Featured Compare LLM evaluation tools for CI gates, reproducible runs, judge control, trace linkage and licensing, with promptfoo and lm-eval configs.
Read the article →
GUI-Primitives tests spatial reasoning in vision-language models for GUI grounding, finding weak localization…

Position, verbosity and self-preference bias in model graders, how to measure judge agreement against humans,…

What MMLU, HumanEval and GSM8K actually contain, how each score is computed, and the point where a headline b…

A comparison of lm-evaluation-harness, OpenAI Evals, HELM, Inspect, promptfoo, DeepEval and Ragas, with the t…

How to design an LLM evaluation: task-grounded eval sets, grader choice, sampling variance, and the contamina…
Evals ML is a reference for measuring language models: what the standard benchmark suites actually contain, how much confidence a published score carries, which harness fits which kind of task, and how to keep a model grader honest. Every figure quoted here is attributed to the paper or repository it came from.
MMLU, HumanEval and GSM8K: what is inside each set, how each is scored, and where the number stops being useful.
Building an eval set from the failures you care about, choosing a grader, and turning it into a regression suite that catches real change.
lm-evaluation-harness, OpenAI Evals, HELM, Inspect, promptfoo, DeepEval and Ragas, compared on the task shapes each one fits.
Position, verbosity and self-preference bias, how to detect each, and how to calibrate a judge against human labels.
Generation count, token spend and worst-case margin of error for a given dataset size and sampling budget. Runs in the browser, no signup.
All guides · Browse by topic · How these articles are researched