Editorial desk
Evals ML Editorial
Evals ML Editorial is the publishing identity for Evals ML. It is a desk, not a person: no named author, no biography, no professional certifications.
Articles published under this byline are researched from primary sources — vendor and project documentation, published standards and specifications, research papers, and measurements published by whoever took them — drafted with AI assistance, and edited against those cited sources before publication. Nothing here is based on first-hand testing in a private lab, and any figure that appears is attributed to the source it came from.
Corrections go to [email protected]. More detail is on the about page and the editorial disclosure.
Posts (6)
- Metrics
Pass@k Metrics Explained: Formula, Estimator, Pitfalls
Pass@k is the chance at least one of k samples is correct. How the unbiased estimator works, why pass@1 is the honest number, and when pass^k wins.
- Eval Tooling
LLM Evaluation Frameworks Compared: Task Fit and CI Gates
Compare LLM evaluation tools for CI gates, reproducible runs, judge control, trace linkage and licensing, with promptfoo and lm-eval configs.
- Research
GUI-Primitives: Spatial Reasoning Failures in Vision Models
GUI-Primitives tests spatial reasoning in vision-language models for GUI grounding, finding weak localization and 32% peak strict point-in-box accuracy.
- Model Grading
LLM-as-a-Judge Bias: How to Detect and Correct It
Position, verbosity and self-preference bias in model graders, how to measure judge agreement against humans, and the calibration steps that reduce each.
- Benchmarks
LLM Benchmarks Explained: MMLU, HumanEval, GSM8K
What MMLU, HumanEval and GSM8K actually contain, how each score is computed, and the point where a headline benchmark number stops being useful.
- Eval Design
LLM Evaluation Design: From Task to Regression Suite
How to design an LLM evaluation: task-grounded eval sets, grader choice, sampling variance, and the contamination traps that mislead benchmark scores.