#model-grading
-
LLM-as-a-Judge Bias: How to Detect and Correct It
Position, verbosity and self-preference bias in model graders, how to measure judge agreement against humans, and the calibration steps that reduce each.
-
LLM Eval Frameworks Compared: Pick the Right One
A comparison of lm-evaluation-harness, OpenAI Evals, HELM, Inspect, promptfoo, DeepEval and Ragas, with the task shapes each harness actually fits.
-
LLM Evaluation Design: From Task to Regression Suite
How to design an LLM evaluation: task-grounded eval sets, grader choice, sampling variance, and the contamination traps that mislead benchmark scores.