← All frameworks

lm-evaluation-harness

The benchmark runner behind most published model scores.

Evals and testing

The benchmark runner behind most published model scores. Built by EleutherAI.

Others in evals and testing

Common questions

Is lm-evaluation-harness free to use?

lm-evaluation-harness is open source and free to run yourself. You still pay whichever model provider you point it at, and this tutorial is free with no signup.

Do I need to know Python to use lm-evaluation-harness?

Basic Python is enough. Functions, dictionaries and imports cover most of what lm-evaluation-harness asks of you.

When should I not use lm-evaluation-harness?

When something else in evals and testing fits the job better. The others in that group are listed below.

Lessons for lm-evaluation-harness are being written. Meanwhile the APIs for AI tutorial covers the same ground: state, tools, loops, memory and human approval. Most of it carries straight over.

Start the APIs for AI tutorial →