openai/evals
steadyEvals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks.
Python
View on GitHub
Stars
19,500
Forks
3,092
Open issues
176
24h
+10
+0.1%
7d
+54
+0.3%
Refresh
2h
Star history (7 days)
Last checked
15m ago
Last pushed
14 Apr 2026
Next check
just now