Back to List
Notice:This resource is provided by a third-party author. Please review the code with AI tools or manually before use to ensure security and compatibility.
Pythonhuggingface/lighteval

lighteval

Lighteval is your all-in-one toolkit for evaluating LLMs across multiple backends

76.4/100
2.5KForks: 553
View on GitHubHomepage →
Loading report...

Similar Projects

deepeval

87

The LLM Evaluation Framework

Python18.2K

continuous-eval

74

Data-Driven Evaluation for LLM-Powered Applications

Python517

mlflow

91

The open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all sizes to debug, evaluate, monitor, and optimize production-quality AI applications while controlling costs and managing access to models and data.

Python27.8K

oumi

90

Easily fine-tune, evaluate and deploy Qwen, Gemma, or any open weight LLM!

Python9.4K
Back to List