Back to List
Notice:This resource is provided by a third-party author. Please review the code with AI tools or manually before use to ensure security and compatibility.
Pythonconfident-ai/deepeval

deepeval

The LLM Evaluation Framework

87.2/100
18.2KForks: 1.9K
View on GitHubHomepage →
Loading report...

Similar Projects

continuous-eval

74

Data-Driven Evaluation for LLM-Powered Applications

Python517

lighteval

76

Lighteval is your all-in-one toolkit for evaluating LLMs across multiple backends

Python2.5K

future-agi

81

Open-source, end-to-end platform for evaluating, observing, and improving LLM and AI agent applications. Tracing · Evals · Simulations · Datasets · Gateway · Guardrails. Self-hostable. Apache 2.0.

Python1.9K

SkillCorpus

80

Open-source infrastructure that turns scattered SKILL.md files into curated, retrieval-ready agent-skill corpora—with retrieval and evaluation tooling included.

Python618
Back to List