# DeepEval

> Open-source LLM evaluation framework

- Category: [Observability & Evals](https://ailandscape.org/category/observability-evals) › LLM Evaluation & Benchmarking
- Homepage: https://deepeval.com
- Repository: https://github.com/confident-ai/deepeval
- Tags: evaluation, testing, llm
- Added to the landscape: 2026-03-18

## Similar tools in LLM Evaluation & Benchmarking

- [Braintrust](https://ailandscape.org/tool/braintrust): LLM evaluation and prompt management platform
- [EleutherAI LM Evaluation Harness](https://ailandscape.org/tool/eleutherai-lm-evaluation-harness): Framework for few-shot language model evaluation
- [HELM](https://ailandscape.org/tool/helm): Holistic evaluation of language models
- [HumanEval](https://ailandscape.org/tool/humaneval): OpenAI's benchmark for evaluating code generation
- [LangSmith](https://ailandscape.org/tool/langsmith): Platform for debugging and evaluating LLM apps
- [MMLU](https://ailandscape.org/tool/mmlu): Massive Multitask Language Understanding benchmark
- [Ragas](https://ailandscape.org/tool/ragas): Evaluation framework for RAG pipelines

---

Part of [AI Landscape](https://ailandscape.org), an open map of the AI ecosystem. Web page: https://ailandscape.org/tool/deepeval · Index for AI assistants: https://ailandscape.org/llms.txt
