# EleutherAI LM Evaluation Harness

> Framework for few-shot language model evaluation

- Category: [Observability & Evals](https://ailandscape.org/category/observability-evals) › LLM Evaluation & Benchmarking
- Homepage: https://github.com/EleutherAI/lm-evaluation-harness
- Repository: https://github.com/EleutherAI/lm-evaluation-harness
- Tags: evaluation, benchmarks, llm
- Added to the landscape: 2026-03-18

## Similar tools in LLM Evaluation & Benchmarking

- [Braintrust](https://ailandscape.org/tool/braintrust): LLM evaluation and prompt management platform
- [DeepEval](https://ailandscape.org/tool/deepeval): Open-source LLM evaluation framework
- [HELM](https://ailandscape.org/tool/helm): Holistic evaluation of language models
- [HumanEval](https://ailandscape.org/tool/humaneval): OpenAI's benchmark for evaluating code generation
- [LangSmith](https://ailandscape.org/tool/langsmith): Platform for debugging and evaluating LLM apps
- [MMLU](https://ailandscape.org/tool/mmlu): Massive Multitask Language Understanding benchmark
- [Ragas](https://ailandscape.org/tool/ragas): Evaluation framework for RAG pipelines

---

Part of [AI Landscape](https://ailandscape.org), an open map of the AI ecosystem. Web page: https://ailandscape.org/tool/eleutherai-lm-evaluation-harness · Index for AI assistants: https://ailandscape.org/llms.txt
