# Braintrust

> LLM evaluation and prompt management platform

- Category: [Observability & Evals](https://ailandscape.org/category/observability-evals) › LLM Evaluation & Benchmarking
- Homepage: https://braintrust.dev
- Tags: evaluation, testing, llm
- Added to the landscape: 2026-03-18

## Similar tools in LLM Evaluation & Benchmarking

- [DeepEval](https://ailandscape.org/tool/deepeval): Open-source LLM evaluation framework
- [EleutherAI LM Evaluation Harness](https://ailandscape.org/tool/eleutherai-lm-evaluation-harness): Framework for few-shot language model evaluation
- [HELM](https://ailandscape.org/tool/helm): Holistic evaluation of language models
- [HumanEval](https://ailandscape.org/tool/humaneval): OpenAI's benchmark for evaluating code generation
- [LangSmith](https://ailandscape.org/tool/langsmith): Platform for debugging and evaluating LLM apps
- [MMLU](https://ailandscape.org/tool/mmlu): Massive Multitask Language Understanding benchmark
- [Ragas](https://ailandscape.org/tool/ragas): Evaluation framework for RAG pipelines

---

Part of [AI Landscape](https://ailandscape.org), an open map of the AI ecosystem. Web page: https://ailandscape.org/tool/braintrust · Index for AI assistants: https://ailandscape.org/llms.txt
