# Observability & Evals

> LLM observability and evaluation: tracing and monitoring with Langfuse, Helicone, and Arize, eval frameworks like Braintrust, DeepEval, and LangSmith, plus explainability tools and prompt management. Tracing shows what your models do in production; evals tell you whether they do it well. Community model rankings live in Benchmarks & Leaderboards.

27 tools in 5 subcategories. Web page: https://ailandscape.org/category/observability-evals · Each tool also has a markdown version at /tool/{slug}.md

## LLM Observability

- [AgentOps](https://ailandscape.org/tool/agentops): Observability and debugging platform purpose-built for AI agents
- [Helicone](https://ailandscape.org/tool/helicone): Open-source LLM observability and proxy
- [Langfuse](https://ailandscape.org/tool/langfuse): Open-source LLM engineering platform for tracing and evals
- [Opik](https://ailandscape.org/tool/opik): Open-source LLM evaluation and tracing by Comet
- [Phoenix (Arize)](https://ailandscape.org/tool/phoenix-arize): Open-source AI observability and evaluation platform
- [Traceloop](https://ailandscape.org/tool/traceloop): OpenTelemetry-based LLM observability

## Model Monitoring

- [Arize AI](https://ailandscape.org/tool/arize-ai): ML observability and monitoring platform
- [Evidently](https://ailandscape.org/tool/evidently): ML model monitoring and data quality checks
- [Fiddler AI](https://ailandscape.org/tool/fiddler-ai): ML model performance management
- [WhyLabs](https://ailandscape.org/tool/whylabs): AI observability and data monitoring

## LLM Evaluation & Benchmarking

- [Braintrust](https://ailandscape.org/tool/braintrust): LLM evaluation and prompt management platform
- [DeepEval](https://ailandscape.org/tool/deepeval): Open-source LLM evaluation framework
- [EleutherAI LM Evaluation Harness](https://ailandscape.org/tool/eleutherai-lm-evaluation-harness): Framework for few-shot language model evaluation
- [HELM](https://ailandscape.org/tool/helm): Holistic evaluation of language models
- [HumanEval](https://ailandscape.org/tool/humaneval): OpenAI's benchmark for evaluating code generation
- [LangSmith](https://ailandscape.org/tool/langsmith): Platform for debugging and evaluating LLM apps
- [MMLU](https://ailandscape.org/tool/mmlu): Massive Multitask Language Understanding benchmark
- [Ragas](https://ailandscape.org/tool/ragas): Evaluation framework for RAG pipelines

## Explainability & Alignment

- [Captum](https://ailandscape.org/tool/captum): Model interpretability library for PyTorch
- [ELI5](https://ailandscape.org/tool/eli5): Debug and explain ML classifiers and regressors
- [InterpretML](https://ailandscape.org/tool/interpretml): Microsoft's toolkit for model interpretability
- [LIME](https://ailandscape.org/tool/lime): Local interpretable model-agnostic explanations
- [SHAP](https://ailandscape.org/tool/shap): Game theoretic approach to explain ML model outputs

## Prompt Management

- [Guidance](https://ailandscape.org/tool/guidance): Constrained generation and prompt control by Microsoft
- [PromptFlow](https://ailandscape.org/tool/promptflow): Suite of development tools for LLM apps
- [Promptfoo](https://ailandscape.org/tool/promptfoo): Test, evaluate, and red-team LLM prompts
- [PromptLayer](https://ailandscape.org/tool/promptlayer): Prompt versioning and management platform

---

Part of [AI Landscape](https://ailandscape.org), an open map of the AI ecosystem. Index for AI assistants: https://ailandscape.org/llms.txt
