TOPIC GUIDE
LLM observability and evaluation
Langfuse, Braintrust, Arize Phoenix, Promptfoo, W&B Weave and Helicone for tracing, evaluating and testing AI applications.
- products
- 11
- things to compare
- 3
- updated
- 2026-10-02
Once an AI feature is in production, you need to see what it did and know whether a change made it better. Observability tools trace each model and tool call; evaluation tools score outputs against test cases; security tools probe for prompt injection and data leaks.
What to compare
- Instrument one real workflow and check that traces show every model and tool step.
- Turn real failures into test cases and run them before each release.
- Compare self-hosting options and licenses if your data cannot leave your infrastructure.
This is a focused selection, not an exhaustive ranking. See each profile for evidence and review scope.
Projects to explore
Langfuse
An open-source platform for tracing, evaluating and improving LLM apps and agents, with prompt management; part of ClickHouse since 2026.
Read the project profileBraintrust
An observability and evaluation platform for AI agents that traces production, runs evals and catches regressions.
Read the project profileArize Phoenix
Arize's source-available platform for tracing, evaluating and debugging LLM applications, which you can run yourself.
Read the project profilePromptfoo
An open-source tool and platform for red teaming, security testing and evaluating LLM applications; being acquired by OpenAI.
Read the project profileW&B Weave
Weights & Biases' toolkit for tracing, evaluating and comparing LLM applications; Weights & Biases is part of CoreWeave.
Read the project profileHelicone
An AI gateway and LLM observability platform for routing and monitoring AI applications.
Read the project profilePortkey
A production stack for generative AI applications across an organization, including a gateway for model access.
Read the project profileLangChain
A Python framework for composing model calls, tools, and agent behavior through a configurable agent harness.
Read the project profileLaminar
An open-source platform to trace, evaluate and debug AI agents, highlighting where an agent failed.
Read the project profileOpik
Comet's open-source platform for tracing, evaluating and testing LLM applications and agents.
Read the project profileDeepEval
An open-source framework for testing LLM applications, with 50+ metrics for agents, RAG and chatbots that run like unit tests.
Read the project profile