Submit projectSubmit

TOPIC GUIDE

LLM observability and evaluation

Langfuse, Braintrust, Arize Phoenix, Promptfoo, W&B Weave and Helicone for tracing, evaluating and testing AI applications.

products
11
things to compare
3
updated
2026-10-02

Once an AI feature is in production, you need to see what it did and know whether a change made it better. Observability tools trace each model and tool call; evaluation tools score outputs against test cases; security tools probe for prompt injection and data leaks.

What to compare

  1. Instrument one real workflow and check that traces show every model and tool step.
  2. Turn real failures into test cases and run them before each release.
  3. Compare self-hosting options and licenses if your data cannot leave your infrastructure.

This is a focused selection, not an exhaustive ranking. See each profile for evidence and review scope.

Projects to explore

Langfuse

An open-source platform for tracing, evaluating and improving LLM apps and agents, with prompt management; part of ClickHouse since 2026.

Read the project profile

Braintrust

An observability and evaluation platform for AI agents that traces production, runs evals and catches regressions.

Read the project profile

Arize Phoenix

Arize's source-available platform for tracing, evaluating and debugging LLM applications, which you can run yourself.

Read the project profile

Promptfoo

An open-source tool and platform for red teaming, security testing and evaluating LLM applications; being acquired by OpenAI.

Read the project profile

W&B Weave

Weights & Biases' toolkit for tracing, evaluating and comparing LLM applications; Weights & Biases is part of CoreWeave.

Read the project profile

Helicone

An AI gateway and LLM observability platform for routing and monitoring AI applications.

Read the project profile

Portkey

A production stack for generative AI applications across an organization, including a gateway for model access.

Read the project profile

LangChain

A Python framework for composing model calls, tools, and agent behavior through a configurable agent harness.

Read the project profile

Laminar

An open-source platform to trace, evaluate and debug AI agents, highlighting where an agent failed.

Read the project profile

Opik

Comet's open-source platform for tracing, evaluating and testing LLM applications and agents.

Read the project profile

DeepEval

An open-source framework for testing LLM applications, with 50+ metrics for agents, RAG and chatbots that run like unit tests.

Read the project profile

Go a little deeper