Arize Phoenix

Listed at https://phoenix.arize.com

Overall Rank #12 ⭐ Consider
✅ Open-source self-host + Phoenix Cloud (China SaaS access requires proxy) | 🌍 International

💰 Token Pricing

TypePriceNote
Input Cloud 免费层: 10 GiB 存储/工作空间;自托管: $0(仅基础设施费) per million tokens
Output Phoenix Cloud 付费: 按量付费(存储计费);Arize AX 企业: 合同制 per million tokens
💡 Free Credits: Phoenix Cloud free tier: 10 GiB storage per workspace (no time limit)

🤖 Supported Models (100)

OpenTelemetry 原生 trace collector,支持任意 LLM 框架Phoenix Cloud 托管 + 自托管 (Apache-2 / Elastic-2.0)

✨ Pros

  • Only Elastic-2.0 open-source (10,600+ stars) OpenTelemetry-native LLM observability platform
  • 17+ LLM frameworks auto-instrumented: OpenAI / Anthropic / LangChain / LlamaIndex / vLLM / Ollama
  • Phoenix Cloud free tier: 10 GiB storage, no time limit
  • Three deployment modes: Phoenix Cloud / self-host / Arize AX enterprise
  • Full trace + evaluations + datasets + experiments + prompt playground + PXI agent
  • PXI (Phoenix Intelligence) AI engineering agent (BETA) for natural-language trace investigation

⚠️ Cons

  • ×No request-side caching or fallback routing (observation-only)
  • ×Cloud free tier capped at 10 GiB storage; high-volume use needs paid tier or self-host
  • ×PXI still in BETA — can produce wrong answers on edge cases
  • ×UI slightly less polished than closed-source peers like LangSmith
  • ×China access to Phoenix Cloud requires proxy (self-host bypasses this)
  • ×Self-host requires PostgreSQL + Python app server ops

🎯 Best For

AI engineering teams wanting maximum control and minimum vendor lock-in; platform teams needing OpenTelemetry-native trace backends; compliance-sensitive (healthcare/finance/gov) self-hosted deployments; teams evaluating many prompt versions / RAG pipelines / agent workflows

💰 Pricing & Plans

PlanPriceStorage / ScaleBest For
Phoenix Cloud Free$010 GiB per workspace, no time limitEvaluation, open-source dev, small teams
Phoenix Cloud ProPay-as-you-go (storage-based)Unlimited storage / higher trace volumeProduction LLM apps, larger teams
Self-hosted (open source)$0 (infrastructure only)Your own PostgreSQL + computeCompliance-sensitive / maximum control
Arize AX EnterpriseCustom contractSSO/RBAC/audit log, dedicated supportLarge orgs needing governance

🔧 API & Developer Experience

  • OpenTelemetry Native: Phoenix is an OpenTelemetry-native trace backend; you can throw OTLP traces from any instrumented app straight at it without vendor-specific SDKs.
  • Auto-instrumentation: 1-line wrappers or environment variables auto-capture traces for 17+ frameworks: OpenAI, Anthropic, LangChain, LlamaIndex, vLLM, Ollama, DSPy, and more.
  • Traces, Spans & Evaluations: Deep span/trace explorer for LLM calls, plus first-class evaluation tooling (LLM-as-judge, custom scorers, human review) that runs against captured traces.
  • Prompt Playground: Iterate prompts and compare model outputs live; experiment tracking records prompt versions alongside evaluation scores.
  • PXI Agent (BETA): Phoenix Intelligence lets you ask natural-language questions about your traces ('where are failures spiking?') and get investigation answers grounded in your data.
  • Datasets & Experiment Hub: Versioned datasets, regression testing, and experiment comparison built into the platform — closest to a CI/CD loop for prompt and RAG changes.
  • Open Licenses & Local UI: Self-host ships Apache-2/Elastic-2.0 licensed server + local web UI (no telemetry), so raw trace data never leaves your VPC.

🧩 OpenTelemetry-Native LLM Observability

Arize Phoenix stands out because it is one of the few LLM observability platforms built directly on the OpenTelemetry standard rather than a proprietary instrumentation format. Teams already emitting OTLP traces can feed Phoenix immediately — or use one-line auto-instrumentation for 17+ frameworks — without re-architecting their stack or adopting a vendor lock-in SDK. This means the same trace pipeline that serves your general microservices can also power deep LLM-specific analysis: span trees, token/cost per call, prompt–completion pairs, and retrieval quality for RAG pipelines. Phoenix folds evaluation directly into the observability loop. Instead of exporting traces to one tool and running evals in another, you capture traces, score them with LLM-as-judge or custom scorers, log datasets, and iterate prompts in one place. The optional PXI agent adds natural-language investigation over the whole corpus. Combined with an Elastic-2.0 open-source self-host option, it delivers production-grade tracing, evals, and experiment tracking with minimal vendor lock-in — a strong fit for platform teams building internal LLM infrastructure.

🌐 China Access & Latency

Arize Phoenix has a clear advantage for China-based teams: the core product is open source and self-hostable, so mainland teams can run the Phoenix server on a domestic cloud or on-premises without proxying to any overseas SaaS. The self-hosted server ships an Apache-2/Elastic-2.0 licensed local web UI with no external telemetry, and trace data stays inside the VPC — ideal for compliance-sensitive workloads subject to Chinese data-residency rules. Because it is an observability backend rather than a hosted model API, there is no per-token egress cost and no China-specific rate-limiting at the provider edge. The trade-off appears only if you use Phoenix Cloud (the managed SaaS hosted overseas): reaching it from mainland China then requires a VPN or an overseas relay, with typical added latency in the hundreds of milliseconds. For teams that want the managed experience with a Chinese-accessible path, self-hosting Phoenix on a domestic cloud and connecting clients locally is the standard, effectively latency-free approach.