BetaThe Langfuse and Arize Phoenix event trackers are in beta. They read directly from each tool’s own database, and those schemas are not stable APIs. See Caveats before relying on the generated SQL in production.
Why combine them
Tracing tools are built for observability, not for experimentation. Langfuse’s own A/B testing guide, for example, picks a prompt label withMath.random() and compares the two groups by eye in the Langfuse UI. That is fine for a quick look, but it gives you no consistent assignment (the same user can get a different prompt on every request), no targeting, no traffic control, and no answer to “is this difference real?”.
GrowthBook adds the pieces a tracing tool leaves out:
- Deterministic assignment. The same user (or session, or request) always gets the same variation, hashed on an attribute you choose.
- Targeting and rollout. Limit the test to a segment, ramp traffic, or stop it from the GrowthBook UI without a deploy.
- Statistics. Frequentist or Bayesian results, sequential testing, CUPED, and guardrails on top of the tracing data.
- Bandits. Let GrowthBook shift traffic toward the winning prompt automatically, using an eval score or cost as the reward.
How it fits together
The tag contract
The only link between GrowthBook and the trace is a tag on the trace of the form:gb.exp:support-prompt=prod-b. The tracing plugin in the JavaScript SDK builds these tags for you every time an experiment is evaluated; you attach them to the trace with your tracing SDK. The generated exposure queries in GrowthBook then split the tag on the first = to recover the experiment and variation for each trace.
Two rules follow from that:
- Experiment keys must not contain
=. The plugin skips any assignment whose experiment key contains it, so such experiments would silently record no exposures. - Evaluate your flags and attach the tags at the top of the traced request, before any child spans are created. The exposure query reads tags from the trace (Langfuse) or the root span (Phoenix); tags set only on a nested span are not seen.
Identifiers must match
GrowthBook joins exposures to metrics on an identifier, so the id your tracing SDK records must be exactly the value GrowthBook hashed on:- The tracer’s user id (Langfuse
userId, OpenInferenceuser.id) must equal the value of the experiment’s hash attribute (labelled Assign variation by attribute on a feature rule, or Hashing attribute on the experiment page). Hash on auser_idattribute so its name matches theuser_ididentifier type in the generated queries, and pass that same value as the user id. - The tracer’s session id (Langfuse
sessionId, OpenInferencesession.id) must likewise equal the GrowthBook attribute you hash on when you experiment per session.
Choosing the unit of randomization
Both trackers give the data source three identifier types. Pick the one that matches what you are testing:user_id— one variation per user across all of their requests. Use this when you care about user-level outcomes (retention, thumbs-up rate per user, total cost per user).session_id— one variation per conversation. Useful for chat products where a whole session should get a consistent prompt but different sessions of the same user can differ.trace_id— a fresh assignment for every request. To use this, set a GrowthBook attribute to the trace id (for exampleattributes: { trace_id: traceId }) and hash the experiment on that attribute. This gives you the most statistical power for per-request metrics like latency, cost, and eval scores, at the cost of inconsistency across requests for the same user.
user_id and still use session_id-keyed metrics (or the reverse).
Where the prompt lives
There are two common ways to wire the variation to an actual prompt. Both work with the tag contract above.Pattern A: the prompt lives in the tracing tool
The GrowthBook flag returns a Langfuse prompt label (or a Phoenix prompt tag), and the app fetches the prompt by that label. Prompt engineers keep iterating on the prompt text in the tracing tool; GrowthBook only decides which label each user gets.Pattern B: the prompt lives in GrowthBook
A JSON feature flag holds the whole configuration, and the variations differ in that JSON:Stamping the trace
Install thetracing plugin, then hand its tags to your tracing SDK at the top of each traced request. The plugin works in the browser, in Node.js, and with the multi-user GrowthBookClient.
JavaScript with Langfuse
Requires the Langfuse JavaScript SDK v4 or later (@langfuse/tracing).
sessionId alongside userId if you also experiment or measure per session.
JavaScript with Arize Phoenix (OpenInference)
Phoenix reads OpenInference context attributes.@arizeai/phoenix-otel exports setUser, setSession, and setTags helpers that write them onto the active OpenTelemetry context.
setSession(ctx, { sessionId }) as well when you want a session id on the trace.
Plain OpenTelemetry
If you are not using either vendor SDK, or want the assignment as a span attribute instead of a tag, use the plugin’sonAssignment hook. It fires once per unique experiment assignment.
gb.exp: tag, so if you rely on span attributes instead you will need to edit the exposure SQL to read your attribute.
Server-side Node.js with GrowthBookClient
Pass the plugin in the client options. The shared client has no assignments of its own, so the plugin re-runs for each createScopedInstance() and every scoped instance keeps its own tag set.
Python
The Python SDK does not have a tracing plugin yet, so build the tag by hand in theon_experiment_viewed callback. The callback receives experiment, result, and user_context as keyword arguments; experiment.key and result.key are the two parts of the tag.
tags list scoped to one request (for example on the request object) rather than module-level, so assignments from one user do not leak onto another user’s trace. For the multi-user GrowthBookClient, the same callback works; use user_context.attributes to find the right per-request list.
Connecting the data source
Once traces carry the tag, add the tracing tool’s database as a GrowthBook data source. On the Add data source screen, pick the event tracker and GrowthBook generates the exposure queries, identifier types, fact tables, filters, and starter metrics for you.Langfuse
Works with self-hosted Langfuse v3, which stores traces in ClickHouse.- Choose the Langfuse event tracker (marked Beta).
- Connection type is ClickHouse. Point it at the Langfuse ClickHouse database (the one with
traces,observations, andscorestables). Create a read-only ClickHouse user for GrowthBook. - Option
Langfuse project id— found in Langfuse under Project Settings. Leave blank to include every project in the database, which is what a single-project self-host usually wants.
Arize Phoenix
Works with Phoenix backed by Postgres (the default for a self-hosted or Docker deployment).- Choose the Arize Phoenix event tracker (marked Beta).
- Connection type is Postgres. Point it at Phoenix’s Postgres database with a read-only role.
- Option
Phoenix project name— the Phoenix project your traces are sent to. Defaults todefault. Leave blank to include every project.
What gets created
For either tracker the wizard generates:- Identifier types
user_id,session_id, andtrace_id. - Exposure queries for each identifier type, parsing the
gb.exp:tag. Langfuse exposures exposetrace_name,release, andversionas experiment dimensions; Phoenix exposures exposetrace_name. - An identifier join between
user_idandsession_id. - Three fact tables with filters and six starter metrics.
Langfuse Traces— one row per trace. Columns includetrace_id,user_id,session_id,timestamp,trace_name,release,version,tag_count. Metric: Traces per user.Langfuse Observations— one row per observation (generation, span, or event). Columns includeobservation_type,observation_name,model,level,latency_ms,time_to_first_token_ms,input_tokens,output_tokens,total_tokens,total_cost,prompt_name,prompt_version. Filters: LLM Generations (observation_type = 'GENERATION') and Errors (level = 'ERROR'). Metrics: LLM calls per user, LLM cost per user, LLM error rate, p95 LLM latency, Tokens per LLM call.Langfuse Scores— one row per score (evals, annotations, user feedback). Columns includescore_name,score_value,score_string_value,score_data_type,score_source,trace_name. No filters or metrics are generated; see below.
Phoenix Traces— one row per trace. Columns includetrace_id,user_id,session_id,timestamp,trace_name,status_code,latency_ms,input_tokens,output_tokens. Filter: Errors (status_code = 'ERROR'). Metric: Traces per user.Phoenix Spans— one row per span. Columns includespan_name,span_kind,status_code,model,latency_ms,input_tokens,output_tokens,total_tokens,total_cost. Filters: LLM Spans (span_kind = 'LLM') and Errors (status_code = 'ERROR'). Metrics: LLM calls per user, LLM cost per user, LLM error rate, p95 LLM latency, Tokens per LLM call.Phoenix Annotations— one row per span annotation (LLM evals, code checks, human feedback). Columns includeannotation_name,annotation_label,annotation_score,annotator_kind,span_name,trace_name. Filter: LLM Evals (annotator_kind = 'LLM'). No metrics are generated; see below.
- Traces per user — mean number of traces (requests) per unit.
- LLM calls per user — mean number of model calls per unit.
- LLM cost per user — mean of
total_costacross model calls per unit. - LLM error rate — ratio of model calls that errored to all model calls.
- p95 LLM latency — 95th percentile of
latency_msacross model calls (event-level quantile). - Tokens per LLM call — ratio of
total_tokensto model calls.
Adding eval-score metrics
Score and annotation names are defined by you (helpfulness, toxicity, thumbs_up, …), so GrowthBook cannot generate metrics for them in advance. Instead, score_name (Langfuse) and annotation_name (Phoenix) are set up as inline-filter columns, which makes a new score metric a two-minute job:
- Go to Metrics and create a new fact metric on the
Langfuse ScoresorPhoenix Annotationsfact table. - Choose the metric type. For an average score use a Ratio metric: the numerator sums
score_value(Langfuse) orannotation_score(Phoenix), and the denominator counts rows. A Mean metric would instead average each unit’s summed score across all units, counting units with no scores as zero, so it rewards how many requests were scored rather than how well they scored. For the share of units with at least one matching score, use a Proportion metric. - Add a row filter on
score_name(orannotation_name) equal to the score you want, for examplehelpfulness. For a ratio, apply it to both the numerator and the denominator.
HUMAN annotator kind.
Running the experiment
From here it is a normal GrowthBook experiment:- Create a feature flag (string flag for Pattern A, JSON flag for Pattern B) and add an Experiment rule with your variations. Set Assign variation by attribute to the attribute that matches the id you send to the tracer, typically a
user_idattribute marked as a unique identifier. - Open the linked experiment, pick the data source you just created, and choose the assignment query whose identifier matches the attribute (
user_id,session_id, ortrace_id). - Add goal metrics (an eval score, thumbs-up rate) and guardrails (LLM cost per user, p95 LLM latency, LLM error rate).
- Start the experiment. Results update on the usual schedule.
Contextual bandits
Nothing extra is needed to run a bandit or contextual bandit on an LLM feature. Any of the generated metrics, or an eval-score metric you added, can be the decision metric (reward): for example, maximize averagehelpfulness while GrowthBook shifts traffic toward the prompt that scores best, or minimize LLM cost per user across several model choices. Because the tag is stamped on every trace regardless of how the variation was chosen, exposure and reward data flow through the same tables.
Caveats
These trackers are in beta and read from databases that were not designed as public interfaces.- Langfuse’s ClickHouse schema is not a stable API. Langfuse v4 moves traces, observations, and scores into a single
eventstable. Re-validate the generated SQL after any Langfuse upgrade; the generated fact tables target the v3traces,observations, andscorestables. - Langfuse Cloud has no ClickHouse access. Use Langfuse’s scheduled blob storage export (Parquet or JSONL to S3, GCS, or Azure) into your own warehouse, connect that warehouse as a custom data source, and adapt the fact table SQL. The exported
observationsrows already carryuser_id,session_id,tags,metadata,latency,total_cost, andusage_details, so the joins back totracesare not needed there. FINALon large Langfuse installs. The Langfuse fact tables useFINALto deduplicate ReplacingMergeTree rows. On very large tables, replace it withORDER BY event_ts DESC LIMIT 1 BY idfor better performance.- Langfuse
environmentcolumn. It is omitted from the generated SQL because older installs lack it. Add it to the fact table SQL by hand if you want to slice by environment. - Phoenix’s Postgres schema is likewise unversioned. The
span_coststable only exists on newer Phoenix versions; on older installs remove theLEFT JOIN ... span_costsand thetotal_costcolumn from thePhoenix Spansfact table.trace_annotationsare not included yet, onlyspan_annotations. - Phoenix tags arrive in two shapes. The Python SDK writes
tag.tagsas a JSON array and the JavaScript SDK writes it as a JSON-encoded string. The generated SQL normalizes both, so no action is needed, but keep this in mind if you write your own queries. - Editing the tracker option later (in the data source settings) only changes the stored option. It does not rewrite SQL that has already been generated. Regenerate the resources or edit the existing SQL by hand.

