Skip to content

Supabase observability surface: trace headers, health advisors, log sources

How a request made by supabase-js becomes a row in the project logs, when the platform’s Health Check Advisors raise an error-rate lint, and how much of that surface a script can read or write from outside. It is for whoever has to wire tracing, alerting or log export on a Supabase project and wants to know which documented behaviour held when it was run.

All measured rows below ran on 2026-10-10 from one macOS machine in Singapore (Bun 1.3.14, a current Node release for the esbuild bundles, supabase CLI 2.120.0), against throwaway projects in ap-southeast-1 on a Pro-plan organisation, driven through the Management API from the same machine. Three runs are labelled run1, run2 and run3 in the lab’s RUNLOG; every cell is n = 1 unless a row says otherwise. Rows marked documented come from the public pages cited at the claim and were not run. The lab’s raw captures carry project refs and were not published, so the figures are quoted from the RUNLOG of the observability-surface experiment, not from a published out/ artifact.

TL;DR:

  • supabase-js 2.111.0, 2.112.0 and 2.117.3 attach traceparent to REST and Edge Function calls made inside an active span, and never tracestate or baggage with the default provider.1 A bundle run with no node_modules loses the header on 2.111.0 only, with no warning; 2.112.0 and 2.117.3 kept it in all five packagings.
  • The client’s trace id lands in edge_logs as the trace_id attribute (38 of 38 rows carry one). function_edge_logs (0 of 39) and function_logs (0 of 78) do not; the id reaches function_logs only if the function prints the traceparent it received.
  • The Health Check Advisors fired on at least 5 failing 5xx requests in each of two consecutive five-minute buckets started at a UTC boundary, at shares from 1% to 100%, and did not fire on 4 or fewer. The lint text says “at least 10%”.2
  • A lint appeared 35-66 s after the second bucket closed, the answer is cached for about 60 s, and it cleared 306-367 s after the last failing request.
  • The logs endpoint returned 36 of 36 marked requests per source within 30 s (15 s polling) in a 36-minute window, and carried 11 sources.
  • supabase notebooks pull and push round-trip by notebook name, keep cell ids, and leave a project notebook with no local file alone under --yes with a closed stdin.
  • Log drains were not measured: the v2 route answered 403 on a Pro and a Team organisation, and the v1 API publishes no route.

You want toSurfaceHeld on 2026-10-10
Correlate a client span with a gateway log rowtraceparent from supabase-js, read as edge_logs.trace_idyes, 14 of 14 sending variants
Correlate a client span with an Edge Function log linetraceparent received by the functiononly if the function logs it
Be told that a service is failingHealth Check Advisors, POST /v2/projects/{ref}/advisors/runafter 5 failures in each of two buckets
Read logs from a scriptanalytics/endpoints/logs, see the logs endpoint guide11 sources, 15 s poll ceiling
Keep dashboard notebooks in gitsupabase notebooks pull / pushyes, by notebook name
Ship logs to your own endpointlog drainsnot measured

With tracing on, supabase-js attaches traceparent, tracestate and baggage to requests for Supabase domains only, and the trace id then shows in API Gateway and Edge Function logs.13 From 2.112.0 the OpenTelemetry integration is the opt-in @supabase/supabase-js/tracing import; without it the client warns once and sends no headers. Releases 2.106.0 to 2.111.x load @opentelemetry/api dynamically and send nothing, silently, when it is missing.1 Before 2.112.3 an unsampled span sends no headers; from 2.112.3 it sends traceparent only.1

The probe used one client built on @opentelemetry/sdk-trace-node with the default W3C propagator, rendered once per variant against one project with a table and an Edge Function that logs the traceparent it receives. Releases came from npm aliases: 2.111.0 (the only stable 2.111.x on the registry on 2026-10-10), 2.112.0 and 2.117.3 (the latest that day). Headers were read at the client’s own fetch boundary, so “sent” means the SDK attached them, not that a server kept them.

A cell is traceparent attached to a REST call and to an Edge Function invoke made inside an active span; both calls gave the same answer in every cell. The same matrix came out in run1, run2 and run3. tracestate and baggage were attached in no cell.

ReleaseUnbundled (Bun)Bun bundle, next to node_modulesBun bundle, no node_modulesesbuild bundle (Node), next to node_modulesesbuild bundle (Node), no node_modules
2.111.0sentsentnone, no warningsentnone, no warning
2.112.0sentsentsentsentsent
2.117.3sentsentsentsentsent

The Bun bundles ran with Bun’s runtime auto-install disabled. The same module without that setting showed 2.111.0 sending from the no-node_modules directory; that Bun fetched @opentelemetry/api from the registry at run time is an inference from the setting changing the outcome, and the fetch was not observed. The reading: the loss on 2.111.0 needs a bundle in which the SDK’s runtime import("@opentelemetry/api") cannot resolve, and 2.112.0 and later carry the import statically through the /tracing subpath and kept the header.

ControlResult
2.112.0 without the /tracing importno headers, one console.warn beginning “tracePropagation is enabled but the tracing runtime is not loaded” (as documented)
2.112.0, unsampled spanno headers (documented for releases before 2.112.3)
2.117.3, unsampled spantraceparent sent (documented for 2.112.3 and later); tracestate and baggage not asked about in this row
2.117.3 with tracePropagation offnone
REST call with no active spannone, in every variant

The wrapped fetch was exercised with a fake transport underneath, so nothing left the machine.

URL host handed to the wrapped fetchHeader attached
third-party.example.testnone, all three releases
xsupabase.conone, all three releases
supabase.co.example.testnone, all three releases
any other host under .supabase.cotraceparent, all three releases
the base URL of a client created against a non-Supabase host (the self-hosted shape)traceparent, on that host

The last row was observed; the explanation, that the SDK adds the base URL’s hostname to the default targets, was read in the installed 2.117.3 dist/index.mjs (getDefaultPropagationTargets) and not tested separately. The docs page lists *.supabase.co, *.supabase.in and localhost as the defaults and does not mention the exception.1 Not measured: a custom domain, redirects, and a custom fetch that rewrites the URL after the SDK’s check (from the source, the check reads the URL before the custom fetch runs; not from a run).

Run run3, 19 variants, 57 trace ids.

supabase-js(traceparent)API gatewayedge_logstrace_id attributeEdge Function(receives traceparent)function_edge_logsno trace idfunction_logsonly if the function logs itconsole.log

Text version: the gateway writes the client’s id to edge_logs; the function receives it; function_edge_logs records the invocation without it; function_logs holds it only when the function’s own console.log wrote it.

SourceRows with a trace_id
edge_logs38 of 38
postgrest_logs0 of 106
function_edge_logs0 of 39
function_logs0 of 78
auth_logs0 of 43
storage_logs0 of 3
realtime_logs0 of 5
pgbouncer_logs0 of 96
postgres_logs0 of 17
  • REST: the client’s trace id came back as the trace_id attribute of an edge_logs row in 14 of 14 variants where the header was sent, and in none of the 5 where it was not. Every edge_logs row has a trace_id, including requests with no client header.
  • Edge Function: the function received a traceparent carrying the client’s trace id in all 14 sending variants (the function echoes it). A search of every attribute key and the message of every source for six client trace ids (the REST id and the function id of each of the three unbundled variants) found the three REST ids as the trace_id attribute of edge_logs rows, and the three function ids only inside the function_logs message that the function’s own console.log wrote.
  • A function call made with no active span still reached the function with a platform-generated traceparent (trace flags 00) and a baggage entry sb-request-id, in every variant.

The Edge Function result is the one row where measurement and the docs page differ: “trace id appears in Edge Function logs” held only for what the function prints itself. It is one project and one run, and the platform may stamp function rows under a key or source this query did not cover.

Not measured: browser bundles, webpack, Vite and Turbopack, Realtime frames, the Swift, Dart and Python SDKs, Sentry and Datadog propagators, a tracestate or baggage value set by the app.


The platform ships four checks, log_data_api_error_rate_high, log_auth_error_rate_high, log_storage_error_rate_high and log_edge_function_error_rate_high, “read from log data”; POST /v2/projects/{ref}/advisors/run returns them, results are cached, and an empty result means every check ran and found nothing.2 The lint text the platform returns says 5xx for at least 10% of requests across two consecutive five-minute periods, and that failures must persist in both.

The probe used one project per arm, arms concurrent. Each arm waited for the next UTC five-minute boundary, sent at fixed rates for two buckets (one bucket for the single-bucket arms), and polled the four health lints every 30 s until the lint cleared or 8 minutes after traffic ended. The 5xx were real service answers (see Inducing a 5xx). run1 had 11 arms, run2 8 and run3 5.

ArmRequests per bucketFailing per bucketShareFired
data, 100%100100100%yes
data, 50%1005050%yes
data, 12%1001212%yes
data, 8%10088%yes
data, 5% (run2, repeated in run3)10055%yes, yes
data, 4%10044%no
data, 3% (run2, repeated in run3)10033%no, no
data, 2%25052%yes
data, 1%about 600 (598 sent)61%yes
data, 1%10011%no
data, 100%, one request a minute55100%yes
data, 20%, one request a minute5120%no
data, 100%, one request per bucket11100%no
data, 404 on every request1000 (all 404)0% 5xxno
data, 100% failing in bucket 1 only, healthy bucket 2100100 then 0-no
data, healthy bucket 1, 100% failing in bucket 21000 then 100-no
data, 100% failing for one bucket, no second bucket100100-no
Auth admin create, 100% 500100100100%yes
Storage list, 100% 500100100100%yes
Edge Function returning 500100100100%yes
Edge Function that throws100100100%yes

The lint fired in every arm with at least 5 failing requests in each of two consecutive five-minute buckets and in no arm with 4 or fewer, at any share from 1% to 100%. On this project the operative rule looks like a failing-request count of about 5 per bucket, not the 10% share in the lint text; 6 of 600 (1%) fired while 1 of 100 (1%) did not. This is an inference from the rows above. Not run: exactly 4 against 5 at other volumes, a count of 5 at thousands of requests per bucket (whether a share rule applies on top of the count at that volume), and a start part-way through a bucket, so the clock alignment of the buckets is the design of the probe rather than a separated finding.

Thirteen firings across the three runs.

QuantityMeasuredResolution
First lint after the second bucket closed35-36 s in run1 (9 arms), 65-66 s in run2 and run3 (4 arms)30 s polling, so each figure is an upper bound within 30 s
Cache refresh (observed_at steps)median 86-91 s per arm at 30 s pollingthe next 30 s poll after a refresh
Cache refresh, one project polled every 10 s62-64 s steps, four refreshes (exploratory, not archived)about 60 s
Lint gone after the last failing request306-337 s in 12 arms, 367 s in the 8% armfirst poll after the next bucket boundary plus the cache

The lint was absent at every poll before the second bucket had closed. An immediate second call at baseline returned the same empty list in all 24 arms. The detail text at firing reads Failing: <service> (<share>% of <n> requests failing), where <n> is the request count of the most recent bucket (100, 250 and 600 in the arms above); for Edge Functions the service field is the function path. level was ERROR and categories was HEALTH. Arms sending 100 requests per bucket reported 100 and the 5-per-bucket arm reported 5; the 600-per-bucket arm reported 600 against 598 sent. In a first exploratory run (60 then 500 failing requests, not archived) the lint reported 538 requests against 560 sent, so a small fraction of requests may go uncounted. No other lint appeared in any arm and no advisor_check_unavailable was seen.

ServiceWay to get a 5xx
Data API (PostgREST)a function that runs raise sqlstate 'PT500'
Autha raising BEFORE INSERT trigger on auth.users, then an admin create
Storagea raising function used by a storage.objects policy, then a list
Edge Functionsa function returning 500, or one that throws

A 404-only load did not fire, so a 5xx has to come from the service itself.

ClaimDocumentedMeasured 2026-10-10Filed upstream
Advisor firing thresholdat least 10% of requests, two consecutive five-minute periods5 failing requests in each of two buckets fired at 1%, 2% and 5%; 4 or fewer did notnot checked; no report is recorded in the lab RUNLOG
tracestate and baggageattached with traceparent1not attached by the default provider in any cellnot checked
Trace id in Edge Function logsappears in Edge Function logs3absent from function_edge_logs and function_logs unless the function prints itnot checked
Default propagation targets*.supabase.co, *.supabase.in, localhostalso the client’s own base URL hostnot checked

Not measured: the other lints in the v2 enum, Realtime, projects with real user traffic, other regions and plans, the Studio Health tab.


The query shape, the source column and the traps of the endpoint are in Query Supabase project logs through the Management API; this section holds only what the observability probe added.

Canary design (run1): one project, one marked REST request (lands in edge_logs) and one marked Edge Function invoke (lands in function_edge_logs by URL and in function_logs by the console line) every minute for 36 minutes, 00:48-01:24 UTC. A poller asked the logs endpoint for every marker every 15 s; the first poll that returned a marker is its first-seen time. The public pages make no claim about ingestion lag.

SourceMarkers returnedLag p50p90p99MaxOver 240 s
edge_logs36 of 3615 s15.1 s30 s30 s0
function_edge_logs36 of 3615 s15.1 s15.1 s15.1 s0
function_logs36 of 3615 s15.1 s15.1 s15.1 s0

The 15 s poll is the resolution: nearly every marker came back on the first poll after it was sent, so these figures are ceilings within one poll; the distribution below 15 s was not resolved. The row’s own timestamp minus the send time was 0 to 1 s. There were 142 polls and none returned an error. A lag above 4 minutes recorded once in a separate lab experiment did not recur in this window, and one 36-minute window supports no alert threshold.

  • Sources on the project after one request of each kind: auth_audit_logs, auth_logs, edge_logs, function_edge_logs, function_logs, pgbouncer_logs, postgres_logs, postgrest_logs, realtime_logs, storage_logs, supavisor_logs (11). The run opened one connection through the shared pooler on 6543 and on 5432 and joined one websocket. Which source appeared when was not measured, and neither was whether pgbouncer_logs rows exist before any connection.
  • Message search of this run’s own markers: the failing pooler statement’s literal in postgres_logs (2 rows, one per port); the storage bucket and object name in storage_logs (2 rows) and edge_logs (1); the Auth user email in auth_logs and auth_audit_logs (1 each); the Realtime topic name in no source.
  • A 300-request REST burst, all answered 200, returned 300 rows from the logs endpoint.
  • usage.api-requests-count returned 341, equal to the project’s edge_logs row count: the REST canary 36, the burst 300 and 5 further rows (not itemised). The 37 function invocations sit in function_edge_logs, in neither figure.
  • usage.api-counts returned per-minute total_rest_requests, total_auth_requests, total_storage_requests and total_realtime_requests buckets (HTTP 200).
  • GET /platform/organizations/{org}/usage with a personal access token answered 401 Unsupported access token, and the v1 OpenAPI document has no organisation usage path, so organisation-level logs GB figures were not reachable by token and no known log volume was compared with them.

Not measured: logs per GB, query-quota billing, the Dashboard Logs Explorer, load above one request a minute other than the burst, other regions. No throttling was met at 4 logs queries a minute; the advisor arms ran alongside without a 429 seen by this module.


The CLI help (2.120.0) and the v2 OpenAPI document say pull writes project notebooks to supabase/notebooks, keeping existing files unless an id is given; push writes local files and asks about project notebooks the directory lacks; an update replaces the body, a cell echoing its id keeps its identity, and a cell without one is added. Measured on run2, one project, one pass:

StepResult
Create through POST /v2/projects/{ref}/notebooks with markdown, database and log cellsserver-assigned ids on all 3 cells
pullone file ob05-alpha.json (file name is the notebook name); top-level keys description, favorite, content, no id and no name; cells equal to the API’s
Edit one sql, append a cell without an id, push”0 created, 1 updated”; the API held 4 cells: the edited sql, the 3 original ids, an id on the new cell, updated_by set
Second pull on the unchanged treewrote nothing (“Kept 1 existing local notebook(s) unchanged”)
pull <id> over a locally modified filerestored the project’s version
New local file, pushcreated on the project (“1 created, 1 updated”)
Local file for a project notebook removed, push --yes, stdin closedupdated the remaining notebook, reported “1 project notebook(s) are not in …: ob05-alpha - Left alone - rerun interactively to resolve”; the notebook stayed on the project

The CLI exit code was 0 in every call, including the one that left a notebook alone, so a script has to read the output to see it. Not measured: the choices the interactive prompt offers, notebooks with a read-replica database_identifier, running a notebook (the v2 OpenAPI document has no run endpoint), concurrent edits from the Dashboard.


Drains are documented on Pro, Team and Enterprise; an HTTP drain posts a JSON array in batches of at most 250 events or one second, with optional gzip and HTTP/1 or HTTP/2.45 On the Pro organisation, GET /v1/organizations/{slug}/entitlements read log_drains hasAccess true and audit_log_drains false. GET and POST /v2/projects/{ref}/analytics/log-drains both answered 403 forbidden, “Your organization does not have access to this API”, on that Pro organisation’s project (archived in the RUNLOG), and on a Team organisation’s project in an exploratory check that was not archived. The v1 API publishes no log-drains route; the v2 document lists /v2/projects/{ref}/analytics/log-drains and its /{id} form.6 What grants access to the v2 route was not determined.

No drain was created, so batch size, flush spacing, the content-encoding the platform sends, HTTP version, event fields, arrival lag and arriving sources are unmeasured. A Worker sink with a SQLite-backed Durable Object was tested on its own (a plain POST and a gzip POST stored, the gzip body decoded to [{"probe":"gzip"}], HTTP/1.1 seen, a POST without the key answered 401) and waits for a drain. Residency consequences are in Supabase data residency and sovereignty.


ClaimMeasured or documentedHow it was checked
2.111.0 loses traceparent in a bundle with no node_modules; 2.112.0 and 2.117.3 do notmeasured, run1 to run3, n = 1 per cellrecording fetch at the client boundary, five packagings per release
Registry fetch explains the Bun 2.111.0 difference when auto-install is left oninferencethe setting changed the outcome; the fetch was not observed
Default propagation targets include the client’s base URL hostsource read, one releasegetDefaultPropagationTargets in 2.117.3 dist/index.mjs; the behaviour was also seen on a self-hosted-shaped base URL
edge_logs.trace_id on 38 of 38 rows; other sources 0measured, run3, one projectper-source counts of rows with a trace_id
5 failing requests in each of two buckets fire the lint, at 1% to 100%measured, 13 firings, n = 1 per arm except where repeatedrunAdvisors polling every 30 s, per-arm projects
Lint cached about 60 smeasured, exploratory, not archived10 s polling on one project; the archived 30 s runs show 86-91 s steps
Ingestion lag at most 30 smeasured, one 36-minute window, 15 s resolutionmarker per minute, first-seen time
Notebook push and pull behaviourmeasured, run2, one project, one passsupabase CLI 2.120.0 against the v2 notebooks API
v2 log-drain route 403measured on a Pro org and a Team org (neither artifact published)GET and POST, PAT
Documented behaviours cited abovedocumented, read 2026-10-10the footnoted pages

Two limits on generalising. Every number is one region, one plan tier and an idle project, so the advisor count threshold and the lag ceiling describe this probe’s load and are no platform contract. The firing rule in particular has no run above 600 requests per bucket.


PracticeEvidenceModule
Use supabase-js 2.112.0 or later where a bundle runs without node_modules.2.111.0 sent no traceparent and no warning from a Bun bundle and an esbuild bundle with no node_modules; 2.112.0 and 2.117.3 sent in all five packagings.OB01, run1 to run3
Import @supabase/supabase-js/tracing on 2.112.0 and later.Without it: no headers and one console.warn.OB01a
Query edge_logs for the trace_id attribute.38 of 38 rows carry one; the client’s value matched in 14 of 14 sending variants.OB01b, run3
Log the received traceparent in the function when function_logs must carry it.function_edge_logs 0 of 39 and function_logs 0 of 78 rows carried a trace id; function ids appeared only in the function’s own console.log message.OB01d, OB01e
Send at least 5 failing 5xx in each of two aligned buckets when testing an advisor.Fired at 5 of 100, 5 of 250 and 6 of 600; did not fire at 4 of 100, 3 of 100 (twice), 1 of 5 or on a 404-only load.OB02
Do not expect an advisor from one failing bucket.A single failing bucket, or failures in only one of two buckets, did not fire.OB02
Poll advisors/run no faster than every 60 s.The returned observed_at stepped 86-91 s at 30 s polling and 62-64 s at 10 s polling; the 10 s figure is exploratory, not archived.OB02
Expect a lint 35-66 s after the second bucket and a tail up to 367 s.Detection 35-66 s at 30 s polling; gone 306-367 s after the last failure.OB02
Name columns in logs queries and filter on source.select * answered a backend error in an exploratory query (not archived); the logs guide records the same trap.OB03, exploratory
Size a logs poller for a ceiling of 30 s at 15 s polling.36 of 36 markers per source returned, max 30 s; one 36-minute window, so no lag below 15 s was resolved.OB03
Read supabase notebooks push output for its result; the exit code carries none.Exit 0 on every call, including “Left alone” for a notebook with no local file.OB05, run2
Expect to set up log drains by hand until the v2 route accepts the organisation; Dashboard setup was not tried here.GET and POST /v2/projects/{ref}/analytics/log-drains answered 403 on a Pro and a Team org; v1 has no route.OB04, one pass

The lab’s captures for these runs carry project refs and were not published to out/, so each row has no artifact; the figures are in the experiment’s RUNLOG.

ModuleExperimentTestArtifact
OB01observability-surfaceob01-trace-propagation.tsnone published
OB02observability-surfaceob02-health-advisors.tsnone published
OB03observability-surfaceob03-log-canary.tsnone published
OB04observability-surfaceob04-log-drain.tsnone published
OB05observability-surfaceob05-notebooks.tsnone published
  1. Supabase, “Client-side tracing,” Supabase Docs. https://supabase.com/docs/guides/observability/client-side-tracing ↩ ↩2 ↩3 ↩4 ↩5 ↩6

  2. Supabase, “Health Check Advisors,” Supabase Changelog, 2026-09-18. https://supabase.com/changelog/50577-health-check-advisors ↩ ↩2

  3. Supabase, “Connect client traces to your logs,” Supabase Blog. https://supabase.com/blog/connect-client-traces-to-your-logs ↩ ↩2

  4. Supabase, “Log drains,” Supabase Docs. https://supabase.com/docs/guides/observability/log-drains ↩

  5. Supabase, “Log drains now available on Pro,” Supabase Blog. https://supabase.com/blog/log-drains-now-available-on-pro ↩

  6. Supabase, “Management API OpenAPI document (v1),” Supabase. https://api.supabase.com/api/v1-json ↩