Running a platform on the Supabase Management API
The Management API makes projects, keys, config, usage and lifecycle programmable. The docs cover the endpoints but not the platform layer on top of them: the provisioning loop, the rate budget, the metering, the OAuth paths and the pause playbook. This guide covers that layer. Every number in it was measured against the live API on 2026-08-17/18 on throwaway projects; where something is gated or untested, it says so. To follow along: a Supabase org on a paid plan, a personal access token with Management API access, and throwaway projects you will delete when done. The placement doc decides where a tenant should live; this guide is the mechanism that puts it there.
The shape being built:
1. Decide who owns the project
Section titled “1. Decide who owns the project”Two integration shapes, differing on who owns the Supabase organization and the billing relationship:
- You provision on their behalf - your org, your PAT, one project per tenant, users never see Supabase. White-label, you hold the billing relationship. The rest of this guide is this shape.
- They bring their own backend - the user’s org, connected to your app through the Management API OAuth2 flow. You get a scoped token; they keep the direct billing relationship. Section 5 has the measured surface.
You can offer both. The project-claim flow is the (gated) bridge from the
first to the second when a tenant outgrows you. On a normal paid org the
claim route answers
404 {"message":"Cannot POST /v1/oauth/authorize/project-claim"}: the route
does not exist for the credential class, which is different from being
forbidden. On a platform-plan org on 2026-08-24 the claim, oauth/apps,
and projects/{ref}/transfer routes all answered 404 as well. The
BYO-backend bridge is off by default regardless of plan, so the two shapes
are not separable in practice until the bridge is enabled for the account.
For a time-boxed, network-restricted, role-scoped grant instead of a full
project handover, the JIT database access surface works on a
platform-plan org: POST /database/jit/invite (email + role + expires_at +
allowed_networks CIDRs) -> 200 with an invite_id, and
DELETE /database/jit/invite/{id} -> 200 (measured 2026-08-24). The identical
invite on a Pro org is rejected with a 500 (A/B measured 2026-08-25), so do not
offer it there.
2. Provision a tenant
Section titled “2. Provision a tenant”One call creates a project; the rest is setup around it.
POST /v1/projectswithorganization_slug, a strong generateddb_pass, and a region. Create ->ACTIVE_HEALTHYmeasured 131-159 s on paid Micro, 12-13 s on a free org.- Smart region selection works on a normal paid org:
region_selection: {type: "smartGroup", code: "apac"}was accepted (201) and the platform pickedap-northeast-2. Defer the city choice instead of pinning per tenant. - Fresh projects refuse their first
/auth/v1/admin/userswrite after reportingACTIVE_HEALTHY: 5 of 5 with500 unexpected_failure, accepting the second call one poll later (2026-08-04), and 2 of 2 with500 "Database error checking email"for ~10 s after passing the per-service health poll (2026-08-03). PollGET /v1/projects/{ref}/health?services=auth&services=rest&services=db, then retry the first write with backoff; do not treat it as a finding. - Fetch keys with
GET /v1/projects/{ref}/api-keys?reveal=trueand select bynameORtype- new projects carry both legacy JWTs and the newsb_publishable_/sb_secret_shapes. - Configure services with the config endpoints (auth, PostgREST, storage,
realtime, functions + secrets). The OAuth server flips on with
PATCH /v1/projects/{ref}/config/auth{oauth_server_enabled: true, oauth_server_authorization_path: ...}- measured 200.
3. Know the sizing floors
Section titled “3. Know the sizing floors”| Org class | Nano | Floor | Pause |
|---|---|---|---|
| Paid (Pro) org | rejected three ways: create 400 Minimum instance size on paid plans is Micro, addon PATCH 400 addon_variant: Invalid input, absent from available_addons | Micro, always-on | 400 Project is not free-tier |
| Free org | accepted (201) - and there is NO compute addon catalogue at all | shared/free compute | full lifecycle: pause -> INACTIVE, restore wakes in 162-204 s, data API answers HTTP 540 Project paused while parked |
| Legacy free-era project inside a paid org | n/a | keeps its paused state after the upgrade | one-way door: once restored it cannot be re-paused - pause follows the org’s current plan, not the project’s lineage |
Scale-to-zero economics are part of the Supabase for Platforms
programme1. Measured 2026-08-24: an SfP organization is the platform
plan, and Nano is the platform plan’s create default - the SfP-prescribed
create (no desired_instance_size) provisions infra_compute_size: nano
(224MB shared_buffers) and the project can pause. Nano is not a select or
resize target on any plan. On a normal paid org there is no idle discount,
which is the cost premise the
tenant placement doc is built
on.
4. Budget the rate limit
Section titled “4. Budget the rate limit”Measured on a normal paid org, on a cheap scoped read:
- Every response carries
x-ratelimit-limit,x-ratelimit-remaining,x-ratelimit-reset. The limit read 120, decrementing 1:1 per call. Read the header; never infer the budget. - A deliberate burst tripped
429 {"message":"ThrottlerException: Too Many Requests"}at request 118 withretry-after: 60and recovered after the window. The breach on this endpoint is the API’s own machine-readable JSON - but aggressive polling elsewhere has produced a non-JSON interstitial from the edge layer, so treat “non-JSON body” and “JSON 429” as the same back-off signal. - The budget is cumulative across a user’s PATs (measured 2026-08-18):
alternating two tokens from the same user, each token’s
x-ratelimit-remainingdrops on the OTHER token’s calls - one shared counter per user per scope. PAT sharding does not multiply the budget.
At fleet scale, build the client as a queue with a token bucket.
5. Let users bring their own backend
Section titled “5. Let users bring their own backend”The OAuth2 flow for the Management API: register an OAuth app, the user approves, you hold a scoped access/refresh pair.2 Measured surface on a normal paid org:
- Unknown
client_idat/v1/oauth/authorize->422 {"message":"Unrecognized client_id"}. Client validation fires before session validation, so the no-session redirect is only observable with a real client_id. - The token lifecycle, measured 2026-08-18 after a real consent: refresh grants return 24-hour access tokens AND a NEW refresh token every time - rotate or lose the grant. The token is org-scoped: it sees exactly the org approved at consent (1 org, its 3 projects), not the account’s other orgs. Revocation answers 204 and the refresh grant fails immediately with 404 - zero measurable lag.
- The
jwt-bearertoken grant validates parameters before any gating (422 Required parameter: client_id).
The project’s OWN OAuth 2.1 server (Supabase-as-IdP for third-party apps) is
a separate surface. It is fully headless-automatable and its tokens carry
client_id, which RLS can read. Measured 2026-08-18:
auth.jwt() ->> 'client_id' in a policy showed a row to the matching
client’s token and hid it from a second client of the same user, so a
shared project can hold per-client permissions as well as per-user ones.
6. Meter what tenants consume
Section titled “6. Meter what tenants consume”Metering has three layers, all measured. Per-project cost attribution is the full build, and also covers the non-tenant case of one consolidated org wanting per-project chargeback.
- The invoice already itemises usage per project ref (compute, disk, egress, storage, PITR, custom domains) with quantity and rate. The plan fee, monthly active users (MAU), function invocations, realtime, and discounts are org-level aggregates, so splitting them per tenant takes an allocation policy.
- Between invoices, estimate from the compute SKU in
billing/addons,pg_database_size()and the Storage listing (exact to the byte),usage.api-counts(exact, ~1 min lag). No per-service egress bytes exist in the public per-project API. - Your own gateway meters in real time: exact for anything transiting it, and scoped keys mean no god-mode PAT in scrapers (measured: 200 with 278 metric families for the allowed key, 403 deny-by-default elsewhere, ~1 s audit flush).
7. Move tenants when they outgrow the placement
Section titled “7. Move tenants when they outgrow the placement”Cross-project trust (third-party JWKS) makes a tenant’s tokens verifiable on
their new project, with two measured edges: refresh only works at the
issuing project (400 refresh_token_not_found verbatim at the trusting
one), and key rotation needs a maintenance window (the trusting project’s
cache does not re-resolve on any observed timeline). The build
recipes:
shared tier,
promotion,
consolidation; the decision
framework is the
placement doc, and what each
operation costs in client-visible downtime is in
platform operation costs.
8. Gotchas and lessons learned
Section titled “8. Gotchas and lessons learned”- Deleting a project has a tail: a delete returns ~2 s but the project lingers in lists briefly; batch deletes need canary batches.
- A parked project loses its public DNS record. The
HTTP 540 Project pausedthe data API returns is a transient, measured seconds after the pause completed while DNS was still live or cached; at 13 and 50 minutes parked the hostname answers NXDOMAIN on independent public resolvers, withdb.<ref>gone too (September 2026). A client has to handle the resolution failure; once teardown finishes there is no 540 left to map to a 503. - Restores are slow: a legacy project exceeded a 20-min wake bound before coming healthy; free-org restores measured 162-204 s. Wake ahead of the user, not on their click.
Verified / tested
Section titled “Verified / tested”| Claim | How it was checked |
|---|---|
| Create -> healthy 131-159 s paid / 12-13 s free | Measured 2026-08-17/18, n=5+ on paid |
Smart region accepted on a paid org, picked ap-northeast-2 | Measured 2026-08-17 |
| Nano rejected three ways on paid, accepted (201) on free | Measured 2026-08-17/18 |
| Legacy project: paused survives upgrade, cannot re-pause | Measured 2026-08-18 - 400 Project is not free-tier |
Rate-limit headers + JSON 429 with retry-after: 60 | Measured 2026-08-17 - burst to the boundary |
| OAuth authorize: 422 client validation first; claim flow 404 | Measured 2026-08-17 |
Project IdP: client_id claim usable in RLS | Measured 2026-08-18 - headless Proof Key for Code Exchange (PKCE) flow, two-client isolation |
| Metering: exact ground truth, exact 13/13 analytics at 61 s lag, no per-service egress in the API | Measured 2026-08-17/18 |
| Scoped-key gateway: 200/278 families, 403 deny-by-default, ~1 s audit | Measured 2026-08-18 against a live credential-proxy deployment |
| Refresh only at the issuing project | Measured 2026-08-18 - 400 refresh_token_not_found verbatim |
What to do about it
Section titled “What to do about it”A practice with no module id is a design choice rather than a
result: canary batch sizes and token-bucket parameters beyond 120 per minute
per user are unmeasured. No Management API operation restarts a parked
project’s compute, so there is no wake signal to send. 52 parameter-free project GETs and 73 write operations were run
against a parked project, including POST /database/query and POST /restart,
and every one left it INACTIVE. The endpoints that need the instance fail
instead - 544 carries a connection timeout and costs the caller the full
wait, and /database/migrations took 20 seconds to return it. POST /pause
and POST /restore are excluded, since they change state by definition. The
reason is structural: with no DNS record there is no hostname for traffic to arrive at, and an automatically paused project lands
in the same state as a manually paused one (both NXDOMAIN, including a control
project nothing had touched).
That result covers PAUSED projects. Scale-to-zero is a separate mechanism -
the control plane models it as its own state, project_hibernating, distinct
from INACTIVE - and none of the above was measured against it. The null
result rests on a state with no DNS record, where traffic has nowhere to
arrive, so do not assume it carries to a state that keeps its hostname. Treat
the two as separate questions.
| Practice | Evidence | Module |
|---|---|---|
| Run the provisioning service under a dedicated automation user. | The rate budget is per user across all its PATs (each token’s x-ratelimit-remaining drops on the other’s calls), so a separate user is a separate 120-per-minute bucket; and a PAT is unscoped (every scope endpoint 404s while /organizations and /profile return 200), so the user’s membership set is the blast radius. | rate-limits L01b (2026-08-18); platform-facts F03 |
Poll GET /v1/projects/{ref}/health?services=auth&services=rest&services=db before the first write on a fresh tenant, then retry the write anyway. | 2 of 2 fresh projects passed the per-service poll and still failed the first admin/users call (2026-08-03); 5 of 5 refused the first write and took the second (2026-08-04). | supabase-lab AGENTS.md, provisioning note (2026-08-03); placement reference, Verified row ‘Create -> healthy, and healthy is not writable’ (2026-08-04, n=5, bash run not in the lab repo) |
Enumerate creatable regions per org with GET /v1/projects/available-regions?organization_slug=<slug> before offering a tenant a region picker. | A Team org answered 200 with {recommendations, all: {smartGroup[], specific[]}}, 17 specific regions and 3 smart groups; the call without the slug answers 400, and /regions and /projects/regions 404. | platform-facts F04c (2026-08-20) |
| Persist the rotated refresh token before acknowledging the refresh. | Every refresh grant returns a new refresh token alongside the 24-hour access token, so a dropped rotation loses the grant (2026-08-18). | byo-oauth O01c, O01e (2026-08-18) |
Map a 404 on the refresh grant to “revoked, re-consent required”. | Revocation answered 204 and the next refresh 404 on the first poll, 0 s lag. | byo-oauth O01c, O01e (2026-08-18) |
| Budget a manual dashboard step per environment for OAuth app registration. | OAuth app registration has no API: the project-claim route answers 404 on Pro (byo-oauth O02a) and oauth/apps 404 on the platform org (sfp-platforms S05, 2026-08-24), and the published spec’s 169 operations hold only two org-scoped writes, organization creation and the project-claim callback. | byo-oauth RUNLOG, the manual drill; sfp-platforms S05 (2026-08-24); platform-facts F05 |
Do not offer JIT database access, POST /database/jit/invite, on a Pro org. | A platform-plan org answers 200 with an invite_id and the identical invite is rejected with 500 on Pro. | sfp-platforms S15 (2026-08-25, A/B) |
Mint tenant-facing or worker credentials with POST /v1/projects/{ref}/api-keys and secret_jwt_template, and read them with ?reveal=true. | The create response redacts api_key otherwise. The templated role and custom claim reach auth.jwt() on Pro and platform orgs. A fresh RPC answers PGRST202 until notify pgrst, 'reload schema'. | sfp-platforms S14 (2026-08-25, both org classes) |
Size a dedicated project at ci_small or above and enable PITR before promising a tenant read replicas via POST /read-replicas/setup. | The setup call answers 400 "Read replicas require a minimum size of small" on Pro and platform orgs, then waits on a completed physical backup; the chain closed on Pro after pitr_7 (setup 204). PITR enable on the platform org is its own refusal, 400 "Organization is not entitled to the selected PITR duration". | sfp-platforms S07, S07e (2026-08-25) |
| Grow disk to 6 GB or more in one step on a fresh gp3 volume. | The fresh volume is 2 GB / 3000 IOPS / 125 MiB/s and the gp3 IOPS floor makes 2 -> 4 GB impossible. The grow is async (201 with an empty body); 8 GB confirmed landed on a later poll. The floor was sent on a Pro 2 GB volume on 2026-10-01: 4 and 5 GB answered 400 Invalid IOPS value for gp3 volume 3000, 6 GB 201. | sfp-platforms S08 (2026-08-24/25); compute-disk D11 (RUNLOG, 2026-10-01) |
Do not promise tenants a backup time of day from GET /database/backups/schedule on Pro or platform orgs. | Pro and platform orgs both answer a structured 402 entitlement_required (feature: backup.schedule); the spec text says Enterprise. | sfp-platforms S11 (2026-08-24/25) |
| Before restoring a paused free-era project inside a paid org, decide whether it should instead be deleted or moved to a Free org. | Once woken it joins the always-on floor (400 Project is not free-tier on the next pause), and the wake exceeded a 20-minute bound. | instance-sizing I03 (2026-08-18) |
After DELETE /v1/projects/{ref}, poll GET /v1/projects until the ref is absent before reusing the name or counting against a quota. | The delete returns in about 2 s and the project lingers in lists. | section 8 (lingers in lists); placement reference, Verified row ‘Create -> healthy, and healthy is not writable’ (2026-08-04, n=5, delete about 2 s, bash run not in the lab repo) |
Verification
Section titled “Verification”The probes re-run on throwaway projects you create and delete; the calls are the ones the sections make. Sections 3 and 8 additionally need a free-tier project for the free-org rows.
- Sizing (3) - create a project and re-take the three Nano rejections
(create, addon PATCH,
available_addons); create withregion_selection: {type: "smartGroup", code: "apac"}and read which region lands. - Rate budget (4) - burst a cheap scoped read until the 429, reading the
budget from the
x-ratelimit-*headers; alternate two of your PATs to re-check the shared counter. - OAuth (5) - authorize with an unknown
client_id, then run the token lifecycle after a real consent: refresh, revoke, re-refresh. - Metering (6) - insert a known payload and read
pg_database_size(); send a counted batch of REST GETs and watchusage.api-countscatch up; scrape per-project metrics through your gateway with a scoped key. - Deletions and pauses (8) - delete the project and re-poll the list for
the tail; pause a free-tier project and read the data API for the
HTTP 540 Project paused.
The measurement harness is public at
supabase-lab - one
RUNLOG.md per experiment under experiments/ - but its destructive probes
run against the author’s orgs and secrets, so a reader re-runs the API calls
above and reads the run logs in the repo.
File reference
Section titled “File reference”| Section | Experiment in supabase-lab, RUNLOG pinned to commit d386347 |
|---|---|
| 1, 5 - the claim flow and the OAuth surface | byo-oauth |
| 1, 2, 3 - the platform-plan A/B: JIT access, templated keys, the replica and disk floors, the backup schedule | sfp-platforms |
| 2, 4 - the unscoped PAT, the region catalogue, the read-only membership surface | platform-facts |
| 3, 8 - sizing floors, the legacy door, pause and restore | instance-sizing |
| 4 - the rate budget | rate-limits |
| 6 - metering and the scoped-key gateway | usage-metering |
| 7 - cross-project trust | cross-project-auth |
Modules
Section titled “Modules”| Module | Experiment | Test | Artifact |
|---|---|---|---|
| F03 | platform-facts | f03-pat-scope.ts | none published |
| F04c | platform-facts | f04-regions.ts | none published |
| F05 | platform-facts | f05-control-plane-writes.ts | none published |
| I03 | instance-sizing | i03-legacy-pause-lifecycle.ts | none published |
| L01b | rate-limits | r01-rate-limit-surface.ts | none published |
| O01c | byo-oauth | o01-oauth-lifecycle.ts | none published |
| O01e | byo-oauth | o01-oauth-lifecycle.ts | none published |
| O02a | byo-oauth | o02-gated-oauth-surface.ts | none published |
| S05 | sfp-platforms | s05-project-claim.ts | none published |
| S07 | sfp-platforms | s07-read-replicas.ts | none published |
| S07e | sfp-platforms | s07-read-replicas.ts | none published |
| S08 | sfp-platforms | s08-disk-modification.ts | none published |
| S11 | sfp-platforms | s11-backup-schedule.ts | none published |
| S14 | sfp-platforms | s14-secret-jwt-template.ts | none published |
| S15 | sfp-platforms | s15-jit-database-access.ts | none published |
Related
Section titled “Related”| If you are asking | Go to |
|---|---|
| Where should a tenant live, and what does it cost? | Tenant placement |
| What does each operation cost in downtime? | Platform operation costs |
| Build the shared tier | Shared tenancy guide |
| Move a tenant to its own project | Tenant promotion |
| Merge many projects into one | Tenant consolidation |
References
Section titled “References”-
Supabase for Platforms - the white-label programme; scale-to-zero is the platform plan’s Nano create default (measured 2026-08-24), not a gated catalogue variant. ↩
-
Build a Supabase OAuth integration - the Management API OAuth2 flow (authorize, token exchange, refresh, revoke). ↩