Skip to content

Multigres OSS failover: what a client sees

What a Postgres client sees when the primary of an open-source Multigres cluster dies, freezes or has a lagging standby: how long writes stall, which errors come back, and whether any acknowledged write is lost. For anyone weighing Multigres against the warm standby tier of a managed project.

The numbers are from one local rig on 2026-10-10: a Docker Desktop VM (Linux kernel 7.0.14-linuxkit, Docker server 29.8.2, 10 CPUs, about 8 GB) on an arm64 Mac. The tests, the load generators and the fault injector (docker exec kill) ran on the macOS host and reached the gateway through a published port. No cloud resources and no Supabase project were involved. The Multigres blog and the v0.1.0 release page are documented sources and were not tested.12

Scope first. This is the upstream repository’s all-in-one image (Dockerfile.cluster), built from a pinned main commit of 2026-10-09 because the v0.1.0 tag has no such Dockerfile (the GitHub contents API answered 404 at that ref). It is not the Kubernetes operator path1 and not the invite-only private alpha, and neither was measured.

TL;DR:

  • 17 failovers, 2,216,532 acknowledged writes checked against a client-side commit log, 0 acknowledged-but-lost. That includes a standby left 22.3 to 23.3 MB behind by SIGSTOP, followed by killing the primary and the other standby: the lagging standby was never promoted (MG08).
  • The longest ack-free gap depends on the fault. Postgres SIGKILL: 1.1 to 4.2 s with 8 writers, 1.9 to 6.7 s with one 10 writes/s client. Whole-cell kill: 14.7 to 15.1 s. Hung primary: 20.0 and 25.1 s in the main run, 20.0 to 56.2 s over six runs in three artifacts.
  • Clients are not shielded. In-flight statements fail (one per worker), new connections get no writable primary is currently available or database is temporarily unavailable; please retry, and 2 to 5 connects per postgres-kill run were refused with password authentication failed using the correct password.
  • Nothing restarted a killed cell in this container: the cluster stayed at 2 nodes for 45 s.
  • The gateway passed 9 of 9 session-feature probes on two cells, with one uncontended client. Pool contention, network partitions and other durability policies were not measured.

One container (child processes)cell zone1Client8 writersmultigatewaypublished portetcdmultiadminmultipoolerpostgres + pgctldcells zone2, zone3same five processesANY 1 of 3 syncmultiorchpromote
  1. The client connects to the zone1 multigateway through one published port.
  2. The gateway routes to the multipooler, which holds the backends on that cell’s postgres.
  3. Each cell runs its own postgres, pgctld, multipooler, multiorch and multigateway; etcd and multiadmin are shared.
  4. Commits are acknowledged under synchronous_commit=on with synchronous_standby_names of the form ANY 1 (<3 ids>), so one standby must flush before the client sees an acknowledgement.

What was run, as measured in MG01 and the run configuration:

ItemValue
Imageupstream main at commit e3cbdbbb0b94453ca2ec43c9c40c05c98ed7bb73 (2026-10-09), built from Dockerfile.cluster
Server bannerMultigres 0.1.0-SNAPSHOT (unknown revision) built with go1.26.9 (image built without the commit stamp)
Cells3 (zone1 to zone3); a single cell never becomes ready
Durabilitysynchronous_commit=on, ANY 1 of 3 poolers, 2 walsenders visible from the primary, max_connections=110
Multiorch timersrecovery-cycle-interval and pooler-health-check-interval 500 ms
Bootstrapallow-unsafe-initial-cohort: true written into every multiorch block by the entrypoint
Load8 workers, one connection each, autocommit INSERT (w, n) through the zone1 gateway, closed loop, 5,398 to 6,046 acknowledged writes per second before the fault

The 500 ms timers and the unsafe-initial-cohort flag are the local provisioner’s values; the operator’s were not read or run, and whether they change the results against a default deployment was not measured.


Each run is one failover; n = 3 per fault unless stated. An INSERT that returns is acknowledged. One that errors is in doubt and is checked against the final table. Acknowledged-but-lost means an acknowledged (w, n) missing from the table read back after recovery. The stall is the longest interval with no acknowledgement from any worker that ends after the fault.

Module and faultnAcknowledged writes checkedAcknowledged-but-lostAck stall per run (ms)
MG02 SIGKILL the primary’s postgres (pgctld and multipooler left up)3442,94402,238 / 1,100 / 4,233
MG03 SIGKILL postgres, multipooler and pgctld of the primary’s cell3263,121014,702 / 15,136 / 14,735
MG04 SIGSTOP the primary’s cell for 45 s, then SIGCONT2395,924020,016 / 25,122
MG08 one standby SIGSTOPped 20 s (22.3 to 23.3 MB behind), then primary and current standby SIGKILLed3759,31906,852 / 6,431 / 8,161
MG09 as MG02 with one writer at 10 writes/s372006,657 / 1,900 / 6,589
MG05 pgbench relaunch loop and two single runs3354,504 logged0n/a / n/a / 2,283

Over the 17 runs, 2,216,532 acknowledged writes were checked and none was missing from the final table. Every run ended with the same row count on every surviving node and on a read through the gateway. The fault clock starts just before the docker exec kill, whose latency is inside every window (largest 154 ms in the printed evidence).

After the postgres SIGKILL, the promote step took 327 to 374 ms in MG02, 343 to 364 ms in MG09 and 345 to 2,181 ms in MG08. The rest of the stall is the time before multiorch starts its leader appointment: the first starting leader appointment line came 570 to 3,786 ms after the kill in MG02, 1,411 to 6,082 ms in MG09 and 5,765 to 6,288 ms in MG08.

The spread is wide, from about 1 s to about 6.7 s, with most runs near 1 to 2 s or near 6 to 7 s, and the cause is unexplained. The one-writer MG09 runs gave (6.7 / 1.9 / 6.6 s) and the 8-writer MG02 runs (2.2 / 1.1 / 4.2 s), but two earlier full runs of the same fault code gave (6.3 / 1.8 / 1.0 s) and (1.4 / 1.7 / 1.3 s). At n = 3 per cell the claim that traffic speeds up detection is not supported and is not made.

Killing postgres, multipooler and pgctld together gave multiorch the reason LeaderUnreachableByCohort. Its first starting leader appointment line came (14.2 to 14.6 s) after the kill in all three runs, and the promote step took 360 to 364 ms. Nothing restarted the killed cell: after a 45 s wait the cluster still had 2 nodes in all three runs, so each MG03 run starts from a freshly created container.

A frozen cell keeps its sockets open. Multiorch’s reason was LeaderQuorumWritesStalled. In the two main-run measurements:

EventRun aRun b
First executing appoint leader action check+18.6 s+18.7 s
starting leader appointment+19.4 s+20.1 s
New primary promoted at the pooler+19.9 s+20.6 s
Multiorch logs the promotion as a success+39.7 s+40.3 s
Client ack stall20,016 ms25,122 ms

Multiorch’s success line lags the pooler-level promote by about 20 s because the recruit RPC to the frozen cell had to time out first (recruit_ms about 20,250). In run a the 20,016 ms stall coincides with the client’s own 20 s statement timeout: all 8 workers ended with Query read timeout at +20 s, so the client could not see recovery earlier. In run b the reconnects after that timeout hit the client’s 5 s connect timeout once (8 x timeout expired at +25.15 s), which accounts for the 25,122 ms on top of a cluster-side recovery at about +20.6 s. That 5 s is client behaviour, not a cluster window.

The stall was not always about 20 s. Across the six MG04 runs in three artifacts it was 20.0 / 25.1 s (main run), 20.0 / 20.0 s and 25.1 / 56.2 s. In the 56.2 s run the first promotion attempt failed at +39.7 s, a second started at +48.3 s and succeeded at +57.8 s, and the client saw 8 x Query read timeout plus 8 x portal execution failed ... EOF. The cause of that failed first attempt was not investigated. The earlier artifacts are unpublished and cited only for this range.

One standby was stopped for 20 s. In all three runs the primary’s pg_stat_replication showed it 22,265,848 to 23,256,456 bytes behind (flush lag) just before the double kill. The lagging standby was not promoted in any run: the winner was the restarted original primary once and the up-to-date killed standby twice. A first version with a 5 s stop (5.4 to 5.8 MB behind) promoted the lagging standby in 1 of 3 runs with 0 lost, and proves nothing, because that much WAL may sit in the stopped standby’s socket receive buffer and arrive after SIGCONT although the primary is dead. The buffer size was not measured. The 20 s stop makes the lag about four times larger, and that is the version reported.


Failover does not hide the fault from the client. Measured during the runs above:

ObservationDetail
In-flight statements failOne failed INSERT per worker in every MG03 and MG04 run (8 of 8), 8 to 11 in MG02. Of the failed INSERTs, 0 to 5 per run were in the table anyway (committed, unacknowledged); the rest were absent
Error text, verbatim fragmentsfailed to read message: EOF, terminating connection due to unexpected postmaster exit, no writable primary is currently available, and the client-side Query read timeout (MG04)
New connections during a postgres kill (MG02, MG08, MG09)database is temporarily unavailable; please retry or no writable primary is currently available. Failed connects per run: 45 to 237 (MG02), 349 to 561 (MG08), 13 to 59 (MG09)
Refused logins with the correct passwordIn every one of those runs, 2 to 5 connects were refused with password authentication failed for user "postgres" although the same password connected before and after
Held connects3 connects per run in MG02 and in two MG08 runs were accepted but held before completing (longest per run 1.2 to 4.6 s)
New connections during cell loss (MG03)1,096 to 1,136 failed per run, all no writable primary is currently available, none held over 200 ms
New connections to a hung primary (MG04)Run a had 0 failed connects: the clients sat on their open sockets until the 20 s timeout

pgbench shows the reconnect burden plainly (MG05). In one long run with the simple protocol and one with -M prepared, all 8 clients aborted 45 to 51 ms after the kill, pgbench exited 2, and no client completed a transaction after the fault, so pgbench alone gives no failover window (reported n/a). A relaunch loop of 41 back-to-back pgbench -T 2 runs, 21 of which exited non-zero, had 8 of 8 client slots complete transactions after the fault, with an ack stall of 2,283 ms and 240,174 transactions logged. pgbench’s per-client log counts matched the per-client row counts: 0 acknowledged commits missing in all three runs, and rows exceeded logged transactions by 3 to 5 (committed, unacknowledged).

The Multigres launch post says “During a failover, multigateway can hold requests until a new primary is promoted, thereby minimizing errors.”1 Statements already sent to the dead primary failed here, and most new connects were refused with an error rather than held; the 3 held connects per run are the part that matches.


MG06 re-ran the 9 probes of the pooler-semantics S01 matrix against the primary’s own postgres (control, 9 of 9), the zone1 gateway and the zone2 gateway. Both gateways passed 9 of 9: pid_stable, prepared_first, prepared_reuse (bind and execute with no re-parse), advisory_lock, listen_notify, session_guc, cursor_with_hold, temp_table and explicit_txn.

One client was connected, so the pool was uncontended. The gateway did not give that client a dedicated backend: autocommit statements ran on backend pid 9171 (both gateways), the explicit transaction on a different pid (9928 and 9931), and the control on 9927. Whether session state survives when other clients compete for the pool was not measured.

MG07 connected roles mg07_a and mg07_b through the gateway; the primary’s pg_stat_activity showed a backend with usename equal to each role, on different pids. That is consistent with the blog’s “separate connection pool per user with no shared pool and no SET ROLE impersonation”1 for two roles and one connection each. Pool sizing and fair-share allocation were not tested.


ClaimHolds forDoes not follow
0 acknowledged-but-lostANY 1 of 3, synchronous_commit=on, process kill and SIGSTOP, 17 runsPartitions, other policies, a contended pool
Stall of 1.1 to 6.7 s after a postgres killThis rig, a 500 ms orchestrator cycleA default deployment; the spread is unexplained
14.7 to 15.1 s for a lost cellThree runs, each on a fresh containerWhat an operator does with a dead pod: here nothing restarted the cell
20.0 to 56.2 s for a hung primarySix runs in three artifacts; two of them in the main runA bound: the 56.2 s run had a failed first promotion of unknown cause
Latency and throughputContext only: closed-loop figures from a Docker Desktop VMAny comparison with a managed project

Not measured: the Kubernetes operator path and real node loss; the invite-only private alpha (no access); network partitions between cells (a frozen process is not a partition); loss of a gateway or of etcd; durability policies other than ANY 1 of 3; whether allow-unsafe-initial-cohort and the 500 ms timers change the results; the cause of the detection spread and of the failed first promotion.

The nearest managed comparison is the hand-rolled warm standby in Supabase disaster recovery tiers, which is logical replication with a rehearsed cutover measured in minutes. This rig fails over inside the cluster in seconds on synchronous replication, with the caveats above. Moving data between two Postgres endpoints with logical replication is the subject of pgmig.


Sizing and client rules from this page’s own numbers. A row that is a design choice says so.

PracticeEvidenceModule
Retry a failed write on a new connection, and make it idempotent.In-flight INSERTs failed, one per worker in every MG03 and MG04 run; 0 to 5 failed INSERTs per run were committed anyway.MG02, MG03, MG04
Retry a refused login with the same password before alarming.2 to 5 connects per postgres-kill run answered password authentication failed with the correct password; cause not investigated.MG02, MG08, MG09
Budget about 15 s for a lost cell and 60 s for a hung primary.14.7 to 15.1 s (3 runs) and 20.0 to 56.2 s (six runs); the budgets are design choices above those ranges.MG03, MG04
Set the client statement timeout to the stall you can absorb.The 20,016 ms stall matched the client’s own 20 s statement timeout; recovery was at about +20 s.MG04
Keep a connect timeout in the retry loop and count it in the stall.A 5 s client connect timeout added 5 s (25,122 ms against recovery at about +20.6 s).MG04
Relaunch pgbench in a loop around a failover.All 8 clients aborted 45 to 51 ms after the kill and pgbench exited 2; a relaunch loop recovered with a 2,283 ms gap.MG05
Replace a lost cell yourself; do not wait for a restart.Nothing restarted the killed cell in this container: 2 nodes after 45 s, all three runs. Unknown under an operator.MG03
Test session state under a contended pool before relying on it.9 of 9 probes passed with one client, and autocommit and explicit-transaction statements ran on different backend pids.MG06

Number in this docStatusHow it was checked
17 failovers, 2,216,532 acknowledged writes, 0 lostmeasured 2026-10-10per-run acked sums (MG05: logged_transactions) against the table read back after recovery, each run’s row count equal on every surviving node and through the gateway; local Docker Desktop, 22 results all pass or info
Postgres kill stall 1.1 to 4.2 s (MG02), 1.9 to 6.7 s (MG09)measured 2026-10-10MG02 and MG09, n = 3 each; cause of the spread not found
Cell kill stall 14.7 to 15.1 smeasured 2026-10-10MG03, n = 3, fresh container each run
Hung primary 20.0 and 25.1 s, and 20.0 to 56.2 s over six runsmeasured 2026-10-10 (main run); earlier artifacts unpublishedMG04, n = 2 in the main run; the range includes three earlier artifacts
Lagging standby 22,265,848 to 23,256,456 bytes behind, never promotedmeasured 2026-10-10MG08, pg_stat_replication flush lag, n = 3
Cluster stayed at 2 nodes for 45 smeasured 2026-10-10MG03 wait after the cell kill, 3 of 3
Gateway 9 of 9 on zone1 and zone2; per-user backendsmeasured 2026-10-10MG06 and MG07, one uncontended client, two roles
Gateway holds requests; per-user pools; operator; 2026-06-04 postdocumentedthe launch post1, read 2026-10-10 through a page-summarising fetch, not tested beyond the measurements above
No Dockerfile.cluster at v0.1.0measured 2026-10-10GitHub contents API answered 404 for it and for docker-compose.yml at that ref; release page2

The one docs-versus-runtime difference is the gateway-holds-requests sentence against statements failing on the dead primary, described above.


ModuleExperimentTestArtifact
MG01multigresmg01-cluster-facts.tsnone published
MG02multigresmg02-kill-postgres.tsnone published
MG03multigresmg03-kill-cell.tsnone published
MG04multigresmg04-freeze-primary.tsnone published
MG05multigresmg05-pgbench-failover.tsnone published
MG06multigresmg06-feature-matrix.tsnone published
MG07multigresmg07-per-user-pools.tsnone published
MG08multigresmg08-lagging-standby.tsnone published
MG09multigresmg09-idle-kill.tsnone published
S01pooler-semanticss01-feature-matrix.tsnone published for the MG06 re-run

The experiment’s RUNLOG holds the per-run figures; its raw evidence directory is gitignored.

  1. Supabase, “Multigres v0.1 Alpha: an operating system for Postgres,” Supabase Blog, 2026-06-04. https://supabase.com/blog/multigres-v0-1-alpha ↩ ↩2 ↩3 ↩4 ↩5

  2. Multigres, “v0.1.0,” GitHub Releases. https://github.com/multigres/multigres/releases/tag/v0.1.0 ↩ ↩2