Skip to content

Scheduling one GPU between LLM, image and training work

A local model stack has one scarce resource and several consumers. This page is the design of the layer that arbitrates it: a proxy that owns container lifecycle for an LLM server, an image-generation server and a training service, all sharing one GPU, and the operating rules its incidents produced.1 It is the scheduling counterpart to the bench numbers, which measure the models rather than the machinery around them.

Provenance: the behaviour below is read from the stack’s own source and its incident notes, not measured by a bench run. The hardware is the same as the bench page states - an RTX 5090 (32 GB, sm_120) in a WSL2 dev box, CUDA containers throughout. The stack’s repository is private, so file paths are named but not linked. Where a claim is an operating rule rather than a code path, it says so.

  • One proxy owns the GPU. It runs exactly one of llama-server, ninfer-server (a second LLM engine), ComfyUI or lora-train, and swaps between them on demand; the GPU services live outside Compose and are spawned through the Docker API.
  • A swap drains first: in-flight requests on the outgoing model finish before its container is killed, up to a 60 s grace deadline. Past the deadline the swap proceeds anyway and long generations die mid-stream.
  • A swap wipes the engine’s context cache. A ~200k-token agent session re-prefills from scratch on the next turn, which clients report as 40-90 s of time-to-first-token and a 0% cache-hit rate.
  • A preset lock pins a model against evicting swaps. It has a 900 s TTL, is refreshed by traffic or explicitly, and a contended lock returns 409 rather than hijacking the running model. --wait joins a FIFO queue instead.
  • The two failure modes that cost the most time were both contention, not engine bugs: a lock that silently lapsed mid-leg, and two clients on different presets swapping the GPU back and forth. A third - the retracted 2.2x decode regression - was contention misread as a regression.
agent harness(pi, Claude Code, curl)model-proxy-go :11434scheduler event loopstate in active.toml/v1/*llama-server :8080GGUF presetsmode llmninfer-server :8080.ninfer artifactsmode llmcomfyui :8188mode comfyuilora-train :8787mode trainone GPUone workload at a time

Text fallback: an agent harness sends OpenAI-compatible and Anthropic-compatible requests to the proxy on port 11434. For mode = llm the proxy spawns and forwards to whichever engine the resolved preset names - llama-server for GGUF presets, ninfer-server for .ninfer artifacts - both on port 8080 in their own container. For mode = comfyui it forwards to ComfyUI on 8188, and for mode = train to the LoRA trainer on 8787. All four GPU containers are candidates for the same card and only one exists at a time; the proxy persists the current mode and model to active.toml so a restart knows what it left running.

PathTargetNotes
/v1/chat/completions, /v1/completionsactive LLM enginebody model selects a preset; auto selects a route alias
/v1/messagesactive LLM engineAnthropic shim, so Claude Code can point ANTHROPIC_BASE_URL here
/v1/modelspreset storelists preset names as id, engine-side ids under meta.model_id
/v1/presets (POST, DELETE)ephemeral preset registryin-memory presets for sweep runs; TOML wins on a model-id collision
/comfyui/*ComfyUIforwards to 8188
/train/*LoRA trainerforwards to 8787
/mode (GET, POST)schedulermode switch, lock, unlock, renew
/lock, /unlock (POST)scheduleraliases for the lock calls
/status, /healthschedulermode, model, lock owners, queue, in-flight counts

Modes are llm, comfyui, train and idle. A write request for a path in an inactive mode triggers the swap; a read-only request does not, and returns 503 service_inactive instead. mode llm is a single mode covering two engines, so a request cannot be routed by mode alone - the preset decides the container.

A swap is a state machine on one goroutine, with the container work handed to a goroutine that reports back on the same channel. The scheduler is adopted from llama-swap’s internal router1, with drain, capability routing and the lock API added on top.

request for preset B(resident: A)drainin-flight A finishnew A requests deferredserve in place(alias or capabilitystill satisfiable by A)optionalgrace deadline60 s (LLMC_DRAIN_GRACE_S)A still busystop A's containerstart B's containerA drainedswapping anywayhealth wait900 s (LLMC_HEALTH_TIMEOUT)grant waitersforward requesthealthytimeout: 503 to waiters

Text fallback: a request for a preset other than the resident one starts a pending swap and records the resident as the drain key. While draining, new requests for the same target join the pending swap as waiters; a request an alias or capability can still satisfy on the resident model is granted in place with no swap; anything else is deferred. The swap starts when the resident’s in-flight count reaches zero, or when the grace deadline fires and logs that it is swapping with requests still in flight. Starting a swap stops the outgoing container before creating the incoming one, then waits for the health endpoint. On success the waiters are granted with the in-flight key under the new model. On failure the state is re-derived from Docker rather than assumed - a health-check timeout leaves a container up but unhealthy, so both “idle” and “still running” would be wrong in one of the two cases - and the waiters get a 503.

ControlDefaultEffect
LLMC_DRAIN_GRACE_S60 show long in-flight requests may hold a swap off before it proceeds anyway
LLMC_HEALTH_TIMEOUT900 sreadiness wait after spawning any GPU service
LLMC_VRAM_LIMIT_GB32card budget used for the pre-swap gate
LLMC_VRAM_RESERVE_GB6headroom subtracted from it, so a preset must fit 26 GB
LLMC_LOCK_TTL_S900 slock and queue-entry lifetime

The VRAM gate is checked twice: on a preset before the scheduler sees the request (422 model_unavailable), and on the chain head of a route alias at swap-decision time. The reserve is the reason a preset that fits the card can still be refused the swap.

Swap latency is dominated by the model read. Cold storage is 30-60 s per swap; a page-cache-warm swap is 5-10 s. An agent session that hits a swap pays that once, plus the re-prefill described below.

A preset is a TOML file in the presets directory, mounted read-only into the proxy. It declares the model (repo plus file, or a local path), an optional multimodal projector and chat template, a [runtime] block, an optional [bench] tokenizer, and a capabilities list.

  • The schema is strict. One unknown key fails the entire reload, not just that file, so a preset added with a newer key silently freezes /v1/models at its last good state. A key that exists in the Python preset loader but not in the Go one is the common way to hit this.
  • Reload is live. /v1/models rescans the directory, so adding or editing a TOML needs no rebuild and no restart. A schema-shape change to the proxy itself does need one.
  • Presets deduplicate by model id, derived from the artifact filename stem. Two presets naming one file is a load error rather than a duplicate entry, so an A/B arm on the same weights needs a second filename - a hardlink costs no disk.
  • Ephemeral presets can be registered at runtime through POST /v1/presets. They survive reloads, are never written to disk, and are lost on restart; an on-disk preset wins a model-id collision.
  • capabilities = ["vision", ...] lets a request name a capability rather than a model, either as an X-LLM-Capability header or as a cap:<name> model. If the resident model has it, the request is served in place and the swap is skipped entirely.
  • The listing carries the metadata clients need to size themselves: effective context, whether reasoning is on, the output cap, the VRAM estimate, and which preset is currently loaded.

A second TOML file, routes.toml, maps model aliases to ordered preset chains. auto resolves the default chain and auto:<name> a named one; a request naming an alias walks the chain, serves in place when the resident preset is already in it, and otherwise swaps to the chain head. An empty chain means “whatever is resident, never swap” - the alias for a client that should never be the reason a model changes. A chain entry that does not resolve fails the request with 404 unknown_route rather than silently degrading.

A lock pins a preset so evicting swaps are refused. It exists because the alternative - trusting every client to know when the GPU is occupied - did not hold: an unattended loop and an interactive session on different presets will swap the card back and forth indefinitely.

CallBody or flagResult
lock, uncontended{"lock": "<preset>", "owner": "<id>"}200, with owners and expiry
lock, contendedsame409 refusing to hijack, with the current owners and queue
lock, contended, waiting{"lock": ..., "owner": ..., "wait": true}202 {"queued": true, "position": N}
lock the resident model{"lock": true}pins whatever is loaded
renew{"renew": true, "owner": "<id>"}extends the TTL; works for a queued waiter too
unlock one owner{"lock": false, "owner": "<id>"}releases that owner, drops its queue entry
unlock all{"lock": false}clears every owner

Rules that matter in practice:

  • A contended lock never hijacks the running model. Before the FIFO queue existed it did, and that killed a running loop mid-iteration. The queue exists to make “wait your turn” the cheap answer.
  • Only the queue head may take a free lock, and only for the model it queued for. Joining an already-locked model’s owner set is never gated, because adding an owner to a lock that is already held carries no eviction risk.
  • The TTL lapses a silent leg. Traffic under the lock refreshes it; a leg that makes no request for longer than the TTL - long local thinking, or waiting in the queue - loses the pin without an error anywhere. Heartbeat from anything that can go quiet for 15 minutes.
  • Requests under a lock are policed. A request for a different preset while a lock is held is rejected 422 model_unavailable rather than swapping, and an unknown model name is rejected rather than passed through to silently run on the locked model.
  • The queue holds names, not swaps. A queued owner waits for the preset; the swap itself happens lazily, on the winner’s first request. Entries carry a timestamp and are pruned once their owner has been silent for a TTL, so a dropped queue heals itself and a waiter re-enqueues on its next poll. The state file carries the queue as well as the lock, so the lock is restart-safe; a waiter’s own polling loop is what makes its place survive.
  • Bench and audit modules take the lock with their own owner name and fail fast rather than queueing, so a measurement either runs on an uncontended card or refuses to run.

This is the part clients feel, and the reason a lock is worth holding. Stopping the engine discards its context cache: the reused-prefix state that makes a long agent session cheap to continue.

SymptomCauseObserved figure
Time to first token balloonsfull re-prefill of the session after the cache is dropped40-90 s on a ~200k-token session
Cache-hit rate reads 0%client-side prefix cache is invalidated by the respawnpi reported cache 0.0% every turn
Container appears to crash-loopeach new request swaps to a different preset, so the container is torn down and rebuilt repeatedlyonce per ~90 s over 13 minutes of two clients fighting
Long generation dies mid-streamthe swap proceeded at the grace deadline with the request still in flightupstream_died_midstream in the proxy log
Decode looks like a regressionthe outgoing and incoming engines coexist briefly during drain, and both hold the GPU55-60 tok/s against a 139 tok/s baseline

The last row is the one that misleads. Draining means the two engines overlap, and the overlap starves decode until the outgoing container is gone. On a locked card the same probes measured 127-131 tok/s on the older engine revision and 129-131 on the newer one; the “regression” was another client POSTing through the proxy during the measurement window.

SituationDo thisWhy
Unattended loop or long buildllmc lock <preset> --owner <session-id> before startingthe loop is the workload that must not be interrupted, and it cannot detect a swap
Two loopssame preset, one owner each; different presets queue with --waitdifferent presets on one card swap against each other forever
Interactive work alongside a looplock the same preset the loop uses, or accept the queuea second preset is a second swap
Measuring throughputlock the preset, and name the locked model in the requestan unlocked probe measures whoever else is using the card
One image or a short training runjust call ComfyUI or /train/*the proxy swaps when it can, and refuses when a lock is held
A client that must never trigger a swapa route alias with an empty chainresolves to whatever is resident, or fails cleanly
does the workneed the GPUexclusivelyfor minutes?lock the preset+ heartbeat the TTLyes, unattendedlock refused?noretry with waitjoin the FIFO queue409another preset isbeing measuredno lockshare the card,expect swapscoordinate, orshare the preset

Text fallback: if the work holds the GPU exclusively for minutes - an unattended loop, a benchmark, a training run - take a preset lock and heartbeat its TTL. If the lock is refused with a 409, either retry with --wait to join the FIFO queue, or coordinate with whoever holds it; for two loops, the answer is to share one preset rather than queue. If the work does not need exclusivity - a short generation, a chat turn - do not lock: send it and accept that the proxy may swap the model under you.

Every row here is a dated entry in the stack’s own incident log, not a hypothetical. They are the reason the current design is shaped this way.

DateWhat happenedWhat changed
2026-08-17A lock for a different preset while one was held took the GPU anyway, killing a loop mid-iterationcontended locks return 409, or queue with wait; never hijack
2026-08-19A 30-minute leg that made no requests found its lock silently lapsedTTL documented as a heartbeat obligation; lock --renew added
2026-08-21A container killed out of band left the proxy 502-loopinga connection-level upstream death flips mode to idle, keeping the model name; the next acquire respawns (verified: kill, 502, idle, respawn served in about 7 s)
2026-09-07A swap failed after the outgoing container was already stopped; the proxy reported a model loaded with no container behind it and refused to respawnswap failure re-derives mode from Docker instead of trusting pre-swap state
2026-09-08A bare unlock while a bench was running handed the GPU away and produced a fake resultlock discipline became a measurement precondition, with the lock held for the whole run
2026-09-14A client abort mid-materialization wedged the engine’s single slot for 89 minutes: 17.6 tok/s prefill, host at 0%, then self-recovery - the wedged request’s own client never disconnected2an engine-side transport watchdog patch, a proxy watchdog enabled in production, and per-request diagnostics logging
2026-09-18A ~2.2x decode slowdown was blamed on an engine bump; it was another client POSTing during the measurement windows, with drain overlap starving decode. Retractedspeed probes must hold the lock; the retraction is recorded in the bench page
2026-09-19A loop and an interactive session on different presets fought for 13 minutes; the container was rebuilt every ~90 s, TTFT reached 40-90 s, generations died at the grace deadlineone winner, or share a preset; the diagnosis is alternating model names in the proxy log
2026-09-28Every swap failed with a dial error on the container runtime socketthe socket is mounted as its parent directory rather than as a file, so a recreated socket is not pinned to a dead inode

Two of these are worth separating in kind. The wedge is an engine bug with a client-side trigger: the damage landed on the request after the one that was aborted, which is why the defence is a probe on the next request rather than a fix to the aborted one. The swap war is a scheduling failure with two ordinary clients - each client’s behaviour was correct in isolation, and reading the proxy log for the pair reproduced it in minutes.

  • Swap latency is a storage number first. Cold 30-60 s and warm 5-10 s per swap says the swap cost is dominated by reading tens of gigabytes of weights, not by container start, and that a second swap five minutes later is cheap. A loop that alternates presets therefore pays the expensive case repeatedly only when something else evicts the cache between turns.
  • TTFT after a swap is a session-size number. 40-90 s on a ~200k-token session is the re-prefill of the whole conversation, so the same swap on a short chat is invisible. Any report of “the local model got slow” is worth reading as a session-size question before an engine question.
  • Throughput under an unlocked card is not a measurement of the engine. 55-60 tok/s versus 127-131 tok/s is the spread a second client can induce, so an unlocked number carries an error bar larger than any engine revision this stack has shipped.
  • The drain overlap is bounded but not free. Both engines hold the card until the outgoing container stops, which is precisely the window in which a concurrent benchmark is invalid.
  • Idle auto-unload is specified but not implemented. A spec describes releasing the GPU after a configured idle period by stopping the engine and keeping the model name so the next request respawns it; the shipped proxy has no handling for the environment variable it names. Until that lands, a model stays resident after its last request, holding its VRAM and power.
  • Residency detection covers one engine. The probe that reports which model is loaded understands the llama.cpp engine and not the second one, so a proxy restart with that engine resident forces one needless swap.
  • The proxy schedules its own three workloads only. A sister transcription stack keeps a model resident outside the proxy on the same card (about 5.6 GiB), which the VRAM budget does not account for and which has twice invalidated measurements taken through the proxy.

The paths below are in the stack’s private repository. Requests are against the proxy’s own HTTP surface, so everything in the locks section can be exercised with curl against 127.0.0.1:11434 without touching the Docker API.

SurfaceWhere in the source
Scheduler event loop, drain, lock and queueproxy-go/internal/proxy/scheduler.go
Persisted state, modes, queue entriesproxy-go/internal/proxy/state.go
Preset schema, live reload, capability list, model-id dedupproxy-go/internal/proxy/presets.go
Routes, ephemeral presets, mode handler, VRAM gateproxy-go/internal/proxy/server.go
Container lifecycle and health waitsproxy-go/internal/proxy/orchestrator.go
Defaults for every setting in the table aboveproxy-go/cmd/proxy/main.go
Alias chains (auto, auto:<name>)routes.toml
Incident log and operating rulesthe repository’s AGENTS.md
ClaimStatusHow it was checked
One workload at a time; four GPU containers, one cardassertedarchitecture section of the repository’s notes and Compose comments
Drain-before-swap, grace deadline, waiter and in-place rulesdesign-readscheduler.go; the grace path logs when it swaps with requests in flight
Swap failure re-derives mode from Dockerdesign-readcomment and code path in handleSwapDone
Lock semantics: 409, 202, TTL, FIFO gate, 422 under a lockdesign-readhandleLock, handleUnlock, expireLockIfNeeded, pruneQueue
Lock is persisted, queue entries are pruned on the TTLdesign-readstate.go plus the pruning path
Preset strictness, live reload, model-id dedup, ephemeral presetsdesign-readpresets.go, handleModels, handlePresetRegister
Route aliases and the VRAM gatedesign-readacquireRoute, CheckVRAMBudget, the gate in server.go
Cold 30-60 s and warm 5-10 s swap latencymeasuredrecorded in the repository’s README
Lock-hijack incidentmeasureddated incident entry, 2026-08-17
TTL lapse on a silent legmeasureddated postmortem, 2026-08-19
Idle-unload spec with no implementationdesign-read (absence)spec present; no handling of its variable anywhere in the proxy source
Drain overlap starving decode; retracted 2.2xmeasured then retracteddated entry 2026-09-18 and the retraction in the bench page
Swap war, container rebuild cadence, TTFT, mid-stream deathsmeasureddated entry 2026-09-19
Engine wedge, 89 minutes, one request latermeasureddated entry 2026-09-14 and the engine’s upstream issue

No figure on this page was measured for it. The throughput and latency numbers that justify the lock rules are in Qwen3.8-27B for agentic work and the bench numbers; this page explains what the proxy was doing while those were taken.

  1. mostlygeek, “llama-swap,” GitHub. https://github.com/mostlygeek/llama-swap ↩ ↩2

  2. Neroued, “NInfer issue #184,” GitHub. https://github.com/Neroued/ninfer/issues/184 ↩