slskarr: one process for Soulseek, the PIA tunnel, and the Lidarr import
slskarr is a single Go binary that runs the Soulseek client, its own PIA WireGuard tunnel, and the bridge that hands finished albums to Lidarr. It replaced three containers - slskd, the wg-pia-slskd tunnel container, and soularr - on the servarr NAS on 2026-09-30. This doc covers why the three-container design was retired, what the in-process tunnel and the root-owned download tree cost, and how the import pipeline is arranged so a failed import never loses a library file.
TL;DR:
- Transfers are native in-memory state. There is no REST API to poll and no in-memory mirror that can desync, which was the failure mode that let a completed download become invisible.
- The PIA WireGuard session, the port-forward lease, and the tunnel-liveness probe all run inside slskarr. The Soulseek session starts only from the tunnel’s up callback, so a failed tunnel means no Soulseek connection at all.
- slskarr runs as root because the in-process WireGuard needs
CAP_NET_ADMIN; run as--user 1000:100the effective capability set is empty andip link add ... type wireguardreturnsOperation not permitted. - Root writes are corrected at the point of writing (
internal/config/perm.go: directories0775, files0664, ownerSLSKARR_FILE_UID/SLSKARR_FILE_GID), with a one-time marker-gated walk over trees written before the owner was configured. - Imports go through Lidarr’s
ManualImportcommand with the artist id passed explicitly. The olderDownloadedAlbumsScancommand parses the artist from the folder name, so a folder namedc<id>imports asUnknown Artist c72. - Lidarr deletes the library file it is replacing before it moves the new one in. When the move failed on a root-owned download, three library tracks were permanently deleted. Every guard in the import path exists because of that ordering.
What replaced what
Section titled “What replaced what”| Retired service | Address | Replaced by | How |
|---|---|---|---|
soularr | .8 | slskarr | slskarr took the slot and the job: Lidarr wanted-list to Soulseek download |
slskd | .16 | slskarr | The Soulseek protocol client is now in-process, not a daemon behind REST |
wg-pia-slskd | .16 | slskarr | PIA session, WireGuard netlink setup, and port-forward lease are all in-process |
beets stayed in the stack. Its fingerprinting half is duplicated by the bundled fpcalc, but its library-wide acoustic dedup has no replacement yet, so the container is still there.
Text fallback for the diagram:
- Lidarr and slskarr sit next to each other on the
servarrbridge network, in the reserved address block the arr stack patterns guide documents. - Inside slskarr: the PIA session (WireGuard plus port-forward lease), the Soulseek client (transfers and share index), the import bridge, and the reconciler.
- The PIA session tells the Soulseek client which port to listen on.
- The reconciler drives the Soulseek client (search, download) and the import bridge (trigger, verify).
- The Soulseek client writes into
/downloads; the import bridge moves folders within it. - The import bridge calls Lidarr’s
ManualImport; Lidarr moves the files into/musicand tags them. /musicis mounted read-only into slskarr, which uses it for sharing and for verifying what it already has.
Which do I pick: one process or the three-container design
Section titled “Which do I pick: one process or the three-container design”| If you have | One process (slskarr) | Separate client behind REST (slskd + soularr) |
|---|---|---|
| A large share index | The scan runs in-process; the API stays up | The HTTP listener serves nothing until the share scan finishes1 |
| Cross-container port-forward | The lease is a goroutine and a session field | A bash sidecar PATCHes a second container; a failed PATCH is silent |
| Quality and delay profiles | Scored natively against Lidarr’s profile mirror | Not modelled unless the bridge implements it |
| Interest in running non-root | Not possible: in-process WireGuard needs CAP_NET_ADMIN | Possible, if you do not mind the extra hop |
| Upgrades | One image, one migration set, one state model | Three projects, three release cadences, three state models |
The tradeoff is real and one-directional: consolidation buys a single state model and costs the privilege separation that separate containers gave. Everything in the ownership section below exists to pay that cost back.
Why the three containers went
Section titled “Why the three containers went”Each retired service added a failure boundary, and the interesting failures happened where two of them met.
- The bridge’s in-memory
grab_listwas the source of truth. If slskd was briefly unreachable when the poll fired, soularr cleared the dict and the actual download state became invisible. Restarting soularr forgot everything in flight. - The share scan blocked the HTTP listener. On a large library the first scan runs for minutes, and the daemon reports an empty reply from the server for the whole window.1 The bridge’s poll saw a connection error, dropped its list, and when the daemon came back the files completed with no record they had ever existed. A 2022 album was orphaned exactly this way.
- The port-forward sync was a bash sidecar in a different container, PATCHing slskd’s runtime config over REST every 15 minutes. Firing that during the scan blackout failed silently, and the published port diverged from the listening port.
- No quality-profile or delay-profile awareness. The bridge searched by text and took the first match, bypassing Lidarr’s per-album quality profile, custom formats, and “wait N minutes before falling back” knob.
- Polling on a 300s tick missed fast state changes, and re-running on the same state did not always converge.
With the client in-process, transfer state is native memory, the scan cannot block its own API, the forwarded port is a field on the session struct, and every reconcile operation is gated on durable rows.
The PIA tunnel, in-process
Section titled “The PIA tunnel, in-process”The control plane runs once at startup and then holds the session open:
| Step | What happens |
|---|---|
| 1 | Authenticate to PIA for a token, then resolve the region (PIA_REGION=sg; the value is a region id, not a city name) |
| 2 | Generate a WireGuard keypair and register the public half with PIA’s addKey endpoint |
| 3 | Bring the tunnel up through Go netlink (wg0), installing the LAN and bypass routes |
| 4 | Request a port-forward lease on that endpoint; a background refresher re-requests it roughly every 15 minutes |
| 5 | On every new forwarded port: re-bind the peer listener first, then tell the Soulseek session the new port (which re-sends the server’s SetWaitPort) |
Step 5’s order is load-bearing. A failed bind keeps the old port serving and the session is not told to publish a port nothing is listening on, which is the in-process replacement for the REST PATCH that used to fail silently.
The keypair lifecycle is the reason the tunnel cannot be left to idle. WireGuard keys expire at PIA’s end after several hours without traffic, so the reference PIA containers set PersistentKeepalive to 25 seconds on an idle link,2 and slskarr holds the same kind of continuous session rather than reconnecting per download.
Kill switch and liveness
Section titled “Kill switch and liveness”- The Soulseek session and its listener start only from the tunnel’s up callback, fired after the netlink configure call returns successfully. If the tunnel never comes up, both stay down for the process lifetime.
SLSKARR_REQUIRE_VPN=truemakes that gate explicit rather than incidental. - Separate from the firewall rules a PIA container would install, this is a kill switch by ordering: there is no window in which the client is connected and the tunnel is not, because the client has no route to the network before then.
- Liveness is read from the kernel: the peer’s latest handshake via
wg show wg0 latest-handshakes. A handshake older than three minutes means the session is stale, because WireGuard rekeys every two minutes and a stale handshake means nothing has crossed for at least one full rekey window. - A port-forward lease that ends while the tunnel is still up is a degraded state, not a failure: the tunnel stays, the client keeps working without an inbound port, and the condition is surfaced on the health endpoint as a PIA port-forward degradation.
Running as root, and keeping the library owned by Lidarr
Section titled “Running as root, and keeping the library owned by Lidarr”Why root
Section titled “Why root”CAP_NET_ADMIN covers interface configuration, routing tables, and firewall administration,3 which is exactly what creating and configuring wg0 needs. Docker’s --user 1000:100 leaves a process with an empty effective capability set, so the netlink call to create the WireGuard interface fails with Operation not permitted. The capability is added to the container (cap_add: [NET_ADMIN]), and the process runs as root so it is in the effective set.
Two sysctls come with that: net.ipv4.conf.all.rp_filter=2 and net.ipv4.conf.all.src_valid_mark=1. The second is what makes the fwmark-based policy routing survive strict reverse-path filtering, and omitting it drops inbound packets on the tunnel.2
Why that is not the end of it
Section titled “Why that is not the end of it”Root writes root-owned files. Lidarr’s containers run with PUID=1000 PGID=100 UMASK=0002, so Lidarr could read what slskarr wrote and could not move it. Its import then failed with UnauthorizedAccessException - after it had already deleted the library file it was replacing. That is the whole reason the ownership rules exist.
internal/config/perm.go applies one policy at every write into the download tree:
| Path kind | Mode | Owner |
|---|---|---|
| Directory | 0775 | SLSKARR_FILE_UID:SLSKARR_FILE_GID |
| File | 0664 | same |
Owner unset (-1, the default) | modes only, no chown | unchanged |
0775/0664 match Lidarr’s UMASK 0002, so a group member can write. The helper covers more than open: MkdirAll chmods each component it creates so the process umask cannot strip the group-write bit, and Rename re-applies the mode and owner to the destination because a rename carries the source inode’s mode with it - a folder staged before the owner was configured stays 0755/root otherwise. The chown is Lchown, so a symlink is never followed.
One-time repairs
Section titled “One-time repairs”Four startup repairs run before the reconciler starts, each guarded by a marker row in a repair_markers table so later starts are a no-op:
| Marker | What it fixes |
|---|---|
repair.false_completed_imports.v1 | Folders Lidarr marked imported while importing nothing: move them back to the pending-import state and append a pending import row to re-run the import |
repair.download_ownership.v1 | Apply owner and modes to every entry under the downloads root, skipping symlinks |
repair.quarantined_missing_artist.v1 | Folders quarantined only because their release had no artist id yet: move back and requeue |
repair.quarantined_other_album.v1 | Folders quarantined only because every item belonged to a different album: move back and requeue |
The ownership repair runs after the false-import repair on purpose, so folders that repair just moved back are covered by the walk. It logs the entry count, deletes nothing, and a missing downloads root is a no-op that is deliberately NOT marked applied - a later start with the mount present still fixes the tree. The requeue repairs rename the sidecar to <dir>.slskarr.json.requeued rather than deleting it, so the reason a folder was quarantined survives its requeue.
The Lidarr import pipeline
Section titled “The Lidarr import pipeline”Folder naming and why the first import attempt failed
Section titled “Folder naming and why the first import attempt failed”A download lands in a folder named c<candidate id> - one safe path segment, derived from the transfer row before any peer connection is opened. That is what makes the import path interesting: Lidarr’s DownloadedAlbumsScan command, given a folder and no artist, parses the artist from the folder name. c72 names no artist, so the log reads Processing path: /downloads/c72, then Parser|Unable to parse c72, then Unknown Artist c72, then Failed to import. Every slskarr import failed this way.
The command also throws ArgumentException("A path must be provided") when its Path is null or whitespace, and its body is only path, downloadClientId, and importMode - there is no artist id to pass.4
The ManualImport path
Section titled “The ManualImport path”The import bridge therefore uses Lidarr’s manual-import API, which takes the artist explicitly:
GETthe manual-import preview for the transfer’s folder with the artist id, which comes from the mirror of Lidarr’s wanted/missing list.- Keep only the items belonging to the transfer’s album, with at least one track and no
permanentrejection. - Queue
ManualImportwithimportMode: moveandreplaceExistingFiles: false, passing the item quality through as raw JSON. - If nothing is importable, or the artist id is missing, that is an error before any command is sent. The trigger refuses; it never falls back to
DownloadedAlbumsScan.
Status is not the outcome
Section titled “Status is not the outcome”Lidarr reports status: completed with result: unsuccessful when an import command matched nothing.4 Reading only the status is what produced 33 folders under imported/ with zero matching trackFileImported events in Lidarr’s history, back when the trigger sent the scan command instead of a manual import. The verification step reads both halves:
| Lidarr status | Result | slskarr outcome |
|---|---|---|
completed | successful, or absent (older Lidarr) | completed |
completed | unsuccessful | failed, lidarr imported nothing (result unsuccessful) |
completed | any unrecognised value | failed, with the value in the reason |
failed, cancelled | any | failed |
queued, started, unknown status | any | still pending, re-checked next tick |
| API error | - | error returned, row stays triggered |
The last row matters: a transient API error must never be recorded as a failed import, or a working transfer gets quarantined because Lidarr restarted.
On-disk layout
Section titled “On-disk layout”The downloads root is a four-state contract, classified by shape only:
| Path | State |
|---|---|
incomplete/<transfer id>/ | in flight; partial files carry a .part suffix |
<transfer id>/ | complete, pending import |
imported/<transfer id>/ | imported successfully |
quarantine/<transfer id>/ | quarantined: failed import, or a rejection |
failed_imports/ | sidecar area, never classified |
Classification reads the first path segment, then falls back to a .part suffix anywhere, then calls the rest pending-import. Hidden files are never classified, which keeps the sidecars out of the state machine.
A resolved import moves the folder to imported/ or quarantine/. The quarantine sidecar is failed_imports/<dir>.slskarr.json with five keys: transfer_id, lidarr_command_id (null when no import row exists), reason, quarantined_at (RFC 3339 UTC), and request_id. The cleanup is idempotent - a missing source folder is a no-op, because the move already happened - and the sidecar is still written in that case. The folder key is validated as a single safe segment (non-empty, not . or .., no path separator, no NUL) before anything touches disk.
Data-loss guards
Section titled “Data-loss guards”The import path never deletes a file, in any state, including the repairs. That rule is here because of the ordering inside Lidarr: it removes the library file it is replacing BEFORE moving the new one in. When the move failed against a root-owned download, the library tracks were already gone - three of them, restored from a ZFS snapshot (tank/media@autosnap_2026-10-01_00:00:37_monthly). Lidarr’s recycle bin is now /data/.lidarr-recycle with 30 days of retention, which is a second copy of what it removes rather than a substitute for the ownership fix.
The four guards, in the order they act:
- Ownership and modes applied at write time, so the move Lidarr wants to make is permitted in the first place.
- The
ManualImporttrigger refusing to run without an artist id, so an import that cannot match an album fails loudly instead of half-importing. - Verification reading
resultas well asstatus, so a completed-but-empty import quarantines instead of being filed as success. - Cleanup that moves and never deletes, with the reason preserved in a sidecar.
The reconciler loop
Section titled “The reconciler loop”One tick converges the whole pipeline, on a 60 second cadence:
| Step | What it does |
|---|---|
| sync | Refresh Lidarr’s wanted/missing list, quality profiles, custom formats, and artist tags into local tables |
| inventory | Walk /downloads and diff it against the transfer ids the Soulseek session says are live; orphans and missing entries are the tick’s input |
| plan searches | Enqueue searches for eligible releases, honoring the delay profile’s soulseekDelay (default 60 minutes after the Usenet delay), capped by SLSKARR_SEARCH_BUDGET (default 10 per tick) |
| start downloads | Turn the best non-rejected candidate into a transfer: write the c<id> row, then request the download per file |
| trigger imports | Trigger Lidarr imports for completed transfers, capped at 10 per tick |
| verify imports | Resolve triggered imports, apply the folder move, record the outcome |
Every step is gated on durable rows rather than memory, so a tick that dies mid-way leaves a consistent state and the next tick resumes. Re-running on unchanged state performs no actions. Each action carries the tick’s request id, and transfer events also stream over a WebSocket for the UI.
Two caps bound the loop. The import trigger cap is 10 per tick because the Lidarr command queue is serial - a large backlog fired at once would stall the resolves owed to earlier triggers behind it, so the rest of the backlog keeps its pending row and waits for the next tick. A re-search after a successful enqueue backs off from 6 hours, doubling per attempt up to a 7 day cap. The 6 hour hygiene tick (filesystem scan, detectors, queued actions) is designed but not built; the fast tick is the whole loop today.
Operational notes
Section titled “Operational notes”| Item | Value |
|---|---|
| Network | servarr bridge, in the stack’s reserved address block |
| Exposed | 8080 for the API and UI, 5030 for the Soulseek server port |
| Capability | cap_add: [NET_ADMIN], plus the two net.ipv4.conf.all sysctls |
| Volumes | SQLite database under /data, downloads root under /downloads, /music read-only for sharing and verification |
| Auth | Login page with a signed session cookie; X-Api-Key and HTTP Basic accepted for scripts; /healthz and /metrics exempt |
| Edge | A name under the internal servarr domain, through the edge Caddy proxy to the container’s API port |
Downloads share a filesystem with the library so imports are renames rather than copies. The downloads root lives on the scratch pool because the bytes are re-derivable; the library is on the redundant pool because they are not.
Reading the numbers. The production defects in the first days were all boundary problems between the two systems, not protocol problems: a SQLite writer pool that allowed concurrent writers (10 database is locked errors in the first hour, fixed by capping the pool at one connection), a listener rebind that compared address strings ([::]:53642) instead of ports after every 15 minute port-forward refresh, a stale artist id erased on every sync (385 folders requeued once fixed), and the false-completed import above. None of them were visible in a unit test; all were visible in a day of live traffic against real Lidarr.
What generalises. Pushing a process’s privilege down is only half the job when it writes to a shared tree - the other half is making the artifacts it writes usable by whoever consumes them next, at the moment of writing, not in a periodic fixup. And a verification step that reads a status without reading the outcome is not verification. Both of these cost a library file or a folder before they were fixed.
Related docs
Section titled “Related docs”- slskd and WireGuard: port forwarding with PIA - the retired three-container design, kept for the PIA control-plane mechanics slskarr reimplements in-process.
- Docker Servarr patterns - the compose conventions, PUID/PGID, and path layout this service sits inside.
- Docker Servarr stack security - the privilege-separation tradeoff that in-process WireGuard forces you to spend.
- Hot vs bulk: placing app state on a two-tier ZFS homelab - why the database and the download tree sit where they do.
References
Section titled “References”-
slskd, “The HTTP server is not active until the share scan is complete,” GitHub issue #1160. https://github.com/slskd/slskd/issues/1160 ↩ ↩2
-
thrnz, “docker-wireguard-pia,” GitHub. https://github.com/thrnz/docker-wireguard-pia ↩ ↩2
-
Linux man-pages project, “capabilities(7),” man7.org. https://man7.org/linux/man-pages/man7/capabilities.7.html ↩
-
Lidarr,
DownloadedAlbumsCommandService.cs(develop), Lidarr source. https://raw.githubusercontent.com/lidarr/lidarr/develop/src/NzbDrone.Core/MediaFiles/DownloadedAlbumsCommandService.cs ↩ ↩2