Skip to content

Documentation over SSH for coding agents

A coding agent looking up framework documentation has three options: a web search it cannot trust, a documentation MCP server that returns whole pages, or a local clone of the upstream’s docs that nobody keeps fresh. This is a fourth: a server that mirrors the sources as markdown, indexes them, and answers over SSH, because every coding agent already has a shell.

Every fact below is read from the docs-ssh repository at 952a72a (2026-09-30) and from the deployment files in it. No SSH session was opened and no endpoint was queried while writing it, so the runtime half is described from the code that implements it. Source counts come from src/application/sources.ts at that commit with the module imported and its built array counted; the token figures in the efficiency table are the repository’s own benchmark baselines, captured 2026-04-07 against third-party documentation MCP servers.

  • The corpus is a plain markdown tree under /docs/. The server is a Docker image with sshd in front of it and a ForceCommand that routes every session, so the query surface is a file tree and rg, not an API.
  • Search runs on a pre-built index, not the corpus. build-index.ts writes one TSV row per markdown file holding path, title and summary, and docs_search matches against that. The repository’s own figure is that the index is about 15 times smaller than the files it describes, which is the reason a search costs hundreds of tokens rather than tens of thousands.
  • Multi-word queries OR their tokens and rank rows by how many distinct tokens each row hits. An AND chain was the earlier behaviour, and it zero-resulted on any query phrased as a sentence.
  • A second transport carries the same six tools for chat clients that cannot open a shell: Streamable HTTP at https://docs.erfi.io/mcp. It is stateless on purpose, and what makes that possible is that the corpus is baked into the image and never written to.
  • 361 sources at the pinned commit: 180 through a git sparse checkout, 180 over HTTP with 143 of those fetched page by page, 1 over rsync.
  • Ingestion is three passes: a format converter chosen per file, an extension-based fallback, then cleanup. A converter can drop a page outright when its converted form carries no documentation.
  • OpenAPI specifications become per-tag markdown at ingestion, one file per tag plus an api/overview.md index.
  • Client configuration is generated by the server (agents, tools) from live container data. The clients that cannot use that path are hand-ported forks, and they drift.
  • Host keys are generated per container start, so a pinned known_hosts entry breaks on every deploy and every client connects with host key checking off.
upstreamfetcher (Node, build time)container (read-only image)git source x180sparse checkoutingest + normalise3 passeshttp source x180discovery methodrsync module x1IETF RFCsbuild-index.ts_index.tsvorigins_origins.tsv/docs markdown treeincl. api/*.mdsshd :2222ForceCommand log-cmdrg / bat / finddocs-mcp :8080/mcp + landing pagenode:child_processcoding agenthas a shellssh execchat clientno shellStreamable HTTP

The same picture as text:

  1. Three fetching paths supply the tree. 180 sources are cloned by sparse checkout of the upstream’s documentation directory, 180 are fetched over HTTP with 143 of those fetched page by page, and one is mirrored over rsync.
  2. Ingestion normalises every file through three passes and writes markdown under /docs/, plus a per-source stamp recording each file’s origin URL.
  3. build-index.ts scans the tree and writes /docs/_index.tsv, one row per markdown file. The origins stamps merge into /docs/_origins.tsv.
  4. The image ships the tree read-only. sshd on port 2222 routes every session through log-cmd, which either runs a builtin or executes the command against /docs/.
  5. The same image runs a second listener on port 8080 carrying the MCP endpoint and a static landing page. Both surfaces read the same tree.
SituationPathWhat it costs
Coding agent that already has a shellThe six docs_* tools over SSH, or the generated extension fileOne SSH round trip per call; the server caches identical commands
Agent that can run rg and bat but has no tool file installedRaw ssh exec against /docs/No install at all; the agent writes its own pipelines
Chat client in a web interfacehttps://docs.erfi.io/mcpNo shell needed; HTTP, and the client must support remote MCP
You, looking something up by handAn interactive ssh, or the same rg/bat/tree commandsNothing. The builtins are conveniences over the tree

The two transports are not alternatives in the sense of different capabilities. They run the same six operations with the same output formatting, over two different connections to one tree.

The argument for SSH is that the client already has one. A shell command needs no connector registration, no OAuth flow, and no per-client configuration, which matters when the same documentation should be reachable from agents on several machines and in several harnesses. The generated agent instructions are an ssh invocation with no wrapper around it.

The server constrains what a session can do in return. sshd_config sets ForceCommand /usr/local/bin/log-cmd for every session, and log-cmd.sh then routes to one of three things: an interactive shell logged in as the docs user, a builtin, or an explicit command string. Authentication is passwordless by design - PasswordAuthentication yes, PermitEmptyPasswords yes, PubkeyAuthentication no, AllowUsers docs - because the corpus is public documentation and there is nothing behind the login to protect. The filesystem is read-only and everything a client can reach is under /docs/.

log-cmd.sh reads SSH_ORIGINAL_COMMAND, takes its first word, and routes. Three cases matter.

help, sources, agents, tools and setup are dispatched to scripts in commands/. The first two force colour because their output goes to a human terminal; agents, tools and setup do not, because their output is piped into a file or read by an agent. tools pi is a sub-dispatch to a separate script.

sources is dynamic. It walks /docs/ and reports a file count and a size per source rather than reading a manifest, so nothing in the image carries a precomputed list that could disagree with the tree.

ToolWhat it readsWhy it is its own operation
docs_search/docs/_index.tsvTitle and summary only, so it is cheap enough to call first
docs_summaryThe headings of one fileThe outline, before committing to a read
docs_readOne file, whole or a line rangeLine ranges keep a large file out of context
docs_greprg inside one file or sourceRegex with context, for a phrase the index cannot match
docs_findFilenames by globWhen the name is known and the content is not
docs_sources/docs/_sources.jsonWhich sources exist, and how large each one is

All six exist in three places that share one output contract: the shell templates in src/commands/tools-template.ts and tools-pi-template.ts (generated into commands/tools.sh and commands/tools-pi.sh), and the MCP server’s src/mcp/docs-service.ts. The shell and MCP versions run the same rg/bat/find/awk pipelines and differ only in how they reach the tree, so a query returns the same shape whichever way it arrived.

The generated shell files are tracked in git. pnpm generate:tools regenerates both, and CI runs git diff --exit-code commands/tools.sh afterwards, which is what keeps that one in step with its template. The Pi file is regenerated by the same command but is not diff-checked, so it can drift without failing anything. Changing the output means editing the template and regenerating, not editing the shell.

Anything that is not a builtin runs as timeout "$CMD_TIMEOUT" /bin/bash -c "$SSH_ORIGINAL_COMMAND", with a 60-second ceiling. That is the escape hatch: rg -i 'RLS' /docs/supabase/, bat /docs/postgres/indexes.md, tree /docs/aws/ -L 2, or rg --json for a structured consumer.

Identical commands are served from a tmpfs cache. The reasoning is that the docs are immutable for the container’s lifetime, so the same command always produces the same bytes and the cache can key on the command string alone. A query the cache does not hold pays the real cost; a repeat pays nothing.

Output is capped at 51,200 characters (MAX_RESULT_CHARS in src/mcp/docs-service.ts and tools-template.ts) rather than truncated silently. The tail of an over-long result is replaced by a hint naming the file and suggesting docs_read with offset and lines, so the caller gets a next step instead of a partial page.

build-index.ts writes one TSV row per markdown file: relative path, title, summary. Titles come from frontmatter or the first heading; summaries are built from up to five leading headings, truncated at 200 characters for a title and 300 for a summary, with tabs and newlines flattened to spaces so a literal frontmatter block cannot produce a continuation line that a downstream awk reads as another row.

The parser does real YAML parsing for frontmatter, so block scalars and an embedded --- inside a string value survive, and it tracks code fences under both delimiters so a heading inside a fence does not become a title. It replaced an awk implementation that handled a narrow subset and mishandled the rest, which produced an index whose rows were wrong in ways the search could not report.

Every *.md under the docs root gets a row, including files whose title could not be extracted. Those keep an empty title field and stay discoverable by filename. That is the correct failure mode: a missing row is invisible to search, an untitled row is not.

docs_search runs rg over _index.tsv, not over /docs/. Two decisions inside that:

  • Tokens OR, they do not AND. A multi-word query is split and any token may match a row. Rows are then ranked by how many distinct tokens they hit, ties stable on the index’s own order. The AND version required every word verbatim on the same title-and-summary line, which returned nothing for a query phrased as a sentence.
  • A miss falls through to the tree, and says so. When the index returns no rows, the tool searches filenames and then file contents, and prefixes the result with [no index matches - found via filename/content search]. The prefix matters: it tells the caller the index had no summary match, so the list is a weaker signal than a normal search.

A hit list can also be truncated. Over the requested limit, the result carries [showing N of M results - refine query or add source filter] rather than silently dropping rows.

The index answers where a file is in the mirror. It cannot answer where the page lives publicly, and an agent citing /docs/supabase/guides/auth.md has cited a path that exists inside one container.

A second file covers that. /docs/_origins.tsv maps <source>/<relative path> to the file’s origin URL, built at ingestion time and persisted per source in a .stamp.json that only a real fetch rewrites, so a cached run keeps it. Then docs_read and docs_summary prepend it as a [url] line under [source], and docs_search appends it as a fourth tab-separated column. Both are an awk lookup guarded by a file-size test, so a mirror built without the file produces byte-identical output to the version before it existed.

The mapping is per source type. A source cloned from GitHub gets a blob/HEAD URL, one from GitLab -/blob/HEAD, one from Gitea or Forgejo src/branch/<ref>, an RFC gets its rfc-editor.org page, and a source whose published site does not mirror the repository layout gets an explicit mapper: src/application/public-urls.ts maps this documentation site, ingested as the erfi-technical-blog source, from src/content/docs/<slug>.mdx to https://erfi.dev/<slug>/. A converter whose output is not one-to-one with repository files carries no origin rather than a wrong one.

That column is the difference between a citation a reader can follow and a path that only exists inside a container. It is also why the corpus can cite itself: the mirror holds this site.

At the pinned commit the built array is 361 sources: 180 git, 180 http, 1 rsync. Formats after normalisation are 190 markdown, 110 HTML, 32 MDX, 23 OpenAPI, 3 text, 2 Go doc, 1 AsciiDoc.

The repository states a preference order over fetch mechanisms, most to least durable, and the counts at the pinned commit follow it:

OrderMechanismSourcesWhy it is preferred
1git sparse checkout180Markdown direct from the source, no rendering
2HTTP bulk archive (discovery: "tarball")0One archive, and survives an upstream redesign
3HTTP llms-full14A single AI-targeted dump, common and stable
4HTTP openapi or openapi-dir23A specification is data, not a page to scrape
5HTTP page-by-page discovery143Last resort; brittle to JavaScript rendering and URL changes
-rsync module1The RFC corpus, the one source with no HTTP or git form

The 143 in row five each pick a discovery method, and no ranking is claimed between them:

DiscoverySourcesDiscoverySources
toc49sitemap-index3
llms-txt38dokuwiki2
sitemap25statuspage1
mediawiki6texinfo1
rss3llms-index1
none set14

A page-by-page source is exposed to upstream JavaScript rendering, URL changes and rate limiting, and the repository’s notes record what that costs. cloudflare-blog sits behind Cloudflare’s own bot management, so the default 15-wide page concurrency trips throttling and collapses the source to zero files, which then misses the 10-minute deadline and lands an empty source in the image. A fresh CI fetch has no prior baseline, so the regression check does not catch it. Four concurrent pages with a 40-minute deadline fetch all roughly 3,500 posts with no retries or 429s in about 25 minutes. The shared default is wrong for one source, and the fix lives on that source.

Ingestion runs fixed passes in the order the normaliser array in src/index.ts declares:

  1. Format conversion, one converter per file. supportsFormat() picks the converter, so MDX and HTML each get the right one and a file that matches neither is left alone.
  2. Extension-based fallback, for a file whose format the first pass did not claim.
  3. Cleanup, every remaining normaliser. Markdown cleaning and content sanitisation run here, because their supportsFormat() returns false, which is what places them in the last pass.

Three behaviours inside that change what a file becomes rather than how it is formatted:

  • A converter may drop a page. HtmlNormaliser returns null when the Turndown output is both under 1 KB and under 1 percent of an input over 1 KB. That shape is a single-page-app shell or a paginated listing with no documentation in it, and keeping it would break the invariant that a markdown-capable source contains only markdown files. Dropped pages are logged per source during the fetch.
  • Markdown arriving as markdown skips conversion. The fetcher sends Accept: text/markdown, text/html;q=0.9, the negotiation described by the acceptmarkdown specification1 and shipped by Cloudflare as Markdown for Agents2; the underlying media type is RFC 7763’s text/markdown3. When an upstream honours it, the file is tagged as already normalised and pass one is bypassed, because running an HTML-to-markdown converter over existing markdown corrupts it through escaping. Pass three still runs. Two fallbacks guard the edges: a markdown body under 256 bytes is retried as HTML, and a 404 or 406 is retried with Accept: text/html.
  • One upstream rejects the header outright. openvpn.net answers 503 to every page carrying the markdown Accept value and 200 without it, so that source sets skipMarkdownNegotiation. The negotiation is opportunistic per source, and the opt-out is what keeps that from being a wasted fetch.

Sanitisation strips ANSI escapes, null bytes and control characters, and .. segments are removed from every path at ingestion, so a hostile upstream cannot write outside its own source directory or put a terminal-escape sequence into an agent’s context.

OpenAPI specifications to per-tag markdown

Section titled “OpenAPI specifications to per-tag markdown”

An OpenAPI specification4 is the one source kind where the input is already structured, and dumping it as JSON would waste the reader’s budget. openapi-converter.ts groups operations by their tag, falling back to the first path segment when an operation carries none, resolves $ref pointers so a schema arrives inline rather than as a reference into a component the file does not carry, and emits one file per tag under api/<tag>.md plus an api/overview.md endpoint index. Swagger 2.0 and OpenAPI 3.x are both accepted.

That layout is what makes an API surface searchable the same way as prose. The index rows are per tag, so docs_search returns one endpoint group rather than a multi-megabyte specification, and docs_read on the file returns exactly that group.

The IETF RFCs are not markdown, not in a git repository, and not behind a sitemap. They arrive over the RFC Editor’s rsync module5, which the repository’s notes record is the channel left after the RFC-all tarball was retired, and which is incremental, so a daily refresh transfers only the files that changed. Each .txt is wrapped in a fence with an H1 taken from the RFC header block and the Abstract’s first paragraph as the summary, which gives the index real titles to match. Around 1 percent of RFCs written before 1990 have free-form header blocks that defeat that extraction; they get a bare heading and stay searchable by number.

https://docs.erfi.io/mcp serves the same six tools over Streamable HTTP, the transport in the Model Context Protocol specification revision 2025-11-256. It exists for clients that cannot open a shell: a chat interface with a custom connector, a desktop client with no terminal. Add the URL as a remote MCP server and the tools appear; nothing is installed.

The implementation shares the operations with the SSH path rather than reimplementing them. src/mcp/docs-service.ts holds the six pipelines with a configurable docs root and a node:child_process runner where the shell version has an SSH hop, and src/mcp/server.ts wraps each one in a registerTool with a Zod schema. The Runner is injectable, so the pipelines are unit-tested without a real rg or bat.

The server holds no session state. Each request builds a fresh server and transport with sessionIdGenerator: undefined, and responses are plain JSON with enableJsonResponse: true rather than an event stream. That is possible because the tree is immutable for the container’s lifetime, which makes all six operations idempotent reads with no ordering between them. The consequences are the ones you want from a read-only service: no session store to lose, instances behind a load balancer with no affinity, and responses a CDN can cache.

Cross-origin protection is on (enableDnsRebindingProtection), with Host and Origin allowlists defaulting to the production hostname, the Fly hostname and localhost for hosts, and the Claude and ChatGPT web origins. Configuration is environment only: DOCS_ROOT, MCP_PORT, MCP_HOST, MCP_STATIC_DIR, MCP_ALLOWED_HOSTS, MCP_ALLOWED_ORIGINS and VERSION.

One listener carries both surfaces on port 8080: POST /mcp for the endpoint, a static landing page for every other GET, and GET /healthz for the platform health check. Serving the landing page from the same process is what avoids the one-service-per-external-port limit on Fly, which is where the design was written and is still one of the hostnames the allowlist accepts. The server compiles to one musl binary via bun build --compile, so the runtime image carries no Node and no node_modules; the Alpine base needs libstdc++ and libgcc for it. If the binary is absent, the entrypoint falls back to a busybox HTTP server for the landing page alone, which is how an older image built before the endpoint existed still serves something.

ssh -p 2222 docs@docs.erfi.io agents prints agent instructions, and the formats are pi, opencode, claude, cursor, gemini, skill (with YAML frontmatter) and a default that describes the raw SSH patterns for any agent. The output is built at request time from live container data: the source list and the file counts are read from /docs/, so an instruction file pulled after a deploy names the sources that are present rather than the ones that were present when it was written. The same generator emits a “Related source groups” section from _source_groups.json, which is derived from the per-source tag table, so an untagged source does not appear in it.

tools and tools pi are the other direction. They print a client file - a Zod schema file for OpenCode, a TypeBox extension for Pi - that wraps the same SSH commands. Nothing executes at install time; the output is a file the operator reads and saves. setup prints a structured guide an agent can follow to choose between those paths.

This repository’s clients are generated. The clients outside it are not.

The Pi extension and the Claude Code MCP toolkit used alongside it run from a hand-maintained port of the same tools in a separate dotfiles repository, and the repository states the consequence plainly: a change to the server’s output format does not reach them until someone ports it. When origin URLs were added to the docs_read header, that port was a separate change made later, and until it landed the two clients returned different headers for the same operation.

That is the general shape of the problem rather than a detail of this server. A generated client cannot lag its server; a forked one can, and nothing reports it. Prefer generation where the client can be generated - and when it cannot, treat the copy as a downstream release with its own change to make.

One workflow, .forgejo/workflows/build.yml, covers tag releases, the 02:00 UTC daily refresh and manual runs, all inside one concurrency group so two builds cannot race for the same image tag.

The workflow does not decide whether to fetch by looking at the calendar. A plan step (src/ci/plan-cli.ts over planBuild in src/ci/build-plan.ts, both unit-tested) reads docs/_build.json from the persistent cache, which records when the docs were fetched and at which commit, and then:

  • A tag build always builds, and refetches only if the docs are 20 hours or older, or the fetcher code changed since that commit.
  • A scheduled run skips the whole job when the docs are under 20 hours old and the fetcher code is unchanged. A release that just fetched them turns the next daily run into a no-op.
  • A manual run with refresh: true forces a full refetch.

When the fetcher itself changed, the per-source trust window drops to zero so every source is regenerated with the new code. Otherwise a warm run trusts sources fetched within 36 hours. The expensive step is skipped for a checkable reason rather than on a timer.

Verification gates tags, not the refresh. The verify job runs the tools-file sync check, typecheck, unit tests and the Docker end-to-end tests; a scheduled build ships main, which CI has already gated.

The cache is a Docker volume on the router’s daemon holding the fetched docs/ tree and the clone work directory, moved in and out with docker cp through a throwaway container so the runner needs no volume configuration. It is saved right after a fetch and before the image build, so a failed build does not cost the fetch. Deleting it means the next run is a cold fetch measured in hours against a three-hour runner job ceiling, which is the failure to avoid.

The deploy step posts to Composer and polls the job; a smoke job then connects to the live server. Both need a job container to reach services on the router, and a Forgejo job container cannot reach the router by its public address. The fix was infrastructure-side: split-horizon records that resolve the service hostnames to the router’s service address, and firewall rules that accept the Docker bridges on the SSH and HTTPS ports. The release that revealed the gap failed at the deploy step with a connection timeout after the image push had already succeeded, and its image was deployed by hand.

That belongs to the GitOps deploy path rather than to this server, and it is written up as a runbook in Wiring a Forgejo push to a Composer GitOps deploy, including the command that reproduces the failure. The part that transfers here is the ordering: an image that pushed successfully can still be undeployed, so a release is not finished when the registry has it.

Versioning is a git tag. pnpm release:patch (or minor, major) bumps package.json, commits, tags and pushes in one command, and the landing page reads its version from the tag at image build time.

The container is built to have as little to attack as possible. The corpus is public by design, so the threat model is the container rather than the content.

ControlSetting
Root filesystemread_only: true; /tmp, /run/sshd and /var/log are tmpfs
Capabilitiescap_drop: ALL, then CHOWN, SETUID, SETGID, SYS_CHROOT and AUDIT_WRITE added back, which is what ssh-keygen and sshd need
Privilege escalationno-new-privileges: true
Resources2 CPU and 512 MB
AuthenticationPasswordless, AllowUsers docs, PermitRootLogin no, public key auth disabled
Key exchangesntrup761x25519-sha512 first, a hybrid post-quantum exchange, with curve25519 as fallback
Ciphers and MACschacha20-poly1305 and AES-GCM only; encrypt-then-MAC only
Command routingForceCommand on every session
AuditEvery command appended to a JSONL log owned by root and group-writable by the docs user, so the session user can append but cannot truncate
IngestANSI and control bytes stripped, .. removed from paths
Output51,200-character cap, with a truncation hint naming the file

The repository publishes a token comparison between these tools and third-party documentation MCP servers. Quoted as the repository has it:

ApproachTokensAgainst the baseline
docs_search, returning file paths~48098 percent smaller
docs_summary, returning headings~200not comparable
docs_grep, returning matches~1,50077 percent smaller
A documentation MCP server returning full pages~30,000baseline

Read that table for its shape rather than its precision. The MCP baselines were captured on 2026-04-07 from real calls to vendor servers, so they describe those servers on that day and not any server today. The mechanism behind the gap is the index: a search returns paths and one-line summaries, a summary call returns headings, and a read returns only the range that matters. Three small responses replace one large one.

  • Publish an index, and search the index. The largest single decision here is that a search never opens a documentation file. Any service answering lookup questions for an agent can make the same trade, and the index is cheap because it is derived from the files it describes.
  • Keep the transport separate from the operations. One implementation of the six operations, two ways to reach it. The alternative, an HTTP server that reimplements what the shell scripts already do, is two things to keep in step.
  • Immutability is what buys statelessness. All six operations are idempotent reads because the corpus never changes inside a container. That property belongs to the deployment; without it the stateless endpoint would need a session store.
  • Generate the client, or expect it to lag. A generated client file cannot disagree with its server. A hand-ported one can, and no test notices.
  • Make the cache decision explicit. The plan step exists because fetching every time wastes hours and never fetching serves stale documentation, and the question of which is right has a checkable answer: the age of the cached tree and whether the fetcher changed.
ClaimHow it was checked
361 sources, split 180 git / 180 http / 1 rsyncRead from src/application/sources.ts at 952a72a, imported and the built array counted
Post-normalisation format splitSame import, grouped by the format field
Discovery method countsSame import, grouped by the discovery field
Index row shape, 200 and 300 character caps, five headingsRead from src/commands/build-index.ts
Three passes and the drop ruleRead from src/index.ts and the normaliser sources
Six tools, the 60-second ceiling, tmpfs cachingRead from log-cmd.sh, entrypoint.sh and src/commands/tools-template.ts
51,200-character output capMAX_RESULT_CHARS in src/mcp/docs-service.ts
Stateless MCP, JSON responses, origin allowlistsRead from src/mcp/server.ts
Container hardening and resource limitsRead from the deployment compose file
The token comparison tableThe repository’s README.md, whose figures trace to tests/benchmark/token-efficiency.test.ts

Not checked: the live server. Neither the SSH endpoint nor the MCP endpoint was queried, so every statement above about runtime behaviour is read from the code that implements it, and the per-source file counts a live sources call would report are not stated here.

  1. acceptmarkdown.com, “Serve Markdown to AI Agents with Accept Headers.” https://acceptmarkdown.com/ ↩

  2. Cloudflare, “Markdown for Agents,” Cloudflare Fundamentals docs. https://developers.cloudflare.com/fundamentals/reference/markdown-for-agents/ ↩

  3. IETF, “The text/markdown Media Type,” RFC 7763. https://www.rfc-editor.org/rfc/rfc7763.html ↩

  4. OpenAPI Initiative, “OpenAPI Specification.” https://spec.openapis.org/oas/latest.html ↩

  5. RFC Editor, “Download RFCs,” RFC Editor. https://www.rfc-editor.org/retrieve/rsync/ ↩

  6. Model Context Protocol, “Transports,” Specification revision 2025-11-25. https://modelcontextprotocol.io/specification/2025-11-25/basic/transports ↩