Skip to content

docs: design remote cache server API - #713

Draft
wan9chi wants to merge 9 commits into
mainfrom
remote-caching-design
Draft

wan9chi wants to merge 9 commits into
mainfrom
remote-caching-design

Conversation

@wan9chi

@wan9chi wan9chi commented Sep 8, 2026 •

Copy link
Copy Markdown
Member

Motivation

Define the remote cache server API before implementation so clients and servers can agree on a contract that preserves the existing cache workflow and keeps stored data opaque to the server.

Document fetch, store, and blob download operations, CBOR metadata with multipart blob uploads, plain-text errors, and key associations with step-by-step examples. Reference the existing task cache flow for context. Authentication and server implementation are out of scope.

wan9chi and others added 8 commits September 8, 2026 16:18
Co-authored-by: GPT-6 <codex@openai.com>
Co-authored-by: GPT-6 <codex@openai.com>
Co-authored-by: GPT-6 <codex@openai.com>
Co-authored-by: GPT-6 <codex@openai.com>
Co-authored-by: GPT-6 <codex@openai.com>
Co-authored-by: GPT-6 <codex@openai.com>
Co-authored-by: GPT-6 <codex@openai.com>
Co-authored-by: GPT-6 <codex@openai.com>
@github-actions

github-actions Bot commented Sep 8, 2026 •

Copy link
Copy Markdown

fspy benchmark

linux

dynamic/launch             change  +1.32%  [ -7.42% .. +10.18%]  overhead  +291.54%
dynamic/access             change  +0.25%  [ -6.51% .. +12.27%]  overhead    +8.79%
dynamic/access-relative    change  +0.36%  [ -2.08% ..  +2.21%]  overhead   +48.37%
dynamic/access-contended   change  +3.78%  [-10.08% .. +21.46%]  overhead   +31.61%
static/launch              change  -1.16%  [ -8.40% ..  +8.56%]  overhead  +632.35%
static/access              change  +0.83%  [ -2.84% ..  +5.07%]  overhead  +832.38%
static/access-relative     change  +0.13%  [ -3.25% ..  +3.82%]  overhead +1332.45%
static/access-contended    change  +0.46%  [ -2.84% ..  +3.19%]  overhead +3425.95%

macos

dynamic/launch             change  -0.19%  [ -3.01% ..  +2.29%]  overhead  +214.20%
dynamic/access             change  -0.33%  [ -2.72% ..  +2.39%]  overhead    +7.44%
dynamic/access-relative    change  -0.40%  [ -2.93% ..  +2.34%]  overhead  +253.97%
dynamic/access-contended   change  +0.19%  [ -4.13% ..  +2.58%]  overhead    +2.43%

windows

dynamic/launch             change  +0.32%  [-10.15% .. +13.45%]  overhead   +24.46%
dynamic/access             change  -1.49%  [-12.10% ..  +6.30%]  overhead    +1.54%
dynamic/access-relative    change  +0.17%  [-30.41% .. +14.83%]  overhead    +1.52%
dynamic/access-contended   change  +0.83%  [ -6.09% .. +44.48%]  overhead    +3.54%

Comment thread docs/remote-cache-server-api.md
Co-authored-by: GPT-6 <codex@openai.com>
wan9chi added a commit that referenced this pull request Sep 25, 2026
## Motivation

The remote cache client PRs stacked on this one need to run `vp` against
the test backend, with uploads from one step visible to the next, for
example to restore outputs after `vp cache clean`. The detached daemon
needs explicit `start` and `stop` steps and loses its state on restart.
Its fetch responses also differ from the server API in #713, which the
client will implement.

`remote-cache-server <command> [args...]` starts the backend on a free
loopback port under the `/projects/test` base path, runs the command
with `VP_REMOTE_CACHE_URL` set to that endpoint, and exits with the
command's exit code after stopping the server. The command inherits
stdio, and the wrapper takes no options. State lives in `remote-cache/`
in the working directory: `state.json` holds entries, associations, and
the next blob ID, and each blob is a file in `blobs/` named by its ID.
Consecutive steps share state, and blob IDs keep counting across
invocations, so snapshots stay deterministic.

This removes `start`, `stop`, the daemon, `cache.url`, `cache.log`, the
`--base-path`, `--max-request-bytes`, and `--max-lifetime-ms` options,
and the request size limit. `cbor-http` joins its path onto
`VP_REMOTE_CACHE_URL`, including the base path. As in #713, a fetch miss
returns 200 `{kind: "not_found"}`, and a fallback returns `{kind:
"fallback", key, value, blob_id}`.

The `remote_cache_backend` steps now run as `remote-cache-server
cbor-http ...`. The `request_limit` case is removed. `daemon_lifecycle`
is replaced by `state_across_invocations`, and a new `command_wrapper`
case covers the endpoint's base path and the exit code.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
wan9chi added a commit that referenced this pull request Sep 27, 2026
## Motivation

#727 resolves a remote cache endpoint and access mode for each
execution, but nothing uses them yet. This is the first client PR for
the server API in #713: in `read-write` mode, `vp run` uploads each
successful, cacheable new execution after recording it locally, so the
following PRs can restore it elsewhere.

The new `vt_remote_cache` crate implements `POST {endpoint}/store` over
opaque bytes. It appends `/store` to the endpoint's path, keeping any
namespace path, and sends a CBOR `metadata` part with the key, secondary
key, and value, plus the output archive streamed as the `blob` part.
Only HTTP 200 counts as success, and the response isn't decoded. The
client uses reqwest with rustls and the ring provider, verifies HTTPS
certificates with the operating system's verifier, and sets 10-second
connect and 60-second read timeouts.

`ExecutionCache` creates one client per endpoint on first use, and
uploads after a successful local update, awaiting the upload inline.
Hits never upload. The key is a header plus wincode(`CacheEntryKey`),
the secondary key is the header plus wincode(`ExecutionCacheKey`), and
the value is wincode(`CacheEntryValue`). The header contains the cache
schema version and the target OS and architecture.

An invalid endpoint or a failed upload never fails the task or changes
the exit status. The run summary shows a warning instead, even when a
single task ran, such as `vp run: build not uploaded to the remote
cache: network error.` The reason names only the kind of failure, such
as an invalid endpoint, a network error, or an HTTP status, so it's the
same on every platform. `--last-details` adds the underlying details for
each task, such as the connection error behind a network error, the
parse error for an endpoint that isn't a URL, or the message in an error
response's body.

After the wrapped command exits, `remote-cache-server` prints one line
for each request it served, such as `[remote-cache] POST /store 200` or
`[remote-cache] POST /fetch 200 not_found`. The `remote_cache_backend`
snapshots gain these lines. A new ignored `remote_cache` fixture shows
that a `read-write` run sends one store request, a rerun is a local hit
with no requests, and a `read` run after an input change reruns the task
without requests. Cases in the default suite cover an invalid endpoint
and port 0, where nothing can listen, on every platform.

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants