Home of the LanceDB documentation. Built using Mintlify.
The published site is assembled from three roots, listed in assemble.yaml:
the open-source pages in lancedb/lancedb and the Enterprise pages in
lancedb/sophon, both under docs/web, and this repository's docs/.
This repository's docs/ holds the pages it still owns: the Geneva pages, the
dataset cards and the REST API reference. The Geneva pages stay published until the official Function
launch, because the Function pages do not replace their APIs. They and their
snippets (docs/snippets/geneva_*.mdx) are kept as published at
deploy-freeze; the tests that generated those snippets need the geneva
package and are not in this repository, so the snippets are not regenerated.
Install the Mintlify CLI at the version CI uses:
npm i -g mint@4.2.888With the three repositories checked out side by side, build and preview the site:
make assemble
cd build/site && mint devTo build from other checkouts or worktrees, point LANCEDB_DOCS_ROOT and
SOPHON_DOCS_ROOT at their docs/web directories. The assembler prints the
commit it read each root from, so every build names its inputs.
Check the assembled site the way CI does, and run the assembler's own tests:
(cd build/site && mint validate && mint broken-links --check-anchors --check-redirects)
make test-assemblecd docs && mint dev previews only this repository's pages. Links into the
open-source and Enterprise pages do not resolve there, so check links on the
assembled site.
Merging publishes nothing. Every pull request and every push to main runs the
Assemble workflow, which builds the site from the three roots, checks it as
above, and keeps the checked tree as the run's candidate artifact, with a
record of the commit each root was read from and a SHA-256 for every file. A
checksum over those hashes names the candidate, and the run's summary shows it.
That workflow's token can only read.
Publishing is a separate, manual step: run the Publish workflow with the ID of a successful Assemble run and its candidate's checksum.
stagingputs the candidate on thestagingbranch, for a Mintlify preview of the combined site. A preview of any other branch of this repository is not one: it holds onlydocs/.productionputs it onassembled. It takes only a candidate built onmainfrom both producers'main, runs only frommain, and pushes with theproductionenvironment's deploy key. It waits for approval only if that environment requires reviewers.
Either way the published files are the candidate's, checked against its record
and the checksum, and the record must name exactly the three source commits and
both producers' refs. Nothing is rebuilt, so newer source commits cannot slip in.
Each publication is a new commit on its branch that names the run, the checksum
and the source commits, and nothing is force-pushed. To roll back, publish an
earlier candidate again: its run ID and checksum are in its commit on
assembled. Once its artifact has expired, after 90 days, that earlier commit's
files are published again, after they are checked against the checksum.
scripts/candidate.py records and publishes; make test-assemble runs its
tests against disposable local repositories.
Publishing relies on settings outside this repository:
- A
productionenvironment with required reviewers, deployments limited tomain, and anASSEMBLED_DEPLOY_KEYsecret: the private half of the only deploy key with write access. - A ruleset on
assembledthat restricts updates and deletions and blocks force pushes, with deploy keys as its only bypass. Without it, any workflow token that can write could changeassembled. - A repository variable
MINTLIFY_CONTENT_DIR: the path Mintlify's Git settings readdocs.jsonfrom, such as/docs, or/for the root. Publishing refuses to guess. - Mintlify's Git settings: repository
lancedb/docs, branchassembled, and the content directory above. Mintlify serves the branch its settings name, so publishing toassembledreaches production only once that is the branch. - The LanceDB Docs Reader GitHub App installed on
lancedb/sophononly, with Contents: read and the required Metadata: read permission. Set its Client ID in theSOPHON_DOCS_APP_CLIENT_IDrepository variable and its PEM private key in theSOPHON_DOCS_APP_PRIVATE_KEYrepository secret. The Assemble and Docs Check workflows generate a short-lived token scoped to Sophon with Contents: read. Checkout does not persist it, and the token is revoked when the job ends. Fork pull requests receive no App secret and cannot run the combined build; the assembler guard tests still run without credentials.
Of these, the Publish workflow checks only that production runs from main,
that the environment supplies the deploy key and that MINTLIFY_CONTENT_DIR is
set. It cannot see whether the environment requires reviewers and admits only
main, or whether the ruleset refuses every other writer: verify those
separately before the first production publication.
The code examples on the open-source pages are tested programs in
lancedb/lancedb, under docs/web-tests/{py,ts,rs}. Their snippets are
generated there, into docs/web/snippets/, and committed beside the tests they
come from; this repository generates none. To change an example, edit its test
in a lancedb checkout and regenerate from that repository's root:
uv run docs/web-tests/mdx_snippets_gen.py -s docs/web-tests/py -s docs/web-tests/ts -s docs/web-tests/rs -o docs/web/snippetsThe Documentation section of lancedb's CONTRIBUTING.md describes the same
workflow. Prefer an example in a test over code written into a page.
The only snippets in this repository are the four Geneva ones,
docs/snippets/geneva_*.mdx. They are frozen copies, as published at
deploy-freeze, and stay unchanged until the Geneva pages are retired at the
official Function launch. make snippets stops and prints these instructions.
The Datasets tab is populated from lance-format/lance-huggingface,
the master repository where each Lance dataset published under the lance-format
Hugging Face organization has its own directory with an HF_DATASET_CARD.md. That same file is what gets pushed to
the Hub as the dataset's README.md via the hf CLI, so the GitHub repo is the single source of truth for the
content of every dataset card.
To avoid maintaining the same content in two places, the per-dataset MDX pages under docs/datasets/ are
generated from those upstream cards via scripts/sync_hf_datasets.py. The script:
- Reads
scripts/hf_datasets.yaml, which lists every dataset to publish and maps the upstream directory name, the URL slug, the HF Hub repo, and the human-readable title. - Fetches each
HF_DATASET_CARD.mdfromlance-format/lance-huggingfaceon GitHub. - Rewrites the frontmatter for Mintlify (sets
title,sidebarTitle,description), strips the upstream H1, injects a "View on Hugging Face" card at the top, and sanitizes known MDX hazards (bibtex citations outside code fences, literal<>in prose). - Writes
docs/datasets/<slug>.mdx, regenerates the card grid indocs/datasets/index.mdxbetween theHF_SYNC:START/HF_SYNC:ENDmarkers, and updates theDatasetstab indocs/docs.nav.jsonto keep the sidebar in sync.
Run it from the repo root:
make hf-sync- Author the new dataset's
HF_DATASET_CARD.mdupstream inlance-format/lance-huggingface(and push it to the Hub as usual). - Add a single line for the dataset under the appropriate category in
scripts/hf_datasets.yaml. The four fields (dir,slug,hf,title) are explicit because the GitHub directory name, the HF Hub repo slug, and the desired URL slug don't follow a derivable convention. - Run
make hf-sync. The script will fetch the new card, generatedocs/datasets/<slug>.mdx, refresh the landing-page card grid, and add the new page to theDatasetstab indocs/docs.nav.json. - Run
make assemble, then preview withcd build/site && mint dev(see Development for the required source checkouts). Commit the MDX page, regeneratedindex.mdx, updateddocs.nav.json, and new yaml entry.
If you remove a dataset from the yaml, the next make hf-sync will delete its MDX file and drop the sidebar
entry. The script hard-fails on any fetch error — partial regeneration would be worse than a clear error.