Documentation Search with SereneDB
Self-hosted documentation search: point it at a Git repo, a website or an S3 bucket and it indexes into SereneDB.
BM25 full-text works with no model at all; hybrid vector search, cited AI answers and an MCP server for agents are each an opt-in. Open source, Apache 2.0.
What Is Documentation Search?
Documentation search is search over structured content — pages, sections, code samples — ranked by relevance and answered while the reader is still typing. It is not a site-wide crawl that returns a list of URLs. The unit of a result is a section, so the reader lands on the paragraph that answers the question rather than at the top of a 4000-word page.
A developer types how to create an index. On the third keystroke the modal already shows matching sections, grouped by area, with the matched terms highlighted in the snippet. Enter jumps to the anchor. That is the whole interaction, and everything a documentation search engine does — tokenizing, stemming, full-text search ranking, snippet extraction — exists to make those three keystrokes land.
Why Documentation Search Is a Solved Problem That Keeps Coming Back
The SaaS lock-in problem
Algolia DocSearch is the default answer, and for public open-source docs it is a good one. But the index lives on someone else's servers, the free tier comes with conditions, and the paid tier grows with your content. Private docs or an enterprise review turn "someone else's servers" into the whole conversation.
The keyword-only problem
Keyword matching does not know that db means database or k8s means kubernetes, so the reader gets zero results and assumes the feature does not exist. Pure semantic search has the opposite failure: it returns approximately the right page and loses the exact match on an error string or a function name.
How self-hosted hybrid search solves both
Serene Docs Search keeps the data and the index under your control. BM25 full-text needs no model and costs nothing to run; hybrid search adds vector similarity fused with RRF. Synonyms (db, database; k8s => kubernetes), stemming and typo correction handle the rest. AI answers stay optional, with any provider — OpenAI or a local Ollama.
How SereneDB Powers Documentation Search
txt · ipynb · pdf
url mapping
synonyms · stemming
claude · codex
Serene Docs Search — the ready-made app
Serene Docs Search is a self-hosted application, not a library you assemble. The stack is SereneDB plus a search backend, shipped as Docker images. The front end is a React widget or a single script tag. A configurator wizard writes your config.json and docker-compose.yml for you, so the first deploy is a download and one command. Apache 2.0.
Sources: Git, folder, website, S3
Four source types: a Git repository, a local folder, a live website via crawl and sitemap, or an S3-compatible bucket. Formats: Markdown, MDX, HTML, RST, plain text, Jupyter notebooks and PDF. Markdown can be split by heading, so a result points at the section that matched instead of the page that contains it. URL mapping rewrites file paths to the domain the docs are served from.
Search modes: full-text, hybrid, AI answers
Full-text. BM25, search-as-you-type with the last term treated as a prefix, stemming, synonyms, typo correction, snippet highlighting. No model, so a docs search engine in this mode has no AI cost at all.
Hybrid. BM25 and vector similarity fused with RRF. Embeddings from OpenAI, Ollama or any OpenAI-compatible API; weight, window and distance threshold are yours to set.
AI answers. A separate opt-in. The model searches the same index and streams an answer with citations, multi-turn. Answer model and embedding model are configured independently.
MCP server. Optional. Lets Claude, Codex and other agents search and read that same index.
Step-by-Step: Deploy Documentation Search in 5 Minutes
Configure the source and the search mode
Open the configurator wizard on the Serene Docs Search page. Pick the source — a Git repo URL — the file formats to index, and the mode: full-text or hybrid.
Generate and download the configuration
Select Generate deploy files. Download serene-search.config.json and copy the docker-compose.yml next to it.
Start the stack
Two containers: SereneDB and the search backend.
docker compose up -dRun the initial index build
Back in the wizard, enter the backend URL, hit Test connection, then start the initial index build.
Add the widget to your site
Install the React widget and point it at your backend. A script tag works too if your docs are not React.
npm install @serenedb/docs-search-react@latest
import { SereneDocsSearch } from "@serenedb/docs-search-react";
import "@serenedb/docs-search-react/styles.css";
<SereneDocsSearch backendUrl="https://search.example.com" />;Use Cases for Documentation Search
Open-source project docs
Markdown in a Git repository. Serene Docs Search watches the branch and refreshes the index on commit. BM25 full-text costs nothing to run; hybrid is there when you want it. The React widget gives readers ⌘K. It is an Algolia DocSearch alternative with no external dependency and no data leaving your infrastructure.
Internal knowledge base
Confluence, Notion or a wiki, exported to Markdown or HTML in a bucket or a local folder. Hybrid search covers both the exact phrase and the vaguely remembered one, and AI answers cite the page they came from. As a knowledge base search engine it keeps everything inside your network: nothing goes to a third-party SaaS.
AI agent access via MCP
The optional MCP server hands the same index to agents. Claude or Codex can search your docs and read full pages through standard MCP tools — AI documentation search for the thing writing the code, not just the person reading it. For agents that query tables rather than docs, see database for AI agents.
claude mcp add docs-search \ --transport http \ https://search.example.com/mcp
Frequently Asked Questions
Search over a product's or project's documentation: instant search-as-you-type, relevance ranking with BM25, section-level linking and snippet highlighting. It can be extended with vector search and AI answers, but the baseline is a ranked list of sections while the reader types.
Through the Serene Docs Search guide and its self-hosted app: SereneDB holds the index, the backend serves queries, and a React widget renders results. Run it in BM25 full-text mode or turn on hybrid.
For a typical documentation search, yes: self-hosted, open source under Apache 2.0, data under your control. As a knowledge base search engine it covers the core scenario rather than every feature of a search SaaS, and we would rather say that plainly.
Yes. Ask AI streams an answer with citations over the same index. The answer model and the embedding model are configured independently, and both can be OpenAI, Ollama or any OpenAI-compatible API.
A Git repository, a local folder, a live website via crawl and sitemap, or an S3, R2 or MinIO bucket. Formats: Markdown, MDX, HTML, RST, plain text, Jupyter notebooks and PDF.
No. Full-text mode uses BM25 with no model at all. Hybrid search and AI answers are optional opt-ins, and each is independent of the other — you can run one, both or neither.
Yes. The optional MCP server lets Claude, Codex and other agents search and read the same index through standard MCP tools — the same documentation search engine your readers use, exposed as tools.
Get Started with SereneDB
Run the configurator, start two containers, drop in the widget. Full-text mode needs no model and no API key.