SereneDB
> use case

RAG Database with SereneDB

One rag database holds the text chunks, the embeddings and the metadata, and answers with BM25 and vector search in the same query.

The langchain-serenedb package speaks the VectorStore interface, so the retrieval half of your pipeline is a connection string rather than two systems.

rag pipeline
Documentswiki · tickets · code
│
▼
Chunks + embeddingscontent · emb · metadata
│
▼
SereneDBone index · bm25 + ann
│
▼
LLMgrounded answer + citations
Documents from wikis, tickets, and code are split into chunks with embeddings and metadata. SereneDB indexes the text and vectors together for BM25 and ANN retrieval. Retrieved context is sent to an LLM to produce a grounded answer with citations.
> overview

What Is a RAG Database?

A language model knows nothing about your internal documents, your tickets or your codebase. Retrieval-augmented generation fixes that by finding the relevant pieces first and handing them to the model as context. A rag database is the store that makes the first half possible: it holds the chunks, the embeddings and the metadata, and it has to retrieve well before the model can answer well.

Someone asks how do I configure CI? The question is embedded, the store returns the handful of wiki sections that match on wording and on meaning, and the model answers from those sections with citations back to them. A rag vector database that only compares embeddings covers half of that retrieval step; the ranked list is only as good as both signals together.

retrieve → generate
“how do I configure CI?”
▼ embed
bm25exact terms
vector annmeaning
▼ fuse · filter metadata
1ci/cd · pipeline setup2build runners3secrets in builds
▼ top chunks → llm
Grounded answer with citations
A user asks how to configure CI. The query is embedded, then BM25 finds exact terms and vector ANN finds similar meaning. Metadata filters and fusion produce ranked chunks about CI/CD pipeline setup, build runners, and secrets. The top chunks become context for the LLM.
> the problem

Why RAG Needs More Than a Vector Store

A rag vector database that stores only vectors leaves two jobs for something else to do — and "something else" is usually a second datastore.

01 · precision

The keyword blind spot

Pure vector retrieval always returns something, and near-neighbours are not always relevant. Ask about the SKU AB-1234, an error code or an exact function name, and the embedding has no idea what that string is — it returns plausible chunks instead of the right one. BM25 catches the literal term.

02 · filters

The metadata gap

RAG chunks carry metadata — source, date, department, access level — and most of it is not decoration: it decides what this user is allowed to see. A store built only for vectors either filters weakly or needs a relational database beside it, which is a second system and a sync job you now own.

03 · one engine

How a unified engine closes both

SereneDB keeps text, embeddings and JSON metadata in one inverted index, so hybrid search and metadata filtering happen in one SQL query. The filter is applied inside both fusion branches, before the ranks are merged — so a restricted chunk never occupies a slot it would then be removed from. No separate vector DB, no separate relational DB.

> architecture

How SereneDB Works as a RAG Database

As a vector database for rag it has one structure to explain and one package to install: a single index covering three column kinds, and a LangChain VectorStore that writes the SQL for you.

loader
Documents
splitter
Chunks
embeddings
ai_embed or your service
serenedb
VectorStore
text · vector
metadata
retriever
LLM
A loader reads documents, a splitter creates chunks, and ai_embed or an external service generates embeddings. A SereneDB VectorStore indexes text, vectors, and metadata together. A retriever passes matching chunks to the LLM.

One index: text + vectors + metadata

One CREATE INDEX … USING inverted() names the text column with a dictionary for BM25, the vector column with ivf for ANN, and the structured or JSON columns you filter on. One inverted index over one table — not three stores to keep aligned.

sql
CREATE INDEX docs_idx ON docs
    USING inverted (id, content en_dict, embedding ivf (metric = 'l2'), metadata);
One SQL inverted index covers document text for BM25, embeddings with IVF for ANN, and metadata used to filter retrieval.

Hybrid retrieval with fusion strategies

Three strategies: RRF by default, because it fuses positions and needs no calibration between BM25 relevance and vector distance; Normalized Scores when the margin between hits matters; Weighted Sum when you want to bias one signal deliberately. In LangChain it is one argument, and the metadata filter runs inside both branches before fusion.

python
from langchain_serenedb import HybridSearchConfig, FusionStrategy

config = HybridSearchConfig(
    fusion=FusionStrategy.RRF,
    primary_top_k=20,
    secondary_top_k=20,
)
results = store.similarity_search(
    "CI configuration", k=5, hybrid_search_config=config
)
LangChain hybrid retrieval combines text and vector signals with a selected fusion strategy. Metadata filters run inside both branches before fusion. Available strategies are Reciprocal Rank Fusion, Normalized Scores, and Weighted Sum.

LangChain integration

The langchain-serenedb package implements LangChain's VectorStore interface, so you never write the SQL: it creates the tables and indexes and builds the queries. add_documents, similarity_search, max_marginal_relevance_search and as_retriever() all work as they do anywhere else in LangChain. RAGFlow also supports SereneDB as a doc store.

python
retriever = store.as_retriever(
    search_kwargs={"k": 5},
)

chain = (
    {"context": retriever, "question": RunnablePassthrough()}
    | prompt
    | llm
)
The SereneDB VectorStore becomes a LangChain retriever using as_retriever. It retrieves matching chunks from the combined text, vector, and metadata index for the RAG pipeline.
> build it

Step-by-Step: Build a RAG Pipeline with SereneDB in 5 Minutes

01

Run SereneDB

One statically linked binary, listening on the PostgreSQL wire protocol at port 7890.

shell
curl https://install.serenedb.com | sh
Step 1: Run SereneDB.
02

Install the package

The LangChain VectorStore implementation, from PyPI.

shell
pip install langchain_serenedb
Step 2: Install the package.
03

Create the table and the index

The engine writes the DDL. Vector size has to match your embedding model.

python
from langchain_serenedb import SereneDBEngine, SereneDBVectorStore, IVFIndex

engine = SereneDBEngine.from_connection_string(
    "host=127.0.0.1 port=7890 user=postgres dbname=postgres"
)
engine.init_vectorstore_table("my_docs", vector_size=768, vector_index=IVFIndex())
Step 3: Create the table and the index.
04

Add documents

Metadata travels with the chunk and stays filterable. Embeddings can come from your own service or from ai_embed inside SQL.

python
from langchain_core.documents import Document

store = SereneDBVectorStore.create_sync(
    engine, embedding_service=my_embeddings, table_name="my_docs"
)

docs = [
    Document(page_content="SereneDB speaks the PostgreSQL wire protocol.",
             metadata={"topic": "intro"}),
    Document(page_content="Hybrid search fuses BM25 keyword ranking with vectors.",
             metadata={"topic": "search"}),
]
store.add_documents(docs)
Step 4: Add documents.
05

Search — vector, then hybrid

Same call, one extra argument. Hybrid adds the BM25 branch and fuses the two rankings.

python
# Vector search
results = store.similarity_search("how do I configure the index?")

# Hybrid search (BM25 + vector)
from langchain_serenedb import HybridSearchConfig, FusionStrategy

results = store.similarity_search(
    "index configuration",
    k=5,
    hybrid_search_config=HybridSearchConfig(fusion=FusionStrategy.RRF),
)
Step 5: Search — vector, then hybrid.
> use cases

Use Cases for a RAG Database

Chat-with-your-docs assistants

Internal wiki, Confluence or Notion pages split into chunks, each holding text, embedding and metadata like source and date. Hybrid search surfaces the chunks that match on wording and on meaning, so the model answers from the right sections and can cite them. Wire it up through LangChain's as_retriever().

Enterprise knowledge retrieval

Tickets, contracts and product catalogs in one database. Metadata filtering on department, date and access level runs inside the same query as the hybrid ranking — one SQL engine instead of a vector DB plus a relational DB plus a search engine. The RAGFlow integration can use SereneDB as its doc store.

Code-aware AI agents

Code and docs chunked side by side with embeddings, searched together. Keyword finds the exact function or class name the agent asked for; vector finds the code that does the same thing under a different name. ai_embed() generates the embeddings from SQL, so indexing a new branch is a query. Agents that query tables through their tools are covered under database for AI agents.

> faq

Frequently Asked Questions

A database built for the retrieval half of retrieval-augmented generation: it stores text chunks, their embeddings and their metadata, and supports both vector search and keyword search so the chunks handed to the model are the right ones.

Because a rag vector database is only part of the stack. SereneDB is a full SQL database: vector search, BM25, metadata filtering and analytics in one engine, so you do not run a vector store and a relational store side by side.

Yes — one inverted index covers text for BM25 and vectors for IVF ANN. Fusion is RRF, Normalized Scores or Weighted Sum; see the hybrid search documentation. A chunk that ranks on both signals ends up above one that ranks on either.

Yes. The langchain-serenedb package implements the VectorStore interface, hybrid search included via HybridSearchConfig. All the SQL is generated for you, so no part of the pipeline needs hand-written queries.

Yes. JSON metadata is stored and filtered in the index itself, and the filter is applied inside both fusion branches before the ranks are merged — not as a post-filter that quietly empties your top-k.

Yes — ai_embed() calls OpenAI, Gemini, Ollama or any OpenAI-compatible endpoint straight from SQL, at write time or inline in a query. No external ETL job to generate vectors.

Yes. Since August 2026 SereneDB is one of RAGFlow's doc stores, supported on both the Go and the Python backends.

> get started

Get Started with SereneDB

Install the binary, install the package, and your retrieval layer is five calls away from a working pipeline.