RAG Database with SereneDB
One rag database holds the text chunks, the embeddings and the metadata, and answers with BM25 and vector search in the same query.
The langchain-serenedb package speaks the VectorStore interface, so the retrieval half of your pipeline is a connection string rather than two systems.
▼
▼
▼
What Is a RAG Database?
A language model knows nothing about your internal documents, your tickets or your codebase. Retrieval-augmented generation fixes that by finding the relevant pieces first and handing them to the model as context. A rag database is the store that makes the first half possible: it holds the chunks, the embeddings and the metadata, and it has to retrieve well before the model can answer well.
Someone asks how do I configure CI? The question is embedded, the store returns the handful of wiki sections that match on wording and on meaning, and the model answers from those sections with citations back to them. A rag vector database that only compares embeddings covers half of that retrieval step; the ranked list is only as good as both signals together.
Why RAG Needs More Than a Vector Store
A rag vector database that stores only vectors leaves two jobs for something else to do — and "something else" is usually a second datastore.
The keyword blind spot
Pure vector retrieval always returns something, and near-neighbours are not always relevant. Ask about the SKU AB-1234, an error code or an exact function name, and the embedding has no idea what that string is — it returns plausible chunks instead of the right one. BM25 catches the literal term.
The metadata gap
RAG chunks carry metadata — source, date, department, access level — and most of it is not decoration: it decides what this user is allowed to see. A store built only for vectors either filters weakly or needs a relational database beside it, which is a second system and a sync job you now own.
How a unified engine closes both
SereneDB keeps text, embeddings and JSON metadata in one inverted index, so hybrid search and metadata filtering happen in one SQL query. The filter is applied inside both fusion branches, before the ranks are merged — so a restricted chunk never occupies a slot it would then be removed from. No separate vector DB, no separate relational DB.
How SereneDB Works as a RAG Database
As a vector database for rag it has one structure to explain and one package to install: a single index covering three column kinds, and a LangChain VectorStore that writes the SQL for you.
metadata
One index: text + vectors + metadata
One CREATE INDEX … USING inverted() names the text column with a dictionary for BM25, the vector column with ivf for ANN, and the structured or JSON columns you filter on. One inverted index over one table — not three stores to keep aligned.
CREATE INDEX docs_idx ON docs
USING inverted (id, content en_dict, embedding ivf (metric = 'l2'), metadata);Hybrid retrieval with fusion strategies
Three strategies: RRF by default, because it fuses positions and needs no calibration between BM25 relevance and vector distance; Normalized Scores when the margin between hits matters; Weighted Sum when you want to bias one signal deliberately. In LangChain it is one argument, and the metadata filter runs inside both branches before fusion.
from langchain_serenedb import HybridSearchConfig, FusionStrategy
config = HybridSearchConfig(
fusion=FusionStrategy.RRF,
primary_top_k=20,
secondary_top_k=20,
)
results = store.similarity_search(
"CI configuration", k=5, hybrid_search_config=config
)LangChain integration
The langchain-serenedb package implements LangChain's VectorStore interface, so you never write the SQL: it creates the tables and indexes and builds the queries. add_documents, similarity_search, max_marginal_relevance_search and as_retriever() all work as they do anywhere else in LangChain. RAGFlow also supports SereneDB as a doc store.
retriever = store.as_retriever(
search_kwargs={"k": 5},
)
chain = (
{"context": retriever, "question": RunnablePassthrough()}
| prompt
| llm
)Step-by-Step: Build a RAG Pipeline with SereneDB in 5 Minutes
Run SereneDB
One statically linked binary, listening on the PostgreSQL wire protocol at port 7890.
curl https://install.serenedb.com | shInstall the package
The LangChain VectorStore implementation, from PyPI.
pip install langchain_serenedbCreate the table and the index
The engine writes the DDL. Vector size has to match your embedding model.
from langchain_serenedb import SereneDBEngine, SereneDBVectorStore, IVFIndex
engine = SereneDBEngine.from_connection_string(
"host=127.0.0.1 port=7890 user=postgres dbname=postgres"
)
engine.init_vectorstore_table("my_docs", vector_size=768, vector_index=IVFIndex())Add documents
Metadata travels with the chunk and stays filterable. Embeddings can come from your own service or from ai_embed inside SQL.
from langchain_core.documents import Document
store = SereneDBVectorStore.create_sync(
engine, embedding_service=my_embeddings, table_name="my_docs"
)
docs = [
Document(page_content="SereneDB speaks the PostgreSQL wire protocol.",
metadata={"topic": "intro"}),
Document(page_content="Hybrid search fuses BM25 keyword ranking with vectors.",
metadata={"topic": "search"}),
]
store.add_documents(docs)Search — vector, then hybrid
Same call, one extra argument. Hybrid adds the BM25 branch and fuses the two rankings.
# Vector search
results = store.similarity_search("how do I configure the index?")
# Hybrid search (BM25 + vector)
from langchain_serenedb import HybridSearchConfig, FusionStrategy
results = store.similarity_search(
"index configuration",
k=5,
hybrid_search_config=HybridSearchConfig(fusion=FusionStrategy.RRF),
)Use Cases for a RAG Database
Chat-with-your-docs assistants
Internal wiki, Confluence or Notion pages split into chunks, each holding text, embedding and metadata like source and date. Hybrid search surfaces the chunks that match on wording and on meaning, so the model answers from the right sections and can cite them. Wire it up through LangChain's as_retriever().
Enterprise knowledge retrieval
Tickets, contracts and product catalogs in one database. Metadata filtering on department, date and access level runs inside the same query as the hybrid ranking — one SQL engine instead of a vector DB plus a relational DB plus a search engine. The RAGFlow integration can use SereneDB as its doc store.
Code-aware AI agents
Code and docs chunked side by side with embeddings, searched together. Keyword finds the exact function or class name the agent asked for; vector finds the code that does the same thing under a different name. ai_embed() generates the embeddings from SQL, so indexing a new branch is a query. Agents that query tables through their tools are covered under database for AI agents.
Frequently Asked Questions
A database built for the retrieval half of retrieval-augmented generation: it stores text chunks, their embeddings and their metadata, and supports both vector search and keyword search so the chunks handed to the model are the right ones.
Because a rag vector database is only part of the stack. SereneDB is a full SQL database: vector search, BM25, metadata filtering and analytics in one engine, so you do not run a vector store and a relational store side by side.
Yes — one inverted index covers text for BM25 and vectors for IVF ANN. Fusion is RRF, Normalized Scores or Weighted Sum; see the hybrid search documentation. A chunk that ranks on both signals ends up above one that ranks on either.
Yes. The langchain-serenedb package implements the VectorStore interface, hybrid search included via HybridSearchConfig. All the SQL is generated for you, so no part of the pipeline needs hand-written queries.
Yes. JSON metadata is stored and filtered in the index itself, and the filter is applied inside both fusion branches before the ranks are merged — not as a post-filter that quietly empties your top-k.
Yes — ai_embed() calls OpenAI, Gemini, Ollama or any OpenAI-compatible endpoint straight from SQL, at write time or inline in a query. No external ETL job to generate vectors.
Yes. Since August 2026 SereneDB is one of RAGFlow's doc stores, supported on both the Go and the Python backends.
Get Started with SereneDB
Install the binary, install the package, and your retrieval layer is five calls away from a working pipeline.