Skip to main content

LangChain

langchain-serenedb is a LangChain vector store backed by SereneDB. SereneDB speaks the PostgreSQL wire protocol, so the package connects with psycopg 3 and maps the VectorStore contract onto SereneDB's native features: fixed-size FLOAT[N] vector columns, the inverted index with an ivf operator class for approximate nearest neighbor search, and BM25 relevance scoring for the lexical half of a hybrid query.

What it maps onto SereneDB

LangChain conceptSereneDB featureReference
Embedding columnFLOAT[N] array columnArray
Distance metrics<->, <=>, <#>, <+> and their named functionsVector functions
ANN indexCREATE INDEX ... USING inverted (embedding ivf (metric = '...'))Vector search
Hybrid retrievalOne inverted index over the content and embedding columns, scored with BM25()Hybrid search
MetadataA JSON column, plus optional typed columnsIndexes and Tuning
Metadata filtersIndex-covered predicates, including @@ full-text matchesMetadata Filtering
Write visibilityVACUUM (REFRESH_TABLE)Maintenance

Install

pip install langchain-serenedb

The distribution is named langchain-serenedb; the import name is langchain_serenedb.

RequirementVersion
Python>= 3.10
langchain-core>= 1.2.11, < 2.0
psycopg[binary]>= 3, < 4
psycopg-pool>= 3.2.1, < 4
numpy>= 1.21, < 3

You also need a running SereneDB server. Note the port: SereneDB listens on 7890 by default, not PostgreSQL's 5432.

Quickstart

from langchain_core.embeddings import DeterministicFakeEmbedding

from langchain_serenedb import IVFIndex, SereneDBEngine, SereneDBVectorStore

embeddings = DeterministicFakeEmbedding(size=768)

engine = SereneDBEngine.from_connection_string(
"host=127.0.0.1 port=7890 user=postgres dbname=postgres"
)

# Create the table and its IVF index in one call, so vector search is accelerated
# from the start. All the DDL is generated for you.
engine.init_vectorstore_table("my_docs", 768, vector_index=IVFIndex())

store = SereneDBVectorStore.create_sync(engine, embeddings, "my_docs")

store.add_texts(
["SereneDB indexes vectors with IVF.", "BM25 ranks full-text matches."],
metadatas=[{"topic": "vector"}, {"topic": "text"}],
)

for doc in store.similarity_search("how are vectors indexed?", k=2):
print(doc.page_content, doc.metadata)

engine.close()

This gives you a my_docs table holding both documents — each with its text, a 768-dimensional embedding and its metadata dict — with the embeddings indexed for approximate nearest neighbor search ranked by cosine similarity. similarity_search returns the nearest documents with their metadata attached. DeterministicFakeEmbedding keeps the example runnable without an API key, though it hashes text rather than modelling meaning, so the ranking it produces is arbitrary. In a real application, swap it for any LangChain embeddings model and set vector_size to that model's dimension.

Where to go next

PageContents
Engine and TablesSereneDBEngine, Column, ColumnDict — connecting, creating the table, publishing writes
Vector StoreSereneDBVectorStore — adding documents, searching, scores, retrievers
Metadata FilteringThe filter dict: comparison, set, pattern, logical and full-text operators
Hybrid SearchHybridIndexConfig, HybridSearchConfig, FusionStrategy — fusing BM25 and vector rankings in one query
Indexes and TuningDistanceStrategy, IVFIndex, IVFQueryOptions, MetadataIndexConfig, MetadataColumnIndex, JsonFieldIndex
Async APIAsyncSereneDBVectorStore — the async core, the event loop model, async-only methods
FAQCommon problems and their fixes, indexed by symptom

Package version

import langchain_serenedb

print(langchain_serenedb.__version__)

__version__ is read from the installed distribution metadata, so it is the empty string when the package is imported from a source tree that was never installed.

The package passes LangChain's standard VectorStoreIntegrationTests compliance suite for both the synchronous and the asynchronous store.

For plain SQL access from Python without LangChain, see Python. For the index that powers all of this, see Inverted Index.