Skip to main content

remove_stopwords

The remove_stopwords template removes the words listed in STOPWORDS, or loaded from STOPWORDS_PATH, from the token stream rather than producing tokens of its own. Dropping very common words (the, a, is) shrinks the index and keeps high-frequency terms from dominating relevance scores.

Because it is a filter, it is normally used on the output of an earlier tokenizer, as a stage inside a pipeline after a template such as split_text, and after the normalize_tokens stage that folds the case the list is spelled in. Set HEX = true when the stop words are supplied as hex-encoded byte strings.

As a function: remove_stopwords(value, stopwords := [], stopwords_path := '', hex := false) — the value first, then the options in the order below. See tokenizer functions for how a value, a list and a chain of calls behave.

Query
SELECT remove_stopwords(['the', 'cat', 'is', 'a', 'hunter'], ['the', 'is', 'a']) AS tokens;
Result
 tokens-------------- {cat,hunter}

Options

OptionTypeDefaultDescription
STOPWORDSstring''Words to drop: a list of strings, ['the', 'a', 'an'], or one string of comma-separated double-quoted words, '"the","a","an"'. In the string form each element is whitespace-stripped and must be wrapped in double quotes, otherwise Invalid format of list of words(should be comma-separated and quoted) is raised; empty entries — "", or nothing at all between two commas — are skipped, and because the value is split on commas before the quotes are removed a stop word cannot itself contain a comma there — use the list form or HEX for that. An empty list filters nothing
STOPWORDS_PATHstring''Path to a stop-word file, or to a directory whose files are all loaded. Each line contributes the text up to its first whitespace; the words add to STOPWORDS. A path that does not exist fails at CREATE with File "<path>" referenced by option "stopwords_path" does not exist, and one that exists but yields no readable word list fails with remove_stopwords: failed to load stopwords
HEXbooleanfalseHex-decode every entry of STOPWORDS before it goes into the set, which lets a stop word hold arbitrary bytes. Hex digits of either case, an even number of them; an entry that is not valid hex fails with invalid hex stopword

Tokenization

remove_stopwords compares each token it receives against the list and drops the ones that match, passing everything else through unchanged. Matching is byte-exact: the incoming token is compared as it arrives, with no case folding, normalization, accent folding or whitespace trimming, so a list containing the does not remove The. Put a case-folding stage before it — split_text or normalize_tokens with case := 'lower' — when you want case-insensitive filtering. A surviving token keeps both its value and its offsets.

On its own the template treats the whole input value as a single token, so the result is either that one token or no tokens at all.

InputSTOPWORDSOutput tokens
the"the","a","an","is"(empty — removed)
cat"the","a","an","is"cat
Query
CREATE TEXT SEARCH DICTIONARY stop_filter AS    remove_stopwords(['the', 'a', 'an', 'is']);
SELECT ts_lexize('stop_filter', 'the');
Result
 ts_lexize----------- {}

STOPWORDS and STOPWORDS_PATH are both optional. With neither, or with a list that has no entries, the set is empty and every token passes through, so the dictionary filters nothing.

Filtering inside a pipeline

In practice remove_stopwords follows a tokenizer. A pipeline that splits on spaces and then filters drops the common words from a phrase while keeping the rest. A dropped token leaves no position gap: the surviving tokens are renumbered consecutively.

InputPipelineOutput tokens
the cat is a animalsplit_text_csv (space) → remove_stopwordscat, animal
Query
CREATE TEXT SEARCH DICTIONARY stop_pipeline AS    split_text_csv(' ') | remove_stopwords(['the', 'a', 'an', 'is']);
SELECT ts_lexize('stop_pipeline', 'the cat is a animal');
Result
 ts_lexize-------------- {cat,animal}

Hex-encoded stopwords

With HEX = true the stop words are decoded from hex before matching, so 616263 filters the token abc:

Query
CREATE TEXT SEARCH DICTIONARY hex_stop AS    remove_stopwords(['616263', '6D6E6F'], hex := true);
SELECT ts_lexize('hex_stop', 'abc');
Result
 ts_lexize----------- {}

HEX changes only how the list is read, never how a token is read: the decoded bytes are compared against the incoming token as-is. Duplicate entries are harmless, since the list becomes a set, and a token that is not valid UTF-8 is matched like any other byte string.

See also