Intelligent Internet's Evoke Brings Hybrid Search Inside PostgreSQL Without Extra Services
Intelligent Internet released Evoke, a Postgres extension that folds sparse semantic retrieval into the BM25 inverted index, no vector store or GPU required.
- Intelligent Internet released Evoke, an Apache-2.0 Postgres extension unifying BM25 and sparse semantic retrieval in one index.
- Built on IBM's 30.3M-parameter Granite-Embedding-30M-Sparse; runs on CPU inside Postgres with no GPU.
- Recall@100 on BEIR15 jumps from 0.563 (BM25) to 0.667, within 0.004 of a 0.6B dense model.
- Writes commit immediately; shared background workers encode semantic atoms asynchronously.
- Ships as Docker image
ghcr.io/intelligent-internet/ii-42:pg18-v0.2.5with model bundled, air-gap friendly. - API uses
CREATE INDEX ... USING ii42withsae = trueandii42_query(...), see the launch post.
Evoke puts hybrid search inside PostgreSQL
Adding semantic search to PostgreSQL usually requires an embedding API, a vector index beside the keyword index, a synchronization job, and a step that reconciles incompatible score scales. Intelligent Internet has released Evoke, an Apache-2.0 Postgres extension that stores lexical and semantic evidence in one inverted index and runs its encoder inside the database on CPU.
The initial release supports PostgreSQL 17 and 18 on Linux x86-64. A Docker image bundles the model and runtime. According to the release announcement, the design eliminates a separate retrieval service and keeps indexed content aligned with the database through shared background workers.
One index, two retrieval signals
Evoke uses learned-sparse retrieval. Typical dense encoders represent text as a fixed-length vector, often with hundreds of dimensions, which requires an approximate nearest-neighbor index. Evoke’s encoder produces weighted terms from a fixed vocabulary. Those terms can include related words absent from the source text, allowing the index to match concepts as well as literal words.
The encoder builds on IBM’s Granite-Embedding-30M-Sparse model, which has about 30.3 million parameters. Evoke stores its semantic terms in a dedicated namespace within the same inverted index that holds BM25 keyword postings.
At query time, the index-local scorer adds the BM25 keyword score and sparse semantic score. This structure removes reciprocal-rank fusion, cross-index normalization, and synchronization between separate keyword and vector indexes.
Company tests show higher recall
Intelligent Internet compared Evoke with BM25 and with pplx-embed-v1-0.6B, a dense model roughly 20 times larger served through VectorChord. Recall@100 measures how many known relevant documents appear among the first 100 results.
| Benchmark suite | BM25 Recall@100 | Evoke Recall@100 | Absolute gain |
|---|---|---|---|
| BEIR15 | 0.563 | 0.667 | 0.104 |
| MTEB10 | 0.595 | 0.703 | 0.108 |
Against the dense baseline, Evoke came within 0.004 on Recall@100 and retrieved more relevant documents by rank 1,000 on both suites. First-stage retrieval determines the candidate pool available to a reranker, which cannot recover documents omitted from that pool.
The published figures cover retrieval quality. Production evaluations still need corpus-specific measurements for query latency, indexing throughput, queue lag, index size, memory use, and recall.
Writes converge in the background
Evoke separates transaction latency from document encoding. Inserts and updates record keyword evidence immediately, then queue semantic encoding for shared background workers. New content can therefore appear in keyword results before its semantic postings are ready.
- Local CPU inference: The bundled setup uses ONNX Runtime 1.29.0 and performs encoding on CPU.
- Shared model memory: Database backends send encoding work to shared workers, which keep one model runtime in memory.
- Local text processing: Text remains inside the database during encoding, supporting private and air-gapped deployments.
- MVCC-aware results: Evoke checks index matches against the current heap version, preventing deleted rows from resurfacing through stale postings.
- Eventual semantic indexing: Background workers must process queued rows before inserts and updates appear with current semantic evidence.
A small SQL surface
The extension uses the internal name ii42. BM25 indexing is the default, while the sae index option enables sparse semantic retrieval.
CREATE EXTENSION ii42;
CREATE INDEX docs_semantic_idx
ON docs USING ii42 (body)
WITH (sae = true);
SELECT d.id,
d.title,
ii42_query(
'docs_semantic_idx'::regclass,
'database search architecture'
) AS score
FROM docs AS d
ORDER BY score DESC
LIMIT 10;Linux x86-64 packages are available for PostgreSQL 17 and 18. The fastest packaged setup uses the versioned Docker image:
docker pull ghcr.io/intelligent-internet/ii-42:pg18-v0.2.5The image includes the model and ONNX runtime. Source installations require a separate model download from Hugging Face.
Before production: language, plans, replicas
The Beta 1 model targets English text and uses a fixed lexical vocabulary with calibration derived from NFCorpus. Teams working with specialized terminology should test recall on their own corpus and evaluate a compatible custom model checkout where needed.
- Query shape: The planner-native semantic path supports one base table, ordinary
WHEREpredicates, descending rank, and a boundedLIMIT. Joins, row locking, and global ranking through a partitioned parent fail closed when the extension cannot preserve ranking semantics. - Parallel execution: This release omits parallel heap builds, parallel access-method scans, and parallel
VACUUMdiscovery. - Physical replication: Standbys require the same extension binary, ONNX Runtime, and model checkout as the primary.
- Logical replication: PostgreSQL replicates table rows rather than index relations, so subscribers must build and maintain their own Evoke indexes.
- Consistency: Keyword evidence is available immediately, while semantic evidence appears after background processing completes.
Evoke fits English-language PostgreSQL applications that want hybrid retrieval without operating a separate vector service. Applications requiring dense multilingual retrieval, specialized vector operations, or sub-millisecond approximate nearest-neighbor search across billions of vectors remain better suited to dedicated vector infrastructure.