Skip to main content
Semantic search has two paths: index content when it changes, then embed and query user text at request time. The embedding model is external to Vectorize; use the same model and dimensions for both paths.

Bind the index

The index is created with fixed dimensions and a distance metric. The example below assumes the embedding client returns the correct dimensions.

Index documents

Use stable IDs so re-indexing replaces the same chunks. Persist the returned ID list with the source document and supply it as previousChunkIds next time; that removes stale vectors when a document becomes shorter.
Derive the namespace from the authenticated principal, never a caller-supplied tenant ID. Metadata filters narrow eligible records before similarity scoring.

Add retrieval-augmented generation

For RAG, take the highest-scoring chunks, enforce an application score threshold, and supply their text as context to a generation model. Keep source IDs and titles in the response so the application can cite the original document.

Operational guidance

  • Pin the embedding model. When a replacement keeps the same dimensions, re-embed and upsert every document. Create a new index when dimensions or the distance metric change.
  • Batch upsert calls, with at most 1,000 IDs per batch operation.
  • Keep sensitive full documents outside vector metadata when snippets are not safe to return.
  • Measure relevance before tuning topK; more matches do not automatically improve answers.
  • Use a Stream and Pipeline ahead of the indexing Worker when document changes must be retained and replayed.