>
Open Source

Decide whether vector search needs a database of its own

A search feature does not need a new database merely because its interface includes an AI model. The useful decision comes earlier: does the product need to find records by meaning, or can ordinary fields and text search answer the question? That distinction determines whether an embedding index belongs in the design at all.

The source describes several open-source projects and a small Qdrant example, but the product names are less important than the job being assigned to them. Start with the query, then decide what you need to operate. Otherwise, it is easy to add another service because a tutorial used one.

Separate exact lookup from similarity

A relational database such as PostgreSQL is built for crisp conditions: a known identifier, a date range, or a price above a threshold. Those queries have explicit fields and rules. A request such as finding documents that mean something similar to a sentence is different. It asks for resemblance, not equality.

An embedding (a numeric representation produced by a model) turns text, images, or other input into a vector (an ordered list of numbers). The source explains the idea as positions in a high-dimensional space: inputs with related meaning are expected to land near one another. A vector database stores those representations and retrieves close matches for a new query vector.

Before choosing a product, write down the search request in ordinary language. These examples illustrate the distinction:

  • Exact identity: Find the order with a known ID. A conventional database query is the natural fit.
  • Structured filtering: Return records within a date range or above a price threshold. Existing fields already express the rule.
  • Meaning-based retrieval: Find a support request or document that resembles a natural-language question. This is where vector search may help.
  • Mixed constraints: Find semantically related material, but only within a language, date range, or category. The source notes that combining metadata filters with similarity search adds its own engineering work.

This is a design test, not a claim that every search box needs embeddings. If the user can state the criteria as exact fields, begin with the database and search tools already in the application. Add vectors only when the query really asks for semantic closeness.

Understand the index trade-off

A direct scan compares the query vector with stored vectors one by one. The source uses the contrast between a few thousand, ten million, and a billion vectors to explain why that approach stops being practical as a collection grows. A vector index reduces the work by organizing candidates so a search can avoid examining every item.

One method the source names is HNSW (Hierarchical Navigable Small World, an index approach for approximate nearest-neighbor search). Approximate nearest-neighbor search (finding close candidates without guaranteeing the mathematically closest result) trades exact ranking for less search work. The source describes HNSW as a popular way to make large searches practical. Editorially, treat that general performance description as a reason to measure speed and recall on your own data and configuration, not as a universal timing promise.

That trade-off is often reasonable when the user needs a useful set of related results rather than a provably exact ranking. It may not be acceptable when the ranking itself is part of a strict business or safety rule. Decide what a wrong result costs before selecting an approximate index.

Metadata filtering creates another decision. A query may need semantic similarity and constraints such as date, language, or a tag. The source points out that performing both efficiently is a separate engineering problem. Include the filters your real query needs in the first evaluation. A demo that searches only by meaning does not establish that the production query will behave well.

Choose from the system you already operate

The source’s project list covers different operating shapes, not interchangeable labels. pgvector is a PostgreSQL extension that adds vector types and similarity operators. Its appeal is operational simplicity when a project already uses PostgreSQL, while the source cautions that purpose-built options can perform better at very large scale.

For a first prototype, the source suggests Chroma or pgvector. It describes Chroma as lightweight and easy to embed in Python applications, while pgvector avoids running a separate database when PostgreSQL is already present. Those are useful starting points for different teams, not a promise that either fits every workload.

The other choices have distinct profiles in the source:

  • Qdrant: A Rust-written service with a clean API and rich filtering, described as a self-hosting option.
  • Milvus: A system aimed at very large collections, with multiple index types and integration with cloud-native infrastructure.
  • Weaviate: A schema-oriented option that supports hybrid search, combining vector similarity with keyword search, and includes modules for media.
  • FAISS, LanceDB, and Vespa: The source distinguishes FAISS as a library, LanceDB as an embedded-use option, and Vespa as capable of hybrid search.

A practical recommendation follows from those descriptions: keep the prototype near the database and deployment model the team already understands. Move to a separate service when the workload needs capabilities or performance that the existing store cannot reasonably provide. That is an editorial decision rule, not a measured result from testing these projects.

Evaluate retrieval before adding generation

Retrieval-augmented generation, or RAG (a pattern that retrieves relevant material and supplies it as context to a language model), is one reason vector search became common. The source describes a flow in which a question triggers vector search, selected data is added to the prompt, and the model answers using that context. The database is one component in the chain, not a substitute for checking whether the retrieved passages are useful.

A small evaluation set should reflect the questions users actually ask. Include queries that should find a close match, queries that should not, and queries that need a metadata filter. Compare the returned material with the expected source documents. This is practical advice derived from the source’s distinction between exact and approximate results; it does not assume that one particular benchmark or score is sufficient.

Keep retrieval quality separate from answer quality. If a generated response is wrong, inspect whether the search returned the wrong context before changing the language model prompt. If search results are sound but the answer is not, investigate the generation stage. The separation makes it easier to identify which component owns the failure.

A useful pre-deployment checklist is:

  • Query fit: Does the feature require meaning-based similarity, or will fields and text search do?
  • Filter fit: Are the date, language, and category constraints represented in the actual test queries?
  • Result quality: Are the retrieved documents relevant enough for the task, even when ranking is approximate?
  • Operational fit: Can the team run and back up the chosen store alongside its existing services?

Trade-offs

A purpose-built vector database adds another component to deploy, secure, monitor, and back up. It can provide indexing and filtering capabilities that are difficult to reproduce in an ordinary table, but those benefits need to outweigh the operational cost. pgvector has a different compromise: the source says it may not match specialized products at very large scale, while avoiding a separate database can make a smaller deployment simpler.

Approximate search also means the returned set may not be the exact nearest set. That is acceptable for some recommendation and retrieval tasks, but it should be checked against the consequence of a miss. As a practical inference, an embedding-model change deserves its own plan because the stored and new representations may not be interchangeable. The source does not specify a migration procedure, so confirm that question before changing models.

The practical trade is not “old database versus AI database.” It is one familiar system with a simpler operating footprint versus a specialized index that may better serve a particular similarity workload. Use the smallest architecture that meets the query and quality requirements you can actually demonstrate.

Bottom line

Start by describing what the user is trying to find. If the request is an ID, range, or explicit field condition, use the conventional database path. If it depends on semantic resemblance, test embeddings and vector retrieval against real queries before choosing a store.

For a prototype, the source’s Chroma and pgvector suggestions are sensible starting points for different deployment contexts. Consider Qdrant, Milvus, Weaviate, or other projects when a demonstrated need points there. The index should earn its place by answering a query your existing stack cannot handle well.

Leave a comment