
One-sentence definition
Sistema que almacena vectores y permite recuperar elementos próximos según una métrica de similitud.
What Is a Vector Database?
A vector database is a system that stores vectors—lists of numbers representing data—and lets you search for vectors that are close to a query vector according to a chosen metric. Its purpose is to make similarity search over numerical representations practical. For example, it can help retrieve texts, products, or images related to a query.
The database works with the vectors it receives. A model or another component typically transforms the original data into a numerical representation called an embedding. Generating an embedding and searching with it are separate operations: storing vectors does not mean the database created those representations, understood their content, or confirmed that two items are equivalent.
In an application, a vector is usually associated with an identifier and may also be stored with metadata, such as a category or date, and a reference to the original item. The complete content may live in the same system or somewhere else. Which data types a product accepts, which features it offers, and how it implements them depend on that product; “vector database” does not describe a single, universal architecture.
What a Record Stores and How a Query Works
A vector record may contain an identifier, the vector itself, metadata, and a reference to the associated item. For a documentation passage, for instance, a record might include its vector, a document identifier, the language, and the passage’s location. The original text could be stored in the database or retrieved from another storage system.
A typical workflow has several stages. The application takes the data it wants to index and sends it to an embedding model; it receives a vector and stores that vector with an identifier. When a query arrives, the system generates a compatible query vector—or receives one from another component—and asks for the nearest neighbors. The application can then fetch the original items, apply business rules, or present the results.
Compatibility between stored vectors and the query vector matters. If they were produced by different models, incompatible configurations, or representation spaces that do not correspond, comparing them may not be useful. The database cannot automatically fix an unsuitable representation.
From data to results
- 01Prepare the data and generate its embedding with a model suited to the task.
- 02Store the vector with an identifier and, when needed, metadata or a reference to the original item.
- 03Convert the query into a compatible vector and search for neighbors using the selected metric.
- 04Retrieve associated data and assess the results using criteria and rules specific to the application.
How Similarity Is Measured
A search needs a rule for comparing vectors. Common measures include cosine similarity, inner product, and Euclidean distance. These terms are not interchangeable: they describe different relationships, and their results may be ranked or interpreted differently.
Cosine similarity compares the orientation of vectors; Euclidean distance measures the separation between their points; and inner product combines vector components and can depend on vector magnitude. The choice should not be made out of habit alone. It should match the model, the way its vectors were generated or normalized, and the application’s goal.
Check the configuration end to end: which metric the implementation offers, which one the index uses, and how results are ranked or presented. A label such as “similarity” does not guarantee that two products calculate the same value or that a high value has the same interpretation in every system.
A guide to choosing a measure
This table is conceptual; it does not replace the documentation for a particular model or implementation.
| Measure | What it compares, in general terms | What to check |
|---|---|---|
| Cosine similarity | The relative orientation of the vectors. | Whether the model and implementation are configured for this measure, and how its results are ranked. |
| Inner product | The sum of products between components; magnitude can have an effect. | Whether magnitude behaves as expected for the vectors being generated. |
| Euclidean distance | The geometric separation between vectors. | Whether the distance and its ranking match the system’s goal and configuration. |
Exact Search and Approximate Search
In an exact search, the system compares the query with every vector in the set being considered and returns the best matches according to the metric. This provides a clear reference for evaluating quality, but the amount of work can grow with the size of the collection.
Approximate nearest neighbor search, usually shortened to ANN, uses index structures to explore part of the space instead of exhaustively comparing every vector. This can reduce query time, but it does not guarantee that the system will always find exactly the same neighbors as an exhaustive search. The degree of approximation and the cost depend on the index, its parameters, and the workload.
HNSW is one example of an approximate-search index based on hierarchical graphs. It is not a requirement for every vector database, nor is it the only way to build an index. Documentation for specific systems may describe different indexes and trade-offs; the name of an index alone is not enough to predict performance in a real application.
A useful comparison is to run representative queries against exact search and an approximate index, then consider retrieval quality and latency together. Measuring speed alone can conceal missed results; measuring only agreement with exact search does not show whether the system meets users’ needs.
An initial choice: exact or approximate
| Situation | Option to evaluate | Main trade-off |
|---|---|---|
| A small collection or a need for a quality baseline | Exact search | Compares all candidates under consideration; the cost may grow with the collection. |
| A large collection or demanding latency requirements | An ANN index, such as HNSW if available and appropriate | It may respond faster, but retrieval can differ from exact search. |
| Requirements are not yet known | Measure both options with representative data and queries | Define a quality measure and latency target before deciding. |
Filters and Combined Search
Metadata can describe aspects of a record that are not necessarily encoded in its vector. A filter might restrict a search to a category, language, date, or available items. If the implementation and application support the combination, similarity search can then run over a more relevant subset.
Filters can also change how a query behaves in practice. For example, if a filter leaves very few candidates or selects a particularly narrow part of an index, the system may need a different retrieval strategy. Do not assume that every product applies filters at the same stage, provides the same guarantees, or has the same effect on its index.
An application can combine vector search with lexical search, which finds word matches, and with explicit rules. This can help when both conceptual similarity and exact terms, identifiers, or constraints matter. Hybrid search should not be assumed to be a universal built-in feature: it may require specific product capabilities or coordination in the application itself.
Three Applied Examples
The following cases describe possible uses, not guaranteed outcomes. In each one, search retrieves candidates based on representations; interpretation and validation depend on the system and its context.
These examples illustrate how a vector database can form one part of a larger workflow. In all cases, the quality of the representation and the way results are checked matter as much as the act of retrieving nearby vectors.
What a Vector Database Does Not Do—and Common Confusions
A vector database does not necessarily create embeddings. That is usually the responsibility of an embedding model or another component. Nor does the database, simply by returning nearby vectors, guarantee that the results are relevant to a user. Relevance depends on the representation, the query, the context, and the application’s evaluation criteria.
A vector database is also not, by itself, a generative system that writes an answer. In a retrieval-augmented generation (RAG) workflow, vector search may help find material for another component to use, but retrieval and answer generation are distinct tasks. A RAG system may include a vector database, but the terms do not mean the same thing.
An embedding is a representation, not the database that stores it. A vector store is a broad term for a system or component used to keep vectors; products differ in their surrounding database features and capabilities. A semantic search system describes a search experience or approach, which may use vectors alongside other techniques. A relational database is a different database model, but some general-purpose databases offer vector support. A specialized vector database is therefore not automatically required.
Limitations and How to Decide What You Need
Vector search is only as useful as the representations and configuration behind it. If a model does not capture the distinctions important to the task, nearby vectors may not correspond to useful results. A metric that does not match the model’s assumptions or the application’s objective can also produce unhelpful rankings. Similarity is a retrieval signal, not a guarantee of relevance, truth, equivalence, or usefulness.
Approximate search introduces a trade-off between latency and retrieval. Index choice, configuration, collection size, and query patterns can affect that trade-off. Filters may alter the candidate set and the practical behavior of a search. Updates, memory costs, permissions, privacy requirements, and the way original data is stored also need to be considered for the full application—not just for the nearest-neighbor operation.
There is no universal collection size or performance threshold that determines when a specialized vector database is necessary. A general-purpose database with vector support may be sufficient, depending on the workload and required capabilities. A specialized system may be worth evaluating when its specific indexing, filtering, operational, or scale features fit the application’s needs. Confirm actual product behavior in its documentation rather than assuming that one product’s features apply to all others.
Evaluate with real or representative queries, data, and constraints. Compare exact and approximate retrieval where possible; define what a useful result means for the task; and measure both quality and latency. Include the effects of filters, updates, and any lexical search or application rules that will be used. A benchmark that does not resemble the expected workload may not predict production behavior.
A practical evaluation checklist
- 01Check which model and configuration generate the stored and query vectors, and whether the representations are compatible.
- 02Confirm that the selected metric matches the model setup and the application’s search goal.
- 03Compare exact and approximate search on representative queries; assess retrieval quality alongside latency.
- 04Test the filters, update patterns, and any lexical or rule-based search that the application will actually use.
- 05Review memory and operational costs, as well as permissions and privacy requirements for the data.
- 06Verify each claimed feature in the documentation for the specific product and version being considered.
Related Concepts
An embedding is the numerical representation that a model or another component generates from data; it is the input to vector search, not a synonym for the database. Semantic search is a search approach that aims to retrieve results related in meaning, and it may use vector search together with other signals.
RAG, or retrieval-augmented generation, is a broader workflow in which a retrieval step can supply material to a generative component. Chunking is the process of dividing content into smaller pieces, such as passages that can be embedded and retrieved separately. Quantization is a technique that can change how vector representations are stored or handled; whether and how it is used depends on the implementation. Inference is the process of running a model, including when generating embeddings. These concepts connect to vector databases, but none should be treated as interchangeable with them.