Imagine you've got ten thousand support articles, product documents, and internal policies. A customer asks a question. You need to find the three or four documents most relevant to that specific question — not by keyword, but by meaning — in under a second.
Traditional databases can't really do this. A keyword search for "my payment failed" won't surface a document titled "troubleshooting transaction errors" unless those exact words appear somewhere in it. The meaning is the same; the words are different. And customers, in our experience, almost never phrase things the way your documentation does.
Vector databases solve this. They store information as mathematical representations of meaning — called embeddings — and can find the most semantically similar content to any query in milliseconds. This is what makes AI agents that answer questions from your own data accurate rather than generic.
This guide explains what a vector database is without the math-heavy jargon: what an embedding actually is, how nearest-neighbour search works, why retrieval-augmented generation depends on it, how the main options (Pinecone, Chroma, Weaviate, Qdrant, and pgvector) compare, the concepts you'll run into during a build, and the mistakes that most often make retrieval quality worse than it should be.
What an Embedding Actually Is
Before you can understand vector databases, you need to understand embeddings.
An embedding is a list of numbers — usually several hundred to several thousand — that represents the meaning of a piece of text. Two pieces of text with similar meaning will have embeddings that are numerically close to each other. Two pieces with very different meanings will have embeddings that are far apart.
Here's what a sentence embedding looks like, simplified:
"My payment failed" → [0.23, -0.87, 0.41, 0.12, ...] (1,536 numbers)
"Transaction error" → [0.21, -0.84, 0.39, 0.15, ...] (1,536 numbers)
"Dog breeds" → [-0.92, 0.34, -0.67, 0.88, ...] (1,536 numbers)
The payment and transaction embeddings are numerically similar. The dog breeds one is very different. A vector database finds the most similar embeddings to a query embedding — and that similarity corresponds, roughly, to semantic relevance.
Embeddings are generated by embedding models. OpenAI's text-embedding-3-small (documented in the OpenAI API docs) and Cohere's embed-v3 are the most commonly used in production AI applications. You pass text to the model and it returns the embedding — a list of numbers representing meaning.
How a Vector Database Works
A vector database stores embeddings alongside the original content they represent. When you query it, it:
- Takes your query and converts it to an embedding using the same model
- Compares your query embedding against all stored embeddings
- Returns the most similar ones — the content most relevant to your query
This is called nearest neighbour search. The database finds the stored embeddings nearest (most similar) to the query embedding in mathematical space.
Because embeddings capture meaning rather than exact words, the search finds content that's semantically relevant even when the words don't match. That's fundamentally different from keyword search, and it's why the AI agents you'd actually trust feel like they "get" what you mean.
Why AI Agents Need Vector Databases
An AI agent that answers questions from your data needs to retrieve relevant content at query time and include it in the prompt sent to the language model. That's Retrieval-Augmented Generation (RAG).
Without a vector database, the agent has two bad options:
- Include all your documents in every prompt (too expensive, too slow, blows past context limits)
- Use keyword search (misses semantically relevant content, returns irrelevant results)
With a vector database, the agent retrieves the three to five most semantically relevant documents per query, includes only those in the prompt, and generates an accurate response grounded in actually-relevant content. This is the approach used in virtually every production AI agent that answers from a custom knowledge base.
The honest caveat: vector search isn't magic. It can still pull the wrong chunk if the embedding model doesn't capture intent well, or if your chunking is awkward. We've spent more debugging time on this than on any other part of RAG systems.
Benefits of Vector Databases
Semantic retrieval changes what an AI system can do with your own content. These are the practical advantages for teams building on it.
Search that understands meaning
Users rarely phrase questions the way documentation does. Because embeddings capture meaning, a question about a failed payment finds the article on transaction errors even when no words overlap. That closes the gap between how customers talk and how content is written, and it's the main reason retrieval-backed agents feel genuinely helpful rather than literal-minded.
Grounded answers instead of generic ones
Retrieving the few most relevant passages and passing them to the model lets it answer from your actual policies, products, and procedures. Responses reflect your content rather than the model's general training, which makes them more accurate and easier to verify against a cited source.
Lower cost and latency than stuffing context
Sending every document with every prompt is slow, expensive, and quickly exceeds context limits. A vector database narrows the input to a handful of relevant chunks, so each request carries only what it needs. That keeps per-query costs predictable as the knowledge base grows from hundreds of documents to many thousands.
Fast retrieval at large scale
Approximate nearest-neighbour indexes find similar vectors among millions of entries in milliseconds. Agents can search a large corpus at conversational speed, which matters for chat and voice interfaces where users expect immediate replies. Retrieval is rarely the slow part of a well-built pipeline; the language model usually is.
Filtering and access control alongside relevance
Most vector databases store metadata with each vector, so a search can be limited to a product line, a date range, or the documents a particular user is allowed to see. Relevance and permissions are applied in the same query rather than bolted on afterwards. That is essential for any agent serving users with different access rights.
Content stays up to date without retraining
Adding or changing knowledge means embedding new chunks and updating the index. There's no need to retrain or fine-tune a model whenever a policy changes, which keeps maintenance practical for content that changes weekly.
Vector Database Use Cases
Wherever a system needs to find relevant content by meaning, a vector database is likely involved. These are the most common applications.
Customer support agents
Support teams want an agent that answers from help-centre articles, product documentation, and policy pages. The agent embeds each customer question, retrieves the most relevant articles, and responds using them, citing the source. The result is accurate answers to the long tail of questions customers ask in their own words, and fewer tickets for human agents. When retrieval finds nothing relevant, the agent can say so and hand off instead of guessing.
Internal knowledge assistants
Employees lose time searching wikis, shared drives, and old messages for policies and how-to guides. An internal assistant backed by a vector database retrieves relevant passages across those sources, filtered by department and permissions. Staff get answers in seconds, and new hires can find information without knowing where it lives. Answers that link back to the source document also help keep the underlying content accurate, since errors get noticed.
Semantic search in products
E-commerce, media, and SaaS products use vector search so users find items by description rather than exact keywords, such as "waterproof jacket for hiking" matching products that never use those exact words. Combined with keyword search for exact terms, it improves discovery without replacing existing search entirely. Product descriptions, reviews, and attributes can all be embedded to broaden what a query matches.
Document analysis and research
Legal, compliance, and research teams work with large document collections. Vector retrieval surfaces relevant clauses, sections, or papers for a question, which a model can then summarise or compare. People still review the conclusions, but the time spent finding the right material drops sharply. Metadata filters by date, jurisdiction, or document type keep results relevant to the specific question.
Recommendations and similarity matching
Embeddings aren't limited to questions and answers. Similar articles, related products, or duplicate support tickets can all be found by comparing vectors, giving teams a simple way to suggest related content or cluster near-duplicates. Support teams use the same approach to spot recurring issues by grouping similar tickets.
The Main Vector Database Options
Pinecone
The most widely used managed vector database for production AI applications. Pinecone handles the infrastructure entirely — you don't manage servers, indices, or scaling. It has a generous free tier and scales predictably.
Best for: Production applications where you want managed infrastructure and don't want to deal with operational complexity. The default choice for teams without dedicated DevOps resource.
Limitations: Costs money at scale; you don't control the underlying infrastructure; data is hosted on Pinecone's servers (worth thinking about for data sovereignty requirements).
Chroma
Open-source, lightweight, easy to run locally. The default choice for development and prototyping.
Best for: Development, testing, and small deployments where you want to run everything locally without external services. Not what you reach for at production scale.
Limitations: Requires self-hosting for production; not designed for high-throughput production workloads.
Weaviate
Open-source with a managed cloud option. Strong filtering capabilities — you can combine vector search with metadata filters ("find semantically similar documents that are also from category X and published after date Y"). Good choice when you need hybrid search with complex filtering.
Best for: Applications where metadata filtering alongside semantic search matters. Self-hosting teams who want open-source with production capabilities.
Qdrant
Open-source, high-performance, built for production. Particularly fast on filtered vector search. Good Python and TypeScript client libraries.
Best for: High-throughput applications where filtering matters and you want open-source with production-grade performance. A strong alternative to Pinecone for teams who'd rather self-host.
pgvector
pgvector is a PostgreSQL extension that adds vector search to an existing Postgres database. If you're already running Postgres, adding vector search without standing up a separate service is attractive.
Best for: Teams already on PostgreSQL who want to add semantic search without managing another service. Not optimal for very large vector collections but works well at moderate scale.
Choosing the Right One
| Factor | Recommendation |
|---|---|
| Getting started quickly | Chroma locally, Pinecone for production |
| Already on PostgreSQL | pgvector |
| Need complex metadata filtering | Weaviate or Qdrant |
| Open-source, self-hosted production | Qdrant |
| Managed, no infrastructure management | Pinecone |
| Data sovereignty requirements | Qdrant or pgvector (self-hosted) |
For most teams building their first production RAG application: start with Chroma locally, deploy with Pinecone. It's the path of least friction and least operational risk, and you can revisit the choice later when you actually know your traffic patterns.
Key Concepts You Will Encounter
Chunking: Before storing documents, you split them into smaller pieces (chunks). A 20-page PDF becomes 50 chunks. Each chunk is embedded and stored separately. Retrieval finds the most relevant chunks, not entire documents.
Chunk size: How big each chunk is. Smaller chunks (200–400 tokens) give more precise retrieval but less context per chunk. Larger chunks (600–1000 tokens) provide more context but less precise matching. Most production systems use 400–600 tokens with some overlap between chunks. You'll likely tune this for your corpus.
Similarity metric: How closeness between embeddings is measured. Cosine similarity is the most common — it measures the angle between embedding vectors rather than their distance, which behaves better for text embeddings.
Hybrid search: Combining vector search with keyword search (BM25). Vector search alone misses exact keyword matches (product codes, proper nouns). BM25 alone misses semantic matches. Hybrid search catches both. Most production RAG systems we've built end up using some form of hybrid search once they're tuned.
Re-ranking: After vector search returns the top 20 results, a re-ranking model scores each one for relevance and reorders them before the top 5 are sent to the LLM. Re-ranking significantly improves quality at the cost of additional latency. Worth turning on once basic retrieval is in place.
Common Mistakes With Vector Databases
Most disappointing RAG systems don't fail because of the database brand. They fail on a handful of avoidable decisions around it.
Picking the database before understanding the data
Teams often debate Pinecone versus Qdrant before anyone has looked at what the documents contain. The shape of the content, such as long PDFs, short FAQs, tables, or product codes, drives chunking and search strategy far more than the choice of store does. Spend an hour reading a sample of the corpus first; it usually settles half the architecture questions.
Mixing embedding models
Embeddings from different models aren't comparable. If you change embedding models, every stored document has to be re-embedded. Querying an index built with one model using another produces results that look plausible and are quietly wrong. Record the model name and version alongside each index so a mismatch is caught before it reaches users.
Relying on vector search alone
Pure semantic search struggles with exact identifiers like SKUs, error codes, and names. Adding keyword search alongside it (hybrid search) usually fixes a class of misses that no amount of embedding tuning will. Users searching for a specific invoice number or error code expect an exact match, not a semantically similar one.
Ignoring metadata and permissions
If every chunk lacks source, date, and access metadata, you can't filter out outdated content or stop the agent from retrieving documents a given user shouldn't see. Retrofitting this later means re-ingesting everything. Permissions matter most: an agent that retrieves a document is effectively showing it to the user.
Never measuring retrieval quality
Teams test the final answers and never check whether the right chunks were retrieved. Build a small set of real questions with known correct source documents, and track how often retrieval finds them, before tuning anything else. Without that baseline, every change to chunking or models is a guess.
Vector Database Best Practices
These practices cover most of the difference between a retrieval pipeline that works in a demo and one that holds up in production.
- Build a retrieval test set first. Collect real user questions and note which documents should answer each. Measure how often the right chunks appear in the top results, and rerun that check after every change to chunking, models, or search settings.
- Chunk by structure, not just by length. Split on headings, sections, and paragraphs where possible, keep tables intact, and add modest overlap between chunks so an answer split across a boundary isn't lost. Then tune size against your test set rather than picking a number from a blog post.
- Store rich metadata with every chunk. Include source document, section, date, version, and access permissions. Metadata lets you filter out outdated content, cite sources in answers, and enforce who can see what at query time.
- Pin one embedding model per index. Use the same model for ingestion and queries, record which model built each index, and plan a full re-embedding when you upgrade rather than mixing vectors from different models.
- Add hybrid search early. Combine vector search with BM25 keyword search so exact identifiers like product codes, error messages, and names are found reliably alongside semantic matches.
- Re-rank once the basics work. Retrieve a wider candidate set, then let a re-ranking model reorder it before passing the best few chunks to the LLM. Check the latency cost against your response-time budget.
- Automate ingestion and updates. Build a pipeline that re-indexes documents when they change and removes them when they're deleted, so the agent never answers from content that no longer exists.
- Start simple, then revisit the store. Prototype with Chroma or pgvector, move to a managed or self-hosted production option once you know your volume and filtering needs, and keep your retrieval code behind an interface so switching later is cheap.
What This Means for Your AI Agent Project
If you're commissioning a RAG-powered AI agent — one that answers questions from your documents, knowledge base, or data — your development team will need to:
- Choose an embedding model
- Choose a vector database
- Design a chunking strategy for your content
- Build an ingestion pipeline to load and index your content
- Build a retrieval layer that queries the vector database at runtime
These decisions significantly affect the quality of the agent's responses. A well-designed retrieval pipeline produces accurate, relevant answers. A poorly designed one produces generic or wrong ones — and the rest of the system can't really fix that downstream.
When evaluating vendors, ask specifically about their approach to each of these decisions — not just which tools they use, but why those over the alternatives. A vendor whose answer is "we just use the defaults" is one to be cautious about.
Related guides
- How to build a RAG chatbot step by step
- LangChain vs LlamaIndex: which framework to use
- What is an LLM? A plain-English guide
- LLM integration guide for business applications
- LLM integration services
Talk to us about your project — RAG system design is one of our core capabilities and we're happy to walk you through the architecture decisions before you commit to a build.
Frequently Asked Questions
What is a vector database in simple terms?
A vector database stores information as lists of numbers (called embeddings) that represent meaning rather than exact words. When you search it, it finds content that is semantically similar to your query — meaning it understands intent, not just keywords. This makes it the core storage layer for AI agents that need to retrieve relevant information from large document collections.
Do I need a vector database to build an AI chatbot?
Not for every chatbot — but if your chatbot needs to answer questions from your own documents, knowledge base, or internal data, then yes, a vector database is almost certainly part of the architecture. Without one, you're either stuffing all your content into every prompt (expensive and slow) or relying on keyword search (inaccurate). Vector databases make retrieval-augmented generation (RAG) practical at scale.
How much does a vector database cost?
Pinecone's free tier covers up to 1 million vectors, which is enough for most early-stage projects. Paid plans start around $70/month and scale with storage and query volume. Open-source options like Qdrant and Chroma are free to use but require you to manage your own hosting, which has its own infrastructure costs. For most small-to-medium production applications, managed vector database costs are a minor line item compared to LLM API costs.
What is the difference between Pinecone and Chroma?
Pinecone is a fully managed cloud service — you don't run any infrastructure, it scales automatically, and you pay per use. Chroma is open-source and runs locally on your own machine or server. Chroma is the standard choice for development and testing; Pinecone is the standard choice for production deployment. Most teams use Chroma during development and switch to Pinecone (or Qdrant for self-hosted production) when they go live.
How many documents can a vector database handle?
Modern vector databases handle tens of millions to billions of vectors in production. Pinecone's paid plans scale to hundreds of millions of vectors. Qdrant and Weaviate are similarly capable. For context, a 1,000-page document corpus chunked at 500 tokens might produce around 5,000 to 10,000 vectors — well within the free tier of most managed services. Scale only becomes a concern for very large knowledge bases or high-volume ingestion pipelines.
What embedding model should I use with a vector database?
OpenAI's text-embedding-3-small is the most common choice for teams already using OpenAI's LLMs — it's fast, affordable, and produces 1,536-dimensional embeddings that work well across most use cases. Cohere's embed-v3 is a strong alternative with good multilingual support. The key constraint: you must use the same embedding model at ingestion time and at query time. Switching models later requires re-embedding your entire corpus.
Can a vector database replace a traditional relational database?
No — they serve different purposes. Vector databases are optimised for semantic similarity search over unstructured content like text, images, and documents. Relational databases are optimised for structured data, exact lookups, joins, and transactional integrity. Most production AI systems use both: a relational database for structured business data (users, orders, products) and a vector database for the semantic retrieval layer that powers the AI agent's knowledge access.
Conclusion
AI agents that answer from your own content need a way to find the right few passages among thousands, by meaning rather than exact wording. Vector databases do that by storing embeddings, numerical representations of meaning, and running fast nearest-neighbour search against them. That retrieval step is the backbone of retrieval-augmented generation, and it largely determines whether an agent's answers are grounded or generic.
The database choice matters less than most teams expect. Chroma is fine for prototyping, Pinecone removes infrastructure work, Weaviate and Qdrant suit self-hosted production with heavy filtering, and pgvector is the pragmatic option if you already run Postgres. What matters more is the work around it: sensible chunking, one consistent embedding model, hybrid search for exact terms, metadata for filtering and permissions, and re-ranking once the basics work. Vector search still retrieves the wrong chunk sometimes, so measure retrieval quality directly rather than judging only the final answers.
A good next step is to collect twenty real questions your users ask, note which documents should answer each, and use that as a test set for any prototype. If you'd like help designing the retrieval pipeline for a production agent, our LLM integration team builds these systems.
