Photo via Unsplash
Every growing company has the same knowledge problem: information is scattered across Notion, Confluence, Google Drive, Slack threads, and the heads of people who’ve been around the longest. New employees spend weeks getting up to speed. Customer-facing teams can’t find answers fast enough. And everyone’s frustrated.
RAG — Retrieval-Augmented Generation — is the right solution. An LLM that can search your company’s knowledge base and synthesize answers from it is genuinely useful. But most RAG implementations we encounter are 60% as good as they should be, and the 40% gap is almost entirely in retrieval quality.
What RAG Is and Why the Retrieval Part Is Hard
The concept is simple: when a user asks a question, search a vector database for relevant documents, insert those documents into the LLM’s context, and let the LLM generate an answer grounded in your actual content.
The implementation challenge is that “search a vector database” hides enormous complexity. The naive implementation — chunk documents, embed them, find the most similar embeddings to the query — works surprisingly well for simple questions. It falls apart badly for:
- Questions that require combining information from multiple documents
- Questions where the user uses different vocabulary than the documents
- Questions that require understanding the structure of information, not just its content
- Ambiguous queries where the right answer depends on context
The Architecture That Works
The difference between naive RAG and a system that actually holds up is a set of deliberate engineering choices in three areas: ingestion, retrieval, and generation.
Ingestion
How you prepare your documents for retrieval matters enormously. Most teams chunk by character count and stop there. Better approaches:
Semantic chunking — split at natural semantic boundaries (paragraphs, sections) rather than fixed character counts. A 500-character split that cuts mid-sentence loses context on both sides.
Hierarchical indexing — index documents at multiple granularities: individual sentences for precision, paragraphs for context, full documents for broad questions. Retrieve at the right level depending on the query.
Metadata enrichment — store document title, author, date, source, and document type alongside each chunk. This enables filtered retrieval (“find policies updated in the last 6 months”) that pure semantic search can’t do.
Retrieval: Where Most Systems Underperform
Vector Database Selection
| Database | Best For | Self-Hosted? | Hybrid Search |
|---|---|---|---|
| Pinecone | Managed, zero ops | No | ✅ Yes |
| Weaviate | Complex schemas, GraphQL API | Yes | ✅ Yes |
| Qdrant | High performance, Rust-based | Yes | ✅ Yes |
| pgvector | Already using PostgreSQL | Yes | 🟡 Limited |
| Chroma | Local development, prototyping | Yes | ❌ No |
| OpenSearch | Already in AWS ecosystem | Yes | ✅ Yes |
For teams starting fresh, Qdrant (self-hosted) or Pinecone (managed) are our most common recommendations
Measuring RAG Quality
Without measurement, you’re flying blind. Define these metrics before you launch:
Retrieval recall — of the documents that should answer this question, what percentage are you retrieving? Build a golden dataset of question-document pairs to measure this.
Answer faithfulness — does the generated answer accurately reflect the retrieved documents? LLM-based evaluation works well here: ask a judge model “does this answer contradict or hallucinate beyond the provided context?”
Answer relevance — does the answer actually address the user’s question? Again, LLM-based evaluation is effective.
User satisfaction proxy — track thumbs up/down feedback, follow-up questions (indicating the first answer didn’t resolve the issue), and session abandonment.
The Minimum Viable RAG Stack for Small Teams
You don’t need infrastructure complexity to get started. Here’s the stack we recommend for a team of 5–20 engineers building their first RAG system:
What Production Actually Requires
Beyond the core pipeline, production RAG systems need:
- Access control — users should only be able to retrieve documents they’re authorized to see. Implement at the retrieval layer with metadata filtering.
- Document freshness — documents change. Implement incremental indexing triggered by document updates, not just a nightly full re-index.
- Citation and sourcing — users need to know where answers come from. Store source metadata and surface it in every response.
- Feedback loops — capture what questions go unanswered or get negative feedback. These are your highest-priority documents to add.
Building a RAG system that works well for your specific use case takes about 4–8 weeks of focused engineering effort. The technology is ready; the work is in tuning retrieval quality, setting up evaluation, and handling the edge cases that only real usage surfaces. Start with a narrow use case, measure relentlessly, and expand from there.