Building RAG Systems That Actually Work for Small Teams

Photo via Unsplash

Every growing company has the same knowledge problem: information is scattered across Notion, Confluence, Google Drive, Slack threads, and the heads of people who’ve been around the longest. New employees spend weeks getting up to speed. Customer-facing teams can’t find answers fast enough. And everyone’s frustrated.

RAG — Retrieval-Augmented Generation — is the right solution. An LLM that can search your company’s knowledge base and synthesize answers from it is genuinely useful. But most RAG implementations we encounter are 60% as good as they should be, and the 40% gap is almost entirely in retrieval quality.

What RAG Is and Why the Retrieval Part Is Hard

The concept is simple: when a user asks a question, search a vector database for relevant documents, insert those documents into the LLM’s context, and let the LLM generate an answer grounded in your actual content.

The implementation challenge is that “search a vector database” hides enormous complexity. The naive implementation — chunk documents, embed them, find the most similar embeddings to the query — works surprisingly well for simple questions. It falls apart badly for:

  • Questions that require combining information from multiple documents
  • Questions where the user uses different vocabulary than the documents
  • Questions that require understanding the structure of information, not just its content
  • Ambiguous queries where the right answer depends on context
60%
Naive RAG accuracy
on complex queries
85%+
Hybrid RAG accuracy
with reranking
3×
Support ticket reduction
typical after deployment
2–4wk
Time to working prototype
with right architecture

The Architecture That Works

The difference between naive RAG and a system that actually holds up is a set of deliberate engineering choices in three areas: ingestion, retrieval, and generation.

RAG pipeline architecture — ingestion, retrieval, and generation stages
Production RAG requires distinct attention to ingestion quality, retrieval strategy, and generation guardrails

Ingestion

How you prepare your documents for retrieval matters enormously. Most teams chunk by character count and stop there. Better approaches:

Semantic chunking — split at natural semantic boundaries (paragraphs, sections) rather than fixed character counts. A 500-character split that cuts mid-sentence loses context on both sides.

Hierarchical indexing — index documents at multiple granularities: individual sentences for precision, paragraphs for context, full documents for broad questions. Retrieve at the right level depending on the query.

Metadata enrichment — store document title, author, date, source, and document type alongside each chunk. This enables filtered retrieval (“find policies updated in the last 6 months”) that pure semantic search can’t do.

Retrieval: Where Most Systems Underperform

01
Hybrid Search (BM25 + Vector)
Pure vector search misses exact matches. BM25 (traditional keyword search) misses semantic relationships. Combining them with a weighted blend consistently outperforms either approach alone — typically 15–25% improvement in recall. Most production vector databases (Pinecone, Weaviate, Qdrant) support hybrid search natively.
02
Query Expansion
Before searching, use the LLM to expand the user's query: generate synonyms, rephrase it multiple ways, and break compound questions into sub-questions. Search with each expanded version and merge results. This dramatically improves recall for questions where user vocabulary doesn't match document vocabulary.
03
Reranking
After initial retrieval, pass the top-20 results through a cross-encoder reranker (Cohere Rerank or an open-source model) that evaluates each document's actual relevance to the query. Reranking typically improves precision by 20–30% at the cost of a small latency increase.
04
Contextual Compression
Retrieved documents often contain relevant sections plus a lot of noise. Use an LLM to extract only the relevant portions before inserting into context. This reduces token usage and improves answer quality by keeping the LLM focused.

Vector Database Selection

Database Best For Self-Hosted? Hybrid Search
Pinecone Managed, zero ops No ✅ Yes
Weaviate Complex schemas, GraphQL API Yes ✅ Yes
Qdrant High performance, Rust-based Yes ✅ Yes
pgvector Already using PostgreSQL Yes 🟡 Limited
Chroma Local development, prototyping Yes ❌ No
OpenSearch Already in AWS ecosystem Yes ✅ Yes

For teams starting fresh, Qdrant (self-hosted) or Pinecone (managed) are our most common recommendations

Measuring RAG Quality

Without measurement, you’re flying blind. Define these metrics before you launch:

Retrieval recall — of the documents that should answer this question, what percentage are you retrieving? Build a golden dataset of question-document pairs to measure this.

Answer faithfulness — does the generated answer accurately reflect the retrieved documents? LLM-based evaluation works well here: ask a judge model “does this answer contradict or hallucinate beyond the provided context?”

Answer relevance — does the answer actually address the user’s question? Again, LLM-based evaluation is effective.

User satisfaction proxy — track thumbs up/down feedback, follow-up questions (indicating the first answer didn’t resolve the issue), and session abandonment.

The Minimum Viable RAG Stack for Small Teams

You don’t need infrastructure complexity to get started. Here’s the stack we recommend for a team of 5–20 engineers building their first RAG system:

01
Document processing: LangChain + custom chunking
LangChain's document loaders handle PDF, Notion, Confluence, Google Drive, and Slack. Add custom semantic chunking on top. Total: ~200 lines of Python.
02
Vector store: Qdrant (Docker for dev, cloud for prod)
Free, open-source, excellent performance, and supports hybrid search. The Docker image is one command. Cloud deployment at production scale is simple and affordable.
03
Embeddings: OpenAI text-embedding-3-small
Currently the best price-to-quality ratio for English content. At $0.02 per million tokens, even a large document corpus is cheap to embed.
04
Reranking: Cohere Rerank API
Simple REST API, pay-per-use, significant quality improvement. Add it as a post-processing step after your initial vector search.
05
Generation: GPT-4o or Claude 3.5 Sonnet
Both perform well for RAG. Claude's longer context window is useful for multi-document synthesis. Run evals on your specific use case to choose.

What Production Actually Requires

Beyond the core pipeline, production RAG systems need:

  • Access control — users should only be able to retrieve documents they’re authorized to see. Implement at the retrieval layer with metadata filtering.
  • Document freshness — documents change. Implement incremental indexing triggered by document updates, not just a nightly full re-index.
  • Citation and sourcing — users need to know where answers come from. Store source metadata and surface it in every response.
  • Feedback loops — capture what questions go unanswered or get negative feedback. These are your highest-priority documents to add.

Building a RAG system that works well for your specific use case takes about 4–8 weeks of focused engineering effort. The technology is ready; the work is in tuning retrieval quality, setting up evaluation, and handling the edge cases that only real usage surfaces. Start with a narrow use case, measure relentlessly, and expand from there.

Free Strategy Session

Want Expert Guidance on Your Project?

Book a free 1-hour session. We'll apply these principles directly to your architecture, codebase, or team challenge.

Free 1-hour strategy session

Walk away with a clear, prioritised action plan — no pitch.

Book free session