Building a RAG Pipeline with Weaviate

Retrieval-Augmented Generation, or RAG, has become the secret sauce for building AI applications that don't just hallucinate answers but actually ground their responses in real data. The concept is elegantly simple: instead of relying solely on the knowledge embedded in a language model during training, you retrieve relevant information from your own data sources and feed it to the model as context. This transforms generic chatbots into domain-specific assistants that can answer questions about your documentation, your contracts, or your knowledge base.
When I first built a RAG system for a legal contracts platform, I quickly learned that the vector database you choose makes all the difference. After experimenting with several options, I settled on Weaviate, and it's become my go-to choice for production RAG applications. Weaviate isn't just a vector database - it's a complete ecosystem for building AI-powered applications with built-in vectorization, hybrid search capabilities, and a GraphQL API that makes querying feel natural.
The best way to get started with Weaviate is to spin up a local instance using Docker. This gives you a fully functional vector database in minutes, perfect for development and experimentation. Weaviate provides official Docker images that include everything you need, including optional modules for automatic vectorization.
Running Weaviate with Docker
# Run Weaviate with Docker Compose
# Create a docker-compose.yml file:
version: '3.8'
services:
weaviate:
image: semitechnologies/weaviate:1.24.0
ports:
- "8080:8080"
- "50051:50051"
environment:
QUERY_DEFAULTS_LIMIT: 25
AUTHENTICATION_ANONYMOUS_ACCESS_ENABLED: 'true'
PERSISTENCE_DATA_PATH: '/var/lib/weaviate'
DEFAULT_VECTORIZER_MODULE: 'none'
ENABLE_MODULES: 'text2vec-openai,text2vec-cohere,text2vec-huggingface'
CLUSTER_HOSTNAME: 'node1'
volumes:
- weaviate_data:/var/lib/weaviate
volumes:
weaviate_data:
# Then run:
docker-compose up -d
If you want to use Weaviate's built-in vectorization with OpenAI, you'll need to configure it with your API key. This allows Weaviate to automatically generate embeddings for your text without you having to call the OpenAI API directly.
Weaviate with OpenAI Vectorization
# docker-compose.yml with OpenAI module
version: '3.8'
services:
weaviate:
image: semitechnologies/weaviate:1.24.0
ports:
- "8080:8080"
- "50051:50051"
environment:
QUERY_DEFAULTS_LIMIT: 25
AUTHENTICATION_ANONYMOUS_ACCESS_ENABLED: 'true'
PERSISTENCE_DATA_PATH: '/var/lib/weaviate'
DEFAULT_VECTORIZER_MODULE: 'text2vec-openai'
ENABLE_MODULES: 'text2vec-openai'
CLUSTER_HOSTNAME: 'node1'
OPENAI_API_KEY: 'your-openai-api-key-here'
volumes:
- weaviate_data:/var/lib/weaviate
volumes:
weaviate_data:
Once your Docker container is running, you can verify it's working by visiting http://localhost:8080/v1/meta in your browser. You should see Weaviate's metadata response confirming it's running. Now let's connect to it from Python and start working with data.
Connecting to Weaviate
import weaviate
from weaviate.classes.config import Configure, Property, DataType
# Connect to local Weaviate instance
client = weaviate.connect_to_local(
host="localhost",
port=8080,
grpc_port=50051
)
# Verify connection
print(client.is_ready())
Before you can insert data, you need to create a collection (think of it as a table in a traditional database). A collection defines the structure of your data - what properties it has, and how vectors should be generated. For a simple document collection, you'll want properties like content, title, and metadata fields.
Creating a Document Collection
from weaviate.classes.config import Configure, Property, DataType
# Create a collection for storing documents
client.collections.create(
name="Document",
vectorizer_config=Configure.Vectorizer.text2vec_openai(
model="ada",
model_version="002",
type_="text"
),
properties=[
Property(name="content", data_type=DataType.TEXT),
Property(name="title", data_type=DataType.TEXT),
Property(name="document_type", data_type=DataType.TEXT),
Property(name="source_url", data_type=DataType.TEXT),
Property(name="chunk_index", data_type=DataType.INT),
]
)
print("Collection created successfully!")
If you're not using OpenAI's vectorizer, you can create a collection without automatic vectorization and provide your own vectors. This is useful when you want to use a different embedding model or generate vectors elsewhere.
Collection Without Auto-Vectorization
# Create collection without automatic vectorization
client.collections.create(
name="Document",
vectorizer_config=None, # No automatic vectorization
properties=[
Property(name="content", data_type=DataType.TEXT),
Property(name="title", data_type=DataType.TEXT),
]
)
Now that you have a collection, inserting data is straightforward. You get a reference to the collection and call the insert method with your data. If you've configured automatic vectorization, Weaviate will generate the vector for you automatically based on the text properties you specify.
Inserting Documents
# Get the collection
documents_collection = client.collections.get("Document")
# Insert a single document
documents_collection.data.insert(
properties={
"content": "Laravel is a PHP web framework that makes it easy to build modern web applications. It provides elegant syntax and powerful features for routing, database management, and more.",
"title": "Introduction to Laravel",
"document_type": "tutorial",
"source_url": "https://example.com/laravel-intro",
"chunk_index": 0
}
)
print("Document inserted!")
You can also insert multiple documents at once for better performance. This is especially useful when you're loading a large dataset.
Batch Inserting Documents
# Prepare multiple documents
documents = [
{
"content": "Multi-tenancy allows you to serve multiple customers from a single application instance. Each tenant's data is isolated from others.",
"title": "Multi-Tenancy Basics",
"document_type": "article",
"source_url": "https://example.com/multi-tenancy",
"chunk_index": 0
},
{
"content": "Vector databases store data as high-dimensional vectors, enabling semantic search and similarity matching for AI applications.",
"title": "Understanding Vector Databases",
"document_type": "article",
"source_url": "https://example.com/vector-db",
"chunk_index": 0
},
{
"content": "RAG combines retrieval of relevant documents with language model generation to create accurate, context-aware responses.",
"title": "RAG Architecture",
"document_type": "tutorial",
"source_url": "https://example.com/rag",
"chunk_index": 0
}
]
# Insert all documents
documents_collection.data.insert_many(documents)
print(f"Inserted {len(documents)} documents!")
Once you have data in Weaviate, querying it is where the magic happens. The most common query type is semantic search using near_text, which finds documents similar in meaning to your query text. Weaviate automatically vectorizes your query and finds the most similar vectors in the database.
Querying with Semantic Search
# Query for documents similar to the query text
response = documents_collection.query.near_text(
query="How do I implement multi-tenancy?",
limit=3,
return_metadata=["distance", "certainty"]
)
# Process the results
for obj in response.objects:
print(f"Title: {obj.properties['title']}")
print(f"Content: {obj.properties['content'][:100]}...")
print(f"Certainty: {obj.metadata.certainty:.2f}")
print("---")
The certainty score tells you how similar the result is to your query, with values closer to 1.0 indicating higher similarity. The distance metric works the opposite way - lower values mean higher similarity. You can use these metrics to filter out results that aren't relevant enough.
Filtering by Certainty
# Only return results with high certainty
response = documents_collection.query.near_text(
query="Laravel framework features",
limit=5,
return_metadata=["certainty"],
return_properties=["title", "content"]
)
# Filter results by certainty threshold
relevant_docs = [
obj for obj in response.objects
if obj.metadata.certainty > 0.7
]
print(f"Found {len(relevant_docs)} highly relevant documents")
You can also combine semantic search with traditional filters. This is incredibly powerful - you can search for documents that are semantically similar to your query but also match specific criteria like document type or date ranges.
Filtered Semantic Search
from weaviate.classes.query import Filter
# Search for documents similar to query, but only of type "tutorial"
response = documents_collection.query.near_text(
query="building web applications",
limit=5,
filters=Filter.by_property("document_type").equal("tutorial"),
return_metadata=["certainty"]
)
for obj in response.objects:
print(f"{obj.properties['title']} - {obj.metadata.certainty:.2f}")
Weaviate also supports hybrid search, which combines vector similarity with keyword matching. This gives you the best of both worlds - semantic understanding from vectors and exact term matching from keywords. The alpha parameter controls the balance between vector and keyword search.
Hybrid Search
# Hybrid search: 70% vector similarity, 30% keyword matching
response = documents_collection.query.hybrid(
query="Laravel multi-tenancy implementation",
alpha=0.7, # 0.7 = 70% vector, 30% keyword
limit=5,
return_metadata=["score", "certainty"]
)
for obj in response.objects:
print(f"{obj.properties['title']}")
print(f"Hybrid score: {obj.metadata.score:.2f}")
Sometimes you want to retrieve all documents or query by specific property values. Weaviate supports traditional database-style queries alongside vector search.
Querying by Properties
# Get all documents of a specific type
response = documents_collection.query.fetch_objects(
limit=10,
filters=Filter.by_property("document_type").equal("article")
)
# Get a specific document by ID (if you know it)
document = documents_collection.query.fetch_object_by_id(
id="your-document-id-here"
)
print(document.properties)
The real challenge in RAG isn't the vector search itself - it's chunking your documents intelligently. You can't just throw entire documents into the vector database and expect good results. Large documents need to be split into smaller chunks that are semantically meaningful. Too small, and you lose context. Too large, and your retrieved chunks become noisy and irrelevant.
I've found that the sweet spot is usually between 200 and 500 tokens per chunk, with some overlap between chunks to preserve context. The overlap is crucial - when you split a document at sentence boundaries, the last few sentences of one chunk should appear at the beginning of the next chunk. This ensures that concepts that span chunk boundaries aren't lost.
Intelligent Document Chunking
from typing import List
import tiktoken
def chunk_document(text: str, chunk_size: int = 500, overlap: int = 50) -> List[str]:
"""
Split a document into overlapping chunks while preserving sentence boundaries.
"""
encoding = tiktoken.encoding_for_model("gpt-3.5-turbo")
sentences = text.split('. ')
chunks = []
current_chunk = []
current_size = 0
for sentence in sentences:
sentence_tokens = len(encoding.encode(sentence))
if current_size + sentence_tokens > chunk_size and current_chunk:
chunks.append('. '.join(current_chunk) + '.')
# Start new chunk with overlap
overlap_sentences = current_chunk[-overlap:] if len(current_chunk) > overlap else current_chunk
current_chunk = overlap_sentences + [sentence]
current_size = sum(len(encoding.encode(s)) for s in current_chunk)
else:
current_chunk.append(sentence)
current_size += sentence_tokens
if current_chunk:
chunks.append('. '.join(current_chunk) + '.')
return chunks
# Use it to chunk a document before inserting
long_document = "Your long document text here..."
chunks = chunk_document(long_document)
# Insert each chunk
for i, chunk in enumerate(chunks):
documents_collection.data.insert(
properties={
"content": chunk,
"title": "My Document",
"document_type": "article",
"source_url": "https://example.com/doc",
"chunk_index": i
}
)
Once you have your retrieved chunks, you need to construct a prompt that includes them as context. The prompt engineering here is crucial - you need to clearly separate the context from the user's question and instruct the model to only answer based on the provided context.
Building a RAG Pipeline
def retrieve_and_generate(query: str, limit: int = 5):
# Step 1: Retrieve relevant chunks
response = documents_collection.query.near_text(
query=query,
limit=limit,
return_metadata=["certainty"]
)
# Step 2: Filter by certainty threshold
relevant_chunks = [
obj for obj in response.objects
if obj.metadata.certainty > 0.7
]
# Step 3: Build context from chunks
context = "\n\n".join([
f"Document: {chunk.properties['title']}\n{chunk.properties['content']}"
for chunk in relevant_chunks
])
# Step 4: Construct RAG prompt
prompt = f"""You are a helpful assistant. Answer the user's question based only on the following context. If the context doesn't contain enough information to answer the question, say so.
Context:
{context}
Question: {query}
Answer:"""
# Step 5: Send to LLM (OpenAI, Anthropic, etc.)
# response = openai.ChatCompletion.create(...)
return prompt, relevant_chunks
# Use it
prompt, chunks = retrieve_and_generate("How do I implement multi-tenancy in Laravel?")
print(prompt)
Building a production RAG system requires thinking about more than just retrieval and generation. You need to handle cases where no relevant chunks are found, implement reranking to improve result quality, and add citation tracking so users know where answers came from. Weaviate makes all of this manageable with its flexible query API and metadata support.
The real power of RAG emerges when you combine it with other techniques. I've built systems that use RAG for initial retrieval, then apply reranking models to further refine results. Others use RAG in combination with graph databases to traverse relationships between documents. Weaviate's GraphQL API makes these advanced patterns possible.
What I've learned from building RAG pipelines is that the vector database is just one piece of the puzzle. The quality of your chunking strategy, the relevance of your retrieved context, and the way you construct your prompts all contribute to the final user experience. Weaviate provides an excellent foundation with its Docker setup, intuitive Python API, and powerful query capabilities, but the real art is in how you orchestrate the entire pipeline to create AI applications that are both intelligent and trustworthy.