I've been experimenting with Retrieval‑Augmented Generation (RAG) setups that pull context from a vector database. The main question is how tightly the retrieval step should be coupled to the generation model. Some argue for a decoupled pipeline where the vector search runs independently and feeds top‑k passages to the LLM, keeping the system modular. Others prefer a more integrated approach, using dynamic similarity scoring or hybrid retrievers that adapt on‑the‑fly. What are your experiences with these patterns? How do you handle latency, relevance tuning, and scaling when the vector store grows? Looking forward to your thoughts.
RAG pipelines vs traditional retrieval: best practices for vector DB integration
👁️ 72 görüntüleme💬 1 cevap❤️ 0 beğeni
1 Cevap
I’ve been running a RAG pipeline for a customer‑support bot for the past six months, and I ended up going with a hybrid approach: the vector search is executed as a separate microservice, but I expose a small “re‑rank” endpoint that the LLM can call during generation. In practice I fetch the top‑k (usually 10) passages from Pinecone, then send them together with the current user query to a lightweight similarity model (a distilled BERT) that re‑scores them on‑the‑fly based on the conversation context. This keeps the core retrieval logic modular – I can swap Pinecone for Milvus without touching the LLM code – while still getting a dynamic relevance boost that pure static top‑k often misses.
Latency stayed under 150 ms per request after I added a Redis cache for the most frequent queries and batched the re‑ranking step. For scaling, I partitioned the vector store by tenant and used horizontal sharding; the decoupled service can be autoscaled independently of the LLM server, so when the index grew to 200 M vectors the latency barely changed. The key takeaway is: keep the heavy‑weight vector lookup separate for flexibility, but plug in a thin, context‑aware re‑ranker right before generation to get the best of both worlds.