This is a sample case study showing the format. Replace it with a real project.
Problem
The support team searched ~4k documents (PDFs, Confluence, procedures). Finding the right answer took well over ten minutes on average, and new hires needed months to find their way around.
Decisions
- Hybrid search (BM25 + pgvector embeddings) with re-ranking — pure vector search missed proper names and procedure numbers.
- Every answer cites its sources with a link to the document fragment.
- Model router: simple questions go to a cheaper model, complex ones to a stronger one. Semantic cache for repeated questions.
- A 200-question test set with RAGAS evaluation run on every prompt or index change.
Outcome
- Faithfulness 0.87 on the test set.
- LLM cost -62% compared to one large model for everything.
- p95 response time under 2 seconds.