Teams accumulate knowledge across PDFs, meeting transcripts, and spreadsheets but have no way to query them intelligently. Searching manually means lost context, duplicated effort, and slow onboarding for new members.
Two-pipeline system: Ingestion (upload → extract text → chunk → generate embeddings → store in pgvector) and Query (user question → embed → cosine search → assemble context → LLM via OpenRouter → response with source citations). Local sentence-transformers for embeddings, OpenRouter for 100+ LLM models.
Built the entire system end-to-end: Django backend with pgvector, document processor supporting 6 file formats, local embedding pipeline (zero API cost), OpenRouter integration for 100+ models, conversation handler for casual queries, analytics dashboard, and full web UI.
The main challenge was balancing search quality with performance. Implemented pgvector cosine similarity with tunable thresholds, rich metadata injection into LLM prompts (participants, dates, topics), and a conversation handler that intercepts casual queries before they hit the document search pipeline.
6 supported file formats (PDF, DOCX, Markdown, JSON/Fireflies transcripts, CSV). Local embedding model (all-mpnet-base-v2, 768 dimensions) — zero API cost. 100+ LLM models via OpenRouter. Conversation history with session tracking. Analytics dashboard with performance metrics.
Instant document Q&A across any uploaded file. Zero embedding cost using local models. 100+ AI models accessible through a single interface. Source citations with similarity scores for every answer. Full conversation history and query analytics for continuous improvement.