I will build a production rag system over your documents


About this gig
I build production RAG systems that answer from YOUR documents with citations, not hallucinations.
What you get
- Document ingestion pipeline (PDF, DOCX, HTML, Notion, Confluence) with layout-aware chunking
- Hybrid retrieval: BM25 + dense vectors + reranking, tuned on your own eval set
- Answer synthesis with inline citations and refusal when evidence is missing
- FastAPI / Node service, Docker, CI, and an admin dashboard for re-indexing
Why me
I ship LLM systems in production. Recent measured results: 70% lower p95 latency, 38% lower serving cost, 49% fewer output tokens through context caching, structured output and model routing.
Stack
Python, TypeScript, LangChain, LlamaIndex, OpenAI / Anthropic / vLLM, pgvector, Qdrant, Elasticsearch, Postgres, Redis, Docker, AWS.
Send me a sample of your documents and the questions you need answered, and I will tell you exactly what is achievable before you order.
Get to know Cheolhee Lee
AI Full Stack Developer specializing in LLM and RAG optimization
- FromSouth Korea
- Member sinceApr 2021
- Avg. response time1 hour
Languages
English, Korean
