I will build a production rag system over your documents

C
chlee9
C
chlee9
Cheolhee Lee

About this gig

I build production RAG systems that answer from YOUR documents with citations, not hallucinations.


What you get

- Document ingestion pipeline (PDF, DOCX, HTML, Notion, Confluence) with layout-aware chunking

- Hybrid retrieval: BM25 + dense vectors + reranking, tuned on your own eval set

- Answer synthesis with inline citations and refusal when evidence is missing

- FastAPI / Node service, Docker, CI, and an admin dashboard for re-indexing


Why me

I ship LLM systems in production. Recent measured results: 70% lower p95 latency, 38% lower serving cost, 49% fewer output tokens through context caching, structured output and model routing.


Stack

Python, TypeScript, LangChain, LlamaIndex, OpenAI / Anthropic / vLLM, pgvector, Qdrant, Elasticsearch, Postgres, Redis, Docker, AWS.


Send me a sample of your documents and the questions you need answered, and I will tell you exactly what is achievable before you order.

Get to know Cheolhee Lee

Cheolhee Lee

AI Full Stack Developer specializing in LLM and RAG optimization

  • FromSouth Korea
  • Member sinceApr 2021
  • Avg. response time1 hour
  • Languages

    English, Korean
I keep my employer and my clients unnamed here. I ship production AI systems end to end at an undisclosed B2B AI SaaS company - a sales-automation SaaS and a public-sector AI evaluation platform. Measured: LLM p95 latency -70%, serving cost -38%, output tokens -49% via context caching and structured output. 125x list speedup, threads query 411ms to 1.6ms, bundle 21.7MB to 2.3MB. Re-homed three LLM models to an on-prem DGX with zero downtime; passed TTA review for Korea's AI Verification program. TypeScript, Python, Rust, Go, React, PostgreSQL, AWS, RAG, vLLM, MCP.

My Portfolio

Other AI Development Services I Offer