I will build an ai chatbot for your pdfs and documents with citations


About this gig
Turn your approved documents into a chatbot that answers questions with links back to the source. I build retrieval-augmented generation (RAG) systems with a defined document set, an agreed chat channel, and tests based on questions your users actually ask.
Basic: up to 30 supplied documents, one model, one web chat interface, citations, and setup notes. Standard: up to 200 documents, hybrid retrieval, a sample evaluation set, and support for two agreed languages. Premium: up to 500 documents from two sources, authentication, one integration, evaluation checks, monitoring, and 14 days of defect fixes.
I will test retrieval and answer quality against agreed examples and show you the results. AI answers can still be wrong; citations and a clear fallback make them easier to check. You receive the source code, configuration notes, and a walkthrough.
Message me with your document types, sample questions, desired channel, hosting preference, and privacy needs before ordering. Third-party hosting and model usage are separate.
Get to know Lucas Loo Tan
AI Video Ads, Product Images and AI Development
- FromMalaysia
- Member sinceJan 2023
- Avg. response time1 hour
Languages
Chinese, English, Malay
My Portfolio
FAQ
What doc formats can you ingest?
PDF, DOCX, HTML, Markdown, plain text, Notion, Confluence, public websites (via scraping), CSV structured data. If you have something unusual, ask - I'll tell you if it's feasible.
Which LLM and vector DB do you use?
Flexible - OpenAI, Anthropic Claude, Gemini, or open-source (Llama, Mistral) for self-hosted. Vector DB: Pinecone, Qdrant, pgvector, or Weaviate. Chosen based on your data volume, latency, and budget.
How accurate will the chatbot be?
Depends on your docs quality, but I include an eval harness so you can measure it, not guess. Premium tier includes grounding checks and citation accuracy scoring. If the docs don't have the answer, it says so instead of hallucinating.
Can it handle non-English docs?
Yes. I work across EN / MS / ZH / ID / TH, and the multi-language extra unlocks cross-lingual retrieval (user asks in English, docs are in Chinese, etc).
Do I own the code and data?
Yes, 100%. Source code is yours. Your docs and embeddings stay in infrastructure you control. No vendor lock-in from my side.

