I will fix rag chatbot hallucinations


About this gig
Your chatbot shipped. But is it actually good? If you can't answer that with numbers, you're flying blind, and one prompt change away from a regression nobody catches until users do.
I build custom eval suites for LLM apps: test sets built from YOUR real use cases, an LLM-judge harness calibrated against human preferences, and regression tests that catch quality drops before production.
What you get:
- Custom test sets from your real queries and edge cases
- Judge calibration so scores match what humans prefer
- Regression testing wired into your workflow
- Metrics that matter: accuracy, hallucination rate, consistency
- Pass/fail reports your whole team can read
Standard tier includes model comparison: GPT vs Claude vs open-source on YOUR workload.
Built on LangSmith, Braintrust, or Promptfoo depending on your stack. You get the harness, docs, and a handoff walkthrough.
Who this is for: teams who shipped a chatbot and don't know if it's improving, founders comparing models before committing, anyone burned by a silent quality drop.
Message me with what your bot does and I'll scope the suite fast.
Get to know Michiel H
Marketing Strategist and Blockchain Consultant
- FromMexico
- Member sinceMar 2019
- Last delivery3 years
Languages
English, Spanish, Dutch
Other AI Development Services I Offer
FAQ
Why is my RAG chatbot hallucinating?
Usually broken chunking or retrieval. The model gets blamed, but in most audits the pipeline feeds it garbage context.
Can you guarantee zero hallucinations?
No one can promise that, and anyone who does is lying. You get a measured before/after rate plus citations so any answer can be checked in seconds.
What access do you need?
Access to your vector database setup, embedding config, and a sample of real user questions with the answers you expected.
Do you work with any vector database?
Yes. Pinecone, Weaviate, Chroma, pgvector, Milvus, and plain Postgres setups.
What do I get at the end?
The fixed system, a before/after accuracy report, and docs explaining what was broken and what changed.

