I will fix rag chatbot hallucinations

M
michielhorstman
M
michielhorstman
Michiel H

About this gig

Your chatbot shipped. But is it actually good? If you can't answer that with numbers, you're flying blind, and one prompt change away from a regression nobody catches until users do.


I build custom eval suites for LLM apps: test sets built from YOUR real use cases, an LLM-judge harness calibrated against human preferences, and regression tests that catch quality drops before production.


What you get:

- Custom test sets from your real queries and edge cases

- Judge calibration so scores match what humans prefer

- Regression testing wired into your workflow

- Metrics that matter: accuracy, hallucination rate, consistency

- Pass/fail reports your whole team can read


Standard tier includes model comparison: GPT vs Claude vs open-source on YOUR workload.


Built on LangSmith, Braintrust, or Promptfoo depending on your stack. You get the harness, docs, and a handoff walkthrough.


Who this is for: teams who shipped a chatbot and don't know if it's improving, founders comparing models before committing, anyone burned by a silent quality drop.


Message me with what your bot does and I'll scope the suite fast.


Get to know Michiel H

Michiel H

Marketing Strategist and Blockchain Consultant

5.0(14)
  • FromMexico
  • Member sinceMar 2019
  • Last delivery3 years
  • Languages

    English, Spanish, Dutch
Marketing strategist & blockchain consultant with 5+ yrs experience growing brands & 6+ yrs in crypto. Let's connect & transform your biz with effective digital marketing & web3 solutions—founder of 2 agencies, award-winning work.

Related tags