Browse categories
Explore
Fiverr Pro
English
$
USD
I will build an llm evaluation framework


Sushma Sharma
About this gig
Your AI system looks fine in demos. But is it still working
correctly in production next week?
Most teams find out their AI broke when a user complains.
I build evaluation infrastructure that catches failures before
users do automatically.
What I build:
- Retrieval evaluation (are you fetching the right content?)
- Faithfulness scoring (is the answer grounded or hallucinated?)
- Regression detection (did your last prompt change break anything?)
- LLM-as-judge pipelines with no ground truth required
- Streamlit dashboard showing quality trends over time
If you are running AI in production without evals, you are
flying blind. Message me and I will tell you where your
system is most likely failing.
Get to know Sushma Sharma
Sushma Sharma
AI Engineer
- FromIndia
- Member sinceJun 2026
- Avg. response time3 hours
Languages
English
AI Engineer with 5+ years of experience designing and building AI systems:
🔹 Multi-agent & agentic AI: Built ReAct-based LLM workflows for automated proposal generation. A Multi-RAG framework I designed identified $0.8M in margin leakage in customer proposals.
🔹 RAG & retrieval optimization: Improved retrieval accuracy by 20% using semantic query enhancement. Built GenAI risk assessment workflows for safety documentation.
🔹 LLMOps: Evaluation pipelines, prompt testing, and model performance monitoring for enterprise AI.
Communicate well with cross-functional teams.
