I will build an llm evaluation harness for your ai agent or chatbot


About this gig
Most AI agents ship with no number attached to them. You change a prompt, swap a model, add a tool, and nobody can say whether it got better or quietly got worse.
I build the evaluation harness that answers that.
I am an AI engineer with 4 years of production experience. I was the sole author of the conversational agent behind a US supply-chain platform's in-product assistant, a multi-agent LangGraph pipeline over a 125-parameter tool surface, and I built its three-pass validation layer: 38 enum-checked fields, LLM-based value correction and a hallucination pass.
WHAT YOU GET
- A test set built from your real traffic or your spec
- Metrics that fit your use case: exact match, semantic similarity, faithfulness, tool-call accuracy, latency and cost
- LLM-as-judge scoring with calibrated rubrics
- One command that runs the suite and produces a scored report
- Regression baselines so you can compare any two versions
- Optional CI gating so a bad prompt never reaches production
Stack: Python, pytest, LangChain, LangGraph, OpenAI, Anthropic, Llama, Ragas, DeepEval, MLflow.
Tell me what your agent does and I will tell you what I would measure.
Get to know Akash L
AI Engineer, LLM Agents, RAG and MLOps
- FromIndia
- Member sinceMar 2020
Languages
Tamil, English
FAQ
I do not have a test set. Can you still help?
Yes. Most clients do not. I build one from your production logs, your docs, or a spec of what the agent is supposed to do, and I show you the cases before I score anything.
Do I have to give you access to my codebase?
No. The harness can run against an API endpoint or a thin wrapper you provide. Repo access only helps if you want CI gating wired in directly. I sign an NDA on request.
Which frameworks and models do you support?
LangChain, LangGraph, LlamaIndex, CrewAI or a plain Python app, running on OpenAI, Anthropic, Llama or any model behind an API. If it takes text in and gives text or tool calls out, I can score it.

