I will build an llm evaluation harness for your ai agent or chatbot

A
akashlp
A
akashlp
Akash L

About this gig

Most AI agents ship with no number attached to them. You change a prompt, swap a model, add a tool, and nobody can say whether it got better or quietly got worse.


I build the evaluation harness that answers that.


I am an AI engineer with 4 years of production experience. I was the sole author of the conversational agent behind a US supply-chain platform's in-product assistant, a multi-agent LangGraph pipeline over a 125-parameter tool surface, and I built its three-pass validation layer: 38 enum-checked fields, LLM-based value correction and a hallucination pass.


WHAT YOU GET

  • A test set built from your real traffic or your spec
  • Metrics that fit your use case: exact match, semantic similarity, faithfulness, tool-call accuracy, latency and cost
  • LLM-as-judge scoring with calibrated rubrics
  • One command that runs the suite and produces a scored report
  • Regression baselines so you can compare any two versions
  • Optional CI gating so a bad prompt never reaches production


Stack: Python, pytest, LangChain, LangGraph, OpenAI, Anthropic, Llama, Ragas, DeepEval, MLflow.


Tell me what your agent does and I will tell you what I would measure.

Get to know Akash L

Akash L

AI Engineer, LLM Agents, RAG and MLOps

  • FromIndia
  • Member sinceMar 2020
  • Languages

    Tamil, English
AI engineer with 4 years building LLM systems that run in production, not just demos. I was sole author of a multi-agent LangGraph assistant for a US supply-chain platform. Plain-English questions turned into validated calls across a 125-parameter tool surface. It survived a re-platform and still runs today. I build the unglamorous parts too: RAG pipelines, tool schemas, validation that degrades gracefully instead of crashing, and eval harnesses that catch regressions before your users do. Python, LangGraph, LangChain, RAG, FastAPI, AWS. Tell me what you are building.