I will fix evaluate and optimize your rag or llm app for accuracy and cost

P
pulkit004
P
pulkit004
Pulkit V.

About this gig

Your AI app works in the demo but breaks with real users: wrong answers, hallucinations, slow responses, surprising API bills. I'll find out why and fix it.


Common fixes:

  • Bad retrieval: chunking, embeddings, hybrid search, reranking
  • Hallucinations: grounding, citations, refusal behavior
  • Prompt and structured-output bugs (JSON that breaks)
  • Latency: streaming, caching, smaller models where they're good enough
  • Cost: token budgets, model routing, prompt caching
  • No evals: I set up a test set so you measure quality instead of guessing


Works with LangChain, LlamaIndex, raw OpenAI/Claude SDKs, Vercel AI SDK, AWS Bedrock and n8n.


Why me: As an Engineering Manager leading AI initiatives, I've shipped production RAG pipelines and AI review agents, and I care about the unglamorous parts: evals, failure cases and cost per request.


Message me with your stack and 3-5 examples of bad outputs.

Get to know Pulkit V.

Pulkit V.

Engineering Manager building production RAG and AI agents

  • FromIndia
  • Member sinceJun 2017
  • Languages

    Hindi, English
I build AI systems that ship, not demos. Engineering Manager at Setu (Pine Labs), leading AI initiatives: production RAG pipelines, AI agents in engineering workflows, and LLM-powered insights. 8+ years building products as an engineer and tech lead. Founder of WeaveAI, an AI engineering studio. M.S. Data Science at the University of Pittsburgh. I deliver: RAG with citations, n8n automations and AI agents, Claude Code setups for dev teams, and fixes for LLM apps that hallucinate, lag, or cost too much. Evals included. Message me before ordering so we can scope it right.

My Portfolio