I will set up a litellm gateway with llm observability

S
spoidey
S
spoidey
ali shah

About this gig

LLM apps without a gateway and traces turn into guesswork: no idea which key spent what, which model is slow, or why costs doubled after launch. I set up LiteLLM + Langfuse so you can see and control all of it.


WHAT I SET UP

- LiteLLM proxy on Docker, VPS or your cloud, with an OpenAI-compatible /v1 endpoint

- Multi-provider routing: OpenAI, Anthropic, Gemini, Azure, Bedrock, OpenRouter, local models

- Virtual keys per app/team with budgets, RPM/TPM limits and per-model access

- Fallback chains and automatic retry/failover on 429, 500 and timeouts

- Cost tracking per key, per model, per customer, with 60/80/100% spend alerts

- Langfuse (or LangSmith/Helicone) traces, latency percentiles and error logs

- One dashboard or weekly report your team can actually read

- config.yaml, deployment notes and a handoff runbook


This is not a chatbot build. It's for teams already running or shipping an LLM app who want production control over it.


Send your providers, traffic estimate and where it runs today. I'll confirm scope in 24h.

Get to know ali shah

ali shah

AI LLM Engineer RAG, Agents, LLM Ops Shopify Hydrogen

  • FromPakistan
  • Member sinceAug 2026
  • Languages

    Urdu, English
5 years shipping production software (Systems Limited, then Contour Software) as a backend + AI engineer. I build the unglamorous parts that make AI products work: RAG retrieval that actually returns the right chunk, evals that catch hallucinations before customers do, LiteLLM gateways and Langfuse traces so you know your per-user cost, and vLLM on your own GPU. Also custom Shopify Hydrogen storefronts. .

My Portfolio

Other AI Development Services I Offer