I will fix, debug, and optimize your rag pipeline or ai chatbot


About this gig
Your RAG pipeline or LLM chatbot works in the demo but fails with real users. Wrong chunks retrieved, hallucinated answers, slow responses, rising API costs. I fix that.
I am an AI engineer maintaining a production RAG system over thousands of documents, live on Google Cloud and serving real users daily. I have dealt with every failure mode this gig covers.
What I fix:
- Poor retrieval: chunking strategy, embeddings, hybrid search, reranking
- Hallucinations: grounding, citations, prompt structure
- Latency and cost: caching, model choice, batching
- Broken ingestion: PDFs, scraped data, messy documents
- No visibility: evaluation setup so quality is measurable
Stacks: LangChain, LlamaIndex, custom pipelines, OpenAI, Gemini, Claude, Pinecone, Qdrant, pgvector, Vertex AI Vector Search, FastAPI, GCP, AWS.
How it works: you share your repo or describe your setup, I audit it, and you get a written diagnosis with prioritized fixes. On Standard and Premium I implement the fixes and prove the improvement with before and after tests.
You own all code. NDA on request. Message me before ordering so I can confirm scope.
Get to know Muhammad Azeem
- FromPakistan
- Member sinceJun 2025
- Avg. response time1 hour
Languages
English

