I will red team your llm agent for prompt injection and memory attacks


About this gig
Is your LLM agent safe to put in front of real users? I will attack it the way a real adversary would and tell you exactly what breaks and how to fix it.
I am an AI safety researcher with papers at ICML and IJCNLP and arXiv work on memory attacks against LLM agents. I have also built LLM-judge evaluation benchmarks, so my testing is systematic and reproducible, not a handful of clever prompts.
What I test:
- Prompt injection (direct and via documents, web pages or tool outputs)
- Jailbreaks and policy bypass
- Tool and function-call misuse
- Memory and context poisoning in agents with long-term memory
- System prompt and data leakage
What you get:
A clear written report with each finding, a reproducible example, a severity rating and a concrete fix. Higher packages add the attack suite as code so you can rerun it, plus a retest after you patch.
Works with OpenAI, Anthropic, open-source models, LangChain, LlamaIndex and custom agent stacks. Message me first with a short description of your agent and I will confirm scope. I only test systems you own or are authorized to test.
Get to know Sidharth P
AI Safety Researcher and LLM Fine Tuning Expert
- FromIndia
- Member sinceApr 2017
Languages
Telugu, English, Hindi
My Portfolio
Other AI Development Services I Offer
FAQ
Do you need access to my production system?
No. I can test a staging copy, an API endpoint with a test key, or a prompt and tool spec you share. I only test systems you own or are authorized to test, and I never need real user data.
What do I need to share to get started?
A test endpoint or staging build of your agent, its system prompt, the tools it can call and what it must never do. If it has long-term memory or reads documents or web pages, tell me, since those are common attack paths.
Will you keep my system and findings confidential?
Yes. Your prompts, code, data and findings are used only for your project and shared only with you. I do not publish or reuse anything specific to your system.
Which models and frameworks do you support?
Agents built on OpenAI, Anthropic, Gemini or open-source models (Llama, Qwen, Mistral and others), using LangChain, LlamaIndex, CrewAI, MCP tools or a custom stack. If it takes text in, I can test it.

