I will regression test your ai chatbot or agent for costly failures

United States

I speak English, Chinese

35 orders completed

AI Systems, Automation and Reliability Engineer

Hi, I’m Dylan Feng, a former OpenAI engineer and YC-backed founder specializing in reliable AI systems, workflow automation, and production infrastructure. I help businesses audit and harden n8n workf...
About this Gig

Changing a prompt can fix one failure and quietly create another. I turn your agent requirements into repeatable pass/fail tests, run normal and adversarial conversations, and show the exact evidence behind every failure.


You receive:

deterministic requirements and acceptance checks

exact prompts, responses, and reproducible failure evidence

critical, high, medium, and low prioritization

a release recommendation

a reusable test pack in Standard and Premium


Coverage can include grounding, hallucinated claims, privacy, prompt injection, unsafe responses, and tool/action boundaries. High- and critical-severity results are manually reviewed. This is not a vague AI-generated score or a promise of zero defects.


Initial scope is text-based support, lead-generation, ecommerce, and appointment agents. Regulated medical, legal, and financial systems are excluded.


A staging environment is strongly preferred. I will not test destructive actions in production. You provide test credentials, source-of-truth documents, required and forbidden behavior, and confirm the environment contains no real customer secrets or personal data.

Testing application:

Web application

Device:

PC

Mac

Linux