As a Machine Learning Engineer specializing in NLP and symbolic AI, I provide comprehensive evaluation, benchmarking, and red-teaming for your Large Language Models (LLMs) and RAG pipelines. I help engineering teams identify edge-case failures, eliminate hallucinations, and enforce strict security guardrails.
What I Offer:
- LLM Benchmark & Performance Auditing: Evaluating model accuracy, latency, context usage, and reasoning capabilities against customized testing rubrics.
- Hallucination & Edge-Case Detection: Stress-testing system prompts, RAG retrieval accuracy, and factual consistency under adversarial conditions.
- Red Teaming & Security Testing: Identifying vulnerabilities like prompt injection attacks, jailbreaks, and system prompt leaks.
- Schema & Output Validation: Verifying strict JSON/Pydantic structure compliance, function calling execution, and API integration reliability.
- Detailed Audit Report & Action Plan: Comprehensive failure analysis complete with actionable prompt engineering fixes and guardrail recommendations