I will fine tune llm, evaluate dataset and train custom ai model for quality and safety

Nigeria

I speak English, French, Spanish

AI Data Annotation, LLM Response Evaluation Specialist

Data Annotation & AI Response Evaluation Specialist. I help AI teams, researchers, and developers build reliable models with clean, guideline-compliant datasets. Core Expertise: • CV Annotation: 2D ...
About this Gig

Need expert, human-in-the-loop validation for your AI models?

I provide rigorous LLM response evaluation, prompt testing, and AI data quality assessment to ensure your large language models deliver factual, safe, and reliable outputs.


Services Offered:

  • LLM response evaluation
  • Factuality evaluation & hallucination detection
  • Instruction-following assessment
  • Relevance & completeness evaluation
  • AI safety evaluation & toxicity review
  • Error classification & severity taxonomy
  • Pairwise preference ranking (RLHF)
  • Prompt testing & red-teaming
  • Healthcare LLM evaluation
  • Data quality assurance & dataset preparation


Tools & Formats:

Microsoft Excel, Google Sheets, CSV, JSON, and custom evaluation rubrics.


Quality Assurance:

Every prompt-response pair undergoes multi-pass manual review against your exact evaluation guidelines, scoring criteria, and ground-truth reference sources.


Contact me to review your guidelines or request a sample evaluation!


Relevant Searches:

llm evaluation, ai evaluation, prompt testing, factuality evaluation, ai safety, rlhf, data annotation, response quality, hallucination detection, text evaluation data annotation data labeling image annotation ai data-annotation

Programming Language:

Python

Keras

Pytorch

R

Tensorflow

AI Model Frameworks & Tools:

TensorFlow

PyTorch

Keras

Data Type:

Text

Images

Multimodal

AI Engine:

GPT

Gemini

DALL-E

DeepSeek

Stable diffusion

Midjourney

My Portfolio