I will do ai training, llm evaluation and rlhf data annotation
AI Training, Data Annotation Specialist and LLM Evaluator
About this Gig
I help AI companies build reliable training data through manual annotation, LLM output evaluation, and RLHF ranking.
What I do:
AI training & LLM evaluation (accuracy, coherence, helpfulness, harmlessness scoring)
Data annotation & labeling text, image, video (YOLO/COCO)
RLHF ranking & prompt/response grading
Inter-annotator agreement & QA auditing (Cohen's Kappa)
Why work with me:
Native fluency in English, Swahili & Kikuyu rare multilingual coverage for global AI training pipelines needing cultural nuance and low-resource language evaluation.
Background in Economics & Finance drives strict data integrity, structured rubrics, and audit-ready documentation on every task not just labeling, but defensible, consistent judgment calls on edge cases (occlusion, ambiguous boundaries, multimodal inconsistency).
Every project is delivered with clear scoring criteria and documentation, so your dataset is clean, consistent, and ready for model training not just "done."
Let's build data you can actually trust.
Technique:
Manual
Tagging type:
Text
•
Image
•
Video
My Portfolio
FAQ
What annotation formats do you support?
I work with YOLO and COCO formats for computer vision, plus text and video annotation. I can also adapt to your team's specific labeling schema or guidelines.
How do you ensure annotation quality and consistency?
Every task follows a structured rubric with documented scoring criteria. For larger projects, I track inter-annotator agreement (Cohen's Kappa) to catch inconsistency early rather than after delivery.
Do you evaluate LLM outputs, or only label raw data?
Both. I handle LLM response evaluation (accuracy, coherence, helpfulness, harmlessness scoring) and RLHF ranking, alongside traditional data annotation and labeling.
Can you handle multilingual content?
Yes — I'm a native speaker of English, Swahili, and Kikuyu, which covers annotation and evaluation work requiring cultural nuance or low-resource language coverage.
What if I need more items than your package includes?
I offer additional-item extras on every package, and I'm happy to discuss custom volume pricing for larger ongoing projects.

