I will evaluate ai and llm writing
About this Gig
An AI response can sound confident and still get the facts wrong, miss instructions, or fail the task. I'll review your model's outputs and show you where they fall short.
I've spent three years working in AI training and evaluation, including work with Mercor and other AI platforms. My experience includes response scoring, rubric development, comparing model outputs, and checking business and accounting tasks.
For each response, I'll check:
Accuracy and unsupported claims
Instruction following
Relevance and completeness
You'll receive a spreadsheet with scores and clear explanations tied to specific parts of each response. Standard and Premium packages also include recurring issues, with improvement recommendations included in Premium.
Each item includes one prompt and one response, up to 500 words each. I can use your rubric with up to five criteria. One revisions included for adjustments to the original evaluation.
Need a custom rubric, rewritten answers, or reviews that require extensive research? Message me first so we can agree on the scope.
Technique:
Automated
Tagging type:
Text
•
Image
•
Video
