d
daiki_mayumi

Daiki M

@daiki_mayumi

Native Japanese AI Evaluator, LLM Data Annotation and AI Automation

Japan
English, Japanese
About me
Native Japanese AI evaluator and AI/LLM engineer in Tokyo. Since Dec 2024 I have rated LLM and chatbot outputs for Invisible Technologies, Scale AI, RWS and Turing: rubric scoring, side-by-side comparison, fact and safety checks, and written rationales in Japanese. Familiar with Japanese style guidelines (honorifics, register, punctuation). I also build LLM apps (Claude, Gemini, OpenAI APIs) and evaluation harnesses, so I know how rater data is used downstream. Full-time AI engineer since 2026. Rubric first, no guessing, consistent scores. JP native, EN business (TOEIC 800).... Read more

Skills

d
daiki_mayumi
Daiki M
Offline • 

See my services

Data Labeling & Annotation
I will annotate and label japanese text and audio data for ai training
Proofreading
I will proofread and polish your japanese text as a native speaker

Work experience

SNAFTY_STUDIO

AI / LLM Application Engineer

SNAFTY STUDIO • Full-time

Mar 2026 - Present6 mos

Requirements definition, design, PoC development, API integration, and testing of business systems that embed generative AI (Claude, Gemini, OpenAI APIs). - Wrote more than 10 requirements and design documents in three months: ID-managed functional and non-functional requirements, phasing, open-decision tracking, and cost structure. - Built and shipped LLM applications with Claude Code: order-intake AI that structures email/Excel/PDF orders into CSV, a multi-tenant outfit recommendation web app (Next.js, Go), billing reconciliation automation (Python), and an AI short-video production pipeline with a QA harness. - Design principle: AI handles extraction and drafting, rule engines handle decisions and calculations, humans stay in the loop, audit logs everywhere.

AI Evaluator / Prompt Engineer (Freelance)

Invisible Technologies, Scale AI, RWS, Turing • Freelance

Dec 2024 - Present1 yr 9 mos

Human evaluation of LLM and chatbot outputs in Japanese for AI training vendors (Invisible Technologies, Scale AI, RWS, Turing). - Reviewed and ranked chatbot and generated-content quality against evaluation guidelines; verified factual accuracy, safety, and harmful-content compliance; annotated data and submitted model-improvement feedback. - Designed and tested prompts to stabilize output accuracy and reproducibility; verified naturalness of Japanese-language data as a native speaker. - Designed persona-consistent dialogue prompts for an AI avatar character on a major social platform, including age-restricted scenarios, balancing character consistency, platform guidelines, and user experience.