o
olegokeev

Oleg Okeev

@olegokeev

I build RAG ready knowledge bases from your texts

Vietnam
English
About me
I transform unstructured texts into structured, RAG-ready knowledge bases in Obsidian. My flagship project: a 4,400+ note TCM vault built from 25 expert sources — textbooks, lectures, web databases, and clinical courses. What I deliver: - Obsidian vaults with YAML frontmatter, wiki-links, and semantic tags - OCR extraction from scanned PDFs (3,000+ pages processed) - Custom Python parsers for web scraping and data normalization - Molecular cross-references (LOTUS, PharmGKB, CPIC) - Multilingual content (English, Russian, Chinese) Tools: Obsidian, Claude AI, Python, PDFgear, Whisper... Read more

Skills

o
olegokeev
Oleg Okeev
Offline • 
Average response time: 1 hour

See my services

Database Development
I will build a structured rag knowledge base in obsidian from your texts

Portfolio

Work experience

Fiverr

Knowledge Base Architect & Data Engineer

Fiverr • Self-employed

Jan 2026 - Present • 9 mos

I build structured, RAG-ready knowledge bases in Obsidian from unstructured source materials — medical textbooks, video lectures, web databases, and clinical courses. Key project: TCM Knowledge Base (4,400+ notes) • Processed 25 expert sources across Traditional Chinese Medicine: acupuncture, herbal medicine, diagnostics, dietary therapy, aromatherapy • OCR extraction from scanned PDFs (3,000+ pages) using PDFgear and pdftotext • Built custom Python parsers for web databases (eledia.ru, kiberis.ru, TCMwiki), extracting 1,280+ acupuncture points and 245 treatment protocols • Structured 303 herbal monographs with molecular cross-references from LOTUS (659K compound-organism pairs), PharmGKB, and CPIC databases • Created 99 food-as-medicine entries with full TCM classification • Transcribed video lectures using Whisper (pywhispercpp) for speech-to-text processing Technical approach: • Consistent YAML frontmatter on every note for automated parsing and vector embedding • Wiki-link knowledge graph connecting herbs, acupoints, diseases, and treatment schemes • Standardized taxonomy: 14 meridian codes, functional categories, body-region tags • Semantic chunking — one concept per note, optimal for RAG retrieval Tools: Obsidian, Claude AI, Python, PDFgear, Whisper, LOTUS/PharmGKB/CPIC molecular databases GitHub portfolio: github.com/Smeilz/tcm-knowledge-base