I will build and deploy a multimodal rag system for you
About this gig
Your documents, images, websites, and live APIs all hold answers but only if something can search across all of them at once. I build multimodal RAG (retrieval-augmented generation) systems that do exactly that: one AI assistant, grounded in every source you have, answering with real citations instead of guesses.
What's included:
- Document, image, and URL ingestion OCR and vision embeddings for images, indexed into a vector database (Pinecone, Chroma, or Weaviate)
- Hybrid semantic search with re-ranking, so retrieval stays accurate as your knowledge base grows
- Live API integration, so answers blend stored knowledge with real-time data, not just static files
- LLM integration (OpenAI, Claude, or open-source models like Qwen) with prompt engineering tuned to reduce hallucination
- A clean chat interface or API endpoint, with source citations on every answer
Why work with me:
I'm a computer vision and machine learning engineer with 5 years of production experience including LLM fine-tuning (QLoRA on Qwen) and vision systems the exact skill overlap multimodal RAG actually needs, not just prompt wrapping.
Tell me your data sources and use case before you order, and I'll map out the right arch
Get to know Rukon Uddin
AI Engineer
- FromBangladesh
- Member sinceJul 2026
- Avg. response time1 hour
Languages
English, Arabic, Spanish
Other AI Development Services I Offer
FAQ
What input types can the system handle?
Documents (PDF/Word), images (via OCR and vision embeddings), web pages, and live API data — all searchable together.
How does the image search work?
Images are converted to vector embeddings (CLIP or similar) so the system retrieves relevant images or image-derived text alongside your documents.
Can it pull real-time data?
Yes — I connect live APIs so answers combine your stored knowledge with fresh, current information.

