I will create massive dataset for ai training
Python Developer, Data Extraction Specialist
Vetted by Fiverr Pro
Kawsar was selected by the Fiverr Pro team for their expertise.
Vetted for
Data Processing
Data Scraping
About this Gig
You need a large, clean dataset to train or fine-tune an AI model, not a messy scrape dump or a tiny sample file.
That is the work I do. I am Kawsar, a Python developer who builds data scraping for massive datasets for AI training from public sources, synthetic generation, and structured CSV or JSON exports.
What I deliver:
- Custom AI training datasets sized to your package and schema
- Public web collection and/or synthetic rows when privacy-safe data is needed
- Clean CSV, JSON, or Excel ready for ML and LLM prep
- Field mapping, basic dedupe, and a short data dictionary
- Optional Python collection script on Premium
- Free sample rows before the full build so you can approve the schema first
- JSONL export ready for LLM fine-tuning (OpenAI, Hugging Face, and Llama-style formats)
I build custom datasets for ML and AI teams. Volume depends on the package and source difficulty. Public data only.
Message me your schema, target volume, and use case before you order.
NOTE:
- Public data only. Buyer provides schema and target URLs when scraping
- Custom build per order
- No personal data harvesting
- No school or university assignments
Platform:
MySQL
•
PL/SQL
•
PostgreSQL
•
SQLite
•
SQL Server
Expertise:
Design
•
Normalization
•
SQL
•
NoSQL
•
Performance
Clients I’ve worked with
MyHeritage
Internet Software & Services
I've been hired to carry out certain web extraction-related jobs in order to build a database for user reviews, which will be used to develop the product and make changes in certain instances.
May 2022-May 2023
Fiverr
Internet Software & Services
The project was to gather a detailed list of 5000 keywords from various profession types and list their Google Rank, Link, and even Local Search Results, etc. Information and make a Report with these!
Jul 2024
My Portfolio
FAQ
What kind of datasets do you create?
Structured AI training datasets for ML and LLM prep: tabular CSV/JSON, text corpora, and public-source collections. Image-heavy CV labeling can be scoped when it fits.
Do you scrape the web or generate synthetic data?
Both. We pick public collection, synthetic generation, or a mix based on your schema, privacy needs, and target volume.
Will you train my model too?
This gig is for creating the dataset. Model training and LLM fine-tuning are separate gigs if you need that next.
What files do I get?
CSV, JSON, JSONL or Excel plus a short data dictionary. Premium can include a Python script used to build or refresh the set.
Do I need to provide a schema?
Yes. Send column names, labels, formats, and example rows if you have them. That keeps the dataset model-ready.
Is LoRA or AI influencer dataset work included?
No. This gig is for ML/AI training data, not influencer LoRA packs.
Is this for school or university homework?
No. I do not take school or university assignments.

