I will automate PDF data extraction to excel CSV or json with python
Machine Learning Engineer for LLM Agents RAG and Python Automation
About this Gig
Stop copying document data by hand. I will build a reusable Python script that extracts agreed fields or tables from text-based PDFs into Excel, CSV or JSON.
I am an AI/ML engineer with 4+ years of experience in Python, NLP and document-processing automation. My work includes key-value extraction, batch ingestion and data validation workflows.
Basic: one consistent layout, up to 20 pages and 10 fields.
Standard: up to 2 layouts, 100 pages and 20 fields, with batch processing and output cleanup.
Premium: up to 3 layouts, 300 pages and 30 fields, with validation checks and error reporting.
Each package includes Python source code, one agreed output format, sample results and setup instructions. Revisions apply to the agreed layouts and fields. Scanned PDFs, handwriting, changing layouts, external integrations and deployment require a custom scope. Third-party OCR or API fees are separate.
Please send a sanitized sample and your desired output columns before ordering. I will check feasibility and confirm the extraction rules.
Technology:
Python
Expertise:
Data extraction
•
Data validation
•
Transformation
My Portfolio
FAQ
Do you support scanned PDFs or handwriting?
The listed packages cover text-based PDFs. OCR and handwriting require sample review and a separate quote; extraction quality depends on the source.
Will I receive a reusable script?
Yes. All packages include Python source code and setup instructions for the agreed layouts, plus sample output in one format: Excel, CSV or JSON.
What counts as a document layout?
A layout is a consistent structure of fields and tables. Different invoice templates or substantially changed page structures count as separate layouts.
