I will extract structured data from pdfs to excel CSV and json
Python and Local AI Automation Developer
About this Gig
I will extract structured data from native-text PDF documents and deliver it in a clean, reusable format.
I can extract specific fields from invoices, reports, forms, statements, and other structured or semi-structured PDFs into Excel, CSV, or JSON depending on the selected package.
My workflow can include:
Field-based data extraction
Excel and CSV output
JSON output on Standard and Premium
Page references for extracted values
Missing or invalid value warnings
Data validation
QA report and organized ZIP delivery
Custom validation rules on Premium
I do not guess missing values. If a value cannot be reliably found, it will be left missing or flagged for review.
This service is designed for native-text PDFs. OCR, handwritten documents, password-protected files, account logins, API integration, and production-system access are not included.
Please send the PDFs, the exact fields you want extracted, and your preferred output format.
All communication and delivery can be handled asynchronously through Fiverr.
Technology:
Excel
•
Python
My Portfolio
FAQ
What types of PDF files do you support?
I support native-text PDF documents where text can be selected or copied. Scanned-image PDFs and handwritten documents requiring OCR are not included.
What output formats do you provide?
Excel and CSV are available in all packages. JSON output is included in Standard and Premium.
Can you extract specific fields I choose?
Yes. Please provide the exact field names you want extracted, such as invoice number, date, company name, total amount, email, or other document-specific fields.
Do you provide page references?
Yes. Extracted values can include the source PDF and page reference so you can trace where each value came from.
What happens if a value cannot be found?
I do not invent or guess missing values. The field will be left missing or flagged for review depending on the extraction and validation result.
Can you validate the extracted data?
Yes. Standard and Premium include validation warnings. Premium can also include custom validation rules based on your requirements.
Do you support scanned PDFs or OCR?
No. This Gig is currently limited to native-text PDFs. Image-only scans and handwritten documents that require OCR are outside the supported scope.
Can you access my account, website, API, or database?
No. This service is file-based only. I do not require account logins, API access, database access, or production-system access.
Do we need a video call or meeting?
No. The entire project can be handled asynchronously through Fiverr messages and file exchange.
What should I send before the order starts?
Please send the source PDFs, the exact fields to extract, your preferred output format, and any validation rules or example output you want me to follow.

