I will extract data from PDF files a clean excel table
About this Gig
Your data exists. It is just stuck somewhere useless - inside a PDF, spread across a website, buried in scanned invoices. I get it out and hand you a table you can sort, filter and count.
What you get:
- a clean .xlsx or Google Sheet, columns exactly as you need them
- - consistent formats: dates, numbers, phones and prices normalized, not left as text
- - rows that looked wrong flagged instead of silently dropped
- - a short summary: how many records found, how many skipped and why
I work with PDFs (including scanned ones, via OCR), websites, and mixed sources where half is one and half the other.
What I do not do: sites that require breaking a login or bypassing protection, and anything built on personal data collected without consent. If your task is close to that line, ask me first and I will tell you honestly.
Not sure it can be extracted? Send one sample page or link before ordering. I will look and tell you what is realistic - that costs you nothing.
Technology:
Python
•
Google Sheets
•
Excel
•
Beautiful soup
•
Playwright
Technique:
Automated
FAQ
Can you handle scanned PDFs?
Yes, through OCR. Accuracy depends on scan quality; send one page and I will tell you what to expect before you order.
Will it break if the website changes?
Any scraper eventually does. In the Premium package I hand you the script and explain which line to fix when the layout changes.
Do you deliver Google Sheets or Excel?
Either, your call.
