Looks Like This Service Is On Hold

I will extract and clean offline HTML into structured text files txt MD jsonl

Mexico

I speak English, Spanish

7 orders completed

I’m Alejandro, a Web Scraping & Data Extraction specialist. I collect, clean, and deliver structured data from websites into CSV/Excel/Google Sheets—ready to use. Tech stack: Python (BeautifulSoup/Sel...
About this Gig

I will convert your downloaded/offline website export (HTTrack, SiteSucker, or similar) into clean, readable text files suitable for analysis, documentation migration, search indexing, or AI/RAG knowledge bases.

What you get:

  • Clean TXT (UTF-8) output (one file per HTML page)
  • Optional Markdown (MD) output
  • Folder structure mirrored (organized like your export)
  • Boilerplate removal (menus, navigation, headers/footers, sidebars, repeated blocks)
  • Main-content detection (main/article or heuristic extraction)
  • Processing log + final report (processed/written/skipped/failed)

What I need from you:

  • A ZIP file with your offline site folder (HTML files)
  • Tell me your preferred output: TXT / MD / JSONL
  • Any sections you want excluded (legal pages, login, cart, etc.)

Important notes:

  • This gig works with offline HTML exports. I do not require hosting access.
  • I do not scrape live websites in this gig (offline processing only).
  • Please ensure you have the rights to process the content you provide.

If you want a RAG-ready dataset, choose Premium (JSONL + metadata + dedupe).

Technology:

Python

Expertise:

Data acquisition

Data extraction

Data manipulation

My Portfolio