ScraperHub
A template-driven scraping platform with a live dashboard: point it at a catalog, give it a JSON template, watch the data flow in real time.
- Client
- Brand products data suite
- Year
- 2025
- Category
- API integration
- Role
- Data engineering, platform build
- Stack
- Python · FastAPI · Playwright · SQLAlchemy 2 · WebSockets · React 18 · Vite
Collecting product data from authenticated dealer catalogs meant brittle one-off scripts that died on disconnects and duplicated data. Every new site meant starting from scratch.
I built a generic scraping engine: FastAPI + async Playwright behind a job orchestration system with checkpoint-based resume. Each site is just a JSON CSS-selector template: login, category discovery, pagination, product pages, related items. A React dashboard streams logs live over WebSockets.
Long scraping runs survive network drops, because a connectivity probe pauses the job and auto-resumes when the connection returns. Output lands as clean, category-organised Excel files with persistent SKU deduplication.
rows lost to disconnects, with checkpointed resume built in
The dashboard shows every job as a live progress story: status, items scraped, files created, errors, all streamed to the screen the moment they happen.
Checkpoints persist as JSON in SQLite after every page, so any crashed or interrupted job resumes exactly where it stopped, even across machine restarts.