At scale, crawling is an operations problem: frontier distribution, consistent dedup, backfill management, error recovery and storage that doesn't collapse under its own weight.
I build and tune crawling systems that reliably process millions of URLs across many hosts.
Scale crawling to millions of URLs across many hosts.
Scaled Cenoskop's crawling system from 5M to 30M+ products daily by adding distributed workers, tuning the URL frontier, and implementing backfill scheduling. The system crawls across millions of URLs with per-host politeness.
We audit your current crawl pipeline, then tune the frontier/workers/storage to your scale. From €3,000 for a high-volume tuning engagement.
High-volume crawl tuning and scale planning.
Built and operated production crawlers processing 30M+ products per day across millions of URLs — frontier, workers and storage tuned for throughput.
Other problems I help with.
Build price monitoring infrastructure for products/sellers.
Details →Relevant production work.
Crawler monitoring millions of products. Processing 30M products daily and comparing them.
Built and maintain the full crawling pipeline — from scraping through Elasticsearch indexing to price comparison.
Price intelligence engine for the US market. Processing large-scale product data.
Built the data pipeline and crawling infrastructure for large-scale US product monitoring.
Describe the problem in a few lines — I'll look at it and tell you what I think.