I build crawling infrastructure that scales: URL frontier management, politeness, retry/backoff, distributed workers and a place to store and queue what's been crawled. Not a one-off scraper — infrastructure you can keep growing.
Build scalable web crawling infrastructure.
Built and operated the crawling infrastructure for Cenoskop — URL frontier with per-host rate limiting, distributed workers processing 30M+ product pages daily, and storage with deduplication. The system crawls 500+ sites on a rotating schedule.
We define the crawl surface and scale targets, build the frontier/workers/storage, and run a first crawl to validate. From €2,000.
Initial crawler infrastructure for a domain/scale tier.
Per-host rate limiting, robots.txt awareness, and backoff on errors — built into the frontier so crawls respect targets.
Whatever you use: I build workers as processes I can run under supervisor, systemd, or container workers — no lock-in.
Other problems I help with.
Relevant production work.
Find a job or offer your services.
Built the Go backend — API, matching logic and deployment.
Chat application with end-to-end encryption.
Built the backend and encryption layer for a secure chat application.
Remote security scanner with ping and certificate checking.
Built the scanning daemon and monitoring backend.
A Cleaner, Safer Internet. Take control of your DNS. Block ads, trackers, and malicious sites with customizable profiles.
Built the DNS resolution backend and profile management system.
Crawler monitoring millions of products. Processing 30M products daily and comparing them.
Built and maintain the full crawling pipeline — from scraping through Elasticsearch indexing to price comparison.
Price intelligence engine for the US market. Processing large-scale product data.
Built the data pipeline and crawling infrastructure for large-scale US product monitoring.
Describe the problem in a few lines — I'll look at it and tell you what I think.