Codexworker

Large-Scale Web Crawling

At scale, crawling is an operations problem: frontier distribution, consistent dedup, backfill management, error recovery and storage that doesn't collapse under its own weight.

I build and tune crawling systems that reliably process millions of URLs across many hosts.

I'm a backend developer with 10+ years of experience across PHP, Go, MySQL and Elasticsearch. Recent work includes Cenoskop (30M+ products daily), Niftycent (US price intelligence), and facha.sk (Go backend). I build, fix and optimize backend systems that handle real data and real traffic.

What I Can Help With

Scale crawling to millions of URLs across many hosts.

Typical Problems

frontier distribution deduplication at scale backfill and recrawl scheduling error recovery and retries distributed storage host politeness at scale

Relevant Experience

Scaled Cenoskop's crawling system from 5M to 30M+ products daily by adding distributed workers, tuning the URL frontier, and implementing backfill scheduling. The system crawls across millions of URLs with per-host politeness.

What Is Included

scale planning
frontier + worker tuning
storage/dedup review
politeness tuning
monitoring and alerting

How It Works

We audit your current crawl pipeline, then tune the frontier/workers/storage to your scale. From €3,000 for a high-volume tuning engagement.

Starting from €3,000

High-volume crawl tuning and scale planning.

Frequently Asked Questions

What's your crawl throughput experience?

Built and operated production crawlers processing 30M+ products per day across millions of URLs — frontier, workers and storage tuned for throughput.

Related Services

Other problems I help with.

Large-Scale Web Crawler Development

Build scalable web crawling infrastructure.

Details →

Custom Web Scraping Development

Build custom web scrapers that run reliably.

Details →

Custom Price Monitoring Systems

Build price monitoring infrastructure for products/sellers.

Details →

Product Data Extraction

Extract and normalize product data from external sources.

Details →

Systems I've Built

Relevant production work.

Price Intelligence Engine

Cenoskop / Levnobot

Crawler monitoring millions of products. Processing 30M products daily and comparing them.

Built and maintain the full crawling pipeline — from scraping through Elasticsearch indexing to price comparison.

Data Aggregation Scraping ElasticSearch
US Price Intelligence

Niftycent

Price intelligence engine for the US market. Processing large-scale product data.

Built the data pipeline and crawling infrastructure for large-scale US product monitoring.

Data Aggregation Scraping Big Data

Need Help With This?

Describe the problem in a few lines — I'll look at it and tell you what I think.

Discuss large-scale crawling