Codexworker

Large-Scale Web Crawler Development

I build crawling infrastructure that scales: URL frontier management, politeness, retry/backoff, distributed workers and a place to store and queue what's been crawled. Not a one-off scraper — infrastructure you can keep growing.

I'm a backend developer with 10+ years of experience across PHP, Go, MySQL and Elasticsearch. Recent work includes Cenoskop (30M+ products daily), Niftycent (US price intelligence), and facha.sk (Go backend). I build, fix and optimize backend systems that handle real data and real traffic.

What I Can Help With

Build scalable web crawling infrastructure.

Typical Problems

URL frontier management politeness and rate limiting retry/backoff strategy distributed crawling storage and deduplication monitoring and failure handling

Relevant Experience

Built and operated the crawling infrastructure for Cenoskop — URL frontier with per-host rate limiting, distributed workers processing 30M+ product pages daily, and storage with deduplication. The system crawls 500+ sites on a rotating schedule.

What Is Included

crawl frontier design
worker architecture
queue/storage integration
politeness + retry
monitoring hooks

How It Works

We define the crawl surface and scale targets, build the frontier/workers/storage, and run a first crawl to validate. From €2,000.

Starting from €2,000

Initial crawler infrastructure for a domain/scale tier.

Frequently Asked Questions

How do you stay polite to target sites?

Per-host rate limiting, robots.txt awareness, and backoff on errors — built into the frontier so crawls respect targets.

What runs the workers?

Whatever you use: I build workers as processes I can run under supervisor, systemd, or container workers — no lock-in.

Related Services

Other problems I help with.

Custom Web Scraping Development

Build custom web scrapers that run reliably.

Details →

Large-Scale Web Crawling

Scale crawling to millions of URLs across many hosts.

Details →

Custom Price Monitoring Systems

Build price monitoring infrastructure for products/sellers.

Details →

Custom Data Pipeline Development

Build ETL/import/transform/export pipelines for large datasets.

Details →

Systems I've Built

Relevant production work.

Job & Services Marketplace

facha.sk

Find a job or offer your services.

Built the Go backend — API, matching logic and deployment.

Job Marketplace Services Go
Secure Communication

Paidshield

Chat application with end-to-end encryption.

Built the backend and encryption layer for a secure chat application.

End-to-End Encryption Real-time Communication Security
Security Scanner

remotedaemon.com

Remote security scanner with ping and certificate checking.

Built the scanning daemon and monitoring backend.

Security Ping Certificate
DNS Security

Synopsee

A Cleaner, Safer Internet. Take control of your DNS. Block ads, trackers, and malicious sites with customizable profiles.

Built the DNS resolution backend and profile management system.

DNS Security Privacy
Price Intelligence Engine

Cenoskop / Levnobot

Crawler monitoring millions of products. Processing 30M products daily and comparing them.

Built and maintain the full crawling pipeline — from scraping through Elasticsearch indexing to price comparison.

Data Aggregation Scraping ElasticSearch
US Price Intelligence

Niftycent

Price intelligence engine for the US market. Processing large-scale product data.

Built the data pipeline and crawling infrastructure for large-scale US product monitoring.

Data Aggregation Scraping Big Data

Need Help With This?

Describe the problem in a few lines — I'll look at it and tell you what I think.

Discuss a crawler project