Skip to content

Navigation Menu

Sign in
Sign up
@datacrawler-edu
datacrawler-edu
Follow

Edu Scraping Data datacrawler-edu

Block or report datacrawler-edu

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
datacrawler-edu /README.md

Hey, I'm Edu πŸ‘‹

I build web scraping infrastructure at scale.

I specialize in high-performance crawlers, data pipelines, and scraping systems that handle hundreds of millions of requests. Currently building tools for web technology detection and large-scale data extraction.


πŸ”§ What I work with

Rust β†’ High-perf crawlers, domain scanners, async networking
Python β†’ aiohttp, httpx, FastAPI, data processing
Data β†’ DuckDB, Parquet, Polars, Pandas
Infra β†’ Cloudflare Workers, Hetzner, Oracle Cloud ARM
Approach β†’ Pure HTTP, no browsers. Speed & scale first.

Rust Python DuckDB Cloudflare FastAPI

πŸš€ What I'm building

  • Domain Scanner β€” Scanning 200M+ domains for web technology detection (Rust)
  • Google Maps Scraper β€” Extracted hundreds of millions of records at ~100 req/s, no proxies
  • Tech Detection SaaS β€” BuiltWith/Wappalyzer alternative powered by DuckDB + Cloudflare Workers
  • Lead Generation Pipeline β€” Crawling β†’ email extraction β†’ SMTP verification, near-zero cost

πŸ“Š Numbers I'm proud of

Metric Value
Domains scanned 200M+
Google Maps records 100M+
Avg. request rate ~100 req/s per instance
Infrastructure cost ~5ドル/month

πŸ“ Writing & content

I write about web scraping, data engineering, and building lean infrastructure β€” in Spanish and English.

  • 🐦 X / Twitter β€” Scraping tips, war stories, and data drops
  • πŸ“– Blog β€” Coming soon

πŸ’‘ My philosophy

Browsers are a last resort. HTTP requests are king. Scale with smart architecture, not bigger servers. Near-zero infra cost is not a limitation β€” it's a design goal.


Open to collaborations on scraping tools, data products, and open source infra.

Popular repositories Loading

  1. datacrawler-edu datacrawler-edu Public

    Web scraping & data engineering at scale Building scraping infrastructure with Rust & Python Scraping the web, one billion requests at a time High-performance crawlers & data pipeline

  2. contratacion-publica-extremadura contratacion-publica-extremadura Public

    Base de datos de contrataciΓ³n pΓΊblica de Extremadura. Datos extraΓ­dos de fuentes oficiales, procesados y listos para reutilizar.

  3. bing-search-scraper-python bing-search-scraper-python Public

    Examples and sample data for extracting structured Bing SERP results through the Apify Bing Search Scraper.

  4. website-technology-lookup-python website-technology-lookup-python Public

    Examples and sample data for looking up CMS, hosting and website technology data through the Apify Website Technology Lookup Actor.

  5. linkedin-public-profile-scraper linkedin-public-profile-scraper Public

    Examples and sample data for scraping LinkedIn profiles through the Apify LinkedIn Public Profile Scraper.

  6. google-search-scraper-python google-search-scraper-python Public

    Examples and sample data for extracting structured Google organic SERP results through the Apify Google Search SERP Scraper.

AltStyle γ«γ‚ˆγ£γ¦ε€‰ζ›γ•γ‚ŒγŸγƒšγƒΌγ‚Έ (->γ‚ͺγƒͺγ‚ΈγƒŠγƒ«) /