netweir 0.1.0
The first release of netweir: Chrome-identical fetching, a Rust parser with Scrapy, XPath and Beautiful Soup queries, crawling that heals itself, and a Chrome driver. What it does today, and what comes next.
netweir 0.1.0 is on PyPI. It’s the first release of a web scraper I’ve been building to be three things at once: ultra-fast, self-healing and undetectable. Rust does the work; you write Python.
pip install netweir
There are wheels for Linux, macOS and Windows, on Python 3.10 and up, free-threaded 3.14 included, so there’s nothing to compile. BoringSSL and the HTML parser are inside the wheel.
What it does today
Fetching that looks like a browser. A request from netweir is what Chrome 154 sends, from the TLS ClientHello to the order of the headers, over HTTP/2 and HTTP/1.1, with cookies and redirects handled the way Chrome handles them. Firefox 156 and Safari 27 are there too, as profile="firefox" and profile="safari". Chrome was recorded doing four navigations against a local server, and a test has netweir do the same four and fails if any request differs. One known difference remains, and the README says what it is.
A parser that keeps pace with selectolax, and queries that beat the rest. Pages are parsed by lexbor, the C parser selectolax uses; the extraction then runs in Rust, so Python gets a finished list of strings. CSS with Scrapy’s ::text and ::attr(), all of XPath 1.0 with parsel’s extras, and Beautiful Soup’s find_all family all work on the same page. The benchmarks have the numbers and the commands to rerun them.
Crawling. Spiders with parse callbacks, Request with callback, errback, meta and priority, pipelines, and items written to JSON Lines, CSV or Parquet in Rust. robots.txt and TDMRep are obeyed by default, requests to each site are paced by how quickly it answers, and duplicates are dropped.
Declarative spiders. Describe the item and the links, and the whole crawl runs in Rust without Python running for each page. For spiders whose own Python is the slow part, workers=4 runs callbacks in worker processes; on a heavy spider that was 3.1 times faster, with the same output.
Self-healing. Block pages from Cloudflare, Akamai, DataDome, HUMAN, Kasada, Imperva and AWS WAF are recognised and recovered from. A crawl with a checkpoint resumes after a crash with no page fetched twice and no item written twice. A selector you name with track= finds its element again after a redesign and tells you it had to.
A Chrome driver. For pages that need JavaScript: click, type and wait like Playwright, then read the page with netweir’s selectors. In a crawl, browser="on_block" gives a request that keeps getting blocked one go in Chrome and hands the cookies it earns back to the HTTP client. netweir install chrome fetches the Chrome it’s tuned for.
Where to start
The getting started guide goes from install to a crawl that writes a file in about ten minutes. There’s a page for each part of netweir, and every setting with its default.
What’s next
- Firefox in the driver. The HTTP client already speaks Firefox; the browser driver only drives Chrome. That changes once there’s a Firefox that doesn’t announce it’s automated.
- POST requests and forms. 0.1.0 fetches with GET. Submitting forms from the HTTP client is next.
- Sitemaps. Starting a crawl from a site’s sitemap, instead of from a list of URLs.
Issues, fixes and new browser profiles are welcome on GitHub. netweir is AGPL-3.0; if that doesn’t work for your company, a commercial licence is available.