Benchmarks
Three tables from the README, as measured, with every caveat and the exact command to run each one yourself.
Every speed claim netweir makes is a script in bench/ that you can run. The numbers below are copied from the README, where they were measured on an Apple M-series laptop. Your machine will give different numbers; the shape should hold. Two of the three run in CI, and the build fails if any library beats netweir.
Parsing and CSS extraction
Pulls every price, title and link out of a generated shop page.
| Library | 1 MB page | 10 MB page |
|---|---|---|
| netweir | 5.0 ms | 61 ms |
| selectolax | 5.3 ms | 64 ms |
| BeautifulSoup (lxml) | 196 ms | 2.3 s |
| lxml + cssselect | 347 ms | 82 s |
| parsel | 353 ms | 82 s |
Measured on an Apple M-series laptop with Python 3.14.
The lead over selectolax is small, because both parse with lexbor, and it comes from doing the extraction in Rust. On other machines the two can swap places by a few percent; on CI’s Linux runners they do.
Run it yourself, from a checkout of the repository:
uv run --group bench python bench/parse.pyXPath and find_all
Parses the same shop page and asks six XPath questions, or six Beautiful Soup ones.
| Library | 1 MB page | 3 MB page |
|---|---|---|
| netweir, XPath | 9.9 ms | 33 ms |
| lxml, XPath | 809 ms | 8.8 s |
| parsel, XPath | 814 ms | 8.8 s |
| netweir, find_all | 9.4 ms | 29 ms |
| Beautiful Soup, find_all | 170 ms | 515 ms |
Look at lxml across the row: three times the page took eleven times as long. netweir builds a flat index of the document on the first query, so //x is a scan over a few arrays, and its time grows with the page and no faster.
Both this benchmark and the parsing one run in CI, and the build fails if any library beats netweir.
Run it yourself, from a checkout of the repository:
uv run --group bench python bench/query.pyA whole crawl
One local server plays 100 sites of 100 pages, each answer 50 ms late, and every crawler fetches all 10,000 pages and pulls a title and price from each, with the same limits (100 requests in flight, 8 per site).
| Library | Pages a second | CPU per page | Peak memory |
|---|---|---|---|
| netweir | 1,770 | 0.22 ms | 68 MB |
| Scrapy 2.19 | 890 | 1.1 ms | 125 MB |
| httpx + selectolax | 220 | 3.2 ms | 190 MB |
| Scrapling 0.4 | 180 | 0.8 ms | 76 MB |
With 100 requests in flight and 50 ms per answer, 2,000 pages a second is the most any crawler could do here, so netweir is waiting on the server, not on itself.
Scrapling is held back by its HTTP session’s default of 10 connections, which it doesn’t let you change; that is how it ships.
Same laptop, Python 3.13. Run it once the others are installed.
Run it yourself, from a checkout of the repository:
uv run python bench/crawl.pySetting up to run them
The benchmarks live in the netweir repository and run against a development build. You need Rust, a C compiler, CMake and uv:
git clone --recurse-submodules https://github.com/netweir/netweir
cd netweir
uv sync --group dev
uv run maturin develop --uvThe parsing and query benchmarks install the libraries they compare against with --group bench. The crawl benchmark expects Scrapy, httpx, selectolax and Scrapling to be installed already.