Sign in

Open Source Web Scrapers for Lead Lists in 2026: $0 a Month Against Firecrawl’s $99, and the 4 Licences That Decide If You Can Sell the Output

scrapers-lead-lists-hero

The short answer. A lead list built on a self-hosted scraper costs $0 in licence fees against $99 a month for 100,000 pages on Firecrawl Standard, which is $1,188 a year avoided. The catch is not technical and it is not price. It is the licence: of the 19 scraper repositories we queried through GitHub’s own API on 9 September 2026, 4 carry terms that bind a company running them as a service inside a product it sells, and the single most starred project in the category is one of them.

Option Plan Cost per 1,000 pages Licence condition on selling
Self-hosted Scrapy, Crawl4AI or Scrapling any $0 in licence, plus your own compute None. BSD-3-Clause and Apache-2.0 ask for attribution, not for your source
Firecrawl, hosted Standard, $99 a month $0.99 None on the hosted plan. The self-hosted repository is AGPL-3.0
Zyte API, hosted $500 monthly commitment $0.06 to $7.68 None. A commercial service billed per successful response
The three shapes this decision actually takes. Prices read from each vendor’s own pricing page on 7 September 2026 and re-read on 9 September 2026. Licences from GitHub’s REST API on both dates. The full ladder, the full grid and the full licence audit are below.

This page is for a founder or a small sales team building a prospect list without paying a monthly crawler bill. It does not rank these tools on features. It prices them on their own published meters, audits every licence against the repository rather than against a listicle, and says plainly which of the 11 projects Google names for this query we cannot link because they have no review page on this site.

Before the tools: who publishes the page you are reading this on

We captured the results page for this query on 7 September 2026 and read the homepage of every domain on it on 9 September 2026. Of the 8 organic results, 5 are published by a company that sells a web scraping or web data product, and 0 is an independent test. The result at position 1 is not an article on a publisher’s site at all: it is a Markdown listicle committed into an unrelated GitHub repository. Google’s AI Mode leads its answer with a reference titled “13 Best Web Scraping Tools in 2026”, published by Firecrawl, which is one of the tools being compared.

Position Domain Who publishes it Sells a scraper
1 github.com A listicle copied into an unrelated repository (krishardee443/X-Register) No
2 rayobyte.com Rayobyte, a public web data gathering platform Yes
3 scrapfly.io Scrapfly, a web scraping API and cloud browser Yes
4 coolparse.com CoolParse, a no-code web scraping platform Yes
5 reddit.com A thread whose title begins “I built an open source library” No
6 olostep.com Olostep, web data infrastructure for AI agents Yes
7 thunderbit.com Thunderbit, an agentic web scraper Yes
8 reddit.com A thread whose title begins “My point-and-click web scraper is launching” No
Every domain’s own homepage title and description read on 9 September 2026. This is not a complaint about vendor content, which is often the most detailed writing in a category. It is the reason this article prices its own subject and shows the arithmetic.

Cost per 1,000 pages, which is the number nobody prints

A lead list is priced in pages fetched, not in seats. Every hosted scraper on this results page publishes a plan price and a page or credit allowance on the same card, so one division gives a comparable rate. Here is the whole Firecrawl ladder, computed from the vendor’s own machine-readable pricing document rather than from a rendered card.

Plan Price a month Pages included Cost per 1,000 pages Concurrency
Free $0 1,000 $0 2 concurrent requests
Hobby $19 5,000 $3.80 5 concurrent requests
Standard $99 100,000 $0.99 25 concurrent requests
Growth $399 500,000 $0.798 50 concurrent requests
Scale $749 1,000,000 $0.749 100 concurrent requests
Firecrawl, read from its own pricing document on 7 September 2026 and again on 9 September 2026. Every figure held exactly. The rate falls 5.1 times from Hobby to Scale, and the free tier grants 1,000 pages a month at $0 with no card.

Two things about that ladder matter more than the headline. The first is that a credit is not always a page: the same balance pays for a search at 2 per 10 results and for browser time at 2 per browser minute, so a list build that leans on search burns the allowance faster than the per-page rate suggests. The second is the rate limit, which is the real ceiling on a one-night list build: 10/min on Free against 500/min on Standard for the scrape, map and search endpoints.

Apify is the third meter on this page and it does not convert. It prices compute units rather than pages, at $0.20 on Free and Starter falling to $0.13 on Business at $999 a month, and the pages a compute unit buys depend on which Actor you run. We print the rate and refuse the conversion. Two things about Apify are worth a founder’s attention anyway: annual billing takes 10 percent off, and the company publishes Crawlee, one of the Apache-2.0 libraries in the licence audit below. The vendor selling the hosted option also maintains the free one.

Plan Price a month Platform credit included Per compute unit
Free $0 $5 $0.20
Starter $19 $19 $0.20
Scale $199 $199 $0.16
Business $999 $999 $0.13
Apify, read on 7 September 2026 and re-read on 9 September 2026 with monthly billing selected. Every row held. Add-ons on the same page: concurrent runs $5 per run, Actor RAM $1 per GB, datacenter proxy from $0.6 per IP. The store held 69,037 Actors on 7 September 2026 and 70,309 on 9 September 2026.

Zyte publishes a range of 268 times, and quotes the floor

The banner on Zyte’s pricing page reads “From $0.06 per 1,000 successful responses”. That rate is real and it is the cheapest cell of an eight-cell grid printed lower on the same page. Reaching it needs all three of a $500 monthly commitment, a tier-1 website and no browser rendering. On the same grid, a browser-rendered response on the hardest website tier at pay as you go is $16.08 per 1,000, which is 268 times the headline.

Minimum commitment a month HTTP response body, per 1,000 Browser rendered, per 1,000
pay as you go $0.13 to $1.27 $1.01 to $16.08
$100/month $0.10 to $0.95 $0.75 to $12.00
$200/month $0.08 to $0.76 $0.60 to $9.60
$500/month $0.06 to $0.61 $0.48 to $7.68
Zyte’s own grid, read in a rendered browser on 7 September 2026 and re-read on 9 September 2026. All eight cells held exactly. Websites are priced in five tiers and you are charged for what a site needs, which is an honest design; the point here is only that the headline is the floor of the range, not a price you should expect. Trial: $5 free credit, no commitment, 30 days.

Zyte no overage penalty above the commitment; you are charged at your current discounted rate. That is worth knowing, because the commitment tiers are the whole saving: the same HTTP response falls from $0.13 to $0.06 per 1,000 between pay as you go and the $500 tier, a 54 percent cut for committing.

Thunderbit meters rows, not pages, and its two billing modes are two different plans

Thunderbit is the one tool in this category whose meter has to be converted before it can sit beside the others, and its pricing page is where our own 7 September 2026 reading was wrong. That read concluded the Monthly control on the page did nothing, because it took the text of the whole document, which holds the monthly and the yearly cards at once. Operated in a browser on 9 September 2026 the control works, and it swaps in genuinely different plans.

Thunderbit pricing page in both billing states, with the annual cost per credit computed under each
Both toggle states captured in one browser session on 9 September 2026. The Monthly view sells 500 credits a month for $15; the Yearly view sells 5,000 credits a year for $9 a month. Per credit that is $0.03 against $0.0216.
Billing Plan Price shown Cost a year Credits a year Per 1,000 rows
Monthly Starter $15 a month $180 6,000 $30.00
Monthly Pro $38 a month $456 36,000 $12.6667
Yearly Starter $9 a month $108 5,000 $21.60
Yearly Pro $16.50 a month $198 30,000 $6.60
Computed from the cards on both toggle states, read 9 September 2026. The yearly cards grant credits a year and the monthly cards grant credits a month, so the only comparable figure is the annual cost divided by the annual allowance.

Switching Starter to yearly billing cuts the price from $180 to $108 a year and cuts the allowance too, from 6,000 credits a year to 5,000. The yearly plan is not the monthly plan discounted, it is a smaller plan at a lower unit rate. Per credit that is $0.03 against $0.0216, a 28 percent cut, against a banner that reads “Save 20% on a yearly subscription”. On Pro the same comparison is $0.0127 against $0.0066.

Now the conversion that matters for a lead list. Thunderbit’s own FAQ defines the meter: “1 row of data = 1 credit”, 2 credits a row when subpage scraping is on, and 30 credits for each successful Personal Data Enrichment query. At that rate the Starter yearly plan’s 5,000 credits buy 166 enriched contacts a year, which is $0.65 each. Pro yearly buys 1,000 at $0.20 each. Our lead generation pricing article puts an Apollo credit that reveals one email at $0.0248, so these are different meters doing different work and the numbers are not interchangeable. Print both and let the reader decide which unit is the one they are buying.

The licence audit: 19 repositories, straight from GitHub’s API

Every repository below was queried at api.github.com on 7 September 2026 and again on 9 September 2026. The identifier column is GitHub’s own detection, not our reading of the file. Between the two dates 0 identifiers changed, so the star counts and push dates below are the ones from 9 September 2026.

Project Licence identifier Stars Last push Language
Firecrawl
firecrawl/firecrawl
AGPL-3.0 178,208 2026-09-09 TypeScript
Playwright
microsoft/playwright
Apache-2.0 95,860 2026-09-09 TypeScript
Puppeteer
puppeteer/puppeteer
Apache-2.0 95,558 2026-09-09 TypeScript
Crawl4AI
unclecode/crawl4ai
Apache-2.0 82,029 2026-09-08 Python
MinerU
opendatalab/MinerU
NOASSERTION 79,563 2026-09-09 Python
Scrapling
D4Vinci/Scrapling
BSD-3-Clause 79,549 2026-09-04 Python
Scrapy
scrapy/scrapy
BSD-3-Clause 64,257 2026-09-09 Python
Lightpanda
lightpanda-io/browser
AGPL-3.0 35,195 2026-09-09 Zig
Selenium
SeleniumHQ/selenium
Apache-2.0 34,481 2026-09-09 Java
ScrapeGraphAI
ScrapeGraphAI/Scrapegraph-ai
MIT 30,759 2026-09-07 Python
Cheerio
cheeriojs/cheerio
MIT 30,481 2026-09-09 TypeScript
Crawlee (JS)
apify/crawlee
Apache-2.0 25,706 2026-09-09 TypeScript
Colly
gocolly/colly
Apache-2.0 25,503 2026-09-02 Go
Katana
projectdiscovery/katana
MIT 17,422 2026-09-07 Go
Maxun
getmaxun/maxun
AGPL-3.0 17,397 2026-09-09 TypeScript
Browserless
browserless/browserless
NOASSERTION 13,677 2026-09-09 TypeScript
Crawlee (Python)
apify/crawlee-python
Apache-2.0 9,497 2026-09-09 Python
Heritrix
internetarchive/heritrix3
NOASSERTION 3,316 2026-09-09 Java
Apache Nutch
apache/nutch
Apache-2.0 3,289 2026-09-08 Java
GitHub REST API, read 9 September 2026. The identifiers in bold are the ones that are not a plain permissive licence and that a company should read before shipping. MinerU is in this table because it appears in our own Web Scraping category and is a document parser rather than a web scraper, which is a catalogue problem we name below rather than hide.

What AGPL-3.0 asks of a lead generation product

3 of the scrapers above are AGPL-3.0, and one of them, Firecrawl at 178,208 stars, is the most starred project in this table. The condition is short enough to state plainly. Under GPL-family terms you owe source only when you distribute software. AGPL-3.0 adds a clause for the case that did not exist when those terms were written: if users interact with a modified version over a network, the licence asks you to offer those users the corresponding source of your version. A lead generation product is exactly that case, because your customers reach the scraper through your service rather than by installing it.

Project Stars What the condition means for a product you sell
Firecrawl 178,208 Running it as the engine behind a hosted lead tool puts your version’s source in scope. Running it on your own machine to build your own lists does not.
Lightpanda 35,195 Running it as the engine behind a hosted lead tool puts your version’s source in scope. Running it on your own machine to build your own lists does not.
Maxun 17,397 Running it as the engine behind a hosted lead tool puts your version’s source in scope. Running it on your own machine to build your own lists does not.
The three AGPL-3.0 scrapers in the audit, stars read 9 September 2026. This is a reading of the licence text, not legal advice, and a lawyer should see it before you build a business on any of the three.

Running an AGPL scraper on your own machine to build your own lists is not the case the clause addresses at all, and that is what most readers of this page are doing. The distinction is between using the tool and offering it to other people as a service. Every page ranking for this query writes “open source, free” and none of them separates the two.

“Other” on a GitHub page means three unrelated things, and we read all three

3 repositories in the audit return no detected licence identifier, which GitHub renders on the repository page as the word “Other”. That reads like a licence rather than like an absence of a detection. Our open source alternatives audit found the same badge on n8n and Open WebUI, where it concealed a real restriction, so the rule that came out of that build was to read the file by hand every time. Read on 9 September 2026, the three here are not the same finding at all.

Project Repository What the LICENSE file declares Verdict
Browserless browserless/browserless SPDX-License-Identifier: SSPL-1.0 OR Browserless Commercial License A restriction the badge hides
Heritrix internetarchive/heritrix3 Apache License, Version 2.0 More permissive than the badge suggests
MinerU opendatalab/MinerU Apache License 2.0 with additional terms Permissive below a published threshold
Each LICENSE file fetched from the repository’s default branch and read in full on 9 September 2026. Links to all three files are in Sources.

Browserless

The file opens “This work is dual-licensed under the MongoDB Server-Side Public License”. A closed-source commercial use needs a purchased commercial licence. The repository says so in its own LICENSE file, and the GitHub badge says “Other”.

Heritrix

The file opens “Licensed under the Apache License, Version 2.0”. Plain Apache-2.0, one of the most permissive licences in this table. GitHub reports no detected licence because the file holds the short notice rather than the full text. The badge understates it.

MinerU

The file opens “MinerU is licensed under Apache License 2.0 and is subject to the additional terms”. Apache-2.0 plus a commercial threshold and a mandatory attribution duty. Below the thresholds it is usable commercially; the attribution duty applies at any size, and breaching either terminates the licence automatically. The two published thresholds are monthly active users (MAU) exceed 100 million and total monthly revenue exceeds USD 20 million. If you provide online services based on MinerU you must indicate that MinerU is used, in the interface or in public documentation.

The lesson is not that “Other” is bad. Heritrix is one of the most permissive entries in a table of 19 and its badge undersells it, while Browserless is one of the most restrictive and its badge hides that. The badge carries no information either way, so the only safe move is to open the file. It takes about twenty seconds.

BeautifulSoup: the GitHub repository everyone stars has been dead since 2022

BeautifulSoup is the second project Google’s AI Overview names for this query, and it has no canonical GitHub repository. Searching GitHub for it lands on wention/BeautifulSoup4, a mirror that GitHub reports as archived since 2022-11-08, with 223 stars and no detected licence. Anyone doing the obvious thing, starring the top GitHub result to come back to it, is starring a mirror that stopped 3 years ago.

Where What is there Last change Licence
GitHub, wention/BeautifulSoup4 An archived mirror, 223 stars 2022-11-08 NOASSERTION
Launchpad, the project’s own git Active development, 9 open merge proposals 2026-08-02 on master See the PyPI record
PyPI, the released package Version 4.15.0 2026-06-07 MIT License
All three read on 9 September 2026. The project is alive; the place most people look for it is not. Uploads to the Launchpad repository are restricted to the project’s author.

The eight we would actually reach for, ordered by licence

Ordered by whether the licence puts a further condition on selling the output of a hosted service, then by stars inside each group. That ordering is deliberate: on a page about building a product, a licence that constrains the product is a more useful sort key than popularity. All figures read 9 September 2026.

1. Playwright

Playwright, microsoft/playwright. Apache-2.0, 95,860 stars, last pushed 2026-09-09. Drives a real Chromium, Firefox or WebKit, so it reaches a list behind a login or an infinite scroll. The licence puts no further condition on selling the output of a hosted service. We have no review page for this project on this site yet, so it is named here without a link rather than pointed at a near-miss.

2. Crawl4AI

Crawl4AI, unclecode/crawl4ai. Apache-2.0, 82,029 stars, last pushed 2026-09-08. Crawls and returns Markdown built for a language model, which is the shape an enrichment step wants. The licence puts no further condition on selling the output of a hosted service.

3. Scrapling

Scrapling, D4Vinci/Scrapling. BSD-3-Clause, 79,549 stars, last pushed 2026-09-04. Adaptive selectors that survive a layout change, which is what breaks a standing list build. The licence puts no further condition on selling the output of a hosted service. We have no review page for this project on this site yet, so it is named here without a link rather than pointed at a near-miss.

4. Scrapy

Scrapy, scrapy/scrapy. BSD-3-Clause, 64,257 stars, last pushed 2026-09-09. The asynchronous Python crawling framework, with request queueing and pipelines already solved. The licence puts no further condition on selling the output of a hosted service. We have no review page for this project on this site yet, so it is named here without a link rather than pointed at a near-miss.

5. ScrapeGraphAI

ScrapeGraphAI, ScrapeGraphAI/Scrapegraph-ai. MIT, 30,759 stars, last pushed 2026-09-07. Extraction driven by a prompt rather than a CSS selector, the one Google’s AI Overview names for AI-assisted parsing. The licence puts no further condition on selling the output of a hosted service.

6. Crawlee

Crawlee, apify/crawlee. Apache-2.0, 25,706 stars, last pushed 2026-09-09. Queue management and automatic switching between HTTP and a headless browser. Published by Apify, whose own compute-unit pricing is in the table above. The licence puts no further condition on selling the output of a hosted service. We have no review page for this project on this site yet, so it is named here without a link rather than pointed at a near-miss.

7. Firecrawl

Firecrawl, firecrawl/firecrawl. AGPL-3.0, 178,208 stars, last pushed 2026-09-09. Crawl and scrape behind one API, self-hostable, and the same product sells a hosted plan from $19 a month. AGPL-3.0: run it as a network service inside a product you sell and the licence asks you to offer that service’s corresponding source to its users.

8. Lightpanda

Lightpanda, lightpanda-io/browser. AGPL-3.0, 35,195 stars, last pushed 2026-09-09. A headless browser written for automation rather than for a person, so a render costs less to run. AGPL-3.0: run it as a network service inside a product you sell and the licence asks you to offer that service’s corresponding source to its users.

Is web scraping illegal?

This is the question at position 1 in People Also Ask for this query, so it deserves a straight answer about what we can and cannot tell you. We are not lawyers and this is not legal advice. What we can source, and have, is what each tool’s licence permits, because that is a document you can read in a minute. Three other things decide the rest, and none of them is a property of the scraper you pick.

  • The site’s own terms and robots file. These are published by the site you are fetching, they differ per site, and they change without notice. Read the ones for the sites on your list.
  • What you collect. A lead list is personal data by definition, because a name, a work email and an employer identify a person. Data protection rules apply to holding and using it whether you scraped it, bought it or were given it.
  • Where you and the person are. The rules differ by jurisdiction, and a list of European contacts built from a server in the United States engages more than one set.

What changes with your tool choice is only the licence column. That is why this article audits licences in detail and stops at the edge of the other three rather than summarising case law we have not read in full.

How we tested, and what our catalogue does not cover

  • Licences and repository metrics come from GitHub’s REST API, queried on 7 September 2026 and again on 9 September 2026. Nothing in the licence table is copied from a repository page or from another article.
  • Prices come from each vendor’s own pricing page, read in a rendered browser on 7 September 2026 and re-read on 9 September 2026. Firecrawl publishes a machine-readable pricing document and that is what we read, so those figures are the vendor’s own strings.
  • Where a licence file returns no identifier we fetch the file and read it in full rather than reporting the badge.
  • We take no payment for a position on this page, and the ordering is by licence and stars, not by anything we earn.

Our catalogue holds 33 published tools under Web Scraping, of which 7 are open source repositories, and 9 of them are not web scrapers at all: OpenDataLoader PDF, MinerU, aiPDF, PDF.ai, Parseur, Parsio, Procys, Monkt and Leexi. Eight are document parsers and one is a meeting recorder. That is why the roster on this page is built from the licence audit rather than from a category export. 8 of the 11 projects Google names first for this query have no review page on this site: Scrapy, BeautifulSoup, Crawlee, Playwright, Puppeteer, Selenium, Cheerio, Scrapling. We name them without a link and we are adding those pages.

What we could not verify, and therefore did not print

  • Throughput or pages per second for any tool; not benchmarked on identical hardware
  • Proxy and residential IP costs beyond Zyte’s and Apify’s published rates
  • Block rates
  • Any claim about which scraper defeats which anti-bot system
  • Legal conclusions; the article states what licences and terms say and stops there

The first of those is the one readers will miss most. Pages per second is the figure that would let you convert a hosted rate into a self-hosted one, and we have not benchmarked these tools on identical hardware, so we do not publish it. What we can say without a benchmark is the licence position and the published rate, and those are what the tables above hold.

Questions people ask about this

Is web scraping illegal?

No single answer covers it, and any page that gives you one is overselling. Three separate things decide it and none is a property of the scraper you choose: the terms and robots file of the site you fetch, the data protection rules that apply because a lead list is personal data, and the jurisdictions you and the person are in. What your tool choice does decide is the licence, and that we have audited: 13 of the 19 repositories in the table above carry a plain permissive licence. We are not lawyers and this is not legal advice.

What is the best free open-source web scraper?

3 of the 19 we audited, depending on the shape of the pages. Scrapy for Python or Crawlee for JavaScript, both permissively licensed and both pushed within days of 9 September 2026. If the pages you need render their content with JavaScript, Playwright at 95,860 stars drives a real browser and is Apache-2.0. All three are $0 in licence and none of them constrains a product you sell.

Can ChatGPT do web scraping?

It can read a page you give it and it can write scraper code for you, and those are different jobs from running a crawl. A lead list needs a queue, retries, rate limiting and a place to put the rows, which is what the frameworks in the table above provide. The pattern that works is a language model writing and repairing the extraction step inside one of those frameworks, which is what Crawl4AI and ScrapeGraphAI package.

Is scraping with BeautifulSoup legal?

The licence question is settled: the released package on PyPI carries the MIT License at version 4.15.0, read 9 September 2026, which puts no restriction on commercial use. Everything else about legality is decided by the site you fetch and the data you keep, not by the parser. Note also that BeautifulSoup parses HTML you already have; the fetching is done by something else, and that is where a site’s terms bite.

Can I use an AGPL scraper in a product I sell?

You can sell products that use AGPL-3.0 software. The condition is about source, not about money: if your users interact with your modified version over a network, the licence asks you to offer them the corresponding source of your version. Building your own lists on your own machine is not that case. Running Firecrawl or Lightpanda as the engine behind a hosted lead tool is. Have a lawyer read it before you commit.

What does 1,000 pages actually cost on a hosted scraper?

On Firecrawl, $0.99 on the $99 Standard plan and $0.749 at Scale, computed from the vendor’s own pricing document on 9 September 2026. On Zyte it depends on three variables: $0.06 for a plain HTTP response on a tier-1 site with a $500 commitment, up to $16.08 for a browser-rendered response on the hardest tier at pay as you go. Self-hosted is $0 in licence plus whatever your machine costs.

Do I need a proxy, and what does one cost?

For a small list off public pages, usually not. Where a vendor publishes a rate we can quote it: Apify lists a datacenter proxy from $0.6 per IP, and Zyte bundles datacenter, residential and rendering into the per-response price in the grid above. Beyond those two published rates we cut proxy pricing from this article rather than estimate it, because the market prices by gigabyte, by IP and by success rate and the three do not convert.

Scrapy or Crawl4AI for a first lead list?

Scrapy if the pages are static and the job is volume: it is BSD-3-Clause, 64,257 stars, and request queueing and pipelines are already solved. Crawl4AI if the next step is a language model reading what you fetched: it is Apache-2.0, 82,029 stars, and it returns Markdown shaped for that. Neither costs anything in licence and neither constrains a product you sell. Both read 9 September 2026.

Related reading on this site

Sources

  1. Firecrawl pricing, machine-readable: firecrawl.dev/pricing.md, read in a rendered browser on 7 September 2026 and again on 9 September 2026.
  2. Zyte pricing: zyte.com/pricing, read in a rendered browser on 7 September 2026 and again on 9 September 2026.
  3. Apify pricing: apify.com/pricing, read in a rendered browser on 7 September 2026 and again on 9 September 2026.
  4. Thunderbit pricing: thunderbit.com/pricing, read in a rendered browser on 7 September 2026 and again on 9 September 2026.
  5. Browserless LICENSE file: raw.githubusercontent.com/…/LICENSE, fetched and read in full on 9 September 2026.
  6. Heritrix LICENSE file: raw.githubusercontent.com/…/LICENSE, fetched and read in full on 9 September 2026.
  7. MinerU LICENSE file: raw.githubusercontent.com/…/LICENSE.md, fetched and read in full on 9 September 2026.
  8. BeautifulSoup development: code.launchpad.net/beautifulsoup, and the released package at pypi.org/project/beautifulsoup4, both read 9 September 2026.
  9. Licences, stars and push dates for all 19 repositories: GitHub REST API, queried 7 September 2026 and 9 September 2026.
  10. Google AI Overview, AI Mode, People Also Ask and the organic results for “open source web scrapers”, captured 7 September 2026. Publisher identities read 9 September 2026.
  11. Reddit thread, An open-source tool for extracting data visually, listed in the discussions module on 7 September 2026. Cited, not quoted: reddit.com returns HTTP 403 to an anonymous read.
  12. Reddit thread, Web Scraper - Open-source site/API to scrape websites, listed in the discussions module on 7 September 2026. Cited, not quoted: reddit.com returns HTTP 403 to an anonymous read.

Comments

Sign in to join the discussion.

Login to comment