Brooke Wright
Brooke Wright · @wright_mode
FREE GUIDE

Collect the data that's blocked

The install, the prompt you hand Claude Code, four use cases worth your time — and an honest section on where the line is, because this one needs it.

3 commands No code to write Ethics included Built by Brooke
Something went wrong. Please try again.

No spam. Unsubscribe anytime.

Install it, the prompt you hand Claude Code, and the ground rules before you point it at anything Jump to the setup → Join the Membership
Wright Mode — Free Guide

The scraper that walks past Cloudflare

Scrapling is a free, open-source Python scraper with 78,900 GitHub stars. Hand it to Claude Code or Codex and your agent collects data from sites that block everything else. Install it in three commands — then read the ground rules, because they matter more than the install.

3 commands to install Claude Code drives it No Python to write Ethics section included Free and open source

What's inside

Read the ground rules before you point this at anything

Scrapling walks past blocks that exist for a reason. There's a section further down called Before you scrape anything — robots.txt, rate limits, terms of service, personal data. It is the most important part of this page and it takes two minutes. Don't skip it because the install worked.

Three commands. You need Python already on your machine — if python3 --version returns a number in your terminal, you're set.

1

Install Scrapling with its browsers

The first line gets the library plus the fetchers. The second downloads the browser engines it drives — that's the part that gets past Cloudflare, and it takes a few minutes.

pip install "scrapling[fetchers]" scrapling install
2

Add the AI extras

This is what gives you the MCP server, so Claude Code or Codex can drive the scraping itself instead of you writing Python.

pip install "scrapling[ai]"
3

Hook it up to Claude Code

One line. Then check it landed with claude mcp list, or type /mcp inside a session — you want scrapling · connected.

claude mcp add scrapling -- scrapling-mcp

If that errors with "command not found"

Your shell can't see where pip put it. Run which scrapling-mcp, copy the full path it prints, and use that instead: claude mcp add scrapling -- /full/path/to/scrapling-mcp.

Codex instead of Claude Code

Same server. Run codex mcp add and point it at the scrapling-mcp command, or add it to ~/.codex/config.toml yourself. Confirm with /mcp in a session.

Undo it

claude mcp remove scrapling unhooks it. pip uninstall scrapling removes the library. The downloaded browsers live in your browser cache folder and can stay — other tools use them too.


You don't write Python. You describe the job and the agent drives Scrapling through the MCP tools. This is the one I'd start with — fill in the two brackets and send.

Use the scrapling MCP server to collect data from this page: [URL] What I need out of it: [the fields you want — product name, price, date, whatever] How to do it: 1. Check the site's robots.txt first and tell me what it allows. If it disallows the path I've given you, stop and tell me instead of working around it. 2. Start with the plain HTTP request tool. Only escalate to a browser fetch, and then to a stealth fetch, if the simpler one genuinely fails. Tell me which one worked. 3. Pull only the fields I listed. Do not collect names, emails, phone numbers or anything else identifying a person unless I have explicitly asked for it. 4. Keep it to one page for now so I can check the output before we do more. 5. Give me the result as a CSV I can open, plus a one-line note on anything that looked unreliable. If the site blocks you at every level, say so plainly. Don't loop, don't hammer it, and don't invent data to fill the gaps.

Why the numbered rules are in there

Rule 1 makes it check permission before it acts. Rule 2 stops it reaching for the stealth browser on a site that would have answered a normal request — which is both politer and about ten times faster. Rule 3 keeps personal data out of your spreadsheet by default. Rule 5 is the one that stops a confident agent quietly making things up.

Once one page works, this is how you scale it without being a menace:

That worked. Now run the same extraction across these pages: [list of URLs, or the pattern] Rules for the batch: - Wait 2-3 seconds between requests. I would rather this take ten minutes than get the site's attention. - Stop immediately and tell me if you start getting 429s, 403s or captchas — that is the site asking you to back off, and we listen. - Cap it at [number] pages this run. - Same fields, same CSV, one row per page, with the source URL in a column so I can spot-check anything that looks wrong.

A one-off, no-code version

For a single page you don't even need the agent. Install the shell extra with pip install "scrapling[shell]", then run the command below — it turns any page into clean markdown you can read or feed to an AI. End the filename in .txt for plain text or .html for the raw HTML instead.

scrapling extract get 'https://example.com' content.md

If that page is behind Cloudflare

Swap get for stealthy-fetch and add --solve-cloudflare. Same command shape, slower, and it opens a real browser to do it.


Scrapling's own README says it's "provided for educational and research purposes only" and tells you to respect terms of service and robots.txt. That's the author covering himself — but it's also genuinely the line. This tool removes the technical barrier. It does not remove the legal or ethical one, and the two were never the same thing.

Here's the honest version, for someone running a business rather than a research lab.

Generally fine

  • Public pages with no login, at a human-ish pace
  • Facts: prices, product names, stock status, published dates, public job titles
  • Your own sites, listings and profiles
  • Sites whose robots.txt allows the path you're pulling
  • Anything you'd be comfortable telling the site owner you did

Don't

  • Anything behind a login — that's a contract you agreed to, and breaking it is a different category of problem
  • Personal data: names, emails, phone numbers, addresses. Scraped email lists are a fast route to a privacy complaint
  • Republishing someone's content as your own. Facts aren't copyright; their words and photos are
  • Hammering a small business's site. Hundreds of requests a minute is an attack, whatever you meant by it
  • Ignoring a 429 or a captcha and forcing your way through anyway

The four checks, every single time

1. robots.txt. Type the site's address then /robots.txt in your browser. It's a plain text file saying which paths the owner is happy for bots to touch. It isn't law, but ignoring it after reading it is the difference between an accident and a decision.

2. Terms of service. Search their terms for "scrape", "crawl", "automated" or "bot". If it's forbidden and you're logged in, you're breaching a contract you signed. Logged out, it's murkier — but you now know their position.

3. Rate limit yourself. A few seconds between requests. If you get a 429, a 403 or a captcha, that's the site saying stop. Stop. The whole reason blocks exist is that someone's server is being hurt.

4. Personal data. In Australia the Privacy Act and the Australian Privacy Principles cover personal information regardless of how you got it — "it was on a public website" is not a defence. EU residents bring GDPR with them. Collect facts about businesses, not details about people, and you sidestep nearly all of it.

The captcha question, answered straight

Scrapling can solve Cloudflare challenges. A captcha is the site's owner saying "I don't want automated traffic here." Walking past it isn't a hack — it's you deciding your convenience beats their stated wish. Sometimes that's genuinely defensible: your own site, a client's site with permission, a public price you're allowed to see anyway. Often it isn't. Ask yourself whether you'd send the site owner a screenshot of what you're doing. If the answer's no, that's your answer.

Three things that make this a non-issue

Check for an official API first — most big sites have one, and it's faster, allowed, and doesn't break when they change their layout. Scrape facts about companies rather than details about people. And keep a note of what you pulled, from where, on what date, so if anyone ever asks you have an answer ready.

Not legal advice

I'm not a lawyer and this isn't advice. If you're building scraping into a product, or pulling anything at scale, or touching personal data at all, that's a conversation with an actual solicitor — not a page on my website.


All four stay on the right side of the section above — public facts about businesses, not details about people.

1

Competitor pricing, tracked over time

Pull the price and what's included from five competitors' pages, once a week, into one sheet. Three months in you can see who's discounting, who's quietly raised, and whether you're the cheapest without meaning to be. Prices on a public page are facts — this is the cleanest use there is.

2

Reviews, read at volume

Pull the review text (not the reviewers' names) for a product or category, hand the file to Claude, and ask what people complain about most. It's the fastest market research going — you're reading your customers' actual words instead of guessing at their objections.

3

Watching your own listings

Your Google Business listing, your product pages, your directory entries. Check weekly that hours, prices and links are still right and nothing's gone stale or broken. Your own data, no ethics question at all, and it catches the thing you'd otherwise find out about from a customer.

4

Turning a site into something AI can read

Scrapling converts any page to clean markdown with the prompt-injection junk stripped out. Point it at a documentation site or a long report and you get a file you can hand to Claude and ask questions of. This is the use that has nothing to do with bypassing anything.

The pattern under all four

Scrapling gets the data. Claude does the thinking. Neither is much use alone — a spreadsheet of scraped prices sits there doing nothing, and an agent with no data just makes plausible-sounding numbers up.


I went and read the benchmark

The repo doesn't publish a 700× number. What it publishes is a parsing benchmark: pulling text out of 5,000 nested HTML elements that are already sitting in memory. Scrapling does it in 1.99ms. BeautifulSoup with lxml takes 1,562ms. That's the ~785× that gets rounded down to "700 times faster" on the internet.

Which is a real result, and also not the result most people think they're hearing. Here's the same table with the comparison that matters:

The published numbers

  • Scrapling — 1.99ms
  • Parsel / Scrapy — 2.06ms (1.03×)
  • Raw lxml — 2.56ms (1.29×)
  • PyQuery — 23.98ms (12×)
  • BeautifulSoup + lxml — 1,562ms (785×)
  • BeautifulSoup + html5lib — 3,413ms (1,715×)

What that means for you

  • Against the nearest serious competitor it's 3% faster, not 700×
  • The giant multiplier is against BeautifulSoup, the slow beginner-friendly one
  • It measures parsing, not fetching — the network is what actually makes scraping slow
  • In a real job the wait is the website answering, and no library changes that
  • Parsing 5,000 elements is a stress test, not a normal page

So is it still worth installing?

Yes — just not for the number. The reasons that hold up: it gets past Cloudflare where a plain request gets a wall, its selectors survive a site redesign instead of silently returning nothing, and it ships an MCP server so your agent can use it without you writing Python. Speed is the headline. Those three are the reason.

While we're checking numbers

The star count is 78,900 as of 07/09/2026, not 77,000 — it's gone up, so if you saw the lower figure it was just out of date. Licence is BSD-3-Clause, so commercial use is fine.


Ready to go deeper?

Where to go from here.