Claude web scraping

How to scrape websites with Claude.

Claude can scrape almost any public website. Not by browsing the way you do. It writes the scraper, runs it, fixes it when it breaks, and hands you a spreadsheet.

I use it every week for market maps and prospect lists. The biggest one so far: every golf course in the US, 12,339 of them, with the booking software each one runs. Here's how it works, start to finish.

Chat vs Claude Code

In the Claude chat app, Claude can read a page you point it at. That's fine for one page. It's the wrong tool for a thousand.

For lists, use Claude Code. It runs on your computer, writes small Python scripts, runs them, and saves the output to files. You describe the list you want in plain English. You never have to touch the code.

The seven techniques, cheapest first

Every scrape starts at the top of this list and only moves down when it has to.

1. Just ask for the page

Most websites hand over their content to a plain request. No browser, no tricks. This is the workhorse. It read thousands of golf course websites without a browser.

2. Use a real browser for pages that load late

Some sites build the page with JavaScript after it loads, so a plain request gets an empty shell. A headless browser (a real browser with no window) waits for the page to finish, then reads it.

3. Catch the data behind the page

A lot of sites load their data from a hidden API and then draw it on screen. Find that request and you skip the page entirely. You get clean, structured data, often thousands of rows at once.

4. Build a competitor's customer list

Software leaves footprints on its customers' websites. "Powered by" lines, widget links, backlinks to the vendor. Search for the footprint and you have a list of who uses it.

5. Find a website's hidden pages

Companies run pages that aren't linked from the homepage: portals, booking subdomains, staging sites. Public certificate logs like crt.sh list every subdomain a company has ever secured.

6. Get past sites that block you

Some sites block anything that looks automated. Pebble Beach's site blocked a plain headless browser. CloakBrowser got through. For the toughest sites, a proxy network like Bright Data is the last rung.

7. See what software any business runs

Read the fingerprints on each site and you know the CRM, the booking system, the platform. This is the step that turns a list into a market map. Full walkthrough: technographic data with Claude.

A real run: every golf course in the US

No list bought from anyone. Free public data and Claude Code.

golf-run · summary
→ list OpenStreetMap / Overpass ... 15,955 courses pulled
→ clean drop unnamed + duplicates ... 12,339 named courses
→ addresses OSM tags + Nominatim ... 12,262 (99%)
→ websites OSM + Serper free tier ... 9,101 (73%)
→ software homepage + booking page ... 3,044 identified
→ junk dead, parked, hijacked domains ... 577 flagged

total data cost: $0

The real numbers from the golf run. Every row came from public data.

The payoff wasn't the list. It was what the list showed. Private clubs almost all run one of three member platforms. Public courses mostly run tee-sheet software. One scraped field tells you how each course makes money, and what to pitch it. See the full golf market map.

How to start

  1. Install Claude Code and open an empty folder.
  2. Describe the list you want. Who, where, and what you want to know about each one.
  3. Ask Claude to find a free source first. OpenStreetMap, government datasets, and industry directories cover more than you'd think.
  4. Have it test on 20 rows, check them yourself, then run the rest.

Tell it to leave a cell blank when it isn't sure. Blank beats wrong.

Playing it straight

Stick to public pages. Keep your request rate modest so you never strain anyone's site. Don't log in to scrape, and don't keep personal data you don't need.

Where the list goes

Everything lands in a CSV. From there it goes into HubSpot, Salesforce, Google Sheets, or straight into a cold email tool like Instantly or Smartlead. Anything that takes a CSV. Weighing this against Clay? See Clay alternatives.

The full method

Get the playbook. Run your first scrape this week.

Seven techniques, 30 free data sources, four copy-paste playbooks, a five-minute setup. Free.

Who's behind this

I'm Chris. I do marketing and GTM engineering, with over 10 years in B2B SaaS. I've used Claude to create over $10,700,000 in pipeline, so I wrote down how I actually do it.

@chris_as_is