Claude web scraping
How to scrape websites with Claude.
Claude can scrape almost any public website. Not by browsing the way you do. It writes the scraper, runs it, fixes it when it breaks, and hands you a spreadsheet.
I use it every week for market maps and prospect lists. The biggest one so far: every golf course in the US, 12,339 of them, with the booking software each one runs. Here's how it works, start to finish.
Chat vs Claude Code
In the Claude chat app, Claude can read a page you point it at. That's fine for one page. It's the wrong tool for a thousand.
For lists, use Claude Code. It runs on your computer, writes small Python scripts, runs them, and saves the output to files. You describe the list you want in plain English. You never have to touch the code.
The seven techniques, cheapest first
Every scrape starts at the top of this list and only moves down when it has to.
1. Just ask for the page
Most websites hand over their content to a plain request. No browser, no tricks. This is the workhorse. It read thousands of golf course websites without a browser.
2. Use a real browser for pages that load late
Some sites build the page with JavaScript after it loads, so a plain request gets an empty shell. A headless browser (a real browser with no window) waits for the page to finish, then reads it.
3. Catch the data behind the page
A lot of sites load their data from a hidden API and then draw it on screen. Find that request and you skip the page entirely. You get clean, structured data, often thousands of rows at once.
4. Build a competitor's customer list
Software leaves footprints on its customers' websites. "Powered by" lines, widget links, backlinks to the vendor. Search for the footprint and you have a list of who uses it.
5. Find a website's hidden pages
Companies run pages that aren't linked from the homepage: portals, booking subdomains, staging sites. Public certificate logs like crt.sh list every subdomain a company has ever secured.
6. Get past sites that block you
Some sites block anything that looks automated. Pebble Beach's site blocked a plain headless browser. CloakBrowser got through. For the toughest sites, a proxy network like Bright Data is the last rung.
7. See what software any business runs
Read the fingerprints on each site and you know the CRM, the booking system, the platform. This is the step that turns a list into a market map. Full walkthrough: technographic data with Claude.
A real run: every golf course in the US
No list bought from anyone. Free public data and Claude Code.
The real numbers from the golf run. Every row came from public data.
The payoff wasn't the list. It was what the list showed. Private clubs almost all run one of three member platforms. Public courses mostly run tee-sheet software. One scraped field tells you how each course makes money, and what to pitch it. See the full golf market map.
How to start
- Install Claude Code and open an empty folder.
- Describe the list you want. Who, where, and what you want to know about each one.
- Ask Claude to find a free source first. OpenStreetMap, government datasets, and industry directories cover more than you'd think.
- Have it test on 20 rows, check them yourself, then run the rest.
Tell it to leave a cell blank when it isn't sure. Blank beats wrong.
Playing it straight
Stick to public pages. Keep your request rate modest so you never strain anyone's site. Don't log in to scrape, and don't keep personal data you don't need.
Where the list goes
Everything lands in a CSV. From there it goes into HubSpot, Salesforce, Google Sheets, or straight into a cold email tool like Instantly or Smartlead. Anything that takes a CSV. Weighing this against Clay? See Clay alternatives.