

Finding every page on a website takes a few methods working together, not one magic button. Start with a Google site: search and the XML sitemap for a quick overview, then run an SEO crawler or a custom script to map internal links, and cross-check with Search Console for anything missing. Proxy-Cheap builds infrastructure for large-scale crawling, so this guide focuses on methods that hold up at scale.
| Method | Best for | Access needed | Technical level |
|---|---|---|---|
| Google site: operators | Fast overview, specific sections | None (public) | Beginner |
| XML sitemap + robots.txt | Structured URL list, quick start | None (public) | Beginner |
| SEO crawler (Screaming Frog, Sitebulb) | Audits, internal-link mapping | None (public) | Intermediate |
| Custom Python script | Large or dynamic sites, automation | None (public) | Advanced |
| Google Search Console | Owner's indexed pages | Site owner | Beginner |
| Google Analytics + server logs | Orphan and traffic-only pages | Site owner | Intermediate |
| Bing Webmaster Tools + Wayback Machine | Second index, historical URLs | Mixed | Beginner |
| No-code / AI map tools (n8n, Make, Firecrawl) | Scheduled, repeatable discovery | None (public) | Beginner to intermediate |

The fastest way to see pages on a domain is a Google site: search. Type site:example.com to list the pages Google has indexed, then narrow with inurl:, intitle:, or filetype:. It is instant and needs no tools, but it only shows what Google has indexed and chosen to display, so treat it as a starting point, not a full inventory.
A plain site:example.com gives a rough count of indexed pages and lets you scroll through the URLs Google is willing to show, then narrow the results with a handful of operators.
| Operator | Example | What it does |
|---|---|---|
| site: | site:example.com | Lists indexed pages on a domain |
| inurl: | site:example.com inurl:blog | Pages with a word or folder in the URL |
| intitle: | site:example.com intitle:pricing | Pages with a word in the page title |
| filetype: | site:example.com filetype:pdf | Indexed files of a type |
| - | site:example.com -inurl:blog | Excludes a section |
Stack operators to narrow further: site:example.com inurl:blog -inurl:tag surfaces blog post URLs while filtering out tag archive pages. Full details on syntax and limits are in Google's search operators reference.
The catch is that site: counts are approximate, not a page total. Google omits some URLs and shows some outdated or redirected ones. And one classic trick no longer works: Google retired the cache: operator in 2023, so you can't pull a cached snapshot that way anymore. For historical URLs, Method 7 covers a workaround.
After a site: search, check the sitemap and robots.txt. Visit example.com/sitemap.xml for a structured list of URLs a site exposes to search engines, and example.com/robots.txt for links to more sitemaps and structure clues. Most sites keep their sitemap at a predictable location: /sitemap.xml, /sitemap_index.xml, or a compressed /sitemap.xml.gz. Larger sites usually publish a sitemap index instead of one flat list. An index file doesn't contain page URLs directly; it points to child sitemaps (often split by language or post type), and you need to open each child to get the actual URLs, per the sitemap protocol.
Inside each sitemap, the <loc> tag holds the page URL and <lastmod> tells you when that page last changed, useful for spotting stale pages during a content refresh. If you're pulling sitemap URLs from many domains at once and want session behavior to stay consistent across requests, it helps to know the different proxy types available for that kind of job.
The robots.txt file often lists multiple sitemaps under separate Sitemap: directives, especially on ecommerce or multi-language sites, and its Disallow: lines hint at site structure, though treat them as a request, not proof that a page exists. One more shortcut: site:example.com filetype:xml sometimes surfaces sitemap files that Google has indexed directly.
The limitation is the same as Method 1: never rely on the sitemap alone. Orphan pages, dynamic routes generated by search filters, and pages excluded from indexing are all missing from a sitemap by design.
<!-- sitemapindex: points to child sitemaps --> <sitemapindex xmlns="http://www.sitemaps.org/schemas/sitemap/0.9"> <sitemap> <loc>https://example.com/sitemap-pages.xml</loc> </sitemap> <sitemap> <loc>https://example.com/sitemap-posts.xml</loc> </sitemap> </sitemapindex> <!-- urlset: the actual page list inside a child sitemap --> <urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9"> <url> <loc>https://example.com/services/example-page</loc> <lastmod>2026-06-01</lastmod> </url> </urlset>
For a point-and-click list, use an SEO crawler. Tools like Screaming Frog and Sitebulb start at the homepage, follow internal links, and export every reachable page with its status code, title, and metadata. This is the easiest way to get a clean list for an audit or migration. The main limit is scale: desktop crawlers and free tiers cap out, so very large sites need a script or an API.
Under the hood, a spider loads a root URL, reads every link on that page, adds new same-domain links to a queue, and repeats this basic loop until the queue empties, whether you're running a paid tool or the script in Method 4.
| Tool | Free tier | Best for |
|---|---|---|
| Screaming Frog | 500 URLs | Technical audits, exports |
| Sitebulb | Trial | Visual site structure |
| Ahrefs Site Audit | Verified sites, credit limits | SEO issue reporting |
| Semrush Site Audit | Limited | Diagnostics, clean UI |
Screaming Frog's free tier crawls up to 500 URLs, enough for many small business sites. After a crawl finishes, filter to HTML only (the raw crawl also picks up images, CSS, and JavaScript files), then export the Internal tab to CSV as your working page list. Sitebulb, Ahrefs Site Audit, and Semrush Site Audit follow the same pattern if you already have a subscription and a verified property.
A few crawler settings change what you get back: crawl speed, the user agent presented to the server, and whether JavaScript rendering is on. That last one matters most, since a site with client-side JavaScript navigation will hide whole sections of internal links from a crawler that doesn't render it.
The shared limitation: a crawler can only find pages it can reach by following links. An orphan page, one that links to nothing, never shows up in a crawl export. Method 6 covers how to catch those. Crawlers used for SEO tasks like site audits tend to hit per-IP rate limits once a crawl reaches thousands of URLs, which is where Method 4's approach to routing requests through more IPs becomes important.

When you need full control or a repeatable job, write a script. A short Python script can fetch the sitemap, extract all <loc> URLs, crawl internal links to catch pages the sitemap misses, and export the combined list to CSV. This is the right method for large sites, dynamic routes, or discovery you will run again, and it's the version of URL discovery most web scraping and data collection projects start with, since it also lets you route requests through proxies so a big crawl stays reliable.
The approach has two parts. First, parse the sitemap, since it's structured data and provides a fast, mostly complete starting list. Second, crawl internal links from the homepage the way an SEO crawler would, to catch pages the sitemap left out. Combine both, deduplicate, and you get a more complete picture than either method alone. requests and BeautifulSoup are enough for a simple crawler, xmltodict makes sitemap parsing straightforward, and Scrapy is worth the setup once you're running this against dozens of sites on a schedule. A sitemapindex needs to be parsed recursively, since the index only holds pointers to child sitemaps, not page URLs directly.
Housekeeping rules keep the output clean: deduplicate and normalize URLs (strip fragments, drop trailing slashes), and stay on the same domain so the crawl doesn't wander into linked third-party sites. Respect robots.txt and add a short delay between requests to stay under rate limits.
On a large site, a single IP address can become a bottleneck. A few thousand requests in a row from a single IP address routinely trigger per-IP rate limits on the target server, and the crawl slows or starts returning errors, so the URL list comes back incomplete. Passing a proxies dictionary to the requests session routes traffic through a pool of IPs rather than a single IP, keeping the crawl's success rate high enough to finish with a complete list.
Rotating residential proxies work well here because they automatically rotate the exit IP, using the gateway format proxy-us.proxy-cheap.com:5959 for the US or proxy-eu.proxy-cheap.com:5959 for the EU, with username and password authentication.
python """ Parse a sitemap (including a sitemapindex), crawl internal links from the homepage to catch pages the sitemap misses, then export the combined, deduplicated URL list to CSV. Tested: Python 3.11.15, requests 2.33.1, beautifulsoup4 4.14.3, xmltodict 0.13.0. Replace USERNAME:PASSWORD with real Proxy-Cheap gateway credentials before running. """ import csv import time from urllib.parse import urljoin, urlparse, urldefrag import requests import xmltodict from bs4 import BeautifulSoup DOMAIN = "https://example.com" SITEMAP_URL = f"{DOMAIN}/sitemap.xml" CRAWL_DELAY = 1.0 # seconds between requests, keeps requests under per-IP rate limits MAX_PAGES = 2000 # safety cap for very large or misconfigured sites # Proxy-Cheap rotating residential gateway. Use proxy-eu.proxy-cheap.com:5959 for the EU pool. PROXIES = { "http": "http://USERNAME:[email protected]:5959", "https": "http://USERNAME:[email protected]:5959", } session = requests.Session() session.headers.update({"User-Agent": "site-mapper/1.0"}) def get_sitemap_urls(sitemap_url, seen=None): """Fetch a sitemap or sitemap index and return every <loc> URL, recursively.""" if seen is None: seen = set() if sitemap_url in seen: return [] seen.add(sitemap_url) resp = session.get(sitemap_url, timeout=15) resp.raise_for_status() parsed = xmltodict.parse(resp.content) urls = [] if "sitemapindex" in parsed: entries = parsed["sitemapindex"].get("sitemap", []) entries = [entries] if isinstance(entries, dict) else entries for entry in entries: urls.extend(get_sitemap_urls(entry["loc"], seen)) elif "urlset" in parsed: entries = parsed["urlset"].get("url", []) entries = [entries] if isinstance(entries, dict) else entries urls.extend(entry["loc"] for entry in entries) return urls def normalize(url): """Strip fragments and trailing slashes so the same page isn't counted twice.""" url, _ = urldefrag(url) return url.rstrip("/") def crawl_internal_links(start_url, known_urls): """Follow internal links from start_url to catch pages the sitemap missed.""" to_visit = {normalize(start_url)} visited = set() discovered = set(normalize(u) for u in known_urls) while to_visit and len(visited) < MAX_PAGES: url = to_visit.pop() if url in visited: continue visited.add(url) discovered.add(url) try: resp = session.get(url, proxies=PROXIES, timeout=10) except requests.exceptions.RequestException: continue # log and move on; a handful of dead links shouldn't stop the run time.sleep(CRAWL_DELAY) if "text/html" not in resp.headers.get("Content-Type", ""): continue soup = BeautifulSoup(resp.text, "html.parser") for link in soup.find_all("a", href=True): absolute = urljoin(url, link["href"]) if urlparse(absolute).netloc == urlparse(DOMAIN).netloc: candidate = normalize(absolute) if candidate not in visited: to_visit.add(candidate) return discovered if __name__ == "__main__": sitemap_urls = get_sitemap_urls(SITEMAP_URL) all_urls = crawl_internal_links(DOMAIN, sitemap_urls) with open("all_urls.csv", "w", newline="", encoding="utf-8") as f: writer = csv.writer(f) writer.writerow(["url"]) for url in sorted(all_urls): writer.writerow([url]) print(f"Sitemap URLs: {len(sitemap_urls)} | Combined unique URLs: {len(all_urls)}")
If you own the site, Google Search Console is the most accurate source for the pages Google knows about. Open the Pages report under Indexing to see indexed and not-indexed URLs, and the Sitemaps section to submit and review your sitemap. It shows exactly which pages Google has, including ones a crawler on your own site might miss.
The Pages report breaks pages into indexed and non-indexed buckets, and for every non-indexed URL, it provides a reason (duplicate content, crawled but not indexed, blocked by robots.txt). That reason is often more useful than the URL list itself. The Sitemaps section, under the same Indexing menu, is where you submit a sitemap and see whether Search Console successfully reads it.
Export the Pages report's indexed URLs and treat that as your master reference. Since it comes directly from Google's index for your verified property, it's more reliable than a public site: search, and it's the natural starting point in the owner-vs-outsider split covered later in this guide, since it's the one method that only works if you own the site.
Orphan pages exist on a site but have no internal links pointing to them, so crawlers never reach them. To find them, compare sources: export URLs from your crawler, from Google Analytics (pages with traffic), from Search Console, and from server logs, then look for URLs that show up in traffic or logs but not in your crawl. Those are your orphan candidates.
A page becomes an orphan for mundane reasons: a navigation redesign dropped a link, an old landing page still gets paid traffic but was cut from the menu, or a page was published and never linked from anywhere on the site. Crawlers only find what they can click through to, so none of these show up in a Method 3 or Method 4 export.
To find them, pull URLs from a few sources and diff the lists: your crawler's export, GA4 landing pages (every URL that received a session, linked or not), Search Console's indexed pages, server logs if you have access, and backlink tools for pages that hold external links despite being orphaned internally. Filter for URLs that appear in traffic, indexing, or log data but not in the crawl. What you do with an orphan depends on whether it's worth keeping: link to it from a relevant page, or redirect and retire it if it's outdated.
Two more sources fill the gaps the others leave. Bing Webmaster Tools is a second search index that sometimes lists pages Google does not. The Wayback Machine keeps historical snapshots and has a site map view, which is useful for finding old URLs after a redesign, especially now that Google's cache: operator is gone.
Bing Webmaster Tools works the same way conceptually as Search Console: verify the property and you get access to Bing's own indexed-pages report. Bing crawls independently, so its index doesn't always match Google's, and pages Google has dropped sometimes still show up there.
For historical URLs, especially after a migration or a redesign that changed the URL structure, the Wayback Machine is a more useful tool. Its site map feature lists archived URLs by path across whatever time range has been crawled, the closest thing left to the old cache: operator now that Google has retired it. Rebuilding a redirect map after a site move usually means cross-checking old URLs against this archive.
If you need discovery on a schedule, no-code and AI tools handle it. Platforms like n8n and Make can fetch a sitemap, parse the <loc> URLs, and drop them into a Google Sheet on a timer. Newer AI scraping tools expose a "map" step (Firecrawl, Olostep) that returns a domain's URLs in one call, which has become the first step in many AI data pipelines.
A typical no-code flow: an HTTP or Search module fetches the sitemap XML, a parsing step extracts the <loc> values, a filter drops anything outside the section you care about, and a final step writes the results into a Google Sheets tab on a weekly schedule.
Worth knowing about: several AI-focused scraping platforms now expose a dedicated "map" endpoint that returns a domain's discoverable URLs in one API call instead of requiring you to assemble sitemap parsing and link crawling yourself, and this has become a first step in a lot of AI-driven data pipelines.
No-code flows fit public sites with a clean sitemap and a use case that benefits from a recurring schedule. Once a site relies on JavaScript-rendered routing, needs custom logic, or the volume is large enough to route requests through a proxy API, a script, or a dedicated crawling service is a better tool.
When you crawl a large site, sending thousands of requests from one IP address quickly hits per-IP rate limits, and the crawl slows down or returns an incomplete list. Routing requests through a pool of rotating residential or datacenter proxies spreads them across many IPs, so a large crawl keeps a high success rate and finishes with a complete URL set.
This isn't unique to any one target site; it's how most servers protect themselves from a high volume of requests arriving from a single source in a short window. A proxy pool fixes this by distributing requests across many IPs instead of one, so no single address sends enough volume to trip a rate limit, keeping a steadier success rate across the whole crawl.
Which product fits depends on the job. Rotating residential proxies provides the broadest IP diversity for tougher targets that monitor traffic closely. Datacenter proxies are well-suited for fast, high-volume crawling of straightforward public content. ISP proxies combine residential-style trust with datacenter-level speed for long crawls on account-bound pages, and unlimited bandwidth proxies remove bandwidth as a constraint on very large, request-heavy jobs.
Each runs on pay-as-you-go or per-IP pricing with no monthly commitment, so it's realistic to scale a proxy pool up for one large crawl and back down once the job is done.
Pick by access and scale. If you own the site, start with Search Console and Analytics, then crawl to fill gaps. If not, start with a site: search and the sitemap, then run a crawler or script. Quick check: use operators. Audit: use a crawler. Large or repeatable job: use a script or API with a proxy pool.
| Your situation | Start here | Then |
|---|---|---|
| You own the site, small | Search Console | Crawler to confirm |
| You own the site, large | Search Console + Analytics | Script + server logs for orphans |
| Outsider, quick check | site: operators | Sitemap |
| Outsider, audit | SEO crawler | site: cross-check |
| Outsider, large or repeat | Custom script or API | Proxy pool for reliability |
Access is the real fork in the road, more than site size or technical comfort. A site owner with Search Console access is already looking at the most accurate list Google has for that property, so the job is mostly filling in orphan pages. Method 5 alone can't surface. An outsider mapping a domain before a scraping or research project starts with public methods instead (site: search, sitemap) and scales up to a script with static residential proxies or another proxy type once the site is too large for a manual pass.
A complete URL inventory earns its keep in a handful of recurring situations. SEO audits use it to surface broken links, duplicate content, and thin or orphaned pages that would otherwise sit unnoticed. Site migrations and redesigns depend on a full list to build clean redirects and confirm nothing important gets dropped.
A content refresh needs the same list to know what already exists before planning updates. Competitor and market research uses a mapped domain to read a competitor's structure and priorities from the outside, and any web scraping or data collection project starts with a URL list, since extraction can't begin until you know what pages exist.