Proxy-Cheap
Tutorials
October 9, 2026
7 min

How to Find All Webpages on a Website (8 Working Methods)

Alex Sadovskij
Alex Sadovskij
CEO Proxy-Cheap
How to Find All Webpages on a Website (8 Working Methods)
Summary
Find all webpages on a website with Google site: search, the XML sitemap and robots.txt, an SEO crawler like Screaming Frog, or a custom Python script.

Finding every page on a website takes a few methods working together, not one magic button. Start with a Google site: search and the XML sitemap for a quick overview, then run an SEO crawler or a custom script to map internal links, and cross-check with Search Console for anything missing. Proxy-Cheap builds infrastructure for large-scale crawling, so this guide focuses on methods that hold up at scale.

  • No single tool lists every page. The complete list comes from combining a site: search, the sitemap, a crawler, and Search Console.
  • Your best method depends on access: site owners should start with Search Console and Analytics, outsiders should start with site: search and a crawler.
  • The sitemap and robots.txt are the fastest structured starting point, but they are never complete on their own (orphan and dynamic pages are missing).
  • On large sites, a single IP hits per-IP rate limits and the crawl returns a partial list. Distributing requests across rotating residential or datacenter IPs keeps the list complete.

Which method fits your situation

MethodBest forAccess neededTechnical level
Google site: operatorsFast overview, specific sectionsNone (public)Beginner
XML sitemap + robots.txtStructured URL list, quick startNone (public)Beginner
SEO crawler (Screaming Frog, Sitebulb)Audits, internal-link mappingNone (public)Intermediate
Custom Python scriptLarge or dynamic sites, automationNone (public)Advanced
Google Search ConsoleOwner's indexed pagesSite ownerBeginner
Google Analytics + server logsOrphan and traffic-only pagesSite ownerIntermediate
Bing Webmaster Tools + Wayback MachineSecond index, historical URLsMixedBeginner
No-code / AI map tools (n8n, Make, Firecrawl)Scheduled, repeatable discoveryNone (public)Beginner to intermediate

Method 1: Use Google search operators

The fastest way to see pages on a domain is a Google site: search. Type site:example.com to list the pages Google has indexed, then narrow with inurl:, intitle:, or filetype:. It is instant and needs no tools, but it only shows what Google has indexed and chosen to display, so treat it as a starting point, not a full inventory.

A plain site:example.com gives a rough count of indexed pages and lets you scroll through the URLs Google is willing to show, then narrow the results with a handful of operators.

OperatorExampleWhat it does
site:site:example.comLists indexed pages on a domain
inurl:site:example.com inurl:blogPages with a word or folder in the URL
intitle:site:example.com intitle:pricingPages with a word in the page title
filetype:site:example.com filetype:pdfIndexed files of a type
-site:example.com -inurl:blogExcludes a section

Stack operators to narrow further: site:example.com inurl:blog -inurl:tag surfaces blog post URLs while filtering out tag archive pages. Full details on syntax and limits are in Google's search operators reference.

The catch is that site: counts are approximate, not a page total. Google omits some URLs and shows some outdated or redirected ones. And one classic trick no longer works: Google retired the cache: operator in 2023, so you can't pull a cached snapshot that way anymore. For historical URLs, Method 7 covers a workaround.

Method 2: check the XML sitemap and robots.txt

After a site: search, check the sitemap and robots.txt. Visit example.com/sitemap.xml for a structured list of URLs a site exposes to search engines, and example.com/robots.txt for links to more sitemaps and structure clues. Most sites keep their sitemap at a predictable location: /sitemap.xml, /sitemap_index.xml, or a compressed /sitemap.xml.gz. Larger sites usually publish a sitemap index instead of one flat list. An index file doesn't contain page URLs directly; it points to child sitemaps (often split by language or post type), and you need to open each child to get the actual URLs, per the sitemap protocol.

Inside each sitemap, the <loc> tag holds the page URL and <lastmod> tells you when that page last changed, useful for spotting stale pages during a content refresh. If you're pulling sitemap URLs from many domains at once and want session behavior to stay consistent across requests, it helps to know the different proxy types available for that kind of job.

The robots.txt file often lists multiple sitemaps under separate Sitemap: directives, especially on ecommerce or multi-language sites, and its Disallow: lines hint at site structure, though treat them as a request, not proof that a page exists. One more shortcut: site:example.com filetype:xml sometimes surfaces sitemap files that Google has indexed directly.

The limitation is the same as Method 1: never rely on the sitemap alone. Orphan pages, dynamic routes generated by search filters, and pages excluded from indexing are all missing from a sitemap by design.

<!-- sitemapindex: points to child sitemaps --> <sitemapindex xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">  <sitemap>    <loc>https://example.com/sitemap-pages.xml</loc>  </sitemap>  <sitemap>    <loc>https://example.com/sitemap-posts.xml</loc>  </sitemap> </sitemapindex> <!-- urlset: the actual page list inside a child sitemap --> <urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">  <url>    <loc>https://example.com/services/example-page</loc>    <lastmod>2026-06-01</lastmod>  </url> </urlset>

Method 3: run an SEO crawler

For a point-and-click list, use an SEO crawler. Tools like Screaming Frog and Sitebulb start at the homepage, follow internal links, and export every reachable page with its status code, title, and metadata. This is the easiest way to get a clean list for an audit or migration. The main limit is scale: desktop crawlers and free tiers cap out, so very large sites need a script or an API.

Under the hood, a spider loads a root URL, reads every link on that page, adds new same-domain links to a queue, and repeats this basic loop until the queue empties, whether you're running a paid tool or the script in Method 4.

ToolFree tierBest for
Screaming Frog500 URLsTechnical audits, exports
SitebulbTrialVisual site structure
Ahrefs Site AuditVerified sites, credit limitsSEO issue reporting
Semrush Site AuditLimitedDiagnostics, clean UI

Screaming Frog's free tier crawls up to 500 URLs, enough for many small business sites. After a crawl finishes, filter to HTML only (the raw crawl also picks up images, CSS, and JavaScript files), then export the Internal tab to CSV as your working page list. Sitebulb, Ahrefs Site Audit, and Semrush Site Audit follow the same pattern if you already have a subscription and a verified property.

A few crawler settings change what you get back: crawl speed, the user agent presented to the server, and whether JavaScript rendering is on. That last one matters most, since a site with client-side JavaScript navigation will hide whole sections of internal links from a crawler that doesn't render it.

The shared limitation: a crawler can only find pages it can reach by following links. An orphan page, one that links to nothing, never shows up in a crawl export. Method 6 covers how to catch those. Crawlers used for SEO tasks like site audits tend to hit per-IP rate limits once a crawl reaches thousands of URLs, which is where Method 4's approach to routing requests through more IPs becomes important.

When you need full control or a repeatable job, write a script. A short Python script can fetch the sitemap, extract all <loc> URLs, crawl internal links to catch pages the sitemap misses, and export the combined list to CSV. This is the right method for large sites, dynamic routes, or discovery you will run again, and it's the version of URL discovery most web scraping and data collection projects start with, since it also lets you route requests through proxies so a big crawl stays reliable.

The approach has two parts. First, parse the sitemap, since it's structured data and provides a fast, mostly complete starting list. Second, crawl internal links from the homepage the way an SEO crawler would, to catch pages the sitemap left out. Combine both, deduplicate, and you get a more complete picture than either method alone. requests and BeautifulSoup are enough for a simple crawler, xmltodict makes sitemap parsing straightforward, and Scrapy is worth the setup once you're running this against dozens of sites on a schedule. A sitemapindex needs to be parsed recursively, since the index only holds pointers to child sitemaps, not page URLs directly.

Housekeeping rules keep the output clean: deduplicate and normalize URLs (strip fragments, drop trailing slashes), and stay on the same domain so the crawl doesn't wander into linked third-party sites. Respect robots.txt and add a short delay between requests to stay under rate limits.

On a large site, a single IP address can become a bottleneck. A few thousand requests in a row from a single IP address routinely trigger per-IP rate limits on the target server, and the crawl slows or starts returning errors, so the URL list comes back incomplete. Passing a proxies dictionary to the requests session routes traffic through a pool of IPs rather than a single IP, keeping the crawl's success rate high enough to finish with a complete list.

Rotating residential proxies work well here because they automatically rotate the exit IP, using the gateway format proxy-us.proxy-cheap.com:5959 for the US or proxy-eu.proxy-cheap.com:5959 for the EU, with username and password authentication.

python """ Parse a sitemap (including a sitemapindex), crawl internal links from the homepage to catch pages the sitemap misses, then export the combined, deduplicated URL list to CSV. Tested: Python 3.11.15, requests 2.33.1, beautifulsoup4 4.14.3, xmltodict 0.13.0. Replace USERNAME:PASSWORD with real Proxy-Cheap gateway credentials before running. """ import csv import time from urllib.parse import urljoin, urlparse, urldefrag import requests import xmltodict from bs4 import BeautifulSoup DOMAIN = "https://example.com" SITEMAP_URL = f"{DOMAIN}/sitemap.xml" CRAWL_DELAY = 1.0          # seconds between requests, keeps requests under per-IP rate limits MAX_PAGES = 2000           # safety cap for very large or misconfigured sites # Proxy-Cheap rotating residential gateway. Use proxy-eu.proxy-cheap.com:5959 for the EU pool. PROXIES = {    "http": "http://USERNAME:[email protected]:5959",    "https": "http://USERNAME:[email protected]:5959", } session = requests.Session() session.headers.update({"User-Agent": "site-mapper/1.0"}) def get_sitemap_urls(sitemap_url, seen=None):    """Fetch a sitemap or sitemap index and return every <loc> URL, recursively."""    if seen is None:        seen = set()    if sitemap_url in seen:        return []    seen.add(sitemap_url)    resp = session.get(sitemap_url, timeout=15)    resp.raise_for_status()    parsed = xmltodict.parse(resp.content)    urls = []    if "sitemapindex" in parsed:        entries = parsed["sitemapindex"].get("sitemap", [])        entries = [entries] if isinstance(entries, dict) else entries        for entry in entries:            urls.extend(get_sitemap_urls(entry["loc"], seen))    elif "urlset" in parsed:        entries = parsed["urlset"].get("url", [])        entries = [entries] if isinstance(entries, dict) else entries        urls.extend(entry["loc"] for entry in entries)    return urls def normalize(url):    """Strip fragments and trailing slashes so the same page isn't counted twice."""    url, _ = urldefrag(url)    return url.rstrip("/") def crawl_internal_links(start_url, known_urls):    """Follow internal links from start_url to catch pages the sitemap missed."""    to_visit = {normalize(start_url)}    visited = set()    discovered = set(normalize(u) for u in known_urls)    while to_visit and len(visited) < MAX_PAGES:        url = to_visit.pop()        if url in visited:            continue        visited.add(url)        discovered.add(url)        try:            resp = session.get(url, proxies=PROXIES, timeout=10)        except requests.exceptions.RequestException:            continue  # log and move on; a handful of dead links shouldn't stop the run        time.sleep(CRAWL_DELAY)        if "text/html" not in resp.headers.get("Content-Type", ""):            continue        soup = BeautifulSoup(resp.text, "html.parser")        for link in soup.find_all("a", href=True):            absolute = urljoin(url, link["href"])            if urlparse(absolute).netloc == urlparse(DOMAIN).netloc:                candidate = normalize(absolute)                if candidate not in visited:                    to_visit.add(candidate)    return discovered if __name__ == "__main__":    sitemap_urls = get_sitemap_urls(SITEMAP_URL)    all_urls = crawl_internal_links(DOMAIN, sitemap_urls)    with open("all_urls.csv", "w", newline="", encoding="utf-8") as f:        writer = csv.writer(f)        writer.writerow(["url"])        for url in sorted(all_urls):            writer.writerow([url])    print(f"Sitemap URLs: {len(sitemap_urls)} | Combined unique URLs: {len(all_urls)}")

Method 5: Use Google Search Console

If you own the site, Google Search Console is the most accurate source for the pages Google knows about. Open the Pages report under Indexing to see indexed and not-indexed URLs, and the Sitemaps section to submit and review your sitemap. It shows exactly which pages Google has, including ones a crawler on your own site might miss.

The Pages report breaks pages into indexed and non-indexed buckets, and for every non-indexed URL, it provides a reason (duplicate content, crawled but not indexed, blocked by robots.txt). That reason is often more useful than the URL list itself. The Sitemaps section, under the same Indexing menu, is where you submit a sitemap and see whether Search Console successfully reads it.

Export the Pages report's indexed URLs and treat that as your master reference. Since it comes directly from Google's index for your verified property, it's more reliable than a public site: search, and it's the natural starting point in the owner-vs-outsider split covered later in this guide, since it's the one method that only works if you own the site.

Method 6: Catch orphan pages with Analytics and server logs

Orphan pages exist on a site but have no internal links pointing to them, so crawlers never reach them. To find them, compare sources: export URLs from your crawler, from Google Analytics (pages with traffic), from Search Console, and from server logs, then look for URLs that show up in traffic or logs but not in your crawl. Those are your orphan candidates.

A page becomes an orphan for mundane reasons: a navigation redesign dropped a link, an old landing page still gets paid traffic but was cut from the menu, or a page was published and never linked from anywhere on the site. Crawlers only find what they can click through to, so none of these show up in a Method 3 or Method 4 export.

To find them, pull URLs from a few sources and diff the lists: your crawler's export, GA4 landing pages (every URL that received a session, linked or not), Search Console's indexed pages, server logs if you have access, and backlink tools for pages that hold external links despite being orphaned internally. Filter for URLs that appear in traffic, indexing, or log data but not in the crawl. What you do with an orphan depends on whether it's worth keeping: link to it from a relevant page, or redirect and retire it if it's outdated.

Method 7: Check Bing Webmaster Tools and the Wayback Machine

Two more sources fill the gaps the others leave. Bing Webmaster Tools is a second search index that sometimes lists pages Google does not. The Wayback Machine keeps historical snapshots and has a site map view, which is useful for finding old URLs after a redesign, especially now that Google's cache: operator is gone.

Bing Webmaster Tools works the same way conceptually as Search Console: verify the property and you get access to Bing's own indexed-pages report. Bing crawls independently, so its index doesn't always match Google's, and pages Google has dropped sometimes still show up there.

For historical URLs, especially after a migration or a redesign that changed the URL structure, the Wayback Machine is a more useful tool. Its site map feature lists archived URLs by path across whatever time range has been crawled, the closest thing left to the old cache: operator now that Google has retired it. Rebuilding a redirect map after a site move usually means cross-checking old URLs against this archive.

Method 8: Automate discovery with no-code and AI map tools

If you need discovery on a schedule, no-code and AI tools handle it. Platforms like n8n and Make can fetch a sitemap, parse the <loc> URLs, and drop them into a Google Sheet on a timer. Newer AI scraping tools expose a "map" step (Firecrawl, Olostep) that returns a domain's URLs in one call, which has become the first step in many AI data pipelines.

A typical no-code flow: an HTTP or Search module fetches the sitemap XML, a parsing step extracts the <loc> values, a filter drops anything outside the section you care about, and a final step writes the results into a Google Sheets tab on a weekly schedule.

Worth knowing about: several AI-focused scraping platforms now expose a dedicated "map" endpoint that returns a domain's discoverable URLs in one API call instead of requiring you to assemble sitemap parsing and link crawling yourself, and this has become a first step in a lot of AI-driven data pipelines.

No-code flows fit public sites with a clean sitemap and a use case that benefits from a recurring schedule. Once a site relies on JavaScript-rendered routing, needs custom logic, or the volume is large enough to route requests through a proxy API, a script, or a dedicated crawling service is a better tool.

How to crawl large sites reliably

When you crawl a large site, sending thousands of requests from one IP address quickly hits per-IP rate limits, and the crawl slows down or returns an incomplete list. Routing requests through a pool of rotating residential or datacenter proxies spreads them across many IPs, so a large crawl keeps a high success rate and finishes with a complete URL set.

This isn't unique to any one target site; it's how most servers protect themselves from a high volume of requests arriving from a single source in a short window. A proxy pool fixes this by distributing requests across many IPs instead of one, so no single address sends enough volume to trip a rate limit, keeping a steadier success rate across the whole crawl.

Which product fits depends on the job. Rotating residential proxies provides the broadest IP diversity for tougher targets that monitor traffic closely. Datacenter proxies are well-suited for fast, high-volume crawling of straightforward public content. ISP proxies combine residential-style trust with datacenter-level speed for long crawls on account-bound pages, and unlimited bandwidth proxies remove bandwidth as a constraint on very large, request-heavy jobs.

Each runs on pay-as-you-go or per-IP pricing with no monthly commitment, so it's realistic to scale a proxy pool up for one large crawl and back down once the job is done.

Which method should you use?

Pick by access and scale. If you own the site, start with Search Console and Analytics, then crawl to fill gaps. If not, start with a site: search and the sitemap, then run a crawler or script. Quick check: use operators. Audit: use a crawler. Large or repeatable job: use a script or API with a proxy pool.

Your situationStart hereThen
You own the site, smallSearch ConsoleCrawler to confirm
You own the site, largeSearch Console + AnalyticsScript + server logs for orphans
Outsider, quick checksite: operatorsSitemap
Outsider, auditSEO crawlersite: cross-check
Outsider, large or repeatCustom script or APIProxy pool for reliability

Access is the real fork in the road, more than site size or technical comfort. A site owner with Search Console access is already looking at the most accurate list Google has for that property, so the job is mostly filling in orphan pages. Method 5 alone can't surface. An outsider mapping a domain before a scraping or research project starts with public methods instead (site: search, sitemap) and scales up to a script with static residential proxies or another proxy type once the site is too large for a manual pass.

Why find all the webpages on a website?

A complete URL inventory earns its keep in a handful of recurring situations. SEO audits use it to surface broken links, duplicate content, and thin or orphaned pages that would otherwise sit unnoticed. Site migrations and redesigns depend on a full list to build clean redirects and confirm nothing important gets dropped.

A content refresh needs the same list to know what already exists before planning updates. Competitor and market research uses a mapped domain to read a competitor's structure and priorities from the outside, and any web scraping or data collection project starts with a URL list, since extraction can't begin until you know what pages exist.

Frequently Asked Questions

A Google site: example.com search is the fastest no-setup method, since it lists indexed pages instantly. For a complete list, pair it with the sitemap and a crawler like Screaming Frog.

Yes. Crawlers such as Screaming Frog and Sitebulb, plus no-code tools like n8n and Make, let you list and export a site's pages without writing any code.

No. It only shows URLs Google has indexed and chosen to display, so it misses orphan pages, unindexed pages, and dynamic routes. Treat it as a starting point.

Any crawler or script can export to CSV, which opens directly in Excel or Google Sheets. No-code flows can also write discovered URLs straight into a Sheet on a schedule.

Compare sources: crawl the site, then diff that list against Analytics, Search Console, and server logs. URLs that appear in traffic or logs but not in the crawl are orphan candidates.

Start with a site: search and the public sitemap, then run a crawler or a custom script. For large sites, route the crawl through a proxy pool so it stays reliable.

A single IP sending many requests hits per-IP rate limits, which slows the crawl and leaves the list incomplete. Distributing requests across rotating residential or datacenter IPs keeps the success rate high.

No, Google retired the cache: operator in 2023. Use the Wayback Machine for historical snapshots instead.

It depends on scale. Screaming Frog or Sitebulb cover most audits up to a few thousand pages, while a custom Python script is the better fit once a site is large, dynamic, or needs a proxy pool to finish reliably.

Crawl the site or parse its sitemap and count the unique URLs. Page-counter tools that read the sitemap give a quick estimate, but a crawl plus Search Console gives the most complete count.