Site index for AI agents: /llms.txt. Documentation pages under /docs are also served as markdown at the same URL plus .md (e.g. /docs/geo-audit.md) or via an Accept: text/markdown request.

All insights
Site Architecture

Site Architecture for SEO Planning for Millions of Pages

Sophia Satapathy·, updated

Managing millions of pages is more than having a tidy menu or a neat URL structure. Filters, templates, redirects, pagination and changing inventory at scale can easily create millions of useless URLs. 

In this blog, we’ll examine how to construct a controlled site architecture to 

  1. Ensure vital pages are findable, minimize duplication, 
  2. Allow efficient crawling 
  3. Be sustainable as the site scales

What Should You Define Before Building a Large Site Architecture?

Begin by finding all the different types of pages the platform may produce. Define the reason for the existence of each page class, whether or not users seek it, what makes it meaningfully distinct, how users and crawlers find it, when it should become indexable and when it should disappear.

This creates an indexable inventory contract that gives SEO, product, engineering, and content teams a common framework to interact with URLs.

Page classIndex whenDiscovery pathControl
Entity detailThe entity is active and has sufficient unique, useful informationCategory hubs, related entities, sitemapSelf-canonical; 404 or 410 after permanent removal
Category or topic hubIt represents durable demand and supports a coherent setGlobal or sectional navigationStable copy, curated links, indexable pagination
Facet landing pageThe combination has proven demand and enough matching inventoryCurated hub links, dedicated sitemapAllow listed combinations only
Internal search resultsUsually neverUser search interfaceKeep out of indexable inventory and avoid crawl traps
Comparison or programmatic pageInputs create a distinct decision-support experienceRelevant detail and hub pagesQuality threshold and duplicate detection
Expired or unavailable itemOnly while it serves a durable user needPrior links and related replacementsClear lifecycle rule; avoid soft 404s

A page class should not be approved simply because the system can generate it. It needs a defensible search purpose and clear ownership.

Programmatic publishing can multiply both useful pages and near-identical pages, so architecture cannot compensate for content that lacks a distinct purpose.

Before approving a page type, check these four areas:

  • Demand: Does it resolve a recurring user task or query family?
  • Distinctness: Does it contain information or choices that materially differ from nearby pages?
  • Depth: Can crawlers reach it through stable links rather than only through a search form?
  • Durability: Will the URL remain useful long enough to earn and retain signals?

How Should You Structure Millions of Pages?

How Should You Structure Millions of Pages?

Large websites need a shallow, layered structure that gives important pages predictable routes from strong hubs. Google has explained that links help it understand site structure and that both the number of links required to reach a page and the links pointing toward it can provide signals about relative importance.

A practical architecture can use four layers:

  1. Entry layer: The homepage and durable sectional gateways.
  2. Hub layer: Category, topic, geography, brand, or use-case pages.
  3. Collection layer: Subcategories and approved facet landing pages.
  4. Detail layer: Products, listings, profiles, documents, or articles.

The goal is to measure click depth by page class instead. Frequently changing pages should have strong hub links, crawlable pagination, and very low orphan rates.

How Should You Design the URL Structure?

Decide the resource model before deciding the exact URL syntax. A readable URL can help people and teams understand a page, but folder names alone do not determine site structure.

First define what makes two resources the same, which attributes can change without creating a new resource, and which variations deserve separate URLs. Google has noted that it generally relies on page linkages rather than URL structure alone to understand site relationships.

For large websites, keep URL rules simple and consistent

  • Use one normalized URL for each indexable resource
  • Keep URLs lowercase with stable separators
  • Avoid session identifiers and analytics parameters in internal links
  • Use deterministic parameter ordering when parameters are necessary
  • Do not rely on URL fragments for content that needs indexing
  • Maintain redirect and lifecycle rules for moved, merged, expired, or deleted resources
  • Avoid changing millions of URLs simply to make them look cleaner

A URL migration should solve a measurable discovery, duplication, maintenance, or user problem. Otherwise, the migration cost may outweigh the benefit.

How Can You Control Faceted Navigation?

How Can You Control Faceted Navigation?

Faceted navigation is one of the easiest ways for a finite catalogue to become an enormous crawl space.

A combination of color, size, brand, price, availability, and sorting can generate far more URL combinations than the underlying inventory. Google identifies faceted navigation as a common cause of overcrawling and slower discovery.

The key is to separate user filtering from search landing pages. Users may need many filter combinations, while search engines may only need a carefully selected subset.

Create an allowlist based on demand, inventory depth, uniqueness, and business value.

Facet outcomeRecommended treatmentWhy
Durable, demanded combinationCreate a stable landing page, self-canonicalize, add unique context, and link from hubsIt behaves like a genuine collection rather than a temporary filter
Useful filter with no search valueKeep available to users but prevent uncontrolled crawl discoveryPreserves the user experience without multiplying indexable inventory
Sort, view, session, or tracking stateNormalize internal links and prevent indexingChanges presentation or attribution rather than the resource
Empty or impossible combinationReturn a true 404 where appropriatePrevents endless traversal of invalid states

Canonical tags should not be treated as the main solution to uncontrolled discovery. Google treats the tag -rel=”canonical” as a strong signal, but duplicate URLs may still need to be crawled before that signal can be processed. Use canonicals to consolidate legitimate variants while preventing unnecessary URLs from being generated and linked in the first place.

How Should Pagination Support Crawling?

Infinite scroll can work as a user interface, but it should not be the only discovery mechanism.

Googlebot does not scroll or click a load-more button. Google recommends linking paginated pages sequentially through crawlable URLs so that crawlers can discover the full set. Each section of a large collection should therefore have a persistent URL and standard links.

Do not automatically canonicalize every paginated page to the first page when later pages contain different items. Each page should normally remain independently addressable and linked.

For very large collections, provide shortcuts so crawlers do not have to move through thousands of pages to reach important inventory. Important pages should also receive links from relevant hubs.

How Should Templates Create Useful Pages at Scale?

On a million-page website, a template becomes an editorial policy executed across thousands or millions of URLs. It determines whether a page answers a real question or simply displays database fields.

Every template should define:

  • The minimum useful information
  • The decision the page supports
  • Its relationship with nearby pages
  • What happens when the required data is missing

A detail page might require a clear entity name, defining attributes, availability or status, original descriptive information, primary media with descriptive alt text, relevant provenance, and links to logical parent pages and alternatives.

A category page needs more than a product grid. It should explain the collection, expose meaningful subgroups, and help users choose where to go next.

How Should You Use XML Sitemaps on a Large Site?

A sitemap supports discovery, but it does not replace site architecture. Google treats sitemaps as hints and limits a single sitemap to 50,000 URLs or 50 MB uncompressed. At a large scale, segmentation becomes important because it makes the sitemap system easier to monitor.

Group sitemap files by meaningful characteristics such as:

  • Page class
  • Locale
  • Region
  • Freshness tier
  • Publication cohort

Useful segments could include products by market, category hubs, editorial content, URLs added in the last seven days, or recently updated entities.

This makes Google Search Console reporting more useful. If one segment performs differently, teams can investigate a specific group rather than an undifferentiated pool of millions of URLs.

Only include canonical, indexable URLs that return successful responses. Use the tag “lastmod” when meaningful content changes rather than updating it with every routine build. Keep sitemap counts aligned with the source inventory and watch for unexpected additions or removals.

When Does Crawl Budget Matter?

Crawl budget is not a universal optimization target. It becomes more relevant when a site is very large, changes frequently, or has a significant number of URLs that are discovered but not indexed.

Google’s current guidance primarily targets sites with more than 1 million unique pages that change moderately frequently, sites with more than 10,000 pages that change very rapidly, and sites with a large share of URLs classified as “Discovered - currently not indexed. 

Instead of asking, “How can we increase crawl budget?” ask:

“Why is the crawler spending requests on URLs we do not value, and why are valuable URLs difficult to reach?”

Common causes include:

  • Duplicate parameters
  • Slow or unreliable infrastructure
  • Redirect chains
  • Soft 404s
  • Unstable content
  • Weak internal linking

Server logs show what crawlers actually request. Search Console helps summarize outcomes. Combining both with a canonical inventory lets teams classify requests as valuable, duplicate, invalid, redirected, or unknown.

What Should a Large-Site Architecture Dashboard Track?

A useful dashboard should connect technical signals with specific questions and actions.

SignalQuestion it answersAction threshold example
Indexable inventoryHow many URLs should exist by page class?Unexpected change beyond the release plan
Orphan rateWhich approved pages have no crawlable internal links?Sustained orphaning of priority classes
Depth distributionHow far are valuable pages from strong hubs?Long tail grows after a navigation or pagination change
Crawler requests by patternWhere is crawl activity being spent?Duplicate or invalid states rise materially
Response-code mixAre errors and redirects consuming requests?5xx spike, recurring 3xx chains, or soft 404 pattern
Indexed-to-submitted ratioWhich sitemap cohorts are being selected?Material divergence between similar cohorts
Canonical disagreementIs Google choosing different representatives?The pattern appears across a template or URL family
Discovery latencyHow long until new priority pages are crawled?SLA missed for high-value cohorts

The purpose of the dashboard is not to collect more metrics. It is to make URL-pattern problems visible early enough for teams to act.

How Should You Launch a New Architecture Safely?

How Should You Launch a New Architecture Safely?

Avoid releasing a new architecture across millions of pages as a single event. Start with a representative cohort that is large enough to reveal systemic issues but small enough to reverse.

Compare the test cohort with a control group across crawl patterns, canonical selection, discovery latency, indexation, organic entrances, and server load.

A practical rollout can follow these steps:

  1. Inventory the current URLs using crawls, sitemaps, analytics, Search Console, and server logs.
  2. Define the target page-class contract and map legacy patterns to keep, merge, redirect, noindex, block, or retire.
  3. Implement the new templates and rules for one representative cohort.
  4. Test rendered output and crawl paths in staging without exposing staging URLs to search engines.
  5. Release the cohort and monitor logs, indexing, and pattern-level failures.
  6. Expand by page class or market with explicit rollback criteria and inventory reconciliation after every wave.

The migration is complete when old URL patterns no longer receive internal links, redirects resolve in a single hop, sitemaps contain the new canonical inventory, and monitoring can explain remaining legacy traffic.

How Can You Keep Site Architecture From Breaking Over Time?

Large websites often develop problems after they launch the original architecture. New filters, markets, templates, and campaign parameters can create URLs without following the original rules.

Architecture should therefore become part of change management, with clear owners across SEO, product, engineering, content, analytics, and infrastructure.

Before releasing anything that can create URLs, answer five questions:

  1. How many URLs can this feature generate?
  2. Which URLs are indexable?
  3. How are those URLs linked?
  4. What are their canonical and lifecycle rules?
  5. How will unexpected URL growth be detected?

If a team cannot estimate the upper limit of a URL-generating feature, the feature is not ready for a large site.

Maintain a URL pattern registry, template requirements, approved facet matrix, redirect standards, sitemap ownership, and automated release checks. Review the inventory quarterly and after major catalog, taxonomy, rendering, or routing changes.

Conclusion 

Clear architecture can also improve the conditions under which brand information is found, interpreted, and cited beyond traditional search.

Strong hubs establish context. Stable entity pages reduce ambiguity. Descriptive internal links clarify relationships. Consistent canonical signals make it easier to identify the preferred version of a resource.

Architecture does not guarantee inclusion in an AI-generated answer. It does, however, create a cleaner and more coherent body of information for systems that need to find and interpret your content.

Brands should be thinking about developing long-term discoverability, not just more pages. A big website is only valuable if its pages are informative, findable, and match the proper search intent. Tools such as AirPulse can help organizations monitor brand awareness across high-intent AI search and social conversations. This provides brands with more insight into how consumers find them and relates SEO efforts to measurable visibility rather than just page count.

Frequently Asked Questions

How Flat Should a Million-Page Site Be?

It should be structured so important page classes have reliable routes from strong hubs without making global navigation unnecessarily complex.

Measure depth by page class rather than applying one arbitrary click-depth target to every URL. Prioritize valuable and frequently changing inventory while keeping long-tail pages accessible through reliable crawl paths.

Do XML Sitemaps Solve Orphan Pages?

No. Sitemaps can help search engines discover URLs, but internal links communicate relationships and importance.

An indexable URL that appears only in a sitemap should therefore be treated as an architectural warning, particularly when the page is important to the site.

Should All Filter Pages Be Blocked?

No. Some filter combinations represent genuine, durable search demand and can become useful landing pages.

The better approach is controlled eligibility. Allow valuable combinations to become curated landing pages while preventing temporary, low-value, or unnecessary filter states from creating uncontrolled indexable inventory.

See what AI says about your brand.

Run the free score, then watch AirPulse fix what it finds. Nothing ships without your approval.