vrid.ai Logo

Programmatic SEO in 2026: what still works

Programmatic SEO still works in 2026 if pages carry unique data. Here is the thin-page risk model, the templates that survive.

23 min read

Programmatic SEO in 2026: what still works

TL;DR: Programmatic SEO still works in 2026, but only for pages built on unique, verifiable data. Google’s scaled content abuse policy, live since March 2024, targets pages generated “for the primary purpose of manipulating search rankings and not helping users,” and it does not care whether a human or an AI wrote them (Google Search Central). Zapier’s 800,000+ integration pages and Wise’s currency-conversion pages both rank because every page carries proprietary or public data a competitor cannot copy in bulk (Ahrefs). The templates that fail are variable-substitution pages with no unique facts. The ones that survive answer one distinct query with data nobody else has assembled.


Table of contents

What programmatic SEO actually is

Programmatic SEO is a template plus a dataset. You build one page structure, feed it rows from a database, and publish one URL per row. A directory of 5,000 cities, a comparison page per product pair, a currency-conversion calculator for every currency pair: all the same pattern, all generated at scale instead of written one page at a time.

The mechanism has not changed since programmatic SEO became a recognized category around 2020. What changed is the enforcement. Google shipped a named spam policy for this exact pattern in March 2024, and it has since become one of the most frequently cited reasons for manual actions and ranking collapses in 2025 and 2026 search forums.

This matters more now that the 9 SEO KPIs that still matter when clicks are disappearing increasingly favor citation share and answer visibility over raw pageviews. A programmatic page that gets suppressed for scaled content abuse loses both the click and the citation opportunity at once.

John Mueller, Google’s Search Relations lead, has been blunt about the category’s reputation: “Programmatic SEO is often a fancy banner for spam,” he said, as reported by Ahrefs. That is not a blanket condemnation. It is a warning that the tactic gets used, more often than not, to disguise thin content as scale. Whether your program survives depends entirely on what fills the template, not on the fact that you used one.

The thin-page risk model

Before you write a single template, score your planned pages against four questions. If you answer no to two or more, the page is high risk.

1. Does the page contain data nobody else has assembled in this exact form? A page listing “SEO agencies in Denver” scraped from a business directory is not unique. A page showing your own verified pricing survey of 40 Denver agencies is.

2. Would a single user reading this one page, with no other page on your site open, get a complete answer? If the page requires jumping to three other pages to make sense, it is thin by definition.

3. Does the template produce meaningfully different content across rows, or just different nouns in the same sentences? “Best [City] plumbers” repeated 3,000 times with the city name swapped is variable substitution. Google’s spam policy names this pattern directly: “using generative AI tools or other similar tools to generate many pages without adding value for users” (Google Search Central).

4. Can you defend the page’s existence to a human reviewer without using the word “scale”? If your only answer for why the page exists is “we needed more URLs,” it fails.

flowchart TD
    A[New programmatic template idea] --> B{Unique data per page?}
    B -->|No| C[Kill it: variable substitution]
    B -->|Yes| D{Answers one query completely alone?}
    D -->|No| E[Merge into a parent page instead]
    D -->|Yes| F{Content differs meaningfully row to row?}
    F -->|No| C
    F -->|Yes| G{Defensible without saying 'scale'?}
    G -->|No| C
    G -->|Yes| H[Build it, with unique data sourcing]

Programs that fail this model at more than 20-30% of planned pages should not launch. Google’s scaled content abuse detection is not page-by-page; it evaluates the pattern across your whole template. A minority of strong pages does not protect a majority of thin ones once the classifier flags the template family.

What Google’s scaled content abuse policy actually says

Google folded scaled content abuse into its formal spam policies in March 2024, alongside two other new categories: expired domain abuse and site reputation abuse (Google Search Central Blog). The policy text, from Google’s own spam policy documentation, defines it this way:

“Scaled content abuse is when many pages are generated for the primary purpose of manipulating search rankings and not helping users” (Google Search Central).

Google lists five specific patterns under this policy:

  • Using generative AI or similar tools to generate many pages without adding value for users
  • Scraping feeds, search results, or other content to generate many pages, including through synonymizing, translating, or other obfuscation
  • Stitching or combining content from different web pages without adding value
  • Creating multiple sites with the intent of hiding the scaled nature of the content
  • Creating many pages where the content makes little or no sense to a reader but contains search keywords

Two details matter for anyone running a programmatic program. First, the policy is explicitly method-agnostic: “no matter whether content is produced through automation, human efforts, or some combination of human and automated processes” (Google Search Central). Hiring writers to fill out a thin template does not exempt you. Second, the policy sits alongside Google’s separate guidance on AI-generated content, which states that using automation is not itself against Google’s guidelines. The violation is the absence of value, not the presence of a script.

The enforcement was not cosmetic. Google reported that by April 19, 2024, the March core update and its accompanying spam policies had cut low-quality, unoriginal content in search results by 45%, ahead of the 40% reduction it had initially targeted (Google Search). Elizabeth Tucker, Google’s Director of Product for Search, described the intent behind the change as ranking-system work, not a one-off penalty sweep: “We’re making algorithmic enhancements to our core ranking systems to ensure we surface the most helpful information on the web and reduce unoriginal content in search results,” as reported by Search Engine Journal. A program built on the assumption that thin pages just rank a little worse is working from an outdated read of the policy’s actual reach.

The templates that survived: real examples with real data

Four programmatic SEO programs get cited repeatedly as evidence the category still works, because each one is built on data or output a competitor cannot trivially replicate.

Zapier runs an estimated 800,000+ programmatic pages, one per app-to-app integration pairing, and pulls roughly 306,000 monthly organic visits from that page set alone (Ahrefs). Each page documents a real, working integration between two specific tools, with setup steps that differ because the actual integration mechanics differ. Nobody else has assembled Zapier’s specific catalogue of app-pair connections at that scale.

Wise (formerly TransferWise) runs an estimated 14,888 pages that together generate 4,667,719 monthly pageviews, largely from currency-conversion and international-transfer pages (Ahrefs). The underlying data, live exchange rates and fee comparisons, updates constantly and is genuinely useful on its own, independent of the surrounding SEO strategy.

Nomadlist runs roughly 25,873 pages generating about 41,200 monthly organic visits, built on a proprietary dataset of cost-of-living, internet speed, and safety scores per city, collected and maintained by the site itself (Ahrefs).

Webflow runs an estimated 31,516 website-template pages that together generate about 27,600 monthly organic visits (Ahrefs). Each page is not a description of a template, it is the template: a working, previewable site a visitor can inspect and clone. The page and the product are the same object, which is a harder pattern to fake at scale than a written description ever could be.

The pattern across all four: proprietary or continuously updated data, one page answering one specific query in full, and a template that produces genuinely different content per row because the underlying facts are genuinely different. None of the four would pass as “scaled content abuse” under Google’s own examples, because none of them fit the “little to no value” test.

For teams running an AI SEO workflow from keyword to published post, the lesson translates directly: the automation layer is fine. The dataset underneath it is what Google is actually evaluating.

The templates that died

The other side of the ledger is just as instructive. Sites publishing thousands of near-identical pages through pure template automation, with no unique dataset behind the variable, have reported ranking losses in the 60-90% range following core updates that specifically targeted scaled content, according to multiple 2026 industry retrospectives on the pattern.

The failure mode is consistent across cases: a template like “[Service] in [City]” repeated across a list of cities pulled from a public database, with no city-specific facts beyond the name itself. Google’s classifier does not need to catch every page individually. Once enough pages in a template family match the “little to no value” pattern, the whole family gets suppressed together.

A related failure: sites that scraped competitor listings or search results and republished them with light rewording. Google’s policy names this explicitly as “scraping feeds, search results, or other content to generate many pages… where little value is provided to users” (Google Search Central). Synonymizing or translating scraped content does not launder it; the policy calls out obfuscation by name.

How to build a programmatic SEO page in 2026

Step 1: Pick a dataset before you pick a keyword. Start from what you can uniquely assemble, not from a keyword list. If your only dataset is “list of cities,” you already have a thin-content program.

Step 2: Require at least three distinct facts per page. A page needs more than one number to differentiate it from its siblings. Wise’s pages carry live rates, historical trend context, and fee breakdowns; that is three independent facts, not one repeated across a template.

Step 3: Write the template to fail gracefully on missing data. If a row is missing two of your three required facts, do not publish that page. A partial dataset produces a partial-value page, and partial value is what the scaled content policy targets.

Step 4: Add a human review gate before publish, not after indexing. Sample 5-10% of generated pages before the batch goes live. Reviewing after Google has already crawled and classified the template family is too late to matter.

Step 5: Launch in batches, not all at once. Publishing 50,000 pages in a single sitemap submission gives Google’s spam classifiers a single, obvious pattern to evaluate. Staged rollouts across weeks give each batch room to be judged on its own signals before the next batch lands.

Step 6: Monitor at the template level, not the page level. Track average position and indexation rate for the whole template family weekly. A template-wide decline is the earliest signal that the pattern, not an individual page, has a value problem.

Step 7: Set a re-verification cadence for time-sensitive data. Pages built on rates, prices, or inventory that change go stale fast, and a stale proprietary dataset degrades into the same “little value” pattern Google’s policy targets, just on a delay. Wise’s currency pages work because the rate feed updates continuously; a currency page with a rate from six months ago is functionally a thin page with extra steps. Put a data-freshness check on a schedule, not on a “someone will notice” basis.

Measuring ROI on a programmatic SEO program

Programmatic pages are cheap per unit and expensive in aggregate once you count engineering time, data licensing or collection, and ongoing maintenance. Judge the program the same way you would judge any content investment: revenue or qualified traffic per page, divided by the fully loaded cost of the template, the dataset pipeline, and the review process, not just the cost of generating the HTML.

A useful gut check: if you removed the bottom 20% of pages by traffic from a template family and total program revenue barely moved, those pages are dead weight carrying real ranking risk for no offsetting return. Cull them or fix their underlying data gap rather than leaving them live because deleting a URL feels like admitting the program failed. Teams running a full AI content ROI calculation on a programmatic set should treat each template family as its own line item, since a strong-performing template and a weak one sitting under the same domain can mask each other’s real numbers in a blended report.

Data sourcing: proprietary vs public vs scraped

Data sourceDefensibility✓ / ✗Example
Proprietary (your own survey, transaction, or usage data)Highest: nobody else has itWise’s live rate feed
Public but assembled uniquely (government data cross-referenced with your own scoring)Moderate: replicable, but you got there first and add synthesisNomadlist’s composite city scores
Public and directly republished (census data with no added synthesis)Low: technically not scraping, but adds littleA raw population table with no commentary
Scraped from competitors or search resultsLowest: named directly in Google’s policyReworded competitor listings

Public data is not disqualifying on its own. The test is whether you added synthesis, verification, or a scoring layer a reader cannot get by visiting the raw source directly. Cross-referencing three public datasets into one comparison a reader would otherwise have to build themselves clears that bar. Copying one public dataset into a page shell does not.

Internal linking and crawl budget at scale

Programmatic programs live or die on how Google’s crawler prioritizes them. A 5,000-page template with no internal linking hub gets crawled sporadically and indexed slowly, if at all. Build a hub page that links to every generated page, and link back from each generated page to its category hub and to 3-5 sibling pages sharing a real attribute (same country, same integration category, same price band).

This is the same discipline covered in building topical authority when everyone publishes daily: a dense, logically organized internal link graph tells Google’s crawler which pages you consider important and helps it discover new pages in the template without waiting for a full site crawl cycle.

Do not orphan pages. A generated page reachable only through a sitemap XML file, with zero internal links pointing to it, sends a weak relevance signal regardless of the content quality on the page itself.

The pre-launch audit checklist

Run this against a sample of your generated pages before publishing at scale:

  • Does each sampled page contain at least 3 facts specific to that row, not shared across the template?
  • Would the page make sense to a reader with zero context from other pages on the site?
  • Is the underlying data updated on a defined cadence, or is it a one-time scrape?
  • Does the template produce content that changes by more than the variable itself (not just city name swapped, but city-specific numbers, facts, or context)?
  • Is there a working internal link path from your homepage to every generated page within 3 clicks?
  • Have you removed rows where the dataset is incomplete, rather than publishing thin versions?
  • Does the meta title and H1 differ meaningfully across pages, not just in the variable slot?
  • Have you checked Google’s own guidance on AI-generated content against your actual generation process?

Any unchecked box on a sample of 20+ pages is a signal to fix the template before scaling further, not after. If the audit turns up a mix of strong and weak pages within one template, treat that the same way you would a content refresh vs new content decision: fix the weak rows with better data where the underlying facts exist, and cut the rows where they do not, rather than leaving a blended template live and hoping the strong pages carry the weak ones.

How AI search changes the calculus

Programmatic pages built on real data have one advantage AI search engines reward more than traditional search does: citability. A page with a specific, sourced number (Wise’s live exchange rate, Nomadlist’s specific city score) is easier for an AI answer engine to lift a fact from than a page that only restates a keyword. If your keyword research process for a programmatic template surfaces queries with a factual, single-answer shape, that is a signal the template fits AI-citation patterns as well as traditional ranking. Vrid.ai’s keyword research tooling can help identify which query clusters in a programmatic set have that factual shape versus which ones are better served by a single comprehensive guide page instead of a template.

This does not change the underlying rule. AI Overviews and chat-based answer engines pull from pages Google’s own index already considers valuable. A thin programmatic page that fails Google’s scaled content policy is not going to get cited by an AI engine that Google itself feeds, either.

Common mistakes that trigger manual actions

Publishing before the dataset is complete. Teams under deadline pressure often launch a template with 60% of rows fully populated and 40% thin. The classifier evaluates the family, so the thin 40% drags the whole set.

Treating AI generation as a shortcut around research. Feeding a language model nothing but a city name and asking it to “write 500 words about SEO in [city]” produces exactly the pattern Google’s policy names: generative AI generating many pages without adding value.

Skipping the internal review sample. Programs that never spot-check a random sample of live pages tend to discover quality drift only after a ranking drop, when the fix is much more expensive than a pre-launch catch.

Reusing the same meta description template verbatim. Duplicate meta descriptions across thousands of pages are a weak signal on their own, but combined with thin body content, they reinforce the pattern classifiers are trained to catch.

Ignoring crawl budget on large sites. A 100,000-page programmatic launch on a site with no internal linking discipline can starve your existing, higher-value pages of crawl frequency while Google works through the new URLs.

Running the whole program at a pace no reviewer can keep up with. This is the same trap covered in content velocity vs quality: how many posts per month actually works: a publishing cadence that outruns your review capacity produces a quality gap regardless of whether a human or a script wrote the content. A programmatic template with 50,000 rows and one part-time reviewer is going to ship thin pages by default.

Frequently asked questions

Is programmatic SEO dead in 2026?

No. Programs built on unique, verifiable data still rank and grow traffic, as Zapier’s 800,000+ integration pages and Wise’s currency pages demonstrate (Ahrefs). What is dead is variable-substitution content with no unique data behind each page. Google’s scaled content abuse policy targets the value gap, not the automation itself.

What is scaled content abuse, exactly?

It is a Google spam policy, live since March 2024, defined as generating many pages primarily to manipulate rankings rather than help users, regardless of whether automation, human writers, or a mix produced them (Google Search Central). It covers AI-generated pages with no added value, scraped-and-reworded content, and stitched content from multiple sources.

Does using AI to generate programmatic pages automatically violate Google’s policy?

No. Google’s own guidance states that using automation, including AI, is not against its guidelines. The violation trigger is producing pages “without adding value for users,” not the method of production (Google Search Central).

How much unique data does each programmatic page need?

There is no official Google number. The pattern across surviving programs like Wise and Nomadlist is 3 or more distinct, row-specific facts per page, beyond the templated variable itself, combined with data that updates or was assembled uniquely rather than copied wholesale from one public source.

Can I use public government data for a programmatic template?

Yes, if you add synthesis, scoring, or cross-referencing the raw source does not already provide. Republishing a public dataset with no added value is a weaker signal than combining several public sources into a comparison a reader could not easily build themselves.

What happened to sites that got hit by scaled content abuse penalties?

Reported outcomes vary by source, but multiple 2025-2026 industry retrospectives describe ranking losses in the 60-90% range for template families flagged under the policy, sometimes with a formal manual action visible in Search Console and sometimes as an algorithmic ranking suppression with no manual action notice.

How do I know if my existing programmatic pages are at risk?

Sample 20-30 pages at random from the template family and run them against the four-question risk model above. If two or more of the four questions get a no answer on more than 20-30% of the sample, the template carries risk regardless of current rankings.

Should I noindex thin pages in an existing programmatic program instead of deleting them?

Noindexing removes the immediate ranking risk from those specific URLs but does not fix the underlying value gap. If the pages genuinely serve users despite thin SEO value, keep them noindexed. If they serve no purpose beyond the URL count, deleting them and redirecting to a relevant hub page is usually the better long-term move.

Is Zapier’s integration page strategy replicable for a smaller company?

The pattern is replicable; the scale is not immediately. Zapier’s advantage is a genuinely proprietary catalogue of working integrations built over years. A smaller company can apply the same principle, proprietary or hard-to-replicate data per page, at a smaller scale: dozens or hundreds of pages instead of hundreds of thousands.

What is the difference between programmatic SEO and a normal content template?

Every blog uses templates (title, body, author byline). Programmatic SEO specifically means generating a large number of URLs from a dataset with minimal manual writing per page. The distinction that matters to Google is not the presence of a template but whether the data filling it is unique and valuable per row.

Does programmatic SEO still work for local business directories?

Only with verified, differentiated listings. A directory scraped from public business registries with no verification, reviews, or added context matches Google’s “scraping… to generate many pages… where little value is provided” pattern directly. A directory with verified contact details, real review aggregation, and location-specific context performs differently.

How long does it take to see if a programmatic template is working?

Give a batch 8-12 weeks of indexation and ranking stabilization before judging it. Google’s classifiers evaluate pattern behavior over time, not on day one, so early rankings on a new template can look promising or alarming without reflecting the eventual outcome.

Can I recover a template family after a scaled content abuse penalty?

Recovery is possible but requires fixing the underlying value gap, not just requesting reconsideration. Google’s guidance ties reconsideration to actually removing or fixing the violating pages, not appealing the classification while the pages remain unchanged.

What is the biggest difference between 2020-era programmatic SEO and 2026-era programmatic SEO?

In 2020, variable-substitution templates with thin content regularly ranked because detection was weaker. In 2026, the named scaled content abuse policy, live since March 2024, gives Google’s classifiers an explicit target, and industry reports describe steep, fast ranking losses once a template family gets flagged.

Do comparison and “X vs Y” programmatic pages still work?

Yes, when the comparison data is real and specific per pair, not a templated sentence with two product names swapped in. A comparison page needs distinct facts about each specific pairing, pricing, feature differences, use-case fit, not a generic comparison structure repeated across every possible pair.

How does keyword research fit into building a programmatic SEO template?

Keyword research should identify the query pattern (the “shape” of demand: “X vs Y”, “[service] in [city]”, “[currency] to [currency]”) before you decide whether a dataset exists to fill it. Vrid.ai’s keyword research tooling can surface these patterned query clusters, which is the first check before committing engineering time to a template.

Is there a minimum page count for programmatic SEO to be worth building?

No fixed minimum. The economics depend on the value of each generated page’s traffic relative to the engineering cost of the template and dataset, not a target page count. Building 200 well-differentiated pages beats building 20,000 thin ones, both for ranking risk and for actual traffic quality.

What tools do I need to build a programmatic SEO program?

At minimum: a structured dataset (a database or spreadsheet with verified, per-row facts), a templating system to generate pages from that data (a CMS with dynamic templates, or a static site generator), and a crawl/index monitoring setup to track the template family’s performance as a group, not just individual pages.

Does programmatic SEO help with AI search visibility, or only traditional Google rankings?

Pages with specific, sourced facts per row are easier for AI answer engines to cite than pages restating a keyword with no concrete data. The underlying requirement is the same: real, verifiable, row-specific information, which both traditional ranking and AI citation reward for the same reason.

What is the single most important factor in whether a programmatic SEO template survives?

Whether each generated page would still be worth visiting if the reader had zero interest in your product and no context from any other page on the site. Every surviving example above, Zapier, Wise, Nomadlist, Webflow, passes that test. Every failed example in industry reports does not.

Key takeaways

  • Google’s scaled content abuse policy, live since March 2024, targets pages generated primarily to manipulate rankings rather than help users, and it applies regardless of whether AI, humans, or both produced the pages.
  • Surviving programmatic programs (Zapier, Wise, Nomadlist, Webflow) share one trait: proprietary or uniquely assembled data behind every page, not a variable-substitution template.
  • Run the thin-page risk model against a sample before launch: unique data, standalone completeness, meaningful content variation, and a defensible reason beyond “we needed more URLs.”
  • Monitor at the template-family level, not the individual page level. Google’s classifiers evaluate the pattern, so a minority of strong pages does not protect a majority of thin ones.
  • Public data is usable if you add synthesis or scoring; republishing or scraping it directly is named explicitly in Google’s policy as a violation pattern.

Building a programmatic SEO template starts with the dataset, not the URL count. Audit your current pages against the risk model above before you write another line of template code, and if the data behind a planned page set does not exist yet, build the data pipeline first.

Related Posts