I Published 144 Templated Pages, Then Audited Them Against Google's Spam Policy
There are 144 tool pages on this site. Every single one of them has the same four headings, in the same order: what it is, key features, who it's for, pricing. The body copy averages 222 words. The shortest is 169. The longest is 286.
I wrote all of them. Not a scraper, not a feed, not a loop over an API response. But if you crawled this domain and looked at nothing except the shape of those URLs, you would see a template repeated 144 times across 26 categories, which is exactly the pattern Google's scaled content abuse policy was built to catch.
So which one is it? Programmatic SEO that works, or scaled content abuse that gets a site quietly demoted eleven months from now? The honest answer is that the distinction doesn't live where most posts about programmatic SEO put it, and the test I ended up running on my own pages is one I'd have wanted before publishing them, not after.
What is programmatic SEO, actually?
Publishing many pages that share a structure, where the substance of each page changes but the skeleton doesn't.
Zapier's integration pages are the canonical example: one template, tens of thousands of URLs, each combining two real apps that genuinely connect. TripAdvisor's hotel pages. Wise's currency pages. A directory's category pages. My tool pages. Same idea at wildly different scales.
The thing people get wrong is treating programmatic SEO as a technology decision. It isn't. Whether you generate the pages with a script, a CMS loop, an LLM, or by typing all 144 of them by hand is irrelevant to how Google treats them. I'll come back to that, because it's the single most misunderstood part of the policy and it's the reason my hand-written pages don't get a free pass.
Does programmatic SEO still work in 2026?
Yes, and the qualifier matters more than the answer.
The pattern that still works is unique data per URL. Not unique wording, unique data: something on that page that a reader can't get by asking an AI assistant the same question, and that doesn't exist in the same form on the other 143 pages. Currency rates. Real reviews. An integration that actually exists. A specific, opinionated read on who a tool is wrong for.
The pattern that stopped working is the one everybody built first, where a single template gets filled with a city name, a tool name, or a keyword variant and shipped a thousand times. That version is dead, and it's dead in a specific way that's worth understanding: not banned, just outcompeted into invisibility by the thing that made it cheap to produce in the first place.
Here's the framing I keep coming back to, from Growth Engineer's writeup on the penalty question: if ChatGPT can reproduce your page in ten seconds, Google has no reason to rank it. That's not a policy argument, it's an economics argument, and it applies whether or not any spam update ever touches your domain. A page whose entire value can be regenerated on demand by a model the reader already has open is not a page anyone needs to visit.
The reason directories still work despite all of this is that they aren't only competing on text. A directory is a browsing surface. Someone landing on a category page is scanning a set, comparing options, and making a shortlist. That's a task, and a page that helps you complete a task is a different kind of object than a page that recites facts. I'll get to whether mine actually does that.
Will programmatic SEO get you penalized? What Google actually says
Google has never banned programmatic SEO, and every post claiming otherwise is arguing with a policy that doesn't exist. What exists is the scaled content abuse policy, and it's worth reading the actual wording rather than someone's summary of it. From Google's spam policies:
Scaled content abuse is when many pages are generated for the primary purpose of manipulating search rankings and not helping users.
Every load-bearing word in that sentence is about purpose. "Primary purpose of manipulating search rankings." "Not helping users." There is no number in it. There is no template-to-unique-content ratio. There's no page count above which you're in trouble.
The examples underneath it are more specific and all point the same direction: generating many pages with AI tools without adding value, scraping feeds or search results into pages, stitching content from other pages together, spinning up multiple sites to hide how scaled the content is, and generating pages that don't make sense to a reader but contain the keywords.
Notice what's missing. Nothing in there is about automation as such. That's deliberate. In the March 2024 core and spam update, Google widened this policy to cover content produced by "automation, humans, or a combination," and the phrase "no matter how it's created" got added specifically to close the loophole people were using, which was to argue that a human touched it so it can't be scaled abuse.
That's the sentence that applies to me. I hand-wrote 144 pages. Under the current policy, that fact buys me exactly nothing. The question is what's on them, not who typed them.
Two Google statements are worth having in your head, and they pull in opposite directions on purpose. John Mueller, in 2023: "Programmatic SEO is often a fancy banner for spam." Danny Sullivan, in 2024: "Any method that you undertake to mass generate content, you should be carefully thinking about it. There's all sorts of programmatic things, maybe they're useful."
Often, not always. Maybe useful. Both of those are hedges, and the hedge is the actual message. Google is telling you it evaluates the output, not the method, and that most of what it sees from this method is bad. That's a statement about the average, and the average is not a verdict on your specific pages.
What the August 2026 spam update actually hit
The theory is fine. What got demoted is more instructive.
Glenn Gabe's case studies from the August 2026 spam update, which rolled out August 18 and finished on the 21st, are the clearest picture I've found of what the algorithm is actually reacting to:
- A YMYL site lost visibility for over 200,000 queries. The mechanism was programmatic content scaled across countries, with AI-generated sections stitched into those pages. Two violations compounding, not one.
- An Amazon affiliate site lost around 14,000 queries. Gabe describes it as programmatically driven and turn-key: categories and product pages generated automatically from Amazon's own data, with almost no human involvement, and the content scraped from Amazon at scale.
- A site with 1.5 million indexed URLs dropped, with roughly 85% of those indexed pages sitting in its programmatic section.
- A site that had scaled to 250,000 URLs lost visibility for nearly 25,000 queries while redirecting users onward to more aggressive sites.
Look at the shape of that list. Every one of them is at a scale where a human could not have meaningfully touched each page, and in at least two cases the underlying data was somebody else's, republished. The affiliate site wasn't punished for using a template. It was punished for having nothing on the page except a template wrapped around Amazon's data.
Recovery, per Google's own guidance, takes months after significant changes, and most sites hit in previous rounds never made it back. Of roughly 400 heavily hit sites Gabe tracked through an earlier core update, only about 22% saw a 20% or greater lift afterward and 13% dropped further. That asymmetry is the real argument for getting this right before you publish, not after. The cost of being wrong isn't a bad quarter, it's a domain you may have to abandon.
My 144 pages, audited against the policy
Here's the part where I stop describing other people's sites.
Growth Engineer's five-question risk test is the sharpest version of this I found, so I ran my own tool pages through it honestly rather than looking for the answers I wanted.
Does each URL contain unique data unavailable elsewhere? Partially. The features and pricing on a page like Supabase's are restatements of public information, and that's the weakest part of every page I've written. What isn't available elsewhere in the same form is the "who it's for" section, which names who the tool is a bad fit for. That's a judgment call from someone who's shipped seven products across seven stacks, and it's the only part of the page a model can't confidently reproduce. It's also, tellingly, the shortest section.
Can users complete a task without leaving? On a single tool page, no. On a category page, yes, and that's where the directory actually earns its keep. Comparing 26 categories worth of tools against each other is a real task. Reading a 222-word summary of one of them is not.
Is the source data original rather than scraped or synonymized? Yes. Nothing here came out of an API or a competitor's listing. That's the question I'd pass most comfortably, and per the March 2024 wording it's also the one that matters least on its own.
Do users engage? I don't know, and I'm not going to pretend otherwise. This site is a few weeks old. I don't have the traffic to produce a meaningful dwell-time or pogo-stick sample, and inventing one would put me in the same category as everyone else quoting numbers they didn't measure. Same reason I won't quote a Domain Rating I haven't earned.
Do the pages link to and from hand-written pillars? Now they do. They didn't at first, and that was a mistake I only noticed while writing this section. A programmatic section floating with no editorial content pointing into it is the exact shape of a doorway structure, regardless of intent.
Three passes, one partial, one unknown. That's a more uncomfortable result than I expected going in, and the uncomfortable part isn't the template. It's the 222 words.
That average is the number I'd change if I were starting over. Not because word count is a ranking factor, it isn't, but because 222 words is roughly the amount of space in which it is impossible to say anything a reader couldn't have gotten faster somewhere else. The template isn't the risk. The template is a container, and mine is too small to hold the only thing that justifies the page existing.
The bigger risk isn't a penalty, it's never getting indexed
This is the part almost nobody writing about programmatic SEO leads with, and for a site your size it's overwhelmingly the more likely outcome.
Getting hit by a spam update requires Google to have indexed your pages, ranked them, and then decided to stop. The far more common failure mode is that it crawls them, shrugs, and never indexes them at all. That shows up in Search Console as "Crawled, currently not indexed," and it's not a bug or a queue you're waiting in. It's a verdict.
Reports through 2026 put 40% to 60% of total pages on many sites in some flavor of not-indexed status. Programmatic sections, thin category archives, and affiliate pages are named repeatedly as where that concentrates. One case worth the mental model: a programmatic glossary with 22,500 pages had 6,300 indexed, about 28%, with more than 16,000 stuck as crawled-not-indexed, before the owner cut and consolidated their way to a 94% indexing rate.
Read that backwards. Sixteen thousand pages of work produced nothing, and the fix was mostly deletion. That's the actual expected outcome of shipping a thousand thin pages: not a penalty, just silence, plus a crawl budget you've spent teaching Google that most of your URLs aren't worth its time.
If you're going to run a programmatic section, the number to watch from week one is your indexed-to-published ratio, not your published count. Published count is a vanity metric with a progress bar attached, which is the same trap I fell into with the hundred-directory submission spreadsheet. Indexed count is the one that tells you whether Google agrees the pages should exist.
Do AI search engines cite programmatic pages?
More than they cite most other things, which is the genuinely new argument for building a directory in 2026 and the one I'd have found least believable a year ago.
Wix's AI Search Lab research on which content types get cited across ChatGPT, Google AI Mode, and Perplexity puts listicle content at 21.9% of all AI citations, the highest of any format, ahead of articles at 16.7% and product pages at 13.7%. Narrowed to ChatGPT specifically, "best X" listicle pages account for something like 43.8% of cited page types. Omniscient Digital's analysis of 23,000-plus AI citations found that for branded citations, 17% come from directories.
The mechanism is boring, which is usually a good sign it's real. When someone asks an assistant for the best tool for a job, the model has to retrieve passages that already contain a comparison set. Category pages and structured listings are shaped exactly like that. Pages with tables get cited noticeably more often than unstructured pages of similar length, across every engine, because a table is a set of extractable claims sitting next to each other.
So a well-built directory category page is close to the ideal retrieval object, and a 222-word summary of a single tool is not. That's the same conclusion the indexing section reached from a different direction, which is when I start to trust it.
Two caveats, because I've been wrong about this genre of advice before. Everything above is correlational: it describes what got cited, not what caused a citation, and the sites in those samples are overwhelmingly larger and older than mine. And AI referral traffic is still a fraction of a percent of the web, so "directories get cited" is a real finding attached to a small channel. It's a reason to structure your pages well. It is not a reason to publish a thousand of them.
The test I'd run before generating a single page
If you're deciding whether to build a programmatic section this week, this is the version I'd hand to myself two months ago.
- Name the unique input before you name the template. If you can't finish the sentence "each page contains ___ that exists nowhere else," you don't have a programmatic SEO project, you have a formatting exercise. Write the sentence first. The template is the easy part and it will seduce you into starting there.
- Write ten pages by hand before you generate any. If ten is boring to write, a thousand will be unreadable. This is the cheapest possible test and almost nobody runs it, because generating a thousand is more fun than writing ten.
- Ask whether an assistant can regenerate the page from its title. Paste your own page's H1 into ChatGPT and compare. If the output is as good as yours, the page is a liability whatever the policy says.
- Point editorial content into the section, and link back out of it. A programmatic block with no hand-written pages connecting it to the rest of the site looks structurally like a doorway. Mine didn't have this at first. It does now.
- Watch indexed pages, not published pages, from day one. Set the expectation early that a chunk of what you publish will never get indexed, and that this is information rather than a bug to escalate.
- Cap the section at the size you can actually maintain. 144 pages is a number one person can revisit. Ten thousand is not, and an unmaintained programmatic section decays into exactly the thing the policy describes without anyone deciding to make it that.
Nothing on that list is about avoiding a penalty. That's on purpose. Every item is about whether the page deserves to exist, and passing that bar handles the policy question as a side effect, which is roughly what Google has been saying in public for two years.
What 144 pages are really for
I built the tools directory because I kept hitting the same wall after every launch, which is that nobody knows your product exists and no amount of shipping fixes that. A directory is a distribution surface. That reasoning still holds.
But there's a second reason I built it that I'm less proud of, and it's the same one that put me behind a hundred-row directory submission spreadsheet and made me ship an llms.txt before checking whether it does anything. Publishing 144 pages feels like progress in a way that writing one genuinely useful page does not. There's a counter. It goes up. You can look at it.
The template made that feeling cheap to buy. Four headings, 222 words, and a category to file it under, repeated until the number is impressive. And 144 pages of that is not 144 times better than one good page. Depending on how much of it Google indexes, it might not be better at all.
I'm not deleting them. They're honest, they're mine, and the category pages do a real job. What I am doing is stopping at 144 instead of running the number up, and spending the next stretch making individual pages worth landing on rather than adding rows. That means longer "who it's for" sections, actual comparison tables on the category pages, and the thing I can say that a model can't, which is what broke when I used the tool.
The uncomfortable version of the lesson is that programmatic SEO is a distribution multiplier, and multiplying by a small number gets you a small number. Distribution is still the hard part, and 144 pages of template is what it looks like when you try to solve it with a for loop.
A solo full-stack developer and product builder with 8 years of experience shipping production software and 2 years as an indie hacker.
alexcloudstar.com