syted

syted blog

Semantic SEO Automation: Build the Cluster With Code

Search for "semantic SEO automation" today and the page sitting at the top of the organic results is about 350 words long, carries no date, and has zero backlinks pointing at it. We pulled the live results for that query from Ahrefs before writing this, and the entry barrier does not exist: position 8 belongs to a domain with a Domain Rating of 0. Of the eight organic results, five are pages published by SEO tool vendors (one of them a landing page selling the software directly), a sixth is an agency listicle of tools, and another is a LinkedIn collection written for digital marketers.

That tells you two things. The query has demand and almost no supply of anything executable. And the people answering it are selling software, which is why the answers stop at the word "automation" and never reach a command line.

This article is the other version. It covers what semantic SEO actually is (including the meaning of "semantic" that trips up every developer who reads the term), which parts of it a script can genuinely decide, which parts a script must only propose, and where Google's own documentation draws a hard line. Every claim about Google behaviour below is quoted from first-party documentation, with the quote in the text so you can check it.

The Word "Semantic" Means Two Different Things, and Only One of Them Is SEO

If you write code, "semantic" already means something to you: using HTML elements according to what they represent. Google wrote that definition down itself, in a Search Central post published on Monday, July 16, 2012, called "On web semantics":

In web development context, semantics refers to semantic markup, which means markup used according to its meaning and purpose.

The same post lists the advantages of doing that. There are exactly three, quoted verbatim: "It's the professional thing to do. It's more accessible. It's more maintainable."

Notice what is missing from that list. Ranking is not on it. Google has never claimed that swapping a div for a section moves a page in search results, and in the July 2024 SEO Office Hours it went further on the related question of heading order: "Having headings in semantic order is fantastic for screen readers, but from Google Search perspective, it doesn't matter if you're using them out of order."

So use article, section, nav and a sane h1 to h6 structure. Do it for accessibility and for the person who maintains the template after you. Do not put it on an SEO roadmap and expect a ranking change, and do not let a tool sell you an "HTML semantics audit" as search work.

Semantic SEO means something else entirely. It is the practice of organising a site around topics and the real-world things those topics refer to, rather than around individual keyword strings. One page owns one question, the pages that share a subject link to each other on purpose, and the markup states plainly what each page is about. That is the thing worth automating, and the Ahrefs results for this very query show that the parent topic groups tool-intent variants like "best seo semantic content analysis tool" under it, so the demand really is one topic and not a dozen.

The tactic most often sold under the semantic label is the one to be most careful with. Padding a page with synonyms and "related terms" to look thematically complete is not semantic SEO, it is the thing Google's spam policies name directly: "Keyword stuffing refers to the practice of filling a web page with keywords or numbers in an attempt to manipulate rankings in Google Search results."

The Unit of Work Is a Cluster, Not a Keyword

A semantic pipeline operates on a set of pages, not on one page at a time. That single change is what makes the automation tractable, because almost every useful decision in SEO is a comparison between pages rather than a property of one.

Which page should own "semantic seo automation" and which should own "what is semantic seo"? You cannot answer that by looking at either page alone. Which internal link is worth adding? Same problem. Is a page decaying, or did the whole topic lose demand? Same problem again.

Define the cluster as a small record you keep in the repository, next to the content. A JSON file is enough, and keeping it in git means every change to the topic map arrives in a diff you can review:

{
  "topic": "semantic seo automation",
  "owner": "content/blog/semantic-seo-automation.md",
  "questions": [
    "what is semantic seo",
    "does semantic html help seo",
    "what can be automated in semantic seo",
    "how do you measure a topic cluster"
  ],
  "entities": ["schema.org", "JSON-LD", "Search Console", "topic cluster"],
  "siblings": ["seo-automation", "seo-content-audit", "seo-workflow"]
}

Four fields, and every automated step below reads one of them. The questions array drives the outline and the FAQ. The entities array drives the gap check and the markup. The siblings array drives internal links. The owner field is what stops two of your own pages competing, which is a failure you create yourself rather than one the market inflicts on you.

Keeping the topic map in the repo also gives you something no tool dashboard does: history. When a cluster's traffic changes you can git log the file and see what you decided and when. That matters more than it sounds, because the usual reason a cluster underperforms is a decision nobody remembers making.

Harvest the Question Space Without a Single API Key

The first automatable step is finding the questions a topic actually contains. You do not need a paid tool for the first pass. Google's autocomplete endpoint answers in JSON over plain HTTP:

curl -s "https://suggestqueries.google.com/complete/search?client=firefox&hl=en&gl=us&q=semantic+seo+automation"

Run that today and you get four suggestions back: semantic seo automation, semantic seo meaning, what is semantic seo, what is seo automation. Every one of them is a question about doing the thing. Not one contains "agency", "services" or "company".

That second observation is the cheapest audience check in existence, and it costs nothing. If the suggestions for your topic come back full of "near me", "company" and "pricing", the people typing that query want to hire someone, and an explainer will collect readers who never convert. We run the same test on our own topics before writing, and it is the reason several high-volume keywords never became articles here.

Two caveats worth stating plainly. This endpoint is not a documented, supported Google API, so treat it as a hint rather than infrastructure: wrap it in a timeout, cache the result, and do not let a pipeline step die when it changes shape. And autocomplete tells you about phrasing, not about volume. For volume and difficulty you still need a keyword tool, and the thresholds you apply to the numbers are your own.

Feed the suggestion crawl from the seed plus each of its first-level suggestions, deduplicate, and you have a question space in a few seconds. Group the result by meaning rather than by string, because that grouping is the actual semantic work. "What is semantic SEO" and "semantic SEO meaning" are one page. "Semantic SEO" and "semantic SEO tools" are two, because one wants an explanation and the other wants a shortlist, and merging them produces a page that half-answers both.

Find the Entity Gap, Not the Keyword Gap

A keyword gap tells you a phrase you have not used. An entity gap tells you a thing you have not mentioned. The second is the one that changes an article, because the missing item is usually a standard, a tool, a spec, or a named mechanism that every competent page on the topic has to name.

Here is the whole idea as a script. Pull the text of the pages currently ranking, count the proper nouns and technical terms they share, and subtract the ones your draft already contains:

#!/usr/bin/env bash
# entity-gap.sh <your-file.md> <competitor-text-dir>
set -euo pipefail
mine=$(tr '[:upper:]' '[:lower:]' < "$1")
cat "$2"/*.txt \
  | grep -oE '\b([A-Z][a-zA-Z0-9.+-]{2,}(\.[a-z]{2,})?|JSON-LD|llms\.txt)\b' \
  | sort | uniq -c | sort -rn | head -40 \
  | while read -r count term; do
      if ! grep -qiF -- "$term" <<< "$mine"; then
        printf '%6d  MISSING  %s\n' "$count" "$term"
      fi
    done

Crude, and useful on the first run. Pointed at the pages ranking for this query, the terms that came back most often were schema.org, JSON-LD, Search Console, and the names of specific tools. The first three belong in any honest treatment of the subject. The tool names mostly do not, which is exactly the judgement a script cannot make for you.

That is the shape of every step in this pipeline. The machine produces a ranked list of candidates. You decide which candidates are real. An entity that appears on nine of ten competing pages is a strong signal that it belongs in yours, and it is still not proof: nine pages can repeat the same received idea, and our own SEO content audit process exists partly to catch inherited mistakes like that.

If you want Google's own view of which entities exist, there is an API, and there is a warning attached to it. The Knowledge Graph Search API documentation now carries this notice: "To support our customers with additional enterprise requirements and high QPS use cases, we are migrating this API to Cloud Enterprise Knowledge Graph." Build on it if you like, but do not put it on the critical path of a daily job.

Generate the Markup From the Content, Not From a Plugin Default

Structured data is the one part of semantic SEO where you are stating the meaning of a page in machine-readable form, and it is the part that automates most cleanly, because the input is a file you control.

Start from what Google actually says it does. The structured data introduction defines it as "a standardized format for providing information about a page and classifying the page content", recommends one syntax ("Google recommends using JSON-LD for structured data if your site's setup allows it"), and sets one rule that kills most generated markup: "Don't add structured data about information that is not visible to the user, even if the information is accurate."

That last line is why generating markup from the content beats generating it from a template. A plugin that stamps a fixed block on every page will eventually assert something the page does not say. A generator that reads the file can only describe what is in the file.

The Article reference is unusually permissive about which fields you supply: "There are no required properties; instead, add the properties that apply to your content." It accepts three types, Article, NewsArticle and BlogPosting, and wants dates in ISO 8601 format with timezone information. It also states the limit of the whole exercise in one sentence: "Google does not guarantee that features that consume structured data will show up in search results."

The semantic additions worth generating are the ones that say what a page is about and how it sits in a set. All of these resolve on schema.org today, which we checked by HTTP status before recommending them: about, mentions, sameAs, isPartOf, hasPart, DefinedTerm and BreadcrumbList.

Property What it states Safe to generate from
about The primary subject of the page The cluster topic field
mentions Secondary things the page names Entities confirmed present in the body
sameAs A canonical identifier for an entity A curated map you maintain by hand
isPartOf The collection this page belongs to The cluster record
BreadcrumbList The path a reader took to get here The route, which the framework knows

In the App Router, this ships as a plain script tag. The current Next.js JSON-LD guide (documented against 16.3.8, while this repository runs 16.2.12) is explicit about not reaching for the script component: "The next/script component is optimized for loading and executing JavaScript. Since JSON-LD is structured data, not executable code, a native <script> tag is the right choice here."

export default async function Page({ params }: { params: Promise<{ slug: string }> }) {
  const { slug } = await params
  const post = await getPost(slug)

  const jsonLd = {
    '@context': 'https://schema.org',
    '@type': 'BlogPosting',
    headline: post.title,
    datePublished: post.publishedAt,
    dateModified: post.updatedAt ?? post.publishedAt,
    about: { '@type': 'Thing', name: post.topic },
    mentions: post.entities.map((name: string) => ({ '@type': 'Thing', name })),
    isPartOf: { '@type': 'Blog', '@id': 'https://example.com/blog' },
  }

  return (
    <article>
      <script
        type="application/ld+json"
        dangerouslySetInnerHTML={{
          __html: JSON.stringify(jsonLd).replace(/</g, '\\u003c'),
        }}
      />
      {/* ... */}
    </article>
  )
}

The .replace(/</g, '\\u003c') is not decoration. The Next.js guide warns that JSON.stringify "does not sanitize malicious strings used in XSS injection" and recommends scrubbing the < character exactly this way. If your titles or descriptions can ever contain user input, that one call is the difference between structured data and a script injection point. We have more on the surrounding metadata APIs in Next.js SEO in the App Router.

One thing not to generate: a file or a markup block invented to please an AI assistant. Google's documentation on its AI features is blunt about it. "There are no additional technical requirements", and "You don't need to create new machine readable files, AI text files, or markup to appear in these features. There's also no special schema.org structured data that you need to add." Our longer treatment of that claim is in LLM SEO and AI answer engine optimization.

Internal linking is the step where naive automation does the most damage, and also the step where a semantic model earns its keep.

A string-matching linker finds the phrase "topic cluster" and links it to your topic cluster page. It also links the sentence where you explain that topic clusters are overrated, and the sentence inside a code comment, and the anchor text in a table of contents. After a hundred articles, nobody can tell which links were a decision.

A semantic linker compares the meaning of a paragraph with the subject of each candidate page and ranks the matches. The output should be a list of suggestions with a location and a score, written to a file, and nothing more:

content/blog/semantic-seo-automation.md:121  ->  /blog/seo-automation          0.81
content/blog/semantic-seo-automation.md:188  ->  /blog/seo-content-audit       0.74
content/blog/semantic-seo-automation.md:246  ->  /blog/subdomain-vs-subfolder  0.58

Then apply two rules by hand. First, link only where the sentence already says the thing, so the anchor text is prose you wrote rather than prose bent around a link. Second, cap the additions per page per pass, because a diff of twenty link insertions is a diff nobody reviews.

We apply exactly that constraint to this blog: a run may add one link line to an existing article and may not rewrite a sentence to make room for it. That rule exists because the alternative is a blog that slowly becomes link bait for itself. Where a cluster lives in your URL structure is a related decision with its own tradeoffs, covered in subdomain vs subfolder.

Measure the Cluster as a Set, With the API

A page-by-page traffic report hides the only thing you want to know: is the topic winning, and which page inside it is doing the work. The Search Console API gives you that in one query, grouped by page, and you can diff the result week over week.

The documented limits are worth knowing before you build on it. The rowLimit parameter has a "Valid range is 1 to 25,000; Default is 1,000", the dimensions you can group by are country, device, page, query, searchAppearance, date and hour, and the reference states a caveat most dashboards never surface: "The API is bounded by internal limitations of Search Console and does not guarantee to return all data rows but rather top ones."

Traffic from AI surfaces lands in the same report rather than a separate one. Google states that sites appearing in AI features are "included in the overall search traffic in Search Console", reported in the Performance report under the "Web" search type. So there is no second pipeline to build for AI visibility in Search Console, though there is a separate measurement problem in whether assistants cite you at all, which we cover in LLM brand visibility.

Here is the honest version of what cluster measurement looks like early on. Our own Search Console data for the 28 days to today shows seo-for-interior-designers at 83 impressions and 0 clicks, seo-keywords-for-photographers at 58 impressions and 0 clicks, and the homepage at 38 impressions for 3 clicks. Impressions without clicks is a title and description problem, not a content problem, and it is the cheapest fix available: the page already ranks somewhere a human saw it.

Two thresholds we actually use, offered as starting points rather than laws. A page with more than 30 impressions and zero clicks over 28 days goes on the rewrite list for its title and description. A page whose impressions fall by more than half across two consecutive 28-day windows goes on the decay list, which is a content question and not a metadata one. How long SEO takes explains why those windows are 28 days and not 7.

Where Automation Has to Stop

Google's spam policies define the far end of this road in one sentence: "Scaled content abuse is when many pages are generated for the primary purpose of manipulating search rankings and not helping users." The examples named under it include "using generative AI or similar tools to generate many pages without user value" and automated transformations like synonymizing and stitching content together.

Read that against the semantic pipeline above and the line is clear. Automating the discovery of what a topic contains is analysis. Automating the production of a page per discovered phrase is the thing the policy names.

The quality self-assessment Google publishes is the practical test, and three of its questions are the ones a generated page always fails. "Does the content provide original information, reporting, research, or analysis?" "Is this content written or reviewed by an expert or enthusiast demonstrably knowing the topic well?" "Does the content have any easily-verified factual errors?"

So split the pipeline by whether a step has an opinion. Here is the division we run:

Step Automate the decision Automate the proposal only
Question harvest Yes
Cluster assignment Yes
Entity gap list Yes
Which entities to add Yes
Schema generation from file Yes
sameAs identifiers Yes
Internal link candidates Yes
Internal link insertion Yes
Impression and decay reports Yes
What to rewrite, merge or delete Yes

Every row in the right-hand column is a judgement about meaning or about a reader. The left-hand column is counting and formatting. The division is not about capability, it is about who is accountable when the output is wrong. Our separate article on SEO automation goes through the same split for the non-semantic parts of the job, and SEO workflow puts it on a weekly schedule.

One honest note on outcomes. Nothing here is a ranking lever you can pull, results vary by site and by topic, and no one can guarantee a position in Google for any of it. What a semantic pipeline does reliably is remove the two failures you cause yourself: two of your own pages fighting over one query, and a page that never says the thing the topic is about.

A Build Order That Pays Off Early

Build these in order, and stop whenever the next one stops being worth it for your site size.

  1. The cluster file. One JSON record per topic, in the repo. An afternoon, and every later step reads it.
  2. The autocomplete harvest. One curl plus a dedupe. An hour, no credentials, and it doubles as an audience check.
  3. The Search Console pull. Grouped by page, stored per week. This is what tells you whether anything else worked.
  4. Schema generated from the file. Mechanical, testable with the Rich Results Test, and it removes a class of plugin-induced lies.
  5. The entity gap report. Cheap and crude first, better later. Useful from the first run.
  6. Semantic link suggestions. Last, because it needs embeddings and because its output is advisory anyway.

Steps 1 to 3 are the ones that change behaviour, and none of them needs a vendor. If your blog is not yet showing up in search at all, none of this is the first problem to solve, and diagnosing why a site is not showing up on Google comes first. If you have the opposite problem and need demand rather than structure, start from how to get traffic to your website. Founders building this inside a product rather than beside one will find the staffing version of the question in SEO for SaaS.

FAQ

What is semantic SEO?

Semantic SEO is organising a site around topics and the real things those topics refer to, instead of around individual keyword strings. In practice it means one page owns one question, pages on the same subject link to each other deliberately, and the markup states what each page is about using properties like about and mentions. It is not the same as semantic HTML, and it is not padding a page with synonyms.

Does semantic HTML help SEO?

Not as a ranking factor, on Google's own account. Google's 2012 post on web semantics lists exactly three advantages of semantic markup: "It's the professional thing to do. It's more accessible. It's more maintainable." On the related question of heading order, Google said in July 2024 that from a Search perspective "it doesn't matter if you're using them out of order". Write semantic HTML for accessibility and maintenance, and do not expect it to move rankings.

Is SEO dead now with AI?

The measurement moved, the mechanism did not. Google states that to appear as a supporting link in AI Overviews or AI Mode "a page must be indexed and eligible to be shown in Google Search with a snippet", and that "There are no additional technical requirements." Traffic from those surfaces is reported inside the normal Search Console Performance report under the "Web" search type, so the same indexing and content work feeds both.

Is SEO still worth it in 2026?

That is an arithmetic question, not a philosophical one: compare what a search visit is worth to you against what producing the page costs, over a window long enough to measure. The thing that has changed is zero-click risk. The results page for this very query carries an AI Overview with seven sitelinks above the first organic result, which means some share of the demand never reaches any website. We work the numbers through in is SEO worth it.

Can ChatGPT do SEO?

It can do the parts that are counting and drafting, and it cannot do the parts that are judgement. A model will happily produce a cluster map, an entity list and a schema block. What it will not reliably do is tell you that an entity nine competitors mention is a received idea, or that a keyword's results are full of people shopping for a provider rather than doing the work. Those are the two calls that decide whether a page earns anything.

Do I need semantic SEO tools to do this?

No, and the results page for "semantic seo automation" is the evidence: five of its eight organic results are pages published by tool vendors, which is a supply signal, not a requirement. The three steps that change behaviour (a cluster file in git, an autocomplete harvest over HTTP, and a Search Console pull) need no vendor. Buy a tool when the measurement says a specific step is the bottleneck, not before.

Get cited by ChatGPT. Rank on Google.

You found this article through search. That is the whole product.

  • One researched article a day
  • Published on your own domain
  • Keywords checked against live results
Start writing

Get cited by ChatGPT. Rank on Google.

You found this article through search. That is the whole product.

  • One researched article a day
  • Published on your own domain
  • Keywords checked against live results
Start writing