Journey Docs
CMS

SEO Details

Per-page SEO metadata, the robots policy and its host and member-gate overrides, canonical hostname resolution, and header/footer HTML rules

Every CMS page has exactly one core.cms_page_seo row, created with the page and cascading on delete. All fields are nullable — an unset field means "no opinion", which is not the same as an empty string.

FieldTypePurpose
titleTagtext<title>
metaDescriptiontextmeta description
featuredImageUrltextsocial share image
headerHtmltextmarkup injected into <head>
footerHtmltextmarkup injected before </body>
allowIndexingbooleanrobots index policy
allowFollowingbooleanrobots follow policy

Writes go through PageSeoService.replace, which is a full replacement, not a patch. Any field omitted from the payload is written as NULL. A successful write invalidates the cached published site for that page.

Robots policy

allowIndexing and allowFollowing are deliberately nullable three-state values, and they apply on custom hosts only. The resolver overrides them:

allowIndexing:  onFallbackHost || locked ? false : page.seo.allowIndexing
allowFollowing: onFallbackHost || locked ? false : page.seo.allowFollowing

Both operands are computed during site resolution, not stored:

locked

const locked = page.memberAuthRequired === true;

True when the page sits behind the member gate — the member_auth_required column on core.cms_pages. A gated page is never indexable, because its public content is a sign-in wall rather than the content a crawler would be indexing. See page types for the gate itself.

onFallbackHost

const fallbackHostname = fallback ? `${fallback.hostLabel}.${suffix}` : hostname;
const onFallbackHost = hostname === fallbackHostname;

True when the request arrived on the owner's fallback host rather than a custom domain. Every organization, brand, or property gets one row in core.cms_fallback_hosts, created automatically the first time a page route is made for that owner. The label is derived from the owner's external id, so the hostname is always:

{owner external id}.sites.journey.com

The suffix comes from CMS_SITE_HOST_SUFFIX and defaults to sites.journey.com. The label is not a vanity name and is not author-chosen — dnsSafeHostLabel lowercases the external id, converts _ to -, strips anything outside [a-z0-9-], collapses repeated hyphens, trims leading and trailing hyphens, and truncates to the 63-character DNS label limit. An external id that normalizes to an empty string is rejected outright.

Fallback hosts are working URLs, not vanity ones. Forcing them to noindex keeps the platform-suffixed copy of a page out of search results so the partner's own domain is the only indexable version.

If the owner has no cms_fallback_hosts row, fallbackHostname falls back to the requested hostname, which makes onFallbackHost true for every request. That page is noindex and nofollow on every host, including an active custom domain. An owner missing its fallback-host row is the first thing to check when a partner reports that a live page will not index.

Resulting behavior

RequestEffective policy
on the owner's fallback hostnoindex, nofollow
page has member_auth_requirednoindex, nofollow
owner has no fallback-host rownoindex, nofollow — on every host
on an active custom domain, not gatedthe page's own allowIndexing / allowFollowing

This is enforced at resolution time in PublicCmsSitesService, not left to the renderer, so a fallback-host URL cannot be indexed even if the author set the page to indexable — and a member-gated page never leaks into search results regardless of its stored setting.

The override is one-directional. It can only remove indexing, never grant it. A page stored as allowIndexing: false stays noindex on its custom domain.

Canonical hostname

A page group can be reachable on two hosts — the owner's fallback host and, once active, its custom domain. Resolution returns both and derives the canonical:

canonicalHostname = customHostname ?? fallbackHostname

Because the fallback host is always noindex, the canonical URL points at the custom domain as soon as one is active, and search engines are never offered two indexable copies of the same page.

headerHtml and footerHtml exist for verification tags, structured data, and third-party snippets. They are restricted by an allowlist:

Allowed
Tagsmeta, link, style, script
Attributesname, content, property, rel, href, type, media, sizes

html, head, and body wrappers are ignored rather than rejected, so pasted markup that includes them still validates.

The two write paths treat violations differently:

  • The admin client strips disallowed markup silently via DOMPurify.
  • The assistant path rejects it and returns findings, so the model can retry against the allowlist rather than having its output quietly altered.

Who writes SEO

Humans edit these fields in PageSeoForm in Core Admin. The CMS Assistant writes the same cms_page_seo row through its replace_seo tool — there is no parallel SEO store and no separate assistant-owned representation. The assistant's cms-seo skill covers exactly the seven fields the admin form exposes.

Source: api/src/core/cms/pages/services/page-seo.service.ts, api/src/core/cms/pages/services/public-cms-sites.service.ts, api/src/core/cms/assistant/operations/seo-html.ts.

On this page