# SEO Details
> Source: /cms/seo
> Per-page SEO metadata, the robots policy and its host and member-gate overrides, canonical hostname resolution, and header/footer HTML rules

# SEO details

Every CMS page has exactly one `core.cms_page_seo` row, created with the page
and cascading on delete. All fields are nullable — an unset field means "no
opinion", which is not the same as an empty string.

| Field | Type | Purpose |
| --- | --- | --- |
| `titleTag` | text | `<title>` |
| `metaDescription` | text | meta description |
| `featuredImageUrl` | text | social share image |
| `headerHtml` | text | markup injected into `<head>` |
| `footerHtml` | text | markup injected before `</body>` |
| `allowIndexing` | boolean | robots index policy |
| `allowFollowing` | boolean | robots follow policy |

Writes go through `PageSeoService.replace`, which is a **full replacement**, not
a patch. Any field omitted from the payload is written as `NULL`. A successful
write invalidates the cached published site for that page.

## Robots policy [#robots-policy]

`allowIndexing` and `allowFollowing` are deliberately nullable three-state
values, and they apply on custom hosts only. The resolver overrides them:

```ts
allowIndexing:  onFallbackHost || locked ? false : page.seo.allowIndexing
allowFollowing: onFallbackHost || locked ? false : page.seo.allowFollowing
```

Both operands are computed during site resolution, not stored:

### `locked` [#locked]

```ts
const locked = page.memberAuthRequired === true;
```

True when the page sits behind the member gate — the `member_auth_required`
column on `core.cms_pages`. A gated page is never indexable, because its public
content is a sign-in wall rather than the content a crawler would be indexing.
See [page types](/cms/page-types) for the gate itself.

### `onFallbackHost` [#onfallbackhost]

```ts
const fallbackHostname = fallback ? `${fallback.hostLabel}.${suffix}` : hostname;
const onFallbackHost = hostname === fallbackHostname;
```

True when the request arrived on the owner's **fallback host** rather than a
custom domain. Every organization, brand, or property gets one row in
`core.cms_fallback_hosts`, created automatically the first time a page route is
made for that owner. The label is derived from the owner's external id, so the
hostname is always:

```
{owner external id}.sites.journey.com
```

The suffix comes from `CMS_SITE_HOST_SUFFIX` and defaults to `sites.journey.com`.
The label is not a vanity name and is not author-chosen — `dnsSafeHostLabel`
lowercases the external id, converts `_` to `-`, strips anything outside
`[a-z0-9-]`, collapses repeated hyphens, trims leading and trailing hyphens, and
truncates to the 63-character DNS label limit. An external id that normalizes to
an empty string is rejected outright.

Fallback hosts are working URLs, not vanity ones. Forcing them to `noindex`
keeps the platform-suffixed copy of a page out of search results so the partner's
own domain is the only indexable version.

<Warning>
  If the owner has **no** `cms_fallback_hosts` row, `fallbackHostname` falls back
  to the requested `hostname`, which makes `onFallbackHost` true for every
  request. That page is `noindex` and `nofollow` on every host, including an
  active custom domain. An owner missing its fallback-host row is the first thing
  to check when a partner reports that a live page will not index.
</Warning>

### Resulting behavior [#resulting-behavior]

| Request | Effective policy |
| --- | --- |
| on the owner's fallback host | `noindex`, `nofollow` |
| page has `member_auth_required` | `noindex`, `nofollow` |
| owner has no fallback-host row | `noindex`, `nofollow` — on every host |
| on an active custom domain, not gated | the page's own `allowIndexing` / `allowFollowing` |

This is enforced at resolution time in `PublicCmsSitesService`, not left to the
renderer, so a fallback-host URL cannot be indexed even if the author set the
page to indexable — and a member-gated page never leaks into search results
regardless of its stored setting.

<Note>
  The override is one-directional. It can only remove indexing, never grant it.
  A page stored as `allowIndexing: false` stays `noindex` on its custom domain.
</Note>

## Canonical hostname [#canonical-hostname]

A page group can be reachable on two hosts — the owner's fallback host and, once
active, its custom domain. Resolution returns both and derives the canonical:

```
canonicalHostname = customHostname ?? fallbackHostname
```

Because the fallback host is always `noindex`, the canonical URL points at the
custom domain as soon as one is active, and search engines are never offered two
indexable copies of the same page.

## Header and footer HTML [#header-and-footer-html]

`headerHtml` and `footerHtml` exist for verification tags, structured data, and
third-party snippets. They are restricted by an allowlist:

| | Allowed |
| --- | --- |
| Tags | `meta`, `link`, `style`, `script` |
| Attributes | `name`, `content`, `property`, `rel`, `href`, `type`, `media`, `sizes` |

`html`, `head`, and `body` wrappers are ignored rather than rejected, so pasted
markup that includes them still validates.

The two write paths treat violations differently:

- The **admin client** strips disallowed markup silently via DOMPurify.
- The **assistant** path rejects it and returns findings, so the model can retry
  against the allowlist rather than having its output quietly altered.

## Who writes SEO [#who-writes-seo]

Humans edit these fields in `PageSeoForm` in Core Admin. The
[CMS Assistant](/cms/assistant) writes the same `cms_page_seo` row through its
`replace_seo` tool — there is no parallel SEO store and no separate
assistant-owned representation. The assistant's `cms-seo` skill covers exactly
the seven fields the admin form exposes.

Source: `api/src/core/cms/pages/services/page-seo.service.ts`,
`api/src/core/cms/pages/services/public-cms-sites.service.ts`,
`api/src/core/cms/assistant/operations/seo-html.ts`.
