DEV Community

Daniel Pertu
Daniel Pertu

Posted on

Our site works out its own URL once, and a preview deploy getting that wrong costs us the index

A Next.js app needs its own absolute URL in more places than you would guess before you count them. On pub-trivia.app the list is: metadataBase, every canonical tag, the 72 entries in the sitemap, the sitemap pointer inside robots.txt, every @id and url in the structured data, the success and cancel URLs handed to Stripe Checkout, the link in a password reset email, the link in a team invite email, and the QR codes printed on the table cards that players scan.

Ten consumers. Each of them could call process.env.NEXT_PUBLIC_WEBSITE_URL directly. That is how you end up serving a sitemap for one origin and emailing links for another, with nothing failing anywhere.

So it is derived once, in a file whose entire job is to answer one question.

const PRODUCTION_URL = 'https://pub-trivia.app'
const DEVELOPMENT_URL = 'http://localhost:3000'

function resolveSiteUrl(): string {
  const configured = process.env.NEXT_PUBLIC_WEBSITE_URL?.trim()
  if (configured) return configured.replace(/\/+$/, '')
  return process.env.NODE_ENV === 'production' ? PRODUCTION_URL : DEVELOPMENT_URL
}

export const SITE_URL = resolveSiteUrl()
Enter fullscreen mode Exit fullscreen mode

Three decisions are packed into those few lines, and all three are about failure rather than about the happy path.

The trailing slash strip is not cosmetic

Everything downstream joins a root relative path onto this value, so a configured URL ending in a slash produces https://pub-trivia.app//pricing. That is a different URL. It serves the same page, but the canonical tag on it would not match the URL in the sitemap, the sitemap entry would redirect, and a crawler comparing the two would be looking at what it reasonably reads as two URLs for one page.

The strip is one replace and it means nobody downstream has to think about whether the configured value was written with a slash. The alternative is ten callers each remembering to normalise, which in practice means two of them remembering.

The fallback follows the environment instead of being one string

This is the part I would get wrong if I were writing it quickly. The obvious fallback is a single default, and either choice of default is a bug.

Default to production, and a local dev server with no environment variable claims to be the live site. Your local sitemap lists production URLs. Your local Stripe test checkout redirects you to the live domain. Worse, your local password reset email contains a link to production, so you click it and reset your password on the real site while testing.

Default to localhost, and a production build that is missing the variable tells Stripe that its post payment redirect is http://localhost:3000, tells Google that the site lives on localhost, and sends every invite recipient a link that only works on the machine that generated it.

Neither default is safe, because the question is not what the URL probably is. It is what this environment must never be allowed to claim. So the fallback is a function of NODE_ENV: a dev server may not claim production, and a production build may not claim localhost. The failure mode of each environment is bounded to that environment.

The derived boolean is where it earns its keep

One line further down, a second constant:

export const IS_PRODUCTION_SITE = SITE_URL === PRODUCTION_URL
Enter fullscreen mode Exit fullscreen mode

An identity comparison, not a check of NODE_ENV, and not a check for the string "vercel" in a hostname. A preview deploy is a production build in every sense that matters to Next.js: NODE_ENV is production, the pages are prerendered, the metadata is real. The only thing that distinguishes it is that it is being served from a URL that is not ours.

That distinction gates four things.

robots.txt becomes a refusal. On a preview deploy the whole generated file is one rule, disallow everything. Not because preview content is secret, it is identical to production, but because a preview deploy serves a complete copy of every page. Let that get indexed and you have two URLs competing for every query you care about, and the duplicate is the one with no links pointing at it and a sitemap full of URLs it does not serve.

The robots meta tag flips. The root layout sets index and follow from the same boolean. robots.txt stops crawling; the meta tag stops indexing, and they are not the same thing. A page can be excluded from crawling and still be indexed from a link somebody else published.

The structured data is not emitted at all. The Organization script in the root layout is wrapped in the boolean. This one took the longest to settle on, because the first instinct is to emit it everywhere and keep the markup consistent. The problem is what the node says: it carries an @id built from the production URL. Serving that from a preview host is a page asserting an identity that belongs to a different origin. Emitting nothing is honest; emitting a production identity from a staging URL is the kind of thing that is fine until the day it is a manual action.

The IndexNow submitter refuses to run. The script that pings search engines when content changes checks the same boolean first and exits with an error that names what it found, because a submission of localhost URLs is both useless and a small embarrassment delivered directly to Bing.

Check the result on the live site

The agreement is visible from outside, which is the only reason I trust it.

curl -s https://pub-trivia.app/sitemap.xml \
  | grep -o '[^<]*' | sed 's|||' \
  | awk -F/ '{print $1"//"$3}' | sort | uniq -c
Enter fullscreen mode Exit fullscreen mode

That prints one line: 72 URLs, one origin. Not a mixture, and no localhost left in by a build that ran without the variable set.

curl -s https://pub-trivia.app/pricing | grep -o ']*>'
Enter fullscreen mode Exit fullscreen mode

Same origin again, and it names /pricing rather than the homepage. Then:

curl -s https://pub-trivia.app/ | grep -o 'application/ld+json'
Enter fullscreen mode Exit fullscreen mode

That one is present on the live site. On a preview deploy of the same commit it is absent, and that absence is the feature.

If you would rather see the consumers than the config, the pricing page is where the Stripe redirect URLs get built from this constant, and the about page is where the structured data that carries the derived @id is emitted. Signing up on the free tier, which takes no card, gets you the other two: the links in the account emails are absolute URLs produced by the same function.

What I would take to the next project

Derive the origin once, in a file that exports it and a boolean saying whether this deployment is the real one. Make the fallback depend on the environment, because the question each environment answers is not "what is my URL" but "what am I not allowed to claim". And when a deployment is not the real site, prefer emitting nothing over emitting the real site's identity, because a missing node is a gap and a wrong node is a contradiction.

Top comments (0)