Google gives a page title roughly 60 characters before it truncates, and a meta description roughly 160. Notifio has 39 pages that need to fit inside both.
The complication is that the number written in the file is not the number a crawler counts. The root layout declares a title template:
title: {
default: "Notifio – Rental Listing Alert Monitor",
template: "%s | Notifio",
},
So every page except the home page gets | Notifio appended, which is exactly 10 characters. A page author writing a 58 character title is writing a 68 character title. You can see it on any page: notifio.app/alerts/wg-gesucht has title: "WG-Gesucht Alerts – Be in the First Five Messages" in its source data, and in the HTML.
That gap is the whole reason there is a test. The comment at the top of the file records what happened without one:
/**
* Search engines truncate a title past roughly 60 characters and a description
* past roughly 160, so every page has to fit inside those budgets. The budget
* that matters is the *rendered* one: the root layout applies a `%s | Notifio`
* template, which silently spends ten of the sixty on every page but the home
* page. Hand-counting that is what let three titles drift to exactly 60, so it
* is counted here instead.
*/
const TITLE_MAX = 60;
const DESCRIPTION_MAX = 160;
The suffix is read out of the layout, not retyped
The obvious implementation is const SUFFIX = " | Notifio". That test passes forever, including on the day somebody changes the template to %s | Notifio – Rental Alerts and every title on the site grows by 17 characters.
So the suffix is extracted from the layout file:
function titleTemplateSuffix(): string {
const layout = readFileSync(LAYOUT, "utf8");
const match = layout.match(/template:\s*"(.*?)"/);
assert.ok(match, "root layout no longer declares a title template");
const [, template] = match;
assert.ok(template.includes("%s"), `unexpected title template: ${template}`);
return template.replace("%s", "");
}
Reading source with a regular expression is not elegant, and the two assertions are the reason it is acceptable. If the template is removed, the test fails with "root layout no longer declares a title template" rather than quietly measuring titles against a budget that no longer applies. If the template exists but has lost its %s, that fails too. The failure mode a scraping test has to design against is not a wrong answer, it is a confident green tick on a file it no longer understands.
Two halves, because the site has two kinds of page
28 of the pages are rows in a TypeScript array. 15 alert-site pages like /alerts/kamernet, 5 comparison pages like /compare/rentbird, 3 audience pages like /for/students, and 5 guides like /guides/how-fast-do-rental-listings-go. Those are checked by iterating the arrays:
const collections = [
["/alerts", ALERT_SITES],
["/compare", COMPARE_TARGETS],
["/for", AUDIENCE_PAGES],
["/guides", GUIDE_PAGES],
] as [string, readonly { slug: string; title: string; description: string }[]][];
for (const [prefix, entries] of collections) {
it(`${prefix}/* fits the title and description budgets`, () => {
assert.ok(entries.length > 0, `${prefix} has no entries to check`);
for (const entry of entries) {
check(`${prefix}/${entry.slug}`, entry.title, entry.description);
}
});
}
Note entries.length > 0. A loop over an empty array is a passing test. If an import silently resolves to nothing, the check it was supposed to run disappears and the suite still goes green.
The fourth one actually did. The alerts barrel re-exports with extensionless specifiers that Node's resolver will not follow, so the site files are loaded by globbing the directory instead:
const SITES_DIR = path.join(import.meta.dirname, "../../src/lib/alerts/sites");
const ALERT_SITES = (
await Promise.all(
readdirSync(SITES_DIR)
.filter((name) => name.endsWith(".ts"))
.map((name) => import(path.join(SITES_DIR, name))),
)
).flatMap((module) => Object.values(module).flat());
That started as a workaround and turned out to be the better behaviour. There are three market files today, uk.ts, netherlands.ts and europe.ts, and when a fourth market is added its pages are covered the day the file lands rather than the day somebody remembers to add an import.
The other half reads page.tsx as text
The remaining pages set export const metadata in their own file, so there is nothing to import. The test walks src/app, skips api, collects every page.tsx, and pulls the fields out with a regex:
function literalField(block: string, key: string): string | null {
const match = block.match(new RegExp(`\\n ${key}:\\s*([\\s\\S]*?),\\n \\w`));
if (!match) return null;
const parts = match[1].match(/"(?:[^"\\]|\\.)*"/g);
if (!parts) return null;
return parts.map((p) => JSON.parse(p) as string).join("");
}
It handles plain string literals and multi-line strings split across lines, which is all any page uses. JSON.parse on each quoted chunk rather than a manual unescape means \" and \\ behave. Anything it cannot read returns null, and null is reported as a failure rather than treated as a pass:
const title = literalField(block, "title");
assert.ok(title, `could not read a string title out of ${file}`);
A description may legitimately be absent, because a page is allowed to inherit the root layout's. So check() takes string | null and skips the description assertion when it is null, while the title is mandatory.
Dynamic routes bail out early:
if (source.includes("generateMetadata")) return;
Their metadata comes from the collections the first half already checked, so measuring the four [param] route files would either duplicate that work or, worse, fail to parse a template literal and look like a real problem.
And the suite asserts that it found pages at all:
it("finds the static pages to check", () => {
assert.ok(pages.length >= 10, `only found ${pages.length} pages under src/app`);
});
A recursive directory walk pointed at a path that moved returns an empty array, generates zero it() blocks, and reports success. That is the single most likely way a test like this stops working, and it costs one assertion to rule out.
What the numbers actually look like
21 tests, and this is what they are protecting:
ℹ tests 21
ℹ pass 21
ℹ fail 0
ℹ duration_ms 101.5
The longest rendered title on the site is 59 characters. The longest description is 158. The median title has six characters of headroom once the suffix is counted.
59 WG-Gesucht Alerts – Be in the First Five Messages | Notifio
58 Spotahome Alerts – Catch Verified Listings First | Notifio
58 Rentbird Alternative – £20 Once, No Subscription | Notifio
58 Rental Scams: The Signals That a Listing Is Fake | Notifio
57 Kamernet Alerts – Get New Rooms Within a Minute | Notifio
So this is not a formality with 20 characters of slack. Every one of those titles is one adjective away from being cut off in a search result, and the person adding the next page has no way to feel that while typing. node --test takes a tenth of a second to feel it for them.
The general version
Three things transfer to any site that generates pages from data:
- Test the rendered string, not the authored one. Title templates, breadcrumb prefixes and locale suffixes all mean the number in your source is not the number being measured. Derive the transformation from the code that applies it.
-
Assert that the test found something. Every collection check has a
length > 0, and the file walk has a>= 10. An empty loop is the quietest possible failure. - Make unparseable mean fail. A regex over source will eventually meet a construct it does not handle. Returning null and asserting on it turns that into a two-minute fix; returning a default turns it into an unnoticed gap.
This suite runs under node --test with no test framework and no build step, which has its own set of costs. I wrote those up separately in node --test runs our TypeScript, and the price is that @/ does not exist, and the extensionless-import workaround above is the same root cause showing up a second time.
If you want to see the output, the pages are at notifio.app/alerts, notifio.app/compare and notifio.app/guides. View source on any of them and count.
Top comments (0)