DEV Community

Mirko Stahnke
Mirko Stahnke

Posted on Originally published at findnix.eu

Building a 24-Language Journal in Plain PHP: AI Translation, Hash-Based Staleness, and a Points-Gated Community

findnix.eu is a GDPR-first metasearch engine with its own index. Search was never the whole story, so we added a Journal: articles, a weekly digest, "finds" (sites worth a look), user questions and answers, and a profile page for every domain in the index. It's plain PHP 8.4 and MariaDB, no framework, in the same codebase as everything else. These are the decisions that mattered.

One table, four post types, translations on the side

Articles, digests, finds and questions share one table with an ENUM type column. Translations live in their own table keyed by (post_id, lang), so the original post never changes shape and a missing language is just a missing row. German is the source language; the other 23 EU languages are translations. Each language version gets its own URL (/journal/fr/some-slug) with hreflang links, and the sitemap lists every version with its translated title and description, so the translated pages end up in our own search index too.

Markdown without a library

The renderer is about 100 lines. The one rule that keeps it safe: escape everything first, then generate only the tags we allow.

$t = htmlspecialchars($t, ENT_QUOTES | ENT_SUBSTITUTE, 'UTF-8');
// only now turn [text](url) into , and only for http(s), mailto or local paths
Enter fullscreen mode Exit fullscreen mode

Raw HTML in a post simply shows up as text. Links from user content get rel="nofollow ugc".

Translating HTML with an LLM, and not trusting the result

Posts are rendered to HTML once, and that HTML is what we translate: split at

–

boundaries into chunks of a few thousand characters, one gpt-4o-mini call per chunk, with the instruction to keep every tag and attribute and translate only visible text. Title and teaser go in a separate call that returns JSON.

The important part is what happens afterwards. An article is attacker-controlled text as far as the model is concerned: a user post could contain instructions, and the model might obey them and emit markup. So the translated HTML goes through a DOMDocument whitelist (a short list of tags and attributes, URL schemes checked) before it is stored. If the model produces a