DEV Community

Cover image for Rebuilding Encarta showed me exactly where AI-written code breaks
Jean-Luc Martel
Jean-Luc Martel

Posted on Originally published at singular-state.com AI-assisted

Rebuilding Encarta showed me exactly where AI-written code breaks

The bug that froze the window was type-correct. It compiled, cargo check passed, the tests passed, and the operating system reported the application as not responding.

It was a Save As dialog called through the Tauri plugin's blocking API from inside a command, which blocks the event loop the WebKit window is running on. Nothing in the type system objects to that. Nothing in a headless harness notices. The only way to find it is to be a person, at a real desktop, clicking the button.

That defect is the whole finding of this project, and the project took four phases to earn the right to state it.

The setup

Wikicarta rebuilds Microsoft Encarta on Wikipedia — the visual browser, the atlas, the timeline, the research organiser, MindMaze — as a Tauri v2 desktop app in Rust and React, with an instant toggle between Live mode against the Wikimedia API and Offline mode against a local ZIM archive.

It is the third of four experiments in what current AI can do with legacy systems, and it occupies the worst position of the four. HAL/S had a formal spec and an independent interpreter to check against. Nautilus had a novel and built its own physics oracle. Wikicarta has neither. The source of truth is what the product felt like, and the only oracle is human judgement. There is no test that tells you whether the category wheel feels like Encarta.

An AI agent wrote essentially all of it. What follows is where that went well, where it went badly, and the fact that those two places were not the ones I expected.

The modern substrate is never the shape its docs claim

The offline story was supposed to be a pure-Rust ZIM reader. The zim crate looked right: random access, deferred loading, Send and Sync. It panicked with an integer overflow in parse_article_list before finishing the open call, against a real modern archive.

That pattern recurred with no relationship between the cases. TextExtracts HTML turned out unusable as a reader format, so article HTML comes from action=parse instead. Live image URLs came back protocol-relative — //upload..., not the https://... everything downstream assumed.

Three unrelated features, one rule: for a revival, the modern content, format or API is always slightly different from its specification and from the model's memory of it. Probe a real sample before building on it. The AI will confidently build against the documented shape, because the documented shape is what it was trained on.

The decision that paid for everything

Before it was provably needed, the reader was decoupled from its content source behind a single trait:

trait ArticleSource {
    fn get_article(title: &str) -> Article;
    fn search(query: &str) -> Vec<SearchResult>;
    fn get_image(key: &str) -> Vec<u8>;
}
Enter fullscreen mode Exit fullscreen mode

Adding a whole second content source later — live Wikipedia alongside offline ZIM — was two commands and a mode flag. No reader rewrite. Swapping the storage layer from JSON files to SQLite was invisible above the command boundary.

For software that assumed one sealed data source — a CD, a bundled database, a mainframe — inserting that seam first is what converts every later modernisation from a rewrite into an adapter. It is also the single decision an AI agent is least likely to make unprompted, because at the moment you make it there is exactly one source and the abstraction looks like overhead.

Two features in the final phase were pre-paid by decisions like it. The attribution aggregator became a GROUP BY rather than a re-fetch, because provenance had been stamped at save time back when the notes feature was built. Bundling the kiwix-serve binary into the package was a two-line change, because binary discovery had long been a prioritised candidate list rather than a hard-coded path.

Packaging work disproportionately cashes in — or punishes — architectural decisions made when the feature that needed them didn't exist yet.

Testability turned into architecture, which turned out to be good

The only layer testable without a desktop window is pure functions. So the code got steadily shaped so that the hard parts are pure functions: URL rewriting, upstream reconstruction, cache naming, database round-tripping and ordering, bookmark deduplication, legacy JSON import. The Rust suite went from nothing to 23 offline tests, and each feature's desktop-only remainder shrank to a short explicit list.

"What can I verify headlessly?" is a useful question about how to structure code, not just about how to test it. Under an AI agent it becomes close to essential, because the agent's verification loop is only as good as the surface it can reach.

Where the defects actually were

The pure logic — parsing, backoff, the media-cache URL rewrite, database ordering, live-title normalisation — compiled correct and passed on the first or second attempt. Consistently.

Every real bug was at the runtime seam:

  • The blocking Save As that froze the WebKit window.
  • DOM nodes inside an