In this post I talk about how dialogues for my Japanese speaking trainer are put together, and why "just ask a model to write a conversation" was a no go.
Some context.
For a few years now I have been learning Japanese and even passed JLPT N2 this summer. At some point speaking became the weakest part. Reading and listening can be done alone with any amount of material. Speaking needs a partner, or at least a script worth repeating out loud.
Shadowing apps exist, and so do AI chat partners. With the chat option I kept running into the same thing: you talk with a model about who knows what, and you have no idea how natural its Japanese is. Plus, as a learner, you are the last person who could tell.
What I wanted was boring in a good way: short model dialogues, like the ones in a decent textbook, for everyday situations, at N5 to N3 level. One role is yours. You read it out loud, the browser recognizes your speech, and you only move on once the line is said correctly.
So I built daihon.app (台本, "script"). The trainer itself is a separate story. This one is about the content.
The problem with dialogues.
There are no native speakers on the project, and no native proofreading is coming. In a speaking trainer an unnatural line isn't just a typo: you drill it until it's automatic, so you memorize the mistake together with the pronunciation.
So "sounds natural to me" can't be the quality bar. Naturalness has to rest on something you can check.
And it breaks on two separate levels:
- the line: do people actually say this?
- the flow: do people say this at this point of a conversation?
Borrowing good sentences only fixes the first one. A dialogue stitched together from perfect sentences is a Frankenstein: every part is real, the whole thing is dead. Which is why most of the work goes into the skeleton, not into the lines.
Start from a real conversation.
A dialogue starts as a sequence of moves taken from an actual conversation. Roughly like this, for a restaurant call:
staff: greeting + asks which area
customer: area and number of people, in one go
staff: two options with prices
customer: picks one + asks about the time
staff: confirms + asks about smoking
customer: short answer, doesn't repeat the question
This is exactly the kind of thing you don't get by inventing. On move four a textbook author writes a long polite question. A real customer says 六時でも大丈夫ですか and stops.
The sources are public research corpora of everyday Japanese and open-licensed sentence collections. All of them are credited. Some of them don't allow taking text at all, and that's fine: the structure of a conversation (who speaks when, what kind of move, how long) is a fact, not someone's text. So the closed corpora work where they matter most, on the skeleton, and actual wording only comes from sources whose license allows it.
One thing learned the hard way: a drill is not a conversation. Textbook exercises train one move, so they all have the same shape: question, answer, clarification, thanks. My first batch used them as skeletons and came out as ten dialogues you could not tell apart. Textbooks are still useful as a list of situations and vocabulary, just not as a source of flow.
What a line has to prove.
There is no such thing as an unchecked line. Each one carries a note of where it came from, and there are three kinds:
- Taken as is from a source that allows it.
- A slot swap in a real sentence pattern. 何名様ですか becomes 何時からですか: the construction is real, only the slot changes. This is how a line gets adjusted to the level without breaking grammar.
- Written by me, but the phrase has to show up in a corpus search. This catches calques and "textbook Japanese": if the phrase doesn't occur anywhere, people probably don't say it.
The third kind is a necessity. At N5 about half the lines need simplifying, and the simplified version still has to be checkable.
Zero hits doesn't always mean "wrong", though. Sometimes the construction is fine and only a specific word pair is rare. Sometimes you searched for the wrong spelling: one of the corpora writes numbers with digits, so 三十分 gives zero and 30分 gives over a hundred. And sometimes zero means exactly what it looks like. One line had every word backed by a source, and still the whole frame (a medicine that 咳が止まる) wasn't something people say, while the neighbouring frame with に効く was common. Checking words one by one can't see that. Checking the whole frame sees it immediately.
Rules a text has to pass.
A few rules came out of this that I now check every dialogue against:
- Short reactions are real lines. そうですか, なるほど, ええ make up a big share of real conversation, so they stay. (With one technical catch: the speech synthesizer returns silence for a lonely ええ。, so a reaction always comes with something after it.)
- Ellipsis where Japanese uses ellipsis. The answer doesn't repeat the question, and 私は only appears for contrast. These two give away textbook Japanese from the very first line. The echo is surprisingly easy to miss by eye, because you only see it in the pair, not in the line. Once you strip the final か from the question and compare the predicates, it shows up right away.
- Not every partner line is a question. Strictly alternating question/answer is the shape of an exercise.
- A turn is a move, not an answer. In real talk people usually say two things per turn: they answer and add a reason, a caveat, a counter question.
The rule I didn't see coming: one reading only.
This one is specific to a trainer that checks your speech.
Your spoken line is compared against the expected reading. If a word has two accepted readings, you can say the second one, be completely right, and still fail the line.
Numbers are the main suspects. 四十分 is fine as both よんじっぷん and よんじゅっぷん. 七時 is both しちじ and ななじ. 船便 is both ふなびん and せんびん. Destination 〜行き is both ゆき and いき, and that one sneaks in with any place name.
Fix: move the time to a neighbour that has one reading (六時半 instead of 七時). The scene doesn't change at all, and the line becomes passable for everyone who says it correctly.
Writing in several passes.
Drafts are written with an LLM, but not in one go. Each dialogue goes through several separate passes: collect the material and the skeleton, write the text, review, translate. Each pass starts fresh.
The fresh start for the review is the whole point. A session that just wrote a line also remembers why it wrote it, and it won't argue with itself. A clean pass reads the dialogue the way a reader does, and checks the flow, the tone and whether the scene is believable. The material from the first pass stays as the reference, so the review can still answer "was this intended?".
Translations come last and only touch translations. The Japanese text is final by then.
Audio.
The partner's lines are voiced with an open source offline TTS, generated once and stored, not requested on the fly. Two speakers, so the roles are easy to tell apart by ear. Getting the loudness even was its own small adventure: the standard one-step normalization didn't reach the target on lines under three seconds, and neighbouring lines differed by about 6 LU. Measuring each file and applying the exact gain fixed it.
Where it is now.
30 dialogues so far, 10 each for N5, N4 and N3. The pipeline is slow on purpose: a dialogue that takes a few sessions to make is still cheaper than a learner memorizing a sentence nobody says.
It's still a portfolio project, and I'm sure some lines can be better. If you're a native speaker or a teacher and something sounds off, I'd really like to hear it.
Try it here: daihon.app. No sign up needed for the first three dialogues.
Thank you for reading. Feedback is highly appreciated!
Top comments (0)