Unit tests can't see a preview that blinks or an export that drifts from what you edited. Here's the end-to-end Appium journey that catches both in Captions Bro.
Captions Bro is an iOS video editor for short vertical videos: a multi-track timeline, on-device transcription, animated karaoke-style captions, text and media overlays, and an export built for TikTok, Reels and Shorts.
Video editors break in ways unit tests can't see. A drag lands a frame off. The preview blinks black for a fraction of a second while a composition rebuilds. The exported file doesn't quite match what you saw while editing. Every one of those can happen with a fully green unit suite.
So on top of the unit tests, Captions Bro has one end-to-end test that uses the app the way a person does. It opens the app on a fresh simulator, picks clips in the Photos picker, trims, splits, adds a transition, generates captions, styles them, adds a text overlay, plays the project, exports it and opens the share screen. Then it checks what came out.
Here's how it works.
One command, a fresh device every time
Everything starts with one script: appium/run.sh.
- It leases a simulator from a pool of two, so at most two runs share the machine, and erases it before every run.
- It pins everything that changes pixels. That means the status bar (always 09:41), the locale, and every preference that changes the UI.
- It builds the app in its own derived-data folder, and skips the build if no input changed.
- It fills the Photos library with test clips rendered by ffmpeg. They're the same bytes every run, and each file is dated so the picker grid always shows them in the same order.
- It starts its own Appium server and hands everything to a pytest journey.
Then Appium drives the real UI: real taps, real drags, real sheets, even Apple's out-of-process Photos picker. The smoke profile is 15 steps and takes about seven minutes. The full journey is 25 steps.
No "tap and hope"
The classic way UI tests go wrong is tapping a button and assuming the thing happened. Or worse, sleeping for two seconds and hoping it did.
Captions Bro avoids that by letting the app describe itself. Under test, the editor mounts a hidden accessibility element called uitest-state. Its value is a compact string with the playhead, whether it's playing, which sheet is open, how many clips, captions and overlays the project has, and one short signature per clip: duration, speed, and a letter for each capability applied to it (T for transition, R for reversed, K for keyframes, and so on).
So the "add a transition" step looks like this:
@step("mix transition A→B")
def _mix_transition(j, ui):
ui.deselect()
s = ui.state()
ui.scrub_to(nominal(s["cl"][0]["dur"]))
seam = ui.rect(ui.find("timeline.clip.0"))
ui.tap_xy(seam.right, seam.cy)
ui.tap("Mix")
ui.wait_state("transition applied",
lambda s: "T" in s["cl"][0]["flags"])
j.checkpoint("transition-panel", mask_names=TRANSITION_TILES)
It taps "Mix", then waits until the first clip actually carries the T flag. Every step works this way: it proves the edit landed, not just that a button was tapped. There isn't a single fixed sleep in the journey. Every wait polls the app's state or an element, with a timeout that fails loudly.
Screens are compared, not eyeballed
At key moments the journey takes a screenshot and compares it with an approved baseline using SSIM. The threshold is 0.985. Getting a screenshot to be the same twice in a row took more work than I expected:
- Settle first. A shot is taken only after two identical captures in a row.
- Exact frame. Every seek aims at the middle of a 30 fps output frame, never its edge. Then the test waits until the preview reports that it's really showing that frame. The model's playhead settles the instant the finger lifts, but the still for it decodes a moment later.
- A little slack where drags are involved. Block widths on the timeline can differ by a point between runs, so the timeline band is compared at its best ±3 px alignment, with its own threshold of 0.95.
- Mask what never stops moving. Transition previews and caption style tiles loop forever, so their area is masked out of the comparison.
The export must look exactly like the preview
This is the rule I care about most. Whatever you see in the editor is what should end up in the video you post.
During the journey, the test pauses the editor at 1, 5.5, 8 and 11.5 seconds and saves what the canvas shows. After the export, it pulls the frames at the same instants out of the .mp4 and compares each pair. If captions, text, or a transition render differently in the export than in the preview, the run fails, and the report shows the two frames side by side.
Then the file itself gets checked
The export is pulled out of the app's container and checked with ffprobe:
- H.264, SDR (bt709), 8-bit 4:2:0
- 1080×1920 at 30 fps
- duration within one frame of the timeline
- an audio track is present
- no flat black frame anywhere in the video
Why so strict about SDR? An HDR or HEVC export plays perfectly on your phone. Then TikTok re-encodes it on the server and the back half of the video freezes. That's the kind of bug you only hear about from users, so the test makes sure it can't ship.
The whole run is filmed, then scanned frame by frame
The simulator screen is recorded for the entire journey. Afterwards, the recording is scanned for four things that should never happen:
- Blink: the canvas drops to flat black for up to six frames and comes back.
- Black canvas: the canvas stays black for more than a second, outside steps that are allowed to be empty.
- Flicker: a single frame that differs from both of its neighbours, while the neighbours match each other.
- Freeze: the picture stops changing for more than 0.6 s during a step that's supposed to be playing. Freeze-frame clips in the project are exempt.
These are thresholds, not baselines. The recording's frame timing varies from run to run, so it's never compared against a reference video. Each detector asks a question whose answer should always be no. On top of that, any alert nobody asked for (like "Playback unavailable") fails the run.
The hard part: making it repeatable
Writing the steps was the easy part. Making every run produce the same pixels was the real work:
- Slow drags. Above roughly 200 pt/s, a scroller keeps momentum after the finger lifts, and the deceleration differs every time. Every drag is capped in speed and holds still before letting go.
- Mid-frame seeks. A seek that lands near a frame edge ends up on either side of it from run to run, so seeks always aim at the middle of a frame.
- Trims are undone. Trims snap to the continuous playhead, not to frames, so a tiny residual would shift every later frame of the project. A trim step drags the real handle, checks it, then undoes it. The lasting cut is made with Split, which lands on an exact frame.
- Generated, not committed, media. The test clips are ffmpeg test sources at 30, 24 and 60 fps, one of them landscape, encoded with a one-second keyframe interval like a phone camera.
- Retry the system, not the app. Apple's Photos picker sometimes sits on "Loading…" over a freshly seeded library. The test cancels and reopens it (up to three tries), and records each retry as a known limitation rather than an app bug.
What a run leaves behind
Every run leaves a folder with report.html, the full screen recording, the exported video, every checkpoint screenshot (with a diff image when one drifts), the side-by-side parity frames, and a verdict.json. The exit code is the verdict.
I also don't run it by hand most of the time. Claude Code runs the whole thing through a project skill, reads the report and looks at the failure frames itself, and tells me whether a failure is an app regression, a harness problem, or an intended UI change that needs new baselines.
Unit tests prove each capability on its own. This one test proves the app works when a person uses it.
If you make short videos, try the result at captionsbro.app. It's on the App Store, with guides and help docs at captionsbro.app/guides.
Originally published on Medium.








Top comments (0)