DEV Community

Cover image for One crop does not fit all: framing vertical clips by video type, on the user's PC
Mert ALP
Mert ALP

Posted on AI-assisted

One crop does not fit all: framing vertical clips by video type, on the user's PC

We build LunarClip, a Windows app that turns long videos into short vertical clips. The obvious way to make a 9:16 clip from a 16:9 video is a center crop, or a crop that follows the biggest face. Both work for a talking head and fail almost everywhere else. This post covers the two decisions that shaped the app most: the crop depends on what the video is about, and the heavy work happens on the user's own computer.

Why one crop rule is not enough

A 9:16 window over a 16:9 frame keeps only about a third of the width. What belongs in that third depends on the genre:

  • In a stream, the important things are in two places at once: the streamer's webcam in a corner and the gameplay in the middle. Any single crop loses one of them.
  • In sports, the subject is the ball and the players around it, not the largest face. A face-following crop happily tracks a spectator.
  • In a podcast, two people talk back and forth. A crop that jumps to whoever spoke last feels nervous, and a crop that sits on one person misses half the conversation.
  • In an everyday or food video, the important object is often not a person at all: the pan, the plate, the product.

So the app has a framing rule per video type.

Four video types, four different crops

Type Rule
Game The streamer's webcam goes on top showing only the person, and the gameplay fills the rest, wherever the webcam sits in the source.
Sports Follows the ball and the players around it, switching between close-up, medium, wide and full-frame shots.
Podcast Split screen for back-and-forth, one speaker for monologues, and a wide shot when everyone talks. Switching is deliberately calm.
Normal No automatic split. Keeps food and products in frame, and still zooms in on a person talking to camera.
Auto Picks the type from the title, tags, description and sound, and falls back to Normal when it is unsure.

Two things we learned while building these rules:

  1. Genre detection should fail safe. When Auto cannot tell what a video is, it falls back to the plainest rule. A wrong split screen on a cooking video looks far worse than a slightly loose crop.
  2. Calm beats clever. On podcasts, calm speaker switching reads better than a crop that chases every short "yeah". Fewer cuts, each one for a reason.

Why it runs on your PC

The second decision was where the work happens. Uploading a two-hour video to a server before anything can start means a long upload, a queue and a copy of the video on someone else's machine. So everything that touches the video runs locally.

What runs on the PC and what goes to the server

On the user's computer:

  • speech-to-text with a local speech model (downloaded once, with the first transcription),
  • shot and face analysis and the framing rules above,
  • captions and hook text,
  • rendering, on a recent NVIDIA, AMD or Intel graphics card if there is one, otherwise on the processor.

Only one step uses the server: picking the strongest moments. For that, the app sends short transcript excerpts of candidate clips (text, timestamps and scores), and the server returns its picks with titles and captions. The video, the audio and the finished clips never leave the computer.

The trade-offs are real and worth stating:

  • the app needs an internet connection while a project is being made, because moment picking happens on the server;
  • processing speed depends on the user's hardware instead of a data-center GPU;
  • it is a Windows app for now.

What the user gets in exchange is no upload, no shared queue, and no copy of their video anywhere else.

Takeaway

If you are building anything that reframes video, decide early what your subject is for each kind of footage. "Follow the biggest face" is a rule for talking heads, not a general one. And if most of the pipeline can run on the user's machine, the privacy story becomes a property of the architecture rather than a promise in a policy.
LunarClip

Top comments (1)

Collapse
 
ahmetozel profile image
Ahmet Özel •

Calm podcast switching needs a shot boundary as well as a speaker-duration threshold. Smoothing a crop across a real camera cut can move the window through irrelevant pixels, even though the speaker identity stayed the same.

I would reset the crop trajectory at detected cuts, then apply hysteresis within each continuous shot. A short fixture with an interjection, an occluded speaker and an abrupt camera change would show whether the rule avoids unnecessary switches without preserving a stale target. For sports, the equivalent fallback should also be explicit when the ball disappears, rather than continuing a confident-looking chase.