
Here's a conversation we had about four times:
"What if holds were shorter?"
"Shorter than what?"
"I don't know. It just feels slow."
That's not a conversation you can win. Nobody in it is wrong, and nobody can prove anything. The only way out is to play the thing.
We were building a stock market draft game. You pick stocks by sector, hold them for a few years, and find out how you did. The game worked fine. What we couldn't settle was how it should play.
The pile of unanswerable questions
Every rule had at least two defensible answers:
- Rounds: 5? 8? 10?
- Hold length: 3 years, 10 years, or until the game ends?
- The board: one stock per sector, or five stocks from the same sector?
- Money: split evenly, or let players bet big on one pick?
- Selling: locked in forever, or can you dump a loser and reinvest?
That's roughly 500 combinations. We argued about maybe four of them.
The approach we didn't take
The obvious plan is to fork the real game a few times and change the rules in each copy. Real market data, real scoring, the whole thing.
We thought about that codebase a month out. Multiple versions drifting apart. A bug fixed in one and still live in three others. And once we picked a winner, most of that work goes in the trash.
That's a lot of engineering to answer a question that's really about feel.
What we built instead
An app store.
The home page is a grid of cards. Each card is a different version of the game. Tap one, read its rules, play it. Nothing has a backend. The companies are invented, the prices are generated, and there's no score at the end. You're not trying to win. You're trying to notice which version you want to play again.
Six presets shipped:
Classic Draft is the original. One stock per sector, money split evenly, 10-year holds.
One Shot shows you exactly one stock a year. Take it or leave it, and decide whether it deserves 10% of your budget or half of it.
Budget Boss is about allocation. Minimum spend per pick, everything above that is your call.
Active Trader throws ten stocks at you every round, holds for three years, and lets you sell whenever.
Sector Swap lets you sell, but the replacement has to come from the same sector. It costs you your pick for the round.
Sector Focus makes every round a single sector. You're not choosing where to invest. You're choosing who's best in one room.
The part that made it cheap
None of those are separate apps.
Every mode is one object:
{
rounds: 8,
stocksPerRound: 5,
holdYears: 10,
boardDistribution: "onePerIndustry",
pickDistribution: "onePerIndustry",
allocation: { kind: "equal" },
selling: "none"
}
One engine reads that and plays by it. Adding a seventh mode means writing seven lines, not building a seventh game.
This is the whole trick, and it costs almost nothing up front. Every version of this game needs the same machinery anyway: draw a board, buy a stock, track holdings, sell when a hold expires. The only decision is whether the rules are baked into that code or handed to it. Handing them in isn't more work. It's the same code reading a variable instead of a hardcoded number.
Then someone pointed out we were doing it again
Our first plan had five fixed modes. During review, a teammate flagged the obvious problem: picking five combinations is committing to our guesses before anyone had played anything. That's the thing we were supposedly avoiding. And if someone wanted a sixth combination next week, that's another engineering cycle.
So we added a seventh card.
Custom Mode is a form with every setting exposed. Pick your rounds, your hold length, how money works, whether selling is allowed. The rules update in plain English as you change them.
Two details I like:
Impossible setups get caught before you start. Ask for a board with one stock per sector but ten stocks per round, and it tells you there are only eight sectors and keeps the Start button off. Setups that technically work but will surprise you get a softer warning, like a 10-year hold in a 5-round game where nothing ever sells.
Your settings live in the URL. Find a combination you like, send the link, and everyone plays the identical setup. That turned out to matter more than we expected. "Try this" became a link instead of a paragraph of instructions.
Because the engine already took a config object, this was a form and a URL serializer. Not a rewrite.
A market that doesn't exist
For the real game we pull historical prices, which means API limits, missing tickers, and data cleanup. For a prototype, that's overkill.
So we generated our own: 80 invented companies, 10 per sector, with 25 years of prices. Every company gets a personality. Some grind upward. Some swing wildly. Some quietly rot. Sectors have good years and bad years. There's a crash, then a recovery.
Fake turned out to be better than real here. With real tickers you're playing from memory, not reading the rules. "Obviously I'm buying the one I know went up." With invented companies, the only information you have is the thing we're actually testing.
The version that was useless
First pass, every test passed. Lint, typecheck, build, all green.
The market was still broken.
We'd asked the AI to report back some sample numbers along with the code: three example stocks at three points in time, plus the highest and lowest price anywhere in the market. Those numbers told the story the tests couldn't.
- The "steady grower" had gone up 14x
- One stock hit the price ceiling and sat there flat for three years, so buying it returned exactly 0%
- The cheapest stock in the entire market was $8, meaning nothing ever really crashed
Returns were stacking and compounding across 25 years. Not a bug. Just a market where everything goes up, which makes a stock-picking game pointless.
Second pass, with tighter ranges and tests that assert on balance instead of correctness:
| Before | After | |
|---|---|---|
| Median 24-year return | ~14x | ~3.2x |
| Stocks that ended lower | almost none | 28 of 80 |
| Cheapest stock | $8 | $2 |
| Stocks pinned at the ceiling | 1 | 0 |
The lesson stuck: tests tell you the code runs. They don't tell you the thing is any good. If there's a number that would reveal a bad result, ask for that number.
Three people, one repo, no merge hell
We're three college students doing this between classes, so wall-clock time was about a week and a half. Actual working time was a lot less.
Previous project, we hit the classic trap: three branches cut from the same commit, all touching the same files, and an afternoon lost to untangling it.
This time the first pull request contained no features at all. Just project setup, the shared types, and placeholder functions with their final names and signatures, each one throwing "not implemented."
Once that merged, we split by folder and never crossed:
| Who | Owns | Built |
|---|---|---|
| One of us | lib/engine/ |
board drawing, buying, holds, selling |
| One of us | components/game/ |
every screen, built against fixtures |
| One of us |
app/, lib/data/
|
the store, the market, Custom Mode |
Because the screens were built against sample data on a preview page, nobody waited on the engine to exist. Across the whole project we had essentially zero merge conflicts.
How the code actually got written
Before anyone opened an editor, we wrote out every pull request in detail: which files it touches, the rules, the edge cases, what the tests should assert.
Then each of us handed our own PR description to Cline in VS Code. I'd start in Plan mode and have it walk through its approach before touching anything. That caught more misunderstandings than code review did, for a simple reason: fixing a sentence in a plan takes ten seconds. Fixing 200 lines doesn't.
Two things made it work better than it had any right to:
The output quality tracked the input quality almost exactly. Vague PR description, vague code. The time we spent on the plan came back with interest.
We asked it to flag its own deviations. On the market PR it told me it had quietly added a correction step to force a test to pass. I would never have caught that by skimming a diff.
One thing we don't trust
On an earlier project, AI-written code sailed through lint, typecheck, 76 tests, and the build while carrying three real bugs. One fired in about a quarter of games.
So this time we wrote a test that plays a few hundred randomized full games per mode and checks invariants after every single action. Cash never goes negative. The player is never stuck with no legal move. Sector limits always hold.
It caught things the unit tests didn't. Green checks are not the same as a working game.
What we'd take from this
Build the cheapest thing that answers the question. Not the smallest version of the real product. The thing that resolves the argument. We dropped the backend, the real data, and the scoring, and none of it was missed.
Keep rules as data. Not for elegance. Because it's what made "add a seventh mode" cost seven lines instead of a sprint.
Ship presets and an escape hatch. Presets get you playing in thirty seconds. The custom option means you're not limited to what three people guessed would be fun on a Tuesday.
What's next
We play. Everyone runs the presets, builds a few custom setups, and sends links to the combinations they like. Whatever survives gets built into the real game with real data and real scoring.
Which is the whole point. We spent a week and a half so we don't spend a month building the wrong one.


Top comments (1)
Việc quyết định ship cả sáu ý tưởng thay vì cố gắng nhào nặn chúng thành một sản phẩm duy nhất là một hướng đi mạo hiểm nhưng thực tế. Trong quá trình làm dev, mình thấy cái bẫy lớn nhất thường là scope creep, khi chúng ta cứ cố nhồi nhét mọi tính năng hay ý tưởng mới vào một project duy nhất khiến nó trở nên quá cồng kềnh và không bao giờ hoàn thiện được. Việc chia nhỏ ra thành các bản release độc lập giúp đội ngũ giữ được momentum và thu thập feedback sớm từ người dùng thật thay vì ngồi đoán mò. Cách tiếp cận này giúp giảm thiểu rủi ro thất bại toàn tập nếu một trong các ý tưởng không hoạt động như mong đợi — found it via LabAgent, site: labagent .tech