You're editing a caption. It reads: "POV: an average day in a startup." And in your head, one word carries the whole joke. Not the whole line. Just startup. It needs to be bigger, heavier, maybe in a punchier typeface, because that's the word the entire caption is actually about. Everything else is just setup.
You open the tool. You select startup and bump up the size. Every word in the caption gets bigger. You try changing the font on just that word, and the entire line switches with it. The caption behaves like a single, locked block, and your one good idea gets flattened into "make the whole sentence bigger," or nothing at all.
Which kills the joke. "POV: an average day in a startup" only works if startup is the word that lands, because that's where the punchline lives. Blow the whole line up equally and the emphasis disappears, it just becomes a bigger, louder sentence with no punchline in it.
If you've ever felt that specific, small, deflating moment, you've already lived the exact problem Varnam started as.
The gap nobody names, because everyone just works around it
Open Canva, CapCut, VN, or almost any auto-caption tool, and this limitation is sitting there the moment you go looking for it. Size and typeface are treated as properties of the entire caption, change either one on a single word, and the whole line changes with it.
That's a bigger problem than it sounds like, because size and typeface are exactly the two properties that create hierarchy. Size is what makes one word read as more important than the words around it. Typeface is what gives a word a different character, sharper, softer, heavier, than its neighbors. Those are the two levers emphasis actually runs on, and they're the two you can't pull on a single word.
Nobody talks about this constraint directly, because everyone's quietly found a workaround: retyping the caption across separate text boxes, splitting one line into overlapping layers, or just giving up on the emphasis entirely. Multiply that small workaround by every video, every caption, every time one word mattered more than the rest, and it adds up to a genuinely large amount of creative intent getting quietly erased.
What we actually built, in two weeks
The first version of Varnam wasn't an AI recommendation engine. It didn't read emotion, tone, or pace. It fixed exactly one thing: every word became its own entity.
Upload a video, and it transcribes automatically. Instead of handing back one editable block of caption text, each word comes back as an independent object, its own size, its own typeface, its own position on the frame. Nothing stayed locked together anymore.
What that actually unlocked
The moment a single word could be edited on its own, people didn't just get a new feature. They got their intent back. The word that was supposed to hit harder finally could, bigger, in a heavier typeface, nudged off the baseline, however it needed to look to match what was already in their head.
For the first time, someone didn't have to describe what they wanted and settle for the closest the tool could offer. They could just make it, word by word, exactly the way they'd imagined it.
That reaction, watching someone realize the tool finally does the thing they'd been quietly working around for years, is the entire reason this was worth building first.
Why this was the right place to start
We didn't start with the hardest problem, reading a video's tone and recommending typography automatically. We started with the smallest, most concrete thing standing between someone and their own idea: a tool that refused to edit at the level people actually think at.
People don't think in captions. They think in words. Any tool that doesn't respect that unit is going to feel limiting, no matter how many templates get bolted on top of it.
That's still the foundation everything else at Varnam is built on.
We're still building toward launch, this is just the part of the story we started from. If this is a frustration you've quietly worked around yourself, the waitlist is open, and I'd love for you to be one of the first to see where it's headed.