You wouldn't hire a Michelin chef to boil the water. Yet most writers route every single generation — the throwaway paragraph, the "what if she just left instead" experiment, the fourth rewrite of an opening they'll scrap anyway — through the priciest model they can reach. A month later the bill lands north of a streaming subscription, and they can't figure out why.
Nobody polishes a first draft. So why pay polish prices to make one?
The fix is dull and obvious the second you see it: use different AI models for drafting and editing. Draft on something cheap or free, then hand the final pass to the smart, expensive model where its judgment actually earns the rate. Same finished chapter, a fraction of the spend. It's the easiest way to cost less without writing worse, and setup runs about two minutes.
Why one model for everything quietly drains you
Look at what each stage of writing actually asks for.
Drafting is a volume game. You want words down — placeholder dialogue, a rough cut of the fight scene, three ways the chapter could open. Most of it gets slashed, rewritten, or folded into something better. The quality of any single generation barely registers, because you're going to take a knife to it regardless.
The final pass is a different animal. Now rhythm, subtext, restraint, and a character voice that holds steady across a long scene suddenly carry the whole thing. This is where a stronger model pulls ahead — it reaches for clichés less, sustains tone over a long passage, and knows when not to add a sentence. The models people trust here are the premium tier: Claude Opus and Sonnet, GPT-5.5, and Anthropic's writing-tuned Fable, which keeps topping creative-writing benchmarks for voice and character. They're genuinely better at prose. They also charge real money per token.
So the waste is baked in. You're paying top rates during the phase where quality is disposable, when you'd capture 90% of the value spending nothing there. The money should chase the moment where quality sticks around.
The two-model workflow, plainly
Cut writing into two jobs and give each its own model.
- The workhorse (drafting). Cheap or free. Your beat generator, your idea machine, your "expand this into 400 words so I can tell if it works" button. You'll fire it constantly and never feel a twinge when you do.
- The finisher (editing). Premium. You reach for it rarely — once a scene is roughly right and you want it to actually land. Its job is the rewrite that tightens, the polish, the last bit of smoothing.
The whole thing works because drafting is high-volume and editing is low-volume. You might generate a scene twenty times and polish it once. Put the cheap model on the twenty and the expensive one on the single pass, and your average cost per finished word falls through the floor.
There's a hard ratio underneath this. A premium model can run ten to a hundred times the per-token price of a solid budget one, depending on the pair you pick. If 80% of your token spend happens during drafting — and for most people it does — shifting that 80% onto a free or near-free model is close to an 80% cut. You're not clipping coupons. You're moving the decimal point.
Which models to actually pick
I'll keep this loose, because the free roster on OpenRouter churns month to month — models arrive and vanish as upstream providers shuffle their hosting. Check the live list before you commit to anything. What holds today:
For the workhorse, the free tier drafts perfectly well. As of mid-2026 that covers DeepSeek R1 and V3, Meta's Llama models, and Qwen — all sitting at $0 on OpenRouter, capped by a rate limit instead of a bill. They aren't the finest prose stylists breathing, and they don't need to be. They need to be fast, free, and good enough to hand you something to react to. They clear that bar without trying.
If free rate limits start blocking you, the cheapest paid budget models run a fraction of a cent per generation — effectively free across a working session. I've gone deeper on that trade-off in the cheapest AI models that still write good prose, because "cheap" and "unreadable" aren't the same word, and a few genuinely inexpensive models punch well above their price.
For the finisher, spend on quality. This is where a top-tier model repays itself — you run it a handful of times per chapter, so the per-token price barely dents the total, and the difference in output is right there on the page. Fewer clichés, better rhythm, a voice that doesn't drift over a long scene.
A rough starting point:
- Free workhorse: whichever strong free model is live on OpenRouter this week — a DeepSeek or Llama variant is usually a safe bet.
- Premium finisher: a top prose model — Claude Opus or Sonnet, or Fable for the creative-tuned one.
Don't agonize over the exact picks. The structure is what saves the money; the specific models swap out anytime.
Where the story bible earns its keep
Here's the worry people raise about mixing models: won't the cheap draft trip up the smart editor? Different model, no memory of what came before, and suddenly your protagonist's eyes change color in the polish pass.
That's a live risk if your only "context" is the raw text you paste in. It evaporates the moment the AI reads a persistent story bible — your characters, world rules, tone, the facts that can't move. When both the workhorse and the finisher pull from the same bible, they work off one source of truth even though they're completely different engines underneath. The cheap one drafts against your canon; the premium one polishes against that same canon. Continuity survives the handoff.
This is also why the two-model split doesn't turn into babysitting. You aren't re-explaining your world to each model by hand. You set the facts once, and every generation — cheap or premium — respects them.
Setting it up in KudoWrite
This is exactly what KudoWrite's model-chain settings exist for. You bring your own OpenRouter key — no subscription, no markup on top of what OpenRouter charges, and no content filters, so genre and theme are your call rather than a policy's. Then you tell the app which model handles what.
The workflow drops cleanly onto the two jobs:
- Draft rough beats — a line, a fragment, "they argue and she walks out" — on your cheap or free workhorse.
- Hit Expand to inflate a beat into real prose, Rewrite to try another angle, Split to break a bloated section apart. Run these constantly on the cheap model. That's where the volume lives.
- Once a section is roughly right, switch to your premium finisher for the final Rewrite or polish, then Commit the keepers into the chapter.
Underneath all of it, the story bible feeds context to whichever model is active, so jumping from the cheap draft to the expensive polish never drops your characters or your world. Set your model chain, open the editor and start drafting, and the billing sorts itself out.
Your work saves straight to your own Google Drive, local-first — the app doesn't push your writing through anyone's servers, and there's no one on the far end reading your manuscript. If you want a hand choosing that premium finisher, I keep a running comparison of the best OpenRouter models for creative writing.
The one habit worth building
If nothing else sticks: stop treating every generation as equal. It isn't. A throwaway draft beat and a final polish pass are two different acts, and pricing them the same is how the bill balloons while you're not looking.
Draft like it's cheap, because it should be. Polish like it matters, because it does. Wire the chain up once, forget it's there, and let the expensive model do the single job it's worth paying for.