DeepSeek V4 Flash Chatbot Conversation Cost
Running a chatbot on DeepSeek V4 Flash? Every turn re-sends the whole transcript, so a long chat costs far more than turns × one message. This computes the real cost and what caching claws back.
For teams running chatbots — computes how a conversation's cost grows quadratically as the transcript is re-sent each turn, and how much prompt caching claws back, not a 'turns × one message' estimate.
Model prices from OpenRouter · updated 2026-10-05
01 Your conversation
$0.03/M in · $1.28/M out · cache reads $0.03/M
02 Naive estimate vs real cost
A 40-turn chat costs 1.6× the naive “turns × one message” estimate — the transcript replay adds up.
Real conversation cost is 1.6× the naive estimate
Without caching this would be $3,413 (1.6×) — caching the transcript is your biggest lever on long chats.
Related cost calculators
Prices from OpenRouter, snapshot 2026-10-05, synced daily. Turn k re-sends system + prior turns + new user message; total input = T×(system+user) + (user+assistant)×T(T−1)/2. With caching, the re-read prefix bills at the cache-read rate. The first-turn cache-write premium and any summarization/truncation aren't modeled. All math runs in your browser.
How the math works
It's tempting to price a 40-turn conversation as 40 × one message. But every turn re-sends the whole transcript so far — turn 40 ships the system prompt plus all 39 previous turns before the model reads a single new word. So the input tokens grow with the SQUARE of the conversation length, not linearly.
Same baseline, one identity: the naive "one message × turns" estimate is $0.02 per conversation. Adding the re-sent history (429,000 tokens of replayed prior turns) brings the real, uncached cost to $0.03 — 1.6× the naive number. That gap is pure transcript replay, and it widens every turn.
Prompt caching is the fix, because the transcript is a stable growing prefix — exactly what caching is built for. With caching on, the replayed history re-reads at $0.03/M instead of $0.03/M, cutting the real cost to $0.03/conversation (1.6× naive) — $3,413/month at 100,000 conversations. Turn caching off and the same chats cost $3,413/month.
The levers all target the history term: a sliding window (keep only the last N turns), summarizing old turns into a short recap, and caching the transcript. Because the cost is quadratic, trimming the oldest turns of a long chat saves far more than trimming the same tokens from a short one — the tail of a long conversation is where the money is.
Not modeled: the one-time cache-write premium on the first turn (small), and any summarization/truncation you apply — this assumes the full transcript is re-sent. Inference prices sync daily from OpenRouter (updated 2026-10-05); this is a token-accounting comparison on the live catalog, not a separate price source. All math runs client-side with tested code.
Frequently asked questions
Why does a long conversation with a DeepSeek V4 Flash chatbot cost more than turns × one message?
Because every turn re-sends the entire transcript. A 40-turn chat replays 429,000 tokens of prior turns on top of the new messages, so the uncached cost is 1.6× the naive estimate ($0.03 vs $0.02). The longer the chat, the wider the gap.
How much does prompt caching save on a chatbot?
A lot for long chats — the transcript is a stable prefix, so caching re-reads it at $0.03/M instead of $0.03/M. Here it cuts the bill from $3,413/month (uncached) to $3,413/month. Caching is close to mandatory once conversations run long.
Does the cost really grow quadratically?
Yes — turn k re-sends roughly k turns of history, so summing over a conversation gives a term proportional to turns². Doubling the conversation length nearly quadruples the history-replay tokens. That's why very long sessions get expensive fast, and why capping history matters.
What's the cheapest way to cut chatbot cost?
Cap the history: a sliding window that keeps only the last N turns turns the quadratic back into a linear cost. Summarizing old turns into a short recap does the same while preserving context. And enable prompt caching so whatever history you do keep re-reads cheaply. Trimming the oldest turns of long chats moves the number most.
Are these prices current?
Inference prices sync daily from OpenRouter (updated 2026-10-05). This mold adds no separate price source — it's a token-accounting model of transcript replay on top of the live catalog, so it stays accurate as prices change automatically.
Should a DeepSeek V4 Flash chatbot worry about this?
Less so at 40 turns (1.6× naive), but the multiplier climbs fast with conversation length — worth watching if your sessions run long.