CalcSays
DEEPSEEK V4 FLASH · CHAT HISTORY

DeepSeek V4 Flash Chatbot Conversation Cost

Running a chatbot on DeepSeek V4 Flash? Every turn re-sends the whole transcript, so a long chat costs far more than turns × one message. This computes the real cost and what caching claws back.

For teams running chatbots — computes how a conversation's cost grows quadratically as the transcript is re-sent each turn, and how much prompt caching claws back, not a 'turns × one message' estimate.

Model prices from OpenRouter · updated 2026-10-05

01 Your conversation

Model

$0.03/M in · $1.28/M out · cache reads $0.03/M

Cache the transcripthistory re-reads at $0.03/M

02 Naive estimate vs real cost

A 40-turn chat costs 1.6× the naive “turns × one message” estimate — the transcript replay adds up.

Naive (turns × one msg)
$2,126
no history counted
Real / month
$3,413
429,000 history tok/convo

Real conversation cost is 1.6× the naive estimate

Without caching this would be $3,413 (1.6×) — caching the transcript is your biggest lever on long chats.

Why do long chats get expensive? Every turn re-sends the whole transcript, so a turn deep in the conversation ships everything before it. The input tokens grow with the square of the length — a 40-turn chat is far more than 40× a single message. Cap history with a sliding window or summary, and cache the transcript.
📋 Full cost audit for this exact setup
Your current inputs, the cost decomposition, every savings lever ranked with its dollar impact, and the alternatives — computed instantly by the same tested engines behind this page. No email, nothing uploaded.

Prices from OpenRouter, snapshot 2026-10-05, synced daily. Turn k re-sends system + prior turns + new user message; total input = T×(system+user) + (user+assistant)×T(T−1)/2. With caching, the re-read prefix bills at the cache-read rate. The first-turn cache-write premium and any summarization/truncation aren't modeled. All math runs in your browser.

How the math works

It's tempting to price a 40-turn conversation as 40 × one message. But every turn re-sends the whole transcript so far — turn 40 ships the system prompt plus all 39 previous turns before the model reads a single new word. So the input tokens grow with the SQUARE of the conversation length, not linearly.

Same baseline, one identity: the naive "one message × turns" estimate is $0.02 per conversation. Adding the re-sent history (429,000 tokens of replayed prior turns) brings the real, uncached cost to $0.03 — 1.6× the naive number. That gap is pure transcript replay, and it widens every turn.

Prompt caching is the fix, because the transcript is a stable growing prefix — exactly what caching is built for. With caching on, the replayed history re-reads at $0.03/M instead of $0.03/M, cutting the real cost to $0.03/conversation (1.6× naive) — $3,413/month at 100,000 conversations. Turn caching off and the same chats cost $3,413/month.

The levers all target the history term: a sliding window (keep only the last N turns), summarizing old turns into a short recap, and caching the transcript. Because the cost is quadratic, trimming the oldest turns of a long chat saves far more than trimming the same tokens from a short one — the tail of a long conversation is where the money is.

Not modeled: the one-time cache-write premium on the first turn (small), and any summarization/truncation you apply — this assumes the full transcript is re-sent. Inference prices sync daily from OpenRouter (updated 2026-10-05); this is a token-accounting comparison on the live catalog, not a separate price source. All math runs client-side with tested code.

Frequently asked questions

Why does a long conversation with a DeepSeek V4 Flash chatbot cost more than turns × one message?

Because every turn re-sends the entire transcript. A 40-turn chat replays 429,000 tokens of prior turns on top of the new messages, so the uncached cost is 1.6× the naive estimate ($0.03 vs $0.02). The longer the chat, the wider the gap.

How much does prompt caching save on a chatbot?

A lot for long chats — the transcript is a stable prefix, so caching re-reads it at $0.03/M instead of $0.03/M. Here it cuts the bill from $3,413/month (uncached) to $3,413/month. Caching is close to mandatory once conversations run long.

Does the cost really grow quadratically?

Yes — turn k re-sends roughly k turns of history, so summing over a conversation gives a term proportional to turns². Doubling the conversation length nearly quadruples the history-replay tokens. That's why very long sessions get expensive fast, and why capping history matters.

What's the cheapest way to cut chatbot cost?

Cap the history: a sliding window that keeps only the last N turns turns the quadratic back into a linear cost. Summarizing old turns into a short recap does the same while preserving context. And enable prompt caching so whatever history you do keep re-reads cheaply. Trimming the oldest turns of long chats moves the number most.

Are these prices current?

Inference prices sync daily from OpenRouter (updated 2026-10-05). This mold adds no separate price source — it's a token-accounting model of transcript replay on top of the live catalog, so it stays accurate as prices change automatically.

Should a DeepSeek V4 Flash chatbot worry about this?

Less so at 40 turns (1.6× naive), but the multiplier climbs fast with conversation length — worth watching if your sessions run long.