Skip to content
Woyce Technologies
AboutTeamCareersContactStart a project →

The Economics of LLM Inference: What a Token Really Costs

A breakdown of what actually drives the cost of running large language models in production, from prefill and decode to KV cache memory and GPU utilization.

The Economics of LLM Inference: What a Token Really Costs — Woyce Technologies

A single API call to a large language model looks like it costs a fixed amount: some number of cents per million tokens, multiplied by however many tokens went in and came out. That number on the pricing page is real, but it hides a much messier truth underneath. The cost of producing that response was assembled from at least half a dozen different resource costs — some paid once per request, some paid once per token, some paid continuously whether or not anyone is using the GPU at that moment — and the ratio between them shifts depending on prompt length, output length, batch size, and how the serving software is built. That's the core of LLM inference economics, and it's messier than the sticker price suggests.

Understanding that breakdown matters if you're building a product on top of an LLM API, running your own inference stack, or just trying to figure out why a "cheap" model can end up costing more than an expensive one for your specific workload. This piece walks through what actually happens between a request landing on a GPU and a response streaming back, and why the industry has spent the last two years re-architecting serving infrastructure around that breakdown.

Inference is two different jobs wearing one name

"Running inference" sounds like a single operation, but every LLM request is really two distinct computational phases stitched together, and they have almost opposite performance characteristics.

Prefill is the phase where the model processes your entire input prompt at once — the system instructions, the conversation history, the document you pasted in, the tool definitions. The model runs this whole sequence through its layers in a single parallel pass to build up an internal representation of everything it has been given. This is compute-bound: it uses a large fraction of the GPU's raw arithmetic throughput, and it takes roughly the same amount of wall-clock time whether the GPU is otherwise idle or fully loaded, because the work is one big parallelizable matrix operation.

Decode is the phase where the model generates output, one token at a time. To produce token N+1, the model needs to attend back over every token that came before it — the entire prompt plus everything it has generated so far. Because each new token depends on the one just produced, this phase is inherently sequential: you cannot generate token 50 before token 49 exists. Each individual decode step touches comparatively little new computation, but it has to repeatedly read a large amount of state from memory. That makes decode memory-bandwidth-bound rather than compute-bound — the GPU's math units sit mostly idle, waiting on data to arrive from memory.

This split is the single most important fact in LLM inference economics, because it means the two phases compete for different resources. A serving system optimized for one is often wasting the other.

PrefillDecode
BottleneckCompute (FLOPs)Memory bandwidth
ParallelismWhole prompt processed at onceStrictly sequential, one token at a time
GPU utilizationHighLow, per individual request
Scales withInput (prompt) lengthOutput length × context length
Typical latency driverTime-to-first-tokenTime-per-output-token

The KV cache: the memory bill nobody put on the pricing page

The mechanism that makes decode possible without redoing prefill work at every step is the KV cache (key-value cache). Once the model has computed the attention keys and values for every token in the prompt and every token it has generated so far, it stores them in GPU memory rather than recomputing them on each new step. Without this cache, generating a 1,000-token response would mean re-processing the entire growing context 1,000 times over; with it, each new token only needs to compute its own keys and values and attend over what's already cached.

That cache is not free. Its size scales with the number of layers in the model, the number of attention heads, the hidden dimension, and — critically — the total sequence length (prompt plus generated tokens so far). For a long conversation, a large system prompt, or a document-heavy RAG pipeline, the KV cache for a single request can occupy gigabytes of GPU memory. Multiply that by dozens or hundreds of concurrent users, and KV cache memory — not model weights — often becomes the binding constraint on how many requests a GPU can serve simultaneously.

This is why "how much memory does the model need" and "how much memory does serving the model need" are different questions. Model weights are fixed and shared across all requests. KV cache is per-request, grows for the duration of the conversation, and directly limits concurrency: a GPU that could technically serve 200 short requests at once might only manage 20 requests with long contexts before it runs out of memory for their caches.

Comparison of model weights, fixed and shared across all requests, with the KV cache, which is per request, grows through a conversation and caps how many requests a GPU can serve.

Why this matters right now

For most of the last few years, production LLM serving ran prefill and decode on the same GPU, back to back, for every request. That's simple to build, but it's wasteful: a GPU busy running the compute-heavy prefill for one user sits underutilized on memory bandwidth, while a GPU stuck decoding for a dozen users has spare compute capacity going unused. Batching requests together helps, but a long prefill for a new request arriving mid-batch can stall the decode steps for everyone else already generating — a phenomenon serving engineers call head-of-line blocking.

The current shift is to stop treating prefill and decode as one job. Serving stacks are increasingly splitting them onto separate GPU pools — a set of GPUs dedicated to prefill, tuned for maximum compute throughput on incoming prompts, and a separate set dedicated to decode, tuned for memory bandwidth and high concurrency. When a request arrives, its prompt is processed on a prefill node, the resulting KV cache is handed off to a decode node, and generation proceeds there independently. This disaggregation lets each pool be sized, and even use different hardware, for the job it's actually doing, instead of forcing one GPU generation to be simultaneously good at two opposing workloads.

Disaggregated LLM serving: a request is processed on a compute-tuned prefill pool, its KV cache is handed off to a bandwidth-tuned decode pool, and idle caches are offloaded to flash.

Alongside that split, serving stacks are also offloading KV caches to flash storage rather than keeping every cache resident in GPU memory. GPU high-bandwidth memory is fast but scarce and expensive; flash storage is far cheaper per gigabyte, if slower to access. Because a conversation's KV cache doesn't need to sit in GPU memory during the gaps between a user's messages, systems can evict it to flash when idle and reload it when the conversation resumes — trading a small latency hit on reactivation for a large increase in how many concurrent conversations a fixed GPU fleet can support. Combined, prefill/decode disaggregation and flash-backed KV cache offload are the two biggest architectural levers serving teams are pulling to bring the effective cost per token down without waiting for new chips.

What actually drives the per-token price you pay

The published price per million tokens on any provider's pricing page is the output of all of the above, compressed into a single number and averaged across the provider's own traffic mix. A few structural factors explain why that number looks the way it does, and why it differs so much between models and providers.

  1. Model size and architecture. A larger model needs more FLOPs per token in prefill and more memory bandwidth per token in decode. Mixture-of-experts architectures change this calculus by activating only a subset of total parameters per token, which is part of why some very large models are priced closer to mid-sized dense models.
  2. Batching efficiency. Serving many requests together lets the compute-heavy parts of a GPU's work amortize across users, which is straightforward for prefill and harder for decode because requests finish at different times. Continuous batching techniques, where new requests join and finished requests leave a running batch without waiting for a fixed-size group to fill or empty, materially improve GPU utilization and therefore lower cost per token.
  3. Context length. Both prefill cost and KV cache memory scale with total context. A request with a 100,000-token prompt is fundamentally more expensive to serve than one with a 1,000-token prompt, even producing the same length of output — which is why providers increasingly price long-context usage differently, and why prompt caching (reusing a previously computed prefill for a repeated prefix) has become a standard cost lever.
  4. Output length. Because decode is sequential, output tokens are more time-expensive than input tokens per token generated, which is reflected in the output-token price on most providers' pricing pages running several times higher than the input-token price.
  5. Hardware generation and utilization. Newer accelerators offer more memory bandwidth and compute per dollar, but only if serving software can keep them busy. A well-optimized serving stack on older hardware can beat a naive deployment on newer hardware.
  6. Idle capacity and reserved buffer. Providers must keep spare GPU capacity available to absorb demand spikes and avoid queuing. That reserved, sometimes-idle capacity is a real cost that gets folded into the average price, not a cost that only appears when the GPU is actively computing.

Benefits of understanding LLM inference economics

Knowing how prefill, decode, and the KV cache drive cost isn't just background reading. It changes what teams can do with the same budget.

Bills become predictable

Teams that price requests by input, cached input, and output separately, and that know their workload shape, can forecast spend with far more confidence than those multiplying total tokens by a blended rate. Predictable costs make it possible to set product pricing, usage limits, and budgets without waiting for the month-end invoice to find out whether a feature is profitable.

The same budget serves more users

Caching repeated prefixes, trimming outputs, and pruning context can each remove a large share of the work per request. Applied together, they let a product handle more traffic, or longer conversations, on the same spend. For products with thin margins, that can be the difference between a feature that is viable and one that has to be cut.

Model choices get grounded in workload, not sticker price

Understanding phase sensitivity lets teams compare models on the cost of their actual traffic rather than the headline per-token rate. A model that looks cheaper may cost more for a decode-heavy workload, and vice versa. Decisions made this way tend to hold up after launch instead of being reversed when real bills arrive. It also makes routing strategies, with easy queries on smaller models, easier to justify with numbers.

Latency trade-offs become explicit

Once it's clear that low latency usually means paying for reserved GPU headroom, product teams can decide where speed matters and where it doesn't. Interactive features get priority tiers; background jobs move to batch processing. Users get fast responses where they notice them, and the business stops paying premium rates for work nobody is waiting on.

Self-hosting decisions get easier to evaluate

For teams considering running their own models, knowing that KV cache memory often limits concurrency more than model weights do makes capacity planning realistic. It helps estimate how many GPUs a workload genuinely needs and whether serving optimizations like continuous batching are worth the engineering effort. Without that understanding, self-hosting estimates often assume far more concurrency per GPU than long-context traffic allows.

LLM inference cost use cases

Different products stress different parts of the cost structure. These common patterns show how the economics play out in practice.

Document question answering and RAG

Products that answer questions over long documents are prefill-heavy: large inputs, short answers. The main cost levers are retrieving only relevant passages instead of whole documents and caching any stable prefix such as instructions or a frequently queried reference text. Done well, the cost per question drops sharply without losing answer quality, because most of the original prompt was context the model didn't need. Retrieval quality then becomes a cost lever as well as an accuracy one.

Customer support and conversational assistants

Chat products accumulate context every turn, so later messages cost more than early ones even when the user types a single line. Summarising older turns, dropping resolved topics, and keeping system prompts cacheable control that growth. The outcome is a conversation whose cost stays roughly flat over its length instead of climbing with each exchange. Users rarely notice when older turns are summarised, but the bill does.

Agents with many tool definitions

Agent frameworks often send long tool schemas and instructions with every call. Because that prefix rarely changes, it is an ideal candidate for prompt caching. Structuring requests so the static prefix comes first and dynamic content comes last lets most of each call hit the cache, which matters for agents that make many model calls per task. Trimming unused tools from the list also shrinks the prefix for every call.

Long-form generation

Reports, code generation, and structured data extraction with large outputs are decode-heavy. Here output tokens dominate the bill, so format constraints, tighter instructions, and lower reasoning-effort settings where acceptable have the biggest effect. Teams often find that much of the output length came from verbosity nobody needed. Streaming doesn't reduce cost, but it does make shorter, more focused outputs feel faster.

Bulk and offline processing

Classification of large backlogs, nightly summarisation, and evaluation runs don't need instant responses. Sending them through discounted batch processing lets the provider fit them into spare capacity. The work completes on a slower schedule at a lower price, freeing budget for interactive features. Evaluation suites in particular can grow large, so routing them to batch early prevents a quiet line item from becoming a big one.

Common LLM inference cost mistakes

Most surprises on LLM bills come from a handful of habits that look harmless when a prototype is small.

Estimating with a single blended rate

Multiplying total tokens by one average price ignores the gap between input, cached input, and output rates. Estimates built this way can miss real bills by a wide margin, especially for decode-heavy or cache-friendly workloads. Pricing each request class separately takes a little longer and produces numbers that survive contact with production.

Choosing a model on headline price alone

A cheaper per-token model is not necessarily cheaper for your traffic. If the workload is decode-heavy, or if one provider's caching fits your prompt structure better, the more expensive model on paper can cost less in practice. Benchmarks should use your own request shapes. Run a sample of real traffic through each candidate and compare the actual cost and quality, not just the rate card.

Resending the whole conversation every turn

Passing the full history and every retrieved document on each call inflates prefill cost and KV cache memory with each turn. Without pruning or summarisation, long sessions become the most expensive part of the product, often without anyone noticing until usage grows. Track average input tokens per turn over a session to see whether context is growing unchecked.

Breaking the cache with prompt structure

Putting timestamps, user names, or other changing values at the start of a prompt means the shared prefix never matches, so nothing is cached. Small ordering decisions in prompt templates can decide whether caching works at all. Check cache hit rates in your provider's usage data after any template change.

Paying for speed nobody needs

Running background jobs through low-latency interactive endpoints wastes money on reserved headroom. Work that can wait minutes or hours belongs in batch processing, where it costs less and doesn't compete with user-facing traffic. Tag each workload with its latency requirement so the routing decision is made deliberately.

LLM Inference Cost Best Practices

None of this is purely academic if you're shipping a product that calls an LLM API. The economics above translate directly into decisions you can make.

  • Prompt caching is the single most effective lever available to API consumers. If your application repeatedly sends the same system prompt, tool definitions, or long reference document as a prefix, structuring requests so that prefix is cached avoids paying full prefill cost on every call. This is frequently a 5-10x cost reduction on the cached portion.
  • Output length costs disproportionately. Because decode is the sequential, memory-bound phase, trimming unnecessary verbosity in generated output (via prompting, format constraints, or lower reasoning-effort settings) saves more than trimming an equivalent amount from the input side.
  • Context management is a cost control, not just a quality control. Aggressively pruning irrelevant history, summarizing old turns, or retrieving only the relevant passage instead of pasting a whole document — a RAG-vs-long-context tradeoff — reduces both the KV cache footprint and the prefill bill on every subsequent turn of a conversation.
  • Batch-friendly workloads should be batched. Asynchronous or non-latency-sensitive workloads (bulk classification, offline summarization, evaluation runs) are usually eligible for discounted batch processing, because the provider can schedule them into whatever slack capacity exists rather than reserving latency-sensitive headroom for them.
  • Model choice should be matched to phase-sensitivity. A workload dominated by long documents and short answers is prefill-heavy; a workload that generates long structured output from short prompts is decode-heavy. The "best" model or provider for one shape of workload is not automatically the best for the other, even at the same headline price per token.
  • Latency and cost are coupled, not independent. Requesting faster response times (through premium latency tiers, dedicated capacity, or higher-priority queuing) generally means paying for reserved GPU headroom that would otherwise be shared, so the fastest option and the cheapest option are rarely the same choice.

How to estimate inference cost for your own workload

Pricing pages give you rates; your bill depends on the shape of your traffic. A simple process gets you most of the way to a realistic estimate.

Step 1: Profile real requests

Log a representative sample of production or staging requests and record input tokens, output tokens, and how much of each input is a repeated prefix (system prompt, tool definitions, reference documents). Averages hide a lot, so look at the distribution, especially the long tail of very long prompts or outputs.

Step 2: Classify the workload shape

Decide whether the workload is prefill-heavy (long inputs, short answers), decode-heavy (short inputs, long outputs), or conversational (growing context every turn). That shape tells you which levers matter most: caching for repeated prefixes, output limits for decode-heavy work, and context pruning for long conversations.

Decision table matching LLM workload shapes to cost levers: prompt caching for repeated prefixes, output caps for decode-heavy work, context pruning for long chats, batching for offline jobs.

Step 3: Apply the provider's actual rate structure

Price each request class with separate rates for uncached input, cached input, and output, and check whether any of the traffic qualifies for batch discounts. Using a single blended per-token number is the most common source of estimates that miss by a wide margin.

Step 4: Test one optimization at a time and measure

Restructure prompts for caching, cap output length, or route easy queries to a smaller model, then compare real bills or usage dashboards before and after. As the limitations below explain, measuring beats calculating from first principles.

Limitations and open questions

The economics described here are directionally solid but genuinely hard to observe from outside a serving provider. A few caveats are worth holding onto.

Published per-token prices are averages, not marginal costs. A provider's price reflects a blend of traffic across many customers, prompt shapes, and load conditions; your specific workload's actual cost to serve could be meaningfully above or below that average, and you generally have no visibility into which. This makes it difficult to reason precisely about whether a given optimization (say, restructuring prompts for better caching) will move your bill by 5% or 40% — the honest answer is usually "test it and measure" with proper observability, not "calculate it from first principles."

There is also no standardized way to compare serving efficiency across providers. Two providers running the same open-weight model at the same published price may have very different actual margins, GPU utilization, and latency guarantees, because the serving stack — batching strategy, disaggregation, cache offload, hardware generation — is invisible from the API surface. Pricing pages tell you what you pay; they don't tell you what it cost to serve you, which is exactly the gap that makes this a genuinely evolving field rather than settled engineering.

Finally, energy and hardware supply constraints sit underneath all of this and are largely exogenous to any individual serving optimization. GPU availability, power delivery to data centers, and high-bandwidth memory supply all shape the floor on inference costs in ways that clever serving software can't fully offset — architectural efficiency gains buy time and headroom, but they don't eliminate the underlying capital and energy intensity of running frontier-scale models at volume.

What to watch next

A few trends are likely to keep reshaping the cost curve over the next couple of years:

  • Wider adoption of prefill/decode disaggregation as a default serving pattern rather than a specialized optimization only the largest providers use, as the open-source serving frameworks that implement it mature.
  • Cheaper, faster tiers of memory between GPU HBM and traditional flash storage, purpose-built for KV cache offload rather than general storage, narrowing the latency penalty of eviction.
  • More granular, usage-shaped pricing — separate rates for cached versus uncached tokens, batch versus real-time, and long-context versus short-context requests are already appearing and are likely to become the norm rather than the exception.
  • Specialized inference hardware built around the memory-bandwidth-bound nature of decode specifically, rather than general-purpose accelerators originally designed for training workloads.
  • Smaller, more efficient models absorbing more production traffic as routing systems get better at sending easy queries to cheap models and only escalating genuinely hard queries to frontier-scale models, changing the effective blended cost of running an AI product even where per-model pricing stays flat.

Teams trying to actually engineer their prompts, context strategy, and model routing around these cost dynamics — rather than guessing at what moves the bill — can get hands-on help from Woyce Technologies.

FAQ

Why does output cost more per token than input on most LLM APIs?

Generating output happens in the decode phase, which is sequential and memory-bandwidth-bound — each token requires a full pass over the growing context before the next one can be produced. Processing input happens in the parallelizable prefill phase, which uses GPU compute far more efficiently. That structural difference in GPU utilization is why output tokens are typically priced several times higher than input tokens.

What is the KV cache and why does it matter for cost?

The KV cache stores the attention keys and values the model has already computed for a conversation, so it doesn't need to recompute them at every new token. It's the main reason decode is fast, but it consumes GPU memory proportional to context length, and that memory footprint — not model weight size — is often the real limit on how many concurrent users a GPU fleet can serve.

Does a longer conversation cost more even if I only send a short new message?

Yes. Every turn re-processes (or re-uses a cached version of) the full conversation history as context, and the KV cache for that history has to be resident or reloaded for the model to generate a response. Longer running conversations carry rising prefill and memory costs even when each individual new message is short.

What is prefill/decode disaggregation?

It's a serving architecture that runs the compute-heavy prefill phase and the memory-bandwidth-heavy decode phase on separate pools of GPUs, tuned differently for each job, instead of running both phases on the same GPU for every request. It reduces resource contention between the two phases and improves overall throughput per GPU.

Why do some providers charge less for cached tokens?

When a request reuses a prompt prefix that was already processed in a recent request, the serving system can skip re-running prefill for that portion and reuse the previously computed KV cache instead. Since that skips the expensive compute step entirely, providers pass some of that savings back as a lower price for the cached portion of the request.

Is a cheaper model always cheaper to run for my use case?

Not necessarily. A lower per-token price doesn't account for how prefill-heavy or decode-heavy your specific workload is, how well the provider's serving stack batches your traffic pattern, or how much of your context is cacheable. Two models with different sticker prices can end up costing similarly — or invert — depending on your prompt and output shapes.

Will inference costs keep dropping over time?

Historically yes, driven by more efficient model architectures, better serving software (disaggregation, caching, batching), and newer hardware generations. But the rate of decline is not guaranteed to continue at the same pace, since it's bounded by real constraints in energy availability, chip supply, and memory bandwidth that software optimization alone can't remove.

Conclusion

The per-token price on an LLM pricing page is a simplification of a more complicated cost structure. Every request is two jobs: a compute-bound prefill pass over the whole prompt and a memory-bound decode loop that produces output one token at a time, with a KV cache that grows with context and occupies scarce GPU memory. Those mechanics explain why output tokens usually cost more, why long conversations get more expensive with every turn, and why cached prefixes and batch workloads earn discounts.

For teams building on LLM APIs, the useful levers follow from that: cache repeated prefixes, keep outputs no longer than they need to be, prune and retrieve context instead of resending everything, batch work that isn't latency-sensitive, and match models to the shape of the workload rather than the headline rate. The caveat is that published prices are averages, serving efficiency is invisible from outside, and hardware and energy constraints set a floor that software can't remove, so measured results will differ from back-of-envelope estimates.

A good next step is to profile a week of real traffic and price it with separate cached, uncached, and output rates. If you want help redesigning prompts, context strategy, or model routing around those numbers, our LLM integration team can help.

WT

Woyce Technologies

AI & Engineering Team · Woyce

Woyce Technologies builds AI chatbots, LLM integrations, voice AI, and full-stack web applications for businesses in the US, UK, Europe & APAC. Based in Rajkot, Gujarat.

READY TO BUILD?

Let's build something
that actually works.

Tell us about your project. We'll be honest about whether we're the right fit — and if we are, we move fast.