Skip to content
Woyce Technologies
AboutTeamCareersContactStart a project →

Low-Precision AI: What FP8, FP4, and INT4 Mean for Speed and Cost

A plain-language guide to how FP8, FP4, and INT4 quantization make AI models faster and cheaper to run, and where the trade-offs bite.

Low-Precision AI: What FP8, FP4, and INT4 Mean for Speed and Cost — Woyce Technologies

A number that used to take 32 bits to store now takes 4. That single change — using fewer bits to represent the weights and activations inside a neural network — is quietly responsible for most of the recent gains in AI inference speed and cost. It is not a new model architecture, a smarter training algorithm, or a bigger cluster. It is arithmetic getting coarser on purpose, and it works because neural networks turn out to be remarkably tolerant of imprecision.

Low-precision computing — running models in FP8, FP4, or INT4 instead of the FP32 or FP16 formats they were historically trained and served in — has moved from a research curiosity to the default assumption baked into new AI silicon. Understanding what these formats actually are, and where they help and hurt, matters for anyone deciding what hardware to buy, what model size to deploy, or how to budget for inference at scale.

What "precision" means in a neural network

Every number inside a neural network — the weights that define what the model learned, and the activations that flow through it during a forward pass — has to be stored and manipulated in some numeric format. That format has a fixed number of bits, and those bits get split between representing the sign, the exponent (how big or small the number can get), and the mantissa (how many significant digits of precision you get).

The formats in common use today:

FormatBitsTypical useRough dynamic rangePrecision character
FP3232Legacy training, scientific computingVery wideHigh precision, high memory cost
FP1616Older mixed-precision training/inferenceModerateGood precision, narrower range
BF1616Standard training format todayWide (matches FP32 exponent)Lower precision, wide range
FP88Modern training and inferenceNarrow-moderateCoarse, needs scaling tricks
FP44Inference on newest acceleratorsVery narrowVery coarse, needs per-block scaling
INT44Inference, especially on edge/mobile NPUsFixed, integer-onlyCoarse, needs calibration

The pattern across the table is simple: fewer bits means less memory per number, less data to move, and less energy per multiply-accumulate operation — the core arithmetic step that GPUs and NPUs spend nearly all their time on. It also means less precision, which is the whole tension this article is about.

Why fewer bits translates directly into speed and cost

AI accelerators are not primarily bottlenecked by how fast they can do arithmetic — they are bottlenecked by how fast they can move data from memory to the compute units. This is usually called being "memory-bandwidth bound" — the same constraint explored in the memory wall. A model's weights have to travel from high-bandwidth memory into the chip's compute cores for every token generated. If those weights are stored in FP4 instead of FP16, there is a quarter as much data to move, and a quarter as much memory needed to hold them.

The effect compounds three ways:

  1. Memory footprint drops, so a model that needed multiple high-end GPUs to fit in memory at FP16 might fit on a single GPU at FP4 or INT4 — or a large model might fit on a laptop or phone at all.
  2. Data movement drops, which directly increases achievable throughput on memory-bound workloads like token generation.
  3. Raw compute throughput often increases too, because chip designers can pack more low-precision arithmetic units into the same silicon area than high-precision ones, and newer accelerators have dedicated low-precision math units that run at multiples of their FP16 throughput.

Three compounding gains of low-precision inference: smaller memory footprint, less data moved per token, and more low-precision math units per chip, with accuracy risk.

None of this is free. Coarser number representations mean each individual value is a rougher approximation of what the model "really" computed, and that approximation error can accumulate in ways that hurt output quality if not handled carefully.

FP8 vs. FP4 vs. INT4: what actually differs

These three formats get lumped together as "low precision," but they solve different problems and carry different risk profiles.

FP8

FP8 is now a mainstream training and inference format, not just an inference shortcut. It keeps a floating-point structure (sign, exponent, mantissa) but squeezes it into 8 bits, typically split as either 4 exponent bits and 3 mantissa bits, or 5 exponent bits and 2 mantissa bits, depending on whether you want more dynamic range or more precision. Because it retains a floating-point exponent, FP8 handles the very large and very small values that show up during training (gradients, in particular) more gracefully than a fixed-point integer format would. This is why FP8 has become viable not just for running trained models but for training them in the first place, with careful per-tensor or per-block scaling to keep values inside a usable range.

FP4

FP4 pushes the same floating-point idea down to 4 bits — usually something like 1 sign bit, 2 exponent bits, and 1 mantissa bit, though exact layouts vary by vendor. At this bit width there are only a handful of representable values, so FP4 is essentially unusable without aggressive per-block scaling: instead of one scale factor for an entire tensor, the values are grouped into small blocks (say, 16 or 32 values at a time), each with its own scale. This lets a block of small numbers and a block of large numbers both be represented reasonably accurately in FP4, even though the format itself can only express a tiny set of distinct values. FP4 is, as of the current generation of hardware, primarily an inference format — training directly in FP4 remains an active research problem rather than a settled practice.

INT4

INT4 is a fixed-point integer format: 4 bits representing a plain integer (0–15, or -8 to 7 signed), with a separate scale and zero-point value used to map those integers back to the real range of numbers the model needs. INT4 has no exponent, so it cannot natively represent a wide dynamic range the way FP4 can — it depends even more heavily on good calibration (choosing the right scale and zero-point per channel or per block) to avoid clipping outlier values or wasting precision on a range the data never uses. INT4 has been around longer than FP4 in practical deployment, especially on mobile and edge neural processing units (NPUs), because integer arithmetic units are simpler and cheaper to build into silicon than floating-point ones.

Bit layouts of low-precision formats: FP8 as E4M3 or E5M2 with 256 values, FP4 with one sign, two exponent, one mantissa bit, and INT4 as four plain integer bits.

Quick comparison

PropertyFP8FP4INT4
StructureFloating-pointFloating-pointFixed-point integer
Distinct values2561616
Handles wide dynamic rangeWellOnly with block scalingPoorly without calibration
Common roleTraining + inferenceInferenceInference (esp. edge)
Calibration/scaling needsModerateHigh (per-block)High (per-block + zero-point)
Hardware supportBroad on recent GPUsNewest flagship acceleratorsWidespread on NPUs

Why this matters right now

Precision has always been a knob researchers could turn, but for most of deep learning's history it wasn't turned very far — FP16 and BF16 covered the vast majority of training, and INT8 covered a modest slice of inference optimization. What has changed is that the hardware itself has shifted to make ultra-low precision the default path rather than an edge case.

The 2026 generation of flagship AI accelerators — surveyed in more depth in the state of AI inference hardware — is built around FP4 as a first-class format, with dedicated tensor cores designed to run FP4 matrix multiplication at throughput multiples well beyond what the same silicon achieves in FP8 or FP16. That is a deliberate design choice by chipmakers, not an incidental capability — it reflects a bet that the industry's appetite for larger models and more inference volume will keep outpacing the growth in raw silicon capacity, and that precision reduction is the most reliable lever left for closing that gap. At the same time, edge NPUs — the low-power AI chips built into phones, laptops, and other on-device hardware — have moved to ship hardware INT4 support as standard, not as a niche feature reserved for specialized SKUs.

That combination — flagship datacenter silicon optimized for FP4 and consumer edge silicon optimized for INT4 — means low-precision inference is no longer something only a handful of frontier labs use for cost-cutting at massive scale. It is becoming the default execution path across the stack, from the largest cloud deployments down to the model running locally on a phone.

Benefits of Low-Precision AI

If you deploy or build on top of AI models, precision choices show up as concrete line items: GPU hours, memory footprint, latency, and the size of model you can practically ship.

Lower inference cost per token

Serving a large language model is dominated by memory bandwidth during token generation, a cost structure broken down in LLM inference economics. Moving from FP16 to FP8 or FP4 weights can meaningfully cut the cost per million tokens served, because the same GPU can hold a bigger batch in memory and move less data per step.

Bigger models on smaller hardware

A model quantized to INT4 or FP4 might fit on a single GPU or even a high-end laptop where the FP16 version required a multi-GPU server. This changes the calculus for teams weighing self-hosting costs against calling an API. Smaller hardware footprints also simplify deployment in regulated or air-gapped environments, where renting large clusters is not an option.

On-device and edge deployment

INT4 support in edge NPUs is what makes it feasible to run a meaningfully capable model directly on a phone or laptop with acceptable battery life, rather than round-tripping every request to a cloud server. Local inference also keeps sensitive input on the device and works without a network connection.

More throughput from the same hardware

For teams with a fixed GPU allocation, lower precision often means more requests served per second on the same hardware, which is a direct lever on both cost and latency for user-facing products. That headroom can absorb traffic growth without an immediate hardware purchase, or let a team serve a larger model within the same latency budget.

Less energy per operation

Each low-precision multiply-accumulate uses less energy than its high-precision equivalent, and moving fewer bits from memory saves power too. At data-centre scale that reduces electricity and cooling costs per request; on phones and laptops it is the difference between a feature that drains the battery and one users leave switched on. As power becomes a binding constraint for AI deployments, this benefit grows in importance alongside raw cost.

Common Low-Precision AI Mistakes

Precision reduction is not a free upgrade, and treating it as one is the most common mistake in production deployments.

Assuming accuracy loss is uniform

A model quantized to INT4 might show negligible quality loss on casual conversation but measurable degradation on tasks that require precise numerical reasoning, code generation, or long multi-step logic — the exact tasks where errors compound. A single quality score averaged across mixed tasks can hide a serious regression in the one task your product depends on.

Treating quantization as a setting

Getting good results at FP4 or INT4 typically requires calibration against representative data, careful choice of block sizes for scaling, and sometimes quantization-aware fine-tuning rather than simply casting weights down after training. Teams that skip calibration often conclude a format "doesn't work" when the real problem was the process.

Quantizing every layer the same way

Production quantization pipelines often keep certain sensitive layers (embeddings, final output layers, or specific attention components) at higher precision while quantizing the bulk of the model, because uniform quantization across every layer tends to concentrate error in the places that matter most.

Planning around hardware you do not have

FP4 tensor core support exists on the newest flagship accelerators but not on older GPU generations still in wide production use, meaning the cost benefits are only available to teams that have upgraded or are renting the newest hardware tier. Budgets built on FP4 savings need to confirm that the target fleet actually accelerates FP4.

Trusting benchmarks over your own evaluation

A quantized model can look fine on standard benchmark suites while quietly underperforming on a business's specific use case, because benchmark tasks rarely stress the exact failure modes low precision introduces. Build a test set from real requests, including the hardest ones, and compare quantized and full-precision outputs side by side before switching.

Choosing a Precision Level

For teams deciding how far to push precision reduction, the tradeoff generally maps like this:

PriorityReasonable starting point
Maximum accuracy, cost is secondaryBF16 or FP16
Balanced cost/accuracy for most production LLM servingFP8
Aggressive cost reduction, willing to validate carefullyFP4 or INT4 with calibration
On-device / edge deploymentINT4, matched to NPU support
Numerically sensitive tasks (finance, code, math)Keep sensitive layers at higher precision even in a quantized model

Low-Precision AI Use Cases

Low precision is already the default in several kinds of deployment. These are the places teams apply it most often, roughly in order of how widely each is used today.

High-volume LLM serving

A company serving a chat assistant or API to many users pays for every token generated. Moving weights to FP8 roughly halves the memory traffic compared with FP16, letting each GPU hold larger batches and serve more requests. Teams usually find FP8 a comfortable default because accuracy loss tends to be small, then consider FP4 on the newest hardware where the extra savings justify careful validation.

Self-hosting an open model on limited hardware

A team wants to run an open-weight model in its own environment for privacy or cost reasons, but the full-precision version needs more GPU memory than it has. Quantizing to INT4 or FP4 can bring the model within reach of a single GPU or a workstation. The outcome is a self-hosted deployment that would otherwise have required renting multi-GPU servers, provided the quality holds up on the team's own tasks.

On-device assistants and features

Phones and laptops with NPUs run INT4 models for features like summarisation, transcription, and local assistants. Running locally removes network latency, keeps data on the device, and avoids per-request cloud costs. The trade-off is a smaller, more heavily quantized model, so these features are usually scoped to tasks where that is acceptable.

Training large models in FP8

Frontier and large enterprise training runs increasingly use FP8 for much of the computation, with scaling applied per tensor or per block and sensitive operations kept at higher precision. That speeds up training and reduces memory pressure, letting the same cluster train larger models or finish runs sooner.

Batch and offline processing

Jobs such as classifying large document archives, tagging support tickets, or generating embeddings run without a user waiting. Lower precision lets teams process more items per GPU hour. Because outputs can be sampled and checked before use, these workloads are often a low-risk place to try aggressive quantization first, and lessons learned there carry over to latency-sensitive services later.

Low-Precision AI Best Practices

Teams that get good results from quantization tend to follow the same disciplined process:

  • Start from a full-precision baseline. Measure quality, latency, and cost for the unquantized model on your workload first, so every quantized variant has a clear point of comparison.
  • Build a task-specific test set. Use real requests, including numerically sensitive and multi-step cases, rather than relying on public benchmark scores alone.
  • Step down one level at a time. Try FP8 before FP4 or INT4, and stop at the lowest precision that still meets your quality bar instead of jumping straight to the most aggressive format.
  • Calibrate with representative data. For post-training quantization, choose calibration samples that look like production traffic, since scale factors tuned on unrepresentative data can clip important values.
  • Keep sensitive layers at higher precision. Measure per-layer sensitivity and leave embeddings, output layers, or other fragile components in a wider format where it makes a difference.
  • Match the format to your hardware. Check which formats your GPUs or NPUs accelerate natively; a format without hardware support may save memory but deliver little speed.
  • Consider QAT for the hardest cases. If post-training quantization loses too much accuracy at the bit width you need, quantization-aware fine-tuning is often worth the extra compute.
  • Version quantized models like code. Record the format, calibration data, block size, and tooling version for each quantized artefact, so you can reproduce it, compare variants, and roll back quickly if a regression appears.
  • Measure cost, not just speed. Track cost per completed request and energy use alongside latency, because the real business case depends on what each request costs at your actual traffic levels and batch sizes.
  • Monitor quality after deployment. Track user feedback, error rates, and sampled outputs over time, since quantization effects can surface on inputs your test set did not cover.

How quantization actually gets done

There are two broad approaches, and the choice between them is mostly about how much compute and data you're willing to spend to preserve quality.

Post-training quantization (PTQ) takes an already-trained model and converts its weights to a lower-precision format after the fact, usually using a small calibration dataset to determine good scale factors per layer or per block. This is cheap and fast — often a matter of minutes to hours — which is why it's the default path for most teams quantizing an off-the-shelf model.

Quantization-aware training (QAT) simulates low-precision arithmetic during training or fine-tuning itself, so the model's weights adapt to the rounding and clipping effects it will experience at inference time. This produces better accuracy at very low bit widths (like INT4 or FP4) than PTQ typically achieves, but it costs substantially more compute and requires access to training infrastructure and representative data, which many teams deploying third-party models simply don't have.

Comparison of post-training quantization, which is fast and needs only calibration data, against quantization-aware training, which keeps more accuracy at FP4 or INT4 but costs more.

Between these two extremes sit a range of lighter-weight techniques — outlier-aware quantization that keeps a small number of unusually large values at higher precision while quantizing the rest, and mixed-precision schemes that vary bit width by layer based on measured sensitivity. Most production quantization today is some blend of these ideas rather than a pure PTQ-versus-QAT choice.

Limitations and open questions

Low-precision AI is genuinely useful, but it is not a solved problem, and several open questions remain live:

  • There is no universal "safe" bit width. What counts as acceptable quality loss depends entirely on the application — a customer support chatbot can tolerate more approximation error than a system generating financial calculations or medical dosage information.
  • Training in FP4 is still immature. While FP4 inference is well-supported on new hardware, training directly in FP4 remains difficult because gradients are far more sensitive to precision loss than trained weights are; most FP4 deployments today are quantized after training in a higher-precision format.
  • Evaluation is harder than it looks. Measuring the real-world impact of quantization requires task-specific evaluation, not just aggregate benchmark scores, and many organizations lack the evaluation infrastructure to catch subtle degradation before it reaches users.
  • Hardware and software support are still catching up to each other. New precision formats often arrive in hardware before software frameworks fully support them, and vice versa, creating a lag where the theoretical benefits of a format like FP4 aren't fully realized in production toolchains yet.
  • Precision reduction interacts with other optimizations. Techniques like speculative decoding, key-value cache compression, and mixture-of-experts routing all interact with quantization in ways that are not yet fully standardized, making it hard to predict combined effects without direct testing.

What to watch next

A few developments are worth tracking if you want to stay ahead of where low-precision AI is headed:

  1. Whether FP4 training moves from research to practice. If reliable FP4 training methods mature, the cost of training frontier models could drop meaningfully, not just the cost of serving them.
  2. Software framework support catching up to hardware. Watch whether mainstream inference frameworks, including NVIDIA's own tooling, make FP4 and INT4 deployment as close to a default, low-effort setting as FP16 is today, rather than requiring specialized tooling.
  3. Standardization of block-scaling schemes. Different vendors currently use different block sizes and scaling conventions for FP4 and INT4, which fragments tooling; convergence on shared standards, like the Open Compute Project's microscaling format work, would make cross-hardware deployment easier.
  4. Sub-4-bit formats. Research into 2-bit and 3-bit quantization is active, and while quality loss is currently steep at those widths, the same forces pushing FP8 to FP4 could eventually push further.
  5. Edge NPU capability growth. As more devices ship hardware INT4 (and eventually FP4) support, expect more capable models to run fully on-device, shifting some workloads away from the cloud entirely.

Teams weighing precision trade-offs against real cost and latency budgets, whether for a new deployment or an existing LLM integration, can get a second opinion from Woyce Technologies.

FAQ

What is the difference between FP8 and INT4?

FP8 is a floating-point format with a sign, exponent, and mantissa, which lets it represent a wide range of magnitudes reasonably well. INT4 is a fixed-point integer format with no exponent, so it depends much more heavily on calibrated scale and zero-point values to map its limited set of integers back onto the real range of numbers a model needs.

Does quantization hurt model accuracy?

It can, and the amount of degradation depends on the bit width, the technique used, and the task. FP8 typically shows minimal accuracy loss compared to FP16, while INT4 and FP4 require more careful calibration or quantization-aware training to avoid noticeable degradation, especially on numerically sensitive tasks like math or code generation.

Why are chipmakers building hardware specifically for FP4?

Because inference at scale is largely bottlenecked by memory bandwidth and by how many low-precision operations a chip can perform per second, and FP4 lets more values be moved and computed in the same amount of memory and silicon area than higher-precision formats. Building dedicated FP4 tensor cores lets flagship hardware serve more requests per dollar of hardware and power.

Can you train a model directly in FP4?

Training directly in FP4 is still an active research area rather than standard practice, because gradients are more sensitive to precision loss than already-trained weights. Most FP4 deployments today take a model trained in a higher-precision format like BF16 or FP8 and quantize it down for inference. Research groups are testing FP4 training with careful block scaling and keeping sensitive operations at higher precision, but until those methods are proven at scale, plan on FP4 as an inference format.

What is the difference between post-training quantization and quantization-aware training?

Post-training quantization (PTQ) converts an already-trained model to lower precision using a calibration dataset, and is fast but can lose more accuracy at very low bit widths. Quantization-aware training (QAT) simulates low-precision effects during training or fine-tuning itself, producing better accuracy at formats like INT4 or FP4, at the cost of significantly more compute and data requirements.

Is INT4 only used on edge devices?

No — INT4 is common on edge NPUs because integer arithmetic is cheaper to implement in low-power silicon, but it is also used in datacenter inference when teams want maximum cost reduction and can tolerate or mitigate the accuracy trade-offs through careful calibration. Whether it fits depends on the workload, so test it on your actual tasks first.

How do I know if my use case can tolerate FP4 or INT4?

The safest approach is task-specific evaluation on your actual workload rather than relying on general benchmark scores, since quality loss from low precision is highly dependent on the specific task. Numerically sensitive applications like financial calculations, code generation, or multi-step reasoning tend to be more affected than conversational or classification tasks.

Conclusion

Most AI inference is limited by how fast weights can move from memory to compute, not by arithmetic itself. Reducing precision from FP16 to FP8, FP4, or INT4 cuts the data that has to move, shrinks the memory a model needs, and lets newer chips use dedicated low-precision math units. That is why the newest datacenter accelerators treat FP4 as a first-class format and edge NPUs ship with INT4 as standard.

The trade-off is approximation error, and it is uneven. FP8 is now a safe default for most production serving and even for training. FP4 and INT4 can deliver much larger savings but depend on good calibration, block scaling, and sometimes quantization-aware fine-tuning, and they tend to hurt precise tasks such as maths, code, and long multi-step reasoning more than casual conversation. Hardware support is also uneven, so the savings are only available on recent silicon, and benchmark scores can hide regressions in your specific workload.

The practical next step is to quantize a candidate model at two or three precision levels and compare them on a test set built from your own real requests, tracking cost per request alongside quality. If you want help running that evaluation or tuning a self-hosted deployment, our AI and machine learning team can work through it with you.

WT

Woyce Technologies

AI & Engineering Team · Woyce

Woyce Technologies builds AI chatbots, LLM integrations, voice AI, and full-stack web applications for businesses in the US, UK, Europe & APAC. Based in Rajkot, Gujarat.

READY TO BUILD?

Let's build something
that actually works.

Tell us about your project. We'll be honest about whether we're the right fit — and if we are, we move fast.