Skip to content
Woyce Technologies
AboutTeamCareersContactStart a project →

Why Regulators Measure AI in FLOPs: Compute Thresholds Explained

A plain-language look at why AI regulators use raw training compute — measured in FLOPs — as the trigger for oversight, and what that means for model builders.

Why Regulators Measure AI in FLOPs: Compute Thresholds Explained — Woyce Technologies

Ask a regulator what makes an AI model dangerous enough to warrant oversight, and you might expect an answer about what the model can do — whether it can help synthesize a pathogen, write working exploit code, or manipulate someone into a bad decision. Instead, the answer increasingly comes down to a number stamped in a training log: how many floating-point operations it took to train the thing. If that number crosses a line — an AI compute threshold, commonly set at 10^25 FLOPs — the model gets a different regulatory status, whether or not anyone has demonstrated it's actually more dangerous than the model just below the line.

This is compute-based AI governance, and it's no longer a theoretical policy proposal. It's operational law in the European Union, it shaped (and briefly defined) US federal policy, and it's the backbone of how several governments currently decide which AI labs get extra scrutiny. Understanding how it works — and where it breaks down — matters for anyone building, deploying, or investing in frontier AI systems.

This explainer covers what a compute threshold is, how training FLOPs are estimated, why regulators chose compute over capability tests, what it means for builders, where the approach breaks down, and the developments most likely to reshape it.

What a compute threshold actually is

A FLOP (floating-point operation) is a single arithmetic step — an addition or multiplication — performed during model training. Training a large model requires an enormous number of these operations, roughly proportional to the size of the model (its parameter count) multiplied by the amount of data it's trained on. Researchers estimate this quantity, often called "training compute," using well-established scaling formulas rather than counting operations one by one.

A compute threshold is a regulatory line drawn at a specific FLOP count. Cross it, and a model is presumed to fall into a stricter oversight category — subject to things like mandatory risk assessments, incident reporting, red-teaming obligations, or advance notification to a government AI safety body. Stay below it, and a developer generally faces lighter or no comparable obligations.

The appeal is that FLOPs are:

  • Measurable before deployment. A lab knows its training compute budget before it ever runs an evaluation, so the threshold can trigger obligations early, not after harm occurs.
  • Hard to fake retroactively. Cloud compute usage, GPU-hours, and cluster specifications leave a paper trail that's harder to obscure than, say, self-reported capability claims.
  • Correlated (imperfectly) with capability. Empirically, more training compute has tended to produce more capable general-purpose models — this is the premise behind "scaling laws" that guided the last several years of frontier AI development.

That correlation is doing a lot of work, and it's also the threshold approach's biggest weakness — more on that below.

How the numbers actually work

FLOP counts for frontier models are almost incomprehensibly large, so it helps to see them written both ways.

NotationApproximate valueRough real-world anchor
10^23 FLOPs100,000,000,000,000,000,000,000Compute scale of several earlier-generation large language models
10^25 FLOPs10,000,000,000,000,000,000,000,000The EU AI Act's "systemic risk" presumption threshold for general-purpose AI models
10^26 FLOPs100,000,000,000,000,000,000,000,000The reporting threshold set in the 2023 US executive order on AI (later rescinded)

A useful way to think about the jump between these numbers: each additional power of ten is roughly an order of magnitude more compute, which in practice means meaningfully larger clusters, longer training runs, and higher electricity and hardware costs. The gap between 10^24 and 10^25 FLOPs isn't a rounding error — it's the difference between a model trainable on a modest cluster and one that requires the kind of infrastructure only a handful of organizations worldwide currently operate.

Regulators don't expect companies to count operations manually. In practice, compliance relies on standardized estimation methods — using known formulas relating parameters, tokens, and hardware utilization — that labs are expected to calculate and, in some jurisdictions, report to a supervising authority.

Why regulators reach for compute in the first place

Government bodies writing AI rules faced a genuine problem: how do you regulate a technology whose capabilities are discovered through use, often after deployment, and where "harm" can range from biased hiring recommendations to assistance with weapons development? Capability-based rules — write law around what a model can actually do — require running extensive evaluations, agreeing on what constitutes a dangerous capability, and doing this fast enough to keep pace with model releases. That's slow, contested, and easy to game by simply not disclosing the capability that triggers a rule.

Compute is available as a shortcut because of a specific historical pattern: over the 2018–2024 period, the most capable general-purpose models tended to be the ones trained with the most compute. Scaling laws research from that era showed a fairly predictable relationship between compute invested and downstream performance on benchmarks. Policymakers concluded that flagging outlier levels of training compute was a reasonable, verifiable proxy for flagging outlier levels of general capability — not perfect, but auditable, and available before a model ever reaches the public.

There's also a practical enforcement logic. Training runs at the frontier require thousands of specialized chips running for weeks or months, procured from a small number of cloud and chip providers. That concentration makes compute a chokepoint regulators can actually monitor, unlike, say, the diffusion of a trained model's weights once released.

Benefits of AI compute thresholds

For all their flaws, compute thresholds give regulators and the industry several things that capability-based rules struggle to provide. Understanding them explains why the approach has spread so quickly.

Obligations that apply before release

Because a lab knows its training budget before it runs a single evaluation, a compute threshold can attach obligations at the planning stage. Risk assessments, red-teaming, and notification can happen before a model reaches users rather than after a problem surfaces. That timing is the main thing regulators wanted and could not get from rules that depend on observing a model's behaviour in the wild.

A number that can be audited

GPU-hours, cluster specifications, and cloud invoices create records that exist independently of the developer's own claims. An auditor or supervising authority can check an estimate against those records using standard formulas. Compared with self-reported capability claims or evaluations a lab designs itself, that is a much firmer basis for enforcement and a harder one to quietly game.

Predictability for developers

A published FLOP line tells a lab, months in advance, which regime its next model will fall under. Teams can plan compliance budgets, governance staffing, and documentation alongside the training run instead of waiting for a regulator's judgement after release. Even developers who dislike the threshold generally prefer a clear rule to an open-ended assessment of what their model "might" be able to do.

A light touch for most of the market

Because the line sits at the extreme end of training scale, the vast majority of developers, fine-tuners, and startups fall well below it. Heavier obligations land on the small number of organisations running frontier-scale training, which are also the ones with the resources to meet them. Smaller players face baseline transparency rules rather than the full systemic-risk regime.

An enforceable chokepoint

Frontier training needs specialised chips in large quantities, supplied by a handful of manufacturers and cloud providers. Tying oversight to compute lets regulators work through that concentrated supply chain, which is far easier to monitor than the spread of model weights after release or the many ways a deployed model can be used.

Why this matters right now

The clearest evidence that compute thresholds have moved from theory to enforcement is the European Union's AI Act. Under the Act, general-purpose AI (GPAI) models are presumed to carry "systemic risk" if the cumulative compute used for their training exceeds 10^25 FLOPs — a threshold that pulls in additional obligations around model evaluation, adversarial testing, incident tracking, and cybersecurity.

From August 2, 2026, the EU's AI Office gains the power to actually fine GPAI providers for non-compliance. That single date matters because it converts the 10^25 FLOP line from a classification exercise into something with real financial consequences attached. Before that point, a provider crossing the threshold faced obligations on paper; after it, the AI Office can act on non-compliance. The threshold itself doesn't change on that date — what changes is that the number now decides, with enforcement teeth behind it, which providers are inside the EU's systemic-risk regime and which are not.

This is the pattern to watch across jurisdictions: the FLOP number is set first, often years ahead of enforcement, and it quietly becomes the dividing line everyone builds their compliance planning around long before any fine is issued.

What this means for builders and businesses

If you're training foundation models, procuring frontier compute, or building products on top of general-purpose AI systems, compute thresholds create concrete planning questions rather than abstract policy debate.

  1. Model developers near the line have a real incentive to stay under it. A lab training a model at 9 × 10^24 FLOPs and one training at 1.1 × 10^25 FLOPs may produce models with barely distinguishable capabilities, but very different compliance obligations. Expect some developers to deliberately architect training runs to land just under published thresholds — a behavior regulators are aware of and actively trying to design around with parallel qualitative triggers.
  2. Downstream deployers inherit exposure indirectly. A business fine-tuning or deploying a systemic-risk-tier model doesn't itself cross a FLOP threshold, but it may face due-diligence expectations, contractual requirements from the model provider, or supply-chain risk questions from its own customers and auditors.
  3. Procurement and infrastructure decisions become compliance decisions. Where you train, how much compute you reserve, and how you document GPU-hours are no longer purely engineering or cost questions — they're inputs to a regulatory filing in some jurisdictions.
  4. Thresholds shape competitive dynamics, not just safety posture. A hard compute ceiling changes the calculus around whether to pursue a larger training run at all, potentially favoring companies pursuing efficiency gains (more capability per FLOP) over companies pursuing raw scale.

A quick comparison of the compliance postures

Position relative to thresholdTypical obligationsPractical business question
Well below thresholdBaseline transparency and documentation requirementsCan efficiency gains deliver the capability we need without crossing the line?
Near thresholdSame as above, plus active monitoring of cumulative training computeDo we track and log compute precisely enough to prove where we sit?
Above thresholdSystemic-risk obligations: evaluations, incident reporting, cybersecurity measures, possible advance notice to regulatorsDo we have the governance function to meet reporting deadlines and respond to inquiries?

The real limitations of measuring risk in FLOPs

Compute thresholds are a pragmatic compromise, not a scientifically validated measure of danger, and the gap between the two is where most serious criticism lands.

  • Capability doesn't scale linearly with compute anymore. Algorithmic improvements, better data curation, architecture changes (like mixture-of-experts designs that activate only a fraction of parameters per inference), and post-training techniques such as fine-tuning and reinforcement learning can produce large capability jumps without a proportional compute increase. A smaller model trained more efficiently can match or exceed the practical capability of a larger, more compute-intensive one — while sitting comfortably under a FLOP threshold.
  • Inference-time compute complicates the picture further. Some of the most capable recent systems get much of their performance not from training compute but from spending more compute at answer time — reasoning through a problem step by step before responding. A threshold pegged purely to training compute doesn't capture this at all.
  • Thresholds create a moving target that ages badly. A number that looked appropriately restrictive when set can become either too loose (as hardware and algorithms improve, yesterday's frontier-defining number becomes tomorrow's mid-tier model) or arbitrary (once a large share of commercially available systems sit above it, the "systemic risk" label stops functioning as meaningful differentiation).
  • It's an imperfect proxy for the harms regulators actually care about. A model trained with modest compute but fine-tuned specifically for a narrow harmful use case may pose more concrete risk than a much larger general-purpose model with strong safety mitigations. Compute measures scale of training investment, not intent, deployment context, or downstream safeguards.
  • Verification is harder than it sounds. While cloud compute leaves records, organizations training on owned infrastructure, using distributed or federated training approaches, or operating outside jurisdictions with reporting requirements can make exact FLOP accounting difficult to independently confirm.

Most regulatory frameworks that use compute thresholds acknowledge these limits by treating the FLOP number as a presumption rather than an absolute determination — the EU AI Act, for instance, allows both upward designation (a regulator can flag a model as systemic-risk even below the threshold, on other evidence) and rebuttal (a developer can argue their above-threshold model doesn't actually pose systemic risk). In practice, though, the bright-line number does most of the day-to-day work, because presumptions are administratively cheap and case-by-case capability review is not.

AI compute threshold use cases beyond a single law

Training-compute thresholds get the most attention because they're the cleanest version of the idea, but the same logic shows up in several places. Separating them keeps the terminology from blurring together.

Classifying general-purpose models under the EU AI Act

The most prominent live use is the EU's systemic-risk presumption for general-purpose AI models trained above 10^25 FLOPs. Providers estimate cumulative training compute, compare it with the line, and take on evaluation, incident reporting, and cybersecurity obligations if they cross it. Regulators can also designate models below the line on other evidence, and providers above it can argue for rebuttal, so the threshold works as a first-pass filter rather than a final verdict.

Federal reporting for large training runs

The 2023 US executive order on AI used a 10^26 FLOP threshold to decide which developers of dual-use foundation models had to report training plans and safety testing to the government. The order was later rescinded, but it showed how a compute line could drive reporting duties without any capability test at all, and it shaped how other governments thought about setting their own numbers.

Export controls on chips

Export controls on chips use compute-adjacent metrics too, restricting the sale of accelerators above certain performance specifications to particular destinations. The theory is that limiting access to the hardware needed for large training runs is an upstream lever on the same risk that training-compute thresholds address downstream. Here the threshold applies to a chip's capability rather than a model's training history.

Cloud compute reporting requirements

Some proposals ask infrastructure providers, rather than model developers, to flag when a customer's usage pattern looks like a large training run. That turns the threshold question into something cloud providers have to answer about their own customers, not something only the lab training the model self-reports. It adds a second, independent source of evidence about who is training at scale.

Tracking cumulative compute across model families

Some frameworks look at the compute used for a single training run; others look at cumulative compute across a family of related models or fine-tuning passes. That distinction matters for organisations that iterate frequently on a base model rather than training one large model from scratch, because repeated post-training can add up. Labs use it internally to decide when a new version needs a fresh compliance review.

None of these approaches solves the core limitation described above — that compute is a proxy, not a direct measure of harm — but together they show that "measure the resource, regulate the resource" has become a default policy instinct across AI governance, not a quirk unique to one law.

How this differs from capability-based regulation

It's worth being explicit about the alternative, because most debates in this space are really debates between two philosophies.

ApproachHow it decides who's regulatedMain advantageMain drawback
Compute thresholdFLOP count crosses a fixed numberMeasurable in advance, hard to fake, simple to auditWeak, decaying link to actual capability or harm
Capability-basedModel performance on defined evaluations or benchmark tasksDirectly targets the risk regulators care aboutRequires running and agreeing on tests; slower; easier to obscure results
Deployment/use-case basedWhere and how the model is actually applied (e.g., hiring, credit, critical infrastructure)Focuses on real-world harm, not model internalsDoesn't address general-purpose models before deployment; requires tracking downstream use
Hybrid (threshold plus rebuttal)Compute sets a presumption; qualitative evidence can override itCombines auditability with flexibilityAdds administrative complexity and litigation risk over classification

Most mature frameworks, including the EU's, are converging on some version of the hybrid row — using compute as the trigger for a first pass, then layering capability and use-case considerations on top for the models that matter most. That's a tacit admission that compute alone was never meant to be a complete answer, just a workable starting filter given how fast the underlying models were shipping relative to how slowly capability-testing standards could be agreed upon.

Common AI compute threshold mistakes

Organisations misread compute thresholds in a few recurring ways, usually by treating a narrow legal trigger as if it said more than it does.

Assuming "below the line" means no obligations

Staying under 10^25 FLOPs keeps a model out of the systemic-risk presumption, but it doesn't exempt the provider from baseline transparency and documentation duties, or from use-case rules elsewhere in the same law. Regulators can also designate a model below the line on other evidence. Teams that plan only around the threshold can find themselves unprepared for obligations that never depended on it.

Treating the threshold as a safety rating

A model above the line is not necessarily more dangerous than one below it, and a model below it is not certified safe. Compute measures training investment, not intent, deployment context, or safeguards. Using threshold status as shorthand for risk in procurement or board reporting gives a false sense of precision, and can lead teams to skip the use-specific risk assessment they actually need.

Logging compute too loosely to prove a position

Developers near the line need to show where they sit, not just believe it. Rough estimates made after the fact, missing records for experimental runs, or no record of fine-tuning passes make that hard. If a regulator asks, a team without contemporaneous logs has to reconstruct numbers it should have captured during training, and may not be believed.

Assuming deployers are unaffected

Businesses building on a systemic-risk-tier model don't cross any threshold themselves, so they often assume the rules are someone else's problem. In practice, provider obligations flow downstream as contract terms, documentation expectations, and questions from customers and auditors. Ignoring that leaves gaps in vendor due diligence that are awkward to close later.

Planning around one jurisdiction

The EU number, the rescinded US number, and whatever other governments adopt don't line up neatly. A company that maps its exposure under one regime and assumes the same applies elsewhere can be caught out when another market sets a different line or measures compute differently.

AI compute threshold best practices

Whether you train models, fine-tune them, or build products on top of them, a few habits make threshold rules much easier to live with.

Log training compute as you go

Record parameters, tokens, hardware, utilisation, and run duration for every significant training and fine-tuning run, using a documented estimation method. Keep the underlying cloud and cluster records alongside. Contemporaneous logs are far more credible than estimates reconstructed months later.

Track cumulative compute, not just the latest run

Maintain a running total for each model family, including post-training work. Set an internal review trigger well below the legal line so that a new version doesn't cross a threshold without anyone noticing until after release.

Map your dependencies on large models

List every foundation model your products rely on, note which providers sit in a systemic-risk tier, and record what documentation and contractual commitments each provides. Revisit the list whenever you switch models or a provider releases a new version, and ask providers directly how they classify each model.

Pair threshold status with your own risk assessment

Use compute status as one input, not the conclusion. Assess risk based on what your system actually does, where it is deployed, and what safeguards are in place, so your governance holds up even if thresholds change.

Watch more than one jurisdiction

Assign someone to track compute rules in every market you operate in, including any recalibration of the numbers themselves. Thresholds are expected to be revisited as training efficiency improves, and early notice gives you time to adjust plans.

Build governance before you need it

If your roadmap could plausibly approach a threshold, set up evaluation, incident reporting, and security processes before the run that crosses it. Retrofitting them under a regulatory deadline is slower and more expensive than building them alongside the work.

What to watch next

A few developments will determine whether compute thresholds remain the dominant regulatory tool or get supplemented — or replaced — by other approaches:

  • Whether thresholds get revised as training efficiency improves. If regulators don't periodically recalibrate the FLOP number, it risks either capturing far more models than intended or becoming irrelevant as a filter.
  • How enforcement actually plays out post-August 2026. The EU AI Office's first enforcement actions against GPAI providers will signal how strictly the compute threshold gets applied in practice, and how much weight is given to qualitative risk factors alongside the raw number.
  • Whether other jurisdictions converge on similar numbers or fragment. A global patchwork of different FLOP thresholds would create real compliance complexity for any lab operating across multiple markets — echoing the broader patchwork of global AI regulation already emerging — and would put pressure on international standard-setting bodies to harmonize definitions.
  • Whether inference-time and post-training compute get folded into future frameworks. As reasoning-heavy, inference-compute-intensive systems become more common, expect proposals to measure "effective compute" or capability-adjusted metrics rather than raw training FLOPs alone.
  • How smaller, highly capable models are treated. If efficient models under the threshold begin to match the practical capabilities of above-threshold systems, expect regulatory and public pressure to either lower the threshold or introduce capability-based triggers as a supplement.

Teams that need help translating thresholds like these into concrete compliance and model-governance workflows can find hands-on support from Woyce Technologies.

FAQ

What is a FLOP in the context of AI regulation?

A FLOP (floating-point operation) is a single basic arithmetic calculation performed during model training, such as an addition or multiplication. Regulators use the total estimated FLOPs used to train a model — its "training compute" — as a proxy for how capable and potentially risky that model might be. It's a measure of scale, not of behavior, which is both its strength and its weakness.

What is the EU AI Act's compute threshold?

The EU AI Act sets a presumption of "systemic risk" for general-purpose AI models trained using cumulative compute greater than 10^25 FLOPs. Models above that line face additional obligations, including risk evaluation, incident reporting, and cybersecurity requirements, and from August 2, 2026 the EU AI Office can fine providers found non-compliant.

Why do regulators use compute instead of testing what a model can actually do?

Compute is measurable before a model is trained or deployed, harder to misrepresent than self-reported capability claims, and historically correlated with general capability due to scaling laws. Capability-based rules require running and agreeing on evaluations, which is slower and easier to game by not disclosing results. Compute was the number regulators could actually check in time.

Can a company avoid a compute threshold by training a smaller but more efficient model?

Yes, and this is one of the main criticisms of the approach. Algorithmic efficiency gains, better data, and techniques like mixture-of-experts architectures can let a model achieve strong capability using less training compute, potentially keeping it under a regulatory threshold while performing comparably to models above it. Regulators are aware of this, which is why many frameworks allow qualitative designation too.

Does the United States have a compute threshold for AI models?

The US set a 10^26 FLOP reporting threshold for dual-use foundation models in a 2023 executive order, but that order was later rescinded, and US federal policy in this area has continued to evolve. Businesses operating across jurisdictions should track requirements separately rather than assuming one country's framework applies elsewhere.

Are compute thresholds the same as capability requirements?

No. A compute threshold measures the scale of resources used to train a model, not what that model can actually do once trained or fine-tuned. Most frameworks treat crossing the threshold as a presumption of risk that can, in principle, be rebutted or supplemented with additional qualitative evidence. Treat it as a filter that decides who gets looked at closely.

Will compute thresholds get replaced by other regulatory approaches?

Not immediately, but most policymakers treat them as a starting point rather than a permanent solution. Expect ongoing proposals — in the spirit of frameworks like the NIST AI Risk Management Framework — to incorporate inference-time compute, capability evaluations, or "effective compute" metrics that adjust for algorithmic efficiency, especially as the link between raw training compute and real-world capability continues to weaken.

Conclusion

Compute thresholds exist because regulators needed a trigger they could apply before a model shipped, and training FLOPs were the most measurable, auditable number available. The 10^25 line in the EU AI Act and the now-rescinded 10^26 US reporting threshold both follow that logic: measure the resource, presume the risk, and layer obligations on top.

The weakness is built into the premise. Compute is a proxy for capability, and that proxy decays as training gets more efficient, as distillation and mixture-of-experts architectures spread, and as more capability comes from inference-time compute rather than training. That's why the more mature frameworks have drifted toward a hybrid model, with compute as a first-pass filter and capability or use-case evidence on top.

For most businesses, the practical point is that thresholds affect you indirectly. Few companies train models anywhere near 10^25 FLOPs, but many build on models that do, and the obligations attached to those providers shape documentation, transparency, and contract terms downstream.

A good next step is to list the foundation models your products depend on and check which providers have been designated as systemic-risk GPAI under the EU regime. If you need help building model governance into your AI stack, our LLM integration team can help you set that up.

WT

Woyce Technologies

AI & Engineering Team · Woyce

Woyce Technologies builds AI chatbots, LLM integrations, voice AI, and full-stack web applications for businesses in the US, UK, Europe & APAC. Based in Rajkot, Gujarat.

READY TO BUILD?

Let's build something
that actually works.

Tell us about your project. We'll be honest about whether we're the right fit — and if we are, we move fast.