For fifty years, if you wanted a faster computer, you mostly just waited. Every couple of years, transistors got smaller, chips got denser, and software got faster without anyone touching a line of code. That free ride is over. Transistors are still shrinking, but the reliable, predictable doubling that defined an entire industry's planning cycle has broken down — and the gains that used to arrive automatically now have to be engineered deliberately, layer by layer, through architecture, packaging, and code.
This isn't a story about the end of Moore's Law bringing computing to a stop. It's a story about where speed comes from changing shape. Understanding that shift matters whether you're picking cloud instances, designing a product roadmap, or just trying to figure out why your new laptop doesn't feel dramatically faster than the one it replaced.
Quick answer: Moore's Law hasn't ended cleanly, but the automatic, predictable doubling is gone. Speed now comes from four places layered together — specialized hardware (GPUs, ASICs, TPUs), chiplet-based packaging, software and algorithmic efficiency, and early-stage new device physics. The practical upshot: match your workload to the right silicon, and treat software optimization as a durable investment rather than something next year's hardware will paper over.
What Moore's Law Actually Said
Gordon Moore's 1965 observation was narrower than most people remember it. He noted that the number of components that could be economically packed onto an integrated circuit was doubling roughly every year (later revised to about every two years), driven by improvements in manufacturing. It was an observation about transistor density and cost, not a law of physics, and not directly a promise about speed.
The reason Moore's Law felt like a speed guarantee for decades is that it traveled alongside a companion principle: Dennard scaling. Robert Dennard and colleagues showed in 1974 that as transistors shrank, you could increase their density, speed, and switching frequency while keeping power density roughly constant. Put those two effects together and you got the golden era of computing: more transistors, running faster, without a proportional increase in heat or power draw. Clock speeds climbed year over year, and software got faster just by running on newer hardware.
Two Separate Trends, One Shared Ending
It helps to separate the two trends explicitly, because they broke down at different times and for different reasons:
| Trend | What it promised | When it slowed |
|---|---|---|
| Moore's Law | Transistor count per chip roughly doubles every ~2 years | Still continuing, but at rising cost per transistor and with longer intervals |
| Dennard scaling | Power density stays constant as transistors shrink | Broke down in the mid-2000s due to leakage current and heat |
Dennard scaling's collapse is why clock speeds plateaued around 3-4 GHz in the mid-2000s and stayed there — chipmakers could no longer crank frequency without cooking the chip. The industry's answer was multi-core processors: instead of one faster core, you got several cores running at a similar speed. That bought another decade of progress, but it shifted the burden onto software, which had to be rewritten to actually use those extra cores. A lot of code never was.
Why We're Nearing the End of Moore's Law
Moore's Law itself hasn't stopped so much as it has gotten expensive and physically strained. A few concrete forces are pushing back:
- Quantum effects at small scales. As transistor features shrink toward a few nanometers, the insulating layers meant to stop current from leaking become thin enough that electrons can tunnel through them anyway. That leakage wastes power and generates heat even when a transistor is supposedly "off."
- Heat density. Packing more switching elements into the same area means more heat generated per square millimeter. Removing that heat fast enough to avoid throttling performance is now a first-order design constraint, not an afterthought.
- Lithography cost. Each new process node requires more advanced (and more expensive) manufacturing equipment. Extreme ultraviolet lithography tools cost a substantial multiple of the previous generation's equipment, and only a handful of foundries in the world can run them at volume.
- Diminishing economic returns. Even when a smaller node is technically achievable, the cost per transistor doesn't always keep falling the way it used to. For some designs, moving to a newer node no longer makes financial sense.
- Design and verification complexity. Chips with tens of billions of transistors take enormous engineering effort to design, verify, and manufacture without defects, which lengthens development cycles and raises the cost of getting it wrong.
None of this means silicon scaling has stopped entirely — leading-edge nodes continue to ship — but the clean, predictable cadence that let entire industries plan around "wait 18 months, get double the performance" no longer holds. Speed now has to be found somewhere else — and that "somewhere else" is really four somewheres, covered next.
Where Speed Comes From Now
If raw transistor scaling isn't enough on its own, where are the performance gains actually coming from? Mostly from four places, often layered together.
Specialized, Domain-Specific Hardware
The general-purpose CPU used to be the default engine for nearly every workload. Increasingly, it's the fallback for whatever doesn't have a purpose-built accelerator. Graphics processing units, originally designed for rendering images, turned out to be extremely good at the matrix multiplication that underlies both graphics and machine learning, which is why GPUs became central to AI training and inference. Beyond GPUs, companies now design application-specific integrated circuits (ASICs) and tensor-processing units tuned for narrow tasks — neural network inference, video encoding, cryptographic operations — that vastly outperform a general-purpose chip on that one job, at a fraction of the power draw.
The tradeoff is flexibility. A specialized chip is fast precisely because it throws away generality. That's a fine trade when the workload is predictable and large-scale (like data center AI inference), and a poor one when workloads are varied or short-lived.
Advanced Packaging and Chiplets
Instead of trying to shrink one monolithic piece of silicon further, chip designers increasingly build systems out of several smaller "chiplets" — separately manufactured dies, sometimes on different process nodes, connected through high-bandwidth interconnects and stacked or placed side by side in a single package. This approach lets a manufacturer use the most expensive, cutting-edge process only for the components that truly benefit from it (like compute cores), while using cheaper, mature processes for components like I/O controllers that don't need the latest node.
It also improves manufacturing yield, since a single large chip has more area where a single defect can ruin the whole part, while smaller chiplets are individually easier to build correctly.
3D stacking — placing memory or cache directly on top of compute logic instead of beside it — is a related technique that shortens the physical distance data has to travel, cutting latency and power use for memory-bound workloads.
Software and Algorithmic Efficiency
When hardware gains slow down, the same performance improvement has to come from writing smarter code. This shows up in a few concrete ways:
- Better algorithms. A more efficient algorithm can outperform years of hardware progress on the same problem; algorithmic improvements in areas like sparse computation and approximate methods, often first published on arXiv, have delivered order-of-magnitude gains in specific domains.
- Compiler and runtime optimization. Modern compilers increasingly tailor code generation to specific hardware quirks — cache layouts, vector instruction sets, memory hierarchies — squeezing out performance that used to be left on the table.
- Precision reduction. Many workloads, especially machine learning, don't need full 32-bit or 64-bit precision. Running computations in lower-precision formats trades a small amount of numerical accuracy for a large gain in throughput and energy efficiency.
- Better parallelization. Software that's rewritten to actually exploit multiple cores, multiple chips, or distributed clusters can scale performance in ways a single faster core never could.
New Device Physics and Materials
Further out, researchers are exploring approaches that don't rely on shrinking a conventional silicon transistor at all. Photonic computing, which uses light instead of electricity to move and sometimes process data, promises much lower energy loss for data movement between chips. Novel transistor geometries — like gate-all-around designs that wrap the gate around the channel on all sides instead of just three — extend traditional scaling a bit further by improving control over current leakage. Neuromorphic chips, which mimic the event-driven, sparse activity pattern of biological neurons, aim to cut power consumption for certain AI workloads dramatically compared to conventional architectures. None of these are wholesale replacements for silicon transistors yet; they're additive tools being explored in parallel.
Benefits of Post-Moore Computing
Losing the automatic speedup is a real cost, but the approaches that replaced it bring advantages the old model never offered. They reward teams that understand their workloads instead of teams that simply wait.
Large Gains for Workloads That Fit
A general-purpose CPU spends much of its silicon and energy on flexibility: branch prediction, deep caches, instruction decoding for every kind of program. An accelerator built around one pattern, such as dense matrix math or video encoding, drops most of that overhead and spends the area on the operation that matters. For workloads that map cleanly onto that pattern, the result is a jump in throughput that no single process-node shrink ever delivered. AI training and inference are the obvious beneficiaries, but media processing, compression, and some search workloads see the same effect.
Better Performance per Watt
When power and heat became the binding limit, efficiency stopped being a secondary metric. Specialized silicon, lower-precision arithmetic, and 3D-stacked memory all reduce the energy spent per useful operation, mostly by avoiding wasted data movement and unnecessary precision. That matters in two very different places: battery-powered devices, where it extends runtime, and dense data centers, where power delivery and cooling cap how much compute fits in a rack. A design that does the same work on less power can often be deployed where a faster but hotter one cannot.
Lower Cost Through Mixed-Node Chiplets
Monolithic chips force every block onto the same process node, including blocks that gain nothing from it. Chiplets let a designer reserve the expensive leading-edge node for compute cores and build I/O, analog, and memory controllers on mature, cheaper processes. Smaller dies also yield better, since one defect scraps a small piece of silicon instead of a large one. The combined effect is that high-end parts can keep improving without every component paying leading-edge prices, and vendors can reuse proven chiplets across several products.
Software Gains That Compound
In the Moore era, a clever optimization was often overtaken by the next hardware generation within a couple of years. Now an algorithmic improvement, a better memory layout, or a move to lower precision keeps paying off for much longer, because the hardware underneath is no longer erasing the difference. That makes engineering effort on efficiency easier to justify, and it means teams with strong profiling and optimization skills hold an advantage that doesn't evaporate with each product cycle.
More Choice in Infrastructure
A world of heterogeneous hardware gives buyers options that didn't exist when one CPU line served everything. Cloud providers offer instances built around GPUs, custom inference chips, and CPUs with wide vector units, each priced differently. Teams that profile their workloads can pick the silicon that fits rather than paying for general capacity they don't use. The flip side is more complexity in the decision, but for organisations with large, steady workloads, the ability to choose is where much of the cost saving comes from.
Post-Moore Computing Use Cases
The shift is already visible in products people use daily. These are the areas where the post-Moore toolkit is most established, each pairing a hard performance ceiling with a specific technique for getting around it.
AI Training and Inference in Data Centers
Training large neural networks is dominated by matrix multiplication across huge datasets, which general-purpose CPUs handle slowly and expensively. The industry's response has been GPUs and custom tensor accelerators, packaged with high-bandwidth memory stacked close to the compute and run at reduced precision where accuracy allows. Every layer of the post-Moore stack appears here at once: specialized silicon, advanced packaging, memory integration, and software frameworks tuned to the hardware. The outcome is that AI workloads have kept getting faster and cheaper per operation even as general-purpose scaling slowed.
Smartphones and Laptops
Modern phone and laptop processors are no longer just CPUs. They combine CPU cores with GPU, image-processing, and neural-processing blocks on one system-on-chip, each handling the tasks it suits. When a phone processes a photo, transcribes speech, or runs on-device AI features, much of that work runs on dedicated blocks rather than the main cores. The result is better battery life and responsiveness than general cores could deliver, which is also why a new laptop can feel only modestly faster on everyday tasks while being far faster on the specific ones its accelerators target.
Video Streaming and Media Processing
Encoding and decoding video is computationally heavy and runs at enormous scale across streaming platforms and video calls. Dedicated media engines, built into GPUs, phone chips, and some server hardware, perform this work far more efficiently than software running on CPU cores. For a streaming service, that translates into more streams processed per server and lower power per stream. For consumers, it is why devices can play high-resolution video for hours without draining the battery or spinning up the fans.
Scientific Simulation and Engineering
Climate models, molecular dynamics, and fluid simulations once relied on ever-faster CPUs in supercomputers. Much of this work now runs on GPU-accelerated systems, and research groups have rewritten simulation codes to exploit massive parallelism. Some fields also lean on algorithmic advances, such as reduced-precision methods or machine-learned approximations of expensive physics, to get answers faster. The pattern here shows how much of the gain depends on software: the hardware only helps once the code has been restructured to use it.
Networking and Security Offload
Encryption, packet processing, and compression are repetitive operations that consume CPU cycles in servers handling heavy traffic. Many processors now include dedicated cryptographic instructions, and data centers increasingly use network cards with their own processors to take this work off the main CPU. The outcome is more server capacity available for the application itself, and lower latency for traffic-heavy services, without requiring a faster general-purpose chip.
Why This Matters Right Now
For most of computing history, performance was something you could take for granted as a background tailwind. A team could ship inefficient code today, reasonably confident that next year's hardware would paper over the waste. That assumption is no longer safe. When transistor-driven speedups slow down, the gap has to be closed by architecture choices, packaging engineering, and — increasingly — the skill of the people writing software.
This shift is also reshaping who benefits from hardware progress. Under the old model, faster general-purpose chips lifted every workload roughly equally. Under the specialized-hardware model, the benefits are uneven: workloads that map well onto GPUs, ASICs, or other accelerators (large-scale AI training, certain scientific simulations, video processing) keep getting dramatically faster, while workloads that don't fit those patterns see much smaller gains and depend far more on manual software optimization. That unevenness is becoming a real strategic variable for anyone building products on top of computing infrastructure, not just a topic for chip designers.
Common Post-Moore Computing Mistakes
Most of the expensive errors in this area come from carrying Moore-era habits into a world where they no longer pay off.
Assuming the Next Hardware Generation Will Fix Slow Code
Teams still defer performance work on the theory that newer servers will absorb the inefficiency. For workloads that don't map onto the latest accelerators, a hardware refresh may deliver only a modest improvement, and the slow code stays slow at a higher price. The safer assumption now is that inefficiency persists until someone fixes it, and that the fix belongs on the roadmap rather than in a vague future upgrade.
Buying Accelerators Before Profiling
Specialized hardware is attractive, and it is easy to provision GPUs for a workload that spends most of its time waiting on a database, network calls, or single-threaded preprocessing. In that case the accelerator sits mostly idle while the bill grows. Profiling first shows where time actually goes, and frequently the bottleneck turns out to be data movement or I/O rather than raw compute.
Comparing Chips on Peak Throughput Alone
Vendor specifications highlight peak operations per second, but real workloads rarely reach peak. Memory bandwidth, interconnect limits, software support, and power draw often decide actual performance and cost. Choosing hardware on the headline number, without benchmarking a representative workload, regularly leads to parts that look fast on paper and disappoint in production.
Ignoring the Software Ecosystem
An accelerator is only as useful as the compilers, libraries, and frameworks that target it. Teams sometimes select hardware with impressive specs but immature tooling, then spend months porting code or working around missing features. The engineering cost of moving off a well-supported platform can outweigh any price-performance advantage, so toolchain maturity deserves as much weight as the silicon itself.
Treating Power and Cooling as Someone Else's Problem
Dense accelerator deployments draw far more power per rack than traditional servers. Projects that plan compute purchases without checking power delivery, cooling capacity, or cloud instance availability can end up with hardware they cannot run at full speed, or at all. Thermal limits are now part of the performance equation and need to be in the plan from the start.
Practical Implications for Businesses and Builders
The shift away from automatic hardware speedups changes how technical decisions should get made.
| Old assumption (Moore's Law era) | New reality (post-Moore era) |
|---|---|
| Wait a generation, get a free performance boost | Performance gains require deliberate architecture and software choices |
| One general-purpose CPU handles most workloads well | Workload-specific accelerators (GPU, ASIC, TPU) often outperform general compute by a wide margin |
| Inefficient code is "forgiven" by faster hardware next cycle | Inefficient code increasingly stays inefficient; efficiency work has lasting payoff |
| Hardware upgrade cycles drive most performance planning | Software, packaging, and hardware choices all need to be co-designed |
| Cost per unit of compute reliably falls over time | Cost per transistor is flattening at the leading edge; efficiency has to be earned |
The table describes the shift; the practices below turn it into decisions.
Post-Moore Computing Best Practices
- Match the workload to the silicon. Before assuming you need more general-purpose compute, ask whether the workload — inference, encoding, search, simulation — has a specialized accelerator that fits it better. The performance and cost difference can be substantial, but only for workloads whose core operation the accelerator was built for.
- Profile before you provision. Measure where time and money actually go, including I/O, memory stalls, and serialization, before choosing hardware. A profile often shows that the fix is in data layout or batching rather than in a bigger instance.
- Treat software efficiency as a durable investment. Optimization work that used to be a nice-to-have, because hardware would eventually mask the cost of skipping it, now has a longer shelf life and a clearer return. Put it on the roadmap with an owner and a measurable target.
- Benchmark representative workloads, not spec sheets. Run your own code, with realistic data sizes, on candidate hardware before committing. Compare cost per unit of useful work rather than peak throughput.
- Plan hardware refresh cycles around real workload gains, not habit. If a workload doesn't map well onto the newest accelerators, upgrading hardware on the old cadence may deliver much less benefit than it used to.
- Watch power and cooling as first-class constraints, not afterthoughts. As chips pack more into less space, power delivery and heat removal increasingly limit what's deployable, especially in dense data center environments. Confirm rack power, cooling, and instance availability before committing to a hardware plan.
- Expect more heterogeneous systems. Products that used to run on a single type of processor increasingly rely on a mix — CPU plus GPU plus specialized accelerator — which raises the complexity of both hardware selection and software design. Favour portable frameworks and abstraction layers so moving between hardware types doesn't mean a rewrite.
Limitations and Open Questions
None of the post-Moore approaches are clean substitutes for what transistor scaling used to provide automatically, and each comes with real constraints.
Specialized hardware only pays off at sufficient scale and workload predictability; designing and fabricating a custom ASIC is expensive and slow, which makes it a poor fit for workloads that change quickly or run at small volume. Chiplet-based packaging solves some yield and cost problems but introduces new ones around interconnect bandwidth, latency, and thermal management between dies that didn't exist in a single monolithic chip. Precision reduction and algorithmic shortcuts can improve throughput, but they trade away accuracy or generality, and not every workload can tolerate that tradeoff.
There's also a broader open question about how long even the current pace of progress across these fronts can be sustained. Each of the post-Moore techniques — packaging, specialization, software efficiency — has its own diminishing-returns curve, and none of them individually offers the multi-decade runway that transistor scaling once did. It's not yet clear whether stacking these techniques together will deliver a comparably long runway, or whether the industry is instead trading one long, smooth curve of progress for a series of shorter, harder-won gains that each require more engineering effort than the last.
Finally, the economics of leading-edge manufacturing are increasingly concentrated. A shrinking number of foundries can produce the most advanced chips, which raises questions about supply chain resilience and pricing power that didn't matter as much when progress felt more evenly distributed across the industry.
What to Watch Next
A few threads are worth tracking if you want a sense of where post-Moore performance gains are heading:
- Packaging standards. As chiplet designs from different vendors need to interoperate, industry standards for chip-to-chip interconnects will shape how mixed-vendor systems get built.
- Power efficiency, not just raw speed. As thermal and power limits become the binding constraint, expect benchmarks and marketing to shift further toward performance-per-watt rather than peak throughput alone.
- Software toolchains that target heterogeneous hardware. Compilers, runtimes, and frameworks that make it easier to target a mix of CPUs, GPUs, and accelerators without hand-tuning each workload will matter more as hardware diversity increases.
- Alternative computing paradigms maturing beyond the lab. Photonic interconnects and neuromorphic designs are still mostly research and niche deployments; watch for signs they're moving into mainstream production use for specific workloads.
- How memory keeps up. As compute gets faster and more specialized, moving data to and from memory increasingly becomes the bottleneck; advances in memory bandwidth and 3D integration will matter as much as compute itself.
For a deeper look at one of the biggest levers here, see our explainer on why GPUs won AI training. Teams navigating this shift — deciding where specialized hardware, better software, or infrastructure redesign will actually move the needle — can find hands-on help from Woyce Technologies.
FAQ
Is Moore's Law dead?
Not entirely — transistor density is still increasing at the leading edge of manufacturing — but the reliable, low-cost doubling that defined the original observation has slowed considerably and costs more per transistor to achieve than it used to. Most engineers now describe it as slowing and becoming economically strained rather than cleanly "ending" on a specific date.
What replaced Moore's Law as the main source of speed gains?
No single thing replaced it. Speed now comes from a combination of specialized hardware (GPUs, ASICs, TPUs), advanced chip packaging like chiplets and 3D stacking, and software-level efficiency gains such as better algorithms and lower-precision computation. These levers stack rather than substitute for each other: the fastest systems today pair a workload-specific accelerator with careful packaging and heavily optimized software, and each layer contributes a share of the gain.
Why did clock speeds stop increasing years ago?
Clock speeds plateaued in the mid-2000s because Dennard scaling broke down — shrinking transistors further no longer kept power density constant, so pushing clock speeds higher generated more heat than could be practically removed. The industry shifted to adding more cores instead of running each core faster, which moved the burden onto software that had to be rewritten to run in parallel.
What are chiplets, and why do they matter?
Chiplets are smaller, separately manufactured pieces of silicon that are combined into a single chip package instead of building one large monolithic die. They improve manufacturing yield, let designers mix process nodes for cost efficiency, and are a major reason modern high-performance chips keep getting faster despite transistor scaling slowing down.
Does the end of easy Moore's Law gains mean AI progress will slow down?
Not directly. Much of recent AI hardware progress has come from specialized accelerators like GPUs and TPUs rather than general transistor scaling, so AI-specific hardware can keep improving even as general-purpose chip scaling slows. That said, AI workloads still depend on advances in packaging, memory bandwidth, and power efficiency to keep scaling.
Should businesses change how they plan hardware upgrades?
Yes, in the sense that assuming automatic performance gains from waiting for newer general-purpose hardware is less reliable than it used to be. It's worth evaluating whether a workload benefits from specialized accelerators and whether software optimization offers a better return than a hardware refresh. A practical habit is to profile the workload first, estimate the gain from each option, and only then commit budget to new hardware or cloud instance types.
What is Dennard scaling, and how is it different from Moore's Law?
Moore's Law is about transistor density — how many transistors fit on a chip. Dennard scaling is about power — the idea that as transistors shrink, power density stays roughly constant, letting chips run faster without overheating. Dennard scaling broke down in the mid-2000s, well before Moore's Law itself started slowing, and that breakdown is why clock speeds stalled while transistor counts kept climbing.
Key Takeaways
If the free hardware lunch is over for your workload, start with these three moves:
- Profile before you provision. Before defaulting to more general-purpose compute, check whether the workload — inference, encoding, search, simulation — has a specialized accelerator that fits it meaningfully better.
- Fund efficiency work like it's permanent. Optimization that used to be optional, because hardware would eventually mask the cost of skipping it, now has a longer shelf life and a clearer ROI.
- Budget for heterogeneity. Plan for systems that mix CPU, GPU, and specialized accelerators rather than one processor type — and factor power and cooling in as first-class constraints, not afterthoughts.
Conclusion
The core problem is simple to state: the automatic speedup that came from shrinking transistors has slowed, and the cost of each new process node keeps rising. Performance still improves, but it no longer arrives on schedule for everyone. It has to be earned through specialized silicon, chiplet packaging, memory and interconnect engineering, and software that wastes less.
The main insight is that these gains are uneven. Workloads that fit an accelerator, such as AI training and inference, video, or certain simulations, keep improving quickly. General-purpose code that nobody has profiled sees much smaller gains from each hardware generation. That makes workload analysis and efficiency work strategic decisions rather than cleanup tasks.
There are real caveats. Custom ASICs only pay off at scale, chiplets bring their own thermal and interconnect trade-offs, and lower-precision math is not acceptable for every application. Newer device physics, from photonics to neuromorphic designs, is promising but mostly not production-ready yet.
The next step is concrete: pick your most expensive or slowest workload, measure where its time and money actually go, and test whether a better-matched accelerator or an optimization pass beats a routine hardware refresh. If you want a second opinion on where the effort will pay off, our tech consulting team can help you map workloads to the right infrastructure.
