Mời bạn đọc theo dõi "Featured Post":

Hoctroviet — bản đồ một năm

10.03.2026

The Five-Layer Cake: What Each Layer Is, How It Works, Why It Matters — and Why Nvidia Bought Hugging Face

Written by: Claude Opus AI.

Curator/Editor: Học Trò.


Jensen Huang has a habit of compressing a whole industry into a single image he can draw on a napkin. For the last year, that image has been a cake. AI, he says, is five layers stacked on top of each other: energy at the bottom, then chips, then the infrastructure he calls "AI factories," then models, and on top the applications that people and companies actually use. He has used the picture at Davos with Larry Fink, in a signed essay on Nvidia's own website, on CBS, at the Korea Society, and as the spine of a long interview with Ezra Klein. It is a good picture. It is also a picture drawn by the company that sells the second layer, and so it is worth taking seriously in two ways at once: as a real map of how AI gets made, and as a map that tells you where its author wants you to look. This essay walks up the cake one layer at a time — what each layer is, how it works, why it carries weight of its own — and then turns to the purchase that makes the most sense when you read it against the cake: Nvidia's $12.9 billion acquisition of Hugging Face, a company that does not sit in Nvidia's layer at all, but at the hinge where the fourth layer turns into the fifth.

1. Where the Cake Came From

The phrase is new, but the thinking behind it is not. In his conversation with Y Combinator, Huang placed the origin about fifteen years back, when deep learning first worked: "this is a way of doing software, with implications for the processor, the middleware, the algorithms, the applications — what I now describe as the five-layer cake. I imagined reinventing that entire industrial stack all together about 15 years ago." What changed in 2026 is that he started saying it out loud, everywhere.

The first big public outing was the World Economic Forum in January 2026, on stage with BlackRock's Larry Fink, where Huang described AI as "a five-layer cake" spanning energy, chips and computing infrastructure, cloud data centers, models and applications, and said of the top layer, "This layer on top, ultimately, is where economic benefit will happen" (NVIDIA Blog, Davos). Seven weeks later, on March 10, he put his name to a short essay on Nvidia's site titled simply "AI Is a 5-Layer Cake" (NVIDIA Blog). That essay is the canonical version, and its central sentence is the one to keep in mind for everything that follows: "Every successful application pulls on every layer beneath it, all the way down to the power plant that keeps it alive."

By the autumn the cake had become his default answer to almost any question. Asked by CBS why he disagreed with restricting chip sales to China, he answered with the cake: "remember that AI is a five-layer cake. It is not just the model." In that same answer he added a wrinkle he does not usually mention — "you could argue there's six layers, or five" — and named the extra one: "the data layer." Asked by Ezra Klein how Nvidia decides where to invest its money, he answered with the cake again: "if you look at my mental model of the A.I. industry, it's a five-layer cake, and we're investing across all of it." When a framework is used to explain export policy, capital allocation and job creation in the same month, it is no longer a slide. It is a worldview.

2. Layer One — Energy

What it is. Electricity: the power plants, transmission lines, substations and on-site generation that feed data centers. Huang's March essay states the point without decoration: "Intelligence generated in real time requires power generated in real time." Every answer a model produces "is the result of electrons moving, heat being managed and energy being converted into computation."

How it works. Older computing could hide its energy bill, because most of what a data center did was store and retrieve files. Generative AI cannot. Each token is produced fresh, and producing it costs a measurable amount of power, plus the power needed to remove the heat. Scale that across hundreds of millions of users and, in Huang's framing, hundreds of billions of software agents, and the power plant stops being background and becomes the binding constraint. The International Energy Agency's base case has global data-center electricity use more than doubling from about 415 terawatt-hours in 2024 to around 945 TWh by 2030, with the United States accounting for the largest share of current consumption and AI the most important driver of growth (IEA, Energy and AI; Carbon Brief). Carbon Brief's reading of the same work notes that grid-connection queues in developed countries were already long before data centers joined them.

Why it matters on its own. Energy is the only layer where money alone cannot buy speed. A chip order can be expedited; a transmission line takes years of permits. That is why Huang, who is otherwise relentlessly optimistic about American capability, concedes this layer to China. "One of the advantages China has right now in A.I. is energy," Klein put to him, and Huang agreed: "they just have a lot more energy than we do, and they plan to build a lot more than we did." His diagnosis of the American side is blunt — "we got ourselves really gummed up in climate change and sustainable energy, and as a result, we just didn't plan enough energy production" — and so is his admission of the political cost: the industry "moved so fast" that it failed to prepare the towns where data centers were going up, and "now there's a fair amount of frustration around the country."

One can argue with his energy politics and still accept his structural point. The bottom of the cake sets the ceiling for everything above it. If the power is not there, the best chip in the world sits dark, the factory does not run, the model is not served, and the application never reaches the hospital. It is the only layer that can stop all the others by itself.

3. Layer Two — Chips

What it is. The processors that turn electricity into computation: GPUs and other accelerators, plus the high-bandwidth memory and the interconnects that tie them together. This is Nvidia's home layer, and Huang's essay describes it in the terms Nvidia has always used: AI demands "enormous parallelism, high-bandwidth memory and fast interconnects," and the job of the chip is to "transform energy into computation efficiently at massive scale."

How it works. The key metric here is not speed in the abstract but intelligence per watt. If energy is the hard constraint, then the chip that produces the most useful tokens from the same megawatt wins, because it lets the same power plant support more customers. Huang's other argument about this layer is less obvious and more important to Nvidia's business. In the Klein interview he described Nvidia hardware as "fungible because we're general purpose": the same machines serve "data processing to pretraining to post-training to eval to inference," so "if a customer no longer needs it, another customer will be more than happy to pick it up." And because Nvidia keeps shipping new software for old hardware, he argued, "the useful life of our compute is much longer." His analogy was the airliner, which starts life carrying passengers and ends it carrying cargo, and which banks are therefore happy to finance as an asset class.

Why it matters on its own. The chip layer is where the cake's profits currently concentrate. Nvidia's most recent quarter, as reported by the New York Times when the Hugging Face deal was announced, brought revenue of $96.22 billion and nearly $60 billion in profit, roughly ten times its quarterly profit three years earlier (New York Times). It is also the layer where national policy bites hardest: export controls are written in terms of chips, not models or apps. And it is the layer with the longest lead times after energy, because every leading chip depends on a short chain of advanced fabrication and packaging plants.

But there is a deeper reason this layer is important on its own terms, and it is the reason the rest of this essay keeps returning to it. A chip has no value until something runs on it. Huang said this himself, in a sentence that is easy to skip: "We can't really create demand because in the end, if the A.I. services have no off take, then obviously, building computers for it is pointless." The second layer is the cake's profit center and its most exposed position at the same time. It earns everything the layers above it are willing to pay for, and nothing more.

4. Layer Three — Infrastructure, or the AI Factory

What it is. Everything required to turn a pile of chips into a working computer of continental size: "land, power delivery, cooling, construction, networking and the systems that orchestrate tens of thousands of processors," in the words of Huang's essay. It is also the cloud services that rent that capacity out.

How it works. Huang insists on a change of vocabulary here, and the change carries his argument. The old building was a data center, "a center that holds data," as he told the Korea Society; the new one is an AI factory, because "we're now producing intelligence." A factory is judged by output, not by cost alone. That is why he frames this layer in terms of yield: "It's $50 billion to build a one-gigawatt data center, one-gigawatt A.I. factory, and you can rent it for $40 to $50 billion per year." That rental figure is his, and it is generous; it is best read as a statement of how the builders of this layer see their own economics, not as an audited return.

Why it matters on its own. This is the layer where the cake turns into steel, concrete and jobs, and it is the one Huang leans on when he wants to talk to people who will never write a line of code. At Davos and again in his March essay he listed the workers it needs — "electricians, plumbers, pipefitters, steelworkers, network technicians, installers and operators" — and called them "skilled, well-paid jobs" in short supply. It is also where the sheer scale of the buildout is visible. The five largest American hyperscalers planned between $660 billion and $690 billion in capital spending for 2026 alone, most of it on AI compute, data centers and networking (Futurum Group). Huang's own estimate in the March essay is that the world is "a few hundred billion dollars into it" and that "trillions of dollars of infrastructure still need to be built."

The infrastructure layer is also where financial risk accumulates. A factory built for demand that arrives two years late is a very expensive building. Huang does not deny the cycle — "at some point, we will likely have more supply than demand. I just don't know when that is" — but he expects only "a period of digestion" of perhaps "six months... nine months... a year." Whether he is right depends entirely on what happens two floors up.

5. Layer Four — Models

What it is. The trained systems themselves. The public thinks of this layer as four or five chatbots. Huang works hard to widen that picture. "There are language models, but there are models of all kinds," he told Klein: "chemical models, biology models, physics models, articulation models, robotics, navigation models, self-driving cars." His March essay names the domains as "language, biology, chemistry, physics, finance, medicine and the physical world itself."

How it works. The model layer is where the cake splits into two economies. Closed models — he lists OpenAI, Anthropic, Grok and Gemini — are sold as services, and he has no quarrel with them: "the people working on them are incredible... they're at what we call the frontier." Open-weight models are downloaded, modified and run by whoever has them. Huang's argument for open models is not ideological; it is operational. A company that wants AI inside its own business needs "open weights so that I can fine-tune them, put them into my data flywheel, make them better and better every day with my intelligence and my domain expertise, and then I need to have control over it because I have a company to run and I can't rely on somebody else's service." He claims the market is moving his way, saying that at the beginning of the year about 70 percent of tokens came from closed models and that the ratio has since run "about 70-30 the other way." That is his figure, not an independent measurement, and it should be read as such; but the direction he describes is consistent with what he says about American start-ups, "80 percent" of which, in his account, now build on Chinese open models.

Why it matters on its own. The model layer is the cake's translator. Everything below it is general-purpose capacity; everything above it is specific use. A model is what turns a megawatt of computation into a radiology reading or a contract review. It is also the layer where the political fight is now fiercest. Washington has debated limits on open models, especially those from China, and Nvidia's own securities filing for the Hugging Face deal warned that "other parties are actively lobbying the U.S. government and other stakeholders worldwide to adopt legislative or regulatory measures that would restrict or disadvantage open-source models" (New York Times). For Nvidia this is not an abstract debate. Huang's March essay gives the reason in a single example: when DeepSeek-R1 made a strong reasoning model freely available, "it accelerated adoption at the application layer." Open models are the cheapest way to multiply the number of applications, and applications are what pull on chips.

6. Layer Five — Applications

What it is. The software that a doctor, a lawyer, a logistics manager or a factory engineer actually touches: drug-discovery platforms, robots, legal copilots, autonomous vehicles, and the AI built into everyday business systems. Huang calls this "the most important layer, and the layer that I care most about and the one that our country takes advantage of."

How it works. The application layer is where AI stops being a technology and becomes a change in how a job is done. Huang's favorite illustration is radiology. Computer vision has been "superhuman" at reading scans for about a decade, and yet radiologists were not replaced, because "for everybody's job, there's the purpose of the job and then there's the task you do as the job." The scan-reading task was automated; the purpose — diagnosing disease and helping patients — was not. In his March essay the same example ends with the line: "radiologists can focus on judgment, communication and care." Whether that pattern holds across other professions is the biggest open question in the whole debate, and it is not one this essay can settle. But it explains why Huang locates the value at the top: an application is the point where AI meets a paying problem.

Why it matters on its own. Three reasons, in ascending order of importance for Nvidia.

First, it is the only layer most people will ever see. "The consumers of the technology are going to enjoy it at the highest level," Huang said; they "don't have to deal with calculus and physics and quantum physics." Public trust in AI will be won or lost here, not in a fab.

Second, it is the layer through which a whole economy benefits, rather than a few companies. "Walmart has to benefit, Safeway has to benefit, Federal Express has to benefit," he told Klein; on CBS, "I love the fact that FedEx is going to be an AI company. I love the fact that Walmart's going to be an AI company." In his view, that is the real race: "That's the layer that touches society — all the layers underneath are technology enablers."

Third, and this is the one his business depends on, it is the source of all demand below it. Recall his own sentence: if the services have "no off take," building computers "is pointless." Recall the March essay: every application "pulls on every layer beneath it." Read together, these two lines explain Nvidia's investment strategy better than any press release. The company sits at layer two, but the oxygen for layer two is produced at layer five. Nvidia is therefore in the unusual position of needing thousands of companies it does not own, in industries it does not understand, to succeed at things it cannot build for them. That is why, as Huang told Klein, "the amount of investments that we're putting into the application layer so that each one of the industries could have the technology diffuse into them... that's probably one of the biggest things that we do."

7. Reading the Cake Critically

A map is also an argument, and this one has three features worth noticing.

It puts the mapmaker in the middle. In the cake, every layer above chips depends on chips, and every layer below chips exists to feed them. That is true as engineering. It is also a flattering picture for the company that holds roughly the middle slice. Klein put it gently — "You've become like a single-company industrial policy for American A.I." — and Huang did not object. By his own estimate, Nvidia's investment across the cake is "all in... $100 billion," "larger than the CHIPS and Science Act."

It separates where value is created from where profit is captured. Huang calls the top layer the most important, and in social terms he is probably right. But in 2026 the money is pooling in the second and third layers. That gap is not a contradiction in his framework; it is the problem his framework is designed to solve. If the application layer stays thin — if AI remains a handful of chatbots rather than ten thousand industrial tools — the profits in the layers below eventually become a bubble. Huang's strategy, read through the cake, is an attempt to thicken the top before the bottom gets ahead of it.

It hides the circularity question inside a word. When Nvidia invests in AI labs, new cloud providers and data centers that then buy Nvidia chips, critics see circular financing, and the Times says so plainly in its report on the Hugging Face deal, describing a "growing roster of financing deals" that "is raising concerns that the A.I. boom is being fueled through increasingly risky bets, with Nvidia at the center of many of them." Huang's reply, in the cake's language, is that he is not creating demand but "supporting" it at every layer, and that some investments simply "open new markets," open "a new route to market for us," or "secure a critical resource for us." Both readings can be true at once. The test is whether the applications show up.

That test is the right frame for the Hugging Face deal.

8. Hugging Face: Buying the Hinge Between the Fourth and Fifth Layers

What Nvidia bought. On September 3, 2026, Huang announced that Nvidia had agreed to acquire Hugging Face for about $12.93 billion (NVIDIA Blog). By the companies' own count, the platform serves more than 18 million developers and hosts more than 3 million models, 500,000 datasets and 1 million applications, and more than 200,000 companies use it to deploy AI (NVIDIA Blog; TechCrunch). In 2021 it hosted 13,590 models (New York Times). TechCrunch reports that Hugging Face had turned down a $500 million offer from Nvidia only a year earlier, and that the company runs at about $150 million in annualized revenue. Huang told Klein that the initiative came from the other side: Clément Delangue "came to the conclusion they need a lot more scale... and we'd really like Nvidia to be our home."

Which layer is it? The honest answer is: not quite the fifth. Hugging Face is mainly a model-layer company — a library, a distribution point and a collaboration space for open weights. It is fair to describe the deal as Nvidia strengthening the top of the cake, but only if one is precise about how. Hugging Face is not an application. It is the place where models become applications. Its million hosted apps, its deployment services, and the 200,000 companies that use it to put models into production make it the busiest on-ramp from layer four to layer five in the open ecosystem. And with half a million datasets, it is also the closest thing the industry has to the "data layer" Huang mentioned on CBS as the possible sixth slice of his own cake.

That is a more interesting purchase than "Nvidia buys an app layer." Nvidia already makes chips, already builds factory designs, and already publishes models — it describes itself in the announcement as the largest contributor of open models and data to Hugging Face, with more than 500 models and 250 datasets posted there. What it lacked was the junction where the world's developers pick a model, adapt it with their own data, and ship it. That junction is exactly the step Huang described when he explained why companies want open weights: download, fine-tune, put into "my data flywheel," and keep "control over it." Hugging Face is where that sentence happens, millions of times a month.

Why it helps Nvidia in the long run. Read against the cake, the deal does five things for Nvidia, and only one of them is about revenue.

It widens the funnel at the top. If every successful application pulls on every layer beneath it, then the number of applications matters more to Nvidia than the size of any single one. A few frontier labs building giant products is a narrow funnel. Hundreds of thousands of companies fine-tuning open models for narrow jobs — a hospital's triage tool, a freight company's routing agent, a bank's compliance checker — is a wide one. Hugging Face is the single most important piece of plumbing for that wide funnel, and Nvidia now owns it. Huang's own blog post put the mission in the cake's terms without using the word: open models let "start-ups, businesses, universities and public institutions build on advanced capabilities without training every model from scratch," and so carry AI to "factories, hospitals, farms, classrooms and Main Street businesses" (New York Times).

It diversifies who Nvidia depends on. A chip company whose demand comes from a handful of model labs is exposed to those labs' choices — including their choice to design their own chips. A chip company whose demand comes from a long tail of enterprises, governments and start-ups running open models on many clouds is much harder to disintermediate. The application layer, for Nvidia, is not just a market. It is insurance.

It gives Nvidia early sight of demand. Jeff Pollard of Forrester told the Times that "Hugging Face is where a large portion of the A.I. ecosystem exchanges models, data sets and code," and that Nvidia "now sits closer to the entire A.I. supply chain." In cake terms: Nvidia can see which models and which kinds of applications are gaining traction at layers four and five months before that traction turns into chip orders at layer two. For a company that, in Huang's words, has to "live in the future 5 to 10 years" because a system takes years to design and ramp, that early signal is worth a great deal.

It puts weight behind the open-model side of a policy fight. Nvidia's filing acknowledged that restrictions on open models "could negatively impact" the acquisition. That cuts both ways. It is a risk, but it also means Nvidia has turned its lobbying position into a balance-sheet position, which makes the argument for open models — the argument that most directly fattens the top of the cake — harder for Washington to dismiss as one start-up's special pleading.

It makes the "fungibility" story true one layer up. Huang's case for Nvidia hardware as an asset class rests on general-purpose machines finding a new customer when an old one leaves. Hugging Face does the same thing for models: a model released by one lab is picked up, adapted and redeployed by thousands of others. The more liquid the model layer is, the more steady the demand for general-purpose compute beneath it.

What could go wrong. The cake also shows the risks clearly.

The first is neutrality. Hugging Face's value depends on every chipmaker and every cloud treating it as common ground. Huang anticipated this in the announcement — "NVIDIA compute will not be required to build on or deploy through Hugging Face," and the platform will "continue to support multi-cloud and multi-accelerator development and deployment." The promise is sincere and probably kept in the near term. The long-run risk is subtler: features that arrive first or work best on Nvidia hardware, and a slow drift of developers onto the best-supported path. If rivals and developers come to suspect the hinge is tilted, they will build another hinge, and the application funnel Nvidia paid to widen will split.

The second is regulatory. Nvidia already dominates layer two; owning the main crossing point between layers four and five invites antitrust attention in the United States and Europe, where its earlier attempt to buy Arm was blocked.

The third is the open-model policy fight itself. Nvidia has bet $12.9 billion — including roughly $1 billion in retention incentives for Hugging Face staff, per its filing — on an outcome in Washington that it does not control.

The fourth goes back to section 7. If the top of the cake does not thicken fast enough, owning the on-ramp will not help. A wide funnel only matters if things are flowing through it.

9. Closing: A Cake Is Eaten From the Top

The image Huang chose is better than it first looks. A cake is built from the bottom up — you cannot frost a layer that is not there — but it is eaten from the top down. Energy, chips and factories are where the effort and, today, the profits sit. But nobody pays for a power plant because they love power plants. They pay because, at the top, a radiologist reads scans faster, a freight company routes trucks better, or a small business answers its customers at night. Every dollar spent in the lower layers is a bet that enough of that will happen soon enough.

Each layer earns its importance differently. Energy is important because it can stop everything. Chips are important because they set the cost of intelligence and currently capture most of its profit. Infrastructure is important because it is where the bet becomes concrete — literally — and where the jobs and the financial risk both pile up. Models are important because they translate general capacity into specific skill, and because the open-versus-closed question there decides how many people get to build on top. Applications are important because they are the only layer that pays for all the others.

Seen that way, Nvidia's purchase of Hugging Face is less a move into a new business than a move to protect its old one. A company at layer two has bought the busiest crossing between layers four and five, because that is where the demand for layer two is ultimately born. If Huang keeps his promise to leave the crossing open to everyone, the deal widens the top of the cake for the whole industry and Nvidia profits as the largest supplier underneath it. If the crossing tilts, he will have paid $12.9 billion to narrow the very funnel he needed to widen. The cake metaphor makes the stakes plain: whoever owns the middle of the cake still lives on what people eat at the top.


References

  1. NVIDIA Blog — Jensen Huang, "AI Is a 5-Layer Cake" (March 10, 2026)
  2. NVIDIA Blog — "'Largest Infrastructure Buildout in Human History': Jensen Huang on AI's 'Five-Layer Cake' at Davos" (January 2026)
  3. NVIDIA Blog — "NVIDIA to Acquire Hugging Face" (September 3, 2026)
  4. New York Times — "Nvidia Extends A.I. Spending Spree With $12.9 Billion Deal for Hugging Face" (September 3, 2026)
  5. TechCrunch — "Nvidia confirms it will buy Hugging Face for $12.9 billion"
  6. New York Times, The Ezra Klein Show — "Jensen Huang Thinks A.I. Alarmism Has Gone Too Far" (September 23, 2026)
  7. CBS News interview with Jensen Huang (Jo Ling Kent), 2026 — "The Five-Layer Cake" segment
  8. Korea Society fireside chat with Jensen Huang (Juju Chang), 2026
  9. Y Combinator conversation with Jensen Huang, 2026
  10. International Energy Agency — Energy and AI, Executive Summary
  11. Carbon Brief — "AI: Five charts that put data-centre energy use and emissions into context"
  12. Futurum Group — "AI Capex 2026: The $690B Infrastructure Sprint" (February 12, 2026)

No comments: