返回知识库
0

title: "顶尖创业者为何涌向 AI 基础设施"
source_url: "https://www.youtube.com/watch?v=Zx1Ec8LWFeM"
author: "a16z"
excerpt: "公共早报 a16z 合伙人认为,推理与智能体工作负载不断攀升,已让物理算力、电力、散热和供应链设计成为 AI 经济中的核心约束与创业机会。"


Brief Description

Ben Horowitz, Martin Casado, and Ragu Yarlagadda join the host to introduce a16z's Machine Age Fund, discussing the massive, unprecedented demand for AI infrastructure and why the industry requires a ground-up redesign of its hardware ecosystem. They explore how the shift from scaling software engineering to scaling physical compute has created supply bottlenecks across GPUs, memory, and power generation, leading to multi-billion-dollar opportunities for new founders. The conversation also covers the evolution of AI agents as autonomous employees, the necessity of building sustainable, gigawatt-scale data centers, and why maintaining American leadership in this physical layer of AI is critical for the next technological era.

Table of Contents

  • Introduction to the Machine Age Fund

  • The Macro Conditions and Unprecedented Demand

  • Supply Constraints and the Cost of Compute

  • The Unending Demand for Tokens

  • AI Agents as Autonomous Employees

  • Redesigning Systems for AI Workloads

  • Power Grids and the Gigawatt Era

  • The Machine Age and New Founder Opportunities

  • A Look Ahead: Why America Needs to Win

Introduction to the Machine Age Fund

Host: Ben, Martin, Ragu, welcome. I want to start with a Marc Andreessen quote to introduce this new fund: "This is the biggest technological revolution of my lifetime. This is clearly bigger than the internet. The comps on this are the microprocessor, the steam engine, and electricity, or maybe the wheel." Guys, the Machine Age Fund. Please introduce it. Ben, start us off.

Ben: Basically, we have a whole new technology that is the most important technology ever. What happens every time there is a dramatic new way of using all of the things that we love is you need a whole new infrastructure. Never has it been more high-impact as it is on this one.

Not only do we need new chips and new system software, we need new ways of doing power. We need to replace copper. It is absolutely everything, so it is a very exciting time. Particularly for the hardware aspects of this new era, we needed a new approach.

Ragu: I would agree. Normally, at least in computing, when we talk about the infrastructure world, we are talking about the servers, the storage, and the network. Here, it goes all the way down to the copper mines. That is how widespread this thing is going to be.

Number two, I think what we have seen over the last three years is the steady increase in the capabilities of the models, where the model is no longer the bottleneck. In fact, using AI, these models are getting better and faster. Now the bottleneck is all what I call "south of the model," and so that is why we need to work on that.

Martin: The only thing I would add very quickly is that we tend to follow founders. What we've been watching over the last couple of years is the number of very strong teams going after complex hardware problems has increased. I was trying to estimate it over the weekend, and previously I think maybe five percent of the deals from top founders would be hardware.

Now I would say it is north of 20 or 30 percent. The founder community, which tends to be much smarter than the VC community, has identified this as a very active area for innovation, and they are responding.

The Macro Conditions and Unprecedented Demand

Host: Explain some of the macro conditions that have led to this change. What are they seeing that enables this surplus of founders pursuing these ideas?

Martin: The obvious is that the demand for AI is basically infinite. As a result, every part of the supply chain is under duress, including materials used to make things like memory. It is also very interesting because the demand is infinite, what you tend to worry about is the margin of companies, which is how efficient it is.

Normally you worry about growth, like can I just get people to buy this stuff? You don't have to worry about that here. The question is, can you do this in a way that's profitable? A lot of the efficiencies are actually strictly a physical limitation of hardware.

Even the business model of the AI wave is putting a lot of stress on the existing systems because they weren't built for AI workloads. There is this global observation that we actually need to change the core components to get that efficiency, to help drive the growth, and to drive the value of the businesses.

Host: How did we know that demand is actually outpacing supply here, rather than this being another hype cycle?

Ragu: Firstly, some of the smartest judges of demand are cutting huge purchase orders. If you look at the hyperscalers, their capex spend has been exploding. Next year, supposedly, it is going to reach a trillion dollars collectively across the big hyperscalers. This year, it is about 700 billion dollars.

If you think about the hyperscalers' position in the industry, they see demand from everywhere. They see the frontier labs wanting their compute, they see AI-native companies, they see the enterprise, and they see international geographies. They've been jacking up their capex like it's never been seen before, so that is a clear sign.

Secondly, if you look at the application companies we see on a day-to-day basis, they are all ripping. The growth is insane, and it's been documented. So on the demand side, the signals have never been clearer that this is not a hype. To top it all off, prices are going up like we've never seen on chips.

Ben: They always go down. If you look at the price curve, it went down and then went back up. We know only five to ten percent of the addressable market is tapped today.

Martin: If you look at the supply across the board, it is basically all booked out to 2028. It is so bad, we've actually seen multi-day auctions for a few thousand GPUs. The other side of that, of course, is demand.

As Ragu said, we've seen the fastest-growing companies in the history of the industry. But also, the unit of work that AI can do, and the value of that unit of work, keeps increasing. Underneath the covers, the number of tokens that are consumed is growing by orders of magnitude.

If it was a hundred tokens for a chat, for an agent it is thousands of tokens. You have expansion on both sides of demand. One is the unit of work becoming more consumptive of tokens. Secondly, it is not just developers; it is going to be all knowledge workers and beyond.

Supply Constraints and the Cost of Compute

Host: You said the key components in supply are sold out to 2027, maybe 2028. What does it mean for an entire industry to be sold out that far? I don't know that this has ever happened before. Remember in the internet days, the massive buildout being put in the ground was speculative dark fiber. Here, basically every GPU being created is already pre-sold.

Ben: There was a lack of bandwidth in the 1998 to 1999 time frame, but there wasn't that much real demand for it because there weren't that many people on the internet. It was a two-sided thing. Companies were rushing there and needed more bandwidth theoretically, but there weren't users on the other side to consume it.

To really consume a lot of bandwidth, you had to do high-bandwidth things like video, which weren't really viable for reasons having nothing to do with data centers. So it smelled similar, but it wasn't. This is flat out. People are reselling GPUs for four times what they bought them for.

We are also out of power and cooling. On top of that, it is really hard to build because there are incredible political headwinds. It is unprecedented in my career that we've had anything like this.

Martin: I want to give you a quick anecdote. I was talking to a CFO of a large public company who had historically been very resistant about going to the cloud. They had a lot of servers and were doing an inventory check. They realized that the memory in their servers had increased in value so much it could fund the entire migration to the cloud. We are in a very unusual situation.

Ragu: That's right, we're out of many things: power, cooling, memory, GPUs. You name it, we're out of it. The flagship conference for the industry, Hot Chips, is going on at Stanford. The leading memory vendor said the demand they have today will take them three years of capacity to supply. It's just today; it's not even future demand.

Host: Is it because people just underestimated how good the models would be, or how useful they would be?

Martin: I don't even think it is that. This stuff came out of nowhere. We are only four years into this, so even if we had a perfect oracle once it started working, I don't think we could have built the capacity.

We are talking about chip cycles which tend to be three to four years. We are talking about breaking ground and building data centers, which takes four to five years. You either have to build your own power or secure a power source, which is not easy.

Ragu: The machine learning industry growing at 20 to 30 percent is a great growth rate. But it is being connected to an AI software industry that is growing at triple digits. You can see the disconnect right there, so it is just wide open.

Host: Why didn't this fund exist five or seven years ago? Why was it not a great category to invest in the same way?

Martin: I would say we are probably just in time, but we probably would have been well-suited to have it a couple of years ago. You could actually point to basically every epoch and see an independent company that came up. The move from mainframe to client-server saw a bunch of companies. The move to the internet gave us Cisco and Juniper.

Even in mega data centers, which were largely driven by incumbent cloud providers verticalizing, you saw the rise of Arista. There has been the ability to invest in silicon and hardware, but it's been relatively minor because the change was minor. Here, absolutely everything is changing.

Ben: The demand for intelligence is so vertical with really no end in sight. Every company that has adopted it is growing very fast in its usage, and most companies haven't adopted it to a high degree. Consumers are just getting started, so the demand for tokens is probably going to grow close to a thousand percent a year, and you cannot grow supply that fast.

The amount of work we are going to have to do across the board to grow infrastructure at that rate is vast, so there's a lot of investing opportunity on the way. Also, all the architectures of the hardware systems were built for a whole different era of computing. We need more than just capacity; we need opportunities to build different kinds of infrastructure.

Ragu: They are all reaching the physics limits for what they were designed for. You could go across every one of these categories and find the limit of this type of technology. Now you have to get technical breakthroughs to get to the next one.

The Unending Demand for Tokens

Host: I want to dive deeper on the demand side for a second. As we've moved from chatbots to reasoning to agents and multi-agents, each step has multiplied the number of tokens a single task takes up by increasing orders of magnitude. Why does that keep happening instead of leveling off? Do you see that happening indefinitely?

Martin: Right now, the way we are achieving scaling is through a lot of inference. If you think about reinforcement learning, it is a lot of inference. Chain of thought is a lot of inference. Long-running agents are a lot of inference. That has basically been one of the approaches we've been using to scale.

If you want to step back and look at the macro trend, it used to be that when you built something, it was an engineering problem. You would throw a bunch of engineers at it, but that doesn't scale. Here, it feels like it really is a resource limitation. We are pouring a ton of money into systems, and right now we are bottlenecked on the system's ability to match the resources we are pouring into them.

Tokens are probably where we are on the scaling curve, but we don't have a natural regulator like engineering did before. I think we should expect this to continue, and we have to build the supply to support it.

Ben: The simple way to think about it is any problem that you have can be solved with enough infrastructure, basically GPUs, power, and money. Until we run out of problems, we are not going to run out of demand. AI's answer to getting better and better is to use more AI. Inference is a basic building block that it keeps using over and over again.

Martin: Even the auto-catalytic effect, like using AI to create more AI or create a GPU kernel, is just using more AI as part of the process. In the past, money would come in for an engineering problem, and there was a natural governor before you got the product on the other end.

Here, there is nothing between the money going in and the hardware creating intelligence. We are just limited by our ability to create supply. As long as you have the money, the GPUs, and the data, you will be able to scale these things.

Host: It is fascinating because over the last decade the pervasive sentiment was that there's too much money going to startups. Say more about that, Ben, because there used to be skepticism about bigger outcomes, but now the market is as big as we collectively contribute to it.

Ben: The one thing we all knew in the startup world is that if I have a two-year lead on you, and you try to catch me by hiring a thousand engineers, you are going to wreck your company. Nine women can't have a baby in a month. Now, that works, but it's not hiring a hundred thousand engineers; it's taking three billion dollars and lighting up a magnificent cluster.

Then all of a sudden, something like Grok can come out of nowhere and be real. You can throw money at almost any problem and that works. That is completely different than anything we've ever lived through, and we are all psychologically adjusting to this.

Ragu: The ChatGPT app has a billion weekly actives. There are about 30 million developers who are using a big portion of compute demands. How do we think about compute needs now and in the future? ChatGPT was a casual app, and now coding is for professionals. We have now built amazing tools for knowledge workers, which is the next frontier.

There are over a billion knowledge workers in line, and with that, it is a long way to go for demand. Then you get to the back office, which is all the agents. Progressively, each of these things unlocks an order of magnitude more demand.

Ben: And now you have bots where what happened with coding is kind of happening with all computer use. We are in a whole another wave of demand, and there is going to be more to come. It seems quite unlimited at the moment, and we haven't even gotten into embodied AI or robots, which are going to be another source of demand.

AI Agents as Autonomous Employees

Martin: Part of this uses computer use, which is like a human being sitting inside the computer typing away. I used it over the weekend to update my credit card with a bunch of services I'd been lazy to do and to cancel subscriptions. This is not coding; this is true computer use. All of a sudden, you are creating half a billion knowledge workers sitting inside the computer doing what I do.

I think Marc Andreessen is right that the correct analog here is the steam engine or electricity. We've introduced this new thing that you can turn to work. There are some very obvious applications now, but there's probably 30 to 40 years of throwing computing at problems. We are looking at science, materials, biology, and creativity. We've removed traditional software engineering as a key bottleneck.

Host: Talk about these agents, Martin, because we were talking at the offsite about what struck you about them.

Martin: We've gone through multiple realizations as an industry for how AI enters our lives. Very early on, you added AI to a product like a search bar, and you would chat with it. Then earlier in the year, AI models for coding showed up, and it was like Google but better. Maybe it becomes an extension of you, sharing your keys and passwords.

What the new wave of agents got really right is treating it as an employee. Now you have an entity with its own computer and browser. Because these are the smartest models in the world, they can do whatever an employee can do. If I want something done, my first thought is whether the AI can do it for me. I'll have it read through my email and do triage, and it will know to check with me before actually acting.

Host: Ben, I know you think a lot about culture. What are your thoughts on how this works in an organization?

Ben: If you look at us, it is like having a new kind of employee, and there are going to be a lot of them. We spent many years figuring out how to work with regular human employees, and now we have these other kinds of employees. There is a learning curve. They can burn a lot of tokens and get nothing productive done. They can forget stuff, make stuff up, and create security problems.

But they can also be super duper productive. Figuring out how to integrate them and have them work nicely with actual humans is something we are learning how to do. I don't want to sit up here and say I've cracked the code and the whole firm is completely automated. We are much more focused on how we make all our humans superhuman without wrecking the place because the bots got out of control.

Host: It's interesting because we tried a couple of different ways to get agents into the system. Eventually, it was Martin's insight to just treat them as people and get it done, which turned out to be the most durable way.

Redesigning Systems for AI Workloads

Host: I want to go back to the supply side and go deeper into the bottlenecks. We were talking about how data centers, chip architectures, system software, and facilities were not designed with AI in mind. What is the mental model for thinking about what that could mean?

Ragu: The original model of infrastructure on any of these fronts has to change. You have to get to a system where you look at what an inference engine does. It takes up a lot of memory, generates new tokens, and requires compute. You have to think about how to optimize all of this: what does the memory need to be, how do they talk to each other, how much power do they need, and how do you cool them?

That is the exercise underway in the industry with a lot of founders right now. They are breaking down the problem into fundamental components. How do I optimize my compute around matrix multiplications? What is the best way to hierarchically arrange memory to generate these tokens? How does it consume power, and what are the ways of connecting it across chips and data centers?

Martin: Let me give you an interesting mental model. Today, to build a frontier model costs let's say three to five billion dollars to train. The inference has to pay back at least that, let's say ten billion. If you can save 20 percent on efficiency, that is two billion dollars. You can easily build an ASIC for two billion dollars.

We've reached an interesting point where it actually makes sense to build an ASIC per model because of the massive capital investment. Unlike traditional software which is very dynamic, these model weights are fixed. It gives you a great mental model of how you would evolve the architecture to be bespoke for these massive investments.

Host: Rack power requirements are moving from five to ten kilowatts up to a hundred or fifty kilowatts. Compute density is climbing 70x, and cooling is moving from air to liquid. What are the investment opportunities as a result?

Ben: When you get to that level of power per rack, AC power doesn't work anymore. Now you are into DC power, which requires its own cooling and is super dangerous. This is what Edison used to electrocute animals to demonstrate how dangerous AC power was, ironically.

With cooling, we are already at liquid cooling for any state-of-the-art data center. But given the political environment, it has to be eco-friendly liquid cooling. DC power must contribute to the power of society, not take away from it. Data centers that waste water or act as parasites on the power grid are going to have to end.

Racks are so dense that the way floors are designed has to support incredible weight. These things are really loud, so you have to build data centers with thicker walls so you don't disturb the neighborhood. A much smaller percentage of the data centers we have today will work. Prices for reinforced concrete are increasing fast.

Also, when sending 800 volts to the rack, it is so dangerous that we don't have enough electrical contractors certified on DC power in the US. Meta has a whole program to train people up for this. AI is taking all the jobs, but it is going to create a lot of new electricians.

Ragu: The big cloud data centers are all furiously experimenting with robots to do the work of assembling or putting servers into the data center. You will see that increasing as a result of AI evolution. To be clear on the fund we are raising, our focus is on computer science infrastructure. Think chips, network interconnects, storage, all the way down to electricity.

Martin: One of the great breakthroughs of AI is that it allows computers to interact with the physical world. It can see, hear, and talk, which means new platforms. We don't do heavy regulated industries, but any computer science platform that will push AI further out into embodied devices, we are quite interested in.

Power Grids and the Gigawatt Era

Host: Going back to data centers, by 2028 new data centers are going to need something like 44 gigawatts of additional power against maybe 25 gigawatts of expected grid additions.

Martin: Hold on. We use that word gigawatt. I mean, how big is it? It's massive. It's the equivalent of 50,000 homes. I grew up in Flagstaff, Arizona, a town of 40 to 60,000 people, and we had less than a gigawatt of power consumption. You can basically light up and air-condition your entire town for a gigawatt.

Everybody talks about the gigawatt, but there are very few gigawatt data centers actually up. We've got a long way to go.

Host: Then why can't utilities and hyperscalers just build faster?

Ben: Right now, you need humans to build them. Much more than that, you need permits, and you need access to power that you can plug into. Securing access to power is a massive regulatory bidding struggle. You also have to build your own power, and we have shortages of transformers and turbines.

This is not a software problem where engineers can just work weekends. There are real bottlenecks, and these lead times are not easy to compress. We have the best minds trying to figure it out, but it's not easy, and demand is growing 10x a year.

Right now, if we have new companies going for GPUs, it's often in Mexico, Australia, or another country just because it is so difficult in the United States. We are creating huge economic opportunities in other countries by practically banning data centers here.

The right answer would be to set a standard where a data center contributes back to the community, makes power better, causes no noise or water issues, and adds jobs. There are data centers that provide their own power, give power to the state during the day, and borrow it at night. It is a symbiotic relationship.

The Machine Age and New Founder Opportunities

Host: Zooming out, why do we think "Machine Age" is a compelling term for what we're doing here?

Martin: Artificial intelligence was the wrong word; we shouldn't have called it that. It is machine intelligence. It is not necessarily how humans think; it is a cache of human thoughts. We don't know how to take an AI with no knowledge and have it reconstruct language. We built something that can learn off everything we've already learned.

AI is a general term that goes back 70 years in computer science with a lot of science fiction baggage. "Machine Age" acknowledges that this really is machine intelligence. There is deep irony that the "software is eating the world" folks have come to a place where you pour money into something and are limited by the actual machines below it. We want to acknowledge that the hardware component is so significant.

Ragu: That's what is going to create the next breakthroughs in intelligence---the quality of the machines underneath.

Host: Given how much has been spent on AI infrastructure to date, are we past the point where new companies can break in? Why not let incumbents like Nvidia just take the lion's share of these markets?

Ragu: They are all doing well, but to keep the pace of improvement continuing on tokens per dollar or tokens per watt, you need fundamental new innovations. New innovation traditionally comes from brilliant founders thinking about solving the problem from first principles in a different way.

Martin: This is the law of markets. If the existing silicon incumbents are multi-trillion dollars in market cap, even five percent of that is a massive private company. Nvidia could do that, but why would they if they are focused on things that are driving 90 percent of the growth? Once you get to a certain scale, there is tremendous opportunity for innovation at the margins.

Ben: Our partner Alex Rampell was trying to sell his startup TrialPay to Facebook. Dan Rose, the head of corp dev, told him, "You can collect a lot of silver bricks, but I have so many gold bricks I can't even pick them all up. The last thing I'm doing is looking at a silver brick." I think Nvidia is in that position.

Martin: As markets expand, they fragment. Ford's Rouge River plant used to make everything vertically, but now the car industry has multiple levels of suppliers. When growth slows down, they tend to consolidate, but there are so many use cases now that it's impossible for the biggest company to optimize everything perfectly. Optimization in hardware is absolutely meaningful to the upside of the business in a way we haven't seen before.

Host: Let's get deeper into the types of companies we'll be investing in. Could you talk about the sub-sectors or some investments?

Ragu: The sub-sectors are every one of these categories: computer chips, memory innovation, networking, and power chips. You need to build a full system now. Around that, there is a layer of software to automate and manage these fleets. Each of these are categories where you can see public-company-style businesses emerging.

Host: Talk about what's different about these kinds of companies. The first rounds have been massive, hundreds of millions of dollars. Is it a different kind of founder?

Ben: A lot of money goes in before they get to a product, which comes with a little more risk. Many of the chip founders have been here from the past. The guys who know how to make memory are not young, which is different but exciting.

Ragu: These founders have all got to be systems founders. You can't just be a researcher or a great computer scientist; you have to architect the chip, figure out how it will get manufactured, and handle the downstream supply chain. Jensen Huang is the Michael Jordan of this. They think about the entire ecosystem from the get-go.

Martin: There are two environmental factors. First, the labs are so desperate that they will engage with startups early and ink deals before hardware is available. Second, capital availability has loosened up a lot for these areas.

Host: Patrick Collison remarked a few years ago that it feels like there are fewer younger founders today, compared to Zuck or Gates. Is that true here?

Ben: If you are building something with a complicated supply chain that has to manufacture things, experience helps. Elon Musk and Travis Kalanick started with software companies when they were young and graduated to more complicated domains. If you don't completely understand the product and have to learn it while you build the company, that is a steep learning curve.

Martin: It has been defocused by academia for the last 20 years. We just haven't had people coming out of universities who have done this. But that is changing now. Elon Musk has created these great companies, but the number of entrepreneurs that have come out of SpaceX and are changing the industrial complex is his greatest legacy.

Ragu: One of our investments was started by two founders in their 20s, but if you go to their office, you see experienced people working with them as well. It is an ideal combination.

Host: This is a big new fund we're launching, and there are no new GPs. We're using your collective experience.

Martin: Our backgrounds are in hardware, and we are drawn to that. We've been clearly investing in hardware over the years with early checks in SpaceX, Anduril, Astranis, and Waymo. This is in our DNA.

A Look Ahead: Why America Needs to Win

Host: If this fund does what we think it will do, how do we see the world changing in 5 to 10 years?

Ben: Hopefully, America wins in the infrastructure game. We have super eco-friendly data centers, an abundance of chips, memory, and power. America is a special place where anybody in the world can come with nothing and do something profound. We'd like to keep that going, and that doesn't continue if we lose our lead in technology. If we do, it'll be another era, and another country with a different set of values will take over.

Host: That's a wrap. Thank you.

AI知识库 / 顶尖创业者为何涌向 AI 基础设施 0 字 0 行 iliuqi
2026-09-04T08:55:45.261381649Z 2026-09-04T09:39:45.869006992Z