At Sierra Ventures' 21st Annual CXO Summit, Sachin Katti, VP of Compute at OpenAI, joined Andrew Feldman, CEO and Founder of Cerebras, and Dylan Patel, CEO and Founder of SemiAnalysis, for a panel on the race to build AI infrastructure.
Sachin's role puts him at the center of that race. His perspective was clear: the models can already do far more than most people use them for, and the hard work now is building the compute, products and systems that bring that capability to everyone.
Sachin described a pattern that keeps repeating. Each time his team thought it had secured enough compute, demand caught up again. He admitted he never expected to see a world where AI providers turn customers away, but that is where the industry is today.
His prescription was simple: build faster. The obstacles are not only technical. They include a long list of real-world problems that come with putting up physical infrastructure. Even with data centers in place, he expects memory to be especially tight over the next several quarters. Smarter choices about the size of compute clusters and how they are networked can help with inference, he noted, but they don't solve the training problem.
For a frontier lab, a delayed data center costs far more than anything saved on construction choices. That changes how OpenAI approaches the build. Sachin explained that his team has to stay involved day to day on site and take ownership and accountability for getting compute online, rather than leaving the schedule entirely to contractors.
On community pushback against data centers, Sachin's view was that all politics is local. Opposition isn't the same everywhere in the U.S. Many rural communities see real benefits from data centers, including property tax revenue and new jobs.
The answer, in his view, is to engage locally: local newspapers, local events and showing up in the community. The narrative can't be changed from Silicon Valley. He pointed to efforts as grassroots as buying livestock at rural fairs as the kind of engagement that matters.
Sachin shared how AI use is changing inside OpenAI. Research has sped up because so much of the work is now done with AI coding tools and the models themselves. But the steepest rise in usage has come from non-technical functions such as HR and operations. He expects the same adoption curve to show up across the broader market as compute constraints ease.
His larger point was that model capabilities far outpace what people know how to use them for. The gap is a product gap and a compute gap, not a limit of the models. As with any technology, product investment has to catch up so everyone can benefit. He described the long-term ambition as giving every person access to dedicated compute all the time, which would mean far more compute than the industry plans to build this decade.
He also sees user expectations shifting. People increasingly want long-running agents that keep building memory and context over time, instead of one-off exchanges. That puts new pressure on memory and on how context is compacted and reused. Sachin's own definition of AGI is a system that keeps learning from its past interactions with you and keeps getting better at working with you.
As AI reaches very large user bases, Sachin sees the infrastructure question becoming one of fit. Model sizes, inference costs and how much "thinking" a model does per task will all vary by use case. Today many users default to maximum effort for everything, and that isn't always necessary.
Some work can run on devices, he noted, but devices vary widely in capability, and network latency adds another variable. At scale, every one of these choices becomes a technical decision shaped by user experience and expectations: what level of intelligence a task needs, how much thinking time to spend, where to run tool calls and what needs to persist. He emphasized that the industry is still very early here.
He was firm that users and enterprises shouldn't have to sort through all these options themselves. Today, people orchestrate their own sub-agents. In Sachin's view, models will inevitably take over that role, deciding which sub-agents to launch, which models to use and how much compute each piece of work deserves. He cautioned that deciding the right amount of intelligence for any given task may be almost as hard as reaching AGI itself.
When asked about smaller, cheaper models that match frontier performance, Sachin wasn't surprised that a specialized model can do as well as a general frontier model on a narrow task at lower cost. But he argued that any single model is just a snapshot. The more important measure is how fast the frontier itself is moving. If it slowed, specialized models would become more important. As long as it moves fast, anything specialized will be playing catch-up.
He sees no sign of a slowdown. Frontier labs have gone from months or quarters between major model releases to roughly every couple of months, and he expects that pace could speed up as compute scales and AI increasingly helps drive research inside the labs. In his view, scaling is continuing on several fronts at once. Pretraining still produces unexpected emergent behavior, and post-training methods such as reinforcement learning and synthetic data keep scaling too. Training runs are already at gigawatt scale and heading to multiple gigawatts.
On running AI locally, Sachin argued that the edge's real role isn't hosting intelligence. It's computer use: agents operating the tools already on your phone or laptop. Pushing intelligence out to devices, he noted, is worse for total energy consumption than running it centrally.
Sachin closed by looking beyond software. All engineering, he observed, comes down to design and optimization. AI is already showing it can run many experiments in parallel and reason toward the best solution, as it has in software engineering. He expects the same to come to physical engineering wherever a reliable simulation environment exists. That opens the door to products that are no longer one-size-fits-all. The limit, in his view, isn't AI's capability but our ability to experiment reliably in the physical world, and the infrastructure built to support that.