From Frontier to Foundation: How Open Models Are Changing the AI Economics From Our Portfolio Event at NVIDIA

From Frontier to Foundation: How Open Models Are Changing the AI Economics From Our Portfolio Event at NVIDIA

Written by

Anne Gherini

Published on

July 27, 2026

SIERRA Ventures Portfolio & EPL Session at NVIDIA HQ

Sierra Ventures recently brought together NVIDIA leaders, portfolio founders, and members of our Engineering & Product Leaders Council at NVIDIA headquarters to discuss how open models are moving into production.

The session was part of Sierra’s broader work with leading AI labs, giving our founders direct access to the technical teams, emerging technologies, and market insights shaping the next generation of enterprise AI.

The conversation centered on a practical question: when do open models make economic sense for companies building real products?

The answer is changing quickly. Open models are becoming competitive on production workloads, easier to deploy, and increasingly attractive at scale.

 

Open models have crossed the production threshold

Models such as NVIDIA Nemotron, Qwen, and Llama are now competitive with frontier models across many reasoning and agentic workloads.

Sierra Ventures and NVIDIA have both invested in this shift through decacorn Reflection AI, which is building frontier-scale open-weight models in the U.S. and using reinforcement learning to advance them from reasoning toward agency.

The more meaningful proof is coming from founders testing models on their own products. In side-by-side comparisons, output quality is often similar. The real differences show up in latency, token consumption, control, and cost per completed task.

This does not mean open models will replace closed models. Most companies will use both. But open models are now strong enough that they should be actively evaluated rather than dismissed as an infrastructure-heavy alternative.

 

The real metric is cost per completed task

Falling token prices do not automatically create sustainable unit economics.

Agentic applications can involve multiple model calls, retrieval steps, planning loops, and tool interactions. A model may be inexpensive while the full workflow remains costly.

The better metric is total cost per completed task, including inference, infrastructure, latency, and reliability.

This is where open models can create an advantage. Companies can use smaller models for simpler workloads, deploy models inside their own environments, and optimize infrastructure around their actual usage patterns.

FriendliAI removes the infrastructure tax

IMG_6522-1

The biggest obstacle to adopting open models has never been access to the weights. It has been the operational burden.

Companies need to deploy GPUs, optimize inference, manage capacity, maintain reliability, and keep pace as new models are released. That can require significant engineering investment.

FriendliAI, a Sierra portfolio company that presented during the session, is built to remove that friction.

Its platform gives companies the speed, reliability, and ease of use they expect from closed-model APIs while preserving the cost, control, and customization advantages of open models.

The founding team helped pioneer modern inference techniques such as continuous batching and has built custom GPU kernels, prefix caching, and speculative decoding into the platform.

In reported workloads, FriendliAI has delivered up to three times the speed of vLLM and cost savings of 50 to 90 percent compared with closed-model APIs.

The practical impact is significant. Companies can adopt open models without building and maintaining the full inference stack themselves. New models can be supported quickly, making it easier for teams to test, compare, and deploy them as the market evolves.

That closes much of the practical gap between calling a closed API and running an open model.

Open models give founders back control

IMG_6536-1

The decision is not open versus closed. It is determining which model is best for each task.

A complex reasoning workflow may still require a frontier model. Classification, extraction, or domain-specific work may be better suited to a smaller open model. Sensitive or regulated workloads may require deployment inside a company’s own environment.

Open models also give teams more freedom to fine-tune and post-train around proprietary data and workflows. As base models become more interchangeable, differentiation increasingly comes from what companies build around them: domain expertise, data, tools, evaluations, and workflow integration.

They also reduce dependency on any single provider. Pricing, performance, and availability will continue to change. Companies that can route workloads across multiple models will be better positioned to adapt.

The founder playbook is changing

IMG_6531-1

 

Founders should start by testing open models against their current stack on real workloads.

Measure quality, latency, reliability, and cost per completed task. Do not rely only on benchmark rankings or published token prices.

Design for a multi-model future, even if the product initially uses one provider. Understand which workloads could benefit from greater control or customization. Before building an internal inference platform, evaluate whether a provider like FriendliAI can deliver the economics without the operational burden.

A year ago, open models were primarily an infrastructure decision. Today, they are becoming a product and business-model decision.

The strongest teams will not commit to one model by default. They will build systems that can continuously take advantage of changes in performance, control, and cost.