Epoch Finds Open-Weight AI Models Trail Closed Frontier by Four Months
Epoch AI's updated Capabilities Index analysis finds open-weight models trail frontier closed models by four months and 8 ECI points.
- Open-weight models lag frontier closed models by 4 months on average since January 2026, per Epoch AI
- The vertical gap is 8 ECI points, comparable to the difference between GPT-5 and GPT-5.5
- Previous Epoch analysis put the gap at 3 months across January 2023 to October 2025, a modest widening
- Requiring strict point-estimate dominance instead of a statistical tie pushes the gap to roughly 6 months
- Estimate likely understates reality: open models may overfit public benchmarks and closed labs withhold top systems
- Chinese labs now dominate open weights, mirroring a 7-month US-China capability gap
The gap between the best open-weight models and the frontier of closed proprietary systems has a new number attached to it. Epoch AI's latest data insight finds that since January 2026, the most capable open-weight models have lagged frontier closed models by an average of four months in the Epoch Capabilities Index (ECI), with an average ECI gap of 8 points, similar to the gap between GPT-5 and GPT-5.5.
That is a modest widening from their previous estimate. The October 2025 analysis found that open models lagged by an average of three months between January 2023 and October 2025. The headline shift from three to four months is small, but the framing matters: the frontier has been moving fast, and open-weight releases are roughly holding their relative position rather than collapsing behind.
What ECI is actually measuring
The Epoch Capabilities Index is the yardstick doing the work here, so it is worth unpacking. The ECI combines scores from many different AI benchmarks into a single general capability scale, allowing comparisons between models even over timespans long enough for single benchmarks to reach saturation. The general ECI uses scores from 40+ distinct benchmarks to generate a single, general capability scale, stitching them together by determining their relative difficulty wherever models are evaluated on multiple benchmarks. Models obtain higher ECI scores if they perform better on harder benchmarks.
The score itself is abstract, but the deltas are interpretable. The ECI scale is linear, so a 10 point jump should be equally impressive moving from 100 to 110 as it is going from 140 to 150, and at launch, a 5 point gain in ECI roughly corresponded to a doubling of the METR Time Horizon. By that yardstick, an 8 point lag is roughly the difference between agents that can handle tasks of one length and agents that can sustain coherence over tasks several times longer.
How the four-month figure is computed
Epoch's method is more careful than a back-of-envelope comparison. For each day between January 1, 2026, and May 28, 2026, they identify the open-weight model with the highest ECI score available by that date, then compare it to the historical state-of-the-art ECI frontier and ask what is the most recent date on which the SOTA model was not significantly better. The time gap for that day is the number of days elapsed since that date.
Because scores carry uncertainty, the comparison uses bootstrap resampling. ECI summarizes a model's capability across benchmarks it has been evaluated on, giving more weight to the benchmarks that carry the most signal, and scores can be compared across models to judge meaningful differences in underlying capabilities. The vertical gap comes out at 8 points, with a 90% confidence interval of 7 to 11 units.
One sensitivity worth flagging: the four-month figure depends on a soft definition of catching up. The Epoch authors note that requiring the open model's point estimate to strictly exceed the closed model would push the gap to roughly six months. So whether open weights trail by a season or half a year depends on how you score a statistical tie.

Where the estimate probably leans optimistic
Two caveats make the real gap likely larger than four months. First, open-weight models may benchmark better than they generalize. Epoch cites evidence that open releases tend to underperform on private benchmarks relative to closed models, plausibly because they are tuned more aggressively against the public ones everyone can see.
Second, leading closed labs do not always release their most capable models, for safety, commercial, or competitive reasons, so this analysis likely understates the true open-vs-closed gap whenever closed labs are sitting on more capable models that have not yet been published. The visible frontier is the one you can hit an API against, but the actual ceiling at OpenAI, Anthropic, and Google DeepMind sits somewhere above it.
The geography behind the gap
The composition of the open-weight tier has shifted dramatically and that shift is now mostly Chinese. The same dynamic shows up in Epoch's parallel US versus China analysis: since 2023, every model at the frontier of AI capabilities as measured by the ECI has been developed in the United States, and over that same period Chinese models have trailed US capabilities by an average of seven months, with a minimum gap of four months and a maximum gap of 14.
This gap closely resembles the broader gap between proprietary and open-weight models, which is unsurprising since nearly all leading Chinese models are open-weight while frontier US models remain closed. In practice, the open-weight conversation is now a story about labs like DeepSeek, Moonshot, Zhipu, and MiniMax keeping pace with whatever the closed US labs ship next.
What this means if you're shipping
For anyone choosing between hosting an open model and paying for a frontier API, four months is a usable planning horizon. It is short enough that an open release within the past quarter is likely close to what you would get from a paid endpoint on routine work, and long enough that frontier-only capabilities, particularly in long-horizon reasoning and agentic coding, will still favor closed models for tasks at the edge of what is possible. The size of the lag has stayed roughly stable across two years of accelerating progress, which is the actual signal worth tracking: the gap is not closing, but it is not running away either.