Google's ERA AI Tops CDC Flu Forecast Challenge Beating 39 Teams

Google's AI-generated forecasting model beat 38 other submissions in the CDC's FluSight evaluation, built by an LLM plus tree-search system now in trusted tester preview.

·
·
Google's ERA AI Tops CDC Flu Forecast Challenge Beating 39 Teams
Read4 min
TypeNews
  • Google's Google_SAI-FluEns model ranked #1 of 39 entries in CDC FluSight 2025-26 evaluation.
  • Forecasts were generated by Empirical Research Assistance (ERA), an LLM-plus-tree-search code generator.
  • ERA research was published in Nature and now also powers COVID-19 and RSV forecasts.
  • Season peaked week ending Dec 27 2025 with over 40,000 weekly hospital admissions.
  • Even winning ensemble missed turning points, under 25% peak-week coverage at two-week horizon.
  • ERA tech available to trusted testers via Google Labs experimental science tools.

Google’s AI-generated flu model leads CDC evaluation

Google_SAI-FluEns ranked first among 39 eligible individual entries in the CDC’s retrospective FluSight evaluation for the 2025-26 season, according to Google’s published report and independent evaluation coverage. The model forecast weekly influenza hospital admissions for every U.S. state, covering the current week and the following three weeks.

Empirical Research Assistance, or ERA, generated and iteratively refined the software behind the submission. Google researchers defined the task, data, and accuracy measure, while ERA used a language model and tree search to produce and test candidate programs. The result provides a public test of AI-generated scientific code under the same scoring rules applied to entries from government, academic, and industry teams.

Evaluation at a glance
Result First among 39 eligible individual entries
Season 2025-26 FluSight season
Target Weekly state-level influenza hospital admissions
Horizons Current week through three weeks ahead
Metric Weighted Interval Score on log-transformed observations
ERA access Trusted testers only, with no public API

The score behind the ranking

FluSight, the CDC’s long-running respiratory disease forecasting challenge, collects weekly submissions from participating teams between October and May. The forecasts estimate hospital admissions across several time horizons, and the hub combines submissions into an ensemble that can inform planning for state-level hospital demand.

The CDC scores forecasts after observed admissions become available, using Weighted Interval Score, or WIS, on log-transformed values. A lower WIS rewards narrow probability ranges when they contain the observed result and penalizes intervals that are too wide or miss the result. The metric therefore evaluates both accuracy and uncertainty, while the logarithmic scale makes proportional errors more comparable across jurisdictions with different admission volumes.

ERA searches code, then tests it

ERA treats scientific modeling as a program-search problem. A language model proposes candidate algorithms, each candidate runs against held-out data excluded from fitting, and an objective metric determines its score. Tree search retains promising approaches, modifies them, and explores their descendants over repeated rounds rather than accepting a single generated program.

Google began submitting ERA-generated FluSight forecasts when the 2025-26 challenge opened in November. The company also uses the system for year-round state-level COVID-19 hospitalization forecasts and the CDC’s newer respiratory syncytial virus forecasting hub. Google describes broader applications in epidemiology and geospatial analysis in the underlying Nature paper.

Peak weeks exposed weak coverage

The CDC classified the 2025-26 flu season as moderate in its preliminary in-season assessment. Weekly hospital admissions exceeded 40,000 at their highest point, rising from mid-November 2025 and peaking nationally during the week ending December 27.

The CDC’s aggregate FluSight ensemble struggled with the season’s sharpest changes. Its 50% and 95% prediction intervals, the ranges assigned those expected coverage levels, failed to anticipate the late-December increase and the mid-January decline. Around the peak, fewer than 25% of two-week-ahead intervals across jurisdictions contained the observed values. Reported ensemble coverage stabilized near 95% beginning in February 2026.

These aggregate-ensemble figures cannot isolate Google_SAI-FluEns’s behavior during the peak; they establish that the operational forecasting system had its weakest coverage around rapid reversals, when hospitals may need forecasts for staffing and capacity decisions.

What one season establishes

A first-place finish under fixed, public scoring rules offers evidence that automated program search can produce competitive epidemiological forecasting code, while the available record leaves several questions open:

  • Winning margin: Google has not published the score difference between its entry and the runner-up, preventing assessment of whether the lead was narrow or substantial.
  • Generalization: One season for one disease provides limited evidence about performance across pathogens, regions, data regimes, and future seasons.
  • Turning points: The published aggregate results show weak coverage near the peak, but they do not provide the winning model’s corresponding interval coverage.
  • Reproducibility: ERA has no public API, which limits independent attempts to reproduce the program-search workflow.
  • Reuse: The COVID-19 and RSV deployments give researchers additional settings in which to evaluate the approach.

Access remains limited

Developers currently need approval to use ERA because Google limits the technology to trusted testers through Google Labs. The program targets research collaborators rather than general application developers, while the Nature paper provides the most detailed public description of the language-model and tree-search loop.

A workable evaluation package

  • Objective: Define a numerical measure that can rank generated programs consistently.
  • Data split: Reserve historical observations for evaluation so candidate programs are tested on data excluded from fitting.
  • Baselines: Establish existing methods and scores for meaningful comparisons.
  • Constraints: Specify runtime, data-access, interpretability, and deployment requirements before beginning the search.
Trending
  • No trending articles

Comments

avatar

Next Reads