Upstage's Solar 31B Hits 2,000 Tokens per Second on Cerebras Hardware
Upstage's Solar 31B now runs at up to 2,000 tokens/sec on Cerebras hardware, making deep research queries across hundreds of sources feel nearly instant.
- Speed milestone: Upstage's Solar 31B now runs on Cerebras hardware at up to 2,000 tokens/sec, demonstrated completing a deep research query across 246 sources.
- Why it matters: Agentic and deep research workloads chain hundreds of inference calls; at 2,000 tok/sec, multi-source research that takes minutes on GPUs finishes in seconds.
- Solar 31B credentials: Upstage's model ranked above GPT-4.1 on the Artificial Analysis Intelligence Index -- Korea's first frontier-class LLM -- while matching 70B-scale models at 31B parameters.
- Cerebras momentum: Fresh off the largest semiconductor IPO ever ($6.4B raised) and a $20B+ OpenAI deal, Cerebras is now expanding its inference platform to international AI labs.
- Upstage's rise: Korea's first generative AI unicorn, backed by $380M in sovereign investment, with 130%+ YoY revenue growth and enterprise deployments at Samsung and Korean insurers.
- Competitive signal: Cerebras runs 6.7x faster than the next-fastest GPU cloud on comparable workloads, putting direct pressure on GPU-based inference providers for agentic use cases.
Korea's Upstage AI and Cerebras Systems have announced a collaboration that puts Solar 31B -- Upstage's flagship language model -- on Cerebras inference hardware, hitting speeds of up to 2,000 tokens per second. To put that in perspective, Cerebras demonstrated the model completing a deep research query across 246 sources. That kind of multi-step, multi-source workload is exactly where raw inference speed changes the user experience from "go grab a coffee" to "watch it happen in real time."
Two companies with a lot to prove
Upstage was founded in 2020 by Sung Kim, a former professor at the Hong Kong University of Science and Technology who previously led Naver's Clova AI team. The company has been on a remarkable trajectory: it recently crossed the $1 billion valuation threshold to become the first generative AI company in Korea to achieve unicorn status, with total cumulative funding reaching approximately $279.7M. More significantly, Korea's financial authorities approved a 560 billion-won ($380.6 million) investment in Upstage -- a direct signal that Seoul sees the company as a national AI champion.
Cerebras, founded in 2016, built its reputation on a genuinely different approach to chip design. Its core product, the Wafer-Scale Engine, is substantially larger than standard GPUs -- rather than cutting a silicon wafer into hundreds of individual chips, Cerebras uses the entire wafer as one massive processor. That architecture is the reason for the speed: Cerebras solves the memory bandwidth bottleneck by building the largest chip in the world and storing the entire model on-chip, integrating 44GB of SRAM on a single chip and eliminating the need for external memory and the slow lanes linking external memory to compute.
Why 2,000 tokens/sec actually matters
For a single-turn chat, speed above ~100 tokens/sec is mostly invisible to a human reader. But deep research is a fundamentally different workload. The best research agents don't just summarize individual documents -- they follow citation trails, cross-reference findings across papers, identify contradictions in the literature, and produce structured reports that would take a human researcher days to compile. Each of those steps is a separate inference call, and the calls chain together.
Agentic workflows chain many token generations across tool calls, and end-to-end latency is dominated by inference throughput. At high speeds, an agent that would take 30 seconds on standard GPU infrastructure completes in under 3 seconds -- the difference between users abandoning a workflow and adopting it as core to their day. The 246-source demo Cerebras showed is a direct illustration of this: the bottleneck is not the model's intelligence, it's how fast it can generate tokens while hopping between sources.
Solar 31B: small model, frontier ambitions
Solar Pro 2 builds on its predecessor with a significant performance upgrade, increasing its parameter size from 22B to 31B while remaining in the small language model category. The model introduces a hybrid architecture with two user-selectable modes: "Chat Mode" for fast responses and "Reasoning Mode" for structured, multi-step logical thinking. Its Chain-of-Thought reasoning approach significantly improves performance in advanced tasks like math, coding, and logic-heavy workflows.
The benchmarks back up the ambition. Solar Pro 2 ranked above GPT-4.1 and other leading big tech models in the Artificial Analysis Intelligence Index, becoming Korea's first-ever global AI frontier model. The achievement puts Korea alongside AI giants like OpenAI, Google, and Meta in the top 10 of the world's most advanced frontier model developers. The key differentiator is efficiency: at 31B parameters, Solar Pro 2 matches the performance of models up to 70B in size, and excels in English, Japanese, and especially Korean -- outperforming ~70B models in multiple benchmarks.
Here is what the model is designed to do well:
- Agentic task execution: Solar Pro 2 has evolved into a fully agentic LLM, capable of executing multi-step tasks by interacting with external tools -- autonomously searching the web, analyzing results, and generating structured outputs.
- Hybrid reasoning: Users can switch between Chat Mode for fast responses and Reasoning Mode for structured, multi-step logical thinking.
- Multilingual strength: Solar Pro 2 showcases remarkable strength in Korean language understanding, matching or outperforming language models by tech giants across key benchmarks, while excelling in high-stakes domains such as finance, medicine, and law.
- Enterprise deployment: Upstage's models are already being used by Fortune 500 companies, including Samsung, and are widely used by Korean insurance companies.
The bigger picture for Cerebras
This partnership is not happening in isolation. Cerebras has been aggressively expanding its model roster on its inference platform. The company raised $6.4 billion in what became the largest semiconductor IPO of all time, and announced a multi-year deal with OpenAI for 750MW valued at more than $20 billion. The Upstage deal is a different kind of signal -- it shows Cerebras is positioning itself as the inference layer for international AI labs, not just US hyperscalers.
Qwen3 235B Instruct already runs more than ten times faster than leading GPU clouds on Cerebras, and Qwen3 Coder 480B reached 2,000 tokens per second -- the same ceiling Solar 31B is now hitting. The pattern is consistent: Cerebras is proving that its wafer-scale architecture can run models from any lab at speeds that GPU clusters simply cannot match for single-user latency. Cerebras is 6.7x faster than the next-fastest GPU cloud and 23x faster than the median on comparable workloads.
Who wins here
The clearest winners are developers building agentic or deep research products who want to use a capable, efficient model without paying frontier-model prices. Solar 31B on Cerebras gives them a 31B model running at speeds previously associated with much smaller models. The Cerebras Inference API is fully compatible with the OpenAI Chat Completions API, making migration seamless with just a few lines of code.
For Upstage, the collaboration solves a real go-to-market problem. Having a strong model is one thing; having it run fast enough to be compelling in a live demo -- 246 sources, visibly instant -- is a different kind of proof point for enterprise sales. The company has delivered revenue growth exceeding 130% year-over-year and was selected as the lead organization for Korea's Ministry of Science and ICT's Independent Foundation Model project. International infrastructure partnerships like this one are how that growth extends beyond Korea.
The competitive pressure lands squarely on GPU-based inference providers. For most of 2025, the AI race was about model intelligence. In the past few months, the race has shifted -- model intelligence is still critical, but across every major frontier lab, inference speed has become a new and urgent focus. A Korean model running at 2,000 tokens/sec on American wafer-scale silicon, completing deep research queries in seconds, is a concrete example of what that shift looks like in practice.