Upstage's Solar Mini 4 Triples Its Benchmark Score With 3B Active Parameters

Upstage's new 35B MoE reasoning model hits a Pareto-optimal sweet spot for intelligence per active parameter, but verbose outputs make tasks slow and costly.

·
·
Upstage's Solar Mini 4 Triples Its Benchmark Score With 3B Active Parameters
  • Upstage released Solar Mini 4, a 35B total / 3B active MoE reasoning model.
  • Scores 24 on Artificial Analysis Intelligence Index, triple Solar Pro 3's score of 8.
  • Pricing: $0.10/$0.40 per 1M input/output tokens, with 50% launch discount through October 22.
  • Strong on long-context (83% AA-LCR) and hallucination control (64% non-hallucination rate).
  • Weak on agentic coding: 1% Terminal-Bench, 22% AutomationBench-AA; task cost 5x GPT-6 Luna.
  • Available via Upstage Console, OpenRouter, and on-premises; proprietary weights, 1M context, Korean/English/Japanese.

Solar Mini 4 nearly triples its benchmark score with 3B active parameters

Korean AI company Upstage has released Solar Mini 4, a proprietary text reasoning model built around sparse mixture-of-experts routing. Artificial Analysis gives it 24 on its composite Intelligence Index, nearly three times Solar Pro 3’s score, while Upstage has reduced per-token prices by roughly one-third. High reasoning-token use and limited prompt-cache reuse push its measured cost per task above several peers.

Three billion parameters fire per token

A mixture-of-experts model stores multiple specialist subnetworks and routes each token through a selected subset. Upstage reports 35 billion total parameters and 3 billion active parameters per token. The active count approximates inference work for each token, while the full parameter count still affects memory, storage, and deployment requirements.

Artificial Analysis’s benchmark comparison places Solar Mini 4 near the leading edge of compact sparse models. Qwen3.6 35B A3B scores 18 with the same active-parameter count, while K2 Horizon MoVA 36B A4B scores 25 with 4 billion active parameters. Solar Mini 4’s proprietary weights prevent independent verification of Upstage’s parameter and routing claims.

Long context leads the scorecard

  • Long-context reasoning: Solar Mini 4 scores 83% on AA-LCR v1.1, matching MiniMax-M3 and GPT-6 Luna (max). Gemini 3.8 Flash and GPT-6 Astra each score 81%.
  • Scientific coding: It reaches 48% on SciCode, one percentage point above MiniMax-M3 and Inkling (xhigh).
  • Abstention and factuality: Its AA-Omniscience non-hallucination rate is 64%, compared with 32% for Inkling (xhigh) and 23% for GPT-6 Luna (max). The model abstains on about half the questions and answers 18% correctly, so the non-hallucination score partly reflects its willingness to withhold an answer.
  • Decode speed: Artificial Analysis reports 204 tokens per second, above the 110-token-per-second median for reasoning models in the same price tier.

Reasoning volume drives the bill

Artificial Analysis records 88,000 output tokens per Intelligence Index task, including 72,000 internal reasoning tokens. That total is about 2.5 times Inkling (xhigh)’s output and five times Gemini 3.5 Flash-Lite’s output. Those models score 25 and 22, respectively.

The measured decode speed still produces an average generation time of 7.1 minutes per task. GPT-6 Luna (max) averages 5.8 minutes, while Inkling (xhigh) averages 2.8 minutes. Solar Mini 4 costs about $0.36 per completed benchmark task, compared with $0.07 for GPT-6 Luna.

Prompt caching reuses computation for unchanged input prefixes, reducing latency and input charges when applications resend conversation history or large documents. Solar Mini 4 serves 48% of repeated context from cache in the benchmark, compared with 99% for GPT-6 Luna. Uncached input contributes about $0.30 to Solar Mini 4’s $0.36 task cost.

Agentic evaluations show weaker results. Solar Mini 4 scores 1% on Terminal-Bench 4.0 and 22% on AutomationBench-AA. It records 1072 Elo on GDPval-AA and 872 on AA-Briefcase, two comparative rankings for agentic knowledge work. These results provide limited support for autonomous terminal and multi-step office workflows.

These task-level figures come from the Artificial Analysis evaluation harness. Production cost will vary with prompt length, reasoning volume, cache behavior, output limits, retries, and the percentage of runs that complete successfully.

Endpoints are live, limits vary

Solar Mini 4 is available through Upstage Console, the Playground, on-premises deployment, OpenRouter, and other gateways, according to the published release details. Upstage lists tool calling and structured outputs among the supported capabilities and targets Korean, English, and Japanese workloads.

Published context and output limits differ across listings. Upstage advertises a 1-million-token context window, while its API documentation lists 512K. Maximum output is listed between 128K and 262K tokens. Developers should verify the limits, pricing, and cache behavior of the specific endpoint selected for production.

Published Solar Mini 4 specifications
Specification Published value
Architecture Mixture of experts, 35B total parameters and 3B active per token
Context window 1 million tokens advertised; Upstage API documentation lists 512K
Maximum output 128K to 262K tokens across published listings
Modalities Text input and text output
Languages Korean, English, and Japanese
Standard input price $0.10 per 1 million uncached tokens
Standard output price $0.40 per 1 million tokens
Cached input price $0.01 per 1 million tokens
Knowledge cutoff February 2026
License Proprietary; weights are unavailable

The cited launch offer advertises a 50% discount through October 22, reducing prices to $0.05 per million input tokens and $0.20 per million output tokens. Ongoing cost estimates should use standard pricing unless the serving provider confirms the promotion.

Korean documents are the clearest target

  • Long-context Korean and Japanese analysis: The supported languages and AA-LCR result justify testing the model on document review, retrieval synthesis, and one-shot analysis.
  • Scientific coding assistance: The SciCode score supports a pilot for bounded code-generation tasks. The Terminal-Bench result provides little evidence for autonomous shell operation.
  • Factual question answering: Applications need an explicit fallback path for abstentions and unanswered requests because the model answers only 18% of AA-Omniscience questions correctly.
  • Long-running agents: Growing conversation histories can magnify uncached-input charges, while the model’s reasoning volume increases latency and output cost.

Production pilots should measure end-to-end cost per successful completion, reasoning-token volume, prompt-cache hit rate, latency, abstention frequency, tool-call success, and endpoint-specific limits. Those measurements will show whether Solar Mini 4’s sparse compute and low token prices translate into an economical workload.

Trending
  • No trending articles

Comments

avatar

Next Reads