Ant Group's Ling 3.1 Flash Doubles Agent Scores With 560B Parameters

Ant Group's new 560B mixture-of-experts hybrid reasoning model more than doubles its predecessor's intelligence score and targets agentic workloads with a 1M token window.

·
·
·
Ant Group's Ling 3.1 Flash Doubles Agent Scores With 560B Parameters
  • Ant Group's InclusionAI released Ling 3.1 Flash, a 560B/25B active MoE hybrid reasoning model
  • Scores 41 on Artificial Analysis Intelligence Index v4.3, more than doubling Ling 3.0 Flash's 20
  • Terminal-Bench jumps from 0% to 33%, AutomationBench-AA from 3% to 62%
  • Priced at $0.30 / $0.90 per million input / output tokens, cached input at $0.06
  • 1M token context target, currently served at 262K via Novita AI
  • Open weights planned soon, following prior Ling-flash-2.0 release pattern

Ling 3.1 Flash scales to 560B parameters and lifts agent scores

Ant Group's InclusionAI lab has released Ling 3.1 Flash, a sparse mixture-of-experts model available through Novita AI. It scored 41 on the Artificial Analysis Intelligence Index v4.3, up from 20 for Ling 3.0 Flash. The largest gains came from tool use, terminal workflows, automation, and factual accuracy.

InclusionAI says the model's weights will be published within weeks, although it has not provided an exact date, license, or checkpoint format. Until then, developers can evaluate the hosted API but cannot reproduce the results or deploy the model on their own infrastructure.

560B weights, 25B at work

Ling 3.1 Flash has roughly 560 billion total parameters and activates about 25 billion for each token, according to a model overview. A mixture-of-experts model routes each token through a selected subset of specialized parameter groups, reducing arithmetic compared with activating the entire network.

Ling 3.0 Flash had 124 billion total parameters and 5.1 billion active parameters. The new model is about 4.5 times larger overall and activates 4.9 times as many parameters per token. Its active share is approximately 4.5%, close to the predecessor's 4.1%.

Ling 3.1 Flash specifications
Specification Current detail
Total parameters Approximately 560B
Active parameters Approximately 25B per token
Hosted context window 262K tokens
Target context window 1 million tokens
Maximum output 32,768 tokens
Reasoning mode Hybrid, with an explicit thinking toggle
Modalities Text and code

InclusionAI's architecture notes describe auxiliary-loss-free sigmoid routing, multi-token prediction, QK-Norm, and Partial-RoPE. In plain terms, those techniques manage expert selection, train the model to predict several future tokens, stabilize attention, and apply positional encoding to part of each attention vector.

InclusionAI claims this design is seven times as efficient as an equivalent dense architecture. The published figure lacks a defined workload, hardware configuration, and efficiency metric, so it cannot yet support direct latency or cost comparisons.

Agent benchmarks supply the biggest gain

Artificial Analysis reports a 33-point gain on Terminal-Bench v4.0 and a 59-point gain on AutomationBench-AA. These evaluations test whether a model can complete multi-step tasks by operating terminals, invoking tools, and responding to intermediate results.

Reported benchmark changes
Metric Ling 3.0 Flash Ling 3.1 Flash Change
Intelligence Index v4.3 20 41 +21
Terminal-Bench v4.0 0% 33% +33 percentage points
AutomationBench-AA 3% 62% +59 percentage points
AA-Omniscience score -18 +2 +20 points
AA-Omniscience accuracy 18% 29% +11 percentage points
AA-Omniscience hallucination rate 44% 38% -6 percentage points
AA-Omniscience attempt rate 56% 58% +2 percentage points

AA-Omniscience is designed to separate correct knowledge, abstention, and confident error. Ling 3.1 Flash's nearly unchanged attempt rate, combined with higher accuracy and fewer hallucinations, indicates that improved answer quality drove the gain. The same evaluation reports a 97% hallucination rate for DeepSeek V4.1 Flash (Max).

Ling 3.1 Flash also recorded a GDPval-AA v2 Elo of 1,622 and an AA-Briefcase Elo of 1,400. Artificial Analysis places those results near GLM-5.3-Flash, Gemini 3.8 Flash (High), and DeepSeek V4.1 Flash (Max). Elo scores are relative to the evaluator's opponent pool and harness, so comparisons should remain within the same benchmark version.

Higher rates meet lower token use

Ling 3.1 Flash costs $0.30 per million input tokens and $0.90 per million output tokens. Cached input costs $0.06 per million tokens, an 80% discount. The predecessor charged $0.075 for input and $0.22 for output, making the new rates approximately four times higher.

At list prices, a request containing one million uncached input tokens and one million output tokens would cost $1.20. The same request would cost $0.96 if the provider accepted every input token as cache-eligible.

The model generated 218 million output tokens during the Intelligence Index evaluation, 16% fewer than Ling 3.0 Flash. That reduction suggests better token efficiency on this benchmark, although it offsets only part of the higher per-token rates.

Artificial Analysis cost per benchmark task
Model Cost per task
DeepSeek V4.1 Flash (Max) $0.32
GLM-5.3-Flash $0.42
Ling 3.1 Flash $0.99
Gemini 3.8 Flash (High) $1.24

Cost per task depends on the evaluation's prompt mix, output length, reasoning settings, and provider pricing. Application-specific tests will produce different totals, particularly for workloads with long contexts or repeated cached prompts.

API first, weights later

Developers can currently access Ling 3.1 Flash through Novita AI. The available material does not specify the provider's model identifier, API compatibility, tool-call schema, rate limits, regional availability, or data-retention policy.

A production integration therefore requires confirmation of several provider-specific details:

  • The exact model ID and endpoint format
  • How the reasoning toggle affects billing and latency
  • Whether tool calls use structured JSON or provider-specific schemas
  • Whether the 32,768-token output limit reduces the usable input window
  • Which prompts qualify for cached-input pricing
  • Rate limits, logging policies, and regional processing options

InclusionAI's earlier repository used safetensors, but the Ling 3.1 Flash format remains unconfirmed. The open-weight announcement also leaves the license, tokenizer files, inference configuration, and official quantizations unspecified.

Active compute does not solve memory

Once weights become available, the 25-billion active-parameter count should place per-token arithmetic closer to a 25B dense model than a 560B dense model. Routing overhead, attention costs, expert communication, and implementation quality will still affect throughput.

Memory capacity follows the full 560-billion parameter count because all experts generally need to remain resident or distributed across devices. Approximate raw weight storage, before runtime buffers and the key-value cache, is substantial:

Approximate raw weight memory
Weight format Approximate memory
BF16 1.12 TB
8-bit 560 GB
4-bit 280 GB

Long contexts add key-value cache memory, while MoE deployments may require expert parallelism across several GPUs or nodes. The active parameter count can reduce computation, but self-hosting still demands a large memory pool, suitable interconnects, and an inference engine that supports the released architecture.

Best suited to tool-heavy workloads

Teams building agents that chain tool calls, operate terminals, inspect large repositories, or run multi-step automation have the clearest reason to test Ling 3.1 Flash. Its benchmark gains align with those workloads, and the current 262K hosted context can accommodate large codebases and lengthy task histories.

Short-form chat and routine generation present a weaker cost case. In the cited evaluation, GLM-5.3-Flash cost $0.42 per task compared with Ling 3.1 Flash's $0.99, while the 1-million-token context remains a target rather than a deployed limit.

Ling 3.1 Flash's measurable advance is agent execution within the compared Flash tier. Its broader value will depend on provider latency, tool-call reliability, the eventual weight license, and reproducible results from independent deployments.

Trending
  • No trending articles

Comments

avatar

Next Reads