Perplexity Ships GLM 5.2 Orchestrator That Matches Claude Opus at One-Third the Cost

Perplexity post-trains GLM 5.2 as a cheaper orchestrator for Computer, hitting near-frontier performance at 34% of Claude Opus cost

·
·
Perplexity Ships GLM 5.2 Orchestrator That Matches Claude Opus at One-Third the Cost
  • Perplexity released a research preview of a new orchestrator model for Perplexity Computer, based on a post-trained version of GLM 5.2.
  • The model runs at 0.344x the cost of Claude Opus, averaging roughly half the cost across all benchmarks tested.
  • A built-in "advisor tool" automatically escalates to a stronger model when the task exceeds the orchestrator's capability.
  • GLM 5.2 is a 744B parameter MoE model with a 1M-token context window, purpose-built for long-horizon agentic coding tasks.
  • The model is hosted in the US on Nvidia B200 GPUs and is live now in research preview; full benchmarks are coming in the next few weeks.
  • Perplexity Computer is available on the Max subscription tier ($200/month) and coordinates up to 20+ frontier AI models per workflow.

Perplexity just shipped a research preview of a new orchestrator model for Perplexity Computer, its cloud-based multi-model agent platform. The model is an adapted version of Z.ai's GLM 5.2, post-trained specifically for the Computer harness, and it is already live. The headline claim: near-frontier agent performance at roughly one-third the cost of Claude Opus.

What Perplexity Computer actually does

Perplexity Computer is a multi-model AI agent that accepts a natural-language prompt and autonomously plans, browses the web, manipulates files, and calls APIs to deliver finished artifacts. Built on Perplexity's internal orchestration framework, it routes each subtask to a specialized frontier LLM and chains tool operations into end-to-end workflows without manual intervention between steps.

Perplexity's bet is that the next frontier isn't "better answers" , it's execution. Instead of replying with text, Computer is designed to take an outcome you describe, turn it into a workflow, and then delegate tasks to sub-agents that do the work in the background. Until now, the central reasoning engine driving all of that orchestration was Claude Opus, which is powerful but expensive to run at scale across every task in a workflow.

The new orchestrator: GLM 5.2, adapted for agents

The new model is not a general-purpose release. It is a version of GLM 5.2 that Perplexity post-trained specifically for the Computer agent harness , meaning the model was further trained on agent-specific tasks, tool use patterns, and the Computer environment's particular demands. Under the hood, GLM 5.2 is a 744B total parameter Mixture-of-Experts architecture with 40B parameters activated per token, supporting configurable thinking modes for step-by-step reasoning.

GLM 5.2 delivers a solid 1M-token context and has undergone months of specialized training for long-horizon coding agent scenarios, covering high-value tasks such as large-scale implementation, automated research, and performance optimization. That makes it a natural fit for an orchestrator role, where the model needs to maintain state across many steps and tools without losing track of earlier decisions.

On benchmarks, the base model is already competitive with closed-source frontier models. On FrontierSWE, which measures whether an agent can complete open-ended technical projects at the scale of hours to tens of hours, GLM 5.2 trails Opus 4.8 by only 1% while edging out GPT-5.5 by 1%. On PostTrainBench, where each agent is given an H100 GPU and judged by how much it improves small models through post-training, GLM 5.2 outperforms both Opus 4.7 and GPT-5.5, ranking second only to Opus 4.8.

The advisor tool: smart escalation

The most architecturally interesting part of this release is not the model itself , it is the advisor tool. The new orchestrator includes a native mechanism that detects when a task exceeds its own capability and automatically escalates to a stronger model. This is a form of dynamic routing built directly into the model's toolset, not bolted on as a separate layer.

In practice, this means the cheaper GLM 5.2-based orchestrator handles the bulk of tasks autonomously, and only calls in a heavier model when it genuinely needs to. The result is a tiered cost structure that keeps inference bills low without sacrificing quality on hard tasks.

The cost story

Perplexity measured cost per task against GLM 5.2 as the baseline. The numbers are striking:

  • GLM 5.2 alone: 1.0x (baseline)
  • GLM 5.2 + advisor (the new orchestrator): 2.1x on WANDR benchmark
  • Claude Opus: 6.1x on WANDR benchmark

Averaging across all benchmarks, the new orchestrator runs at roughly 0.344x the cost of Opus , less than half. WANDR is Perplexity's internal benchmark for evaluating agentic computer-use task performance, measuring how well a model completes real-world computer tasks end-to-end.

GLM 5.2 uses 43,000 output tokens per task with 37,000 dedicated to reasoning, built for exactly that kind of extended execution. The integration is OpenAI-compatible and priced at first-party rates, meaning developers can route to GLM 5.2 through the same interface they already use for GPT-5.5 or Claude without changing SDKs or managing separate credentials.

Infrastructure and availability

The model is hosted in the United States by Perplexity on Nvidia B200 GPUs , a detail worth noting for teams with data residency requirements. It is available now as a research preview inside Perplexity Computer, accessible on the Max subscription tier ($200/month). Perplexity Computer originally ran entirely in the cloud on the Perplexity Max subscription tier. Perplexity says it will publish full benchmarks in the coming weeks as the model improves through the research preview phase.

Why this matters beyond the cost headline

The competitive dynamics have shifted from bigger and faster models to systems that can coordinate activities between many different models and agents to accomplish more complicated tasks. Rather than relying on a single frontier model, Perplexity is leaning into orchestration. And orchestration, not scale alone, may define the next phase of AI.

What Perplexity is demonstrating here is that the orchestrator role in a multi-agent system does not need to be the most capable model , it needs to be the most efficient model that can still make good routing decisions. By post-training an open-weight model specifically for this role and giving it a built-in escalation mechanism, they have effectively decoupled orchestration cost from task complexity. Simple tasks stay cheap; hard tasks still get frontier-level reasoning when needed.

Perplexity's Search as Code architecture treats web search as a programmable function within an agent's reasoning loop rather than a lookup bolted onto the side. For long-horizon agentic tasks, that architecture rewards models that can plan across many steps and use retrieval strategically. GLM 5.2's long-context strengths make it a natural fit for exactly this pattern.

What this is good for

  • Long-horizon research workflows , tasks that require many web searches, cross-referencing sources, and assembling structured outputs
  • Automated report and document generation , Deep Research is now directly available in Computer work threads, and results can be turned into a report, spreadsheet, presentation, dashboard, website, or workflow without leaving the environment
  • Cost-sensitive production pipelines , teams running many agent tasks per day where Opus-level pricing is prohibitive
  • Workflows with variable complexity , the advisor tool means you get cheap inference on easy tasks and automatic escalation on hard ones, without manual routing logic

What to watch

Perplexity is still in research preview and has not yet published full benchmark numbers for the adapted orchestrator model. The WANDR cost comparisons are self-reported, and independent evaluations will be needed to confirm the performance-cost tradeoff in practice. That said, the architectural approach , a purpose-trained, cost-efficient orchestrator with native escalation , is a real and replicable pattern. This decoupling means teams can swap models as better alternatives emerge without redesigning the entire system. The open-weight foundation of GLM 5.2 also means the broader community can study and build on the same base model that Perplexity is adapting.

Trending
  • No trending articles

Comments

avatar

Next Reads