Perplexity Adds Kimi K3, the World's Largest Open-Weight Model, on US Servers

Perplexity brings Moonshot AI's record-breaking 2.8T-parameter open-weight model to its search and agentic platform, with a US-only hosting pledge to sidestep data sovereignty concerns.

·
·
Read6 min
TypeNews
TopicLlms · Api
SubtopicLong Context
  • Perplexity added Kimi K3 to its platform and Perplexity Computer for Pro ($20/mo) and Max ($200/mo) subscribers.
  • Kimi K3 is the world's largest open-weight model at 2.8 trillion parameters, with a 1M-token context window and native vision from Beijing-based Moonshot AI.
  • Perplexity is hosting the model exclusively on US-based servers, directly addressing data sovereignty concerns tied to Chinese AI law.
  • K3 benchmarks neck-and-neck with Anthropic and OpenAI's top models on coding and reasoning tasks, at roughly one-third the output token cost of Claude Fable 5.
  • The US government is actively considering banning Chinese AI models post-K3 launch; Moonshot has denied allegations of using restricted Nvidia chips or distilling US model outputs.
  • Independent testing flagged a 51% hallucination rate and a confirmed April 2026 cross-user data leak from Moonshot's own API -- Perplexity's US-hosted inference layer mitigates but does not eliminate all risk.

Perplexity has quietly dropped one of the most consequential model additions to its platform in months: Kimi K3 is now available inside both Perplexity and Perplexity Computer for Pro and Max subscribers. The move brings a frontier-class Chinese open-weight model into one of the most widely used AI search products in the US -- and Perplexity is making a pointed promise about where your data actually lives.

The model behind the headline

Kimi K3 is the flagship release from Beijing-based Moonshot AI, and its specs are hard to ignore. It is a 2.8-trillion-parameter Mixture-of-Experts model with native vision and a 1-million-token context window. To put that in perspective, Kimi K3 is roughly 75 percent larger than DeepSeek's V4 Pro, which sits at approximately 1.6 trillion parameters.

It is the world's first open-source model in the 3-trillion-parameter class, designed for frontier intelligence scenarios including long-horizon coding, knowledge work, and reasoning. The architecture leans on two internal innovations: Kimi Delta Attention, a hybrid linear attention mechanism, and Attention Residuals, which the company describes as a drop-in replacement for residual connections that delivers consistent scaling gains. In plain terms, these techniques help information flow more efficiently through very long sequences without the usual degradation that plagues large models.

It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at navigating large repositories, using tools, debugging, and iterating against images, logs, tests, and runtime feedback. K3 also ships with always-on reasoning -- you cannot turn it off, but you can dial down the reasoning effort with a reasoning_effort flag set to low, high, or max (default).

Where Perplexity fits in

Perplexity is not just a chat interface. Perplexity Computer runs multi-step workflows across research, coding, design, and deployment from a single prompt, routing tasks across 20 specialized models and connecting to 400+ applications. Kimi K3 now joins that model roster, giving Computer a new option for the heavy-lifting tasks where its 1M context window and coding chops shine most.

The addition lands across two tiers:

  • Pro ($20/month): Pro gives you unlimited Pro Searches and model switching. Kimi K3 is now selectable as one of those models.
  • Max ($200/month): The $200 monthly fee includes the agent itself, 10,000 monthly credits, unlimited Pro searches, access to advanced models, Sora 2 Pro video generation, the Comet AI browser, and unlimited Labs usage. Max subscribers get higher credit limits and earlier access to new features.

This is not the first time Perplexity has reached for a Moonshot model. Perplexity previously added Kimi K2.5, hosting it on Perplexity's own inference stack in the US for tighter control over latency, reliability, and security. K3 follows the same playbook.

The US-hosting detail is not a footnote

Perplexity's announcement specifically called out that Kimi K3 is hosted exclusively on US-based servers. That single line carries a lot of weight right now. Moonshot AI is a Beijing-based company subject to Article 7 of China's National Intelligence Law, which requires all Chinese organizations to support and cooperate with national intelligence efforts. China's Data Security Law and Cybersecurity Law impose additional data localization requirements and government inspection authority that explicitly extend to AI systems.

These obligations apply regardless of where inference runs, regardless of Moonshot's Singapore incorporation, and regardless of the company's stated privacy policy -- developers routing sensitive queries through the Kimi API are sending data through servers subject to this framework. By running inference on its own US infrastructure, Perplexity insulates users from that exposure at the data layer.

The geopolitical backdrop makes this framing even more deliberate. The US government has been moving to ban leading Chinese AI models following the release of Kimi K3, citing cybersecurity concerns, while critics argue such a ban would stifle innovation and encourage monopolies. Kimi K3 has triggered fresh claims from US officials about access to restricted NVIDIA chips and the possible use of American model outputs during training -- Moonshot has denied the claims.

Why this model, why now

The timing is not accidental. Moonshot AI, the Beijing-based startup backed by Alibaba, released Kimi K3 as a model that benchmarks show performs neck-and-neck with the most powerful proprietary systems from Anthropic and OpenAI. Kimi K3 still trails Anthropic's Claude Fable 5 and OpenAI's GPT 5.6 Sol on overall performance, but consistently outperforms other tested models.

The cost story is compelling too. K3 costs $15 per million output tokens, compared to $4.40 for GLM-5.2 and $0.87 for DeepSeek V4 -- but it is still cheaper than the equivalent US models, with Fable running at $50 for the same output. For Perplexity, hosting K3 on its own stack means it can offer frontier-level coding and reasoning at a price point that undercuts what users would pay going directly to Anthropic or OpenAI.

Moonshot AI raised $2 billion in funding in May, valuing the company at over $20 billion, with annual recurring revenue exceeding $200 million. The model's launch was so well-received that Moonshot stopped accepting new subscriptions within two days after computing demand exceeded available capacity.

Who wins, who watches carefully

For Perplexity subscribers, the calculus is straightforward: you get access to what is arguably the most capable open-weight model ever released, wrapped in a US-hosted, Perplexity-managed inference layer. The 1M context window alone opens up use cases -- think full codebase analysis, long document reasoning, and multi-session agentic tasks -- that simply were not practical before.

For Anthropic and OpenAI, the pressure is real. Frontier labs like Anthropic and OpenAI have more to worry about, because the Kimi models perform at or near the frontier. An open-weight model that matches closed proprietary systems on key benchmarks, available through a popular consumer product at a fraction of the cost, chips away at the premium pricing those labs depend on.

The broader industry implication is a shift in how frontier AI gets distributed. The open-weight nature of Chinese models such as DeepSeek V4 and Kimi K3 -- which allows companies to download the models locally and host them on private servers -- is driving adoption by giving enterprises data privacy and slashing API costs compared to closed Western alternatives. Perplexity's US-hosting approach is essentially a third path: you get the model's capabilities without the data sovereignty headache, and without the GPU cluster bill of self-hosting.

The risks are not zero. Independent testing found a 51 percent hallucination rate in Kimi K3's outputs, and Moonshot experienced a confirmed cross-user data isolation failure in April 2026 in which one user's personal data was disclosed to another. For production use cases involving sensitive data, those numbers warrant caution regardless of where inference runs.

Still, for developers and researchers who want to push the limits of what a 1M-context, multimodal reasoning model can do inside a familiar search-and-agent interface, Kimi K3 on Perplexity is now the most accessible on-ramp available.

Comments

avatar