Together AI's Link Cuts Coding Agent Model Costs by 80%

Together AI's new CLI drops frontier open models like Kimi K3 and GLM 5.3 into Claude Code, Codex, and OpenCode with one install command.

·
·
·
Read5 min
TypeNews
  • Together AI launched Together Link, a free MIT-licensed CLI connecting coding agents to open models.
  • Supports Claude Code, Claude Desktop, Codex, ChatGPT Desktop, OpenCode, and Pi via one install command.
  • Auto Router picks GLM 5.3, Kimi K3, MiniMax M3, or DeepSeek V4.1 Flash per session.
  • With an Anthropic key, hard tasks escalate to Opus 5.5 in Claude harnesses only.
  • Claims 50 to 80 percent spend reduction versus running every session on Opus 5.5.
  • Session receipts and togetherlink usage show token totals and seven-day spend history.

Together Link routes coding agents to open-weight models

Together AI has released Together Link, a free, MIT-licensed CLI that connects existing coding agents to models on Together’s hosted inference platform. The beta supports Claude Code, Claude Desktop, Codex, ChatGPT Desktop, OpenCode, and Pi.

For teams already standardized on one of those clients, Link offers a way to change the model provider while preserving the agent interface, project setup, and tool-calling workflow. Together claims this can reduce model spending by more than 50%, though actual savings depend on routing, token usage, caching, and model choice.

Six clients, one gateway

  • Platforms: The curl-based installer supports macOS and Linux.
  • Required credential: A Together API key is enough to start. An Anthropic key enables optional Opus routing in supported Claude clients.
  • Cost: The CLI is free. Inference remains metered through Together or Anthropic, depending on the selected model.
  • Status: Together Link is in beta, so teams should test compatibility before adopting it broadly.

Keep the harness, change the endpoint

In an agentic coding tool, the harness is the surrounding agent loop, interface, tool permissions, shell integration, and project context. Together Link preserves that layer and redirects model requests to Together’s gateway.

Together says the CLI runs no local proxy or background daemon. Terminal agents receive a temporary configuration for each launch, which is removed when the session ends. Each client communicates directly with the hosted gateway.

Together Link CLI listing Claude Code, Codex, and Claude Desktop as featured agents
Together Link configures supported agents from a command-line interface.

The router commits once

Together Link uses a virtual auto model by default. The Auto Router examines the first task in a session, selects a backend, and keeps that model for the rest of the session. Together describes quick fixes as candidates for faster, cheaper models and more demanding tasks as candidates for higher-capability models.

The selected backend remains fixed even if later tasks become more complex. That session-level decision allows repeated prompt prefixes to reach the same provider, preserving prompt-cache hits and avoiding the added latency that per-request routing can create.

Automatic routing by credential and client
Client and credentials Possible routes Billing
Claude Code or Claude Desktop with an Anthropic key Opus 5.5 or GLM 5.3 Opus usage goes to the Anthropic account; Together models use Together billing
Claude Code or Claude Desktop without an Anthropic key GLM 5.3 or GLM 5.3 Flash Together
Codex, OpenCode, Pi, or ChatGPT Desktop Together-hosted models Together

Developers can bypass automatic routing and pin a model with the --main flag:

code
togetherlink --main zai-org/GLM-5.3 claude

Four models, wide price spread

Together Link exposes four open-weight models positioned for agentic coding. The following serverless rates are the published figures cited for the release and may change over time.

Published serverless pricing in US dollars per one million tokens
Model Input Output
Kimi K3 $3.00 $15.00
GLM 5.3 $1.40 $4.40
MiniMax M3 $0.30 $1.20
DeepSeek V4.1 Flash $0.30 $1.20

Together claims teams can cut model spending by more than 50% by combining existing agent harnesses with its hosted models. Its FAQ estimates savings of up to 80% compared with sending every session to Opus 5.5. Workload complexity, output length, cache performance, and the share of sessions routed to Opus will determine the actual result.

Receipts expose the running bill

Each completed session prints a receipt with token counts and dollar totals. The togetherlink usage command reports spending from the previous seven days, giving teams a built-in view of costs from long-running agents and repository-heavy tasks.

Together Link session receipt with token totals and costs for Kimi K3 and GLM 5.3 Flash
Session receipts break down model usage, tokens, and cost.

Profiles, privacy, and rollback

Claude Desktop and ChatGPT Desktop use separate, reversible profiles. A command such as togetherlink chatgpt off restores the standard configuration. Together says existing ~/.claude settings, memory, and history remain untouched.

Because model requests reach hosted services, prompts and any repository content included by the agent leave the local machine. Opus requests may involve Anthropic when that route is enabled. Teams handling regulated or confidential code should review the providers’ retention, privacy, residency, and access-control terms before deployment.

Tool permissions remain governed by the selected agent. Together Link changes model routing and does not reduce the shell, filesystem, or network access already granted to the harness.

Together wants the routing layer

Together is betting that open-weight models now handle enough coding work to make cost-based routing practical. Its serverless platform provides managed, usage-based inference, while the optional Opus path covers sessions that the router classifies as more demanding.

In its launch materials, Together cites the largest shares of OpenRouter tokens for DeepSeek V4.1 Flash at 40.8%, GLM 5.3 Flash at 28.2%, and Kimi K3 at 23.1%. Those figures measure provider traffic through OpenRouter and support Together’s claim that its infrastructure already serves substantial workloads for these models.

A shared integration across six agent clients gives Together a central place to add or change model options as new releases arrive. The MIT license applies to the Link wrapper, while hosted inference remains a metered service. The GitHub repository contains the source, and the command reference covers setup, headless runs, model selection, usage reporting, and image generation.

Trending
  • No trending articles

Comments

avatar

Next Reads