Together AI introduced Together Link, a free, MIT-licensed CLI in beta that connects developers’ existing coding agents to open models hosted on Together AI. It supports Claude Code, Claude Desktop, Codex, ChatGPT Desktop, OpenCode, and Pi. The idea is simple: keep the harness, swap the model, and shrink the bill. You can deploy it today on macOS or Linux with one command and a Together API key. Note the beta status: commands, routing, and the model list may change.
What is Together Link, and is it ready today?
It is a free, MIT-licensed CLI in beta that you can install and use now. It connects popular coding agents to open models hosted on Together AI. Installation takes one command on macOS or Linux and only needs a Together API key. Because it is in beta, commands, routing behavior, and model availability may shift.
Compatibility covers Claude Code, Claude Desktop, Codex, ChatGPT Desktop, OpenCode, and Pi. You keep working in familiar tools while switching models to open ones. The goal is to preserve the agent harness and lower costs without changing daily workflows.
The concept is explicit: keep your agent pipelines intact and let Together Link steer models. That lets you trial open alternatives without reconfiguring environments. There is no extra desktop app or proxy, and your agent configs remain unchanged.
“Keep the harness, swap the model, and shrink the bill” is the product’s pitch.
You need a Together API key and macOS or Linux to get started. In practice, it is a quick way to try open models across your everyday toolchain. The team also notes the beta status, so details can evolve.
What problem does it target in coding agents?
It tackles a common inefficiency: coding agents send every task to the same premium model, from one-line fixes to full rewrites. In its launch post, the Together team says engineering orgs spend tens of thousands to millions monthly on closed models. Its argument is that open models have closed much of the gap.
According to the team, Kimi K3 and GLM 5.3 target hard coding work. GLM 5.3 Flash and DeepSeek V4.1 Flash handle everyday tasks. One setup can separate “hard and expensive” from “fast and cheap” without user micromanagement.
This powers Together Link’s logic: route easy requests to fast, low-cost models, and send hard ones to frontier capability. You do not change agents, and session context stays intact. The result is fewer overpayments on small tasks.
Does that guarantee constant wins? The claim is careful: open models have closed much of the gap. Hence the flexible routing rather than a rigid, single premium default for all jobs.
How do you install Together Link and run your tools?
Installation happens with one curl command. The installer adds Bun if needed and places commands in ~/.local/bin. You can then open a launcher using togetherlink or start a tool directly: togetherlink claude, togetherlink codex, togetherlink opencode, or togetherlink pi. Shortcuts like tclaude also work.
Per the official docs, no local proxy or daemon runs. Each tool talks directly to Together’s hosted gateway. Terminal agents receive a temporary per-launch configuration that is removed at session end.
Claude Desktop and ChatGPT Desktop use separate, reversible profiles. Switching back to your normal config is a single command, like togetherlink chatgpt off. Your usual agent configuration files remain untouched, which simplifies trials.
So you install once and immediately connect familiar agents to open models. There is no local service and no manual config edits—just invoke the tool with the desired path. Next question: how does Together Link choose a model?
How does the Auto Router work, and when does Opus 5.5 escalate?
By default, sessions start on a virtual auto model. According to the launch post, the router reads the session’s first task. Quick fixes go to fast, low-cost models, and hard problems go to frontier capability. With an Anthropic API key, it routes between Opus 5.5 and GLM 5.3. Without one, it routes between GLM 5.3 and GLM 5.3 Flash.
Routing occurs once per session, so prompt caching keeps working. That matters for context stability and predictable costs. The Opus path applies only to Claude Code and Claude Desktop and is billed to your Anthropic account.
Codex, OpenCode, Pi, and ChatGPT Desktop always stay on Together models. To pin one model, place the flag before the tool name, for example togetherlink --main zai-org/GLM-5.3 claude. This provides stable, repeatable behavior across sessions.
Inside Claude Code, the /model menu maps tiers to open models. Opus runs Kimi K3, Fable runs GLM 5.3, Sonnet runs GLM 5.3 Flash, and Haiku runs DeepSeek V4.1 Flash. You think in tiers while operating open counterparts.
“Routing happens once per session, so prompt caching keeps working.”
Which models, what pricing, and how does it compare to alternatives?
The docs list Kimi K3, GLM 5.3, GLM 5.3 Flash, and DeepSeek V4.1 Flash—each with 1M context. The product page lists Kimi K3 at $3.00 in and $15.00 out; GLM 5.3 at $1.40 in and $4.40 out; and DeepSeek V4.1 Flash and MiniMax M3 at $0.30 in and $1.20 out. Current rates live on Together’s pricing page.
Billing runs on your existing Together key, via pay-as-you-go or credit packs. Each session prints token and dollar totals on exit. In Claude Code, the status line shows estimated spend beside the equivalent Opus cost. The command togetherlink usage --last 7d shows gateway-tracked spend over the last seven days.
The Together team also notes it serves the largest OpenRouter token share for DeepSeek V4.1 Flash (40.8%), GLM 5.3 Flash (28.2%), and Kimi K3 (23.1%), as of 9/30/2026. That underscores demand for these open models in cost-sensitive production scenarios.
How does it stack up against close alternatives? OpenRouter needs environment variables in your shell profile; runs without a local proxy; provides guides for Claude Code, Codex CLI, OpenCode, Cursor, and more; offers an openrouter/auto model for auto routing; includes an activity dashboard and a statusline script; provides a catalog of models; runs wherever Claude Code runs; and is a hosted service.
Claude Code Router offers a desktop app or npm CLI and a local gateway you manage on port 3456; supports 10 agents, including Claude Code, Codex, OpenCode, and Pi; uses rule-based routing with fallbacks; shows token usage and cost estimates in logs; works with any provider you configure; supports macOS, Windows, and Linux; and is MIT-licensed and free.
Ollama launch starts with ollama launch claude; can run a local server or connect directly to Ollama Cloud; supports Claude Code, OpenCode, and Claude and ChatGPT Desktop (macOS); auto routing and cost tracking are not documented; offers local and Ollama Cloud models; desktop app connect is macOS; it is free locally, and cloud needs an API key. Together Link’s edge is a curated, one-command path with built-in savings receipts. Claude Code Router offers broader provider control but requires a local gateway you manage.
Based on Together AI.