Alibaba’s Qwen team released Qwen-Audio-3.1, a five-model audio stack for ASR, TTS, and realtime interaction. The headline model is Qwen-Audio-3.1-Realtime, a full-duplex speech model for agents that call tools. Prices dropped sharply: about 85% on Realtime, about 70% on TTS, and up to 95% on ASR. Access is managed via QwenCloud over WebSocket. No open weights were announced, so self-hosting is not available.
What exactly did Qwen ship, and what sits at the core?
Qwen shipped Qwen-Audio-3.1, an audio stack spanning ASR, TTS, and full-duplex interaction. The main model, Qwen-Audio-3.1-Realtime, targets voice agents that reason, call tools, and manage turn-taking. It aims for two-way dialogue without rigid pauses or strict reply ordering.
The full-duplex focus lets the model listen and speak in parallel. It ingests speech cues, decides when to answer, and streams voice back. This design fits scenarios with interruptions, clarifications, and quick feedback loops.
The stack covers speech recognition, speech synthesis, and interactivity. That reduces friction between speech-to-text, reasoning, and text-to-speech. The result is consistent agent behavior across the entire conversation.
Qwen also strengthened tool use. It includes function calling, web search, structured outputs, and a context cache. Fine-tuning is available, aligning responses with domain processes and voice scenes.
Why does this matter now? Voice interfaces are quickly becoming standard. Full-duplex reduces cognitive load and keeps a natural conversation tempo. The agent not only hears but also acts, reads context, and stays goal-directed.
“A 5-model audio stack spanning ASR, TTS and realtime interaction.”
Is it deployable, and how is pricing structured?
Qwen-Audio-3.1 is available as a managed API. The qwen-audio-3.1-realtime-plus model runs on QwenCloud over WebSocket. No open weights were announced, so deployment is limited to cloud access through a managed stack.
The model page lists both text and audio as inputs and outputs. Context is 262K tokens, with 245K max input and 16K max output. Default limits are 60 requests and 100K tokens per minute. These bounds clarify performance for interactive use.
Pricing is $6.4 per 1M audio input tokens and $0.8 per 1M text input tokens. Text and audio output costs $24 per 1M tokens, with output text not charged. This schedule makes voice integrations cost-predictable.
Key features include function calling, web search, structured outputs, context cache, and fine-tuning. They enable programmatic scenarios, tool orchestration, policy grounding, and typed replies. That helps with state updates and permissioned actions.
A companion model targets offline long-audio transcription: Qwen-Audio-3.1-ASR-Flash-Filetrans. It supports hot words, speaker separation, punctuation, and multilingual plus Chinese dialect recognition. It costs $0.15 input and $0.47 output per 1M tokens.
“Default limits are 60 requests and 100K tokens per minute.”
Architecture: two models behind one voice — how does it work?
Under the hood, two models share the same Audio Encoder and LLM design. The first is a full-duplex decision model. It predicts whether to keep listening, speak, stop, or resume. The second is a speech-to-text model that composes the response.
Once text is ready, a context-aware voice renderer turns it into streaming speech. The renderer conditions on conversation history, voice cues, and acoustic context. That preserves tone, rhythm, and relevance expected in live talk.
This split yields a clean decomposition. Turn-taking choices stay separate from content composition, and TTS does not dictate logic. Instead, the Think–Act–Speak loop runs transparently, controllably, and reproducibly.
The architecture enforces voice discipline. The agent can cut a phrase, wait for clarification, and return to topic. It then renders speech quickly, retaining tone and conversational pace.
“Architecture: 2 Models Behind 1 Voice.”
How it was trained: Think, Act, Speak and Coordinate
Training is organized into three layers: Think, Act, and Speak and Coordinate. First, the model reasons and realigns with its source text LLM. Next, it acts in executable environments and receives rewards. Finally, it decides when and how to speak, coordinating duplex.
The Think layer uses M²-OPD. Core-Cocktail SFT re-anchors the audio model to the source text LLM using million-hour-scale paired data. Multimodality OPD follows, scoring each token of the student’s own trajectory.
A Text Teacher and a frozen Audio Reference score each token. This is on-policy distillation, not imitation of pre-written answers. Domain experts for empathy, pragmatic intent, and acoustic scenes are then trained with GRPO. Multi-Teacher OPD merges them into one deployable model.
The Act layer uses Executable Environments. Each domain bundles a tool pool, a stateful JSON database, and a natural-language business policy. Domains are seeded from open-source tool and MCP server definitions.
Every task defines one of three outcomes: a write, a justified refusal, or an unsupported request. Scoring checks terminal state, then permitted writes, then behavioral assertions. A fluent reply cannot rescue a failed state check.
“A fluent reply cannot rescue a failed state check.”
GRPO receives rewards at dialogue, milestone, and turn level. Search training penalizes redundant queries with r_query = q min(1, n_ref / n_pred). Mean queries per search call fell from 4.37 to 1.05, while Trigger F1 slipped from 60.87% to 58.61%.
The Speak and Coordinate layer decides whether, when, and how to speak. On Full-Duplex-Bench v1.5, replies to people talking to someone else fell from 0.13 to 0.03. On v3.0, the filler rate dropped from 0.7590 to 0.2960.
There are trade-offs. After interruptions, the unwanted resume rate rose from 0.035 to 0.130. Interruption stop latency is 1.116 seconds, versus 0.383 for GPT-Realtime-2. That frames expectations for specific deployment contexts.
“Interruption stop latency is 1.116 seconds, versus 0.383 for GPT-Realtime-2.”
Comparisons, benchmarks, and takeaways — where does Qwen-Audio-3.1 stand?
Gains are visible across multiple metrics. Versus 3.0, Audio MultiChallenge rises from 47.12 to 52.21. The 14-language BBA average climbs from 81.7% to 88.1%. FLEURS WER falls from 9.01 to 3.98. That signals better understanding and multilingual range.
There is an important caveat. The τ-Voice figures use a half-duplex speech-to-text adaptation and are not comparable to official full-duplex results. In a 50-session human red-team study, GPT-Realtime-2 still leads, 96.00% versus 92.00%. That sets a performance bar for field tests.
An interactive explainer lets you explore the Think–Act–Speak loop. It shows duplex decisions, a scored training episode, and the search reward. That helps reveal mechanics beyond headline numbers.
How does it stack up? Qwen-Audio-3.1-Realtime-Plus accepts text and audio and outputs text and audio. OpenAI GPT-Realtime-2 accepts text, audio, and image and outputs text and audio. Google Gemini 3.8 Live accepts text, images, audio, and video and outputs text and audio.
Context limits also differ. Qwen offers 262K context and 16K max output. GPT-Realtime-2 has 128K context and 32K output. Gemini 3.8 Live provides 131,072 context and 65,536 output. That impacts dialogue length and output breadth.
Function calling is supported across all three. Qwen includes built-in web search, while GPT-Realtime-2 does not list it. Gemini uses Google Search grounding. For reasoning, Qwen offers a Thinking mode with 2K max reasoning; GPT allows configurable effort; Gemini uses interleaved reasoning.
Pricing contrasts are clear. Audio input is $6.4 per 1M tokens on Qwen, $32 on OpenAI, and $3.00 on Google. Audio output is $24 on Qwen, $64 on OpenAI, and $12.00 on Google per 1M. No open weights are available on all three. Sources are QwenCloud, OpenAI docs, and Model page.
Note that prices are list rates checked September 28, 2026. Audio tokenization differs by provider, so rates are not directly comparable. Treat them as directional context rather than accounting equivalence.
Key takeaways condense into several points. On a τ-Voice adaptation, task success rises from 78.4% to 82.0% over 3.0. Replies to background speech drop from 73.0% to 13.0% on Full-Duplex-Bench v1.5. That shows stricter listening discipline.
Next, scale and cost. A 262K context, function calling, and web search come at $6.4 per 1M audio input tokens. Tool use is trained with GRPO inside self-evolving executable environments. Multi-turn attack success falls to 26.0% (Chinese) and 23.5% (English).
What is Qwen-Audio-3.1-Realtime? It is a full-duplex speech model from Qwen for agents that reason, call tools, and manage turn-taking. It listens, speaks, stops, and resumes in realtime. The focus is productive, controllable conversation.
Can I self-host it? No open weights were announced. Access is through the QwenCloud API as qwen-audio-3.1-realtime-plus. Connectivity is over WebSocket, simplifying interactive channels.
How much does it cost? $6.4 per 1M audio input tokens and $24 per 1M output tokens for text and audio. Output text is not charged. These rules set clear budget expectations.
Based on the provided source material.