Alibaba’s Qwen team has released Qwen-Audio-3.1 , a 5-model audio stack spanning ASR, TTS and realtime interaction. The main model is Qwen-Audio-3.1-Realtime , a full-duplex speech model built for voice agents that call tools. Qwen also cut prices : about 85% on Realtime, about 70% on TTS and up to 95% on ASR. Is it deployable? Yes, as a managed API. qwen-audio-3.1-realtime-plus is live on QwenCloud over WebSocket. No open weights were announced. What Ships on QwenCloud The model page lists text and audio as both input and output. Context is 262K tokens, with 245K max input and 16K max output. Default limits are 60 requests and 100K tokens per minute. Pricing is $6.4 per 1M audio input tokens and $0.8 per 1M text input tokens. Text and audio output costs $24 per 1M tokens, with output text not charged. Key features include function calling, web search, structured outputs, context cache and fine-tuning. A companion model, Qwen-Audio-3.1-ASR-Flash-Filetrans , targets offline long-audio transcription. It supports hot words, speaker separation, punctuation and multilingual plus Chinese dialect recognition. It costs $0.15 input and $0.47 output per 1M tokens. Architecture: 2 Models Behind 1 Voice The system runs 2 models with the same Audio Encoder and LLM design. A full-duplex decision model predicts whether to keep listening, speak, stop or resume. A speech-to-text model writes the response content as text. A context-aware voice renderer then turns that text into streaming speech. It conditions on conversation history, voice cues and acoustic context. Training is organized int
Source: MarkTechPost
