StepAudio 3 Released: Voice AI Advances to Full-Cycle Interaction and Creation

Deep News
3 hours ago

On September 15, the latest generation of voice large models, the StepAudio 3 series, was unveiled, introducing five models at once: StepAudio 3 Realtime, StepAudio 3 ASR, StepAudio 3 TTS, StepAudio 3 Gen, and StepAudio 3 Music. These models span scenarios including real-time voice interaction, speech recognition, lifelike voice generation, and music creation, with several ranking first globally on the leaderboards of Artificial Analysis, an authoritative third-party AI model evaluation institution.

Unlike earlier approaches that largely competed on isolated capabilities such as speech recognition or generation, StepAudio 3 extends its scope to a full-stack portfolio encompassing voice understanding and generation, real-time reasoning, task execution, and audio creation. This enables AI to not only listen and speak, but also grasp the semantics and emotions behind sound, while participating in more complete interaction and content production cycles.

Designed for real-time voice interaction, StepAudio 3 Realtime supports native full-duplex conversations, allowing the system to determine when to respond, when to pause, and how to handle interruptions or continuous feedback during ongoing exchanges. On the Artificial Analysis Conversational Dynamics leaderboard, this model achieved a composite score of 98.9%, securing the top ranking worldwide.

Distinct from conventional real-time voice models that focus primarily on low latency and natural dialogue, StepAudio 3 Realtime adds voice reasoning and task execution capabilities. It can simultaneously interpret semantics, tone, emotion, and ambient sounds, with support for parallel reasoning and speech generation. Tool calls and long-running tasks are executed asynchronously, ensuring that ongoing conversations are not interrupted during task processing. The model also holds the global number one position on the Artificial Analysis Speech Reasoning leaderboard.

On the speech recognition front, StepAudio 3 ASR merges high-accuracy recognition with the contextual understanding and knowledge capabilities of large language models. Beyond ensuring precise hearing, it strengthens comprehension of professional terminology, personal names, place names, dialects, and complex contextual cues. The model supports multiple scenarios, including Chinese, English, dialects, code-switching between Chinese and English, long audio files, and specialized fields, while being optimized for challenging inputs such as low-volume speech, rapid speaking, mumbling, and background music. On the Artificial Analysis non-streaming speech recognition accuracy leaderboard, StepAudio 3 ASR achieved a Word Error Rate (WER) of 1.7%, tying for the top spot globally. A lower WER indicates fewer transcription errors, placing the model’s recognition accuracy within the world’s leading tier in third-party comprehensive testing.

Beyond real-time interaction and recognition, this release further pushes the boundaries of voice generation models. StepAudio 3 TTS targets lifelike voice production, delivering not only timbre, intonation, rhythm, pauses, and emotional expression, but also generating paralinguistic phenomena common in genuine communication, such as laughter, hesitation, repetition, and speech corrections. Built on a streaming generation architecture, it enables simultaneous generation and playback, meeting the demands of real-time voice applications.

Meanwhile, StepAudio 3 Gen consolidates previously fragmented generation tasks—such as vocals, sound effects, ambient noise, and background music—into a single unified process. Users can describe characters, scenes, and plots in natural language to produce complete audio containing multiple sound elements in one pass, with control over character timbre, emotion, and the timing and sequence of different sounds. This is broadly applicable to content production scenarios, including film, animation, gaming, radio dramas, audiobooks, and short videos.

For music creation, StepAudio 3 Music moves beyond the simple “one-sentence-to-one-song” approach, supporting songwriting, a cappella accompaniment, song covers, and multi-turn interactive creation based on ABC notation. Users can collaborate with the model through lyrics, singing, reference tracks, or other input modes to jointly refine arrangements and iterate on works, enabling AI to play a more substantial role in real-world music production workflows.

Over the past year, voice large models have transitioned from single-task recognition and generation toward a convergence of real-time interaction, sound comprehension, and content creation. With the release of the StepAudio 3 family, the audio model product landscape now comprehensively spans listening, speaking, understanding, reasoning, and creation. All five models in the StepAudio 3 series are currently available on the StepFun open platform.

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Most Discussed

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10