StepAudio 3 — StepFun's Full-Duplex Real-Time Voice Model Family

StepAudio 3 is StepFun's full-duplex, real-time voice model family, covering Realtime, ASR, TTS, Gen, and Music, built for developers who want voice agents that think while they speak. StepFun released the series on September 15, 2026, with several variants reaching the top of Artificial Analysis audio leaderboards.

Core Features

  • Think-While-Speaking architecture runs private reasoning in parallel with streaming speech, cutting latency without dropping deliberation.
  • Deep Perception captures acoustic cues like intent, emotion, and environment, while Seamless Duplex handles pauses, backchannels, and interruptions.
  • An integrated Voice Agent executes tool calls asynchronously without breaking the dialogue flow.
  • API access through StepFun's Open Platform, with the Realtime report citing a self-reported 98.9 on the Artificial Analysis Full-Duplex benchmark and 90.6 on MMSU.

Use Cases

  • Real-time voice agents that must respond naturally during overlapping speech.
  • Chinese-language and bilingual voice AI where StepFun has closed ground on Western providers.
  • Audio content creation spanning TTS, music, and complete audio generation.

Pricing

StepAudio 3 is available via the StepFun Open Platform API. A free evaluation quota is offered, with paid per-call credits for production traffic; exact rates are published on the platform rather than as a single flat list price.

Pros and Cons

  • Pros: Think-While-Speaking keeps latency low; strong self-reported duplex scores; full model family.
  • Cons: Headline benchmarks are self-reported; pricing is quota-based, not a flat list.

Our Take: Best for low-latency bilingual voice agents thanks to the Think-While-Speaking design; the trade-off is that the headline benchmarks are self-reported, so validate them independently before committing production infrastructure. Browse more in our AI Audio category.

FacebookXWhatsAppEmail