StepAudio 3: A Voice Model That Thinks While It Speaks, and a Music Model That Plans the Song First

StepFun’s StepAudio 3 Realtime reasons privately while it talks, and StepAudio 3 Music writes an arrangement before generating audio. The voice model reports 98.9 on a full-duplex speech test; the music model ranks fourth on an independent leaderboard.
artificial-intelligence
Author

Kabui, Charles

Published

2026-09-22

Keywords

realtime-voice-agents, full-duplex-speech, ai-music-generation, abc-notation, speech-to-speech