According to Beating, Step Audio 2.5 Realtime, an end-to-end real-time voice model by Step Cosmos, launched on its open platform API in April 2026. The model emphasizes natural conversation with customizable character personas and paralinguistic perception (tone, pauses, sighs).
In official testing across five dimensions, Step Audio 2.5 Realtime ranked first in all categories. The subjective evaluation score (real user phone app conversations) reached 80.41, compared to 68.01 for GPT-Realtime-1.5 and 67.16 for Gemini Live. Voice Q&A benchmark scored 79.80, nearly 1.5 times GPT-Realtime-1.5's 53.20. API pricing: 10 yuan per million input tokens (2 yuan with cache hits), 70 yuan per million output tokens, with continuous voice calls estimated at 3.8 yuan per hour.