Models

OpenAI GPT-Live-1 Tops Speech-to-Speech Index

OpenAI’s new GPT-Live-1 model has claimed the top spot on the Artificial Analysis Speech-to-Speech Index by separating real-time voice processing from backend reasoning tasks.

AlphaSignal5 hrs agoModels
Image: AlphaSignal

OpenAI has launched GPT-Live-1, a full-duplex voice model that has secured first place on the Artificial Analysis Speech-to-Speech Index with a score of 81.5, edging out Grok Voice Think Fast 2.0 High by 0.2 points. The model achieves this by separating conversational mechanics from cognitive tasks. While GPT-Live-1 manages speech, turn-taking, and interruptions, it routes complex reasoning and tool calls to backend text models like Astra or Sol. This hybrid architecture eliminates the need for developers to build fragile, multi-step pipelines combining speech-to-text, language models, and text-to-speech, offering 12 available voices.

The model's performance varies depending on the chosen backend. When paired with Astra at medium reasoning effort, GPT-Live-1 scored 67.9% on the Tau-Voice agentic benchmark, beating Grok's 56.5% by 11.4 percentage points. However, it lagged on Big Bench Audio, scoring 90.1% with Astra and 89.0% with Sol, behind Grok's 97.2% and Qwen's 99.2%. On the Full Duplex Bench, the Sol configuration scored 97.3%, placing second behind Qwen Audio 3.0 Realtime Plus, though Grok still led on the Speech Agent Arena.

For practitioners, this architecture introduces trade-offs in latency and cost. Delegating tasks to a backend model results in a slower time to first audio, averaging 1.34 seconds with Astra and 1.24 seconds with Sol, compared to Grok's 0.70 seconds. The GPT-Live-1 voice layer costs $0.05 per minute via the API, with total hourly costs ranging from $4.47 to $5.83 once backend tokens are included. Despite the latency, the full-duplex capabilities are highly effective; language-learning platform Speak reported that early testing reduced false interruptions during user pauses by nearly 80%.

This is our own summary of reporting by AlphaSignal

More in Models