Step-1o Audio
First-generation end-to-end voice model for low-latency voice conversation, with preset voice styles and tool calling.
Specs
- Context
- —
- Max output
- —
- Input price
- $3.57/ 1M tokens
- Output price
- $8.57/ 1M tokens
- Released
- —
Context
- Not published
Pricing · USD / 1M tokens
- Input
- $3.57
- Output
- $8.57
- Cache read
- $0.71
- Cache write
- —
Modalities
- Input
- AudioText
- Output
- AudioText
API
- API types
Reasoning
- Not published
Info
- Status
- Active
- Released
- —
- Knowledge cutoff
- —
Features
- Confirmed
How to call
1 provider
stepfunmodel = step-1o-audio
JSON
Standard format
{
"id": "step-1o-audio",
"object": "model",
"created": null,
"owned_by": "stepfun",
"name": "Step-1o Audio",
"api": {
"types": [
"audio"
]
},
"limits": {
"context": null,
"input": null,
"output": null
},
"modalities": {
"input": [
"audio",
"text"
],
"output": [
"audio",
"text"
]
},
"reasoning": {
"supported": null,
"efforts": []
},
"pricing": {
"currency": "USD",
"unit": "1M_tokens",
"input": 3.57,
"output": 8.57,
"cache_read": 0.71,
"cache_write": null
},
"features": [
"tools"
],
"info": {
"status": "active",
"release_date": null,
"knowledge_cutoff": null,
"description": "First-generation end-to-end voice model for low-latency voice conversation, with preset voice styles and tool calling.",
"docs": "https://platform.stepfun.com/docs/zh/guides/models/audio",
"verified_at": "2026-10-11"
}
}Official sources
3