StepFun releases Step 5 Preview with a 1M-token context window
StepFun's multimodal Step 5 Preview is available through OpenRouter and coding tools. OpenCode offers a free week with zero data retention, and Cline also offers free access.

TL;DR
- Step 5 Preview brings a 1M-token context window and multimodal input to a 600B-parameter MoE with 27B active parameters, according to OpenRouter's launch announcement.
- API pricing is $1 per million input tokens and $2.70 per million output tokens, with launch-time cache reads at $0.05 per million, per OpenRouter's pricing announcement.
- OpenCode offers a free week with zero data retention in its launch announcement, while Cline also announced free access.
- StepFun reports wins over Kimi K3 and GLM-5.3 on six of eight highlighted benchmarks, according to OpenRouter's benchmark summary.
A 24-hour H100 kernel-optimization experiment gave each model four attempts and reported its best run, with Step 5 reaching 508 TFLOPS after roughly 22 hours. An hour-long snow-globe demo generated a histogram to check where snow starts falling, a gloriously nerdy acceptance test.
600B MoE and multimodal input
StepFun's native API specifications add concrete limits beneath the 1M-context headline:
- Maximum output: 64k tokens, with text output.
- Images: Up to 60 per request, supplied by URL or Base64. Supported formats are JPEG, PNG, WebP, and static GIF, with
loworhighdetail. - Video: URL, Base64, or a
stepfile://reference from the Files API. Supported containers are MP4, QuickTime, and Matroska. - Video URLs: Individual MP4 files must be under 128 MB; the documentation recommends videos shorter than five minutes.
- Native model ID:
step-5-preview.
Three reasoning effort levels
StepFun's configuration guide exposes three effort settings and two API syntaxes:
- Effort levels:
low,medium, andhigh. - Chat Completions: Set
reasoning_effort. - Messages API: Set
output_config.effort.
Other supported features include:
- Streaming responses.
- Tool calling.
- Structured output through JSON Mode and JSON Schema.
- Prompt caching.
The guide documents max_tokens as defaulting to INF, leaving output length to the model, alongside the separate 64k model output maximum. Search, code execution, and access to external services come from the application integrating the model.
Coding benchmarks
StepFun's launch table compares Step 5 Preview at High effort with Kimi K3 and GLM-5.3 at Max effort:
| Benchmark | Step 5 Preview | Kimi K3 | GLM-5.3 |
| --- | ---: | ---: | ---: |
| DeepSWE v1.1 | 67.7% | 67.5% | 66.9% |
| ProgramBench pass rate | 80.5% | 77.8% | 72.0% |
| StepCodeBench | 49.0% | 43.9% | 40.2% |
DeepSWE evaluations used the SWE-agent harness with temperature=1.0 and top_p=0.95. StepCodeBench is StepFun-built, spans 553 independent repositories and 33 programming languages, and reports avg@4, rather than a single-attempt result.
Artificial Analysis places Step 5 at 44 in its launch-week charts, tied with Kimi K3 and one point below GLM-5.3. That provides an independent aggregate comparison alongside the vendor's task-specific results.
Terminal and document-work gaps
StepFun's broader results table includes weaker results outside the highlighted coding wins:
- Terminal-Bench v4: 33.3%, versus GLM-5.3's 41.9%, 8.6 points behind.
- SWE-Marathon v1.1 partial score: 72.7%, versus Kimi K3's 84.4%, 11.7 points behind.
- GDP.pdf: 14.8%, versus Kimi K3's 22.0%, 7.2 points behind.
HLE with tools also has a methodology mismatch: Step 5 and GLM-5.3 were evaluated on the text-only subset, while the other models used the full dataset. StepFun explicitly says those results are not directly comparable.
Finance benchmarks
StepFun reports 66.4% on FrontierFinance in its launch report, versus 64.1% for GLM-5.3 and 62.6% for Kimi K3. The external benchmark covers six investment use cases through 220 expert-crafted questions and 11,543 evaluation criteria; these scores are reported by StepFun.
Three internal FinStepBench evaluations separate the financial workflows:
- LiveSearch: Retrieving and verifying timely financial information. Step 5 scores 74.5%, matching GPT-6 Astra in the table.
- CorporateValuation: Building internally consistent forecasts and reproducible valuations. Step 5 scores 60.6%, tied with Kimi K3.
- DeepResearch: Gathering evidence, analyzing it, and producing supported research reports. Step 5 scores 55.8%, versus 53.3% for GLM-5.3 and 48.9% for Kimi K3.
The internal suite evaluates source traceability, explicit assumptions, and reproducible calculations alongside accuracy.
Pricing and free access
OpenRouter describes the $0.05-per-million cache-read rate as a 50% launch discount. Its model listing says requests currently go to a single provider.
The free-access offers have different stated terms:
- OpenCode: Free for one week, with zero data retention.
- Cline: Free access is available; the announcement does not specify a duration. Its CLI announcement gives
npm i -g clineand links to the Desktop beta.
AI Gateway and agent integrations
Vercel's AI Gateway announcement uses stepfun/step-5-preview, the same slug as OpenRouter. Vercel documents text and image input, while StepFun's native guide and OpenRouter's announcement also list video.
Kilo Code and Hermes Agent also support the model, according to an early-access report, alongside the OpenCode and Cline integrations.
Visual checks in long-running tasks
The snow-globe demo came from a disclosed paid StepFun sponsorship. Its author reported a $2.26 API bill for approximately an hour of work.
In an early Cline test, ai_for_success described multi-file execution, test verification, and repeated terminal checks until changes passed. That report used an early build tested a couple of weeks before release.
Open weights on October 15
StepFun schedules the open-weights release for October 15 in its launch post, one week after the October 8 hosted release. Cline already used the term “open weights” in its promotion.
For the announced 600B total parameters, a four-bit representation works out to roughly 300 GB of raw weights, before quantization metadata, KV cache, or runtime buffers. The 27B active figure describes how many parameters participate per token.