Skip to content
AI Primer
release

StepFun releases Step 5 Preview with a 1M-token context window

StepFun's multimodal Step 5 Preview is available through OpenRouter and coding tools. OpenCode offers a free week with zero data retention, and Cline also offers free access.

5 min read
StepFun releases Step 5 Preview with a 1M-token context window
StepFun releases Step 5 Preview with a 1M-token context window

TL;DR

A 24-hour H100 kernel-optimization experiment gave each model four attempts and reported its best run, with Step 5 reaching 508 TFLOPS after roughly 22 hours. An hour-long snow-globe demo generated a histogram to check where snow starts falling, a gloriously nerdy acceptance test.

600B MoE and multimodal input

StepFun's native API specifications add concrete limits beneath the 1M-context headline:

  • Maximum output: 64k tokens, with text output.
  • Images: Up to 60 per request, supplied by URL or Base64. Supported formats are JPEG, PNG, WebP, and static GIF, with low or high detail.
  • Video: URL, Base64, or a stepfile:// reference from the Files API. Supported containers are MP4, QuickTime, and Matroska.
  • Video URLs: Individual MP4 files must be under 128 MB; the documentation recommends videos shorter than five minutes.
  • Native model ID: step-5-preview.

Three reasoning effort levels

StepFun's configuration guide exposes three effort settings and two API syntaxes:

  • Effort levels: low, medium, and high.
  • Chat Completions: Set reasoning_effort.
  • Messages API: Set output_config.effort.

Other supported features include:

  • Streaming responses.
  • Tool calling.
  • Structured output through JSON Mode and JSON Schema.
  • Prompt caching.

The guide documents max_tokens as defaulting to INF, leaving output length to the model, alongside the separate 64k model output maximum. Search, code execution, and access to external services come from the application integrating the model.

Coding benchmarks

StepFun's launch table compares Step 5 Preview at High effort with Kimi K3 and GLM-5.3 at Max effort:

| Benchmark | Step 5 Preview | Kimi K3 | GLM-5.3 |
| --- | ---: | ---: | ---: |
| DeepSWE v1.1 | 67.7% | 67.5% | 66.9% |
| ProgramBench pass rate | 80.5% | 77.8% | 72.0% |
| StepCodeBench | 49.0% | 43.9% | 40.2% |

DeepSWE evaluations used the SWE-agent harness with temperature=1.0 and top_p=0.95. StepCodeBench is StepFun-built, spans 553 independent repositories and 33 programming languages, and reports avg@4, rather than a single-attempt result.

Artificial Analysis places Step 5 at 44 in its launch-week charts, tied with Kimi K3 and one point below GLM-5.3. That provides an independent aggregate comparison alongside the vendor's task-specific results.

Terminal and document-work gaps

StepFun's broader results table includes weaker results outside the highlighted coding wins:

  • Terminal-Bench v4: 33.3%, versus GLM-5.3's 41.9%, 8.6 points behind.
  • SWE-Marathon v1.1 partial score: 72.7%, versus Kimi K3's 84.4%, 11.7 points behind.
  • GDP.pdf: 14.8%, versus Kimi K3's 22.0%, 7.2 points behind.

HLE with tools also has a methodology mismatch: Step 5 and GLM-5.3 were evaluated on the text-only subset, while the other models used the full dataset. StepFun explicitly says those results are not directly comparable.

Finance benchmarks

StepFun reports 66.4% on FrontierFinance in its launch report, versus 64.1% for GLM-5.3 and 62.6% for Kimi K3. The external benchmark covers six investment use cases through 220 expert-crafted questions and 11,543 evaluation criteria; these scores are reported by StepFun.

Three internal FinStepBench evaluations separate the financial workflows:

  • LiveSearch: Retrieving and verifying timely financial information. Step 5 scores 74.5%, matching GPT-6 Astra in the table.
  • CorporateValuation: Building internally consistent forecasts and reproducible valuations. Step 5 scores 60.6%, tied with Kimi K3.
  • DeepResearch: Gathering evidence, analyzing it, and producing supported research reports. Step 5 scores 55.8%, versus 53.3% for GLM-5.3 and 48.9% for Kimi K3.

The internal suite evaluates source traceability, explicit assumptions, and reproducible calculations alongside accuracy.

Pricing and free access

OpenRouter describes the $0.05-per-million cache-read rate as a 50% launch discount. Its model listing says requests currently go to a single provider.

The free-access offers have different stated terms:

AI Gateway and agent integrations

Vercel's AI Gateway announcement uses stepfun/step-5-preview, the same slug as OpenRouter. Vercel documents text and image input, while StepFun's native guide and OpenRouter's announcement also list video.

Kilo Code and Hermes Agent also support the model, according to an early-access report, alongside the OpenCode and Cline integrations.

Visual checks in long-running tasks

The snow-globe demo came from a disclosed paid StepFun sponsorship. Its author reported a $2.26 API bill for approximately an hour of work.

In an early Cline test, ai_for_success described multi-file execution, test verification, and repeated terminal checks until changes passed. That report used an early build tested a couple of weeks before release.

Open weights on October 15

StepFun schedules the open-weights release for October 15 in its launch post, one week after the October 8 hosted release. Cline already used the term “open weights” in its promotion.

For the announced 600B total parameters, a four-bit representation works out to roughly 300 GB of raw weights, before quantization metadata, KV cache, or runtime buffers. The 27B active figure describes how many parameters participate per token.

Further reading

Discussion across the web

Where this story is being discussed, in original context.

On X· 6 threads
TL;DR1 post
Coding benchmarks1 post
Pricing and free access2 posts
AI Gateway and agent integrations1 post
Visual checks in long-running tasks2 posts
Open weights on October 151 post
Share on X