Nous Research Offers Step 5 Preview Free for a Week on Its Portal

StepFun's 600B/27B mixture of experts model is free on Nous Portal for a week, scoring in GPT-6 Luna territory on agent benchmarks.

·
·
·
Nous Research Offers Step 5 Preview Free for a Week on Its Portal
  • StepFun's Step 5 Preview is free on Nous Portal for one week
  • 600B total / 27B active sparse MoE, 1M-token context, vision input, 64k output
  • Scored 32.83 on Hermes Index, just behind GPT-6 Luna (33.89) and GLM 5.3 Flash (34.95)
  • Normal API pricing is $1 input / $2.70 output per million tokens
  • Vendor-reported 67.7 on DeepSWE v1.1 high reasoning, competitive with Kimi K3 Max
  • Open weights promised October 15, requiring ~1.2TB BF16 for self-hosting

Step 5 Preview gets a free week on Nous Portal

Nous Research has made StepFun’s Step 5 Preview available at no charge on Nous Portal for one week. The flagship model combines a 600-billion-parameter sparse architecture, vision support, and a 1-million-token context window. Its corrected Hermes Index score of 32.83 trails GPT-6 Luna by 1.06 points.

The promotion lets developers test the model without StepFun’s usual token charges, listed in its published API pricing as $1 per million input tokens and $2.70 per million output tokens. Nous Portal also provides the Hermes Agent harness used for the public leaderboard, allowing teams to evaluate the model under a comparable agent workflow.

Only 27B weights fire per token

Step 5 Preview uses a sparse mixture-of-experts architecture. The model stores roughly 600 billion parameters but activates about 27 billion, or 4.5%, for each token. This design reduces computation per request, though deployments still need enough memory to hold or stream the full set of weights.

The StepFun model page lists the following capabilities:

  • Model core: 600B total parameters, 27B active parameters per token, and a 92-layer narrow, deep architecture.
  • Context and output: A 1M-token context window and up to 64,000 output tokens.
  • Modalities: Text, image, and video input.
  • API features: Streaming, tool calling, JSON mode, JSON Schema, prompt caching, and low, medium, or high reasoning effort.

Million-token prompts make conventional attention expensive because the model must compare large numbers of token relationships. StepFun addresses that cost with Sparse Grouped-Query Attention and block-wise token merging. The company says this reduces indexing and top-k selection costs to about one-eighth of a denser baseline while combining overlapping selections from neighboring blocks.

Training focused on long-running agent tasks through on-policy reinforcement learning, where the model generates trajectories and learns from the results of its current policy. StepFun also reports bit-for-bit alignment between training and inference for expert routing. Its serving stack uses load-aware scheduling, FP8 arithmetic, MTP-3 speculative decoding, and KV-cache offload, which moves stored attention state out of scarce accelerator memory. The company claims these changes produced more than a threefold end-to-end speedup for long-horizon reinforcement learning.

A correction puts Luna ahead

The Hermes Index evaluates models across Hermes Bench, TerminalBench 4, TerminalBench Science, and SkillsBench. Every model runs through the same Hermes Agent harness, with reasoning effort set to high when supported. Results use pass@1, meaning each task receives one scored attempt, and the final index averages the four suite scores.

Nous initially reported a 33.89 score for Step 5 Preview, matching GPT-6 Luna. It later corrected the result to 32.83. Selected leaderboard results are shown below:

Model Hermes Index Mean cost per task
Claude Opus 5.5 63.31 $4.99
GPT-6 Astra 56.25 $11.61
Claude Sonnet 5.5 53.14 $2.82
GLM 5.3 Flash 34.95 Not cited
GPT-6 Luna 33.89 Not cited
Step 5 Preview 32.83 Not cited

The leaderboard controls the harness and task suites, which makes its results useful for agent comparisons within that setup. It does not establish performance for every repository, toolchain, or production environment, so teams still need workload-specific tests.

Coding strength carries a verbosity cost

StepFun positions Step 5 Preview for agentic coding and professional knowledge work, including finance. Its published DeepSWE v1.1 comparison places the model near several competing coding systems:

Model DeepSWE v1.1 score
GPT-6 Astra Max 74.1
Claude Opus 5 Max 74.0
Step 5 Preview, high reasoning 67.7
Kimi K3 Max 67.5
GLM-5.3 Max 66.9

These are vendor-reported results, and the model names and evaluation settings differ from those used by the Hermes Index. They support StepFun’s coding claims but do not provide a direct comparison with the independent leaderboard. The supplied figures also contain no separate finance benchmark.

Output volume complicates the model’s low token price. Step 5 Preview generated 160 million output tokens during its Hermes Index run, compared with a 92 million median across evaluated models. Long reasoning traces can therefore offset part of the per-token savings. Production evaluations should track cost per successful task alongside token price, latency, and completion rate.

Three workloads for the free window

  1. Repository-scale coding agents. Test multi-file refactors, issue reproduction, patch generation, test execution, and recovery from failed tool calls. The 1M-token limit provides room for large repositories, while the trial can reveal how reliably the model retrieves relevant details from that context.
  2. Multimodal document pipelines. Evaluate workflows that combine reports, diagrams, screenshots, and video. Native support for text, images, and video can reduce the need for separate extraction and vision services.
  3. Self-hosting feasibility. StepFun says the model weights are scheduled for release on October 15, 2026. A raw BF16 checkpoint for 600 billion parameters would require roughly 1.2 TB before runtime overhead, KV cache, and redundancy. Sparse activation lowers computation per token, while quantization and offloading will determine practical hardware requirements. Licensing terms and checkpoint support will also need review when the weights arrive.

StepFun’s pricing and planned weights release follow a broader push by Chinese labs such as Kimi, GLM, and DeepSeek toward lower API costs and more deployable models. Teams using the free window can produce a useful comparison by fixing repository snapshots, prompts, tool permissions, reasoning settings, and token budgets across models. Recording task completion, wall-clock latency, tool failures, and output-token volume will show whether Step 5 Preview’s low rates translate into lower production cost per completed task.

Trending
  • No trending articles

Comments

avatar

Next Reads