News Weekly
LV 10 XP
0% read
Your progress · 0/5 chapters
About 6 min total
ModelsISSUE #2 · STORY 10 OF 20Sep 20, 2026CONFIRMED

StepFun launches a low-cost flagship model and promises open weights

StepFun's Step 5 Preview scored 44 on Artificial Analysis' Intelligence Index at $0.72 per task, the cheapest measured offer in its class. The catch: it is API-only today, with open weights promised for October 15.

Illustration: a floating staircase of clean modular blocks rises across a wide abstract sky, an endless glowing measuring ribbon unspooling beside it — an artistic impression of a model with an active mixture-of…

Read it your way

CHAPTER 1 · THE 60-SECOND VERSIONPicked for Explorers

A fourth Chinese contender

On Sept 20, 2026, Shanghai's StepFun introduced Step 5 Preview, a declared 600-billion-parameter mixture-of-experts model with 27 billion active, and opened a paid API the same day. Artificial Analysis scored it at 44 on its Intelligence Index, matching Kimi K3. StepFun also promised open weights on October 15, a company commitment with no license named yet, not a delivered fact.

The independent scoreAA rates Step 5 Preview at 44, far above the comparable-priced median of 25.
Cheap but not topThe same week's MiMo-V2.6-Pro is both cheaper per token and higher on the index.
Weights are a promiseAs of Sept 23 the weights were not downloadable, and AA still classified the model as proprietary.
Big day oneLaunch included a documented Claude Code integration, video input and a 1 million token context.
Finish this chapter for +15 XP
Flip the switch

The day a new price floor landed

YOU GETA live cheap flagshipStep 5 Preview runs today at $1/$2.70 with 1M context and an AA score of 44.
YOU GETA dated promiseOpen weights now have a public Oct 15 deadline, license still unnamed.
YOU GETAn auditable claimAA's per-run costs make the roughly 1/8-versus-Opus claim checkable, not rhetorical.
Play with the numbers · +10 XP

What your output bill looks like at scale

DRAG THE SLIDER
Step 5 Preview
$5,400
$2.70 per million output tokens, official pricing
Claude Opus 5
$50,000
$25 per million output tokens, per AA pricing cited in the file
You'd save
$44,600
every month

Output includes the model's reasoning trace, and Step 5 Preview is measurably verbose: 160 million output tokens on AA's index run versus an 88 million median.

Your next move · as a Explorer

Wait for the drop

1Check the Oct 15 Hugging Face release for weights, license and config parity.
2Replay the cost math yourself: 5.9x flat versus 7.88x on AA run costs.
3Compare with MiMo-V2.6-Pro, cheaper and higher on the index today.

Switch your reading mode at the top to see a different next move.

Tap to open

Things to keep an eye on

Pop quiz · unlock the Cost Cutter badge

Did it stick?

0/3
When does StepFun promise to release the open weights?+20 XP
What is capped at 64k tokens?+20 XP
Where does the roughly 1/8-of-Opus cost claim hold best?+20 XP
Your call · +5 XP

Should we judge a model launch by promised open weights on a date, or only by what ships today?

Deep dive

The full research, labeled and sourced

CONFIRMED17 sources · 80 min
Story identity
FieldValue
Story IDS10
TitleStepFun introduces Step 5 Preview: 27B-active MoE, 1M context, weights promised October 15
OrganizationStepFun (阶跃星辰, Shanghai; discovery record's "步知" form was not found in any consulted source — all official pages, docs and coverage use 阶跃星辰 / "StepFun")
CategoryModel Release
Event date2026-09-20 (official announcement + live paid API; StepFun's page: "Today, we're introducing Step 5 Preview")
Announcement date2026-09-20 (KuCoin flash timestamp 2026-09-20 02:23 UTC; MarkTechPost article dated 2026-09-20)
Article dates2026-09-20 (MarkTechPost, OrcaRouter, 36kr), 2026-09-21 (eesel, cnboat, BitsMinds, WorldAttention), 2026-09-22 (Pandaily, Toolbit, eesel review)
Alternative datingArtificial Analysis dates the model "released September 18, 2026" (two days before StepFun's own announcement — its first measured/leaked result date; OrcaRouter flags the same). Both dates are inside the window, so eligibility is unaffected
Window check2026-09-20 ∈ [2026-09-18, 2026-09-22] — eligible (also 2026-09-18 per AA, still eligible)
Evidence statusCONFIRMED (official announcement text, official platform docs — model spec + pricing pages fetched in full; Artificial Analysis model page fetched in full; five independent analyses fetched in full; two more secondary recaps)

Corrections / refinements to the discovery record (important):

  • "DeepSWE v1.1 66.4 (High)" — numeric mix-up. The official page's software-engineering row is DeepSWE v1.1: 67.7 (High) (vs Kimi K3 Max 67.5, GLM-5.3 Max 66.9, GPT-6 Astra 74.1, Claude Opus 5 74.0). 66.4 is the FrontierFinance score (vs GLM-5.3 64.1, DeepSeek V4.1 Flash 63.0, Kimi K3 62.6, Opus 5 69.7, GPT-6 Astra 55.0). The discovery record seems to have merged two rows. All are COMPANY CLAIMS anyway (StepFun's own harnesses; Step ran at High effort while rivals ran at Max — flagged by MarkTechPost and BitsMinds).
  • "Roughly 1/8 the task cost of Claude Opus 5" — methodology-dependent, largely corroborated by AA's own index-run costs. Flat token-price math on an identical 1M-input/0.1M-output workload gives ~5.9× (not 8×) (Toolbit's arithmetic, reproduced in labs/S10.md: $1.27 vs $7.50). But Artificial Analysis' published cost to complete the full Intelligence Index v4.3.2 run is $922.84 for Step 5 Preview vs $7,274.74 for Claude Opus 5 — a 7.88× ratio, ≈ 1/8 (OrcaRouter, citing AA's run-cost data at maximum reasoning effort). The company claim is therefore consistent with AA's task-level costs, and the 1/8 line is what the Chinese media relayed (36kr, 财联社, 新浪财经 via WorldAttention); it is still a vendor-framed claim with an undisclosed methodology and should be reported as "≈1/8 on AA's own per-run costs; ~5.9× on flat token arithmetic; treat as a cost ratio of roughly 6–8× depending on methodology."
  • "Text/image/video input" — CONFIRMED per official docs; Artificial Analysis lists only text and image. The StepFun docs page explicitly lists Text, images, and video (video: MP4/QuickTime/Matroska, <128MB, ~5 min recommended). AA's model page says "Supports: text and image." Minor discrepancy (AA may not yet exercise video input), but worth pinning before citing.
  • "Weights promised October 15" — COMPANY COMMITMENT, not a released fact. Official page: "The model will be released with open weights on October 15." As of Sep 22–23: no weights, no model card, no license named. The official HF repo stepfun-ai/Step-5-Preview-BF16 exists but is an empty placeholder (eesel) / unavailable — removed or gated (Toolbit). Direct HF API probe during this research returned HTTP 401 for both that repo and a fabricated path (so the probe alone is inconclusive; the independent descriptions carry the finding). Artificial Analysis still classes the model Proprietary today: FAQ "Is Step 5 Preview open source? No. The model weights are not publicly available." License expectation: StepFun's prior open Flash releases were permissive (Step-3.5-Flash Apache-2.0 per multiple sources; Toolbit and eesel both state the Flash line shipped Apache-2.0; one directory source, LLMReference, lists Step 3.7 Flash as proprietary — minor source conflict — but no license has been named for Step 5 at all).
  • "Par with Kimi K3 Max" — INDEPENDENTLY VERIFIED as a number, with a class caveat. AA Intelligence Index 44 for Step 5 Preview equals Kimi K3 (max)'s 44 (S09 research, AA open-source leaderboard). Caveat: AA currently classes Step 5 Preview as proprietary (no weights yet), while Kimi K3/GLM-5.3/MiMo-V2.6-Pro are downloadable — so "top-three open-weight standing" is numeric only until Oct 15. Ordering in the open-weight family on AA's data as of Sep 22: MiMo-V2.6-Pro 46 > GLM-5.3 (max) 45 > Kimi K3 (max) 44 = Step 5 Preview 44 — i.e., Step 5 ties for third in the open-weight class behind Xiaomi and Z AI.
  • KuCoin headline error: the crypto-aggregator title says "100K Token Context" while its own body says 1M — headline does not survive contact with the article; an example of why crypto-news recaps need cross-checking.
  • Rank/speed figures drift: AA's live page (research date 2026-09-23) shows #32/212 on the Intelligence Index, 83.6 t/s, $0.72 per index task, 160M output tokens; launch-week recaps cited #26/202 and ~99.8–100 t/s and a $0.71 task cost. AA index-vs-price medians on the live page: intelligence median 25, input median $2.00, output median $10.00, speed median 74 t/s, verbosity median 88M. Capture dates when citing.

✓

What happened?

🎓 For Explorer

On September 20, 2026, StepFun (阶跃星辰, Shanghai) introduced Step 5 Preview — its new flagship model for agentic work, framed as "advancing the Pareto frontier": frontier-tier performance at substantially lower cost, with particular strength in software engineering and finance. The launch skipped the Step 4.x line entirely (Step-3.7-Flash → Step 5), which the company presents as a deliberate step-change (FACT; noted by OrcaRouter and BitsMinds).

What shipped (FACT, official docs + announcement):

  • Sparse MoE, 600B total / 27B active per token (~4.5% of weights per forward pass) — config is vendor-declared and cannot be independently checked until weights exist; AA lists "600 billion parameters."
  • 1M-token context window (max input); max output 64k tokens — the 1M/64k asymmetry was widely missed in headlines.
  • Input: text, images (up to 60 per request), video (MP4/QuickTime/Matroska <128MB, ~5 min recommended); output: text only (FACT, docs; AA lists text+image).
  • Live same-day paid API + StepFun AI Studio, model ID step-5-preview; streaming, tool calling, JSON Mode/Schema, prompt caching, selectable reasoning_effort (low/medium/high), an OpenAI-compatible Chat Completions API + Messages API, and a documented Claude Code integration ("Step Plan", 1M context) (FACT, docs).
  • Pricing (FACT, official pricing page + AA): $1.00/M input (cache miss), $0.05/M input (cache hit — a 95% discount), $2.70/M output (reasoning trace included). Rate tiers V0–V4 by cumulative top-up (5 → 130 concurrency; 100 → 2,600 RPM; 0.5M → 13M TPM). One API provider at launch (StepFun; also surfaced via Vercel AI Gateway).
  • Weights promised for October 15 (official page; countdown on homepage per eesel) — a commitment, with no license named and no model card published as of this research (Sep 23).

Independent measurement (INDEPENDENTLY VERIFIED, Artificial Analysis, page fetched 2026-09-23): Intelligence Index v4.3.2 = 44 (#32/212 models; median of comparable-priced reasoning models: 25), well above class median; $0.72 cost per Intelligence Index task; 160M output tokens on the index run vs 88M median (very verbose — 4/4 verbosity units); output 83.6 t/s (above class median 74); TTFT 3.17 s (better than median 3.83 s); 95% cache discount; 1M context (~1,500 A4 pages); 1 API provider; classified Proprietary (weights not public). AA dates the model to September 18, 2026.

Company-published performance claims (COMPANY CLAIM until independently replicated): DeepSWE v1.1 67.7 (High) vs Kimi K3 67.5 / GLM-5.3 66.9 / GPT-6 Astra 74.1 / Claude Opus 5 74.0; StepCodeBench (in-house, 553 repos, 33 languages) 49.0 vs Kimi 43.9 / GLM 40.2; ProgramBench 80.5; Terminal-Bench v2.1 85.0; BrowseComp 88.7; FrontierFinance 66.4 (edges all open peers; beats GPT-6 Astra 55.0); DRACO 83.3; GPQA Diamond 93.5%; HLE 46.5%; "leads open-weight models" across six StepCodeBench categories; ~70% of expert evaluators judged it capable of autonomously solving moderately high-complexity coding tasks; a single agent action coordinated 950 web fetches; 65% lower cost than GLM-5.3 and Kimi K3 at comparable intelligence (launch chart).

Two 24-hour agent experiments (COMPANY CLAIM — demos, not benchmark evals): (1) autonomous MLA GPU-kernel optimization on an H100 head-dim 512, batch 1/64 heads/8,192 tokens — reached 508 TFLOP/s after ~22 hours, compared with Claude Opus 5's 493 TFLOP/s in the same task setup (official page, MarkTechPost, KuCoin, OrcaRouter); (2) autonomous post-training of Qwen3-30B-A3B on AIME24 via an API annotator — 53.3% → 60% accuracy, matching Claude Opus 5's result with fewer annotator tokens.


Δ

What changed?

  • The Chinese open-weight price/intelligence war got a fourth major entrant in a week. Xiaomi MiMo-V2.6-Pro (46, $0.435/$0.87) and GLM-5.3 (45) opened the week; Step 5 Preview lands at 44 with $1.00/$2.70 — cheaper per token than GLM-5.3 ($1.40/$4.40) and Kimi K3 ($3.00/$15.00), while MiMo remains cheaper and higher-indexed. (INDEPENDENT EVIDENCE on index/price rows; the strategic framing is INTERPRETATION.)
  • "State-of-the-art open-weight" leadership now comes with a deferred-weights caveat. StepFun claims a top-three open-source global ranking, but as of Sep 23 the weights are not downloadable; AA classifies the model as proprietary. The 44 score is real; the "open" part is an October promise. (INDEPENDENT EVIDENCE + COMPANY CLAIM.)
  • The per-task cost debate became checkable. AA publishes real cost-to-run data; the 1/8-of-Opus-5 claim approximately survives contact with AA's numbers (7.88× on index-run cost) even though flat token arithmetic says ~5.9×. The verbosity tax is now quantified: 160M output tokens on the index run — at $2.70/M that is $432 of output alone; the same verbosity at Opus 5's $25/M would be $4,000. (INDEPENDENT EVIDENCE + arithmetic; see labs/S10.md.)
  • 1M-in / 64k-out context asymmetry entered the public spec sheet. A 1M-token context is a headline feature, but maximum output is 64k — a constraint that alters long-form generation planning and was absent from most launch coverage. (FACT, docs.)
  • Agentic-model economics shifted to the cache. The 95% cache-hit discount makes re-reading a ~1M-token repo over an agent loop dramatically cheaper: first read $1.00 + 9 cached reads at $0.05 = $1.45 vs $10 uncached. (FACT, arithmetic; INDEPENDENT in docs pricing.)

↔

Before → Change → After

🎓 For Explorer

Before (as of mid-September 2026):

  • StepFun's latest open release was Step-3.5/3.7-Flash (~200B/11B-active class, 256k context, $0.20/$1.15, Apache-2.0 per most accounts); the Step-4.x line never shipped.
  • Open-weight family on AA's data: MiMo-V2.6-Pro 46, GLM-5.3 (max) 45, Kimi K3 (max) 44 — all downloadable. Only Kimi K3 sat at 44 in the open tier.
  • Closed frontier pricing: Claude Opus 5 $5/$25 (AA index 51), GPT-6 Astra $10/$50 (index 53); Grok 4.7 $2/$6 the same week.
  • A 600B-parameter API at $1/$2.70 with 1M context did not exist in this price band; no StepFun flagship-class model was available at any price.

Change (Sep 20, 2026):

  • Step 5 Preview announced and API opened same-day: 600B/27B sparse MoE, 1M context, text/image/video input, $1.00/$0.05/$2.70, 64k max output, reasoning-effort dial, Claude Code integration, AA index 44 with a $0.72/task cost.
  • Weights announced for Oct 15 (promise only; placeholder HF repo; no license).
  • Independent record updated: AA page live, index 44 (#32/212), verbosity 160M tokens flagged, TB4.0 measured at 33.3 (vs the vendor's Terminal-Bench v2.1 85.0 headline).

After:

  • The open-weight price-performance frontier has a new price-setter at the "capable mid-tier" end; the top of the open field (46/45/44) is now Chinese-lab-dominated, with StepFun staking the finance/agentic-coding vertical at 44.
  • Every enterprise and developer model-selection table now carries "Step 5 Preview (44, $1/$2.70, API-only, weights 10/15 pending license)" as a row to watch — and MiMo-V2.6-Pro as the both-cheaper-and-smarter open alternative.
  • The "Pareto frontier" framing that StepFun names is now measurable: AA's cost-per-index-task is published, so "cheaper per achieved intelligence" claims are checkable rather than rhetorical. (INTERPRETATION)

⚙

How it works

  • Sparse MoE (FACT as declared spec; architecture details are vendor description until weights land): 600B total parameters, 27B active per token — a router activates ~4.5% of weights per forward pass, which is the structural basis for the low token price. 27B active is heavier than DeepSeek's/GLM's "flash" tiers (a Hacker News criticism), so the efficiency argument is "intelligence-per-active-parameter," not raw active count. (COMPANY CLAIM + community critique via eesel.)
  • 92-layer narrow-deep stack (per Pandaily, repeated by MarkTechPost): StepFun's stated rationale is that deeper stacks give longer paths for implicit multi-hop reasoning during long prefill — relevant when an agent searches, executes code, and reads tool returns in one long context. (COMPANY CLAIM as to design intent.)
  • Training/systems claims (COMPANY CLAIM): on-policy, long-horizon RL; bit-wise train-inference alignment across MoE routing; MTP-3 speculative decoding; FP8 MoE; KV-cache offload; reported >3× end-to-end speedup for long-horizon RL. No technical report or paper is published yet (noted by eesel).
  • Reasoning control (FACT): reasoning_effort = low/medium/high; output tokens include the full reasoning trace, so effort is a direct cost dial on an output-priced API.
  • Serving stack (FACT, docs): OpenAI-compatible Chat Completions + Messages APIs, streaming, tool calling, JSON Mode/Schema, prompt caching, image-by-URL-or-Base64 (up to 60 images, low/high detail), video via URL/Base64/stepfile:// Files API, and a Claude Code ("Step Plan") integration exposing the 1M context. Retrieval/tools are provided by the integrating application, not the model itself (docs are explicit on this).
  • Agentic positioning (COMPANY CLAIM): the 24-hour H100 MLA-kernel experiment and the Qwen3-30B-A3B AIME24 post-training loop are demonstrations of long-horizon autonomy: the model edits, runs, measures, and iterates on its own for 24 hours.

!

Why it matters

🎓 For Explorer
  1. The price-per-intelligence curve moved again — and this time the claim is checkable. AA independently measures 44 at $0.72 per index task (median for comparable-priced reasoning models: 25 index). The 1/8-vs-Opus-5 claim approximately holds on AA's own run-cost data (7.88×) — a rare case where a launch cost claim survives contact with an independent benchmark's bill. (INDEPENDENT EVIDENCE with the 5.9× flat-price caveat.)
  2. It completes a four-way Chinese open-weight-price plate (MiMo 46, GLM 45, Step 5 = Kimi 44) in one window. StepFun's differentiation is not raw open-class price-performance (MiMo is cheaper and smarter); it is the finance/coding agentic vertical, multimodal input, 1M context, and day-one Claude Code integration at $1/$2.70. (FACT on rows; INTERPRETATION on differentiation.)
  3. The "open weights" promise is now a dated commitment with real governance weight. October 15 is a public deadline: license, config parity (600B/27B), and BF16 artifacts are all checkable items. Enterprises planning budgets around it should treat it as a promise with a strong permissive precedent (Apache-2.0 Flash line) and a non-zero restrictive-license risk. (COMPANY COMMITMENT + INTERPRETATION.)
  4. Verbosity is now a first-class economic variable in model selection. AA's 160M-token index run (vs 88M median) is a 4/4-verbosity reading: at $2.70/M output, verbosity eats a meaningful share of the price advantage. Per-task cost, not per-token price, is the correct decision metric (and AA publishes it). (INDEPENDENT EVIDENCE.)
  5. A Chinese lab demonstrated 24-hour autonomous systems work as a product demo (kernel tuning to 508 TFLOP/s; AIME24 post-training to 60%). These are company demos, not independent evals — but they set an expectation bar for agentic long-horizon claims across the industry. (COMPANY CLAIM, direction is INTERPRETATION.)

✦

What became possible?

🎓 For Explorer
  • Running a frontier-sized (600B-declared) reasoning model via API at $1/$2.70 with a 95% cache discount — the cheapest "44-class" reasoning offer yet measured on AA, live on launch day.
  • Long-context agentic coding and finance workflows that were previously uneconomical: re-reading ~1M-token contexts repeatedly is the pattern the 95% cache discount targets (first read $1.00, each cached re-read $0.05).
  • Claude Code-native evaluation with zero toolchain migration: the documented "Step Plan" integration means teams can A/B a second model inside their existing Claude Code flow this week.
  • Visual-input agentic work at low cost: up to 60 images/request and short video understanding for screenshots, charts, and UI flows — at text-tier pricing with per-image token accounting.
  • Self-hosting — soon, conditionally. Once Oct 15 delivers weights + license + config, organizations can plan self-hosted deployment (600B BF16 ≈ 1.2 TB before KV cache — multi-node, not casual; MarkTechPost/BitsMinds arithmetic). Not possible today.

◎

Implications

Technical

  • Benchmark versioning traps: the vendor headlines Terminal-Bench v2.1 85.0 while AA's Terminal-Bench 4.0 reads 33.3; eesel's independent table also shows TB v4 at 33.3 (Kimi 12.6, GLM 41.9, Opus 5 52.3). Different suites measure different things — vendor-vs-independent comparisons must pin the benchmark version.
  • The 1M-in/64k-out asymmetry is a real design constraint for long-document assertions and codegen pipelines: maximum input is 1M, but single-completion output caps at 64k — plan chunked/windowed generation accordingly. (FACT, docs.)
  • Reasoning-effort dialing is a token-economics control: output is billed including the reasoning trace, so low/medium/high is a direct cost lever; workflows should tune per task complexity (FACT; cost impact is arithmetic).
  • Cache design (95% hit discount) rewards stable prefixes — fixed system prompts, session reuse, repository snapshots — the optimal prompting pattern for agent loops. (FACT on pricing; the pattern guidance is derived.)
  • No independent check of the 600B/27B configuration is possible yet — AA lists "600 billion parameters" (vendor-supplied); verification of MoE topology, layer count, expert count requires the weight drop. The 92-layer narrow-deep claim and MTP-3/FP8/KV-offload stack are vendor descriptions until the card/technical report ships.
  • Speed figures drift: launch-week captures ~99.8–100 t/s; AA's live page now reads 83.6 t/s — measurement-window sensitivity (same pattern as MiMo-V2.6 in S09). TTFT 3.17 s at 1M prefill is a long-context serving data point to monitor.

Developer

  • Add step-5-preview to the routing matrix this week (API-only): index 44, $1/$2.70, $0.05 cached, 95% cache discount, finite rate tiers (V0: 5 concurrent / 100 RPM / 0.5M TPM until you top up). It is a cheap, live pilot candidate for coding agents and finance research — not a production dependency until Oct 15 resolves license and weight availability.
  • Budget per-task, not per-token. Use AA's $0.72/index-task figure as the reference; adjust for your own verbosity profile — this model is measurably verbose (4/4 on AA), and output includes reasoning traces.
  • Exploit the cache: stable system prompts, conversation reuse, and repo snapshots turn the 95% hit discount into the real economic story for long agent loops.
  • Reasons to hold off building product on it: 64k max output; preview pricing may not survive the "Preview" suffix; no weights/license yet; official HF repo is a placeholder/gated; third-party "1.21TB shard" repos with license:other tags are not official — do not treat them as weights (Toolbit).
  • Claude Code integration is live and documented ("Step Plan", 1M context) — lowest-friction way to test without leaving an existing agentic toolchain.
  • Governance checks for procurement: license unknown until Oct 15 (Flash-line precedent is Apache-2.0); data-residency applies (China-based vendor API); rate-limit tiers require cash top-ups ($1,500 cumulative for V4-class concurrency); model does not itself access external services (docs are explicit — tools come from the integrating app).

Enterprise

  • A new benchmarked reference price for "capable mid-tier" agentic AI: enterprises negotiating agentic-coding and finance workloads now have AA data showing ~44-index capability at $0.72/task — a strong lever in vendor conversations regardless of whether StepFun is adopted.
  • Candidate for finance-heavy, vision-including, long-context agent workflows (FrontierFinance 66.4 leads open peers and beats GPT-6 Astra 55.0 on vendor data; FinStepBench DeepResearch 55.8 second only to Opus 5 59.1) — with the required caveat that these are self-reported.
  • Deployment posture today vs Oct 15: today = API pilot only (one provider, preview status, no SLA-level redundancy — OrcaRouter notes preview endpoints are capacity-constrained); Oct 15 = self-host option if license is permissive; if restrictive, MiMo-V2.6-Pro (MIT) and GLM-5.3 remain the deployable open alternatives. Model-selection decisions should be explicitly date-staged.
  • Cost-simulation benefit: the 95% cache discount makes the economics of repo-scale agent loops modelable in advance (see labs/S10.md arithmetic) — useful for CFO/architecture sign-off before any commit.
  • Risk posture to document: Chinese-lab API (data residency, export/regulatory context), preview pricing volatility, benchmark-version gaps (TB2.1 85.0 vs TB4.0 33.3), and the unverified 600B config.

Strategic

  • StepFun's play is vertical + price, not top-slot: best-of-the-cheaper-tier positioning, explicitly "Pareto frontier" framing rather than "we beat GPT-6 Astra"; the finance-and-coding agentic wedge plus multimodal input plus day-one Claude Code integration is the differentiation against a MiMo that is cheaper and higher-indexed in the same week. (INTERPRETATION grounded in the launch materials.)
  • The Oct 15 open-weights date is the strategic payload. A permissive Apache-2.0-class release would add a third/fourth serious self-hostable agentic option from China in a single quarter (MiMo MIT, GLM, Step) and further compress what closed vendors can charge for mid-tier agentic work. A restrictive or delayed release would cost StepFun credibility it explicitly staked with a public countdown. (PREDICTION + INTERPRETATION.)
  • Cost per achieved intelligence is becoming the battleground metric — AA's per-task costs let buyers compare across vendors on one methodology; "1/8 of Opus 5" is exactly the kind of claim the market can now audit, and vendors know it (StepFun bet the launch on it). (INTERPRETATION.)
  • The WSJ-reported Hong Kong IPO (~$500M raise at up to a $12B valuation — press-reported, company-unconfirmed) frames the stakes: this launch cadence and price war is plausibly pre-IPO positioning for the Chinese labs (RUMOR/secondary; treat accordingly).
  • Expect a Kimi and Qwen response within weeks (36kr notes "the answer for Kimi is almost out in the open") — the Chinese release cadence (Step, Qwen, Zhipu, Kimi) is now a weekly drumbeat against which Western closed labs must price. (PREDICTION, partly grounded in 36kr commentary.)

⚠

Risks & limitations

Risks
  • The headline benchmarks are self-graded (own harnesses; StepCodeBench is in-house; Step ran at High while rivals ran at Max; no technical report) — independent re-runs pending; treat DeepSWE/FrontierFinance/StepCodeBench rows as COMPANY CLAIM. (Risks of over-commitment on unverified rows.)
  • Open-weights promise may slip or land restrictive. No license named, HF repo placeholder/gated, weights in BF16 (~1.2 TB) with no serving recipes yet; precedent is permissive (Apache-2.0 Flash line) but not guaranteed. Until Oct 15 this model is a rental with a dated promise (Toolbit framing).
  • Verbosity tax: AA 4/4 verbosity (160M vs 88M median tokens) inflates real spend vs terser models at the same index on output-priced APIs — per-task costs will vary by workload; $0.72 is AA-methodology-specific.
  • Preview volatility: pricing may not survive the Preview suffix; the HF repo already changed state (placeholder → removed/gated per Toolbit's observation); preview endpoints are capacity-constrained and single-provider at launch.
  • 1M-context reality at production scale: TTFT 3.17 s and long-prefill pressure; no published degradation data for full-1M operations; max output 64k caps single-shot generation.
  • Supply-chain hygiene: third-party unverified "Step-5-Preview-BF16" shard repos exist under license:other — provenance-confusion risk for teams downloading "weights" before the official drop.
  • Geopolitical/data-governance exposure: Chinese-lab API with data flowing to Shanghai infrastructure; relevant for regulated enterprises regardless of model quality.

Limitations
  • Weights are not out. The "open-weight" standing is numeric (44 in the open-weight family comparison) but not physical; AA classes the model Proprietary today. All architecture details (600B/27B, 92-layer narrow-deep, MTP-3, FP8) are vendor-declared and unverifiable until the release.
  • All flagship benchmark rows are vendor-reported; only AA's composite index, price, speed, latency, and verbosity are independent. Terminal-Bench claims are version-sensitive (v2.1 85.0 vendor vs v4.0 33.3 independent).
  • "1/8 of Claude Opus 5" does not reproduce under flat token-price arithmetic (~5.9×); it only approximately holds on AA's methodology-weighted run costs (7.88×). It is a marketing-grade figure with an undisclosed methodology, consistent with — but not proven by — independent data.
  • VIDEO input is documented but AA measures text+image only, and no independent multimodal evals (MMMU-Pro etc. rows visible on AA) were extracted for this story.
  • Not frontier-top: 44 vs GPT-6 Astra 53 and Opus 5 51 on the index; on DeepSWE the vendor itself shows 67.7 vs 74.0/74.1 for the closed flagships.
  • 64k max output limits very long single completions; one API provider and tiered rate limits cap throughput without top-ups.
  • No technical report, paper, or model card published as of Sep 23, 2026.

?

Open questions

  1. What exactly drops October 15 — BF16 weights only, or also a model card, license, serving recipes, and the technical report? Does the released config match the declared 600B/27B and the 92-layer narrow-deep spec?
  2. Which license? Apache-2.0 (Flash-line precedent) vs MIT vs something restrictive — and does license language allow commercial fine-tuning/distillation?
  3. Do independent re-runs confirm DeepSWE 67.7, FrontierFinance 66.4, and the six-category StepCodeBench leadership — with equal effort settings (Step ran at High, rivals at Max)?
  4. Does the ~1/8-vs-Opus-5 per-task ratio hold on real enterprise workloads, or is it an index-methodology artifact (AA's 10-eval composite with its own weights)?
  5. Why does AA date the model September 18 while StepFun announced September 20 — early test access, or a leaked-tag artifact (Toolbit's "test-tag leak")?
  6. Will the 508 TFLOP/s MLA-kernel result and the AIME24 53.3%→60% post-training loop reproduce independently, and are those demos closer to marketing or to capability?
  7. Does the 95% cache discount survive heavy 1M-context multi-turn usage in production, and what do real TTFT/latency numbers look like at full context under load across providers?
  8. What is the roadmap after the "Preview" suffix drops — a non-reasoning variant (AA notes one "may exist"), a Step-5-Flash-class smaller tier, and pricing stability?
  9. Will the WSJ-reported Hong Kong IPO (~$500M at up to $12B valuation) be confirmed, and does the open-weights strategy survive the positioning it implies?

↗

What happens next?

🎓 For Explorer
  • Days (to Sep 25): AA's live page keeps updating speed/TTFT/rank as more index runs land (83.6 t/s today vs 99.8 t/s in launch-week captures); first community re-runs of DeepSWE/StepCodeBench; Hacker News discussion continues around verbosity and the demo's thinking traces.
  • Weeks (to Oct 15): the countdown runs its course; watch for a Kimi K response (36kr anticipates it "almost out in the open") and Qwen/GLM pricing replies to the $1/$2.70 line; expect Chinese media recaps to lock in the "top three open-source at 1/8 of Opus 5" narrative (already the dominant framing per WorldAttention).
  • Oct 15 and after: the decisive event — weights live or not; license named or not; config parity (600B/27B, 92-layer) or not. Outcomes: permissive license → Step 5 joins MiMo/GLM in the self-hostable agentic tier and pressure on mid-tier closed pricing intensifies; slip/restrictive license → MiMo-V2.6-Pro and GLM-5.3 consolidate as the deployable open alternatives and StepFun's credibility takes a hit it explicitly staked. Meanwhile the WSJ-reported HK IPO story either confirms (strategic context: this is pre-IPO positioning) or dies. (PREDICTION)

★

Editorial takeaway

🎓 For Explorer

Step 5 Preview is the week's most honest cost story and its most conditional open-weights story at once. The verified core is real: a 44 on Artificial Analysis' Intelligence Index at $0.72 per index task — the cheapest independently-measured "44-class" reasoning offer on the board, with a 95% cache discount that makes long agent loops genuinely cheap, plus a live $1/$2.70 API, 1M context, vision/video input, and a documented Claude Code integration. The honest caveats are equally real: every headline benchmark (DeepSWE 67.7, FrontierFinance 66.4, StepCodeBench 49.0) is self-graded with Step running at High against rivals at Max — and the "top-three open-source at 1/8 of Claude Opus 5" claims point at an open-weight standing that is, as of today, a countdown and a placeholder repo, not weights. The 1/8 claim survives contact with AA's own run-cost data (7.88×) but not flat token arithmetic (~5.9×) — a perfect teaching example of per-task vs per-token economics, made worse by a verified 4/4 verbosity reading. The through-line for the weekly: the frontier-for-less pricing war has a fourth Chinese entrant in one window (MiMo 46 / GLM 45 / Step 5 44 / Kimi 44), and October 15 is now a public test of whether "we will open-source the model" on a specific date is a commitment the market can bank on.

Illustration: frame: a staircase of clean modular blocks rises through an abstract sky, an unspooling glowing ribbon suggesting an immense context length — an artistic impression of a previewed expert-mix model.
⌘

Lab: VERIFY

Step 1 — Probe the official Hugging Face repo state (HTTP status probe, executed 2026-09-23)
for r in "stepfun-ai/Step-5-Preview-BF16" "stepfun-ai/definitely-not-a-real-repo-xyz" "stepfun-ai/Step-3.7-Flash" "stepfun-ai/Step-3.5-Flash"; do
  code=$(curl -s -o /dev/null -w "%{http_code}" --max-time 20 "https://huggingface.co/api/models/$r"); echo "$r -> $code"
done

Actual output:

stepfun-ai/Step-5-Preview-BF16 -> 401
stepfun-ai/definitely-not-a-real-repo-xyz -> 401
stepfun-ai/Step-3.7-Flash -> 200
stepfun-ai/Step-3.5-Flash -> 200

Findings — and an honest caveat: HTTP 401 on the official Step-5-Preview-BF16 repo cannot alone prove "gated", because a fabricated path also returned 401 in this environment (auth-wall behavior), while the older public Step-Flash repos return 200 (public). The probe is therefore inconclusive on its own. The operative finding — no official weights, model card, or license are downloadable as of Sep 23 — is established by triangulation: eesel ("empty placeholder with no model card, weights, or license"), Toolbit ("unavailable (removed or gated) as of September 22; it was a placeholder at launch"), and Artificial Analysis' live classification: "Is Step 5 Preview open source? No. The model weights are not publicly available." Lesson: an HTTP code alone is not evidence; pair it with two independent descriptions.

Step 2 — Verify the official spec and pricing from the provider's own docs (both pages fetched in full 2026-09-23)

From https://platform.stepfun.ai/docs/en/guides/models/step-5-preview and https://platform.stepfun.ai/docs/en/guides/pricing/details:

Claim in circulationOfficial docs sayVerdict
600B total / 27B active MoEDeclared in announcement (docs reference the model; spec config is vendor-stated, unverifiable until weights)PARTIAL — vendor-figure
1M-token contextContext window 1M; max input 1M; max output 64kVERIFIED (with the 64k output nuance most headlines missed)
Text/image/video inputText, images (≤60/request), video (MP4/QuickTime/Matroska, <128MB, ~5 min recommended)VERIFIED per docs (AA lists text+image only)
$1.00 / $2.70 per MInput cache-miss $1.00, cache-hit $0.05, output $2.70 (trace included)VERIFIED
95% cache discount$0.05 vs $1.00 = 95%VERIFIED (arithmetic on official rows)
Reasoning effort selectablereasoning_effort: low / medium / highVERIFIED
Open weights Oct 15Announcement + homepage countdown; docs do not list weightsPARTIAL — commitment only
Step 3 — Reproduce the cost claims (pure arithmetic on published prices; script executed 2026-09-23)
# step5_cost.py — run with plain python3, no deps
models = [
    # name, in_$/M, out_$/M, cache_$/M
    ("Step 5 Preview",   1.00, 2.70, 0.05),
    ("Claude Opus 5",    5.00, 25.00, 0.50),
    ("GPT-6 Astra",     10.00, 50.00, None),
    ("MiMo-V2.6-Pro",    0.435, 0.87, 0.00435),  # 99% cache discount
    ("GLM-5.3 (open)",   1.40, 4.40, None),
    ("Kimi K3 (open)",   3.00, 15.00, None),
]
work_in, work_out = 1_000_000, 100_000
print(f"{'model':<16}{'flat-workload $':>16}{'ratio vs Opus5':>16}")
opus5_cost = None
for name, p_in, p_out, p_cache in models:
    cost = work_in/1e6*p_in + work_out/1e6*p_out
    if name == "Claude Opus 5": opus5_cost = cost
    ratio = cost/opus5_cost if opus5_cost else float('nan')
    print(f"{name:<16}{cost:>16.2f}{ratio:>16.2f}")

Actual output:

model            flat-workload $  ratio vs Opus5
Step 5 Preview              1.27             nan
Claude Opus 5               7.50            1.00
GPT-6 Astra                15.00            2.00
MiMo-V2.6-Pro               0.52            0.07
GLM-5.3 (open)              1.84            0.25
Kimi K3 (open)              4.50            0.60

Agent-loop cache scenario (same script):

Step 5 agent loop (1M ctx, 10 reads): first read $1.00 + 9 cached reads $0.45 = $1.45 (vs $10.00 uncached)
AA index-run cost: Step5 $922.84 vs Opus5 $7274.74 -> ratio 7.88x ~= 1/7.88
Step 5 output-token cost on index alone: 160M x $2.70/M = $432

Findings — this is the story's most interesting numerics:

  • Flat token-price math on an identical 1M-input + 0.1M-output workload: Step 5 is ~5.9× cheaper than Claude Opus 5 (not 8×). The vendor's "1/8 of Opus 5" claim does NOT reproduce on per-token arithmetic alone.
  • On Artificial Analysis' actual index-run costs (v4.3.2): Step 5 = $922.84, Opus 5 = $7,274.74 → 7.88× ≈ 1/8. The claim is consistent with AA's task-level costs, because Step 5's extreme verbosity (160M output tokens on the index vs 88M median) eats into its per-token advantage — but its output price is so low that the full run still costs ~1/8. The honest one-line: the "1/8" figure is methodology-dependent — confirm at ~5.9× on flat token math and ~7.88× on AA's methodology-weighted run costs; never quote it as a bare fact.
  • The cache is the practical win for agent loops: a 10-read spreadsheet of a 1M-token repo costs $1.45 vs $10 uncached.
Step 4 — COMPARE: where Step 5 Preview sits (AA data, research date 2026-09-23; launch-week recaps in brackets)
ModelAA Index (v4.3.2)In $/MOut $/MCache discountOpen weights today?
GPT-6 Astra5310.0050.00—No
Claude Opus 5515.0025.0090%No
MiMo-V2.6-Pro460.4350.8799%Yes (MIT)
GLM-5.3 (max)451.404.40—Yes
Kimi K3 (max)443.0015.00—Yes
Step 5 Preview441.002.7095%No — promised Oct 15, no license named

Reading of the table (INTERPRETATION): Step 5 Preview ties Kimi K3 at 44 and undercuts GLM-5.3 and Kimi K3 on price per token, but Xiaomi's MiMo-V2.6-Pro is simultaneously cheaper AND higher-indexed (46 vs 44) and actually downloadable under MIT — Step 5's real differentiation is the finance/coding agentic positioning, multimodal (video) input, 1M context at $1, and the day-one Claude Code integration, not raw open-class price-performance. And per AA's own current classification, Step 5 Preview is proprietary until Oct 15 actually delivers weights.

Step 5 — Verification checklist (reusable pattern for any "open weights promised" launch)
ClaimVerdictEvidence
Announced Sep 20, 2026 (event date)VERIFIEDOfficial page + KuCoin timestamp (Sep 20 02:23 UTC); AA dates model to Sep 18 — both inside window
600B total / 27B active MoEPARTIALAnnounced spec; unverifiable until weights/config released (AA lists "600 billion parameters")
1M context / 64k max outputVERIFIEDOfficial docs (source of the 1M-in/64k-out nuance)
Text/image/video inputVERIFIED per docsOfficial docs; AA lists text+image (minor discrepancy noted)
Pricing $1.00/$0.05/$2.70VERIFIEDOfficial pricing page + AA model page
AA Intelligence Index 44, $0.72/task, 160M tokens (verbose)VERIFIED (independent)AA model page fetched in full (#32/212)
"1/8 of Claude Opus 5 per-task cost"METHODOLOGY-DEPENDENT~5.9× flat token math; ~7.88× on AA run costs; vendor methodology undisclosed
"Top-three open-source globally"NUMERIC44 ties Kimi K3 for #3 in the open-weight family (MiMo 46 > GLM 45 > Kimi/Step 44) — but weights not out yet
DeepSWE 67.7 / FrontierFinance 66.4 / StepCodeBench 49.0NOT INDEPENDENTLY RERUNVendor-reported; own harnesses; Step ran at High vs rivals at Max; no technical report
Kernel demo 508 TFLOP/s vs Opus 5's 493; AIME24 53.3%→60%COMPANY DEMOReported by company + 3 independent outlets; not independent evals
Weights Oct 15, BF16, licensePROMISE ONLYNo weights/card/license as of Sep 23 (eeesel, Toolbit, AA proprietary class; HF probe inconclusive alone)

What this lab does NOT do (honest limits)

  • It does not call the API (no paid account) and does not run the model on real workloads.
  • It treats every vendor benchmark row and both agent demos as COMPANY CLAIM (evidence discipline).
  • AA's $0.72/task and the $922.84/$7,274.74 run-cost ratio are AA-methodology-specific (their 10-eval composite with its own weighting); real workloads will differ (prompt mix, caching behavior, effort levels, verbosity).
  • Leaderboard rows are live data as of 2026-09-23 (speed readouts already drifted: 99.8 → 83.6 t/s between launch-week and the live page).

Result

A one-page, evidence-labelled verification pack proving: (1) no official weights are downloadable today and none will be before Oct 15 (triangulated, with the probe's own limitation recorded); (2) official specs/pricing ($1.00/$0.05/$2.70; 1M-in/64k-out) confirm the launch sheet; (3) the 1/8-of-Opus-5 cost claim is methodology-dependent — ~5.9× flat, ~7.88× on AA run costs, never a bare fact; and (4) Step 5 Preview ties Kimi K3 at 44 in the open-weight family while MiMo-V2.6-Pro is cheaper and smarter and already downloadable — so the Oct 15 weight drop, license, and config parity are the three decisive checkable events.

≡

Research sources

Primary Sources (4)
Primary
Hugging Face — `stepfun-ai/Step-5-Preview-BF16` (official repo, placeholder/gated state) - **URL:** https://huggingface.co/stepfun-ai/Step-5-Preview-BF16** The official weights repository exists as a **placeholder without weights, model card, or license** (eesel: "empty placeholder with no model card, weights, or license as of writing"; Toolbit: "unavailable (removed or gated) as of September 22; it was a placeholder at launch"). Direct probe executed 2026-09-23: HTTP 401 on both the rendered page and the REST API for this repo — note the probe was inconclusive on its own because a fabricated path (`stepfun-ai/definitely-not-a-real-repo-xyz`) also returned 401, while the older public repos `stepfun-ai/Step-3.7-Flash` and `stepfun-ai/Step-3.5-Flash` returned 200 (see labs/S10.md Step 1). The operative fact for the story — no official weights downloadable as of research date — is cross-confirmed by eesel, Toolbit, and Artificial Analysis' "Proprietary / weights not publicly available" classification. — ** FACT (no weights/no license as of Sep 23 — the basis for "weights promised, not released"). ---Date: ** probed 2026-09-23 via HTTP (page and API); repo referenced at launch (Sep 20)
URL unavailable
Primary
StepFun Open Platform — official pricing and rate limits page (fetched in full) - **URL:** https://platform.stepfun.ai/docs/en/guides/pricing/details** Official `step-5-preview` pricing per 1M tokens: **input (cache miss) $1.00, input (cache hit) $0.05, output $2.70** (output includes the reasoning process and final answer; cache-miss price includes writing new content to cache); comparison rows `step-3.7-flash` $0.20/$0.04/$1.15; tiered rate limits V0–V4 by cumulative cash top-up (V0 under $15: 5 concurrency / 100 RPM / 500k TPM; V4 $1,500+: 130 / 2,600 / 13M); enterprise tiers via sales; image input billed as tokens (default 169 tokens/image). — ** FACT (pricing and rate-limit structure — the anchor for all cost arithmetic in this research).Date: ** consulted 2026-09-23 (live docs; pricing as published around Sep 20–22, 2026)
URL unavailable
Primary
StepFun Open Platform — official model documentation: "Step 5 Preview" (spec page, fetched in full) - **URL:** https://platform.stepfun.ai/docs/en/guides/models/step-5-preview** Canonical spec sheet: model ID `step-5-preview`; context window **1M tokens**; max input **1M tokens**; **max output 64k tokens**; input types text/images/video; output text; reasoning effort `low/medium/high`; streaming, tool calling, JSON Mode/JSON Schema, prompt caching; image input via URL/Base64, JPG/JPEG/PNG/WebP/static GIF, **up to 60 images per request**, detail level low/high; video via URL/Base64/`stepfile://` Files API, MP4/QuickTime/Matroska, individual files under 128 MB, ~5 min recommended; explicit statement that tools/search/code execution are provided by the integrating application, not the model; model billed on actual token usage; Claude Code integration ("Step Plan", 1M context). — ** FACT (the authoritative spec — resolves the 1M-in/64k-out asymmetry and the text/image/video input question).Date: ** consulted 2026-09-23 (live docs; model launched 2026-09-20)
URL unavailable
Primary
StepFun — official announcement: "Step 5 Preview: Advancing the Pareto Frontier" - **URL:** https://www.stepfun.com/step-5-preview** Official release narrative ("Today, we're introducing Step 5 Preview… our new flagship model for agentic work"); 600B total / 27B active sparse MoE; 1M-token context + vision input; company benchmark tables (DeepSWE v1.1 **67.7** High vs Kimi K3 67.5 / GLM-5.3 66.9 / GPT-6 Astra 74.1 / Claude Opus 5 74.0; FrontierFinance **66.4** vs GLM-5.3 64.1 / DeepSeek V4.1 Flash 63.0 / Kimi K3 62.6 / Opus 5 69.7 / GPT-6 Astra 55.0; FinStepBench sub-scores LiveSearch 74.5 / CorporateValuation 60.6 / DeepResearch 55.8; GPQA Diamond 93.5%; HLE 46.5%; DRACO 83.3 — with the note that GDPval-AA v2.1 rows are "based on the latest results from Artificial Analysis, as of Sep. 20, 2026"); StepCodeBench (own benchmark: 553 repos, 9 task categories, 20 domains, 33 languages; "leads open-weight models" across feature modification, bug repair, refactoring, documentation generation, performance tuning, code generation); the two 24-hour agent experiments (MLA GPU-kernel optimization on an H100 — head dim 512, batch 1/64 heads/8,192 tokens — and the Qwen3-30B-A3B AIME24 post-training loop, 53.3% → 60%, matching Claude Opus 5 with fewer annotator tokens); Blender/Three.js web-dev case; availability through products and API; the sentence "The model will be released with open weights on October 15." — ** FACT baseline (release date, declared specs, pricing direction, the Oct 15 commitment as a promise) + COMPANY CLAIM (all benchmark and demo performance figures are vendor-published; the Oct 15 weight release is a commitment, not a released fact).Date: ** 2026-09-20 (announcement; page verified during research; full text captured via search-index excerpts — the page shell returns the company name "阶跃星辰" when fetched directly, confirming a JS-rendered SPA)
URL unavailable
Independent Sources (7)
Independent
OrcaRouter — "Step 5 Preview: what StepFun's 600B flagship really ships" - **URL:** https://www.orcarouter.ai/blog/step-5-preview-open-weights** Independent recap: StepFun skipped the Step 4.x line (Step-3.7-Flash → Step 5 "is itself a statement"); AA dates the model September 18, 2026, two days before StepFun's announcement (procurement-calendar warning); vendor-claims-vs-third-party-numbers separation (1/8-of-Opus-5 claim, second place behind Astra/Opus 5 on ALE-CLI/FrontierFinance/DRACO, 24-hour kernel run to 508 TFLOPS vs 493, AIME24 53.3% → 60%, DeepSWE 67.7%, StepCodeBench 49.0%); "the date that matters is October 15." — ** Independent reporting — corroborates the demo figures and the Sep 18 dating nuance. ---Date: ** 2026-09-20
URL unavailable
Independent
OrcaRouter — "Step 5 Preview vs Claude Opus 5: 7 points, 7.9x the bill" (Alistair Wren) - **URL:** https://www.orcarouter.ai/blog/step-5-preview-vs-claude-opus-5** Independent cost-per-run readout citing AA v4.3.2 at maximum reasoning effort: **Step 5 Preview index-run cost $922.84 vs Claude Opus 5 index-run cost $7,274.74 — a 7.88× ratio ≈ 1/8** ("7.9x the cost for 7 index points"); Step 5: AA index 44, $1/$2.70, 95% cache discount, output speed 99.8 t/s (their capture date), 1M context, weights due Oct 15; Opus 5: index 51, $5/$25, 90% cache discount, 55.3 t/s, proprietary; the footnote that "StepFun agentic rows are vendor-run"; the practical decision rule (routine calls to the cheap model, escalate on complexity/failover). — ** Independent evaluation-vendor analysis — the source that lets the research state the ≈1/8 claim is *consistent with* AA's published run-cost data (while not reproduced by flat token arithmetic per source 8).Date: ** 2026-09-20
URL unavailable
Independent
Pandaily — "StepFun Launches Step 5 Preview: 600B Sparse MoE, 1M Context, Weights Open Oct 15" - **URL:** https://pandaily.com/stepfun-step-5-preview-600b-moe-1m-context** Independent English-language recap from the China-tech beat: 600B/27B sparse MoE; **92-layer narrow-deep stack** and 1M context "for long-horizon agents"; text and image inputs; targeting AI coding, software engineering, financial analysis, professional knowledge work; API live now; weights scheduled to open October 15. — ** Independent reporting — the origin of the widely repeated 92-layer architectural detail.Date: ** 2026-09-20 (article; page metadata "Published: September 20, 2026"; surfaced Sep 22)
URL unavailable
Independent
Toolbit AI — "StepFun Step 5 Preview: $1/M Coding Model, Verified vs Promised" - **URL:** https://www.toolbit.ai/blog/stepfun-step-5-preview-analysis** Independent verification-oriented analysis — the key critical source for this story: AA index **44, rank #26 of 202** (launch-week capture; AA's live page now shows #32/212), ties **Kimi K3 at #3 among open-weight-class models** while those rivals are downloadable and Step 5 is API-only; AA page dates the model to **September 18 — a test-tag leak two days before the Sep 20 announcement**; vendor-reported bucket (DeepSWE 67.7%, StepCodeBench 49.0%, GPQA 93.5%, Terminal-Bench v2.1 85.0%, BrowseComp 88.7% — none independently confirmed); the independent-leaning TB 4.0 figure of 33.3 from AA data vs the announcement's v2.1 85.0 (version-gap analysis); **the "1/8 of Claude Opus 5" claim does not reproduce under flat token-price math (~5.9× on a 1M-in + 0.1M-out workload: $1.27 vs $7.50)** and has an undisclosed methodology; agent-loop cache arithmetic (1M-token repo read 10×: $1.00 + 9×$0.05 = $1.45 vs $10 uncached); verbosity tax (160M tokens, 70% more than median); rate tiers (V0 5 concurrent / 10 rpm; 10,000-concurrent-class scaled tiers around $1,500 cumulative top-up — note this research's official docs list V4 as 130 concurrent / 2,600 RPM); **official HF repo unavailable (removed or gated) as of Sep 22**; a third-party **1.21 TB shard repo tagged `license:other` with unproven provenance — not official weights**; WSJ-reported Hong Kong IPO of roughly $500M at up to a $12B valuation (press-reported, not company-confirmed); Step-3.5/3.7-Flash Apache-2.0 precedent; comparison table (Step 5 44 / GPT-6 Astra 53 / Opus 5 51 / MiMo-V2.6-Pro 46 / GLM-5.3 45 / Kimi K3 44 with prices); the pilot-today-vs-wait-for-checkpoint decision rule; three Oct 15 checks (weights live, license, 600B/27B config match). — ** Independent evaluation — the adversarial check on the 1/8 claim, the HF-repo state, the license vacuum, and the MiMo comparison.Date: ** 2026-09-22 (published; "Last updated: September 23, 2026")
URL unavailable
Independent
eesel AI — "StepFun Step 5 Preview: specs, pricing, benchmarks, and my take" - **URL:** https://www.eesel.ai/blog/stepfun-step-5** Independent analysis: spec table (model ID, 600B/27B MoE, 1M context, 64k max output, text/images/video in, text out, reasoning_effort low/medium/high); pricing table ($1.00/$0.05/$2.70 for step-5-preview vs $0.20/$0.04/$1.15 for step-3.7-flash); free V0 tier (5 concurrency, 10 requests/min); the vendor benchmark table reproduced independently (DeepSWE 67.7, StepCodeBench 49.0, ProgramBench 80.5, Terminal-Bench v4 **33.3** — the independent-leaning version of the claim, vs Kimi 12.6/GLM 41.9/Opus 5 52.3/Astra 57.9 — and FrontierFinance 66.4); "65% lower cost than GLM-5.3 and Kimi K3" vendor-chart claim; ~70% of expert participants judging moderately-high-complexity coding autonomy (company claim); the documented countdown-on-homepage to open weights Oct 15; **HF repo exists but is an empty placeholder** with no card/weights/license; StepFun's Flash line previously Apache 2.0 ("permissive open release is likely"); Hacker News community critique (verbosity "waffly padding"; demo thinking-trace controversy; the fair challenge that 27B active is heavier than flash-tier rivals); Claude Code integration through "Step Plan" exposing the full 1M context. — ** Independent technical reporting — corroborates specs/pricing and supplies the Terminal-Bench v4 cross-check, the HF placeholder finding, and community pushback.Date: ** 2026-09-21 (authored; "last edited September 21, 2026"; expert-reviewed)
URL unavailable
Independent
MarkTechPost — "StepFun Launches Step 5 Preview: A 600B-Total, 27B-Active MoE Model With 1M Context for Long-Horizon Agentic Work" (Michal Sutter) - **URL:** https://www.marktechpost.com/2026/09/20/stepfun-launches-step-5-preview** Independent recap: ~4.5% of weights active per token; official docs spec list (text/images/video input, output text, reasoning effort low/medium/high, streaming, tool calling, JSON Mode/Schema, prompt caching); company claims (950 web fetches in a single agent action; Claude Code integration via Step Plan); **92-layer narrow-deep architecture** (via Pandaily) and the deeper-stacks-for-multi-hop-reasoning rationale; training/systems claims (on-policy long-horizon RL, bit-wise train/inference alignment for MoE routing, MTP-3 speculative decoding, FP8 MoE, KV-cache offload, >3× end-to-end long-horizon RL speedup); company benchmark rows (FrontierFinance 66.4 vs 69.7/55.0; DRACO 83.3 vs 87.6/76.8; DeepSWE 67.7; StepCodeBench 49.0; ProgramBench 80.5) with the "High vs Max" effort asymmetry flagged; the two 24-hour agent experiments (H100 kernel to **508 TFLOPS** vs Claude Opus 5's 493; AIME24 53.3% → 60%); AA index **44** (median for comparable price tier: 24 per this article; AA's live page now says 25) and output 99.8 t/s (launch capture; live page now 83.6); pricing table $1.00/$0.05/$2.70 with AA medians $1.88 in / $10.00 out; verbosity 160M vs 92M median; BF16 ≈ 1.2 TB for 600B before KV cache; open weights scheduled Oct 15. — ** Independent reporting — corroborates the official claims, adds the architecture/systems framing, and supplies launch-week AA readouts.Date: ** 2026-09-20 (published; deployment-verified block updated Sep 22, 2026)
URL unavailable
Independent
Artificial Analysis — "Step 5 Preview — Intelligence, Performance & Price Analysis" (model page, fetched in full) - **URL:** https://artificialanalysis.ai/models/step-5** INDEPENDENTLY VERIFIED headline numbers: **Intelligence Index 44, #32 of 212** models (median of comparable-priced reasoning models: 25 — "well above average"); **$0.72 cost per Intelligence Index task**; **160M output tokens on the index run vs 88M median — very verbose (4/4 verbosity units)**; output **83.6 t/s** (#71/212, above class median 74); **TTFT 3.17 s** (median 3.83 s); pricing $1.00 in / $2.70 out / **95% cache discount**; blended price $0.51/M; 1M context (~1,500 A4 pages); text+image input, text output (AA's listing — narrower than the docs' video claim); reasoning variant ("a non-reasoning variant may also exist"); 1 API provider; classification **Proprietary** with FAQ answers directly relevant to the story ("Is Step 5 Preview open source? No. The model weights are not publicly available"; "How intelligent…? 44… median: 25"; "Released September 18, 2026"); AA Intelligence Index v4.3.2 methodology note (10 evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1). — ** Independent technical evaluation — the principal third-party anchor for score, task cost, verbosity, speed, latency, and the proprietary-today classification.Date: ** consulted 2026-09-23; AA dates release September 18, 2026; page data is live
URL unavailable
Secondary Sources (5)
Secondary
WorldAttention — "StepFun's Step 5 Preview ranks top three globally in open-source AI at one-eighth Claude Opus 5 cost" - **URL:** https://worldattention.com/stories/stepfun-releases-step-5-preview-open-source-ai-model-3de4c7aae0** Multi-source aggregation showing the spread of the launch narrative across Chinese media (新浪财经 / 财联社 / 证券之星 with Sep 19–20 timestamps) and neutral testingcatalog; reader-quick-take restatement of the 1/8-of-Opus-5 and top-three framings. Useful as a map of *claims in circulation* — the research treats the underlying claims per their own evidence status. — ** Secondary aggregation — corroborates that the "1/8 cost / top-three open-source" framing dominated launch-week coverage. ---Date: ** first published 2026-09-21 00:13 UTC (story over coverage Sep 19–20)
URL unavailable
Secondary
36Kr English — "Global Foundational Model Showdown in Full Swing…" (StepFun release passage) - **URL:** https://eu.36kr.com/en/p/3991347438173186** Chinese mainstream-media framing: "In the globally authoritative Artificial Analysis Intelligence Index, Step 5 Preview scored 44 points, ranking among the top three open-source models in the world, and the cost per task is only 1/8 of Claude Opus 5"; light on the competitive field ("the answer for Kimi is almost out in the open"). — ** Secondary reporting — documents the dominant Chinese-media narrative the announcement generated (relevant to evaluating the "top-three / 1/8" claims as widely-relayed claims).Date: ** 2026-09-20/21 window
URL unavailable
Secondary
CloudPrice — "Step 5 Preview pricing & specs — StepFun" - **URL:** https://cloudprice.net/models/step-5-preview** Secondary specs/pricing cross-check: release date 2026-09-20; context 1.0M; input $1.00 / output $2.70 / cache read $0.05; provider listing including Vercel AI Gateway (`stepfun/step-5-preview`); capabilities list (reasoning, adaptive reasoning, function calling, parallel function calling, structured outputs, native JSON schema, prompt caching); version table showing Step 3.7 Flash ($0.20/$1.15, 262K) below Step 5 Preview; a separate `step-5` (non-preview) listing with an Intelligence Index of 43.7 aggregate. — ** Secondary directory/API source — independent cross-check of pricing and the release date.Date: ** data updated 2026-09-20 (release date listed 2026-09-20)
URL unavailable
Secondary
cnboat.com (China Supply Info) — "StepFun Releases Step 5 Preview: 600B MoE Model with 1M Context, Weights to Open Oct 15" - **URL:** https://www.cnboat.com/news/show.php?itemid=1312** Chinese-tech secondary recap: AA scored it ~44, "on par with the 2.8-trillion-parameter Kimi K3, placing it among the global top three open-weight models"; pricing $1 per M input / $2.7 per M output at ~100 tokens/s; the domestic release-cadence observation ("Step, Qwen, Zhipu, Kimi" — "price-performance and speed, not just raw scores, decide adoption"); the "shifting from commercial export to ecosystem and standards export" framing. Source credits: StepFun / SegmentFault / Artificial Analysis. — ** Secondary reporting — corroborates the AA 44/top-three framing and provides Chinese-market framing.Date: ** 2026-09-21 (article)
URL unavailable
Secondary
KuCoin flash news — "Step 5 Preview Launches with 600B Parameters and 100K Token Context" (币界网; labeled "CoinDesk reports…") - **URL:** https://www.kucoin.com/news/flash/step-5-preview-launches-with-600b-parameters-and-100k-token-context** Secondary recap of the launch: sparse MoE 600B total / 27B active; 1M-token context (body; the headline says "100K Token Context" — a headline error flagged in this research); text and image inputs; API and Studio fully open; weights set for Oct 15; AA early evaluation — index 44 matching Kimi K3 Max, ~100 tokens/s, task cost $0.71 vs Kimi K3 Max's $2; the 24-hour kernel demo reaching **508 TFlops after 22 hours on H100**, surpassing Claude Opus 5's 493 TFlops; the AIME24 experiment (Qwen3-30B-A3B 53.3% → 60%). This was the discovery record's listed independent source and is corroborated by sources 6, 8, and 11. — ** Secondary reporting (crypto-aggregator relay of Chinese crypto/tech press) — useful time-stamp and demo-figure corroboration; headline quality caveat documented.Date: ** 2026-09-20 02:23 UTC
URL unavailable
Unverified Sources (1)
Unverified
WSJ-reported StepFun Hong Kong IPO (~$500M raise, up to $12B valuation) - **No URL recorded** (not fetched directly; relayed second-hand)** Only context: per Toolbit, "the Wall Street Journal has reported StepFun is preparing a Hong Kong IPO of roughly $500M at up to a $12B valuation; that is press-reported, not company-confirmed." Included solely to frame strategic stakes (Section 11/17 of research/S10.md). — ** RUMOR / press-reported, not company-confirmed — flagged as such; no project claim depends on it. --- **Verification note (per the mandatory URL rule):** every URL above begins with `https://` and is a complete absolute URL. The single limitation without a URL is phrased as "No URL available," not as a URL field.Date: ** reported as circulating around Sep 20–22, 2026 (per Toolbit, source 8)
URL unavailable