macintosh.world | Log In | Register

Today | News | Books | Recipes
Notes | QuickTake | Wiki | Browse
Maps | Reference | Reddit | YouTube
Chat | Games | About

Back to HN

Step 5 Preview: Advancing the Pareto Frontier

by nateb2022 | 137 points | 33 comments | 2026-09-19 23:35:59 Central

Open Source Link | Read Source Here

Open on Hacker News

Comments

BoppreH
In their first demo video, to make a 3D render of the
photo, the thinking trace gives away the game:>
Interesting! It turns out there's already an existing
project here [...] The project is fully built [...]I'm
always astounded how little effort is put into checking
the AI answers displayed in these announcements. Back when
I paid more attention, I remember OpenAI's and Google's
demos constantly showed their AIs giving wrong answers.

  > BoppreH
Since people seem interested in this comment of mine,
here's another fun line I noticed in the same video,
at the very top of the logs, just before the Step 5
agent found the already-completed project:> Error:
OpenAI API error (403): {"message":"model water18-new
is not available for user i-yuliang
[trace_id=bfcfdd6bcdc236ca18d009c65cca52e4
code=40004]", "type":"invalid_request_error",
"param":null, "code":null}Here's a still frame for
posterity (apologies for the quality, the original
video is tiny): https://boppreh.com/room.jpgDemos are
demos and recreating scenarios is to be expected, but
oh boy, somebody should review these things before
publishing.

  > Bolwin
I think it's more likely they recorded the video when
the project was already done than cheated

    > > BoppreH
Yeah, I agree that's the most likely explanation.
My point is that these demo hiccups should not
show up in your announcement, it casts doubt on
the product for no good reason. It's sloppy.

nh43215rgb
> Built on a sparse Mixture-of-Experts architecture, Step
5 Preview has 600B total parameters, with 27B active per
token, and supports a 1M-token context window and vision
input.

> Step 5 Preview scores 44 on the Artificial Analysis
Intelligence Index.

> The model will be released with open weights on October
15.

I guess being Chinese company they decided to skip version
4, while also giving impression to be on the similar
iteration with leading companies (claude opus 5).
I wonder if other Chinese labs like Kimi/Moonshot will
follow suit.

  > JohnsonZou
Another possible reason is that the number 4 is
considered unlucky in traditional Chinese culture.

    > > zozbot234
Parent commenter hinted at that. Yet DeepSeek has
released their V4 which was hugely successful, and
even their new architecture is marked V4.1. Qwen
internals mark their Flash-Next model, also very
compelling, as "qwen4exp". So both of them are
bucking the negative stereotype.

  > Tepix
600b-a27b doesn't sound enticing. Also with the higher
number of active parameters compared to GLM 5.3 flash
and DeepSeek V4/4.1 flash, I don't see how they want
to be more efficient at inference.

  > NooneAtAll3
I kinda wish everyone just used dates instead...
    > > NetOpWibby
Using dates immediately makes you look dated and
everyone is in this constant race to be
first!!1!!1!I agree with you though, ChronVer all
the way.

  > Bolwin
Moonshot has already teased K3.1 so not likely
    > > dannyw
K3.1 would likely be a deeper/longer post-train
from K3, so that'd make sense.It's all marketing
anyways, but that's at least how a lot of labs
have been naming things (sometimes).

  > torginus
Dunno, if US labs would embrace this silly logic, then
Anthropic would be compelled to release Fable/Opus 6
instead of a .1 release

bethekind
> Without any Pokémon-specific optimization, Step 5
Preview has so far sustained progress for more than 3,000
turns and 6 million tokens of interaction. By turn 3,082,
it had unlocked Cut, earned three Gym Badges, and defeated
Lt. Surge. The run is now roughly one-third of the way
through the main story.Finally FireRed is being used as a
benchmark again! I believe Astra can beat it in 18 hours.
Not sure how that compares.

garo-pro
IT's Artificial Analysis Index is the same as Kimi K3,
which is about 4.6x bigger, and GLM 5.3, which is about
1.25x bigger. Pricing is $1/$2.70 i/o. Openweights on
October 15.

  > tomComb
Most of these models, both open and closed, are now so
over tuned to agentic and coding tasks that they no
longer work well for general purpose. Kimi K3 is an
exception to that - maybe you need that larger size
todo well on a broader range of tasks.

ghoshbishakh
Their posisitoning is nice. Instead of saying they are
cheaper and a bit less performant (in terms of
intelligence), they say they are best among the cheaper
and a bit less performant ones.

segmondy
The previous Step models were pretty decent, but
unfortunately for them, their model reasons too much and
too long and the Qwen/Kimi/DeepSeek/GLM have been
stronger. Hopefully this doesn't reason too long to get to
the answer. I welcome any open model, the more the
merrier.

InsideOutSanta
GLM-5.3 and Kimi K3 are just below where I can use them to
completely replace frontier models. Oddly,* SWE-2 is there
for me.If this performs similarly in the real world, we're
approaching a level of capability where for most devs, it
only makes sense to pay for Anthropic or OpenAI
subscriptions if they are heavily subsidized and actually
cheaper than these alternative options.* Oddly, because I
perceived Devin as being kind of a joke before trying
SWE-2.

Jacques2Marais
For anyone else looking for the pricing:
https://platform.stepfun.ai/docs/en/guides/pricing/details
#p...

  > sieve
I regularly hit 200-300M cached reads every day on
some of the models I use. It has exceeded 7-800M on a
couple of occasions. At $0.04/M, that is $8-12 per day
only for cached reads.

    > > ignoramous
> At $0.04/MUnless you meant step-3.7-flash, the
input cache hits are $0.05 per mil for
step-5-preview.> $8-12 per day only for cached
readsPretty decent "API" rates for ~500M+ tokens
on Step Fun 5, a Kimi K3 / GLM 5.3 level
model?Their "Step Plan" is ridiculous, by
comparison: ~$60 usage on $6.99/mo; ~$220 on
$9.99/mo.
https://platform.stepfun.ai/docs/en/step-plan/over
view

      > > > sieve
Yes, I meant the Flash version.I have used
Kimi 2.5 and GLM 5.3 (& 5.3 Flash). Do not
need them for what I do outside of spec
hardening (basically, a lot of chatting).I
tend to know exactly what I want and most of
the weaker models are enough to get me there.
I have mainly been using MiMo, DeepSeek V4
Flash and MuseSpark Contributor over the last
month or so.

      > > > esafak
Not generous enough? How token efficient and
fast is it compared with American models?

wrs
Based on the examples, It's like the thing took a writing
class from Claude, but isn't quite as smart - kind of
worst of both worlds. The "Interactive Reporting" one is
particularly terrible. A huge amount of waffly padding
around a thesis that may or may not exist. So if the goal
is to make long fancy reports that nobody will read, then,
yes, very cost-effective."the question it raises matters
more than the answer" ???"So 'the river drifts from cool
to warm' is not a figure of speech." Oh really?

conception
Huh wonder why they skipped 4?
  > Alpha3031
Sometimes 4 is skipped due to being considered
unlucky.

    > > adrian_b
In China and in places influenced by Chinese
culture, due to homonymy between "4" and death.

  > cynerx
Sounds like death in Chinese.
    > > slekker
Do they rename it like in Japanese 7?
dofm
"Pareto frontier" really is the new "web-scale".
derliebej
How about adding a contested historical facts benchmark?