macintosh.world | Log In | Register

Today | News | Books | Recipes
Notes | QuickTake | Wiki | Browse
Maps | Reference | Reddit | YouTube
Chat | Games | About

Back to HN

Qwen-Image-2.1: Compact, efficient, and unified image creation

by jmillikin | 113 points | 37 comments | 2026-09-20 08:09:25 Central

Open Source Link | Read Source Here

Open on Hacker News

Comments

jfoster
A lot of the previous Qwen models seem to have used Apache
licenses, among
others:https://en.wikipedia.org/wiki/Qwen#List_of_modelsUn
fortunately, it looks like this model is using a much more
restrictive
license:https://github.com/QwenLM/Qwen-Image-2.1/blob/main
/LICENSE

  > gregoriol
Was going to post about this: the last image models
with Apache 2.0 license seem to be from 2025, recent
Qwen models are "non-commercial use".

    > > Luker88
Companies can use llm to license-wash open source
code regardless of license.How difficult would it
be to use this model to create a second model
without licensing issues?

    > > bloaf
I love the non-commercial clauses because of how
many people are using these for deceptive ads and
"virtual staging" and fake social media accounts.
Anything that makes those guys lives harder while
still letting me make silly pictures for my kids
and tapestries for my D&D campaign feel fine by
me.

gunalx
Its happy to see a new open image model from qwen. But the
license is a let down. And it dosent even beat their
closed qwen3 image wich is already a bit old.

trains39472
A 7B diffusion model can now render CJK text better than
Microsoft Windows.

  > tomjen3
Just think about how recently we got that feature in
the official ChatGPT image gen. And now we have that
running locally - assuming that is, I can figure out
how to get this running on my Mac - blows my mind.

    > > Havoc
Pretty sure comfyui has a mac executable
d2kx
God I love the Qwen team. Easily the most diverse set of
models from all the Chinese labs. Only Gemini/DeepMind
comes close.

mdp2021
How do you use this model locally, similarly to using
`llama-server -m <model>`?(Of course I mean: outside
direct use of Python, and in the most efficient way.)

  > Iolaum
There is difussion.cpp which is intended for those
types of models. I set up krea-2-turbo with the help
of ChatGPT 2 months ago, if you have a capable
computer that's what I would suggest once it becomes
supported.

  > fp64
on the linked GitHub page they list support Diffusers,
ComfyUI, vLLM-Omni, SGLang, and LightX2V with links to
each

  > utopiah
why not just as you suggested i.e.
https://qwen.readthedocs.io/en/latest/run_locally/llam
a.cpp.... then get the result either via a UI or
wget/curl it back?

    > > mdp2021
I am not sure that llama.cpp also supports image
generation models.

      > > > utopiah
it's multimodal, see
https://github.com/ggml-org/llama.cpp/blob/mas
ter/docs/multi...

        > > > > exe34
Multimodal doesn't guarantee input and
output.> Currently, we support image,
audio and video input.

  > embedding-shape
Probably ComfyUI is one of the easiest way to get
started with local image/video models. Or perhaps
vLLM, if they have support for it already, would be
something like `vllm serve <model> --omni --port 9080`

hgufj
I am really grateful to the Chinese Labs for open sourcing
their best models. If it was left to the Americans, we
would be forced to pay obscene API fees to use them.

  > jfoster
Note that the license on this has this in it:> You
shall not use the Materials for any commercial purpose
without obtaining a separate commercial license from
us.It probably will be much cheaper to use than other
image models, but it seems that will be up to the
whims of Qwen/Alibaba rather than just being the cost
of putting it in a cloud
provider.https://github.com/QwenLM/Qwen-Image-2.1/blob
/main/LICENSE

fishfasell
The capabilities of local LLM text-to-image is honestly
pretty damn impressive. IMO, I think local image
generation is currently ahead of local code generation. I
can get an image in seconds locally with the quality being
way higher than what I'd expect from a local model.
However with coding it's much slower and much less
impressive. I'm sure there's a reason for this and I'm not
an AI expert so I'll let the smarter folks tell me why,
but that's just been my observation thus far.

  > mft_
I've played with diffusion models on and off since the
first release of Stable Diffusion - just for
amusement, without a particular goal.Recently, I've
been helping a friend's wife with some basic vector
images for her sewing hobby (she has what is
essentially a CNC sewing machine) and have been
super-impressed with FLUX.1-Kontext, which I've been
running on my Macbook Pro with mflux. Its ability to
(for example) take a photo of a human or an animal and
return a line drawing which is recognisably them
(rather than just a generic similarish image as I've
experienced with other models) is excellent.It's an
older model now, but (AIUI) has the text-handling
features baked in, and in my various testing is very
reliable at giving me the outputs that I want, without
the randomness I've experienced previously. It's big
and relatively slow (~3 mins per 512x512 image edit on
my M1 Max Mac) but excellent to work with. It's also
very straightforward to set up, without the harness
complexity of e.g. comfyui.

    > > jLaForest
is the cnc sewing machine an off the shelf model
or something DIY? I'd love to hear more

      > > > mft_
Off the shelf - it's a Brother. It prints via
a proprietary file format (.PES) but there's
an extension for Inkscape that supports
creation and export.

  > victorbjorklund
I mean I'm sure it's the reverse for an artist. They
would be less impressed with the image and more
impressed with the code quality

    > > gedy
To generalize, LLMs are great at what you are not
skilled at.

    > > fishfasell
That's a fair statement, I agree. I'm quite an
abysmal artist so I could be a victim of my own
bias here

trentor
They finally fixed their VAE. It really held back their
models over the last 2 years.

Hard_Space
Interesting in the example of assembling the Cheers team
how the otherwise great result genericizes Shelley Long.

  > hughc
The result seems a pretty good representation given
the source image wasn't that great. I think that Woody
Harrelson comes across much worse.

TomGarden
Very impressive, and kind of worrying a 7B model can have
such capabilities. The implications are huge. And Qwen
does no watermarking (yet) yeah?

  > trentor
They always had a fourier space mark in their models
even without the VAEs are usually pretty easy to
detect.

    > > TomGarden
Ah I wasn't aware, thank you
spottedmarley
Boy do I love waking up to find a new awesome toy from the
Qwen team waiting for me to play with! Pulling it now

bknight1983
While I'm impressed with the Bluey example, the lack of
Muffin disappoints me.

hn45e7pbij
Image gen you eyeball one frame and stop, code needs
hundreds of tokens all correct in sequence, one bad line
and the whole thing fails.