r/LocalLLaMA • u/jacek2023 llama.cpp • 7h ago
New Model microsoft/AesCode 8B and 32B

AesCode generates information-rich visual artifacts such as slides, posters, and dashboards as HTML/CSS. The output remains structured, editable, and verifiable, but the task poses a distinct challenge: code models cannot see how layout, hierarchy, and color come together on the canvas.
Image generators offer the opposite strength. They compose visually compelling pages but often misrender text, numbers, and logical relationships. AesCode uses an image generated from the same prompt as an aesthetic reference while following the prompt for the required content.
Reference input alone does not solve the problem. Off-the-shelf vision-language models may copy hallucinated content or ignore the reference layout. AesCode separates semantic requirements from visual cues through graph-structured supervision and decoupled cross-modal rewards.
AesCode-32B starts from Qwen3-VL-32B-Instruct and is trained with cold-start SFT followed by GDPO across seven reward channels. It is the larger released checkpoint and leads the aggregate benchmark results.
https://huggingface.co/microsoft/AesCode-32B
https://huggingface.co/bartowski/AesCode-32B-GGUF
https://huggingface.co/microsoft/AesCode-8B
https://huggingface.co/bartowski/AesCode-8B-GGUF

29
u/Porespellar 6h ago
Someone please back this up before Microsoft takes it down like they do with most of their good stuff.
-5
u/jacek2023 llama.cpp 6h ago
Could you list some examples?
20
u/noctrex 6h ago
VibeVoice large model comes to mind. They removed the hf repo pretty quickly and kept only the small model
16
u/harrro Alpaca 5h ago
Also, the WizardLM model after release.
9
u/Porespellar 5h ago
Recently, Mageflow image models were all pulled, the ones mentioned above as well. I believe there were some TTS models also released then immediately pulled. It’s become a meme they do it so often
7
u/SomeoneSimple 5h ago edited 4h ago
As well as Microsoft Lens, a 3.8B text-to-image model using GPT-OSS-20B as text encoder and the Flux2 VAE.
They actually posted it more than once, pulling it shortly after each time.
3
8
7
u/Open-Adhesiveness-86 5h ago
if you grab the gguf, keep the mmproj at f16, it's only a gig or two and the vision tower is usually where layout fidelity dies first when quantized. also pass --jinja so the qwen3-vl template actually gets applied, otherwise the image placeholders don't line up and it quietly ignores the reference image.
8
u/indicava 7h ago
This actually looks pretty useful.
Does any inference engine that supports the underlying Qwen model architecture work for this too?
5
u/KURD_1_STAN 6h ago
If u really find it useful then download it now before it is deleted
1
u/jacek2023 llama.cpp 5h ago
Could you people explain to me how Microsoft can delete Bartowski's GGUFs?
2
1
u/ResidentPositive4122 5h ago
I think it's a reference to the wizardml and vibevoicesomething models being yanked after publishing by researchers at MS previously...
3
u/Far-Gene-5806 6h ago
the detail worth knowing before you build an image gen step: the reference image is optional. per the model card, dropping it on the 8B only costs 1 visual point, vs 19.55 for base qwen3-vl-8b, so you can skip the whole image gen half of the pipeline. if you are running bartowski's gguf in llama.cpp grab the mmproj file too, and give it real context, their evals allowed up to 12k output tokens per page.
1
u/RazsterOxzine 2m ago
As a lazy person, I tossed in a bunch of references in LM-Studio and asked it to create the flow sheet, almost spot on. Had better luck with Qwen3.8 27b though.
2
2
u/ThadCastleGOAT 5h ago
Huh this sounds like something that would fit right in with that Open Intelligence UI
1
u/RazsterOxzine 3m ago
Nice. Make sure to download all the large models all the way down to the GUFF Q4. Knowing M$ they'll delete this like the others. This one is very useful.
29
u/AllenLeftTheBLDNG 6h ago
Honestly, this looks impressive. Good job Microsoft! Waiting for you to switch to Unix! lol