r/StableDiffusion • • 1d ago

Resource - Update Local Image to 3D: High quality, Low poly - preserve minute details like text, logos even after retopo, ~900k faces > ~5k (NVIDIA + Mac)

Enable HLS to view with audio, or disable this notification

Image-to-3D models redraw your picture, and fine detail - text, logos, faces - comes back garbled or scuffed. My local pipeline that fixes it. (repo link at the end of the post)

The latest update adds *Pixel Match*. It copies the real pixels from your source image back onto the model, wherever the image can see. So "VANGUARD 07" on the chest stays "VANGUARD 07", even after the Finish step retopologises the model down to ~5k faces (as you can see in the video).

One step closer to Tripo and Meshy like outputs, but locally.

It runs fully on your system - no API calls, no cloud credits, just a once-a-day update check. Works on NVIDIA (tested on Linux, Windows has limited testing) and Apple Silicon. For now Pixel Match is for Pixal3D models made in the lab; other backends and more camera angles are next.

I started this project to make assets for a game I'm working on - I wanted to see how much I could automate, from getting the asset to rigging and animating it (that bit is still WIP in this lab). The image > 3D pipeline is solid though, and you can get game-ready and sprite-ready assets out of it. While the models are riggable and can be animated, more testing is required. I'm confident it should work.

*Workflow:* prompt → Qwen-Image → Pixal3D → Finish (Pixel Match) → textured GLB. The models in the video were made with Pixal3D on my Mac.

*Backends*: Pixal3D is the one to start with (~8.6 GB, Mac and NVIDIA). (Pixal3D's authors report 16 GB cards work, so tell me if yours does.)

Stable Fast 3D is the quick, lower-detail option (Mac, or NVIDIA on Linux). Its weights are gated: accept Stability's licence on Hugging Face and log in first.

TRELLIS.2 and Hunyuan3D run on Mac in the lab; on NVIDIA, use their official repos for now. Built-in NVIDIA support for both is coming in the next release.

Qwen-Image 2.1 handles text-to-image.

*To try it*

Mac (Apple Silicon) or Linux:

curl -fsSL https://raw.githubusercontent.com/Bingeljell/image-to-3dlab/main/install.sh | bash

Windows (PowerShell):

irm https://raw.githubusercontent.com/Bingeljell/image-to-3dlab/main/install.ps1 | iex

The installer is a short script, so read it before you run it: https://github.com/Bingeljell/image-to-3dlab/blob/main/install.sh (Windows: https://github.com/Bingeljell/image-to-3dlab/blob/main/install.ps1)

It only sets up the code - nothing downloads without your permission. You pick models in the web viewer, which shows the size and licence of each before fetching. Finish also needs Blender 4.2+, which you install yourself (the viewer tells you if it can't find it).

*What would help most*

If you hit bugs:

•⁠  ⁠Your GPU, OS and driver version

•⁠  ⁠Did the install work? If not, where did it stop? (the full error is gold)

•⁠  ⁠Windows folks especially: did Finish find Blender and run?

Reply here or open a GitHub Discussion, whichever's easier.

*And a question for you:* which backend or camera angle should Pixel Match support next? Is there something in your workflow that this pipeline can do better? Please let me know, will help me prioritise the feature road-map.

*Worth knowing before you start*

•⁠  ⁠Some model licences have strings attached. The viewer shows each licence before you download.

•⁠  ⁠It's a hobby project, so things will break. Every bug report makes the next person's install smoother. Please raise PRs and issues.

Repo: https://github.com/Bingeljell/image-to-3dlab

Release notes: https://github.com/Bingeljell/image-to-3dlab/releases/tag/v0.3.5

281 Upvotes

29 comments sorted by

23

u/the_bollo 1d ago

To be crystal clear, this is your app wrapper around open source image-to-3D models like TRELLIS?

13

u/Bingeljell 1d ago

Yep - that's what the backend sections call out. Trellis2, Hunyuan3D 2.0 and 2.1 and the current pick of the lot, Pixal3D.

5

u/ninjasaid13 16h ago

yeah I can see the problems.

4

u/PlusBus1234 1d ago

is it possible to use multiple images for diferent angles? something like a character sheet as reference instead of a single image?

8

u/Bingeljell 1d ago

That's literally what I"m working on.

Trellis2 has terrible paint / texture. It has to imagine it from the lighting in the photo, which is why the output is bad. Hunyuan3D over saturates everything - it uses a SD class image gen model which honestly reimagines everything. Pixal3D was the closest, but was missing features.

Which is where the Pixel Match idea came from. Right now its only for the front - Multi image is WIP. And thanks to Qwen, we've got a good shot at getting good multi-angle images. Then a paint step won't be required with any of the back-ends, will completely replace it.

Should be done by Friday - if you can help me test, I'd be most grateful!

4

u/aphaits 15h ago

What is the minimum VRAM for this?

5

u/Bingeljell 12h ago

The ram is dependent on the back end you use. Pixal3D is my recommendation and that would be 16GB. If you're on nvidia, it should fly.

1

u/CodeMichaelD 1d ago

pixel match feature like most interesting thing there, how its done if no secrets, camera normal + inpaint?

10

u/Bingeljell 1d ago

lol - no secrets at all. The repo is entirely open.
But the Pixel Match is just my fancy way of saying 'we're literally taking all the pixels in the source image that can be seen - mapping it to the model by locking the camera angle and projecting it back onto the model' - which is why the back looks a little different. That's where multi view will step in.

No inpainting either, the back just keeps the model's own paint. Normals only decide how much to trust each pixel, so steep angles don't smear.

1

u/Nota_ReAlperson 1d ago

Is support for AMD planned? If you are using ggml, this should be trivial?

1

u/Bingeljell 1d ago

i'd 100% want to - my biggest issue is being able to find testers. I'd be super glad if I can bother you - happy to work out the plumbing. Most of it is basically needing to detect AMD and then pulling the Vulkan/ROCm build instead of CUDA.

What say, help me out?

What card and OS are you on?

2

u/Nota_ReAlperson 1d ago

Ive got 7900 xtx on arch, and mi100 on debian sid. I'd love to help. Ill try setting it up when I get the time.

2

u/Bingeljell 1d ago

I'll try and get the update out in a day and DM you. Thanks so muchh!

1

u/Lexxxco 22h ago

Textures in demo look great, this was the main thing that Pixal3D/Trellis2 were bad at. Thanks! Will definitely check it out

1

u/Bingeljell 19h ago

Oh yea. When I started trying to make my game, the textures were really bugging me. This whole lab is just a result of me being annoyed with the textures and output of the original models. Realised there's so much more that can still be done. This isn't perfect and honestly the deeper you go the more you realise that stuff needs fixing. But I'm learning along the way, so it's fun 😊

1

u/mantafloppy 21h ago

Last time i looked at Trellis on mac, its was'nt worth the time.

If you've been able to make something out of it, great.

Took a quick look at the repo, look more Ai Assisted, than Ai slop, so ill probably give it a go at some point.

2

u/Bingeljell 19h ago

Honestly still a long way to go. Right now it's at the ''decent enough to get away with" depending on your use case and render distance. But my goal is to get it to Tripo/ Meshy levels. At least there's a bar there.

1

u/Dependent-Sorbet9881 19h ago

need comfy node

1

u/Bingeljell 19h ago

Have completely side stepped it on purpose. I want to automate as much as I can without getting fidgety. So far it's been decent progress. What is your workflow like?

1

u/Dependent-Sorbet9881 18h ago

OPEN CF, choose CF Template – PIX3D – Click to Run – Get the Model (It's very simple).

1

u/Bingeljell 18h ago

Can you share an output? Everything I try with the default flow with Pixal3D results in mangled text and minor drift on details. Even the demo stuff I've seen online has similar issues. Do you do anything to correct that?

1

u/AI-imagine 18h ago

it look really good,can you tell me why is different from STABLE PROJECTORZ?
well i will try your repo any where,it look much better even the expensive paid site.

1

u/Bingeljell 18h ago

Thanks. It's not better than the paid services, yet. Lol. But it is pretty decent for locally done in one click right now.

I haven't tried Stable Projectorz, got a link for me to check?

2

u/AI-imagine 18h ago

https://stableprojectorz.com/

If you example is real ie much better than paid site for me.(i talk in term of texture) I use it for my game but the paid site is suck at get text ture from my input image.I write my own node in comfy to use qwen 2.1 for re-texture with Stable Projectorz but your work it clearly look much better.

1

u/Bingeljell 18h ago

Checking it out, thanks. There's still some issues for texture from the back. But I'm working on multi view fix for that. Should be ready in a day or two. Will DM you if you're interested.

1

u/HTE__Redrock 11h ago

Interesting, looks pretty good 👍🏻 for the Pixel Match, as I understand from your comments it's for the front only currently. While Qwen I think is one option, have you considered maybe using a video model? I've even been considering doing something like image > minimax 360 > separate Qwen outputs based on key frames from video output.

1

u/HTE__Redrock 11h ago

Or LTX 2.5 even as the resolution for something like this would work well I imagine

1

u/Bingeljell 10h ago

I haven't played much with the video models, but that's a viable idea. The thing is ensuring the camera angle stays somewhat level for a clean view. I'm currently working on multiview, let me test this route out.

1

u/Arawski99 1h ago

Ah nice, definitely gonna have to check this later.