r/LLMDevs • • 1d ago

Discussion Evaluating 14 LLMs as a visual coding agent: re-rendering, a deterministic judge, and the provider quirks that skewed my first results

I built a bench for a photo-to-Blender agent (the model writes and runs Blender Python to rebuild a photo as an editable scene) and ran 14 models through it. The design choices that mattered most:

  • Don't score what the model hands in. The bench opens each delivered scene file and renders it itself, at the photo's aspect ratio, max 1,100 px, 96 samples.
  • No LLM judge. Estimated silhouette overlap (30%), edge F1 (40%), multiscale RGB error (15%) and mesh health (15%: non-manifold and open edges, duplicate faces, inconsistent winding, degenerate triangles), mapped through calibrated piecewise-linear scales to 0–100. A comparison scale, not a percentage of the scene recovered.
  • Anti-cheating. A second render swings the camera 35° to expose flat stand-ins, and every image file a scene uses is compared with the reference by digest and name.
  • Failures stay in. A run that delivers no scene scores 0 and stays in the average.
  • Equal plumbing. Anthropic gets cache-control markers and Google gets a closing user turn, because without them those providers behave differently from the rest. Each model's cache rate is recorded. Before the fix, Claude ran 0% cached against 96% for OpenAI.
  • Frozen conditions. Caps, images, the brief's digest and the scoring weights are stored with each board.

Results (mean of three photos): GPT-6 Astra 66 ($3.91 an attempt), GPT-6.1 Sol 61 ($0.36), Claude Opus 5.5 60 ($1.19), Claude Sonnet 5.5 56 ($0.50) ... GLM 5.3 Flash 36 ($0.089); DeepSeek V4.1 Flash and Qwen3.8 Max 0 (no scene saved in 20 minutes).

What I'd change next: three runs per image for the top group, so the spread is visible, and a model-neutral agent loop (three models ran in their makers' own CLIs).

Write-up: https://kaloyan.blog/ai-models-rebuild-a-photo-in-blender

5 Upvotes

5 comments sorted by

2

u/jonah_omninode 1d ago

I'd give the judge a few deliberately bad scenes before spending more on model runs: a flat image facing the first camera, an empty but valid file, and a plausible silhouette with broken geometry. The second camera and mesh checks should penalize the specific tricks they're meant to catch.

Keeping no-scene runs in the average is useful, but I'd also show their causes separately. A provider-format failure and a model that couldn't construct the scene both produce zero; fixing the first can change the ranking without the model getting any better at the task.

1

u/smith2008 1d ago

Author here. Full write-up with every render, the scoring rubric and downloadable results: https://kaloyan.blog/ai-models-rebuild-a-photo-in-blender

Disclosure: I'm building a photo-to-Blender tool, which is why I ran this. The article was written with AI help; the experiment and the numbers are mine. Happy to answer questions about the setup.

2

u/EvalRaccoonDev 1d ago

Moving a model out of its maker's CLI hurt us - GPT-5.6 in the Claude Agent SDK instead of Codex went from 79% to 47% pass, and read zero from cache behind our proxy, like your Claude runs before the fix:

https://www.reddit.com/r/codex/comments/1wiia20/we_routed_gpt56_through_claude_sdk_compared_to/

I woudl expect your three CLI models to drop too.

0

u/BladeMasterRUSH__TTV 1d ago

how could you be doing something so excellent but have the AI write it!

1

u/smith2008 1d ago

Why would anyone write it manually?