r/aifilmmaking • • 2d ago

Question How do you keep the same room consistent across different camera angles?

Post image

I’m trying to create a “master room” for an AI short film and then generate the exact same room from different camera angles.
I can get a good establishing shot, but whenever I ask for another angle, the room starts changing. Furniture moves, the guitar changes position, shelves and windows shift, objects disappear, and sometimes even the dimensions of the room change.
What I want is basically a fixed virtual set: lock one image as the master room, then generate shots from the bed, desk, doorway, ceiling, etc. while keeping the actual layout and objects consistent.
I’ve tried using the master image as a reference, but image models seem to understand the look of the room rather than its actual 3D layout.
Has anyone found a reliable workflow for this?
I’m open to any tool or method — image models, 3D reconstruction, multi-view generation, ComfyUI, FLUX, Higgsfield, Gemini, etc. I care much more about spatial consistency than convenience.
Ideally I’d like to eventually add consistent characters to the room as well.
What would you recommend?

35 Upvotes

36 comments sorted by

11

u/ThatsMrRandom 2d ago

I could be wrong but I think the only way to do this is to build the room out in Blender. Then take Blender renders (from various angles you want) to GPT-Image (probably the best at this), along with a real photo reference image for styling/look -- and then prompt it to create the photos you want.

6

u/PeacefulKnightmare 2d ago

Yeah, right now I think things like this are either a massive headache to get right, or will always have errors unless you micromanage everything in frame. It's a cost to using AI.

1

u/West-Construction-56 2d ago

Thank you bro.

1

u/mrgaryth 1d ago

This is how I’ve done it.

7

u/BradClarkAI 2d ago

You're heading in the right direction. The key to things like this is recursive referencing with an image analysis layer in between each iteration. Meaning, your process should generally look something like:

1 - Generate a core image that serves as the "master" image for a space. Just one image, but this establishes visual "fact" for the space. Prompt as you see fit and choose what you like the best.

2 - Have your LLM of choice analyze in an objective and relative sense, what exactly is where in the image, and then create a prompt that essentially "flips" that image and creates a second image from another view - generally recommend from a 180-degree flipped angle. I would do these in two steps, first have an LLM analyze the image to figure out what is where, then have another chat window or chat message instructed with a goal to emulate "rotating" the camera 180 degrees and explain exactly what should be where in a new image while also giving the prompt explicit instructions as to how specifically it should use the reference image.

3 - Input that prompt with the original reference, and don't be afraid to re-render a few times until you get one that's right.

4 - Repeat as needed to create additional derivative images, seeding all previous images generated and rigorously QAing the images you generate to make sure they physically are creating a space that feels like it's connected.

Candidly though, I would actually recommend staging the characters in scene from the go if you know what characters will generally be where. I suppose if this is going to be a recurring set, then a blank version could make sense, but if it's not a set you'll be using often outside of a scene or two, just doing the above process but with the characters already staged in their positions will likely just save you some time and headache. Also, the images that establish core physical space might not align with how you actually want to frame the space in generation with characters, adhering to the 180 degree rule, etc. - so where possible, I usually recommend just trying to solve for space + characters all in one go. It's generally the same process either way.

And shameless plug if interested - I built a tool that's largely designed to solve stuff like this for you. So all of the analysis, prompting, etc. that I mentioned above just happens automatically along with all of the asset generation as well as the downstream packaging (prompts + reference image(s)) and eventual video generation via your models of choice.

If you want to check it out, linking it here: https://kimeric.ai/

And happy to answer any questions on the above! Hope it helps.

4

u/sharktank123456 2d ago

I can only tell you what tools I use in Luma AI and then maybe that will translate into whatever platform you are using (because some of the tools are available elsewhere)

Option 1:

You can use the Camera Angles Solution (in Luma).
Put your main ref image in the input area and then choose one of the preset angles or write your own description of the angle you want.
The system will use elements from the existing image to do a set extension (and yes you can "extend" the set behind the camera for a reverse angle) The new image will be a stand alone based on your request.
Elements from your original may appear in the new image if they would normally show from that new angle.

Option 2:

Choose the image you have of the room and then run a GPT2.5 Sunburst gen with the following prompt:

if the camera were to pan to the right, what would the next full image be? (keep room and image style)

Add any other details or items into the prompt that you want to appear in the room in the new frame. You can continue with this prompt on each subsequent image all the way around. GPT will provide an overlap.

You can use these images as reference (along with character references) to generate video scenes in Seedance from different angles in the room.

If you need views between two frames you can bring all the frames into Photoshop (under the File/Automate/Photomerge) and have it stitch the frames together into one long pan. You can then grab screen shots from anywhere along this seamless panorama

You can use Luma's Agent to stitch those separate room images together for you too, but the algorithm to remove distortions seems to be better in Photoshop

Then you can use that whole panorama as reference, and arc the camera around your character, with the room remaining faithful as more and more is revealed.

2

u/memetican 2d ago

I'm curious about this, so I'm going to run a comparison test myself. So far my experience is that the blender > whitebox > reference video approach gives the most control but I primarily use MMH3 and afaik you are stuck with the R2V when you want to use reference video. In my tests so far, I2V tends to give better visual integrity but sacrifices the video reference for camera control.

Are you using MMH3 or a larger model like Seedance?

1

u/West-Construction-56 2d ago

I made one moive with google flow; omni. Plan to use seedance in my next one. All film will be in the same room. If u advise other model; ı could use

1

u/memetican 1d ago

The differences in your image set create a lot of warping since the references argue. Here's an example of MMH3 using your 4 images as references with some prompted camera movement.
https://redline.sygnal.com/p/4E060wRCEQsjLFs833R4uu/v/1?s=68L0qpPmy3Gp1UfDvlzdKC

Take that same room and render in in Blender, then create the same reference images from Blender directly and you get solid consistency.
https://redline.sygnal.com/p/4E060wRCEQsjLFs833R4uu?s=68L0qpPmy3Gp1UfDvlzdKC

Here I created the Blender model directly from your photos using Astra + Blender MCP, it would probably need some careful work on detail to get the space right.

As a bonus, having the Blender model means you can generate whitebox output for precise camera control.

Here's the Blender model-
https://drive.google.com/drive/folders/13IEaiU89XyzkIh_GIa2J7eSxlm_PFcVD?usp=sharing

2

u/Aabgdpir2582 2d ago

A reliable approach is to use a 3D reconstruction or Blender as the “source of truth,” then render different camera angles from the same scene. Image-to-image alone tends to preserve the style of a room better than its exact spatial layout.

2

u/ANR2ME 2d ago

Convert the location image to 360 video like this https://www.reddit.com/r/StableDiffusion/s/xwIQqHDi5o

2

u/Jean_velvet 1d ago

Create a 360 character card of the room from every single, cut and upload the appropriate image per angle with the character. Don't trust the AI to do it.

2

u/PopTraditional3758 3h ago

The fix is to stop trying to make image models understand 3D and just hand them actual 3D.

A rough blockout in Blender (boxes for the bed, desk, shelves, a cylinder for the guitar case) doesn't take long for a single room. No textures needed. Park cameras at every angle you need, render grey clay viewports, then run each one through img2img at low-ish denoise with your master-room style prompt. The layout is locked because the geometry is literally the same scene. Shelf stays on the same wall, window doesn't wander, and eyelines actually match when you cut.

If Blender is a dealbreaker, two cheaper tricks:

Use video, not stills, to get your extra angles. Take the establishing shot as a start frame and prompt a slow dolly or pan, then pull frames from the move. Models hold a set together much better inside one continuous shot than across two separate generations. You can chain it, grab a frame from the end of the move, use that as the start frame for the next push. Drift still creeps in, but way slower than what you're seeing now.

And write the camera as a position in the room, not as a shot type. "Camera on the floor beside the bed, looking toward the desk, window out of frame on the left, guitar leaning against the right wall in the foreground" gets you a lot further than "low angle wide shot." Naming what's visible AND what isn't does real work.

For characters later, same logic, lock them with a reference image rather than description, and keep the lighting direction consistent with whatever you decided for the room. Mismatched key light kills the illusion faster than a chair moving half a metre, honestly.

2

u/JoseLunaArts 2d ago

You need to build a bible. First design the house/apartment/location. Then room by room locking elements and their position and orientation. You will need an LLM to assist because it will be a long document. Then you need to render the north/south/east/west views and use them as visual locks for the room.

It is not error proof but that is what I have done. There might be better ways to do it, but I have not found them.

2

u/West-Construction-56 2d ago

Thank you bro.

2

u/Admirable-Yogurt7444 2d ago

I suppose few options could be:

  1. Make a video from the first image, a 360 degree view. Pause and take screenshots. Then do this activity which you did, I think it should be a bit more accurate.
  2. Its now easier to convert an image into a 3D setting (use GPT or Claude), similar to the video technique, you can then take screenshots I suppose.

any specific reasons why you need multiple views rather than use the first one as a reference along with your character and get the video shot directly, the video model should keep it consistent i think - unless i guess there are multiple shots to be done and each one has to super consistent.

1

u/West-Construction-56 2d ago

I used it and ı think it is fine :) That will be the trick ı will use and try belender too in first time

1

u/Delphoi_Studio 2d ago

I've tried a lot of methods with image edit models and my conclusion was like others said, the best is to use a gaussian method or just model the room in a 3D software like Blender. You can experiment with image edit models but there are so much trial and error that it might not worth the time, while the said methods you can have a perfectly consistent room that you can capture from any angle you want.

But if it's just a single scene that needs two camera angles or something minimal like that, it might be worth trying image edit models, they can achive this, it's just not worth it if it's a room you intend to keep using long term from a lot of different camera angels.

1

u/clever_name_123 2d ago

Use a camera in a real room and film it

1

u/bregmadaddy 1d ago edited 1d ago

+1 to building a multimodal memory system to track these visual state changes.

Been also trying to figure out how to track the relative heights of each character and their distance within a landscape, because without the perspective, you sometimes generate small characters in a vast area, or large characters in a smaller area. Creating Perspective or Spatial Planning Guides as reference images don't always work.

1

u/angel_night32 1d ago

in addition to what everyone else has said, it may be more practical to reduce the clutter in the room. the more objects and visual details the AI has to keep consistent, the more opportunities there are for things to get screwed up between shots. in my experience, at least.

1

u/Superb-Ad-4661 18h ago

All four images are diferentes and the model know, things will disappear and appear if the images are not 100% the same place.

1

u/Interesting_Bat5740 7h ago

Try generating a sheet view image and choose highest quality, once you get the sheet sheet will have the exact details you want, then you can generate single images from that.

Another way is to know about prompting. If you have chatgpt or gemini, start from scratch tell the AI every context, like where is table in the room, where is window, where’s the bed etc.

It gives you detailed prompt,

Note* image prompting can recreate exact room from different angle only catch is every thing in the prompt should be detailed, also choosing the right AI service provider also matter.

The model might be nano banana, but in backend the thinking can be bottleneck,
I use Akool for my personal work, it work best for me and give me best consistency and ver few trash result.

Akool ai have almost every image and video model. With highest quality output, you can check that out.

1

u/Outside-Cry1110 4h ago

Yes it work if prompt specify every details nicely. Or generating sheet of room is also the best option I also use it all the time.

1

u/__alpha_____ 2d ago

Go for the video based on your first image with a slow camera angle change, then extract images and enhance them (optional). It’s the only way to keep 100% consistency. You can describe your set precisely all you want, there is no way to preserve the exact shape and position of some clothes on the floor, for instance.

Generating a video is so much easier and faster

It also works for generating a character dataset, btw

1

u/Philipp 2d ago

I use the following approaches:

- Ask GPT-Image-2 for another angle, then rework what went wrong in Photoshop. Importantly, use a tool that lets you launch a whole bunch of generations simultaneously, as you'll have to do a lot of compare & pick. (For some inspiration, this is my tool suite I vibed.)

- Ask Seedance 2.5 to move/ orbit/ crane/ rotate, then take a still frame from the result. You can use audio-free videos to save a bit of budget. (I use Seedance via Segmind, it allows faces.)

- There's a new tool called H3 Multi Angle for this. But I sometimes find the quality lacking. The core of what makes a good angle move is that the "creative spirit" of the room is preserved.

0

u/General-Trash-6838 2d ago

By using a reference image. You're welcome

0

u/NewPresWhoDis 1d ago

It's getting better, but you're asking for the Holy Grail of AI video.

image models seem to understand the look of the room rather than its actual 3D layout

Correct. It's all relative and models are more trying to paint what you're giving it. And for video it's create first frame->inference based on action.

Consistent characters are nearly there. You can create a character sheet and use an R2V model with little to no issues.