The Challenge
If you’ve ever asked an AI image tool for a scene with several people in it, you’ll know the problem. The overall image looks great until you zoom in, and then the faces fall apart. Eyes that don’t quite match, a smile that’s slightly melted, features that look almost right but sit in the uncanny valley. The more people you ask for, the worse each face gets.
This isn’t a flaw you can prompt your way out of, and it isn’t a sign you’ve picked the wrong tool. It’s the direct result of how these models spend their attention, which we’ve written about separately. The short version is that every generation has a fixed budget, and when you ask for six faces at once, each face gets a sixth of the attention it would have received on its own. A face needs a lot of detail to look right, and a sixth of the budget simply isn’t enough.
The Fix
The fix is to stop asking the model to do the impossible thing, and instead give each face the full budget one at a time. Here’s the exact method we use. The examples below use Photoshop and Nano Banana, Google’s image model, because that’s the combination we reach for most often, but the technique works with any image editor that supports layers and masks, and any AI image tool that can regenerate a section to a high standard.
The principle underneath all of it is simple. A face that fills the whole frame gets the model’s entire attention. A face that’s one small part of a crowded scene gets a fraction of it. So you generate each detailed element as if it were the only thing in the picture, then bring it back into the scene at the right size. You’re not fighting the budget. You’re spending it six times instead of splitting it six ways.
Here’s how it goes in practice. Generate your base image first, and don’t worry about the faces yet. Get the composition right, the number of people right, the poses, the lighting, the overall mood. Accept that the faces will be poor. They’re placeholders at this stage, and trying to perfect them here is wasted effort. Then isolate the first face and crop in tight, bringing the face area into its own working space so the face fills the frame. Regenerate that face as a full-frame image, asking the AI to recreate the face at this larger size with the same face, the same pose, the same direction of gaze, and the same lighting as the original scene. With the face filling the frame, the model pours its full attention into getting the features right. Bring the new face back in as a smart object, scale it down into position, and mask and feather the edges so it dissolves into the surrounding image rather than sitting on top of it. Repeat for every element that matters.
When you’re done, you have an image where every detailed part was generated as if it were the only thing in the frame, then assembled into a single scene. The faces are sharp. The hands have the right number of fingers. The whole thing holds up under the zoom that would have exposed a single-shot generation immediately.
The trade off
It takes longer than pressing the button once and hoping, which is the honest trade. But it’s the difference between an image that looks impressive in a thumbnail and falls apart on inspection, and one that holds up when a client leans in. For anything that’s going to carry a brand, that difference is the entire point.
The insights
There’s a broader lesson buried in the method, and it’s the same one that runs through most good AI work. The tool is rarely the limiting factor. The way you direct it almost always is. The people getting consistently excellent results from these models aren’t the ones with secret access to better technology. They’re the ones who understand how the technology actually spends its effort, and who build their process around working with that rather than against it. The crowded-faces problem looks like a limitation of the model. It’s really a prompt-and-process problem with a process answer.
If you’re producing visual content at any kind of volume and the quality isn’t holding up the way you need it to, this is usually where the issue sits, and usually where the fix does too.
The method here uses Photoshop for the compositing and Nano Banana for the regeneration, but nothing about it is specific to those. Any editor with layers and layer masks will do the compositing. Any image model that can regenerate a tightly cropped element to a high standard will do the generation. The technique is about sequence and attention, not about a particular product.