Blog
September 21, 2026

AI Photo Editing: What Works, What Fails, How to Prompt It

What AI actually does well with photos, and what it still does badly

Short version: models are reliable at replacing backgrounds, removing unwanted objects, recoloring items, changing lighting "by eye" and raising resolution. These are tasks where one or two runs give an acceptable result without masks or layers.

Things get worse when you need pixel-level precision: keeping a face untouched during a heavy edit, fixing one element without disturbing the rest of the frame, bringing detail back into a very small image. A model does not "edit" a photo the way Photoshop does. It re-synthesizes the image from your description plus the source. So the skin tone can drift along with the background, and the shape of the nose can drift along with the wrinkles.

The practical conclusion is simple. Break the job into small edits, check each one, and do not expect a single sentence to fix the whole frame.

Three kinds of tools: editors, image-to-image models, upscalers

Editors with AI features. A desktop or browser photo editor where the model hides behind a button: "remove object", "expand frame", "denoise". The upside is control: you outline the area yourself and nothing else changes. The downside is that you cannot go beyond the features that exist, and you often have to install software.

Image-to-image models. You upload a photo and describe the change in words. There is almost no limit on the task: swap the background, put an object in someone's hand, change the time of day. The price you pay is that nothing guarantees the untouched part of the frame stays as it was. The Nano Banana family is a typical example of this class.

Upscalers. A separate category: they do not change content, they raise resolution and fill in small detail. They work without a prompt, which makes them the most predictable of the three.

The workflow that holds up in practice: meaning first (image-to-image), upscale second. The reverse makes no sense, because the model will rebuild the picture at its own resolution anyway.

Five things to check when picking a tool

Five properties separate photo editing tools from each other. Look at these rather than at generic "quality".

  • Image input. The model has to accept images. A pure text-to-image model will never see your photo, no matter how well you describe it.
  • Number of input images. One shot is enough for a simple fix. Several matter when you have a background reference, a product shot, a style sample or a second angle of a face. Limits differ: some models take up to 10 images, a few up to 14.
  • Fidelity to the source. How well the model holds composition, light and facial features. You can only test this empirically, on your own frame.
  • Output resolution. Print and large banners need an upscale; some models advertise 4K support.
  • Cost per run. The gap in credits between a cheap model and a flagship can be several times over, and that directly decides how many attempts you can afford.

Photo editing models available on Moleculs

Every model listed below accepts an image as input, which means it can work on an existing frame and not only generate from scratch. Credit prices are as of publication; the current number is always shown next to the model in the interface.

ModelCredits per runInput imagesWhat people use it for
Nano Banana Editor8accepts imagesreference-based edits in image-to-image mode
ChatGPT 5 Image Mini8up to 10fast cheap tests, drafting the prompt
Nano Banana 2 Lite14up to 10the fastest of the family, around 4 seconds per frame
Nano Banana20up to 10edits with flexible aspect ratios
Nano Banana 221up to 104K, text inside the image, multi-image editing
Midjourney23up to 5artistic restyling of a shot
Nano Banana Pro64up to 14complex scenes, many references at once
ChatGPT Image 267up to 10careful adherence to a long, detailed prompt
Topaz Image Upscalesee catalog1up to 8x enlargement without changing content

[[МОДЕЛИ: Nano Banana Editor | Model list scrolled to the Nano Banana family]]

A note on text models. They do not draw, but they can look at a picture and describe it in words: Gemini 3.1 Pro accepts up to 16 images, Claude Sonnet 5 up to 20. Useful when you want a description of the frame, a list of defects, or a prompt you will then hand to an image model.

A step-by-step workflow

An order of operations that saves both time and credits. Cheap first, expensive later.

  1. Prepare the source. Crop the excess, straighten the horizon, pick a frame where the subject takes up enough space. The model cannot invent what is not in the pixels.
  2. Formulate one edit. Not "make it look good", but "replace the background with light gray". Everything else is the next step.
  3. Run a draft on a cheap model. An 8-credit run tells you whether the model understood the prompt, before you spend 60-plus.
  4. Move the working prompt to a stronger model. The wording is already tested, so you need fewer attempts.
  5. Compare with the original at 100% zoom. Watch the face, hands, any lettering, object edges, fabric patterns.
  6. Upscale last. Only once you are happy with the content of the frame.
  7. Save the prompt that worked. In a week you will not remember it, and it will come in handy on a similar shot.

You can test the whole path on a single photo in a couple of minutes: pick an image model, upload your shot, paste a background replacement prompt, run it.

[[СКРИНШОТ: /dashboard?type=image | Prompt field and image model picker]]

How to write an editing prompt

An editing prompt differs from a generation prompt in that half of the text is about what must not change. By default the model treats any part of the frame you stay silent about as fair game.

A four-line frame that works:

Source: the attached photo, edit it, do not create a new image.
Task: [one specific change].
Keep unchanged: [face, pose, lighting, skin tone, clothing, background].
Format: same aspect ratio and crop as the original.

Filled in for a background swap:

Source: the attached photo, edit it, do not create a new image.
Task: replace the background with a flat light gray, no gradient, no texture.
Keep unchanged: face, hairstyle, clothing, pose, direction of light,
shadows on the face and neck, skin tone, edges of the silhouette and hair.
Format: same aspect ratio and crop as the original.

And for object removal:

Source: the attached photo.
Task: remove the cable in the upper left part of the frame, rebuild the sky behind it.
Keep unchanged: everything else, including the tree canopy on the right,
the clouds, brightness and contrast.
Format: same frame dimensions.

Three habits that visibly raise the hit rate. Name the object by its position in the frame ("to the left of the shoulder"), not only by type. Ask for one change per run. Write your constraints as concrete nouns, because "do not change anything else" lands worse than an explicit list.

Portrait retouching: where to be careful

Portraits are the touchiest case. The human eye is trained on faces, so a two percent shift in the distance between the pupils reads as "that is not the same person", even though the frame is technically clean.

What lowers the risk:

  • A shot where the face takes up at least a quarter of the height. On a small face the model simply redraws the features from its own idea of a face.
  • Soft wording: "remove isolated blemishes and skin shine", not "give perfect skin".
  • An explicit ban: "do not change the shape of the eyes, nose, lips, face oval, eyebrows, hairline".
  • Side-by-side checking: original and result next to each other, eyes and teeth first.

Order matters too. Technical fixes first (noise, exposure, background), then skin, then style. If you ask for a new background, retouching and different clothes in one go, the model averages all of it at once.

One more thing: if you have no source of usable quality at all, it is sometimes more honest to build a new image from scratch instead of rescuing the old one. Editing a dead frame usually eats more attempts than generating the scene fresh.

Upscaling and "recovering" detail

An upscaler does not restore lost information, it statistically invents something plausible. On a photo 2000 pixels wide this is nearly invisible. On a 300-pixel frame the result will be sharp and partly made up.

Where upscaling helps:

  • The shot is fine, but you need it for print or a large screen.
  • Resolution dropped after an image-to-image edit.
  • A scan or a film frame: noise goes away, grain smooths out.

Where it only adds noise: heavy motion blur, small faces in a crowd, fine print on packaging, a badly compressed JPEG with block artifacts. Topaz Image Upscale goes up to 8x, but 8x does not mean "eight times better": from 4x onward, check eyes, teeth, lettering and the edges of hair. Often 2x looks more believable than the maximum.

Common mistakes and what to do when a model ruins the shot

  • Ten edits in one prompt. Split them into steps and look at the result between each.
  • No constraints in the prompt. Add a "keep unchanged" line with a list.
  • Going straight to an expensive model. Run drafts on cheap models.
  • Upscaling before editing. The order is the other way around.
  • The target area is too small. Crop in, fix the fragment, put it back in an editor.
  • Judging the result from a preview. Defects show up at 100% zoom, not before.

If a frame is ruined, do not rewrite the prompt in a loop. Go back to the original and start again with sharper wording: edits stacked on edits accumulate distortion, because every next run starts from an already damaged image. Sometimes it is enough to switch models and keep the same text, since different models read the same sentence differently.

What it costs and how not to burn your quota

A subscription spends credits, each model has its own cost per run, and that cost is shown next to the model. The spread is wide: an 8-credit edit and a 67-credit edit differ by a factor of eight, and for testing a prompt there is almost no difference between them. Quotas and what is included are listed on the pricing page.

The arithmetic is easy. A monthly quota of 2,000 credits is roughly 250 runs of a cheap editing model, or around 30 runs of a flagship. The quota refreshes every paid period.

[[СХЕМА: three draft runs on a cheap model, then one final run on an expensive one, then an upscale | A workflow with many attempts and modest spend]]

Three habits that genuinely save quota:

  1. Work out the wording on the cheapest model that accepts images.
  2. Keep your own library of prompts for recurring jobs: background, object, light.
  3. Do not launch an expensive model until you know what result you want. Half of all overspending goes on "just show me something" attempts.

All of the models above sit behind a single interface and one subscription on Moleculs, so switching between a cheap draft and a flagship run is a matter of picking another model from the list.

Corporate access to AI models

Invoice for legal entities, centralised payment, priority support

  • Access to ChatGPT, Gemini, Grok, Claude and DeepSeek
  • Prompt library and shared access inside the team
Request an invoice