Upload both pictures, say how they should meet, and the model draws a single frame. You can combine two images the blunt way — put these people side by side — or the interesting way, where one picture becomes the material the other is made of.
Enter a prompt on the left to start generating your AI image.




Notes from our own runs, failures included. The instruction matters more than the upload order.
"Use the first image as the subject and the second as the background" is the single highest-value sentence. Without it the model decides, and it changes its mind between runs — the main reason people think the tool is inconsistent.
"Standing in front of", "reflected in", "filled with the texture of" all give clean results. "Merge", "blend" and "combine" on their own are too abstract, and you get an average of the two pictures rather than a composition.
Combine two images shot in similar light and they merge convincingly even if the subjects are unrelated. Two pictures of the same person in very different light do not. Check the light before you check anything else.
Facial structure carries through both inputs well. Logos, watch faces, signage and any readable text get rewritten into something that looks like text but is not. Plan around it rather than fighting it.
This is generative, not a layer stack. Run the same pair twice and the composition shifts, as the four passes above show. If you need one exact result, generate a batch and pick, rather than expecting reproducibility.
Three or more sources start losing whichever one you described least. When you need a whole group, the family photo page handles that case with a prompt built for it.
Pre-filled in the generator above and printed here so you can edit rather than guess.
"Combine the two uploaded images into one photograph: use the first image's subject as the silhouette and the second image's scene as the texture filling that silhouette, clean white background, soft edge falloff where the silhouette meets the background, cohesive colour grade across the whole frame, high-detail double exposure."
"Cohesive colour grade across the whole frame." Two photographs almost never share a colour temperature, and without this line the seam between them stays visible no matter how well the shapes line up.
"Use the first image's subject as the silhouette." Swap it for "place the person from the first image into the scene from the second image, standing on the path, matched to its lighting" and the same tool does an entirely different job.
The soft edge falloff. It looks decorative and it is the thing preventing the hard cut-out look that gives away a bad composite.
The same generator, three different instructions. Knowing which one you want saves the most credits.
The most common request: a portrait plus a place. Name where they stand and what the light is doing, and the result reads as a photograph. This is also the case covered step by step on our add-a-person guide.
Two separate portraits into one frame. When you combine two images of people, state the arrangement and the scale — "both at the same distance from the camera" prevents the classic result where one person is subtly a giant.
The double-exposure job shown above, where a silhouette is filled with a landscape. It is the most forgiving of mismatched inputs and the most likely to produce something you would actually print.
Two things worth stating plainly.
The four results above came from our own pipeline, but we have not published the two input photographs alongside them. Until we do, read the gallery as evidence of output quality and run-to-run variation, not as a verified before-and-after.
The person in the portrait does not exist. We would rather generate a face than use somebody's real photograph to advertise a tool.
typical time per result in our runs
~20s
inputs is the reliable maximum
2
free credits on signup
50
one sentence replaces the masking work
No layers
The step-by-step walkthrough of the most common two-input job.
When it is more than two people — built for whole groups.
Repair a damaged source before merging; damage carries through.
Push the finished frame to print resolution.
Direct answers, including the cases that do not work.
Yes. To combine two images the model draws a new picture rather than stacking layers. You upload both, describe how they should relate — one in front of the other, one reflected in the other, one filling the other's silhouette — and get a single frame back in about twenty seconds.
Upload both, then write one sentence that names which image is the subject and which is the setting, plus the relationship between them. Naming the roles explicitly is what stops the model reassigning them differently on every run.
Facial structure carries through both inputs reliably when the source photos are reasonably sharp. Fine detail does not: jewellery, watch faces, logos and any readable text get rewritten into approximations. Judge the result on the face, not the lettering.
Because it is generative. Each attempt to combine two images redraws the frame rather than reapplying a saved operation, which is why our four passes above differ. If you want one specific result, produce several and choose, rather than forcing reproducibility.
You can try, but quality falls away past two, and whichever source you described least is the one that gets dropped. For groups of people specifically, the family photo generator uses a prompt written for multiple inputs and holds up much better.
Signing up here gives 50 free credits, enough to combine two images several times and compare. After that it is credit-based, and the pricing page states the per-run cost rather than hiding it behind a trial.
Mismatched light, more than anything else. Combine two images shot in very different light — a flash-lit subject dropped into a golden-hour scene — and it never quite settles, no matter how good the prompt is.
That is exactly what the gallery at the top of this page shows — a portrait silhouette filled with a mountain landscape. Ask for the silhouette-and-texture relationship explicitly, because the model does not choose that reading on its own.
Only because the prompt refers to it. Say first and second image explicitly in your instruction and the order becomes meaningful; leave it out and the order is effectively ignored.
The image is yours to use. The usual limit applies to the inputs rather than the output: use photographs you have the right to use, and do not merge identifiable people who have not agreed to appear in the result.
Upload both pictures, name which is the subject and which is the setting, and get a single frame back in about twenty seconds. Free credits on signup, prompt already filled in.