You upload a photo of three friends and you wanted the middle one's coat. What comes back is somebody else's shirt, somebody else's bag, and one of the three coats, possibly not the one you meant.
Nothing malfunctioned. The search read the picture and returned the garments in it, which is what it does. The problem is that you knew which person you meant and the picture does not say.
Why can you not point at a person?
Because the image is the entire instruction.
Snagfit takes a photo and returns the garments in it. There is no field for indicating a subject, no way to tap a person before the search runs, and no way to reweight the results afterwards. The frame you upload is the whole of what the search knows about what you want.
That sounds like a gap and it is a consequence of the approach. The pipeline in how outfit search actually works starts with reading the image and ends with a shopping search built from the words it produced. Nothing in that chain has a concept of "the one on the left".
Which leaves the crop as the only steering control, the same conclusion reached in why the search picks the wrong garment for a different reason.
How do the result slots get shared?
Snagfit returns up to eight garments per photo, ordered by prominence.
With one person in frame, eight slots covers a full outfit comfortably: a coat, a top, a trouser, shoes, a bag, and still room for accessories.
With two people, those eight slots are split across two outfits. Each person effectively gets four. With three people, roughly two and a half each, and the smaller items on everybody drop off the list entirely.
Prominence is what decides who gets more, and prominence mostly tracks size in the frame and clarity. So the person standing closest to the camera, or in the best light, or wearing the largest and flattest garments, takes more of the list. That ordering has nothing to do with who you were looking at.
The practical effect: in a group photo, the item you came for is often simply not in the results. Not described badly, not there.
Why does each garment get described worse?
Because resolution is shared the same way the slots are.
Snagfit downscales every upload to 1280 pixels on the longest side. That is the budget, and it is fixed regardless of how many people are in the picture.
One person filling a frame gets all of it. Three people across a frame get about a third each. A coat that would have been described down to its lapel shape and button count becomes "black coat", because a third of the available pixels is what the description was written from.
Thinner description, wider shopping search, more results that are less precise. The same arithmetic that makes sunglasses hard to find in a full-length shot applies to every garment in a group photo.
What does overlap do?
It fragments garments, and fragmented garments read poorly.
When one person stands partly in front of another, the person behind is divided into disconnected patches: a shoulder here, an arm there, a slice of coat between two heads. Snagfit reads continuous regions of fabric, so a garment broken into three separate patches presents far less evidence than the same garment whole.
The result is predictable. The person in front is described well and the person behind is described badly, or not at all, regardless of which one you cared about.
Arms around shoulders make it worse, because they cut across the torso of the person being held. Crowded shots where people are shoulder to shoulder make it worse again.
Can it merge two people's clothes into one description?
It can, in a specific situation.
When two adjacent garments are similar in colour and there is no clear boundary between them, for example two people in black standing shoulder to shoulder against a dark background, the region reads as one continuous area of fabric. The description that comes out describes a garment that does not exist, because it is half of one thing and half of another.
This is rare in good light and common in bad light. Contrast is what defines the edge between two garments, and a dim photo has less of it.
Same reason a garment against a background of the same colour gets read as larger than it is: the search sees a shape and the shape includes things that are not the shape.
What about clothes nobody is wearing?
They count, and this catches people out.
A garment in the frame is a garment whether or not someone is in it. A coat over the back of a chair, a jumper draped on a shoulder, a rail of clothes behind the subject, a folded pile on a bed, a mannequin in a shop window: all of these are read.
The clearest case is a photo taken inside a shop or in front of a wardrobe. The frame can contain twenty garments and only one of them is the outfit. There is nothing in the picture that indicates which.
A jacket held in the hand or slung over an arm is the trickiest version, because it is genuinely part of the outfit and it is also crumpled into a shape that does not look like a jacket. Those tend to come back either missing or described as something else.
What is the fix?
One person per upload, and crop everything else out.
Crop so the person you want fills the frame from head to foot, or from head to whatever the lowest garment you care about is. Everything outside that can go: the other people, the rail behind, the chair with the coat on it.
That single change gives you the full eight slots, the full 1280 pixels and no competition. It is the same crop discipline in how to crop a photo for a better match, applied to the case where it makes the biggest difference.
If you want pieces from two different people's outfits, that is two uploads. Each one gets a full read. Worth deciding before you start, since each upload uses a look.
Is a mirror selfie a group photo?
In effect, often yes.
A mirror shot can contain the person, their reflection, and part of the room behind the camera. That is two copies of the same outfit at different sizes and sharpness, plus whatever is on the walls and the floor.
If the phone is in the frame it covers part of the outfit, usually the chest or the face, which fragments whatever garment is behind it.
The fix is the same. Crop to the sharper of the two copies of the outfit, which is usually the reflection rather than the direct view, and drop the room.
Does a photo of one person with a pet or a child count?
Yes, and animals are the odd case.
A dog in a coat is wearing a garment, and it reads as one. A pet without clothing does not produce a garment, but it does occupy frame and it does provide texture that can be read as fabric, particularly with long fur against a similarly coloured background.
Children in the frame are straightforwardly a second person, with the added complication that childrenswear and adult clothing look alike at small sizes. A child's coat in the corner of a frame can produce a search that lands in the children's section of a retailer, which is confusing when you were looking for something for yourself.
Neither is a problem worth thinking about much. Both are solved by the same crop that solves everything else in this post.
Does the number of people change how many looks it costs?
No. One upload is one look regardless of how many people are in the picture, how many garments come back, or how good the result is.
Which cuts both ways. A group photo costs the same as a single-person photo and returns worse information per person, so it is poor value. Two cropped uploads cost two looks and return far more.
If what you want is a quick sense of what is in a picture, one upload of the whole group is a reasonable use of a look. If you have decided you want a specific garment, crop first, because the cropped upload is the one that will actually find it and the uncropped one may cost you a look for nothing.
That is worth knowing before you work through a backlog of saved photos, since a camera roll is full of group shots. Shopping your camera roll covers how to triage a batch, and the first cut is usually deciding which photos have one person in them.
What if the photo cannot be cropped?
Sometimes the people are genuinely too close together to separate.
Then you are working with what the frame gives you, and the honest expectation is a wider result. Two things still help.
Crop as tight as the picture allows, even if the crop includes part of someone else. Half a neighbouring person is better than a whole one.
And look for a different frame. If the photo came from a video, there is almost always a moment where the group is not overlapping, or where one person is alone in shot. That frame is worth scrubbing for, and it will produce a better result than any amount of work on the crowded one.