You upload a photo of someone in a coat. What comes back describes the shirt underneath it. Nothing about the search was broken, and running it again will give you the same thing.
This is worth understanding because two different failures look identical from the outside, and only one of them has a fix on the search side.
What is the difference between detection and matching?
Snagfit does two separate jobs in sequence, and knowing which one went wrong tells you what to change.
Detection is reading the photo and deciding what garments are in it. The output is a list, up to eight garments, each with a short written description covering the type of garment, its colour, its material, its pattern and its detail.
Matching is taking each of those descriptions, turning it into a four to eight word shopping search, and running that against shopping listings. The output is products.
The whole pipeline is laid out in how outfit search actually works, and the important part here is that the second step never revisits the photo. It only ever sees the words the first step wrote.
So if the description says "cream cotton shirt, long sleeves, button front" when you wanted the coat, no amount of good matching will save it. The photo has already been read and the coat is not in the sentence.
Why does a jacket lose to the shirt underneath it?
Because of how much continuous fabric each one shows.
Picture an open coat over a shirt. To your eye that is obviously a coat, worn over a shirt, and the coat is the outer and more important layer. To something reading the picture as regions of fabric, it is two narrow vertical strips at the left and right edges of the torso, with one large unbroken panel in the middle.
The shirt is a single continuous shape. The coat is two disconnected ones. The shirt shows more surface, more of it uninterrupted, and it sits in the centre of the frame. It reads as the more prominent garment, so it comes back first.
The same thing happens in other layering combinations:
- A scarf across a jumper splits the jumper into an upper and a lower section.
- A bag strap running diagonally across a top divides the top in two.
- Crossed arms cover the middle of a shirt and leave two sleeves.
- A long cardigan worn open does exactly what the coat does.
- A blazer over a dress can turn one dress into a skirt-shaped fragment.
None of these confuse a person for a second. All of them make a garment harder to read as one thing.
Why does the same garment get two different names?
Because the description follows what is visible, and different views show different things.
A garment photographed from the front, with a row of buttons down it, reads as a shirt. The same garment from behind, with no fastening in view and no collar detail, reads as a top. Neither reading is wrong about the picture. They are descriptions of two different pictures of one object.
This is a consequence of the whole approach. What an AI sees in an outfit photo sets out the five fields recorded per garment, and every one of them is a fact about the image rather than a fact about the product. Snagfit does not know what the garment is called in a catalogue. It knows what it looks like from here.
The practical effect: if you have a choice of frames, the front view with the fastening visible carries more information than the back view, and the search that comes out of it will be narrower.
Does the order of the results mean anything?
It does, and the ordering surprises people.
Results come back ordered by prominence, which tracks roughly how much of the frame each garment occupies and how clearly it reads. The first item is usually the biggest, clearest garment in the picture.
Prominence is a measure of the picture rather than of the outfit. A plain white t-shirt filling the middle of a frame will outrank a distinctive pair of boots at the bottom edge every time, because the t-shirt is larger and flatter and easier to read.
Which means the item you came for is often third or fifth in the list rather than first, and that is normal. Working through the whole list rather than stopping at the top one is the habit that shopping a whole outfit is built around.
How do you steer it towards the piece you want?
There is one control, and it is the crop.
Snagfit reads the image it is given. It has no way to know which garment you care about, no field where you can say "the coat", and no way to reweight the results after the fact. What it sees is the frame. So the way to make a garment more prominent is to remove the competition from the picture.
Crop so the piece you want fills most of the frame. A photo cropped to the coat, with the shirt mostly outside the edges, cannot come back describing the shirt, because the shirt is not really in the picture any more.
Two practical notes on this. First, leave a small margin. Cropping so tight that the shoulder line and the hem are both cut off removes the shape of the garment, and shape is one of the things being read. Second, keep something for scale where you can, particularly on small items.
The general habits are in how to crop a photo for a better match. The specific point here is that the crop is the only steering wheel there is.
Should one outfit be several uploads?
Yes, when you care about more than one piece.
One upload of a whole outfit is the right move for an overview: you get a list of everything visible and you find out what is there. It is fast and it costs one look.
But every garment in that upload is sharing the same 1280 pixels, and the description of each one is written from its share. A pair of shoes in a full-length shot gets a fraction of the frame, and the description reflects that.
If you have decided you want the coat and the boots specifically, two cropped uploads will describe both far better than one wide upload described either. That is the trade to think about, and it is why working through a camera roll is faster when you decide what you want before you start uploading.
Why does a busy background make this worse?
Because a background can look like fabric.
A patterned wall, a rail of other clothes, a sofa, a curtain, or another person standing behind the subject all present large regions of colour and texture. Some of that reads as garment-like, and the more of it there is, the more competition the actual clothes have for attention.
The clearest version of this is a shop or a wardrobe as a backdrop. A photo taken in front of a clothing rail contains dozens of garments, and only one of them is being worn. There is nothing in the picture that says which one matters.
Mirrors add their own version. A mirror selfie contains the person, the reflection of the person, and often the reflection of the room behind the camera, so the frame carries two copies of the outfit at different sizes and sharpness.
The fix is the same one as everything else in this post. Crop the background out. A garment against a plain wall reads better than the same garment against a rail of other clothes, and the crop is how you get from one to the other when you cannot change the photo.
Does a print or a graphic change what gets picked?
It makes a garment more prominent, sometimes at the expense of a plainer piece next to it.
A garment with a strong pattern presents more visual information than a plain one of the same size. A printed shirt under a plain black coat can outrank the coat even when the coat is larger, because there is more to describe on the shirt.
That works for you when the printed piece is what you want, and against you when it is not. If you are after the plain garment, the crop matters more than usual, because the printed piece will win any competition it is allowed to enter.
There is a separate problem with print, which is that patterns are harder to match than plain fabric once the search runs. A description like "green floral print" covers thousands of unrelated garments, so even correct detection can produce a loose result. That is a matching problem rather than a detection one, and it needs a different fix.
What if the garment is genuinely mostly hidden?
Then the honest answer is that it is not really in the photo.
If a jumper is worn under a coat and all you can see is the collar and the cuffs, there is not enough there to describe. Snagfit may register something at the neck, but the description will be thin, because thin is what the evidence is. A garment that is entirely covered will not appear at all.
This is the point where the photo has stopped being the right tool and something else takes over. Sometimes there is a better frame from the same source, which is usually true for video. Sometimes the person posted the same outfit elsewhere in a different pose. Sometimes the answer is to search the words yourself, which works better than people expect when you know what you are looking at.
Snagfit shows you what a photo contains and where those things are sold. It does not reconstruct what a photo left out, and a search that claimed to would be guessing.