The Thread: notes on shopping an outfit from a photo

How it works

Why does the search pick the wrong garment?

You wanted the jacket and it described the shirt underneath. That is a detection problem rather than a matching problem, and the two have completely different fixes.

Save to Pinterest

You upload a photo of someone in a coat. What comes back describes the shirt underneath it. Nothing about the search was broken, and running it again will give you the same thing.

This is worth understanding because two different failures look identical from the outside, and only one of them has a fix on the search side.

What is the difference between detection and matching?

Snagfit does two separate jobs in sequence, and knowing which one went wrong tells you what to change.

Detection is reading the photo and deciding what garments are in it. The output is a list, up to eight garments, each with a short written description covering the type of garment, its colour, its material, its pattern and its detail.

Matching is taking each of those descriptions, turning it into a four to eight word shopping search, and running that against shopping listings. The output is products.

The whole pipeline is laid out in how outfit search actually works, and the important part here is that the second step never revisits the photo. It only ever sees the words the first step wrote.

So if the description says "cream cotton shirt, long sleeves, button front" when you wanted the coat, no amount of good matching will save it. The photo has already been read and the coat is not in the sentence.

Why does a jacket lose to the shirt underneath it?

Because of how much continuous fabric each one shows.

Picture an open coat over a shirt. To your eye that is obviously a coat, worn over a shirt, and the coat is the outer and more important layer. To something reading the picture as regions of fabric, it is two narrow vertical strips at the left and right edges of the torso, with one large unbroken panel in the middle.

The shirt is a single continuous shape. The coat is two disconnected ones. The shirt shows more surface, more of it uninterrupted, and it sits in the centre of the frame. It reads as the more prominent garment, so it comes back first.

The same thing happens in other layering combinations:

None of these confuse a person for a second. All of them make a garment harder to read as one thing.

Why does the same garment get two different names?

Because the description follows what is visible, and different views show different things.

A garment photographed from the front, with a row of buttons down it, reads as a shirt. The same garment from behind, with no fastening in view and no collar detail, reads as a top. Neither reading is wrong about the picture. They are descriptions of two different pictures of one object.

This is a consequence of the whole approach. What an AI sees in an outfit photo sets out the five fields recorded per garment, and every one of them is a fact about the image rather than a fact about the product. Snagfit does not know what the garment is called in a catalogue. It knows what it looks like from here.

The practical effect: if you have a choice of frames, the front view with the fastening visible carries more information than the back view, and the search that comes out of it will be narrower.

Does the order of the results mean anything?

It does, and the ordering surprises people.

Results come back ordered by prominence, which tracks roughly how much of the frame each garment occupies and how clearly it reads. The first item is usually the biggest, clearest garment in the picture.

Prominence is a measure of the picture rather than of the outfit. A plain white t-shirt filling the middle of a frame will outrank a distinctive pair of boots at the bottom edge every time, because the t-shirt is larger and flatter and easier to read.

Which means the item you came for is often third or fifth in the list rather than first, and that is normal. Working through the whole list rather than stopping at the top one is the habit that shopping a whole outfit is built around.

How do you steer it towards the piece you want?

There is one control, and it is the crop.

Snagfit reads the image it is given. It has no way to know which garment you care about, no field where you can say "the coat", and no way to reweight the results after the fact. What it sees is the frame. So the way to make a garment more prominent is to remove the competition from the picture.

Crop so the piece you want fills most of the frame. A photo cropped to the coat, with the shirt mostly outside the edges, cannot come back describing the shirt, because the shirt is not really in the picture any more.

Two practical notes on this. First, leave a small margin. Cropping so tight that the shoulder line and the hem are both cut off removes the shape of the garment, and shape is one of the things being read. Second, keep something for scale where you can, particularly on small items.

The general habits are in how to crop a photo for a better match. The specific point here is that the crop is the only steering wheel there is.

Should one outfit be several uploads?

Yes, when you care about more than one piece.

One upload of a whole outfit is the right move for an overview: you get a list of everything visible and you find out what is there. It is fast and it costs one look.

But every garment in that upload is sharing the same 1280 pixels, and the description of each one is written from its share. A pair of shoes in a full-length shot gets a fraction of the frame, and the description reflects that.

If you have decided you want the coat and the boots specifically, two cropped uploads will describe both far better than one wide upload described either. That is the trade to think about, and it is why working through a camera roll is faster when you decide what you want before you start uploading.

Why does a busy background make this worse?

Because a background can look like fabric.

A patterned wall, a rail of other clothes, a sofa, a curtain, or another person standing behind the subject all present large regions of colour and texture. Some of that reads as garment-like, and the more of it there is, the more competition the actual clothes have for attention.

The clearest version of this is a shop or a wardrobe as a backdrop. A photo taken in front of a clothing rail contains dozens of garments, and only one of them is being worn. There is nothing in the picture that says which one matters.

Mirrors add their own version. A mirror selfie contains the person, the reflection of the person, and often the reflection of the room behind the camera, so the frame carries two copies of the outfit at different sizes and sharpness.

The fix is the same one as everything else in this post. Crop the background out. A garment against a plain wall reads better than the same garment against a rail of other clothes, and the crop is how you get from one to the other when you cannot change the photo.

Does a print or a graphic change what gets picked?

It makes a garment more prominent, sometimes at the expense of a plainer piece next to it.

A garment with a strong pattern presents more visual information than a plain one of the same size. A printed shirt under a plain black coat can outrank the coat even when the coat is larger, because there is more to describe on the shirt.

That works for you when the printed piece is what you want, and against you when it is not. If you are after the plain garment, the crop matters more than usual, because the printed piece will win any competition it is allowed to enter.

There is a separate problem with print, which is that patterns are harder to match than plain fabric once the search runs. A description like "green floral print" covers thousands of unrelated garments, so even correct detection can produce a loose result. That is a matching problem rather than a detection one, and it needs a different fix.

What if the garment is genuinely mostly hidden?

Then the honest answer is that it is not really in the photo.

If a jumper is worn under a coat and all you can see is the collar and the cuffs, there is not enough there to describe. Snagfit may register something at the neck, but the description will be thin, because thin is what the evidence is. A garment that is entirely covered will not appear at all.

This is the point where the photo has stopped being the right tool and something else takes over. Sometimes there is a better frame from the same source, which is usually true for video. Sometimes the person posted the same outfit elsewhere in a different pose. Sometimes the answer is to search the words yourself, which works better than people expect when you know what you are looking at.

Snagfit shows you what a photo contains and where those things are sold. It does not reconstruct what a photo left out, and a search that claimed to would be guessing.

Common questions

Why did Snagfit describe the wrong piece of clothing?

Almost always because the piece you wanted was partly hidden, small in the frame, or overlapping something else. Snagfit reads the garments it can see and returns up to eight, most prominent first. A jacket worn open over a shirt shows less continuous fabric than the shirt does, so the shirt can read as the more prominent garment even though the jacket is the outer layer.

What is the difference between a detection problem and a matching problem?

Detection is Snagfit reading the photo and deciding what garments are in it. Matching is taking each of those descriptions and searching shopping listings for it. If the description is of the wrong garment, that is a detection problem and the fix is the photo. If the description is right and the products are wrong, that is a matching problem and the fix is the search terms.

How do you make Snagfit look at the garment you actually want?

Crop the photo so the garment you want fills most of the frame. Snagfit reads what it is given, so removing everything else from the image is the only steering control there is. A photo cropped to a coat, with the trousers and the bag outside the frame, cannot produce a description of the trousers.

Why does a layered outfit confuse the detection?

Because layering breaks garments into pieces. An open coat is two vertical strips of fabric with a shirt between them, and a scarf across a jumper divides the jumper into an upper and lower section. Snagfit reads continuous regions of fabric, so a garment split into disconnected parts by whatever is worn over it presents less evidence than an unbroken one.

Does the order of the results mean anything?

Yes. Snagfit returns up to eight garments per photo, ordered by prominence, which mostly tracks how much of the frame each one occupies and how clearly it reads. The first result is usually the largest, clearest garment rather than the one you personally care about. That ordering is about the picture, not about what you were looking for.

Why does the same garment get named differently in two photos?

Because the name follows what the photo shows. A garment photographed from the front with the fastening visible may read as a shirt, while the same garment from behind, with no fastening in view, reads as a top. Snagfit describes what is visible rather than identifying a specific product, so two views of one garment can produce two descriptions and two different searches.

Can Snagfit tell what someone is wearing underneath?

Only what is visible. If a jumper is worn under a coat and 10 percent of it shows at the collar and cuffs, Snagfit may register something there, but the description will be thin because there is little to describe. A garment that is fully covered is not in the photo in any useful sense and will not appear in the results.

Is it better to upload one photo of a whole outfit or several crops?

Several crops, if you care about specific pieces. One upload of a whole outfit gives you a fast overview and a list of everything visible. A separate cropped upload per garment gives each one the full frame and a much more detailed description. Snagfit downscales every upload to 1280 pixels on the longest side, so a crop spends all of those pixels on the piece you want.

Keep reading

How it works

How do you find an outfit from an iPhone screenshot?

An iPhone screenshot is capped at your screen's resolution, not the video's, and it saves as HEIC by default. Both of those change what Snagfit, or any tool, can actually read.

6 min read
How it works

What does an AI see in an outfit photo?

Five fields per garment, eleven possible categories, and a search query four to eight words long. Here is the exact shape of what comes back, with a real example.

7 min read
How it works

Can you find clothes from a screenshot?

Yes, usually. Whether you find the exact piece or something close depends on three things about the screenshot, and you can check all three before you upload it.

7 min read

← Back to Snagfit