← The Thread: notes on shopping an outfit from a photo

How it works

What does an AI see in an outfit photo?

Five fields per garment, eleven possible categories, and a search query four to eight words long. Here is the exact shape of what comes back, with a real example.

Save to Pinterest

When Snagfit looks at a photo of an outfit, it does not come back with "a person wearing clothes". It comes back with a structured list, and every item on that list has exactly five fields.

Knowing what those fields are tells you why some photos work well and others do not.

What five things does the AI record per garment?

For every garment it finds, Snagfit records:

Name. A short, shoppable name. Not "shirt" but "camp collar linen shirt".

Category. One of eleven fixed options: Top, Bottom, Outerwear, Footwear, Dress, Sleepwear, Headwear, Bag, Accessory, Jewelry, Eyewear.

Colour. The main colour, in plain words. "Faded black" rather than a hex code.

Details. One short phrase on material, pattern, cut or a notable feature. "Heavyweight cotton, dropped shoulders."

Search query. Four to eight words, built from all of the above. This is the one you never see, and it is the one that gets sent to the shops.

What does one detected item look like?

Here is what one detected item actually looks like:

Name: chunky retro runner sneakers. Category: Footwear. Colour: cream and gum. Details: suede overlays, dad-shoe silhouette. Search query: chunky cream retro runner sneakers.

You see the first four on screen. The fifth runs the shopping search when you tap Snag it.

Why is the search phrase four to eight words?

Short enough to match real listings, long enough to be about one specific thing.

Two words returns thousands of unrelated results. A fifteen word sentence is so specific that no listing title contains all of it, and you get nothing. Four to eight is the range where a shopping search still has something to grip.

Why are the categories a fixed list?

Free-form categories drift. One photo returns "jacket", the next returns "outerwear", the next returns "coat", and the results stop being comparable.

Eleven fixed options keep it consistent. It also forces a decision on the pieces people most often get wrong, which is why the next section exists.

Which garments does the AI have to be careful not to confuse?

Some garment types are routinely flattened into whatever is closest, and the flattening ruins the search.

Sleepwear is the clearest example. A flannel pajama set is not a sweater, and searching for a sweater will not find it. Snagfit is instructed to separate sweaters, hoodies, sweatshirts, robes and pajamas by their actual construction: a matching top and bottom set, a drawstring or elastic waist, a button-front sleep shirt, a satin or fleece texture.

Prints get the same treatment. A candy cane print pajama pant stays a candy cane print pajama pant. It never becomes "red and white pants", because the print is the most searchable thing about it.

Will the AI guess a brand?

It will not guess a brand. If there is no visible logo and the piece is not distinctive enough to name, the brand is not in the photo, and inventing one would send you to the wrong shop.

It will not return more than eight items, and it puts the most prominent first. So a group photo returns the main pieces on the main subject rather than forty half-visible garments.

It will not invent anything when it sees no clothing. An empty result is the correct answer to a picture of a landscape.

Why is the AI told to give the same answer every time?

There is a setting called temperature that controls how much a model varies its answers. High temperature produces creative, unpredictable writing. Low temperature produces the same answer to the same input, over and over.

Snagfit runs the description step at a low temperature, because creativity is the opposite of what this job needs. You do not want a poetic description of a coat. You want the same coat described the same way every time, so the shopping search behind it is stable.

It is not perfectly repeatable even so. A model can still phrase a garment differently on two runs, "chunky retro runner sneakers" one time and "chunky retro trainers" another. The category and the colour stay put; the exact wording can move. That is one of three reasons the same photo can return different results.

Why must the AI answer in a fixed shape?

The model does not reply in sentences. It is required to answer as structured data, with every one of the five fields present for every item.

That requirement is enforced at the API level rather than by hoping the model complies. There is a schema attached to the request that says: return an object, it must contain a list called items, each item must contain name, category, color, details and search_query, and all five are required.

Two things follow from this. There is no paragraph of prose to parse, so nothing can be lost in interpretation. And an item cannot arrive half-formed with a name but no search query, because the shape would not validate.

It also means Snagfit never has to guess what the model meant. Either the data is in the field, or the field is not there.

Why is the most prominent item listed first?

The instructions ask for the most prominent items at the top of the list, and that ordering carries real information.

Prominence here means how much of the photo the garment occupies and how central it is to the subject. In practice, that lines up closely with what makes the outfit look the way it does. The oversized coat comes before the socks. The statement bag comes before the plain belt.

So if you are only going to search one thing, the first two items are where to look. If you are trying to work out what actually creates a look you liked, the top of the list is your answer, and shopping a whole outfit properly starts from there.

Why does it stop at eight items?

Eight is roughly where a real outfit ends and noise begins.

Past eight you start getting the second earring, a visible strap, a watch face, a button, and things half visible on someone standing behind the subject. None of those are things you can meaningfully shop for, and a list of twenty items where twelve are noise is worse than a list of six that are all real.

The cap also keeps the result readable on a phone. A list you have to scroll through three times to understand is not a better result, it is a longer one.

What is the details field for?

Of the five fields, details is the one people skim, and it is doing more work than it looks.

The name gives you the garment type. The category files it. The colour narrows it. Details is where the thing that makes this garment specific ends up: the suede overlays, the dropped shoulders, the utility pockets, the flat brim, the mesh back.

That phrase is what turns a search from "sneakers" into a search for one recognisable style of sneaker. When two garments in a photo are both black jackets, details is usually the only field that separates them.

If details comes back thin, that is a signal about your photo rather than about the garment. It means the model could see a shape and a colour but not the construction, which almost always means the piece was too small in the frame, too dark, or partly hidden.

How do the five fields diagnose your photo?

Once you know the five fields, the result becomes a diagnostic rather than just a list.

Too few items. Snagfit found two garments in a photo of a full outfit. Almost always a framing problem: the rest of the outfit was outside the crop, or the subject was small enough that only the largest pieces registered.

A garment in the wrong category. A jacket filed as a Top usually means the closure and the length were not visible, so it read as a layer rather than as outerwear. A different frame fixes this more reliably than searching again.

A vague colour. "Dark" instead of a real colour name means the lighting took the colour out. Find a brighter frame of the same outfit.

Thin details. As above, this is the clearest signal that the garment was too small in the frame.

Items you did not mean. In a group photo, Snagfit reads the main subject. If it came back describing someone else's outfit, crop to the person you meant.

None of these are failures of the search. They are the search accurately reporting what your photo contained, which is useful information if you read it that way.

Is the result useful if you buy nothing?

The garment name is plain English, and it transfers.

Once you know the piece is a "camp collar linen shirt", you can paste that into any search engine, a retailer's own site search, or a resale app. Plenty of people use Snagfit just to find out what the thing is called, then go shopping somewhere else entirely. That is a completely reasonable way to use it.

Common questions

What does Snagfit record about each garment it finds?

Five fields. The name of the item, its category, its main colour in plain words, a short note on material, pattern or cut, and a search query of four to eight words. The first four are what you see on screen. The fifth is what actually gets sent to the shops.

What clothing categories can Snagfit use?

Eleven. Top, Bottom, Outerwear, Footwear, Dress, Sleepwear, Headwear, Bag, Accessory, Jewelry and Eyewear. Every detected item is filed under exactly one of them, which keeps the search specific rather than returning general clothing. The fixed list also stops one photo returning "jacket" and the next returning "coat" for the same kind of garment.

How long is the search query Snagfit builds?

Four to eight words. That length is deliberate. Two words is too general and returns thousands of unrelated listings, while a fifteen word sentence is so specific that shopping listings match nothing at all. The query is built from the garment's name, category, colour and details, and it is the one field you never see on screen.

Does Snagfit identify the brand of a garment?

Only when the brand is genuinely visible or the piece is distinctive enough to be named. If there is no logo and the garment is not unusual, the brand is not in the photo, so Snagfit describes the garment rather than guessing a label. A guessed brand would send you to the wrong shop.

Why does Snagfit describe a print or theme so carefully?

Because a print is the most searchable thing about a garment. A candy cane print pajama set finds a much better result than a red and white pajama set. Snagfit is instructed to preserve any print, character or holiday theme exactly rather than flattening it into a plain item.

Does Snagfit tell the difference between a hoodie and a sweatshirt?

Yes, and it is told to do so specifically. Sweaters, hoodies, sweatshirts, robes and pajamas are separated by their actual construction rather than being lumped together as tops, because the wrong word sends the shopping search to the wrong shelf. A matching top and bottom set, a drawstring waist and a button-front sleep shirt are all read as different constructions.

What does Snagfit do if it cannot see any clothing?

It returns nothing and says so. Snagfit is instructed to return an empty result when there is no visible person or garment, because inventing an item would send you shopping for something that was never in your photo. An empty result is the correct answer to a picture of a landscape.

Can you use the descriptions somewhere other than Snagfit?

Yes, and it is worth doing. The garment name is plain English, so you can paste it into any search engine, a retailer's own site search, or a resale app. Getting the item named is often the most valuable part of the search, and plenty of people use Snagfit only to learn what a piece is called.

Keep reading

How it works

Why do two people in a photo confuse a search?

Group photos are the most common upload that comes back wrong. The search reads every garment it can see and has no way of knowing which person you meant.

8 min read
How it works

Why does the search pick the wrong garment?

You wanted the jacket and it described the shirt underneath. That is a detection problem rather than a matching problem, and the two have completely different fixes.

7 min read
How it works

Why does lighting change what a search finds?

Colour is one of the fields Snagfit records, and light decides what colour a garment appears to be. A navy coat under warm light reads as brown, and the search goes looking for a brown coat.

7 min read

← Back to Snagfit