Trendo Fashion

Methodology

What an AI outfit check can — and cannot — judge

An AI outfit check describes what a photograph shows and what that arrangement tends to do to the eye. It does not establish whether an outfit is correct, because correctness is not a property a photograph contains.

The distinction sets the honest boundary of the category. A model can observe that two colors sit close together in value, that a hemline cuts across the widest point of a silhouette, that one garment is the darkest mass in the frame. It cannot observe whether you like any of it, whether the room you are walking into expects something else, or how the fabric feels after an hour.

What follows is a description of what our own outfit check actually evaluates, what it cannot know, and where we have deliberately stopped it from answering. We are publishing it partly because we had to work the boundary out in order to build the thing, and partly because two of the more useful things we learned were failures.

Photo, context, analysis, output, decision

It helps to separate four different kinds of information, because they carry very different reliability.

  • The photo supplies visible evidence: garments, colors, where lines fall, how volume is distributed, how the pieces sit relative to each other.
  • The context is supplied rather than inferred. The occasion is chosen by you — the system does not guess where you are going. Weather is gathered by the system, not read from the image. The calendar season, month and date are passed in. A style profile from onboarding carries your style personality, the colors you gravitate to, the colors you avoid, any proportion or comfort considerations you have recorded, and your lifestyle needs. If you attach a specific question, that travels too: “is this smart enough?” is a different request from “does this work at all?”
  • The analysis interprets the photo through that context across four areas.
  • The output is an explanation of what produces an effect.
  • The decision is yours, and the output is written to keep it there.

The context layer matters more than it sounds. For a period, the stylist received no notion of time whatsoever — no date, no season, no year. It judged a heavy coat in July exactly as it would in January, and could not self-correct even in principle, because nothing in its input told it what month it was. Adding hemisphere-adjusted season, month and date did not refine the analysis. It closed a hole. Where those values are not available, the seasonal line is left out rather than guessed.

What the check runs on

One still photograph. Not a video, not several angles — a single image, taken with the camera, chosen from the gallery, or rendered from a saved wardrobe outfit.

That is a real boundary and worth stating plainly rather than apologetically. A still frame carries a great deal about color, proportion and arrangement. It carries nothing about movement, fabric hand, comfort, hidden construction or true sizing. Everything the analysis can legitimately say has to survive being derived from one frame plus the supplied context.

A separate recognition pass identifies the individual garments — category, subcategory, dominant colors, and a brand where one is legible. Categories come from a fixed vocabulary rather than free text, which matters shortly. If recognition fails, the check still returns; it simply knows less about the individual pieces.

The four areas

The analysis returns an overall written summary plus four areas, each as a set of named observations with explanations:

  • Style and aesthetic — whether the pieces read as belonging to one idea, and what register the outfit sits in.
  • Color and pattern — the relationships between the colors present: what contrasts, what recedes, what competes, how a pattern behaves against a plain field.
  • Visible fit and proportion — what the photograph actually shows. Where lines break, where volume sits, how the silhouette divides. Not fit in the tailoring sense, which a photograph cannot establish.
  • Weather and occasion suitability — whether the clothing is plausible for the temperature supplied and the occasion chosen.

All four are describable from pixels plus supplied context, which is why they are the four. On the free tier the summary and category-level assessment are returned without the per-observation explanations; the fuller written reasoning and the styling suggestion belong to Trendo Fashion Pro.

Where we took a judgement away from the model

The weather assessment used to be generated as free text along with everything else, and it produced a percentage. That number read as precise, but nothing in the model’s input constrained it — there was no mechanism forcing the figure to correspond to the actual relationship between a garment’s warmth and a real temperature. A heavy layer in genuine heat could still be described as a good match.

So that number is no longer the model’s to give. It is computed instead, from the recognized garments’ warmth against the supplied temperature. Warmth sits on a fixed scale and is read from the recognizer’s fixed category vocabulary rather than from free text, so the result cannot drift with phrasing.

Two decisions inside that calculation are worth naming, because both are refusals to fabricate:

  • Warmth is judged from the upper and core layer only. Trousers, skirts, shorts, shoes and hats are excluded. A sweater is what makes you hot; jeans are worn year-round, and letting them contribute would make a T-shirt read as too warm for summer.
  • Where the inputs do not support a number, none is produced. If no weather was supplied, the metric reads unavailable. If weather is present but no upper-body garment could be classified, the model’s own wording is left untouched rather than replaced by a worse guess.

The transferable principle: when a language model produces a figure that nothing in its input constrains, the fix is usually not a better prompt. It is to stop asking it for that figure and compute it from something real.

What a photograph cannot carry

Each of these follows from the input rather than from caution.

  • Whether it fits. One pose, one moment. A photograph cannot establish size, cut, or how a shoulder seam behaves when you move.
  • How it feels. Fabric hand, weight, whether a waistband digs in after an hour — none of it survives being turned into pixels.
  • Material and quality. An image can suggest texture. It cannot confirm fiber content or construction.
  • Whether a dress code is genuinely met. The system knows the occasion you selected. It does not know your workplace, the venue, or the unwritten rule everyone there already understands.
  • Cultural and personal meaning. What a color signifies at a particular event, or what a garment means to the person wearing it.
  • Your taste. A recorded preference is not the same as what you will want to wear on a given morning.

Trendo Fashion is also designed to analyze the visible outfit rather than to judge the person wearing it. The subject of a check is the clothing and how it reads — not attractiveness, not body shape as something to be corrected, not confidence or status. That is a design choice about what we ask for and what we show, and we would rather state it as a choice than dress it up as a technical impossibility.

The photograph is part of the result

Input quality is not a footnote. Lighting shifts color, and color relationships are one of the four areas — a warm indoor cast can push a cool gray towards beige. Camera angle and distance change apparent proportion, which is what a fit observation reads. A crop can remove the shoe that completed a line. A garment hidden behind an arm may not be recognized at all, and an unrecognized piece is simply one the analysis did not see.

This does not make the output unreliable. It makes it an analysis of a photograph — a smaller and more accurate claim than an analysis of an outfit.

Three worked examples

  • A dark structured jacket worn open over a pale shirt and pale trousers. The analysis can reasonably say that the jacket forms two long vertical dark panels against a light field, so the eye tends to read that vertical line first and the lighter pieces become the ground it sits against. It cannot conclude that the look is improved by the jacket, that it suits the wearer, or that it stays on.
  • A loose top ending level with the widest point of the hip, over full-length trousers. The analysis can reasonably say that the hem places a horizontal line at that point, dividing the silhouette there, so the leg reads as beginning lower than it does when the same top is tucked. It cannot conclude that either version is more appropriate for this person, or which one they would prefer.
  • A heavy knit layer, with a supplied temperature well above room warmth. The analysis can reasonably say that the upper layer is warmer than the stated temperature supports — and that judgement comes from a calculation over recognized garments, not from the model’s impression. It cannot conclude that the wearer will feel uncomfortable, which depends on the room, the fabric and the person.

The more useful question

“Is this outfit right?” is difficult to answer well, because right depends on a person, a place and an intention that no photograph carries.

“What is creating this effect?” is usually answerable from an image. Why does the top half read heavier. Why do these two colors compete. Why does this feel more formal than the same pieces did last week. Those are questions about observable relationships.

The second question also tends to be the more useful one, because its answer transfers. A verdict applies to one outfit. Noticing that a horizontal hem at the widest point divides a silhouette is something you can use tomorrow, with clothes no model has ever seen.

More questions about how Trendo Fashion works →