Tech Teardown · Vol.05

We never trained it on these clothes,
yet it styled them like the brand would

Train a model on a brand's lookbook, and it picks up even the way that brand dresses its clothes.

The last two parts were about fit. Describing it in words doesn't work, and splitting the vocabulary finer doesn't work either. So we tried training a model directly on a brand's clothes.

Fit wasn't why we ran this experiment. We had already seen fit carry over with an earlier brand. What we wanted to check this time was how the model dresses those clothes.

The pitch this experiment set out to test

"We've already trained on your brand."

"From now on, whenever a new piece drops, just send us a single photo of the garment. We'll style it in your brand's mood and deliver shots ready for your product detail page."

Brands shoot a lookbook every time new pieces come out. They book models, reserve a studio, and a stylist matches the shoes and accessories. Every lookbook a brand has shot so far already contains how that brand dresses its clothes.

Here's what happened. Put in a single pair of pants, and a top gets created and the whole outfit comes out styled. We never put in a top.

01 · The Setup

What we fed it

Training works by feeding in questions paired with answers. The question is one product shot of a pair of pants; the answer is one lookbook photo of someone wearing those pants.

Four of the training pairs, with column labels reading 'Input · flat' and 'Answer · worn shot'. On the left is a flat product shot of the pants; on the right, a worn shot of the same pants photographed together with a top, shoes, socks and a bag.
Four of the pairs we trained on. In each pair, the left column (labeled "Input · flat") is the question and the right column (labeled "Answer · worn shot") is the answer. We fed in 114 pairs like these.

We built 114 of these pairs and fed them in. That works out to 114 garments.

And we deliberately held back eleven garments. Those eleven never went into training. Once training was done, we fed in only their product shots and set the results side by side with the lookbook shots the brand had actually taken.

114Garments used in training
11Garments held back, used to check results
0Garments in both sets

The reason for holding them back is simple. If you check results on garments that were in training, you can't tell whether the model memorized those garments or learned the brand.

That said, we don't call the lookbook shot next to it the correct answer. The brand shot several photos with each pair of pants, and some put a different top with the same pants. Our result isn't wrong just because it differs from that one photo.

02 · What Came Back

It copies the brand's look and feel
Above the divider, twelve of the brand lookbook photos used in training. Below it, the results for all eleven pants that were never in training, each shown as a small product shot next to a larger generated full outfit.
Above the divider are twelve of the brand lookbook photos we used for training. This is how the brand dresses its clothes. Below the divider are all the results from feeding in eleven pants that never went into training. In each cell, the small image is the product shot we put in and the large one is the result.

The input is one pair of pants. We never put in a top, yet all eleven came out with a top. Not just any top, either: what comes out is the one outfit you'd expect this brand to put together with those pants.

A few of them are hard to put down to chance.

Two rows, labeled 'Red track pants: angular panels on the outer thigh' and 'Navy track pants: the same thing, once more'. Columns show the product shot we put in, our result, and the brand's actual photo. In both, the angular panel on the outer thigh moves up to the shoulders and upper sleeves of the generated jacket.
Top row: red track pants. Bottom row: navy track pants. Columns, left to right: the product shot we put in, our result, the brand's actual photo. The angular panels set into the outer thigh go straight onto the jacket's shoulders and upper sleeves. The photo the brand actually took has the same panels in the same place. It didn't just match the color; it got where on the body the pattern goes right. This happened once each, separately, for red and for navy.
Labeled 'Gray pinstripe pants: we only put in the pants'. Columns show the product shot we put in, our result, and the brand's actual photo. Feeding in one pair of gray pinstripe pants produces a blazer in the same fabric, worn open over a white inner layer.
Gray pinstripe pants; we only put in the pants. Columns, left to right: the product shot we put in, our result, the brand's actual photo. Here it's not just the clothes that match. It invented a blazer in the same fabric, and it's worn open, the hem of a white inner layer hangs below the jacket, and both hands are in the pants pockets. That's exactly the photo the brand actually took.
Titled 'The same pants, four runs with only the conditions changed: all four came out with something on the head'. Panels show the product shot we put in, then results labeled small · default, small · strong, large · default and large · strong, then the brand's actual photo, labeled 'no hat'. All four results wear a beanie or hood.
We ran the same pair of pants four times, changing only the conditions: image size (small or large) and strength (default or strong), as labeled in the panels. All four came out wearing a beanie or a hood. It's not in the input, not in the caption, and not in the photo the brand took with this garment (labeled "no hat"). The photos we trained on include eleven with a knit beanie.

This isn't the right answer for this garment. It's just something the brand does often. But the fact that it keeps showing up even when the conditions change means it isn't a one-off accident.

That said, it uses this habit more often than the brand does. In the photos we trained on, 26 of 114 match top and bottom in the same fabric, while seven of the eleven results do.

03 · Where It Came From

It's all right there in the lookbook photos we trained on

Look again at the training pairs from earlier and the answer is there. What we put in was one pair of pants, but the photo we used as the answer also shows a jacket, shoes, socks and a bag. And yet the caption attached to that photo describes only the pants.

Here, the caption is a one-line description attached to each training photo. It states in a sentence what the garment is, and goes in together with the photo.

Titled 'One pair we trained on'. Left, the training photo; right, 'The caption attached to this photo', which describes only the pants, followed by 'In the photo but not in the caption': camel blazer, printed shirt, brown dress shoes, black bag, sunglasses, standing pose and hand position, gray studio background.
On the left is a photo we trained on; on the right is the caption attached to it. The caption describes only the pants. Not a word about the jacket, the shoes, the bag or the background. The list under the caption in the image, "In the photo but not in the caption", reads: camel blazer, printed shirt, brown dress shoes, black bag, sunglasses, standing pose and hand position, gray studio background.

It learns what the caption doesn't say

Training pairs a photo with a sentence describing that photo. Whatever is in the photo but isn't explained by the sentence, the model folds into the brand's characteristics.

In this training set, the tails of all 114 captions were identical, down to the last character. Only the garment description differed; everything else was the same sentence. So the styling, the background and even the pose got learned along with it.

04 · Alpha

The clothes looked like they were floating

The brand feel was there. But the clothes looked like they were floating. The fabric came out flat, as if it had been painted on: convincing from a distance, but up close it didn't look like clothing.

Our first thought was this: what if we just apply what it learned more strongly?

What "applying it more strongly" means

In Part 1, we compared this approach to sticking a thin memo on top of a thick manual. The original model stays as it is, and a thin memo made from the brand's lookbook is layered on top.

We introduced a few values you can adjust there, and one of them decides how strongly that memo is applied. The default is 1.0. Raise it and the output moves further from the original and gets pulled more toward the memo. Set it to 0 and the memo isn't applied at all, so you get the original as is.

We raised it from 1.0 to 1.3.

It didn't work. Volume didn't grow consistently, and some garments actually shrank. The direction wasn't consistent either. Strength isn't the dial to turn. What training failed to capture doesn't appear no matter how strongly you apply it.

05 · Image Size

Raising the resolution fixed it

The floating look came down to texture. We had been shrinking the photos before feeding them into training, and fine information like the grain of the fabric was already lost at that stage.

We switched to feeding the photos in large and trained again. The missing detail came back. The pinstripes sharpened, the crepe texture came alive, and the side-seam tape started to show.

Titled 'Close-ups of the same garments'. Four garments zoomed in and compared across four rows labeled original product shot, trained small, trained large, and brand's actual photo.
Close-ups of the same garments. Rows, top to bottom: original product shot, trained small, trained large, brand's actual photo. The second row is the result of training small, and the fabric is flat. The third row was trained large, and the creases and grain come alive. The bottom row is the photo the brand actually took.

It wasn't the strength value that changed the result; it was how large we showed it the images.

06 · Too Much

Overtrain it and the output hardens

Training means showing the model photos and tweaking it a little, over and over. We count one of those repetitions as one step. This training ran for 2,052 steps, and along the way we saved the intermediate state twenty-one times.

More steps don't make it better. Past a certain point, it starts producing the same thing every time. You see it most clearly in the shoes.

Below-the-calf crops of the same five garments compared at two points: the left column is labeled 'after 1,375 steps: different shoes for each garment' and the right 'after 2,280 steps: the same shoes on all'. On the left the shoes vary; on the right they converge on one shape.
Crops of the feet for the same five garments. The left column is labeled "after 1,375 steps: different shoes for each garment" and the right "after 2,280 steps: the same shoes on all". The left splits into white sneakers, mauve loafers, black pointed shoes and cream slip-ons, while the right, after roughly 900 more steps, converges on a single rounded mule.

On the left, each garment got different shoes: white sneakers, loafers, pointed shoes and slip-ons all mixed in. On the right, every garment gets the same shape, whatever the clothes. In those extra steps, the shoes narrowed down to one kind.

The same thing happens to the background. The caption only called the background "plain". It never said what color. The model fills that blank with the gray it saw in training. Not the white background we'd been using, but the gray of the brand's studio.

Titled 'The same outfits, with training run longer: look only at the background'. Three outfits compared side by side at 1,600 steps (left) and 2,052 steps (right). On the right, the background has sunk to gray.
The same outfits at 1,600 steps (left) and 2,052 steps (right), as labeled under each image. The clothes are nearly identical. The background is different. On the right, it isn't white but gray.
·

Background brightness (255 = pure white)

Actual worn-shot background 255.0 Midpoint checkpoints 254.8 to 254.9 Final checkpoint 248.5 ← 7 of the 11 garments at 242 to 246

Both times, the secondary elements broke down before the fit did. Fit keeps improving right to the end, so if you only watch fit, you miss this. The point where fit is best is not the point where the output is best.

So you shouldn't just use whatever state training ends in. You need to keep several intermediate states, line them up side by side, and pick one. In every training run we've done so far, the final state has never once been the best. And we added background brightness to the list of things we check when reviewing results. The eye misses it, but the numbers catch it.

07 · What We Take

To sum up
QuestionAnswer
Does training on a brand's lookbook copy its look and feel?Yes. Color, color blocking, even the styling
Does it copy things we never taught it?Yes. It creates and adds tops and shoes we never put in
Does it work on new pieces that weren't in training?Yes. All eleven test garments were products outside the training set
Does raising the strength make it more pronounced?No. We dropped that hypothesis
Does training longer make it better?No. Past a certain point the output hardens

The lookbooks a brand has shot so far contain the way that brand dresses its clothes. Training lets us draw that out and put it to use.

We also learned what to adjust: image size and where to stop training.

The next question is who writes those captions. The captions used in this training weren't written by people either. Gemini looked at the photos and wrote the sentences.

And we do the same thing in our service. When a seller uploads a garment photo, Gemini writes the product description. Then the seller edits it. Over fifteen thousand of those edits have piled up.

Everything Gemini gets wrong, and everything people fix, is in there. That's what the final part is about.