I logged 240 generations of the same character over seven months and then measured how far her face actually moved
I logged 240 generations of the same character over seven months and then measured how far her face actually moved

I logged 240 generations of the same character over seven months and then measured how far her face actually moved

This is a writeup of a small measurement project rather than a question. Since the first week of January I have been running a weekly picture story about a fictional coastal town, and the same character narrates it. Around April I got the feeling she did not look like herself any more. Not in a way I could point at. The January images and the April images simply felt like two different women photographed under similar lighting. So I spent the first two weeks of August trying to find out whether that was real or whether I had talked myself into it.

Two things up front. The character is generated. There is no woman behind her, no reference photo of a real face, nothing about this touches a person who exists. And this was not a controlled experiment, for reasons that turn out to be the entire point of the post.

The corpus is 240 images made between the first week of January and the last week of July. By month that is 52 in January, 34 in February, 28 in March, 41 in April, 30 in May, 33 in June and 22 in July. Every single one has a row in a spreadsheet with the date, which model I picked, the full prompt text, the seed where there was one, and a keep or bin flag. Of those 240, exactly 38 are marked as anchors, meaning front on, neutral expression, even lighting, no hat, nothing crossing the face. Those 38 are what I feed back in when I want her to look right, and they are the only images in the whole set that are honestly comparable to each other.

Twelve of the 38 anchors come from January. Two come from July. That distribution was never a decision and staring at it now it is probably half the answer already.

The comparison worked like this. In the first week of August I took my January prompt, the actual text copied out of the spreadsheet, unedited, and generated twenty fresh images with it. Then I scaled every image, the old anchors and the new twenty, so the distance between the pupils was exactly 500 pixels, converted them to grayscale, and aligned them on the midpoint between the eyes. After that everything is just pixels and pixels can be compared.

Before measuring the character I measured myself. I took eight of the January anchors and read every dimension off them three times, on three separate days, without looking at what I had written previously. My own readings moved by as much as four pixels. So I decided in advance that anything under eight pixels is me and not the model, and I have held to that even where it was inconvenient.

Here are the results, everything normalized to that 500 pixel pupil distance. Mouth width went from an average of 236 pixels across the January anchors to 261 in the August set. That is 25 pixels wider, about eleven percent, and it is the largest single thing that moved. Hairline height above the pupil line went from 402 to 431, so the forehead grew by 29 pixels. Jaw width measured at the level of the earlobes went the other way, 604 down to 578, a loss of 26 pixels. Nose width at the nostrils went from 147 to 152, five pixels, which under my own rule counts as nothing at all.

Taken together those three real changes describe one face getting narrower at the bottom and both taller and wider at the top. Which is more or less exactly the sense I had in April and could not put a word to.

The other finding is the one that convinced me something had genuinely changed rather than my taste changing. The character was designed with a small mole below the right eye. It shows up in 214 of the 240 images across the seven months. In the twenty August images made with the untouched January prompt, it shows up in six.

Logging went into a spreadsheet, the character came out of APOB, and the comparisons happened in GIMP, stacked as layers in difference mode. The difference stacks were useful for the forehead and the jaw. The thing that actually made the mouth jump out was far dumber than that. I put January on one layer and August on another and flipped the top one on and off about once a second, and I saw it inside four seconds. Everything else I only found afterwards by measuring.

Now the part that ruins the whole exercise. I cannot tell whether any of this is the tool changing or me changing.

My January prompt was 41 words. The version I was running at the end of July was 96 words. I never rewrote it. I added a clause here, tightened a description there, and the spreadsheet shows about seventeen separate edits across seven months, not one of which felt like a change at the time. Several were me trying to correct something I disliked in one specific image, which means the prompt was quietly absorbing my corrections one at a time. Somewhere in there the eye description went from grey green to green with grey in it, and I have no memory whatsoever of typing that.

The models are worse. I did not hold the model constant either, because I kept trying new ones as they appeared. Of the 240 images, 118 were made on one image model, 74 on a second and 48 on a third. July is almost entirely the third one. So my August comparison set was generated on a model that was not in my account in January, using a prompt I had edited seventeen times, and I am then presenting a measured difference as though it says something clean about drift.

It says the face moved. It does not say why, and no further measurement of the images I already have is going to fix that, because the confound is baked into how they were made in the first place.

What I should have run from day one is a frozen reference. One prompt in a text file, never edited, one fixed model, the same seed where seeds are supported, ten images on the first of every month, filed and ignored. That costs almost nothing and it is the only way that question ever becomes answerable. I started that on the first of August, so I will have something worth saying around March.

The rest of the workflow I am leaving alone. The anchors work well enough that the weekly posts read as consistent to a reader scrolling past, and a reader scrolling past is a far lower bar than a difference layer at 500 pixel eye spacing. I have stopped editing the character prompt entirely though. If something is wrong in an image now I either fix it in the image or I bin the image, and the prompt stays exactly where it is.

The measurement I trust most is the mole, and I found it by accident while looking for something else. I wrote all of this down mostly so that in six months, when I try to remember what her face looked like in January, I am not relying on my memory of it, because I already know how that goes.

submitted by /u/No_Issue_8224
[link] [comments]