How a Model Learns a Face

What LoRA model training is, and how to do it
๐Ÿ‘ค Aaron Kushner ๐Ÿ“… August 2026 ๐Ÿท fine-tuning ยท LoRA ยท evaluation
Hiro standing in a neon street at night, facing the camera
one trained character
The rat thing crouched head-on, about to lunge at the camera
another trained character
Hiro kneeling over the wounded rat thing in a blast crater
both, directed into one frame

As an average consumer with a visual story idea, using an image model for the same character twice and you get two different faces. This annoying fact has a fix, with a name โ€” model fine-tuning โ€” that one would think is lots of complex math, but it hides how simple the underlying idea is. This paper builds that idea from the bottom, one small step at a time, in the manner of the 1 + 1 to Attention ladder: no mathematics beyond arithmetic, every step anchored in something you can already picture. Then it shows what happened when we actually did it โ€” the pictures, the numbers, and the bill. The outcome: any consumer willing to talk to a machine can now do this.

Before any of this: the book, the context, the bible, the prompt

Nothing here begins with a picture. It begins with a novel, and everything downstream โ€” the description, the training set, the trained model, the mistake โ€” is a consequence of how the novel was read.

The book is Snow Crash, Neal Stephenson, 1992.

The context โ€” this work is one piece of a personal ambition: to see Snow Crash adapted for the screen. The primary obstacle has always been that special effects were not up to the challenge of giving the material what it deserved (see Johnny Mnemonic, Tank Girl). Once that problem was emphatically solved, two new ones were revealed. The first is that a deep, expensive special-effects spectacle aimed at a niche demographic cannot be made in today's Hollywood. We argue that AI solves this by democratizing and commoditizing accurate, compelling prototypes that can be made ideal on the cheap before shooting even begins. Goodbye cost and schedule overruns. (See the author's note below if this makes you nervous.)

The second is a matter of audience immersion and suspension of disbelief: the novel's central invention needs explaining. In 1992 the Metaverse cost paragraphs โ€” what an avatar is, why a street exists inside a computer, what it means to go goggled. Prose can afford that. A screenplay cannot, not without turning its characters into narrators, and eating up valuable time that could be spent on story and character enrichment.

That obstacle, too, has largely dissolved, because the audience arrived. Avatars, virtual worlds, goggles, a persistent street you log into โ€” none of it needs a line of explanation to a viewer today. Which points at a specific adaptation stance: play it not as a period piece and not as the future, but as an alternate present โ€” diverging from ours at some branch point the adaptation is free to choose and never obliged to state.

The hypothesis: with the explaining budget freed, the story fits. A feature could carry it; a limited series certainly. What was an adaptation problem becomes a runtime decision. (AI as we now know it is not predicted in the novel โ€” but that is a different adaptation problem.)

The second half of the hypothesis is about who is able to attempt it. Storyboarding used to require the ability to draw. Now the drawing can be generated. If the tools cannot hold one character's face steady across a few hundred frames, they cannot control the other elements needed to let natural language deliver genuinely usable control over shot composition.

Author's Note

You could make the whole film with AI. In a sense that is what you are doing with shot-for-shot prototyping. But put together as a movie, it literally has no soul, and I don't think people are going to buy it. Yes, some parts of the work ecosystem will be replaced โ€” but people can retrain, and less expensive movies means more movies. This pitch is about the unfunded part: spec work, done at a loss, gated by whether you could afford an illustrator or already knew someone. Automating that doesn't displace artists โ€” it displaces the gate. The film still needs the warm bodies once real production starts. My theory, but it does make doomerism a choice rather than a fact-based posture.

The character bible was built by reading the whole novel in a single pass โ€” all 71 chapters inside one context window, no chunking, no retries, no repair passes โ€” and writing the result out in one go: 38 characters, 20 locations, 18 entities. Each character carries a face sketch, a body description, wardrobe by story phase, and a compressed keyword line meant for prompting. One caveat was recorded at the time and still stands: the available text is a re-translation. Spanish artifacts survive in it, and two chapters trail off into prose Stephenson did not write. Physical description came through faithfully; the rhythm of his sentences did not.

Here is what the bible says about Hiro's face, verbatim. It is the primary source for every image in this paper:

He is a man of about thirty with warm cappuccino-colored skin and short, spiky black hair that has begun to recede at the temples, exposing a high forehead. His face is dominated by prominent high cheekbones and dark, narrow Asian eyes inherited from his Korean mother. His nose is straight and his jaw is clean-lined and beardless, giving the whole face a sharp, alert geometry.

That description is clear, specific and internally consistent โ€” a man of about thirty, and everything else follows without contradiction. Hold on to about thirty anyway. Two different models are going to read this same paragraph, and neither of them is going to draw a thirty-year-old.

The prompts are not hand-written. A composer assembles each one from four parts โ€” the character's bible block, the staging of that particular moment, the dialect of the target model, and a style pack โ€” and stamps out twelve scenes across five models and four styles, 240 prompts in all, every one carrying the same identity text so that consistency is at least attempted. Each model declares a character budget, 50 words for one and 25 for another, and blocks are priority-truncated to fit.

This is the complete prompt behind Figure 1, the first test in this paper. The identity portion is marked:

the prompt, before any trainingPhotorealistic cinematic still, NIGHT. A male, 30s, mixed-race African-Korean, cappuccino skin, short spiky black hair, receding hairline, high cheekbones, Asian eyes, lean athletic swordsman build, wearing a black bulletproof motorcycle jumpsuit with armored padding, standing in a neon-lit city street, facing the camera, looking directly into the camera, head and upper body filling the frame, face fully lit and unobstructed. 35mm anamorphic, film grain.

Sixty-four words, twenty-nine of them spent on who he is. That is the state of the art before any training: a description derived from the book by machine, and re-interpreted from scratch on every single render. What that produces is the next section.

The problem, in three pictures

That prompt, run three times, changing nothing at all. Not the scene, not the wardrobe, not a single word โ€” and this time he faces the camera, so there is nowhere for the differences to hide.

Man in black armored jacket facing the camera in a neon street
run 1
A different man, same prompt, facing the camera
run 2
A third man, same prompt, facing the camera
run 3
Figure 1 โ€” the same words, three runs, three men. This is better than it has any right to be: the age, the build, the spiky black hair and the wardrobe hold every time, because a long and specific description does real work. But look at the faces โ€” the jaw, the nose, the set of the eyes belong to a different man in each run. What a description buys you is a type, reliably. It does not buy you a person.

And a type is not enough, for two reasons. The small one is that readers recognize characters by face, not by build. The large one is that the type itself is only this model's reading of those words โ€” change the art style and it shifts, change the model and it shifts further. We will watch exactly that happen later in this paper, at some expense.

Nor can you fix it by writing a better paragraph. Every extra adjective narrows the region a little and costs you a little control over the scene, and you never reach a single point. The description is the wrong kind of instrument. What we want is not a better description of the man โ€” it is the man himself, stored in the machine.

Part I โ€” The ladder

Eight steps from something you already know to what a LoRA actually is. Each one is small.

Rung 1 A picture is a pile of numbers, and a face is a direction in the pile

Descartes' move, in 1637, was to say that a point is just two numbers โ€” (x, y) โ€” and with that, every shape became something you could calculate. A photograph is the same move at scale: three numbers per pixel, a million pixels, one long list. And a long list of numbers, as the ladder's Vectors rung has it, is an arrow in space.

[in the figure below, it's not clear unless you know vectors that those all can be combine into one vector... which allows them to be easily compared]

Three panels: a photograph, a column of RGB numbers, and an arrow among many axes
Figure 2 โ€” the same thing, said three ways. A picture is a grid of pixels; each pixel is three numbers; so the picture is one very long list โ€” 1024 ร— 768 ร— 3 is 2,359,296 numbers. And a list that long is an arrow in a space with that many directions. The swatches and the numbers here are read from the marked patch of his cheek, not invented.

That is the whole reason any of this is possible: "how similar are these two pictures" becomes "how nearly do these two arrows point the same way," and that is arithmetic โ€” multiply the two lists slot by slot and add up the results.

One honest caveat, because it sets up everything that follows. Measured on raw pixels, that similarity is a poor judge of identity: move the light or turn the head and the arrow swings hugely while the person stays the same. Getting to a space where direction really does mean identity is not a given โ€” it is precisely the thing the rest of the machine is built to do.

Rung 2 The model is a wall of dials

Inside, the machine does one thing over and over: multiply each number by how much it should count, and add up the results. A recipe โ€” two parts flour, one part water โ€” is that operation, and so is a grade book, and so is every layer of every model ever trained. The interesting part was never the arithmetic. It is the weights: how much each input counts.

Those weights are the dials. A modern image model has roughly twelve billion of them. Every one is a number somebody's computer chose, and the settings are the model โ€” there is nothing else in there.

Three panels: a recipe as a weighted sum, the same sum as three dials, and a wall of twelve billion tiny dials
Figure 3 โ€” the whole machine is one idea, repeated. A recipe is already a weighted sum: multiply each ingredient by how much it counts, add it up. Draw the amounts as dials and the sum has knobs. A model is that picture at scale โ€” twelve billion dials, and the settings are the model; there is nothing else inside. Training, whenever it appears in this paper, means turning/tuning them... until the output tastes (looks) just right.

Rung 3 Training is a blindfolded hiker

How do the dials get set? Show the machine a picture, let it guess, measure how wrong it was, then ask of every single dial: if I nudge this one a hair, does the wrongness go up or down? Nudge them all a little in the direction of "down." Look again. Repeat a few million times.

That is gradient descent, and the image to hold is a hiker in fog, feeling the tilt of the ground underfoot and stepping downhill without ever seeing the valley. No cleverness, no insight โ€” just tilt, step, repeat, at billions of dials per step. Every trained model you have ever used is the residue of that loop.

This is not a machine-learning trick โ€” it is how hard problems get computed at all

The same loop runs under work that has nothing to do with models. The CS9 aerodynamics results are a flow solver guessing a field, measuring how badly it violates the equations of motion, correcting, and repeating until the corrections stop mattering. The Titanic rivet analysis is finite elements doing the same thing to a stress field โ€” and finite elements are literally descent, because the method works by minimizing an energy and the answer is the bottom of the valley. Schrรถdinger's equation has no exact solution for any atom past hydrogen, so quantum chemistry proposes a wavefunction, computes its energy, and descends. In every case the equation cannot be solved perfectly, and in every case you get an answer good enough to build on. Training a model is one more member of this family, not a departure from it.

There is a second great family, and it is worth knowing the difference because they are cousins rather than twins. Descent needs a slope to follow. Monte Carlo [make this a hyperlink that adds: coming soon! and all the links track number of clicks] needs nothing of the kind: propose a random move, keep it if it improves matters, and โ€” the crucial part โ€” keep it sometimes even when it doesn't, with a probability set by Boltzmann's eโˆ’ฮ”E/T. That rule is the Metropolis algorithm, worked out at Los Alamos in the 1940s and 50s to compute neutron diffusion, where the physics was a branching mess of chance no equation could be marched through directly. Descent walks downhill; Metropolis wanders, and is allowed to climb, which is what lets it escape a valley that is merely local.

The two families meet in a place that is physically visualizable, grain boundaries rearranging in cooling steel: Run Metropolis while slowly lowering that T and the wandering gradually settles into descending โ€” simulated annealing, named for the way a cooled metal finds its lowest-stress arrangement. And the acceptance rule it cools is the same Boltzmann factor that appears on the 1 + 1 ladder as softmax: the formula deciding which word a chat model says next was written in 1868 to describe how molecules share out energy. Two centuries of people trying to compute the uncomputable, arriving repeatedly at the same handful of moves.

Three panels: a hot energy landscape with jumps in every direction, a cooling one where the wandering settles, and a cold one with the ball seated in the deepest valley and the steel grains aligned
Figure 4 โ€” annealing, watched happen. The same landscape three times, only the temperature changes. Hot: the ball jumps anywhere, uphill included, and the steel's grains point every direction. Cooling: uphill jumps get rarer and the grains begin to agree. Cold: the wandering has become descent, the ball sits in the deepest valley โ€” not the shallow one it started in โ€” and the metal is annealed. The acceptance rule doing the cooling is Boltzmann's eโˆ’ฮ”E/T, written in 1868; hold T fixed instead and the same formula is the softmax that picks a chat model's next word.

Rung 4 Why you cannot simply retrain

So to teach the model a new face, just run the loop again on pictures of that face. This works, and it is a disaster. You have now produced a second twelve-billion-dial model โ€” many gigabytes โ€” that knows one character. A cast of thirty characters is thirty copies. Worse, the loop happily wanders dials that had nothing to do with faces, and the model quietly gets worse at everything else while it learns your one thing.

It is repainting the entire house because you wanted a different picture in the hall.

Rung 5 The move: a big change written in a few strokes

Here is the idea the whole technique rests on, and it is a genuinely beautiful one.

Picture a hundred-by-hundred multiplication table. Ten thousand numbers on the page. But it is not ten thousand independent facts โ€” you can rebuild every one of them from the numbers 1 to 100 down the side, 1 to 100 across the top, and one rule. Two hundred numbers and a rule regenerate ten thousand. The grid looked big; its actual content was small.

Enormous grids that are secretly describable by a handful of lists are said to be low rank, and it turns out that the change you want to make to a model when teaching it one new thing is almost always of this kind. The adjustment looks like a wall of twelve billion tweaks. Its real content is a few patterns.

So: freeze the model. Do not touch a single one of the original dials. Beside each big grid of them, hang a small pair of lists, and train only those. When the model runs, it adds the little pair's contribution to the frozen grid's output and carries on. This is LoRA โ€” a Low-Rank Adaptation. The word "rank" is just how many of those pattern-pairs you allow. We used sixteen.

What it buys

Our finished adapter is 131 megabytes against a base model many times larger; it trained in 8.6 minutes for $2. Because the original was never touched, the adapter snaps on and off like a lens โ€” and tomorrow's character gets its own, with no copy of the model in between. This is the same manoeuvre used to specialize large language models. When you read that a model was "post-trained" or "adapted," this is very often the literal mechanism.

Three panels: a stack of full model copies one per character, a multiplication table rebuilt from its two edges, and a frozen model with a small trained pair beside it
Figure 5 โ€” the whole of Rungs 4 and 5 in one line. Left, the disaster: retraining means a full twelve-billion-dial copy per character โ€” thirty characters, thirty copies, and the loop degrading everything else as it goes. Middle, the insight: a 100 ร— 100 times table holds ten thousand numbers, but two hundred and a rule regenerate all of them โ€” the change you want to make to a model is low-rank in exactly this way. Right, the move: freeze every original dial, train only the small pair, and the whole adaptation ships in 131 MB that snaps on and off like a lens.

Rung 6 The trigger word: teaching the machine a name

The adapter needs a handle. You invent a string of letters the model has never seen โ€” ours was hiro_sc_v1 โ€” and put it in the caption of every training picture. Because the machine has no prior associations for that nonsense word, the only meaning it can acquire is whatever is common to the pictures it labels.

Which cuts both ways, and this is the trap. Anything constant across your pictures gets welded to the name. If the character wears the same jacket in all twenty-four, the name now means "this man in that jacket," and you will never get him out of it. The defence is to name the removable things explicitly in the caption โ€” the jacket, the sword, the motorcycle โ€” so the model files them under those words instead of under him.

The name buys something else that is easy to miss and turns out to matter as much as the identity: room in the prompt. A prompt is not free-form โ€” every model has a budget, and attention thins as you spend it.

What the description was costing

The scene prompts written for this project average 85 words, of which 37 โ€” nearly half โ€” say who he is, leaving the remainder for camera, action, setting and art style combined. The wider storyboard pack shows the same shape: subject blocks average 53 words inside 141-word prompts. And the budgets are hard limits, not suggestions โ€” one model declares a 50-word character allowance with the note "front-loads attention; outfit must appear in the first ~20 words"; another allows 25, at which point the composer is already truncating the description to make it fit. We were losing the man in order to afford the scene.

The trigger word is two words. The same prompt with the adapter runs 50 words instead of 85 โ€” thirty-five words handed straight back to scene direction, and identity's share of the prompt falling from 44% to 4%. That is the difference between specifying several key details of a shot at once and hoping the model happens to honour them.

Rung 7 The bootstrap: photographs of a man who does not exist

Ordinarily you train on photographs. We had a paragraph in a novel. So the reference set was manufactured: sixty candidate images generated from the character bible, then curated down by eye to twenty-four.

Contact sheet of twelve generated head-and-shoulders reference portraits
 
Contact sheet of twelve generated action and expression references
Figure 6 โ€” the manufactured reference set. Two of the five review sheets. Lighting changes, clothing changes, pose changes, style changes; the person does not. That sentence is the entire art of building one of these sets โ€” whatever you hold constant is what the name will come to mean. Final split: six head, six half-body, six full-body, four action, two expression.

The curation is not housekeeping; it is the training signal, and rejecting is more important than accepting. A model trained on images that disagree about a face learns the average of them โ€” and the average of six nearly-right men is a seventh man who resembles nobody. Near-misses do more damage than obvious rejects, because they survive review.

Rung 8 The strength dial

The adapter arrives with a volume knob: how much of its contribution to add. At full strength it is a very heavy accent โ€” you know exactly who is speaking, but the words start to go. Turned down, the accent stays audible and the sentence comes back. Remember this dial; it decides the entire result.

Part II โ€” The experiment

Two commitments made this a measurement rather than a demonstration.

The thresholds were written down before the run. What would count as success โ€” how many points of improvement, on which axes โ€” was fixed the day before, in the design document. This sounds like bureaucracy and is in fact the whole difference between an experiment and a sales pitch: it stops you from grading your own homework once you have seen the answers.

Everything was held constant except identity. Twelve renders: six scenes, each drawn twice from the same random seed, in the same style, with the same scene description. In one version the character is described in words; in the other he is summoned by his trigger word and the adapter. Same everything else. Any difference you see is the adapter.

Each render was then scored 1โ€“5 on five axes โ€” does the face match, does the build match, are his canonical traits right, did the model obey the scene, and did the requested art style survive.

What it cost

StepWhereCost
60 candidate referencesimage API~$3.00
Training, 1000 stepshosted GPU$2.00
12 test rendershosted GPU~$0.35
6 correction rendershosted GPU~$0.20
6 pencil-style rendershosted GPU~$0.20
Totalone afternoon~$5.75

Worth stating plainly, because the number is the point: training a custom model is no longer an industrial act. The expensive parts of this project were choosing twenty-four pictures by eye and writing down the thresholds beforehand. The machine learning cost two dollars.

What generalizes

The dataset's constant is the lesson. Whatever every training example shares โ€” the person, the lighting, the jacket, the photographic look โ€” is what the model actually learns. It cannot distinguish the thing you meant from the thing that merely happened to be there every time. Vary everything you do not want taught.

Rejecting is the craft. Twenty-four agreeing images beat forty that nearly agree, because disagreement averages into mush.

Approve the cast before you train it. Curation picks among candidates; it cannot conjure one that was never generated. So the decisive review is the one before curation โ€” someone who knows what the character is supposed to look like, looking at the raw sheet and saying yes or no. Skip that and you can run a flawless training pass and end up with a faithful, consistent, high-scoring portrait of the wrong man. Ours came out fifteen years too old and nothing downstream could have noticed, because every number we measured was about consistency, and it was perfectly consistent.

Two models will read the same clear words differently. This needs no ambiguity in the writing to happen โ€” ours said a man of about thirty and got a twenty-something from one model and a forty-something from the other. So when the pictures you train on come from one model and the pictures you make come from another, the adapter is a courier carrying the first model's interpretation into the second one's world, permanently and silently. Decide which model's reading you actually want before you train on it.

A trained identity is also a prompt-budget saving. Thirty-seven words of description became two, and the whole prompt shrank from 85 words to 50. Everything a description spends on who is taken from what is happening, and prompt budgets are small and front-loaded โ€” so the adapter buys control of the shot as surely as it buys the face.

Strength is a setting, not a property. A capability can be present and still fail its evaluation because the knob is wrong. Test the dial before you condemn the training โ€” it costs cents and it moved three of four gates here.

Write the thresholds down first. This one costs nothing and is the only reason the sentence "it improved face identity by 1.17 points" means anything at all.

What this is becoming

The ladder above describes a process. That process is now being built into an application, in a second workbench running on OpenAI's Codex, and everything below this line was made with it.

It ships in two builds because it has no model of its own. It is a director, not an engine: it holds the character bible, runs the casting gate, assembles the reference set, submits the training, and keeps every selection bound to the file it came from โ€” but the actual reading, writing and rendering is done by the coding agent you already pay for. So there is a build that drives Claude Code and a build that drives Codex, and you install whichever one your subscription covers.

Requires an active subscription that grants Claude Code or Codex privileges. The app supplies the process and the memory; your agent does the work, on your account, on your machine.

Placeholder — the interface

A screen of the application goes here once the interface settles.

What it does, in the order it does it. You give it the book. It reads the whole thing and derives a character bible — who exists, what they look like, where the text actually says so. For any character you want to be able to draw twice, it generates a casting sheet and stops: you pick the face, because the rung that matters most is the one no metric can stand in for. From your pick it builds the reference set, captions it so the trigger word carries identity and nothing else, submits the training run, and reports back what the adapter cost and whether it holds. Then you direct scenes in plain language, with the character's name doing the work that thirty-seven words of description used to do, and it keeps every accepted frame bound to its source and its hash so a selected picture can never be quietly replaced by a regenerated one.

That last part is not a detail. Halfway through the sequence below, an automated compositor was asked to assemble four accepted panels and silently substituted an unselected fourth. The fix was to stop generating at assembly time altogether — selected frames are now cut and copied deterministically and checked against their hashes. It is the same lesson as the casting gate, one layer down: the machine is good at making pictures and bad at knowing which one you chose.

Two characters, and a sequence

Raven was the second character through the process, and the first one run with the casting gate in place from the beginning. The portrait on the left is the frame that was approved before any training happened — the one the whole reference set was built from. The frame on the right is the trained adapter doing its job: a different night, a different lens, a scene the portrait never showed.

Head-on portrait of a heavy-set Aleut man in his forties with waist-length black hair, a thin Fu Manchu moustache, and POOR IMPULSE CONTROL tattooed across his forehead
the approved face
The same man at night on a black motorcycle at the edge of an overpass concert, crowds behind him
the face, directed
Figure 7 — the casting gate, then the scene. The left frame is a human decision: four refinements were generated and this one was chosen, before a single training dollar was spent. The right frame was rendered later from the trained adapter at strength 0.7, with the motorcycle, the crowd and the night light all specified in the prompt and the face specified only by name. Neither the forehead lettering nor the cold half-smile is exact yet — both are recorded as open, which is what an honest casting record looks like. The model also invented a brand across the fuel tank; that is painted out here, and nothing else in the frame is retouched.

The last piece is what all of it is for. This is a single scene from the novel — a machine that thinks it is a dog crosses a street to catch a grenade — cut as eight beats. Two characters carry it, both of them trained: the man, and the machine. It opens on the man, because the stakes have to land before the rescue can mean anything. Two of the beats are the same seconds from inside the animal's head; the rest is the street. And the axis of the camera crosses it only once the machine has passed him, because a storyboard that breaks that rule stops reading as one place.

A man in a black coverall freezing on a wet neon street as a grenade bounces toward him
1 he sees it land
Ground-level macro of a grenade wobbling to a stop on wet asphalt
2 the last half-turn
A dark four-legged machine bursting from a lighted nook in a wall, sparks and condensation trailing
3 the nook opens
A golden retriever running flat out on an open highway, cars streaking past
4 the dog-mind: eighty miles an hour
Extreme close-up of the running dog's muzzle and nostrils, whiskers pinned back by wind
5 breath and wind, nothing else
The machine charging down the wet street from behind, the man small in the distance
6 the intercept
The man crouching and turning away, one forearm across his face, an explosion behind him
7 the axis crosses
The man kneeling beside the disabled machine at the crater, its cooling vanes glowing white-hot
8 after
Figure 8 — eight beats, two trained characters, one street. Beats 3, 4, 5 and 6 were selected from generated boards; 1, 2 and 7 were shot to fill the bridge once the shape of the sequence was clear; 8 is the established hero frame, edited only to drive the cooling vanes from orange to white-hot. Every frame here is the frame that was chosen — each one cut from its source and verified against a hash rather than regenerated at assembly.

Method and thresholds from the project's character-consistency design document (August 2026). Base model FLUX.1 [dev]; rank-16 LoRA, 1000 steps, learning rate 5e-4, hosted trainer. All twenty-four reference images and all twenty-four renders (twelve paired, six at reduced strength, six in pencil) were generated for this study; the character is from Neal Stephenson's Snow Crash and this work is an internal methods prototype, not a licensed adaptation. Full evidence record, per-axis scores, and the reproducible pipeline live with the project artifacts.