As an average consumer with a visual story idea, using an image model for the same character twice and you get two different faces. This annoying fact has a fix, with a name โ model fine-tuning โ that one would think is lots of complex math, but it hides how simple the underlying idea is. This paper builds that idea from the bottom, one small step at a time, in the manner of the 1 + 1 to Attention ladder: no mathematics beyond arithmetic, every step anchored in something you can already picture. Then it shows what happened when we actually did it โ the pictures, the numbers, and the bill. The outcome: any consumer willing to talk to a machine can now do this.
Nothing here begins with a picture. It begins with a novel, and everything downstream โ the description, the training set, the trained model, the mistake โ is a consequence of how the novel was read.
The book is Snow Crash, Neal Stephenson, 1992.
The context โ this work is one piece of a personal ambition: to see Snow Crash adapted for the screen. The primary obstacle has always been that special effects were not up to the challenge of giving the material what it deserved (see Johnny Mnemonic, Tank Girl). Once that problem was emphatically solved, two new ones were revealed. The first is that a deep, expensive special-effects spectacle aimed at a niche demographic cannot be made in today's Hollywood. We argue that AI solves this by democratizing and commoditizing accurate, compelling prototypes that can be made ideal on the cheap before shooting even begins. Goodbye cost and schedule overruns. (See the author's note below if this makes you nervous.)
The second is a matter of audience immersion and suspension of disbelief: the novel's central invention needs explaining. In 1992 the Metaverse cost paragraphs โ what an avatar is, why a street exists inside a computer, what it means to go goggled. Prose can afford that. A screenplay cannot, not without turning its characters into narrators, and eating up valuable time that could be spent on story and character enrichment.
That obstacle, too, has largely dissolved, because the audience arrived. Avatars, virtual worlds, goggles, a persistent street you log into โ none of it needs a line of explanation to a viewer today. Which points at a specific adaptation stance: play it not as a period piece and not as the future, but as an alternate present โ diverging from ours at some branch point the adaptation is free to choose and never obliged to state.
The hypothesis: with the explaining budget freed, the story fits. A feature could carry it; a limited series certainly. What was an adaptation problem becomes a runtime decision. (AI as we now know it is not predicted in the novel โ but that is a different adaptation problem.)
The second half of the hypothesis is about who is able to attempt it. Storyboarding used to require the ability to draw. Now the drawing can be generated. If the tools cannot hold one character's face steady across a few hundred frames, they cannot control the other elements needed to let natural language deliver genuinely usable control over shot composition.
The character bible was built by reading the whole novel in a single pass โ all 71 chapters inside one context window, no chunking, no retries, no repair passes โ and writing the result out in one go: 38 characters, 20 locations, 18 entities. Each character carries a face sketch, a body description, wardrobe by story phase, and a compressed keyword line meant for prompting. One caveat was recorded at the time and still stands: the available text is a re-translation. Spanish artifacts survive in it, and two chapters trail off into prose Stephenson did not write. Physical description came through faithfully; the rhythm of his sentences did not.
Here is what the bible says about Hiro's face, verbatim. It is the primary source for every image in this paper:
He is a man of about thirty with warm cappuccino-colored skin and short, spiky black hair that has begun to recede at the temples, exposing a high forehead. His face is dominated by prominent high cheekbones and dark, narrow Asian eyes inherited from his Korean mother. His nose is straight and his jaw is clean-lined and beardless, giving the whole face a sharp, alert geometry.
That description is clear, specific and internally consistent โ a man of about thirty, and everything else follows without contradiction. Hold on to about thirty anyway. Two different models are going to read this same paragraph, and neither of them is going to draw a thirty-year-old.
The prompts are not hand-written. A composer assembles each one from four parts โ the character's bible block, the staging of that particular moment, the dialect of the target model, and a style pack โ and stamps out twelve scenes across five models and four styles, 240 prompts in all, every one carrying the same identity text so that consistency is at least attempted. Each model declares a character budget, 50 words for one and 25 for another, and blocks are priority-truncated to fit.
This is the complete prompt behind Figure 1, the first test in this paper. The identity portion is marked:
Sixty-four words, twenty-nine of them spent on who he is. That is the state of the art before any training: a description derived from the book by machine, and re-interpreted from scratch on every single render. What that produces is the next section.
That prompt, run three times, changing nothing at all. Not the scene, not the wardrobe, not a single word โ and this time he faces the camera, so there is nowhere for the differences to hide.
And a type is not enough, for two reasons. The small one is that readers recognize characters by face, not by build. The large one is that the type itself is only this model's reading of those words โ change the art style and it shifts, change the model and it shifts further. We will watch exactly that happen later in this paper, at some expense.
Nor can you fix it by writing a better paragraph. Every extra adjective narrows the region a little and costs you a little control over the scene, and you never reach a single point. The description is the wrong kind of instrument. What we want is not a better description of the man โ it is the man himself, stored in the machine.
Eight steps from something you already know to what a LoRA actually is. Each one is small.
Descartes' move, in 1637, was to say that a point is just two numbers โ (x, y) โ and with that, every shape became something you could calculate. A photograph is the same move at scale: three numbers per pixel, a million pixels, one long list. And a long list of numbers, as the ladder's Vectors rung has it, is an arrow in space.
[in the figure below, it's not clear unless you know vectors that those all can be combine into one vector... which allows them to be easily compared]
That is the whole reason any of this is possible: "how similar are these two pictures" becomes "how nearly do these two arrows point the same way," and that is arithmetic โ multiply the two lists slot by slot and add up the results.
One honest caveat, because it sets up everything that follows. Measured on raw pixels, that similarity is a poor judge of identity: move the light or turn the head and the arrow swings hugely while the person stays the same. Getting to a space where direction really does mean identity is not a given โ it is precisely the thing the rest of the machine is built to do.
Inside, the machine does one thing over and over: multiply each number by how much it should count, and add up the results. A recipe โ two parts flour, one part water โ is that operation, and so is a grade book, and so is every layer of every model ever trained. The interesting part was never the arithmetic. It is the weights: how much each input counts.
Those weights are the dials. A modern image model has roughly twelve billion of them. Every one is a number somebody's computer chose, and the settings are the model โ there is nothing else in there.
How do the dials get set? Show the machine a picture, let it guess, measure how wrong it was, then ask of every single dial: if I nudge this one a hair, does the wrongness go up or down? Nudge them all a little in the direction of "down." Look again. Repeat a few million times.
That is gradient descent, and the image to hold is a hiker in fog, feeling the tilt of the ground underfoot and stepping downhill without ever seeing the valley. No cleverness, no insight โ just tilt, step, repeat, at billions of dials per step. Every trained model you have ever used is the residue of that loop.
The same loop runs under work that has nothing to do with models. The CS9 aerodynamics results are a flow solver guessing a field, measuring how badly it violates the equations of motion, correcting, and repeating until the corrections stop mattering. The Titanic rivet analysis is finite elements doing the same thing to a stress field โ and finite elements are literally descent, because the method works by minimizing an energy and the answer is the bottom of the valley. Schrรถdinger's equation has no exact solution for any atom past hydrogen, so quantum chemistry proposes a wavefunction, computes its energy, and descends. In every case the equation cannot be solved perfectly, and in every case you get an answer good enough to build on. Training a model is one more member of this family, not a departure from it.
There is a second great family, and it is worth knowing the difference because they are cousins rather than twins. Descent needs a slope to follow. Monte Carlo [make this a hyperlink that adds: coming soon! and all the links track number of clicks] needs nothing of the kind: propose a random move, keep it if it improves matters, and โ the crucial part โ keep it sometimes even when it doesn't, with a probability set by Boltzmann's eโฮE/T. That rule is the Metropolis algorithm, worked out at Los Alamos in the 1940s and 50s to compute neutron diffusion, where the physics was a branching mess of chance no equation could be marched through directly. Descent walks downhill; Metropolis wanders, and is allowed to climb, which is what lets it escape a valley that is merely local.
The two families meet in a place that is physically visualizable, grain boundaries rearranging in cooling steel: Run Metropolis while slowly lowering that T and the wandering gradually settles into descending โ simulated annealing, named for the way a cooled metal finds its lowest-stress arrangement. And the acceptance rule it cools is the same Boltzmann factor that appears on the 1 + 1 ladder as softmax: the formula deciding which word a chat model says next was written in 1868 to describe how molecules share out energy. Two centuries of people trying to compute the uncomputable, arriving repeatedly at the same handful of moves.
So to teach the model a new face, just run the loop again on pictures of that face. This works, and it is a disaster. You have now produced a second twelve-billion-dial model โ many gigabytes โ that knows one character. A cast of thirty characters is thirty copies. Worse, the loop happily wanders dials that had nothing to do with faces, and the model quietly gets worse at everything else while it learns your one thing.
It is repainting the entire house because you wanted a different picture in the hall.
Here is the idea the whole technique rests on, and it is a genuinely beautiful one.
Picture a hundred-by-hundred multiplication table. Ten thousand numbers on the page. But it is not ten thousand independent facts โ you can rebuild every one of them from the numbers 1 to 100 down the side, 1 to 100 across the top, and one rule. Two hundred numbers and a rule regenerate ten thousand. The grid looked big; its actual content was small.
Enormous grids that are secretly describable by a handful of lists are said to be low rank, and it turns out that the change you want to make to a model when teaching it one new thing is almost always of this kind. The adjustment looks like a wall of twelve billion tweaks. Its real content is a few patterns.
So: freeze the model. Do not touch a single one of the original dials. Beside each big grid of them, hang a small pair of lists, and train only those. When the model runs, it adds the little pair's contribution to the frozen grid's output and carries on. This is LoRA โ a Low-Rank Adaptation. The word "rank" is just how many of those pattern-pairs you allow. We used sixteen.
Our finished adapter is 131 megabytes against a base model many times larger; it trained in 8.6 minutes for $2. Because the original was never touched, the adapter snaps on and off like a lens โ and tomorrow's character gets its own, with no copy of the model in between. This is the same manoeuvre used to specialize large language models. When you read that a model was "post-trained" or "adapted," this is very often the literal mechanism.
The adapter needs a handle. You invent a string of letters the model has never seen โ ours was hiro_sc_v1 โ and put it in the caption of every training picture. Because the machine has no prior associations for that nonsense word, the only meaning it can acquire is whatever is common to the pictures it labels.
Which cuts both ways, and this is the trap. Anything constant across your pictures gets welded to the name. If the character wears the same jacket in all twenty-four, the name now means "this man in that jacket," and you will never get him out of it. The defence is to name the removable things explicitly in the caption โ the jacket, the sword, the motorcycle โ so the model files them under those words instead of under him.
The name buys something else that is easy to miss and turns out to matter as much as the identity: room in the prompt. A prompt is not free-form โ every model has a budget, and attention thins as you spend it.
The scene prompts written for this project average 85 words, of which 37 โ nearly half โ say who he is, leaving the remainder for camera, action, setting and art style combined. The wider storyboard pack shows the same shape: subject blocks average 53 words inside 141-word prompts. And the budgets are hard limits, not suggestions โ one model declares a 50-word character allowance with the note "front-loads attention; outfit must appear in the first ~20 words"; another allows 25, at which point the composer is already truncating the description to make it fit. We were losing the man in order to afford the scene.
The trigger word is two words. The same prompt with the adapter runs 50 words instead of 85 โ thirty-five words handed straight back to scene direction, and identity's share of the prompt falling from 44% to 4%. That is the difference between specifying several key details of a shot at once and hoping the model happens to honour them.
Ordinarily you train on photographs. We had a paragraph in a novel. So the reference set was manufactured: sixty candidate images generated from the character bible, then curated down by eye to twenty-four.
The curation is not housekeeping; it is the training signal, and rejecting is more important than accepting. A model trained on images that disagree about a face learns the average of them โ and the average of six nearly-right men is a seventh man who resembles nobody. Near-misses do more damage than obvious rejects, because they survive review.
The adapter arrives with a volume knob: how much of its contribution to add. At full strength it is a very heavy accent โ you know exactly who is speaking, but the words start to go. Turned down, the accent stays audible and the sentence comes back. Remember this dial; it decides the entire result.
Two commitments made this a measurement rather than a demonstration.
The thresholds were written down before the run. What would count as success โ how many points of improvement, on which axes โ was fixed the day before, in the design document. This sounds like bureaucracy and is in fact the whole difference between an experiment and a sales pitch: it stops you from grading your own homework once you have seen the answers.
Everything was held constant except identity. Twelve renders: six scenes, each drawn twice from the same random seed, in the same style, with the same scene description. In one version the character is described in words; in the other he is summoned by his trigger word and the adapter. Same everything else. Any difference you see is the adapter.
Each render was then scored 1โ5 on five axes โ does the face match, does the build match, are his canonical traits right, did the model obey the scene, and did the requested art style survive.
| Step | Where | Cost |
|---|---|---|
| 60 candidate references | image API | ~$3.00 |
| Training, 1000 steps | hosted GPU | $2.00 |
| 12 test renders | hosted GPU | ~$0.35 |
| 6 correction renders | hosted GPU | ~$0.20 |
| 6 pencil-style renders | hosted GPU | ~$0.20 |
| Total | one afternoon | ~$5.75 |
Worth stating plainly, because the number is the point: training a custom model is no longer an industrial act. The expensive parts of this project were choosing twenty-four pictures by eye and writing down the thresholds beforehand. The machine learning cost two dollars.
The dataset's constant is the lesson. Whatever every training example shares โ the person, the lighting, the jacket, the photographic look โ is what the model actually learns. It cannot distinguish the thing you meant from the thing that merely happened to be there every time. Vary everything you do not want taught.
Rejecting is the craft. Twenty-four agreeing images beat forty that nearly agree, because disagreement averages into mush.
Approve the cast before you train it. Curation picks among candidates; it cannot conjure one that was never generated. So the decisive review is the one before curation โ someone who knows what the character is supposed to look like, looking at the raw sheet and saying yes or no. Skip that and you can run a flawless training pass and end up with a faithful, consistent, high-scoring portrait of the wrong man. Ours came out fifteen years too old and nothing downstream could have noticed, because every number we measured was about consistency, and it was perfectly consistent.
Two models will read the same clear words differently. This needs no ambiguity in the writing to happen โ ours said a man of about thirty and got a twenty-something from one model and a forty-something from the other. So when the pictures you train on come from one model and the pictures you make come from another, the adapter is a courier carrying the first model's interpretation into the second one's world, permanently and silently. Decide which model's reading you actually want before you train on it.
A trained identity is also a prompt-budget saving. Thirty-seven words of description became two, and the whole prompt shrank from 85 words to 50. Everything a description spends on who is taken from what is happening, and prompt budgets are small and front-loaded โ so the adapter buys control of the shot as surely as it buys the face.
Strength is a setting, not a property. A capability can be present and still fail its evaluation because the knob is wrong. Test the dial before you condemn the training โ it costs cents and it moved three of four gates here.
Write the thresholds down first. This one costs nothing and is the only reason the sentence "it improved face identity by 1.17 points" means anything at all.
The ladder above describes a process. That process is now being built into an application, in a second workbench running on OpenAI's Codex, and everything below this line was made with it.
It ships in two builds because it has no model of its own. It is a director, not an engine: it holds the character bible, runs the casting gate, assembles the reference set, submits the training, and keeps every selection bound to the file it came from โ but the actual reading, writing and rendering is done by the coding agent you already pay for. So there is a build that drives Claude Code and a build that drives Codex, and you install whichever one your subscription covers.
Requires an active subscription that grants Claude Code or Codex privileges. The app supplies the process and the memory; your agent does the work, on your account, on your machine.
A screen of the application goes here once the interface settles.
What it does, in the order it does it. You give it the book. It reads the whole thing and derives a character bible — who exists, what they look like, where the text actually says so. For any character you want to be able to draw twice, it generates a casting sheet and stops: you pick the face, because the rung that matters most is the one no metric can stand in for. From your pick it builds the reference set, captions it so the trigger word carries identity and nothing else, submits the training run, and reports back what the adapter cost and whether it holds. Then you direct scenes in plain language, with the character's name doing the work that thirty-seven words of description used to do, and it keeps every accepted frame bound to its source and its hash so a selected picture can never be quietly replaced by a regenerated one.
That last part is not a detail. Halfway through the sequence below, an automated compositor was asked to assemble four accepted panels and silently substituted an unselected fourth. The fix was to stop generating at assembly time altogether — selected frames are now cut and copied deterministically and checked against their hashes. It is the same lesson as the casting gate, one layer down: the machine is good at making pictures and bad at knowing which one you chose.
Raven was the second character through the process, and the first one run with the casting gate in place from the beginning. The portrait on the left is the frame that was approved before any training happened — the one the whole reference set was built from. The frame on the right is the trained adapter doing its job: a different night, a different lens, a scene the portrait never showed.
The last piece is what all of it is for. This is a single scene from the novel — a machine that thinks it is a dog crosses a street to catch a grenade — cut as eight beats. Two characters carry it, both of them trained: the man, and the machine. It opens on the man, because the stakes have to land before the rescue can mean anything. Two of the beats are the same seconds from inside the animal's head; the rest is the street. And the axis of the camera crosses it only once the machine has passed him, because a storyboard that breaks that rule stops reading as one place.
Method and thresholds from the project's character-consistency design document (August 2026). Base model FLUX.1 [dev]; rank-16 LoRA, 1000 steps, learning rate 5e-4, hosted trainer. All twenty-four reference images and all twenty-four renders (twelve paired, six at reduced strength, six in pencil) were generated for this study; the character is from Neal Stephenson's Snow Crash and this work is an internal methods prototype, not a licensed adaptation. Full evidence record, per-axis scores, and the reproducible pipeline live with the project artifacts.