The direction language
A pose is not a mood, it is geometry: where the weight sits, what the hands do, which way the head turns. FLAM's lookbook variations are now a named posture library where every entry carries a real photograph and compiles to physical phrases — because we measured photographs against pose diagrams as controllers, and it was not close: 10 out of 12 against 4.5.

A pose is not a mood. It is geometry and intent: where the weight sits, what the hands are doing, which way the torso and the head turn. Ask an image model for "relaxed" and it fills every unstated joint with a default, and defaults look like mannequins. So the lookbook stopped taking adjectives and started taking direction — a library of named postures, each carrying a real photograph and a body of physical phrases.
Nobody names poses, which is strange
While designing this we surveyed the field: some fifteen vendors, nine Shopify apps, five help centres, four API schemas. Not one shows a user a named pose. The single tool we found with pose names at all keeps them in its developer documentation. Even outside software the ground is unclaimed — Coco Rocha's study of a thousand poses names none of them.
That gap is the feature. A posture with a name can be chosen twice the same way, saved as a house default, handed to the automatic lookbook, corrected in one word. "The Lean, but composure at angular hauteur" is direction. "Make it less stiff" is a wish.
Every stop compiles to geometry
Under each named posture is the rule that makes it work: every control compiles to a physical phrase, never a mood word. The anti-stiffness formula, learned from production work rather than theory, is a weight shift plus a hand task plus a fixed framing — not the word "natural" anywhere.
The phrases are camera-relative on purpose: "weight on the foot nearer frame left," never "her left leg." Models fail her-left against her-right almost every time, and an instruction the model reliably misreads is worse than none.
The measurement: photograph 10, diagram 4.5
The academic evidence leans toward pose skeletons, the stick-figure diagrams, as the strongest published pose controller. The founder's instinct said photographs. Nobody had published the head-to-head, so we ran it: one identity, four postures, both reference kinds, every job a fresh submission, scored one point per geometry phrase visibly obeyed in the output.
The photograph took 10 of 12 points. The diagram took 4.5.
The diagram did not lose at the control step; it lost at the derivation step. The diagram for Over-the-Shoulder came back as a generic front-facing pictogram with no turn in it at all, and the control faithfully rendered her front-on, back never turned: 0 of 3. The Contrapposto diagram came back mirrored, and the control copied the mirror. The published benchmark that scored skeletons so highly used hand-authored diagrams; ours had to be derived from the pose, and the derivation is where the geometry died. A photograph needs no derivation. And it means the frame the director sees on the card is exactly the frame the model gets, with nothing riding hidden.
What the judge had to learn, twice
Generating the library's twelve seed frames taught us that our own judge could not see a pose miss. It scored a clean 100 on a Back Walk-Away whose head turned over the wrong shoulder, and another 100 on a contrapposto with the counter-tilt flattened out. The rubric graded craft — hands, seams, anatomy — and pose fidelity was simply not one of its questions. So the judge learned poses: a posed frame is now graded against the posture's own phrases, the weight and the hands and the turns.
Then it learned again, because the founder's eye caught what the new question still missed: a slumped Profile frame that graded well. A probe confirmed the mechanism: the pose question enumerated legs, hands, feet, torso and head-turn, and a model checks what it is told to check, so an unenumerated spine was invisible. The enumeration now names the carriage: spine long or slumped, head stacked over the shoulders or carried forward, chin level or dropped.
We are honest about the ceiling. Measured on the same slumped frame, current production-lane vision models read carriage unreliably: one fires the flag in roughly one draw in four, another not at all. The enumeration stays anyway, because a sometimes-question beats a never-question and it names the miss for the retry brief. But no score is ever decided on carriage alone, and the honest fix, a stronger vision model for pose-bearing frames, is on the roadmap by name.
The drag toward the conventional
One more thing the measurement settled that we had not asked. Image models normalise: they straighten what should bend. The counter-tilt of a contrapposto came back mild in every take, in both lanes. "Forearms resting on the thighs" became hands on the knees in both seated frames. A first Back Walk-Away replaced the stride with a stand.
The remedy is the loop, not the wording: judge the frame against the posture's reference, name the specific miss in the retry, run it again. Four of the twelve seed frames needed exactly one such retry, and the named miss fixed each one.
Related
- Where postures are chosen: the lookbook.
- Why the reference frame is a photograph of a person and not a mannequin: the model you already have.
- What shipped, and when: /changelog#the-direction-language.
Questions
- How do I control the model's pose in FLAM?
- You pick a named posture from a library — Contrapposto, Mid-Stride, Over-the-Shoulder, The Lean — and every entry shows you a real photograph of the pose, not a diagram. Behind the card, the entry compiles to physical phrases about weight, hands and head, and the photograph itself rides the develop as a reference.
- Why photographs instead of stick figures or pose skeletons?
- Because we measured both as pose controllers over the same identity and the same four postures, and the photograph scored 10 out of 12 geometry points against the diagram's 4.5. The diagrams failed at the derivation step — one came back mirrored, one as a generic front-facing pictogram with no turn at all — and the model faithfully copied the errors. The frame you see is the frame the model gets.
- What is a geometry phrase?
- A pose instruction that names something physical and checkable: which leg carries the weight, where each hand is, the turn of the torso and head. Mood words do not survive contact with an image model — 'relaxed' and 'not stiff' produce mannequins, because the model fills unstated geometry with defaults, and defaults look like mannequins.
- Can the judge tell when a pose is wrong?
- It grades a posed frame against the posture's own phrases — the weight, the hands, the turns — and, since the latest rubric, the carriage: spine long or slumped, head stacked or carried forward, chin level or dropped. We are honest about the ceiling: current vision models read carriage unreliably, so it names the miss for the retry without ever deciding a score on its own.