Field guide
The Prompting Guide
Everything here comes from what actually works in the Studio — no magic words, no filler. Ten minutes of reading buys you better generations on the first try instead of the fifth.
Six principles before you type anything
- 01
Keep it to one or two people. The subject should fill 60%+ of the frame with a clearly visible face. Crowds fall apart.
- 02
Name the shot and the angle. Close-up, medium shot, or full body — plus facing the camera, three-quarter view, or profile. If you don’t pick, the model picks for you.
- 03
Describe the face you want. Expression, lighting on the skin, one or two distinctive features. Faces carry the photo.
- 04
Two people need geometry. Say who is where: side by side, facing each other, one behind the other — and give each their own pose and gaze.
- 05
Keep hands simple. Resting on thighs, at her sides, holding a glass. Complex gestures and interlocked fingers are where hands go wrong.
- 06
One change of mind costs a re-roll. Decide the scene before you generate; vague prompts produce generic results, not options.
The ten-slot prompt formula
Fill the slots in order. SpicyGen Images is built to read normal sentences — write it as plain English, no keyword soup. Z Image Spicy and SpicyGen Images Flash prefer the classic stacked-descriptor register: same slots, same order, but comma-separated phrases instead of full sentences. Skip an optional slot rather than padding it.
Assembled, it reads like this
A young woman with long dark hair in a luxury hotel bedroom with soft white sheets, medium close-up at eye level. She wears a white blouse unbuttoned with a black lace bra visible, sitting on the edge of the bed with her legs crossed, leaning back, a shy smile, looking up through her lashes. Soft window light from the side with a warm golden glow, shot on Kodak Portra 400, visible skin pores, sharp focus on her face with shallow depth of field.
Short on patience? Type four words and hit Enhance in the Studio — it expands rough ideas into a complete, model-ready prompt. Each model's Enhance writes in its own register: SpicyGen Images gets the plain-English structure above; Z Image Spicy and SpicyGen Images Flash get a stacked tag-style prompt with weights (details below).
Generating on Z Image Spicy or SpicyGen Images Flash?
- End every prompt with a quality tail. masterpiece, best quality, 8k, hyperrealistic, sharp focus, shallow depth of field
- Weight what matters. Wrap a phrase as (phrase:1.5) to push it harder or (phrase:0.7) to soften it — the same syntax as Qwen Edit Image Spicy.
- No idea where to start? Hit Enhance. It rewrites a rough line into a full tag-style prompt — stacked descriptors, weights, quality tail — in this model's exact register.
- Your prompt runs exactly as written. Nothing is rewritten behind the scenes — what you type (or what Enhance builds) is what generates. Use Enhance or this guide to flesh out a rough idea.
- English only. Other languages degrade quality on every model — this one included.
Expression & gaze cheat sheet
Pair one face phrase with one gaze phrase. Mixing moods reads as uncanny.
Shy
shy smile, soft blush
looking down · looking up through lashes
Alluring
seductive smile, half-smile
direct eye contact · bedroom eyes, half-lidded
Pleasure
blissful expression, parted lips, flushed cheeks
eyes half-closed · dreamy gaze
Playful
playful grin, tongue out slightly
side glance · looking over shoulder
Cool
neutral expression, cold beauty
cold stare · distant gaze
Lazy
relaxed expression, sleepy smile
heavy-lidded · lazy gaze
Light & film looks
Naming a film stock is the fastest way to change a photo's whole mood. One stock per prompt; pair it with a matching light. Golden hour + Portra 400 is the classic portrait combo.
Realism details
Two or three of these per prompt push a render from "AI-smooth" to photographic. All six at once cancels itself out — restraint is the trick.
Combinations that fight each other
When two halves of a prompt contradict, the model splits the difference and both halves lose.
✗ iPhone selfie + professional studio quality
→ casual phone-photo look OR polished professional — pick the world the photo lives in
✗ wearing panties + fully visible genitals
→ pick one — the model can’t do both
✗ bright sunlight + candlelight
→ one primary light source per image
✗ standing pose + lying on the bed
→ one clear pose only
✗ eyes closed + looking at the camera
→ use half-closed eyes looking at the camera
Editing a photo: instructions, not descriptions
Edit prompts are a different language. You're not describing a scene — you're telling the model what to change and what to leave alone.
✗ a beautiful woman with long hair wearing a red dress in a bedroom
✓ Change her dress from blue to red. Keep everything else exactly the same.
- 01
One change per edit. Adjust the outfit, generate, then fix the lighting on the result. Three small edits beat one big one — every time.
- 02
Say what to preserve. "Keep her face, expression, pose and the background exactly the same" is the most valuable sentence in editing.
- 03
Weight syntax — Spicy models only. Wrap a phrase as (phrase:1.4) to push it harder, or (phrase:0.7) to pull it back. Above 1.0 strengthens, below 1.0 softens; leave it off for normal weight. Works on Qwen Edit Image Spicy, Z Image Spicy, and SpicyGen Images Flash; our other models read plain English — no syntax needed.
- 04
Pick the right model. The model card in the editor says what each one is best at — read it before a long session. English prompts work best on every model.
Clothing
- remove all clothing
- change outfit to black lace lingerie
- add a white silk robe loosely draped
Body & expression
- adjust pose to lying on back
- change expression to seductive smile
- add sweat droplets on skin
- enhance muscle definition
Scene & light
- change background to luxury hotel room
- add warm candlelight atmosphere
- add steam effect
Video: animating a photo
Every video model here is image-to-video: you bring a photo, the prompt brings the motion. The single biggest lever is the source image — sharp, well-lit, subject large in frame, hands visible. The second biggest is matching your prompt style to the model, because the three video models speak three different languages. Adults only, always — the same verification rules as everywhere else in the Studio.
SpicyGen Vision: motion prompts & trigger words
SpicyGen Vision reads a motion prompt — describe the action, not the scene (your photo already supplies the scene). Its motion presets fire on trigger words: pick one in the editor and the trigger is added to your prompt automatically, or type it yourself exactly as shown on the chip. Then write the action in blunt, present-tense English with strong verbs:
d0gg1e. He grips her hips with both hands and pounds her hard from behind, her ass jiggling with each thrust, she moans and looks back over her shoulder.
One subject, one act per clip — five seconds is one beat of motion, not a storyline.
Wan 2.2 Spicy: the five-second formula
One paragraph, one template, one act. This model is silent and runs your prompt exactly as written — what you type is what generates.
- 01
Open with the strip — non-negotiable. The model won’t reliably start her naked. Show the clothes coming off: "She rips her top off, then shoves her pants down and kicks them away." Keep the outfit vague ("casual wear") so the model doesn’t fixate on garments.
- 02
Position change? Jumpcut. Never describe her turning around — "A jumpcut instantly places her already naked, on all fours, facing away from the camera, looking back over her shoulder." Use it for doggy, reverse cowgirl, spooning, scene changes. Facing forward the whole time (cowgirl, blowjob POV, solo) needs none.
- 03
One act per clip. Five seconds fits exactly one sex act. Two acts means neither lands. Pacing: strip in the first second, jumpcut if needed, then one continuous action to the end — no more cuts.
- 04
Demand eye contact, every scene. "She keeps her eyes locked on the camera the entire time." Facing away: "looking back over her shoulder at the camera, eyes locked." Unspecified gaze drifts.
- 05
Be anatomically blunt. Vague anatomy renders plastic. Write "bare breasts, erect nipples, vulva with trimmed pubic hair" — the pubic-hair line is what keeps skin looking real. Breasts in motion: add "bouncing" or "swaying".
- 06
Close with light + quality. The model defaults to flat, ugly light. End every prompt with: "Static shot. Soft lighting, detailed skin texture, 4K, photorealistic." Swap in "moody red-tinted", "golden hour" or "candlelit" for mood.
- 07
It’s silent. No dialogue, no moaning, no sound effects in the prompt — audio descriptions waste your prompt budget on this model.
The master template
A 5-second video. The scene opens with a medium shot of a woman fully clothed in casual wear, standing in [LOCATION], [EXPRESSION] at the camera. She rips her top off, then shoves her pants down and kicks them away. Now completely naked, bare breasts, erect nipples, and vulva with trimmed pubic hair fully exposed. [JUMPCUT IF NEEDED]. [ONE SEX ACT]. She keeps her eyes locked on the camera the entire time. Static shot. Soft lighting, detailed skin texture, 4K, photorealistic.
Wan 2.7 Spicy: directing frame by frame
The flagship runs 4–15 seconds with generated audio and gives you per-second control through "At 00:0X" lines. The 00:00 line is the anchor — it must carry all eight elements, in order:
- 01
The first frame is everything. This model animates YOUR image — pose, nudity and scene must already match the prompt. No clothes-ripping opener here: nudify with the editor first, then animate. Subject at 60–70% of frame, hands fully visible, clean background, even light.
- 02
Describe changes, not state. The 00:00 line anchors the full scene (the eight elements above). Every later "At 00:0X" line describes only what CHANGES — repeating the scene wastes the frame and causes drift.
- 03
Write the audio in. It generates sound with the picture. Prefix lines with an emotion: "[breathlessly] Come closer" — 5–10 words per line; longer lines break lip sync. Actions imply sound effects ("her heels click sharply on the marble floor"); name ambient sound explicitly if you want it.
- 04
Hold the camera still. "Camera remains completely static throughout, no zoom, pan, or tilt." Slow dolly in/out is safe; pans are risky; orbiting fails more often than it works. One camera idea per clip.
- 05
Match the action to the length. 4–6s: micro-expressions and hands. 8s: add hand-and-body moves. 10s+: full body. Overambitious action tears frames. Avoid speed words (fast, quick, suddenly) — pace with "slowly", "gently", rhythm.
- 06
50–200 words. Shorter lacks control; longer overloads and the ending gets truncated. This model also runs a built-in prompt optimizer over your text before generating — unlike our other models, where your prompt runs exactly as written.
The 5-second skeleton
At 00:00, the frame is a [SHOT] from a [ANGLE], capturing a nude woman with [hair, makeup,
expression] — [pose and partner]; [anatomical detail]; [partner: torso visible, face out of
frame]; [background]; [lighting]; 4K, sharp focus, no censorship, photorealistic.
At 00:01, [subtle change: breathing, fingers, eyes].
At 00:02, [motion intensifies; skin sheen].
At 00:03, [peak: head jerks back, mouth opens, "[moaning] yes"].
At 00:04, [settle or shift]. Camera remains completely static. No distortion. Face stable.
Longer clips: build in arcs — establish (0–3) → build (4–8) → peak (9–12) → release (13–15). Never copy earlier lines forward; every frame line needs a unique change.
Why results behave the way they do
Same prompt, different image every time?
That's how diffusion models work — every run starts from fresh noise. Generate two or three versions and keep the best. The more specific the prompt, the tighter the variation.
Hands and fingers look wrong?
Hands are the hardest thing in the field. Ask for simple, natural placement ("hands resting on her thighs"), avoid interlocked fingers and complex gestures, and re-roll — a clean-hands result usually lands within a couple of tries.
An edit changed things it shouldn't have?
Add an explicit preserve clause and shrink the ask. "Make her top more revealing. Keep her face, expression, pose, and background exactly the same" gives the model nowhere to improvise.
A generation failed?
Edit failures are usually about the source photo — try a sharper, better-lit image where the subject is large in frame, or simplify the instruction. If a generation keeps failing, contact us at contact@spicygen.xyz.