character-consistent-media-transformation
Overview
Use this class-level workflow when a user wants authorized source photographs or videos transformed to feature a different fictional character while retaining controlled aspects of the source.
Three contracts must be separated before generation:
- Creative remix: preserve broad composition, mood, action, palette, or pacing while changing wardrobe, props, text, and micro-composition.
- Sibling-image reimagination: preserve roughly 85–90% of source structure while changing the visible character and a small, explicitly selected set of role-preserving existing elements so the output feels like a companion photograph from the same visual world.
- Exact owned-source replacement: preserve the source photograph or frame as an immutable base and alter only visible identity-bearing regions.
Do not apply remix changes after the user has established ownership/authorization and explicitly requested exact reproduction. Do not mistake sibling-image variation for either pixel-exact replacement or a broad restaging.
For the live multi-model benchmark, provider contradictions, scores, costs, and source-preservation evidence that informed this workflow, read references/exact-character-replacement-benchmark-2026-07.md.
For the follow-up benchmark covering full-person, partial-person, shadow-only, hand-only, and no-person sources—and the correction from over-specified prompting to source-aware input routing—read references/dynamic-source-presence-benchmark-2026-07.md.
For the account-scale sibling-image workflow, character/world thesis pattern, semantic-drift budget, model evidence, and messaging-delivery lesson, read references/sibling-image-reimagination-2026-07.md. Copy templates/sibling-image-prompt.txt as the generic starter prompt.
Trigger Conditions
Load this skill when the user asks for any of the following:
- “same image, different character”
- “same composition, but slightly different details”
- “reimagine this account through a new character/world thesis”
- “replace me with our avatar/character”
- “keep everything exact except the person”
- identity-consistent image batches derived from source media
- keyframe-first character-consistent video production
- multi-model comparison for photorealistic character replacement
- repair of identity drift, source-person leakage, or over-regenerated backgrounds
Establish whether the source is owned/authorized or merely inspirational.
Remix contract
Use broad structural attributes only:
- framing and camera language
- setting and palette
- action category and pacing
- fashion/prop category
- lighting and mood
Change identity, wardrobe details, props, copy, and scene specifics. Avoid disguised duplication.
Sibling-image contract
Use this when the user wants each source to remain immediately recognizable in composition and narrative role, but not literally duplicated.
- Build a 60–100 word character/world thesis describing the character's worldview, recurring rituals, materials, palette, environments, and camera intimacy. The thesis should explain empty landscapes and object details as well as portraits; continuity must not depend on forcing the character into every frame.
- Preserve roughly 85–90% of the source's structural blueprint: camera position, framing, visual hierarchy, subject count, visible anatomy, gesture, light, time of day, and mood.
- Run a candidate-substitution pass before prompting. Separate immutable anchors (camera, crop, action, subject count, pose, light, major geometry) from existing replaceable nouns (animal breed, plant/tree species, vehicle model, textile construction, flower/fruit/book, gear, cargo, ground material, or one landscape detail). Choose 2–4 candidates on visually rich scenes and 1–2 on sparse scenes; never add objects merely to meet a quota.
- Make substitutions perceptible rather than vague. Each replacement should be clearly different in type, species, model, material, pattern, equipment design, or wording—not merely a tiny recolor—while preserving narrative role, approximate silhouette, scale, position, visual weight, and color mass. When the source has identifiable background vegetation or architecture, prefer at least one background candidate so the result is not just a character/foreground swap.
- If text is part of the visual, either preserve it or rewrite concise original copy in the target character's voice while keeping placement, line count, hierarchy, and typographic weight. Remove or replace third-party branding when the output is a sibling work rather than an exact owned-source edit.
- Route identity references by visible source content exactly as in Section 2. Empty, shadow-only, hand-only, torso-only, and back-turned hidden-face frames must not receive irrelevant face portraits.
Use templates/sibling-image-prompt.txt as the starter and generate one concise CANDIDATE_SWAPS list per source. A vision model can propose candidates, but QA must reject changes to primary action, camera angle, subject count, panel layout, or major geometry. For account-scale work, document whether each selected item is a still, carousel primary slide, selected carousel slide, or video-cover image; “one complete selected image per post” is not the same as downloading every carousel asset. In collages, preserve panel count/order and apply repeated identity, clothing, gear, and environment substitutions consistently across panels.
Exact owned-source contract
Preserve:
- canvas, aspect ratio, crop, and camera position
- lens perspective, head angle, gaze, and expression
- body pose, silhouette, hands, fingers, and occlusions
- clothing, folds, jewelry, props, flowers, and food
- architecture, foliage, shadows, highlights, reflections, and depth of field
- sensor grain, compression texture, color grade, and typography
Change only visible identity-bearing regions. If the original crop omits a face or body part, do not reveal or invent it.
2. Build a Clean Fictional Identity Anchor
Before the batch, create and validate:
- front portrait
- three-quarter portrait
- profile portrait
- full-body view
Lock age, facial geometry, eye color, freckles/skin details, hair color and texture, body proportions, and neutral styling.
For multi-reference editing, prefer clean single-person crops over a dense contact sheet. Remove labels, panel bleed, and distracting backgrounds. A model should not have to guess which panel defines the identity.
Do not automatically attach face portraits to every source. First classify what the source visibly contains, then route references accordingly:
- No visible person: source only; identity references are more likely to hallucinate a character than help.
- Shadow/reflection only: usually source only; use a non-face silhouette reference only when identity must be encoded.
- Hand, legs, torso, back, or cropped head only: use at most one crop-matched body/hair/skin reference that does not prominently show a face.
- Face visible: use one angle-matched face reference; add a second only when live testing proves it improves identity without overriding crop or pose.
The source controls presence, anatomy, crop, and occlusion. References control only identity traits that already exist in-frame.
Some multi-reference editors let portrait-reference geometry override the source canvas. Before a paid batch, test a square and a wide source with the real references. If a square source returns portrait output, create non-destructively padded identity references matching the source aspect ratio, explicitly request the source canvas, rerun only the failed item, and verify decoded dimensions. Do not crop the identity reference merely to force its ratio; preserve the full identity content inside a neutral padded canvas.
3. Select Models Through HMI, Then Live-Validate
Use fresh Hermes Model Intelligence before choosing image/video deployments.
Hard-filter by the actual request:
- image editing rather than text-to-image
- required source/reference count
- mask support when exactness is required
- first-frame/last-frame/reference support for video
- duration, resolution, aspect ratio, privacy, and provider access
HMI/catalog capability is a shortlist, not execution proof. Before a paid grid:
- inspect the exact provider schema
- submit one representative request using the real number and type of inputs
- record output dimensions and response behavior
- replace any lane whose live validator contradicts normalized capability metadata
Do not silently weaken a benchmark by giving one model fewer identity references unless the comparison is explicitly labeled asymmetric.
4. Prefer Masked Editing for Exact Reproduction
A whole-frame prompt saying “preserve every pixel” is not a real mask. Full-frame generative editors can change crop, pose, typography, hair mass, hidden anatomy, or the background.
Production workflow:
- preserve the original as the immutable base
- segment only visible face, hair, neck, and exposed identity-bearing skin
- exclude hidden/out-of-frame anatomy from the mask
- generate the target identity only inside the mask
- composite the edited region over untouched source pixels
- preserve typography and background deterministically outside the mask
- perform localized color/texture repair only when needed
- verify mask edges at 100% for halos, seams, hair bleed, and pore/noise mismatch
Whole-frame multi-edit is acceptable for model reconnaissance, not as proof of exactness.
5. Use Short, Source-Aware Prompts
Do not try to encode every visible object, facial trait, camera property, and preservation requirement into one long prompt. Repeated face/identity language plus prominent portrait references can make an editor render a recognizable face even when the source only shows a torso, hand, shadow, or empty scene.
Prefer one visibility rule and let the model reason from the source:
Recreate IMAGE 1 faithfully. The reference image(s) show the target character. Wherever IMAGE 1 shows all or part of a person, make only that visible person or body part belong to the target character while keeping exactly the same crop and visibility. Never add or reveal a face or body part that IMAGE 1 does not show; if no person is visible, do not add one. Trust IMAGE 1 for everything else: the same scene, composition, pose, text, objects, lighting, color, texture, and natural photographic realism.
Adapt the first sentence to the provider's input notation, but keep the prompt compact. More importantly, pair it with the visibility-based reference routing in Section 2. A prompt cannot fully counteract irrelevant face references.
For masked production, the prompt can be even shorter because the mask—not prose—defines the editable region:
Make the visible masked region belong naturally to the target character. Preserve its source crop, pose, occlusion, lighting, texture, and photographic realism. Do not add anatomy outside the mask.
6. Benchmark Design
For a useful multi-model test, include a source-presence ladder rather than only portraits:
- clear frontal/three-quarter face
- difficult profile or rotated face
- partial-face crop
- torso/legs/back with no visible face
- hand-only or other body-fragment crop
- shadow/reflection only
- no-person scene
- full person with typography or complex props
A character-recreation pipeline must prove that it knows when not to render the character. On no-person sources, test source-only routing; on partial sources, test crop-matched non-face references. Do not force identical reference counts when that would violate the production routing contract—compare models on equivalent source semantics and label reference inputs explicitly.
Record:
- exact deployment ID and provider
- input schema and reference ordering
- elapsed time and estimated/quoted cost
- output dimensions, size, and checksum
- policy/validator failures
- full-resolution outputs and source/model contact sheets
7. Score Exactness and Identity Separately
Never collapse everything into “looks good.” Score at least:
- composition/crop/pose/background preservation
- typography preservation
- target identity transfer
- photorealism, texture, and lighting
- anatomical integrity
Supporting whole-frame metrics can include luminance correlation, edge correlation, RGB error, SSIM, or perceptual similarity. These are not final truth: a model can achieve excellent source similarity by leaving the original identity unchanged.
Output geometry is a hard gate. A model that changes 4:5 into a tall portrait has already failed exact reproduction even if the result is attractive.
8. Failure Recovery
Live validator contradicts registry
Treat the endpoint response as authoritative. Replace the lane and capture the mismatch as provider evidence; do not generalize that the entire tool or provider is broken.
One source triggers provider policy
Ownership and safe_mode: false do not override upstream policy. Retry once with benign authorized-retouch wording only if the classification may be prompt-sensitive. If equivalent input is rejected again, stop retrying, keep successful outputs as supplemental evidence, and run a complete replacement lane.
Hidden-face invention
Do not accept a model-created face when the original crop concealed it. First remove face-dominant identity references and use a crop-matched non-face reference. For exact work, use a mask that excludes nonexistent face area or treat that source as a preservation-only case.
Character hallucinated into a no-person source
Do not keep rewriting the negative prompt while continuing to send portrait references. Remove all identity references and rerun with the source alone. For mixed batches, classify and route each source before generation; do not use one unconditional reference bundle across the whole batch.
Transient empty image result
Preserve completed outputs and retry only the failed source once or twice with bounded backoff. If the same semantic request repeatedly fails, record the failure and stop rather than restarting the entire lane.
Identity transfer fails but source preservation is high
Use the result as evidence of conservative editing, not a successful character swap. Strengthen the localized mask/identity conditioning rather than regenerating the whole image.
Repeated blur or anatomy artifacts
Use one targeted edit. If artifacts persist, regenerate from clean identity refs rather than recursively editing a flawed face. Deterministic cropping is acceptable only when it preserves the requested composition contract.
9. Video Extension
For video, lock the character in approved still keyframes before animation.
- animate short 5–10 second shots rather than reproducing a long reel in one call
- use stable poses, separated hands, clear faces, and simple biomechanics
- inspect multiple chronological frames for identity drift, anatomy, flicker, and background warping
- separate generated/native audio from exact supplied-audio preservation
- keep the original approved keyframe and provider job metadata with each clip
10. Delivery
Deliver actual media, not only prompts or plans. On messaging channels, put the direct visual artifacts before a long explanation so the user can immediately see that generation finished:
- primary-model contact sheet;
- source/model comparison matrix;
- concise verdict;
- full-resolution model-specific archives and report.
Package:
- source/model comparison matrix
- full-resolution outputs
- identity references
- quantitative metrics and concise report
- ZIP archive with a manifest and checksums
If a user asks “where are the images?”, resend the concrete contact-sheet image(s) immediately; do not answer with another path-only explanation. If a large archive fails on the messaging channel, split it into logical model-specific parts below the practical attachment threshold while retaining unsplit local masters.
Verification Checklist
- [ ] Source ownership/authorization and desired contract are explicit.
- [ ] Sibling-image jobs have one concise character/world thesis and an explicit candidate-substitution list per source (normally 2–4 on rich scenes, 1–2 on sparse scenes).
- [ ] Candidate substitutions are visibly different in type/species/model/material/pattern, not merely recolored; identifiable background elements were considered; narrative role, approximate silhouette, scale, position, visual weight, and color mass remain preserved.
- [ ] Every source is classified as face-visible, partial-body, shadow/reflection, or no-person.
- [ ] Identity references are clean, internally consistent, and routed by visible source content.
- [ ] No-person sources are run source-only; partial-body sources avoid face-dominant references.
- [ ] The prompt is short and source-aware rather than a face-heavy preservation inventory.
- [ ] HMI selection is fresh and exact schemas are inspected.
- [ ] Every model passes one live request-shape preflight.
- [ ] Exact tests use the same inputs and prompt across lanes.
- [ ] Output dimensions and aspect ratios are recorded.
- [ ] Composition and identity are scored separately.
- [ ] Partial/hidden faces are not invented.
- [ ] Typography/background outside the mask remains untouched.
- [ ] Final images are inspected at 100%; videos are checked across time.
- [ ] Costs, policy blocks, retries, and supplemental outputs are disclosed.
Common Pitfalls
- Writing a long prompt that repeatedly emphasizes face/identity details and accidentally turns the reference portrait into the composition target.
- Sending the same face-heavy reference bundle to full-person, partial-person, shadow-only, and no-person sources.
- Treating an empty-scene recreation as an identity-transfer problem instead of running it source-only.
- Treating “same model family” as the same provider contract.
- Trusting normalized multi-reference capability without a live probe.
- Using an identity contact sheet when one crop-matched reference—or no reference—would condition better.
- Rewarding an attractive redesign in an exact-reproduction benchmark.
- Rewarding unchanged source pixels when the target identity was never transferred.
- Expecting a text prompt to enforce pixel-level preservation without a mask.
- Continuing equivalent policy-triggering retries.
- Restarting a completed batch because one transient result was empty instead of retrying only the failed source.
- Comparing models with different input counts without labeling the asymmetry.
- Treating a sibling-image request as either a broad redesign or an identity-only recolor. Vague “slightly different” prompting often changes the dog or shirt but leaves obvious background candidates such as tree species untouched; select and name the existing in-place substitutions before generation.
- Strengthening identity language until the model zooms/recomposes a full-body source into a portrait; this is a geometry failure, not a successful identity tradeoff.
- Forcing the target character into empty account images instead of using the shared world thesis to vary animals, plants, objects, vehicles, or landscapes.
- Burying contact sheets and comparison images after a long report, leaving the user unsure whether images were actually delivered.