character-consistent-media-transformation

/home/avalon/.hermes/skills/media/character-consistent-media-transformation/SKILL.md · raw

Character-Consistent Media Transformation

Overview

Use this class-level workflow when a user wants authorized source photographs or videos transformed to feature a different fictional character while retaining controlled aspects of the source.

Three contracts must be separated before generation:

  1. Creative remix: preserve broad composition, mood, action, palette, or pacing while changing wardrobe, props, text, and micro-composition.
  2. Sibling-image reimagination: preserve roughly 85–90% of source structure while changing the visible character and a small, explicitly selected set of role-preserving existing elements so the output feels like a companion photograph from the same visual world.
  3. Exact owned-source replacement: preserve the source photograph or frame as an immutable base and alter only visible identity-bearing regions.

Do not apply remix changes after the user has established ownership/authorization and explicitly requested exact reproduction. Do not mistake sibling-image variation for either pixel-exact replacement or a broad restaging.

For the live multi-model benchmark, provider contradictions, scores, costs, and source-preservation evidence that informed this workflow, read references/exact-character-replacement-benchmark-2026-07.md.

For the follow-up benchmark covering full-person, partial-person, shadow-only, hand-only, and no-person sources—and the correction from over-specified prompting to source-aware input routing—read references/dynamic-source-presence-benchmark-2026-07.md.

For the account-scale sibling-image workflow, character/world thesis pattern, semantic-drift budget, model evidence, and messaging-delivery lesson, read references/sibling-image-reimagination-2026-07.md. Copy templates/sibling-image-prompt.txt as the generic starter prompt.

Trigger Conditions

Load this skill when the user asks for any of the following:

1. Confirm the Transformation Contract

Establish whether the source is owned/authorized or merely inspirational.

Remix contract

Use broad structural attributes only:

Change identity, wardrobe details, props, copy, and scene specifics. Avoid disguised duplication.

Sibling-image contract

Use this when the user wants each source to remain immediately recognizable in composition and narrative role, but not literally duplicated.

  1. Build a 60–100 word character/world thesis describing the character's worldview, recurring rituals, materials, palette, environments, and camera intimacy. The thesis should explain empty landscapes and object details as well as portraits; continuity must not depend on forcing the character into every frame.
  2. Preserve roughly 85–90% of the source's structural blueprint: camera position, framing, visual hierarchy, subject count, visible anatomy, gesture, light, time of day, and mood.
  3. Run a candidate-substitution pass before prompting. Separate immutable anchors (camera, crop, action, subject count, pose, light, major geometry) from existing replaceable nouns (animal breed, plant/tree species, vehicle model, textile construction, flower/fruit/book, gear, cargo, ground material, or one landscape detail). Choose 2–4 candidates on visually rich scenes and 1–2 on sparse scenes; never add objects merely to meet a quota.
  4. Make substitutions perceptible rather than vague. Each replacement should be clearly different in type, species, model, material, pattern, equipment design, or wording—not merely a tiny recolor—while preserving narrative role, approximate silhouette, scale, position, visual weight, and color mass. When the source has identifiable background vegetation or architecture, prefer at least one background candidate so the result is not just a character/foreground swap.
  5. If text is part of the visual, either preserve it or rewrite concise original copy in the target character's voice while keeping placement, line count, hierarchy, and typographic weight. Remove or replace third-party branding when the output is a sibling work rather than an exact owned-source edit.
  6. Route identity references by visible source content exactly as in Section 2. Empty, shadow-only, hand-only, torso-only, and back-turned hidden-face frames must not receive irrelevant face portraits.

Use templates/sibling-image-prompt.txt as the starter and generate one concise CANDIDATE_SWAPS list per source. A vision model can propose candidates, but QA must reject changes to primary action, camera angle, subject count, panel layout, or major geometry. For account-scale work, document whether each selected item is a still, carousel primary slide, selected carousel slide, or video-cover image; “one complete selected image per post” is not the same as downloading every carousel asset. In collages, preserve panel count/order and apply repeated identity, clothing, gear, and environment substitutions consistently across panels.

Exact owned-source contract

Preserve:

Change only visible identity-bearing regions. If the original crop omits a face or body part, do not reveal or invent it.

2. Build a Clean Fictional Identity Anchor

Before the batch, create and validate:

Lock age, facial geometry, eye color, freckles/skin details, hair color and texture, body proportions, and neutral styling.

For multi-reference editing, prefer clean single-person crops over a dense contact sheet. Remove labels, panel bleed, and distracting backgrounds. A model should not have to guess which panel defines the identity.

Do not automatically attach face portraits to every source. First classify what the source visibly contains, then route references accordingly:

The source controls presence, anatomy, crop, and occlusion. References control only identity traits that already exist in-frame.

Some multi-reference editors let portrait-reference geometry override the source canvas. Before a paid batch, test a square and a wide source with the real references. If a square source returns portrait output, create non-destructively padded identity references matching the source aspect ratio, explicitly request the source canvas, rerun only the failed item, and verify decoded dimensions. Do not crop the identity reference merely to force its ratio; preserve the full identity content inside a neutral padded canvas.

3. Select Models Through HMI, Then Live-Validate

Use fresh Hermes Model Intelligence before choosing image/video deployments.

Hard-filter by the actual request:

HMI/catalog capability is a shortlist, not execution proof. Before a paid grid:

  1. inspect the exact provider schema
  2. submit one representative request using the real number and type of inputs
  3. record output dimensions and response behavior
  4. replace any lane whose live validator contradicts normalized capability metadata

Do not silently weaken a benchmark by giving one model fewer identity references unless the comparison is explicitly labeled asymmetric.

4. Prefer Masked Editing for Exact Reproduction

A whole-frame prompt saying “preserve every pixel” is not a real mask. Full-frame generative editors can change crop, pose, typography, hair mass, hidden anatomy, or the background.

Production workflow:

  1. preserve the original as the immutable base
  2. segment only visible face, hair, neck, and exposed identity-bearing skin
  3. exclude hidden/out-of-frame anatomy from the mask
  4. generate the target identity only inside the mask
  5. composite the edited region over untouched source pixels
  6. preserve typography and background deterministically outside the mask
  7. perform localized color/texture repair only when needed
  8. verify mask edges at 100% for halos, seams, hair bleed, and pore/noise mismatch

Whole-frame multi-edit is acceptable for model reconnaissance, not as proof of exactness.

5. Use Short, Source-Aware Prompts

Do not try to encode every visible object, facial trait, camera property, and preservation requirement into one long prompt. Repeated face/identity language plus prominent portrait references can make an editor render a recognizable face even when the source only shows a torso, hand, shadow, or empty scene.

Prefer one visibility rule and let the model reason from the source:

Recreate IMAGE 1 faithfully. The reference image(s) show the target character. Wherever IMAGE 1 shows all or part of a person, make only that visible person or body part belong to the target character while keeping exactly the same crop and visibility. Never add or reveal a face or body part that IMAGE 1 does not show; if no person is visible, do not add one. Trust IMAGE 1 for everything else: the same scene, composition, pose, text, objects, lighting, color, texture, and natural photographic realism.

Adapt the first sentence to the provider's input notation, but keep the prompt compact. More importantly, pair it with the visibility-based reference routing in Section 2. A prompt cannot fully counteract irrelevant face references.

For masked production, the prompt can be even shorter because the mask—not prose—defines the editable region:

Make the visible masked region belong naturally to the target character. Preserve its source crop, pose, occlusion, lighting, texture, and photographic realism. Do not add anatomy outside the mask.

6. Benchmark Design

For a useful multi-model test, include a source-presence ladder rather than only portraits:

A character-recreation pipeline must prove that it knows when not to render the character. On no-person sources, test source-only routing; on partial sources, test crop-matched non-face references. Do not force identical reference counts when that would violate the production routing contract—compare models on equivalent source semantics and label reference inputs explicitly.

Record:

7. Score Exactness and Identity Separately

Never collapse everything into “looks good.” Score at least:

  1. composition/crop/pose/background preservation
  2. typography preservation
  3. target identity transfer
  4. photorealism, texture, and lighting
  5. anatomical integrity

Supporting whole-frame metrics can include luminance correlation, edge correlation, RGB error, SSIM, or perceptual similarity. These are not final truth: a model can achieve excellent source similarity by leaving the original identity unchanged.

Output geometry is a hard gate. A model that changes 4:5 into a tall portrait has already failed exact reproduction even if the result is attractive.

8. Failure Recovery

Live validator contradicts registry

Treat the endpoint response as authoritative. Replace the lane and capture the mismatch as provider evidence; do not generalize that the entire tool or provider is broken.

One source triggers provider policy

Ownership and safe_mode: false do not override upstream policy. Retry once with benign authorized-retouch wording only if the classification may be prompt-sensitive. If equivalent input is rejected again, stop retrying, keep successful outputs as supplemental evidence, and run a complete replacement lane.

Hidden-face invention

Do not accept a model-created face when the original crop concealed it. First remove face-dominant identity references and use a crop-matched non-face reference. For exact work, use a mask that excludes nonexistent face area or treat that source as a preservation-only case.

Character hallucinated into a no-person source

Do not keep rewriting the negative prompt while continuing to send portrait references. Remove all identity references and rerun with the source alone. For mixed batches, classify and route each source before generation; do not use one unconditional reference bundle across the whole batch.

Transient empty image result

Preserve completed outputs and retry only the failed source once or twice with bounded backoff. If the same semantic request repeatedly fails, record the failure and stop rather than restarting the entire lane.

Identity transfer fails but source preservation is high

Use the result as evidence of conservative editing, not a successful character swap. Strengthen the localized mask/identity conditioning rather than regenerating the whole image.

Repeated blur or anatomy artifacts

Use one targeted edit. If artifacts persist, regenerate from clean identity refs rather than recursively editing a flawed face. Deterministic cropping is acceptable only when it preserves the requested composition contract.

9. Video Extension

For video, lock the character in approved still keyframes before animation.

10. Delivery

Deliver actual media, not only prompts or plans. On messaging channels, put the direct visual artifacts before a long explanation so the user can immediately see that generation finished:

  1. primary-model contact sheet;
  2. source/model comparison matrix;
  3. concise verdict;
  4. full-resolution model-specific archives and report.

Package:

If a user asks “where are the images?”, resend the concrete contact-sheet image(s) immediately; do not answer with another path-only explanation. If a large archive fails on the messaging channel, split it into logical model-specific parts below the practical attachment threshold while retaining unsplit local masters.

Verification Checklist

Common Pitfalls