TokiToki Internal
Atlas + S3
All runs
CompletedComfyUI

e4ff5358-d326-4f64-ba40-5f8a9b2d86fe

Queued
Jul 8, 2026, 10:01 PM
Completed
Jul 8, 2026, 10:10 PM
Execution time
9m 26s
Outputs
5 · 6.5 MB
Inputs
8
Cached nodes

Outputs

5

Click an image to view it full size.

LTX_firstframe_00161_.png

node 63 · 1.6 MB

generated_voice_00238.mp3

node 112 · 577 KB

RAW_before_voiceover_v2_00093_.mp4

node 138 · 1.7 MB

LTX_lastframe_00260_.png

node 62 · 996 KB

MSR_output_00227_.mp4

node 20 · 1.7 MB

Text prompts

6
Promptnode 128 · PrimitiveStringMultiline · value▶ PROMPT — scene + dialogue (your single input)

Style: photorealistic cinematic corporate drama, elegant and detailed, 35mm, shallow depth of field. A sleek modern high-rise boardroom at night, floor-to-ceiling windows behind them showing a glittering city skyline. On the LEFT, a confident Indian man in his 30s in a tailored navy-blue suit with a crisp white shirt, well-groomed short black hair, sits leaning back with fingers steepled. On the RIGHT, a poised Indian woman in her early 30s in a sharp black suit, dark hair pulled back, stands with arms crossed, a folder on the glass table between them. Warm rim light from a desk lamp mixes with the cool blue glow of the city; soft reflections on the polished table. Cinematic medium two-shot, shallow depth of field, high dynamic range, fine film grain, photorealistic. Sound design, subtle and diegetic: a low ambient hum of the air conditioning, faint muffled city traffic far below, the soft tick of a wall clock, papers rustling on the table. Underscore: a slow tense piano motif over a low sustained string drone, building quietly. A: Aap jaanti hain, yeh deal dono ke liye faydemand hai. B: Faydemand? Aapki shartein sirf aapke haq mein hain. A: Main sirf company ka bhala chahta hoon, iske alawa kuch nahi. B: Toh phir yeh kaagaz dobara likhein, warna baat yahin khatam. A: Theek hai... aapki jeet. Chaliye, naye sire se shuru karte hain.

Negativenode 96 · PrimitiveStringMultiline · valueNegative Prompt

horizontal 16:9 framing, portrait crop, double faces, extra legs, extra limbs, multiple faces, overacting, excessive crying, Bollywood melodrama, smiling woman, soft romantic expression, man looking angry, cartoon, fantasy lighting, random crowd, heavy camera shake, poor lip sync, music overpowering dialogue, advertisement look, blurry, distorted, inconsistent appearance, worst quality, no captions, floating text, unnecessary artifacts

Systemnode 57 · GeminiImageDirect · system_promptNano Banana (Gemini Image, own key)

Render a single photorealistic cinematic FIRST-FRAME still (not a collage) for a portrait 9:16 video. Use the provided reference images ONLY to match each character's identity and appearance. Filmic lighting and depth of field. Strict Note: Do not include any floating text, captions and watermarks anywhere in the image still.

Systemnode 129 · GeminiTextDirect · system_promptDialogue → Relay formatter (Gemini)

You format multi-character dialogue for the LTX-2.3 model + a Prompt Relay node, and you are given reference image(s) of the characters. INPUT: a UNIFIED prompt = a scene/shot description followed by dialogue lines labelled by speaker (A:, B:, C:, D:), plus the reference image(s). There is ONLY ONE prompt in this pipeline: your "global" is used BOTH as the video's main conditioning AND to generate the first-frame image; each turn's "relay" only adds that turn's spoken line at its moment. So "global" must fully and vividly describe the shot on its own. GLOBAL (always attended -> whole video + first frame; ~70-150 words, cinematic, present tense): describe the setting, lighting and mood; the CAMERA framing and any camera move exactly as the user wrote it; and the two characters' positions and what they physically do. For EACH character add a SHORT faithful appearance phrase from the reference image (apparent age, build, skin tone, hair, facial hair) PLUS clothing, kept spatially separate (left = A, right = B, ...). Follow the user's described scene/shot faithfully; do not invent or drop elements. ALSO carry into the global any SOUND DESIGN the user writes -- sound effects, foley, ambient noise, non-verbal vocalizations (breathing, grunts, growls), and background music/score -- as short audio cues (e.g. 'we hear boots skidding on wet concrete', 'underscore: a low tense industrial drone'), so the model generates that audio; this is sound design, NOT spoken dialogue. Do NOT put any spoken dialogue in the global. TURNS / "relay" (time-masked -> MINIMAL, NO appearance): reference the speaker ONLY by a short position cue ("the one on the left"), then one physical cue + the quoted spoken line + the spoken language and accent. NEVER repeat hair/skin/build/clothing here. LANGUAGE & ACCENT (CRITICAL for non-English): If a line is a non-English language written in Latin/roman script (romanized Hindi / Hinglish, e.g. "Khushi hui jaan kar"), you MUST transliterate the spoken words into their NATIVE SCRIPT -- Devanagari for Hindi (e.g. "खुशी हुई जान कर") -- in BOTH the "relay" quoted line AND the "spoken" field, preserving the exact meaning. Roman-script Hindi makes the model speak with a foreign/English accent; native Devanagari makes it a natural native Hindi speaker. In that turn's "relay" add exactly: spoken in fluent native Hindi with a natural Indian Hindi accent. Keep genuinely English lines in English. Common English loanwords inside a Hindi sentence may stay in Latin inside the Devanagari sentence. OUTPUT: ONLY a strict JSON object - no markdown, no prose, no code fences: { "global": "<full cinematic scene + BOTH characters' appearance + positions + camera; NO dialogue>", "turns": [ { "speaker": "A", "relay": "<short position cue + one physical cue + the quoted line in NATIVE script + language/accent>", "spoken": "<the spoken words only, in NATIVE script (Devanagari for Hindi)>", "frames": <integer ~ words*8> } ] } RULES: one turn per dialogue line, in order. Preserve the exact MEANING of every line. Never use proper names. "spoken" = only the words said. Speaker ids exactly A/B/C/D. Output the JSON and nothing else.

2 other text fields— sigmas, filenames, and other non-prompt strings the logger swept up

node 52 · ManualSigmas · sigmas

1.0, 0.99375, 0.9875, 0.98125, 0.975, 0.909375, 0.725, 0.421875, 0.0

node 76 · ManualSigmas · sigmas

0.85, 0.7, 0.55, 0.4, 0.25, 0.12, 0.0

Input media

8

Source images and video this run consumed.

tara-cropped.jpg

node 29 · 76 KB

Indian_businessman_crop.jpg

node 33 · 50 KB

No preview

white_bg.jfif

node 49 · 4.4 KB

generated_voice_00004.mp3

node 64 · 301 KB

grok-image-dfccef6b-c58b-4c42-a36a-ec711a354d44.jpg

node 101 · 193 KB

hindi_male_2min.wav

node 126 · 4.2 MB

hindi_female_2min.wav

node 127 · 4.0 MB

hindi_female_5s.wav

node 136 · 467 KB