65bdb542-830d-4186-8bfd-e894321b0210
- Queued
- Jul 17, 2026, 09:44 AM
- Completed
- Jul 17, 2026, 09:45 AM
- Execution time
- 11.2s
- Outputs
- 1 · 1.4 MB
- Inputs
- 8
- Cached nodes
- 11
Failure
GeminiImageDirect: prompt is empty.
Stack trace
File "/home/ec2-user/ComfyUI/execution.py", line 542, in execute
output_data, output_ui, has_subgraph, has_pending_tasks = await get_output_data(prompt_id, unique_id, obj, input_data_all, execution_block_cb=execution_block_cb, pre_execute_cb=pre_execute_cb, v3_data=v3_data)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/home/ec2-user/ComfyUI/execution.py", line 341, in get_output_data
return_values = await _async_map_node_over_list(prompt_id, unique_id, obj, input_data_all, obj.FUNCTION, allow_interrupt=True, execution_block_cb=execution_block_cb, pre_execute_cb=pre_execute_cb, v3_data=v3_data)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/home/ec2-user/ComfyUI/execution.py", line 315, in _async_map_node_over_list
await process_inputs(input_dict, i)
File "/home/ec2-user/ComfyUI/execution.py", line 303, in process_inputs
result = f(**inputs)
^^^^^^^^^^^
File "/home/ec2-user/ComfyUI/custom_nodes/ComfyUI-GeminiDirect/__init__.py", line 138, in generate
raise RuntimeError("GeminiImageDirect: prompt is empty.")
Outputs
1Click an image to view it full size.
Text prompts
10The same interior facing the front door, now opening inward; ARJUN steps through in, a warm easy smile on his face, backlit by brighter daylight from outside spilling in around him. Photorealistic, 35mm, warm golden amber grade, shallow depth of field, fine film grain.
Two business colleagues in a modern glass office at golden hour, city skyline through tall windows, polished floor. On the left, Aryan, a man in his 30s in a navy business suit. On the right, Tara, a woman in her late 20s in a grey blazer. They stand facing each other near the window. Realistic, cinematic, soft warm light. Vertical 9:16 framing, medium two-shot.
SCENE: Modern glass office at golden hour, city skyline behind. On the left, a man in his 30s in a navy business suit. On the right, a woman in her late 20s in a grey blazer. Calm, professional mood. Static camera, slow push-in. A: Did you review the final numbers? B: I did. We're slightly ahead of target. A: That's great news. Let's tell the team. B: Agreed. I'll set up the call.
horizontal 16:9 framing, portrait crop, double faces, extra legs, extra limbs, multiple faces, overacting, excessive crying, Bollywood melodrama, smiling woman, soft romantic expression, man looking angry, cartoon, fantasy lighting, random crowd, heavy camera shake, poor lip sync, music overpowering dialogue, advertisement look, blurry, distorted, inconsistent appearance, worst quality, no captions, floating text, unnecessary artifacts
place a very subtle very thin semi transparent 6*6 white grid overlay perfectly algned over the face to break biometric landmark
Render a single photorealistic cinematic FIRST-FRAME still (not a collage) for a portrait 9:16 video. Use the provided reference images ONLY to match each character's identity and appearance. Filmic lighting and depth of field. Strict Note: Do not include any floating text, captions and watermarks anywhere in the image still.
You are a prompt engineer for the LTX-2.3 audio+video model. You are given the user's raw input AND reference image(s) of the characters. Rewrite the input into ONE production-grade video prompt. Output ONLY the final prompt text - no preamble, quotes, labels, headings, markdown, or explanation. RULES (follow exactly): 1. ONE flowing paragraph, natural English, present tense, ~120-200 words. No lists, no timestamps, no 'the scene opens with'. 2. Begin with 'Style:' then a brief style (e.g. 'Style: realistic, cinematic.'), then the scene. 3. STUDY THE REFERENCE IMAGES. For each character write a SHORT distinct visual phrase drawn from the reference (apparent age, build, skin tone, hair, facial hair) PLUS the clothing from the text, so each character is unmistakable in one phrase (e.g. 'a clean-shaven man with short black hair in a black suit', 'an older bearded man in a blue suit'). 4. NEVER use proper names - not in narration, not in dialogue tags. The model binds identity to APPEARANCE, not names. Map each named character to the matching reference image and replace the name with its appearance phrase EVERY time it occurs. 5. Attribute EVERY dialogue line explicitly by that appearance phrase, e.g. 'the man in the black suit says, in a respectful Hindi voice: "..."'. Keep all dialogue word-for-word and unchanged, including non-English/Devanagari text; state the spoken language and accent. 6. Use the SAME appearance phrase for a character every time so the model never merges, swaps, or duplicates the two people. Keep them spatially separate (e.g. one on the left, one on the right). 7. Describe ONLY what is seen and heard. NEVER state emotions, intentions or thoughts - show feeling as physical action (tight grip on the keys, a slow breath, glistening eyes, a small smile, a hand on a shoulder). No smell/taste/touch. 8. Restrained, plain language: plain colours, soft neutral lighting, no melodrama or hype words. 9. Camera: concrete moves only if implied (slow push-in, slow dolly, slow orbit, static) + 'natural motion blur'. 10. Audio: weave specific sounds beside the action (keys jingling, footsteps, a car door); give exact quoted dialogue + speaker appearance + language/accent; never invent speech unless the user mentions talking. 11. Enforce vertical 9:16 portrait framing; remove any horizontal/landscape/16:9 wording. 12. Keep ONE coherent continuous action; if many beats are listed, keep them brief and in order, but never let the two characters' appearance phrases drift.
You format multi-character dialogue for the LTX-2.3 audio+video model plus a Prompt Relay node. INPUT: a SCENE line plus dialogue lines each labelled by speaker (A:, B:, C:, D:). OUTPUT: ONLY a strict JSON object - no markdown, no prose, no code fences. Schema: { "global": "<one scene + identity description for the whole shot; describe each character by APPEARANCE not name, keep them spatially separate>", "turns": [ { "speaker": "A", "relay": "<one segment: the speaking character by appearance phrase + one physical cue + the EXACT quoted line + spoken language/accent + a camera hint>", "spoken": "<the exact spoken words only, verbatim, including any Hindi/Devanagari>", "frames": <integer estimate of this line's length in video frames at 24fps, about words*8> } ] } RULES: one turn per dialogue line, in order. Preserve every spoken word EXACTLY (including non-English). Never use proper names anywhere; map each name to its appearance phrase and reuse it every time. Attribute each line to its speaker by appearance. The "spoken" field must contain ONLY the words actually said (no narration). Use speaker ids exactly A/B/C/D matching the labels. Output the JSON and nothing else.
2 other text fields— sigmas, filenames, and other non-prompt strings the logger swept up
node 17 · ManualSigmas · sigmas
1.0, 0.99375, 0.9875, 0.98125, 0.975, 0.909375, 0.725, 0.421875, 0.0
node 33 · ManualSigmas · sigmas
0.85, 0.7, 0.55, 0.4, 0.25, 0.12, 0.0
Input media
8Source images and video this run consumed.
detective-vance-smoking.jpeg
node 4 · 2.0 MB
LTX_firstframe_00032_.png
node 60 · 1.4 MB
generated_voice_00004.mp3
node 61 · 301 KB
glassy_apratment.webp
node 66 · 39 KB
Indian_businessman_crop.jpg
node 68 · 50 KB
tara-cropped.jpg
node 77 · 76 KB
man_v_sample.mp3
node 85 · 179 KB
female_v_sample.mp3
node 86 · 224 KB