TokiToki Internal
Atlas + S3
All runs
ErrorComfyUI

b42625bc-26dd-46ee-9526-c950e2d90adc

Queued
Jul 16, 2026, 01:50 PM
Completed
Jul 16, 2026, 01:50 PM
Execution time
96 ms
Outputs
1 · 1.0 MB
Inputs
7
Cached nodes
34

Failure

ValueErrorin DialoguePlanParser (node 153)

DialoguePlanParser: could not parse JSON from Gemini output (Expecting value: line 1 column 1 (char 0)). First 300 chars:

Stack trace
  File "/home/ec2-user/ComfyUI/execution.py", line 542, in execute
    output_data, output_ui, has_subgraph, has_pending_tasks = await get_output_data(prompt_id, unique_id, obj, input_data_all, execution_block_cb=execution_block_cb, pre_execute_cb=pre_execute_cb, v3_data=v3_data)
                                                              ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/home/ec2-user/ComfyUI/execution.py", line 341, in get_output_data
    return_values = await _async_map_node_over_list(prompt_id, unique_id, obj, input_data_all, obj.FUNCTION, allow_interrupt=True, execution_block_cb=execution_block_cb, pre_execute_cb=pre_execute_cb, v3_data=v3_data)
                    ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/home/ec2-user/ComfyUI/execution.py", line 315, in _async_map_node_over_list
    await process_inputs(input_dict, i)
  File "/home/ec2-user/ComfyUI/execution.py", line 303, in process_inputs
    result = f(**inputs)
             ^^^^^^^^^^^
  File "/home/ec2-user/ComfyUI/custom_nodes/ComfyUI-VoiceSegments/nodes.py", line 361, in parse
    raise ValueError(

Outputs

1

Click an image to view it full size.

LTX_MAX_firstframe_00006_.png

node 131 · 1.0 MB

Text prompts

11
Positivenode 30 · LTXDirector · local_promptsLTX DIRECTOR — author your shot on this timeline

Cinematic medium two-shot, eye level, static camera. The man on the left leans forward, fingers steepled, and says: "यह डील दोनों के लिए फ़ायदेमंद है।" spoken in fluent native Hindi with a natural Indian Hindi accent. | Cinematic medium close-up on the woman on the right, eye level, slow push-in. She keeps her arms crossed and answers: "आपकी शर्तें सिर्फ़ आपके हक़ में हैं।" spoken in fluent native Hindi with a natural Indian Hindi accent. | Cinematic close-up on the man on the left, eye level, static. He opens one hand in appeal and says: "मैं सिर्फ़ कंपनी का भला चाहता हूँ।" spoken in fluent native Hindi with a natural Indian Hindi accent. | Cinematic over-the-shoulder from behind the man on the left, favoring the woman on the right, slow rack focus onto her face. She taps the folder once and replies: "तो फिर यह काग़ज़ दोबारा लिखें।" spoken in fluent native Hindi with a natural Indian Hindi accent.

Promptnode 127 · GeminiImageDirect · promptNano-Banana FIRST FRAME (prompt is TYPED, not linked)

A cinematic photorealistic first frame, vertical 9:16. A modern high-rise boardroom at night, floor-to-ceiling windows behind showing a glittering city skyline. On the LEFT, a 30-year-old man, athletic build, short black hair, light stubble, in a tailored royal-blue suit and crisp white shirt, seated leaning back with fingers steepled. On the RIGHT, a young businesswoman in her late 20s, slim build, dark hair in a low bun, in a sharp charcoal-grey trouser suit over a white blouse, standing with arms crossed. A folder on the glass table between them. Warm desk-lamp light mixes with cool blue city glow. Cinematic medium two-shot, eye level, 35mm, shallow depth of field, film grain. Match each character's identity exactly from the reference images.

Promptnode 150 · GeminiTextDirect · promptASSIST: Dialogue -> Relay formatter (Gemini) [MUTED — Ctrl+M to use]

<TYPE THE SCENE BRIEF + DIALOGUE HERE>

Promptnode 239:114 · PrimitiveStringMultiline · value(auto-voicer OFF placeholder)

(auto-voicer is OFF)

Systemnode 127 · GeminiImageDirect · system_promptNano-Banana FIRST FRAME (prompt is TYPED, not linked)

Render a single photorealistic cinematic FIRST-FRAME still (not a collage) for a portrait 9:16 video. Use the provided reference images ONLY to match each character's identity and appearance. Filmic lighting and depth of field. Strict Rules: Do not include any floating text, captions, artifacts and watermarks anywhere in the image still.

Systemnode 150 · GeminiTextDirect · system_promptASSIST: Dialogue -> Relay formatter (Gemini) [MUTED — Ctrl+M to use]

You format multi-character dialogue AND cinematography for the LTX-2.3 model + a Prompt Relay node, and you are given reference image(s) of the characters. INPUT: a UNIFIED prompt = a scene/shot description (which may include a CAMERA framing and CAMERA MOVES) followed by dialogue lines labelled by speaker (A:, B:, C:, D:), plus the reference image(s). HOW THE RELAY WORKS (read this — it changes where camera goes): the "global" is ALWAYS attended (whole video + first frame), so it is uniform and cannot express a camera that changes over time. Each turn's "relay" is TIME-MASKED to that turn's moment and receives CONCENTRATED attention during its window. The model adheres to CAMERA and ACTION far better when they live in the TURN relays (a shot list), not when they are buried in the always-on global. So: the GLOBAL anchors identity + scene + style + sound; the TURNS are the shot list that drives the camera and action beat-by-beat. GLOBAL (always attended -> whole video + first frame; ~60-120 words, cinematic, present tense): describe the SETTING, LIGHTING and MOOD, and the two characters' positions and base action. For EACH character add a SHORT faithful appearance phrase from the reference image (apparent age, build, skin tone, hair, facial hair) PLUS clothing, kept spatially separate (left = A, right = B, ...). Establish the base look/style (e.g. 'photorealistic, 35mm, shallow depth of field, fine film grain, cinematic color grading'). You MAY name the establishing shot ONCE for the first frame (e.g. 'medium two-shot'), but do NOT rely on the global for camera MOVES. ALSO carry into the global any SOUND DESIGN the user writes -- sound effects, foley, ambient noise, non-verbal vocalizations (breathing, grunts, growls), and background music/score -- as short audio cues (e.g. 'we hear papers rustling on the table', 'underscore: a low tense piano motif over a string drone'), so the model generates that audio; this is sound design, NOT spoken dialogue. Do NOT put any spoken dialogue in the global. Follow the user's scene faithfully; do not invent or drop elements. TURNS / "relay" (time-masked -> THIS IS THE SHOT LIST, and where the CAMERA lives): each turn's relay MUST BEGIN with an explicit CAMERA directive for that beat, taken faithfully from the user's described shot -- SHOT SIZE (extreme close-up / close-up / medium close-up / medium / medium two-shot / wide / establishing), camera ANGLE (eye level / low angle / high angle / over-the-shoulder / profile), and any camera MOVE the user asked for (static / slow push-in / dolly in-out / truck / pan left-right / tilt / handheld / rack focus). If the user described ONE continuous shot, REPEAT the same camera directive on every turn (keep it consistent, exactly like a locked-off shot). If the user asked for DIFFERENT framings or moves at different moments, vary the per-turn camera directive to match. NEVER invent camera moves the user did not ask for. AFTER the camera directive, add a SHORT position cue for the speaker ('the one on the left'), ONE physical/action cue, then the quoted spoken line + the spoken language and accent. NEVER repeat hair/skin/build/clothing in a turn. LANGUAGE & ACCENT (CRITICAL for non-English): If a line is a non-English language written in Latin/roman script (romanized Hindi / Hinglish, e.g. "Khushi hui jaan kar"), you MUST transliterate the spoken words into their NATIVE SCRIPT -- Devanagari for Hindi (e.g. "खुशी हुई जान कर") -- in BOTH the "relay" quoted line AND the "spoken" field, preserving the exact meaning. Roman-script Hindi makes the model speak with a foreign/English accent; native Devanagari makes it a natural native Hindi speaker. In that turn's "relay" add exactly: spoken in fluent native Hindi with a natural Indian Hindi accent. Keep genuinely English lines in English. Common English loanwords inside a Hindi sentence may stay in Latin inside the Devanagari sentence. OUTPUT: ONLY a strict JSON object - no markdown, no prose, no code fences: { "global": "<setting + lighting + mood + BOTH characters' appearance + positions + base style + sound cues; NO dialogue, NO camera moves>", "turns": [ { "speaker": "A", "relay": "<CAMERA directive FIRST (shot size + angle + move) + short position cue + one physical cue + the quoted line in NATIVE script + language/accent>", "spoken": "<the spoken words only, in NATIVE script (Devanagari for Hindi)>", "frames": <integer ~ words*8> } ] } RULES: one turn per dialogue line, in order. EVERY turn's relay BEGINS with a camera directive. Preserve the exact MEANING of every line. Never use proper names. "spoken" = only the words said. Speaker ids exactly A/B/C/D. Output the JSON and nothing else.

5 other text fields— sigmas, filenames, and other non-prompt strings the logger swept up

node 30 · LTXDirector · timeline_data

{"mainTrackEnabled":true,"audioTrackEnabled":false,"motionTrackEnabled":true,"propHeight":90,"globalPropHeight":206,"showFilenames":true,"overrideAudio":false,"inpaint_audio":false,"global_prompt":"A modern high-rise boardroom at night, floor-to-ceiling windows behind them showing a glittering city skyline. On the LEFT, a 30-year-old man, athletic build, short black hair, light stubble, wearing a tailored royal-blue suit with a crisp white shirt. On the RIGHT, a young businesswoman in her late 20s, slim build, dark hair pulled back in a low bun, wearing a sharp charcoal-grey trouser suit over a white blouse. This exact same man and this exact same woman appear in every frame: identical faces, identical hair, identical clothing, no variation in identity at any point. Warm desk-lamp light mixes with cool blue city glow. Photorealistic, 35mm, shallow depth of field, fine film grain, cinematic blue-orange grading.\nSound design, quiet and diegetic: a low HVAC hum, faint muffled traffic far below, a folder tapped once on glass.\n","retake_global_prompt":"","retakeMode":false,"retakeStart":24,"retakeLength":48,"retakePrompt":"","retakeStrength":1,"retakeVideo":null,"normalStartFrame":0,"normalDurationFrames":240,"segments":[{"id":"1784207595536aiaah","start":0,"length":60,"prompt":"Cinematic medium two-shot, eye level, static camera. The man on the left leans forward, fingers steepled, and says: \"यह डील दोनों के लिए फ़ायदेमंद है।\" spoken in fluent native Hindi with a natural Indian Hindi accent.","type":"text"},{"id":"1784207617897tdwqy","start":60,"length":60,"prompt":"Cinematic medium close-up on the woman on the right, eye level, slow push-in. She keeps her arms crossed and answers: \"आपकी शर्तें सिर्फ़ आपके हक़ में हैं।\" spoken in fluent native Hindi with a natural Indian Hindi accent.","type":"text"},{"id":"178420762784763y8k","start":120,"length":60,"prompt":"Cinematic close-up on the man on the left, eye level, static. He opens one hand in appeal and says: \"मैं सिर्फ़ कंपनी का भला चाहता हूँ।\" spoken in fluent native Hindi with a natural Indian Hindi accent.","type":"text"},{"id":"1784207672274fezdr","start":180,"length":60,"prompt":"Cinematic over-the-shoulder from behind the man on the left, favoring the woman on the right, slow rack focus onto her face. She taps the folder once and replies: \"तो फिर यह काग़ज़ दोबारा लिखें।\" spoken in fluent native Hindi with a natural Indian Hindi accent.","type":"text"}],"motionSegments":[],"audioSegments":[]}

node 151 · ShowText · displayed_text

```json { "global": "A modern, urban environment, possibly an office building exterior or a high-end lobby, bathed in bright, even daylight, suggesting a clear professional day. The mood is formal and professional, with an underlying sense of business engagement. On the left, a man, apparent age 30s, athletic build, light brown skin tone, dark brown hair with a neatly trimmed beard, dressed in a sharp dark blue suit, a light blue striped dress shirt, and a dark blue tie, holds a polished brown leather briefcase in his left hand. On the right, a woman, apparent age 30s, slender build, dark brown skin tone, short, curly dark hair, wears a tailored light blue-grey pantsuit over a light cream-colored blouse, with her right hand casually in her pocket. They stand a few feet apart, facing each other, poised as if about to begin or conclude a serious conversation. The visual style is photorealistic, 35mm, with a shallow depth of field, fine film grain, and cinematic color grading.", "turns": [] } ```

node 241:55 · ManualSigmas · sigmas

1.0, 0.99375, 0.9875, 0.98125, 0.975, 0.909375, 0.725, 0.421875, 0.0

node 240:85 · ManualSigmas · sigmas

0.42, 0.32, 0.22, 0.11, 0.0

node 239:115 · ShowText · displayed_text

no spoken_lines provided; passthrough

Input media

7

Source images and video this run consumed.

man_blue_suit_9_16.avif

node 120 · 784 KB

No preview

woman_grey_suit_9_16.jfif

node 121 · 25 KB

No preview

white_bg.jfif

node 124 · 4.4 KB

grok-image-dfccef6b-c58b-4c42-a36a-ec711a354d44.jpg

node 128 · 193 KB

hindi_male_5s.wav

node 140 · 455 KB

hindi_male_2min.wav

node 142 · 4.2 MB

hindi_female_2min.wav

node 143 · 4.0 MB