通天閣にクマ出没ニュース

台本

通天閣にクマ出没ニュース

制作の流れ

0台本Claude Code
→
1セリフ音声ElevenLabs v4
→
2各カットの画像GPT Image 2.5
→
3動画化Seedance 2.5
Higgsfield MCP
→
4テロップ・結合ffmpeg

ポイント:先に音声を作り、その音声を Seedance に渡して口の動きと間を合わせる

台本

通天閣にクマ出没ニュース

  1. 1アナウンサー

    「ニュースです。きょう午後2時ごろ、大阪市なにわ区の通天閣の周辺で、クマが目撃されました。」

  2. 2アナウンサー

    「警察によりますと、クマは近くの天王寺動物園から逃げ出したとみられています。」

  3. 3撮影者

    「え、ちょ、クマやん!ほんまにクマやって!黄色いで!」

  4. 4関西のおばちゃん

    「びっくりしたわぁ。黄色いから、ぬいぐるみかと思たら、ほんまもんやねん。」

  5. 5若い男性

    「串カツ食べてたら、横をふつうに歩いてて…二度見しました。」

  6. 6リポーター

    「現在も、警察と動物園の職員が、クマの行方を捜しています。付近の方は、外出を控えてください。」

  7. 7テロップ

    ※この映像は、AIで生成したフィクションです。実在の事件・放送局・人物とは関係ありません。

台本

各シーンの音声

ElevenLabs v4 でセリフを1本ずつ生成

  1. 1

    スタジオアナウンサー

    「ニュースです。きょう午後2時ごろ、大阪市なにわ区の通天閣の周辺で、クマが目撃されました。」

    🔊 ElevenLabs v4:Akari(落ち着いた女性)
  2. 2

    現場の引き映像アナウンサー

    「警察によりますと、クマは近くの天王寺動物園から逃げ出したとみられています。」

    🔊 ElevenLabs v4:Akari(落ち着いた女性)
  3. 3

    視聴者提供のスマホ映像撮影者

    「え、ちょ、クマやん!ほんまにクマやって!黄色いで!」

    🔊 ElevenLabs v4:Haruta(本人の声クローン)
  4. 4

    街頭インタビュー①関西のおばちゃん

    「びっくりしたわぁ。黄色いから、ぬいぐるみかと思たら、ほんまもんやねん。」

    🔊 ElevenLabs v4:Aiko(関西弁・女性)
  5. 5

    街頭インタビュー②若い男性

    「串カツ食べてたら、横をふつうに歩いてて…二度見しました。」

    🔊 ElevenLabs v4:Riku(関西弁・男性)
  6. 6

    現場リポーター中継リポーター

    「現在も、警察と動物園の職員が、クマの行方を捜しています。付近の方は、外出を控えてください。」

    🔊 ElevenLabs v4:Masafumi(きびきびした男性)

台本

各カットの画像

GPT Image 2.5 で作った、動画化する前の1枚目

カット1
1スタジオニュースデスクの女性アナウンサー
📝 画像生成プロンプト(GPT Image 2.5)
Frame grab from a Japanese evening television news broadcast. A calm Japanese female news anchor in her early 40s, short neat dark bob hair, navy blazer over a white blouse, sits behind a glossy news desk with a few paper script pages in front of her, medium shot from chest up, centered, looking straight into the camera, lips slightly parted mid-sentence, composed professional expression. Background: a modern generic news studio set with soft deep-green and warm amber panels and a large out-of-focus screen showing an abstract city at dusk. Even, flat broadcast studio lighting, realistic skin texture, shot on a studio broadcast camera, slightly compressed broadcast video quality, natural moderate depth of field. Absolutely no text, no captions, no logos, no channel bugs, no watermarks anywhere in the image.
カット2
2現場の引き映像新世界の通り・通天閣・群衆の向こうを歩くクマ
📝 画像生成プロンプト(GPT Image 2.5)
Frame grab from a live Japanese TV news helicopter-free ground camera, long telephoto lens from an elevated position. The Shinsekai shopping street in Osaka in the afternoon: a straight street lined with old-fashioned colorful restaurant facades, paper lanterns and oversized decorative signboards, and at the far end of the street the steel lattice tower Tsutenkaku rising against a hazy sky. In the foreground, backs of a crowd of onlookers holding up smartphones, two police officers in dark uniforms spreading their arms to hold people back. In the middle distance, in the center of the street, the same realistic golden honey-yellow bear from the reference image walks slowly across the pavement on all fours. All signboards are fictional and illegible from the distance, no readable words. Telephoto compression, slight heat haze, natural overcast daylight, broadcast ENG camera look, slightly compressed broadcast video quality, realistic, not cinematic. No text, no logos, no captions, no watermarks.
カット3
3視聴者提供のスマホ映像串カツ屋の路地を横切るクマ(縦動画)
📝 画像生成プロンプト(GPT Image 2.5)
Vertical amateur smartphone video frame filmed by a bystander, handheld and slightly tilted, from about 10 meters away at eye level. A narrow old Osaka Shinsekai alley lined with small kushikatsu restaurants: red paper lanterns, noren curtains, plastic stools and a beer crate outside, menu boards with illegible scribbles. The same realistic golden honey-yellow bear from the reference image lumbers across the alley from left to right on all fours, mid-stride, while two pedestrians in the background step back startled. Phone camera look: slightly overexposed sky, mild motion blur, digital noise, auto-exposure, ordinary afternoon light, not cinematic. Signs must be illegible or blurred, no readable text, no logos, no UI overlay, no watermarks.
カット4
4街頭インタビュー①新世界で取材を受けるおばちゃん
📝 画像生成プロンプト(GPT Image 2.5)
Frame grab from a Japanese TV news street interview (man-on-the-street). A friendly Osaka woman in her early 60s with short permed reddish-brown hair, wearing a leopard-print blouse and a light cardigan, holding a shopping bag, stands on a busy old-fashioned downtown Osaka shopping street, medium close-up from chest up, positioned slightly right of center, looking just off-camera to the left toward an interviewer, mid-speech with an amused surprised expression, hand near her chest. The tip of a plain grey foam-covered handheld microphone without any logo enters from the lower left of the frame. Background: blurred colorful restaurant facades and lanterns, passersby, signs illegible. Natural afternoon daylight, broadcast ENG camera look, realistic skin texture, slightly compressed broadcast video quality, not cinematic. No text, no logos, no captions, no watermarks.
カット5
5街頭インタビュー②串カツを持った若い男性
📝 画像生成プロンプト(GPT Image 2.5)
Frame grab from a Japanese TV news street interview (man-on-the-street). A Japanese man in his early 20s with short black hair, wearing a grey hoodie under a dark jacket and holding a wooden skewer, stands on an old-fashioned downtown Osaka shopping street lined with small kushikatsu restaurants, medium close-up from chest up, positioned slightly left of center, looking just off-camera to the right toward an interviewer, mid-speech with a slightly stunned half-smile. The tip of a plain grey foam-covered handheld microphone without any logo enters from the lower right of the frame. Background: blurred red lanterns, restaurant facades, passersby, signs illegible. Natural afternoon daylight, broadcast ENG camera look, realistic skin texture, slightly compressed broadcast video quality, not cinematic. No text, no logos, no captions, no watermarks.
カット6
6現場リポーター中継規制線・パトカー・通天閣を背にしたリポーター
📝 画像生成プロンプト(GPT Image 2.5)
Frame grab from a live Japanese TV news field report. A Japanese male field reporter in his mid 30s, short neat hair, dark navy suit with a plain dark windbreaker over it, holding a plain black handheld microphone without any logo, stands facing the camera, medium shot from waist up, slightly left of center, serious focused expression, mid-sentence. Behind him: a plain yellow barrier tape stretched across the street with no readable text, a Japanese police patrol car with its red rotating roof light glowing, a few uniformed police officers and zoo staff in green work uniforms, and in the background down the street the steel lattice Tsutenkaku tower in Osaka's Shinsekai district. Late afternoon overcast daylight, broadcast ENG camera look, realistic, slightly compressed broadcast video quality, not cinematic. No text, no logos, no captions, no watermarks.

台本

各カットの動画

Seedance 2.5(Higgsfield MCP)で、画像と音声から動画化

1スタジオ画像+音声 → 口パクを音声に合わせる
📝 動画生成プロンプト(Seedance 2.5)
Locked-off studio camera, Japanese evening TV news broadcast. The female news anchor at the desk reads the news straight to camera, speaking exactly the Japanese line from the audio reference with precise natural lip sync, calm professional delivery, small natural head movements and blinks, glances down at her script once briefly, hands resting on the papers. Background city screen subtly animated. No camera movement, no zoom. Quiet studio room tone only, no music. Realistic broadcast video, no text, no captions, no logos.
2現場の引き映像画像+音声 → クマがのそのそ歩く
📝 動画生成プロンプト(Seedance 2.5)
Live news camera on a tripod with long telephoto lens, slight handheld-free micro shake and a slow small pan to follow the bear. The realistic golden honey-yellow bear walks slowly and heavily across the Shinsekai street on all fours with natural bear gait, shoulder muscles rolling, briefly sniffing the ground. Police officers keep their arms spread holding back the crowd; onlookers in the foreground raise smartphones and shift nervously. Tsutenkaku tower stays fixed in the background. The audio reference is an off-screen female news anchor voice-over; nobody on screen speaks or lip-syncs. Ambient sound: distant crowd murmur, faint police radio, street noise. Realistic broadcast video, no text, no captions, no logos.
3視聴者提供のスマホ映像画像+音声 → 手ブレのスマホ撮影
📝 動画生成プロンプト(Seedance 2.5)
Vertical amateur smartphone footage, shaky handheld, the bystander filming steps back a little and jerkily re-frames to keep the bear in view, brief autofocus hunting. The realistic golden honey-yellow bear lumbers across the narrow kushikatsu alley from left to right on all fours with a natural heavy bear gait, then turns its head toward the camera for a moment. The two pedestrians in the background step back startled. The audio reference is the excited voice of the person holding the phone, off-camera; nobody on screen speaks. Ambient sound: alley chatter, a startled shout in the distance, phone-mic wind noise. Realistic phone video, no text, no UI, no logos.
4街頭インタビュー①画像+音声 → 口パクを音声に合わせる
📝 動画生成プロンプト(Seedance 2.5)
Handheld TV news street interview, gentle natural camera sway. The Osaka woman in the leopard-print blouse speaks exactly the Kansai-dialect Japanese line from the audio reference to the interviewer just off-camera, with precise natural lip sync, animated expressive face, a surprised look then a laugh, hand patting her chest and a small wave of her hand. The interviewer's grey microphone stays in the lower left of frame. Passersby walk in the blurred background. Ambient sound: busy shopping street murmur. Realistic broadcast video, no text, no captions, no logos.
5街頭インタビュー②画像+音声 → 口パクを音声に合わせる
📝 動画生成プロンプト(Seedance 2.5)
Handheld TV news street interview, gentle natural camera sway. The young man in the grey hoodie speaks exactly the Japanese line from the audio reference to the interviewer just off-camera, with precise natural lip sync, slightly stunned then a short embarrassed laugh, he gestures with the kushikatsu skewer to his side as if showing where the bear walked past. The interviewer's grey microphone stays in the lower right of frame. Passersby walk in the blurred background. Ambient sound: busy restaurant street murmur. Realistic broadcast video, no text, no captions, no logos.
6現場リポーター中継画像+音声 → 口パクを音声に合わせる
📝 動画生成プロンプト(Seedance 2.5)
Live TV news field report, handheld camera on the cameraman's shoulder with slight natural sway. The male reporter holding the microphone speaks exactly the Japanese line from the audio reference straight to camera with precise natural lip sync, serious urgent but composed delivery, a short glance back over his shoulder toward the police line and then back to camera. Behind him the police car's red rotating light flashes, officers and zoo staff in green uniforms move around behind the yellow tape, Tsutenkaku tower in the background. Ambient sound: police radio chatter, distant siren, street noise. Realistic broadcast video, no text, no captions, no logos.

台本

ffmpeg で仕上げ

カットごとに仕上げてから、最後に7本をつなぐ

①カットごとに仕上げる × 6本

  1. 切り出すSeedanceの動画から、使う範囲だけを切り取る
  2. テロップを重ねる局ロゴ・LIVE・時刻・ニュースの見出しを、透明な画像として上に重ねる
  3. 音声を差し替えるElevenLabsで作った元の音声に入れ替え、口の動きに合わせて位置を調整
生成したままの動画
生成したままの動画(文字なし)
+
テロップ画像
テロップ画像(背景は透明)
=
テロップ入り
テロップ入りのカット

カット3(スマホ映像)だけは、縦長の映像を画面の真ん中に置いて、左右をぼかした背景で埋める

②最後の注記カードを作る 2秒

「※この映像はAIで生成したフィクションです」

③7本を順番につなぐ

カット1完成
1
+
カット2完成
2
+
カット3完成
3
+
カット4完成
4
+
カット5完成
5
+
カット6完成
6
+
カット7完成
7

カットの間はハードカット(ニュースっぽく)。最後に全体の音量をそろえて書き出す

→ 完成:約43秒・1080p

bear-news-hook_v1_clean.mp4 / 43.15秒 / 💎535.25 / ElevenLabs 367文字

  • 局名:かもめテレビ
  • 動物園:天王寺動物園
  • クマ:リアルなヒグマ体型・淡い黄金色・服なし
  • 音声:ElevenLabsの声だけ(声のダブりを除去したクリーン版。ダブりありの元ファイルは bear-news-hook_v1.mp4)
クマ基準画像 work/v1_backup/bear_ref.jpg 💎2.75
プロンプト
Documentary wildlife reference photograph of a single real adult bear standing on all fours on an asphalt city street in daylight, full body side three-quarter view. The bear is a completely realistic animal with the anatomy of a brown bear (Ursus arctos): heavy shoulder hump, long snout, small rounded ears, dark nose, natural animal eyes, long claws. Its fur is an unusual soft pale golden honey-yellow color, thick and fluffy, slightly lighter on the face and darker golden on the back, natural fur texture with dust and clumps, realistic wildlife look. No clothing, no shirt, no accessories, no cartoon features, not a teddy bear, not anthropomorphic. Shot on a broadcast news ENG video camera, natural overcast afternoon light, neutral color, slightly compressed broadcast image quality, moderate depth of field. No text, no logos, no watermarks.
1

カット1:スタジオ/アナウンサー

images/c1.jpg(GPT Image 2.5 sunburst)💎2.75
clips/c1.mp4(Seedance 2.5 / 9s)💎108

セリフ(Akari(女性・落ち着き) / eleven_v4 / 8.08s)
[calm, professional news anchor] ニュースです。きょう午後2時ごろ、大阪市なにわ区の、つうてんかくの周辺で、クマが目撃されました。

画像プロンプト
Frame grab from a Japanese evening television news broadcast. A calm Japanese female news anchor in her early 40s, short neat dark bob hair, navy blazer over a white blouse, sits behind a glossy news desk with a few paper script pages in front of her, medium shot from chest up, centered, looking straight into the camera, lips slightly parted mid-sentence, composed professional expression. Background: a modern generic news studio set with soft deep-green and warm amber panels and a large out-of-focus screen showing an abstract city at dusk. Even, flat broadcast studio lighting, realistic skin texture, shot on a studio broadcast camera, slightly compressed broadcast video quality, natural moderate depth of field. Absolutely no text, no captions, no logos, no channel bugs, no watermarks anywhere in the image.
動画プロンプト
Locked-off studio camera, Japanese evening TV news broadcast. The female news anchor at the desk reads the news straight to camera, speaking exactly the Japanese line from the audio reference with precise natural lip sync, calm professional delivery, small natural head movements and blinks, glances down at her script once briefly, hands resting on the papers. Background city screen subtly animated. No camera movement, no zoom. Quiet studio room tone only, no music. Realistic broadcast video, no text, no captions, no logos.
2

カット2:中継カメラ風の引き(VO)

work/v1_backup/c2.jpg(GPT Image 2.5 sunburst)💎2.75
work/v1_backup/c2.mp4(Seedance 2.5 / 7s)💎84

セリフ(Akari / eleven_v4 / 6.00s)
[calm, professional news anchor] 警察によりますと、クマは近くの、てんのうじ動物園から逃げ出したとみられています。

画像プロンプト
Frame grab from a live Japanese TV news helicopter-free ground camera, long telephoto lens from an elevated position. The Shinsekai shopping street in Osaka in the afternoon: a straight street lined with old-fashioned colorful restaurant facades, paper lanterns and oversized decorative signboards, and at the far end of the street the steel lattice tower Tsutenkaku rising against a hazy sky. In the foreground, backs of a crowd of onlookers holding up smartphones, two police officers in dark uniforms spreading their arms to hold people back. In the middle distance, in the center of the street, the same realistic golden honey-yellow bear from the reference image walks slowly across the pavement on all fours. All signboards are fictional and illegible from the distance, no readable words. Telephoto compression, slight heat haze, natural overcast daylight, broadcast ENG camera look, slightly compressed broadcast video quality, realistic, not cinematic. No text, no logos, no captions, no watermarks.
動画プロンプト
Live news camera on a tripod with long telephoto lens, slight handheld-free micro shake and a slow small pan to follow the bear. The realistic golden honey-yellow bear walks slowly and heavily across the Shinsekai street on all fours with natural bear gait, shoulder muscles rolling, briefly sniffing the ground. Police officers keep their arms spread holding back the crowd; onlookers in the foreground raise smartphones and shift nervously. Tsutenkaku tower stays fixed in the background. The audio reference is an off-screen female news anchor voice-over; nobody on screen speaks or lip-syncs. Ambient sound: distant crowd murmur, faint police radio, street noise. Realistic broadcast video, no text, no captions, no logos.
3

カット3:視聴者提供スマホ映像

work/v1_backup/c3.jpg(GPT Image 2.5 sunburst)💎2.75
work/v1_backup/c3.mp4(Seedance 2.5 / 5s)💎60

セリフ(Haruta PVC (JA) / eleven_v4 / 3.84s)
[surprised] え、ちょ、クマやん! [excited] ほんまにクマやって!黄色いで!

画像プロンプト
Vertical amateur smartphone video frame filmed by a bystander, handheld and slightly tilted, from about 10 meters away at eye level. A narrow old Osaka Shinsekai alley lined with small kushikatsu restaurants: red paper lanterns, noren curtains, plastic stools and a beer crate outside, menu boards with illegible scribbles. The same realistic golden honey-yellow bear from the reference image lumbers across the alley from left to right on all fours, mid-stride, while two pedestrians in the background step back startled. Phone camera look: slightly overexposed sky, mild motion blur, digital noise, auto-exposure, ordinary afternoon light, not cinematic. Signs must be illegible or blurred, no readable text, no logos, no UI overlay, no watermarks.
動画プロンプト
Vertical amateur smartphone footage, shaky handheld, the bystander filming steps back a little and jerkily re-frames to keep the bear in view, brief autofocus hunting. The realistic golden honey-yellow bear lumbers across the narrow kushikatsu alley from left to right on all fours with a natural heavy bear gait, then turns its head toward the camera for a moment. The two pedestrians in the background step back startled. The audio reference is the excited voice of the person holding the phone, off-camera; nobody on screen speaks. Ambient sound: alley chatter, a startled shout in the distance, phone-mic wind noise. Realistic phone video, no text, no UI, no logos.
4

カット4:街頭インタビュー1(おばちゃん)

images/c4.jpg(GPT Image 2.5 sunburst)💎2.75
clips/c4.mp4(Seedance 2.5 / 7s)💎84

セリフ(Aiko(関西・中年女性) / eleven_v4 / 6.08s)
[surprised] びっくりしたわぁ。黄色いから、ぬいぐるみかと思たら、[laughs] ほんまもんやねん。

画像プロンプト
Frame grab from a Japanese TV news street interview (man-on-the-street). A friendly Osaka woman in her early 60s with short permed reddish-brown hair, wearing a leopard-print blouse and a light cardigan, holding a shopping bag, stands on a busy old-fashioned downtown Osaka shopping street, medium close-up from chest up, positioned slightly right of center, looking just off-camera to the left toward an interviewer, mid-speech with an amused surprised expression, hand near her chest. The tip of a plain grey foam-covered handheld microphone without any logo enters from the lower left of the frame. Background: blurred colorful restaurant facades and lanterns, passersby, signs illegible. Natural afternoon daylight, broadcast ENG camera look, realistic skin texture, slightly compressed broadcast video quality, not cinematic. No text, no logos, no captions, no watermarks.
動画プロンプト
Handheld TV news street interview, gentle natural camera sway. The Osaka woman in the leopard-print blouse speaks exactly the Kansai-dialect Japanese line from the audio reference to the interviewer just off-camera, with precise natural lip sync, animated expressive face, a surprised look then a laugh, hand patting her chest and a small wave of her hand. The interviewer's grey microphone stays in the lower left of frame. Passersby walk in the blurred background. Ambient sound: busy shopping street murmur. Realistic broadcast video, no text, no captions, no logos.
5

カット5:街頭インタビュー2(若い男性)

images/c5.jpg(GPT Image 2.5 sunburst)💎2.75
clips/c5.mp4(Seedance 2.5 / 6s)💎72

セリフ(Riku(関西・若い男性) / eleven_v4 / 5.28s)
串カツ食べてたら、横をふつうに歩いてて… [laughs] 二度見しました。

画像プロンプト
Frame grab from a Japanese TV news street interview (man-on-the-street). A Japanese man in his early 20s with short black hair, wearing a grey hoodie under a dark jacket and holding a wooden skewer, stands on an old-fashioned downtown Osaka shopping street lined with small kushikatsu restaurants, medium close-up from chest up, positioned slightly left of center, looking just off-camera to the right toward an interviewer, mid-speech with a slightly stunned half-smile. The tip of a plain grey foam-covered handheld microphone without any logo enters from the lower right of the frame. Background: blurred red lanterns, restaurant facades, passersby, signs illegible. Natural afternoon daylight, broadcast ENG camera look, realistic skin texture, slightly compressed broadcast video quality, not cinematic. No text, no logos, no captions, no watermarks.
動画プロンプト
Handheld TV news street interview, gentle natural camera sway. The young man in the grey hoodie speaks exactly the Japanese line from the audio reference to the interviewer just off-camera, with precise natural lip sync, slightly stunned then a short embarrassed laugh, he gestures with the kushikatsu skewer to his side as if showing where the bear walked past. The interviewer's grey microphone stays in the lower right of frame. Passersby walk in the blurred background. Ambient sound: busy restaurant street murmur. Realistic broadcast video, no text, no captions, no logos.
6

カット6:現場リポーター中継

images/c6.jpg(GPT Image 2.5 sunburst)💎2.75
clips/c6.mp4(Seedance 2.5 / 9s)💎108

セリフ(Masafumi(男性・きびきび) / eleven_v4 / 7.68s)
[serious, field reporter] 現在も、警察と動物園の職員が、クマの行方を捜しています。付近の方は、外出を控えてください。

画像プロンプト
Frame grab from a live Japanese TV news field report. A Japanese male field reporter in his mid 30s, short neat hair, dark navy suit with a plain dark windbreaker over it, holding a plain black handheld microphone without any logo, stands facing the camera, medium shot from waist up, slightly left of center, serious focused expression, mid-sentence. Behind him: a plain yellow barrier tape stretched across the street with no readable text, a Japanese police patrol car with its red rotating roof light glowing, a few uniformed police officers and zoo staff in green work uniforms, and in the background down the street the steel lattice Tsutenkaku tower in Osaka's Shinsekai district. Late afternoon overcast daylight, broadcast ENG camera look, realistic, slightly compressed broadcast video quality, not cinematic. No text, no logos, no captions, no watermarks.
動画プロンプト
Live TV news field report, handheld camera on the cameraman's shoulder with slight natural sway. The male reporter holding the microphone speaks exactly the Japanese line from the audio reference straight to camera with precise natural lip sync, serious urgent but composed delivery, a short glance back over his shoulder toward the police line and then back to camera. Behind him the police car's red rotating light flashes, officers and zoo staff in green uniforms move around behind the yellow tape, Tsutenkaku tower in the background. Ambient sound: police radio chatter, distant siren, street noise. Realistic broadcast video, no text, no captions, no logos.
7

カット7:注記カード(2秒)

bear-news-hook_v2.mp4 / 42.96秒 / 追加💎152.25(累計💎687.5)/ ElevenLabs +72文字

  • 局名:なにわテレビ
  • 動物園:たこ焼き動物園(カット2の音声とテロップ)
  • クマ:黄色を強めに、ずんぐり丸い体型で赤いチョッキ(カット2・3を再生成)
  • 音声:Seedanceの音を外し、ElevenLabsの声だけにした(ダブり解消)。屋外カットには合成ノイズの薄い街の環境音
クマ基準画像 images/bear_ref.jpg 💎2.75
プロンプト
Documentary news-camera photograph of a single real adult bear standing on all fours on an asphalt city street in daylight, full body three-quarter view. The bear is a realistic living animal but with a chubby, round, cuddly build like a beloved storybook bear: plump round belly, rounded face with a short muzzle, small round ears, dark nose, gentle dark eyes. Its thick fluffy fur is a vivid warm golden yellow, clearly yellow like honey and sunflowers, not brown, evenly yellow all over. The bear is wearing a short bright red sleeveless vest (a cropped red knitted waistcoat) snugly over its upper body and chest, its yellow round belly showing below the vest. Realistic fur texture and real animal anatomy, photographed as a real event, not a costume, not a toy, not a cartoon. No honey pot, no other accessories. Shot on a broadcast news ENG video camera, natural overcast afternoon light, slightly compressed broadcast image quality, moderate depth of field. No text, no logos, no watermarks.
1

カット1:スタジオ/アナウンサー

images/c1.jpg(GPT Image 2.5 sunburst)💎2.75
clips/c1.mp4(Seedance 2.5 / 9s)💎108

セリフ(Akari(女性・落ち着き) / eleven_v4 / 8.08s)
[calm, professional news anchor] ニュースです。きょう午後2時ごろ、大阪市なにわ区の、つうてんかくの周辺で、クマが目撃されました。

画像プロンプト
Frame grab from a Japanese evening television news broadcast. A calm Japanese female news anchor in her early 40s, short neat dark bob hair, navy blazer over a white blouse, sits behind a glossy news desk with a few paper script pages in front of her, medium shot from chest up, centered, looking straight into the camera, lips slightly parted mid-sentence, composed professional expression. Background: a modern generic news studio set with soft deep-green and warm amber panels and a large out-of-focus screen showing an abstract city at dusk. Even, flat broadcast studio lighting, realistic skin texture, shot on a studio broadcast camera, slightly compressed broadcast video quality, natural moderate depth of field. Absolutely no text, no captions, no logos, no channel bugs, no watermarks anywhere in the image.
動画プロンプト
Locked-off studio camera, Japanese evening TV news broadcast. The female news anchor at the desk reads the news straight to camera, speaking exactly the Japanese line from the audio reference with precise natural lip sync, calm professional delivery, small natural head movements and blinks, glances down at her script once briefly, hands resting on the papers. Background city screen subtly animated. No camera movement, no zoom. Quiet studio room tone only, no music. Realistic broadcast video, no text, no captions, no logos.
2

カット2:中継カメラ風の引き(VO)

images/c2.jpg(GPT Image 2.5 sunburst)💎2.75
clips/c2.mp4(Seedance 2.5 / 7s)💎84

セリフ(Akari / eleven_v4 / 5.76s)
[calm, professional news anchor] 警察によりますと、クマは近くの、たこ焼き動物園から逃げ出したとみられています。

画像プロンプト
[v2: c2 v1画像を編集] Edit the first image. Keep the entire scene exactly the same: the Shinsekai street, Tsutenkaku tower, the crowd with smartphones, the two police officers, camera angle, framing, lighting and broadcast news camera look. Only replace the brown bear in the middle of the street with the bear from the second image: a chubby round bear with vivid golden yellow fur wearing a short bright red sleeveless vest, same size and position and walking pose on all fours. Realistic, matching the lighting and telephoto compression of the scene. No text, no logos, no watermarks.
動画プロンプト
Live news camera on a tripod with long telephoto lens, slight handheld-free micro shake and a slow small pan to follow the bear. The chubby golden yellow bear wearing a short red sleeveless vest walks slowly and heavily across the Shinsekai street on all fours with a natural bear gait, belly swaying, briefly sniffing the ground. The red vest stays on and keeps its shape. Police officers keep their arms spread holding back the crowd; onlookers in the foreground raise smartphones and shift nervously. Tsutenkaku tower stays fixed in the background. The audio reference is an off-screen female news anchor voice-over; nobody on screen speaks or lip-syncs. Ambient sound: distant crowd murmur, faint police radio, street noise. Realistic broadcast video, no text, no captions, no logos.
3

カット3:視聴者提供スマホ映像

images/c3.jpg(GPT Image 2.5 sunburst)💎2.75
clips/c3.mp4(Seedance 2.5 / 5s)💎60

セリフ(Haruta PVC (JA) / eleven_v4 / 3.84s)
[surprised] え、ちょ、クマやん! [excited] ほんまにクマやって!黄色いで!

画像プロンプト
[v2: c3 v1画像を編集] Edit the first image. Keep the entire scene exactly the same: the narrow kushikatsu alley, red lanterns, signs, beer crate, the two startled pedestrians, vertical smartphone framing, camera angle and amateur phone video look. Only replace the brown bear with the bear from the second image: a chubby round bear with vivid golden yellow fur wearing a short bright red sleeveless vest, same size, position and mid-stride walking pose on all fours, head turned slightly toward the camera. Realistic, matching the scene's lighting and phone camera noise. No text, no UI, no logos, no watermarks.
動画プロンプト
Vertical amateur smartphone footage, shaky handheld, the bystander filming steps back a little and jerkily re-frames to keep the bear in view, brief autofocus hunting. The chubby golden yellow bear wearing a short red sleeveless vest lumbers across the narrow kushikatsu alley from left to right on all fours with a natural heavy bear gait, then turns its head toward the camera for a moment. The red vest stays on and keeps its shape. The two pedestrians in the background step back startled. The audio reference is the excited voice of the person holding the phone, off-camera; nobody on screen speaks. Ambient sound: alley chatter, a startled shout in the distance, phone-mic wind noise. Realistic phone video, no text, no UI, no logos.
4

カット4:街頭インタビュー1(おばちゃん)

images/c4.jpg(GPT Image 2.5 sunburst)💎2.75
clips/c4.mp4(Seedance 2.5 / 7s)💎84

セリフ(Aiko(関西・中年女性) / eleven_v4 / 6.08s)
[surprised] びっくりしたわぁ。黄色いから、ぬいぐるみかと思たら、[laughs] ほんまもんやねん。

画像プロンプト
Frame grab from a Japanese TV news street interview (man-on-the-street). A friendly Osaka woman in her early 60s with short permed reddish-brown hair, wearing a leopard-print blouse and a light cardigan, holding a shopping bag, stands on a busy old-fashioned downtown Osaka shopping street, medium close-up from chest up, positioned slightly right of center, looking just off-camera to the left toward an interviewer, mid-speech with an amused surprised expression, hand near her chest. The tip of a plain grey foam-covered handheld microphone without any logo enters from the lower left of the frame. Background: blurred colorful restaurant facades and lanterns, passersby, signs illegible. Natural afternoon daylight, broadcast ENG camera look, realistic skin texture, slightly compressed broadcast video quality, not cinematic. No text, no logos, no captions, no watermarks.
動画プロンプト
Handheld TV news street interview, gentle natural camera sway. The Osaka woman in the leopard-print blouse speaks exactly the Kansai-dialect Japanese line from the audio reference to the interviewer just off-camera, with precise natural lip sync, animated expressive face, a surprised look then a laugh, hand patting her chest and a small wave of her hand. The interviewer's grey microphone stays in the lower left of frame. Passersby walk in the blurred background. Ambient sound: busy shopping street murmur. Realistic broadcast video, no text, no captions, no logos.
5

カット5:街頭インタビュー2(若い男性)

images/c5.jpg(GPT Image 2.5 sunburst)💎2.75
clips/c5.mp4(Seedance 2.5 / 6s)💎72

セリフ(Riku(関西・若い男性) / eleven_v4 / 5.28s)
串カツ食べてたら、横をふつうに歩いてて… [laughs] 二度見しました。

画像プロンプト
Frame grab from a Japanese TV news street interview (man-on-the-street). A Japanese man in his early 20s with short black hair, wearing a grey hoodie under a dark jacket and holding a wooden skewer, stands on an old-fashioned downtown Osaka shopping street lined with small kushikatsu restaurants, medium close-up from chest up, positioned slightly left of center, looking just off-camera to the right toward an interviewer, mid-speech with a slightly stunned half-smile. The tip of a plain grey foam-covered handheld microphone without any logo enters from the lower right of the frame. Background: blurred red lanterns, restaurant facades, passersby, signs illegible. Natural afternoon daylight, broadcast ENG camera look, realistic skin texture, slightly compressed broadcast video quality, not cinematic. No text, no logos, no captions, no watermarks.
動画プロンプト
Handheld TV news street interview, gentle natural camera sway. The young man in the grey hoodie speaks exactly the Japanese line from the audio reference to the interviewer just off-camera, with precise natural lip sync, slightly stunned then a short embarrassed laugh, he gestures with the kushikatsu skewer to his side as if showing where the bear walked past. The interviewer's grey microphone stays in the lower right of frame. Passersby walk in the blurred background. Ambient sound: busy restaurant street murmur. Realistic broadcast video, no text, no captions, no logos.
6

カット6:現場リポーター中継

images/c6.jpg(GPT Image 2.5 sunburst)💎2.75
clips/c6.mp4(Seedance 2.5 / 9s)💎108

セリフ(Masafumi(男性・きびきび) / eleven_v4 / 7.68s)
[serious, field reporter] 現在も、警察と動物園の職員が、クマの行方を捜しています。付近の方は、外出を控えてください。

画像プロンプト
Frame grab from a live Japanese TV news field report. A Japanese male field reporter in his mid 30s, short neat hair, dark navy suit with a plain dark windbreaker over it, holding a plain black handheld microphone without any logo, stands facing the camera, medium shot from waist up, slightly left of center, serious focused expression, mid-sentence. Behind him: a plain yellow barrier tape stretched across the street with no readable text, a Japanese police patrol car with its red rotating roof light glowing, a few uniformed police officers and zoo staff in green work uniforms, and in the background down the street the steel lattice Tsutenkaku tower in Osaka's Shinsekai district. Late afternoon overcast daylight, broadcast ENG camera look, realistic, slightly compressed broadcast video quality, not cinematic. No text, no logos, no captions, no watermarks.
動画プロンプト
Live TV news field report, handheld camera on the cameraman's shoulder with slight natural sway. The male reporter holding the microphone speaks exactly the Japanese line from the audio reference straight to camera with precise natural lip sync, serious urgent but composed delivery, a short glance back over his shoulder toward the police line and then back to camera. Behind him the police car's red rotating light flashes, officers and zoo staff in green uniforms move around behind the yellow tape, Tsutenkaku tower in the background. Ambient sound: police radio chatter, distant siren, street noise. Realistic broadcast video, no text, no captions, no logos.
7

カット7:注記カード(2秒)

まとめ

今回つくったもの

10画像GPT Image 2.5
7音声ElevenLabs v4
8動画Seedance 2.5
2完成動画ffmpeg
約43秒1本あたり

生成時間とクレジット Higgsfield MCP

💎687.5使ったクレジット画像 27.5 + 動画 660≈ 4,812円
⏱11分29秒実際の待ち時間同時に並べて生成した待ち時間の合計
18生成ジョブ画像 10 / 動画 8

バージョン1

💎535.25 ≈ 3,747円

  • 画像7枚:⏱1分26秒(クマ1枚 → 残り6枚を同時に)
  • 動画6本:⏱5分07秒(6本を同時に)

バージョン2(作り直し分)

💎152.25 ≈ 1,066円

  • 画像3枚:⏱1分07秒(クマ1枚 → 2枚を同時に)
  • 動画2本:⏱3分49秒(2本を同時に)

💴 1クレジット = 7円 で換算 / 画像1枚 💎2.75 ≈ 19.25円 / 動画1秒 💎12 ≈ 84円

単価:GPT Image 2.5(sunburst・high・2K)= 1枚 💎2.75 / Seedance 2.5(1080p)= 1秒あたり 💎12。生成時間は各ジョブの開始〜完了の実測

📋 ジョブごとの内訳(18件)
版種類対象モデルクレジット円換算生成時間
v1画像クマ基準GPT Image 2.5💎2.7519.25円⏱33秒
共通画像カット1GPT Image 2.5💎2.7519.25円⏱35秒
v1画像カット2GPT Image 2.5💎2.7519.25円⏱46秒
v1画像カット3GPT Image 2.5💎2.7519.25円⏱52秒
共通画像カット4GPT Image 2.5💎2.7519.25円⏱38秒
共通画像カット5GPT Image 2.5💎2.7519.25円⏱37秒
共通画像カット6GPT Image 2.5💎2.7519.25円⏱37秒
共通動画カット1(9秒)Seedance 2.5💎108756円⏱4分41秒
v1動画カット2(7秒)Seedance 2.5💎84588円⏱4分49秒
v1動画カット3(5秒)Seedance 2.5💎60420円⏱4分16秒
共通動画カット4(7秒)Seedance 2.5💎84588円⏱3分45秒
共通動画カット5(6秒)Seedance 2.5💎72504円⏱2分57秒
共通動画カット6(9秒)Seedance 2.5💎108756円⏱3分31秒
v2画像クマ基準GPT Image 2.5💎2.7519.25円⏱32秒
v2画像カット2GPT Image 2.5💎2.7519.25円⏱35秒
v2画像カット3GPT Image 2.5💎2.7519.25円⏱34秒
v2動画カット2(7秒)Seedance 2.5💎84588円⏱3分49秒
v2動画カット3(5秒)Seedance 2.5💎60420円⏱3分02秒

完成動画

バージョン1:かもめテレビ/天王寺動物園
バージョン2:なにわテレビ/たこ焼き動物園/赤いチョッキ

カットごとの生成物

画像 GPT Image 2.5音声 ElevenLabs v4動画 Seedance 2.5仕上げ ffmpeg
1スタジオ8.6秒
画像 💎2.75(19.25円)・ ⏱35秒
カット1画像
音声

「ニュースです。きょう午後2時ごろ、大阪市なにわ区の通天閣の周辺で、クマが目撃されました。」

動画 💎108(756円)・ ⏱4分41秒
テロップ入り
2現場の引き映像6.4秒
画像 💎2.75(19.25円)・ ⏱46秒
カット2画像
音声

「警察によりますと、クマは近くの天王寺動物園から逃げ出したとみられています。」

動画 💎84(588円)・ ⏱4分49秒
テロップ入り
3視聴者提供のスマホ映像4.6秒
画像 💎2.75(19.25円)・ ⏱52秒
カット3画像
音声

「え、ちょ、クマやん!ほんまにクマやって!黄色いで!」

動画 💎60(420円)・ ⏱4分16秒
テロップ入り
4街頭インタビュー①6.5秒
画像 💎2.75(19.25円)・ ⏱38秒
カット4画像
音声

「びっくりしたわぁ。黄色いから、ぬいぐるみかと思たら、ほんまもんやねん。」

動画 💎84(588円)・ ⏱3分45秒
テロップ入り
5街頭インタビュー②6.0秒
画像 💎2.75(19.25円)・ ⏱37秒
カット5画像
音声

「串カツ食べてたら、横をふつうに歩いてて…二度見しました。」

動画 💎72(504円)・ ⏱2分57秒
テロップ入り
6現場リポーター中継8.9秒
画像 💎2.75(19.25円)・ ⏱37秒
カット6画像
音声

「現在も、警察と動物園の職員が、クマの行方を捜しています。付近の方は、外出を控えてください。」

動画 💎108(756円)・ ⏱3分31秒
テロップ入り

生成した画像ぜんぶ

クマ基準(v1)
クマ基準(v1)
カット1(v1)
カット1(v1)
カット2(v1)
カット2(v1)
カット3(v1)
カット3(v1)
カット4(v1)
カット4(v1)
カット5(v1)
カット5(v1)
カット6(v1)
カット6(v1)
クマ基準(v2)
クマ基準(v2)
カット2(v2)
カット2(v2)
カット3(v2)
カット3(v2)

映画CM

ハルタとクマ

動物園から逃げ出したクマと、少年ハルタの物語 / 約30秒

💎253.75使ったクレジット画像 13.75 + 動画 240≈ 1,776円
⏱6分00秒実際の待ち時間同時に並べて生成した待ち時間の合計
9生成ジョブ画像 5 / 動画 4

主役のふたり

ハルタ
ハルタ(AIオリジナルの少年)💎2.75(19.25円)・ ⏱37秒
クマ
クマ(バージョン2の基準画像をリファレンス)

カットごとの生成物

画像 GPT Image 2.5動画 Seedance 2.5声 ElevenLabs v4(ナレーション:Koichi Yashiro / ハルタ:Toru)
1大阪に現れたクマ
画像
大阪に現れたクマ
動画

ナレーション

「その日、大阪の街に… 一頭のクマが、迷い込んだ。」

ニュース映像(バージョン2のカット2)を再利用。映画風の色と黒帯

2出会い
画像 💎2.75(19.25円)・ ⏱42秒
出会い
動画 💎72(504円)・ ⏱3分07秒

ナレーション

「誰もが、逃げ出した。…ひとりの少年を、除いて。」

夕暮れの路地で、ハルタとクマが目を合わせる

3夜の新世界を駆け抜ける
画像 💎2.75(19.25円)・ ⏱45秒
夜の新世界を駆け抜ける
動画 💎48(336円)・ ⏱3分27秒

(音楽のみ)

ネオンと提灯の中を、ふたりで疾走

4雨の中の決意
画像 💎2.75(19.25円)・ ⏱42秒
雨の中の決意
動画 💎60(420円)・ ⏱4分38秒

ハルタ

「だいじょうぶ。ぼくが、おうちに帰したる。」

クマをかばうハルタ(口パクを音声に合わせる)

5夕焼けの屋上
画像 💎2.75(19.25円)・ ⏱43秒
夕焼けの屋上
動画 💎60(420円)・ ⏱4分09秒

ナレーション

「ふたりの、小さな大冒険が… はじまる。」

通天閣を望む屋上に並ぶふたり

6タイトル
画像
タイトル
動画

ナレーション

「ハルタとクマ。この夏、公開。」

タイトル『ハルタとクマ』+注記

BGMは、Seedance 2.5 が動画と一緒に生成したオーケストラの音を使用。ナレーション中は自動で音量を下げています

Bear Spotted Near Tsutenkaku — News Clip

Script

Bear Spotted Near Tsutenkaku — News Clip

Production flow

0ScriptClaude Code
→
1Voice linesElevenLabs v4
→
2Image for each cutGPT Image 2.5
→
3AnimateSeedance 2.5
Higgsfield MCP
→
4Captions & joinffmpeg

Key point: make the voice first, then pass it to Seedance so the lip movement and timing match

Script

Bear Spotted Near Tsutenkaku — News Clip

  1. 1Anchor

    “News now. At around 2 p.m. today, a bear was spotted near Tsutenkaku Tower in Naniwa Ward, Osaka.”

  2. 2Anchor

    “According to police, the bear is believed to have escaped from the nearby Tennoji Zoo.”

  3. 3Person filming

    “Wait, whoa—it's a bear! It's actually a bear! And it's yellow!”

  4. 4Local Osaka lady

    “I was so surprised. It was yellow, so I thought it was a stuffed animal—but it was the real thing!”

  5. 5Young man

    “I was eating kushikatsu, and it just walked right past me… I did a double take.”

  6. 6Reporter

    “Police and zoo staff are still searching for the bear. Residents in the area are advised to stay indoors.”

  7. 7On-screen text

    ※ This video is AI-generated fiction. It has no connection to any real event, broadcaster, or person.

Script

Voice for each scene

Each line generated one by one with ElevenLabs v4

  1. 1

    StudioAnchor

    “News now. At around 2 p.m. today, a bear was spotted near Tsutenkaku Tower in Naniwa Ward, Osaka.”

    🔊 ElevenLabs v4: Akari (calm female)
  2. 2

    Wide shot of the sceneAnchor

    “According to police, the bear is believed to have escaped from the nearby Tennoji Zoo.”

    🔊 ElevenLabs v4: Akari (calm female)
  3. 3

    Viewer's smartphone videoPerson filming

    “Wait, whoa—it's a bear! It's actually a bear! And it's yellow!”

    🔊 ElevenLabs v4: Haruta (his own voice clone)
  4. 4

    Street interview ①Local Osaka lady

    “I was so surprised. It was yellow, so I thought it was a stuffed animal—but it was the real thing!”

    🔊 ElevenLabs v4: Aiko (Kansai dialect, female)
  5. 5

    Street interview ②Young man

    “I was eating kushikatsu, and it just walked right past me… I did a double take.”

    🔊 ElevenLabs v4: Riku (Kansai dialect, male)
  6. 6

    Live field reportReporter

    “Police and zoo staff are still searching for the bear. Residents in the area are advised to stay indoors.”

    🔊 ElevenLabs v4: Masafumi (crisp male)

Script

Image for each cut

The first frame made with GPT Image 2.5, before animating

Cut 1
1StudioFemale anchor at the news desk
📝 Image prompt (GPT Image 2.5)
Frame grab from a Japanese evening television news broadcast. A calm Japanese female news anchor in her early 40s, short neat dark bob hair, navy blazer over a white blouse, sits behind a glossy news desk with a few paper script pages in front of her, medium shot from chest up, centered, looking straight into the camera, lips slightly parted mid-sentence, composed professional expression. Background: a modern generic news studio set with soft deep-green and warm amber panels and a large out-of-focus screen showing an abstract city at dusk. Even, flat broadcast studio lighting, realistic skin texture, shot on a studio broadcast camera, slightly compressed broadcast video quality, natural moderate depth of field. Absolutely no text, no captions, no logos, no channel bugs, no watermarks anywhere in the image.
Cut 2
2Wide shot of the sceneThe bear walking beyond the crowd on a Shinsekai street, Tsutenkaku behind
📝 Image prompt (GPT Image 2.5)
Frame grab from a live Japanese TV news helicopter-free ground camera, long telephoto lens from an elevated position. The Shinsekai shopping street in Osaka in the afternoon: a straight street lined with old-fashioned colorful restaurant facades, paper lanterns and oversized decorative signboards, and at the far end of the street the steel lattice tower Tsutenkaku rising against a hazy sky. In the foreground, backs of a crowd of onlookers holding up smartphones, two police officers in dark uniforms spreading their arms to hold people back. In the middle distance, in the center of the street, the same realistic golden honey-yellow bear from the reference image walks slowly across the pavement on all fours. All signboards are fictional and illegible from the distance, no readable words. Telephoto compression, slight heat haze, natural overcast daylight, broadcast ENG camera look, slightly compressed broadcast video quality, realistic, not cinematic. No text, no logos, no captions, no watermarks.
Cut 3
3Viewer's smartphone videoThe bear crossing an alley of kushikatsu shops (vertical video)
📝 Image prompt (GPT Image 2.5)
Vertical amateur smartphone video frame filmed by a bystander, handheld and slightly tilted, from about 10 meters away at eye level. A narrow old Osaka Shinsekai alley lined with small kushikatsu restaurants: red paper lanterns, noren curtains, plastic stools and a beer crate outside, menu boards with illegible scribbles. The same realistic golden honey-yellow bear from the reference image lumbers across the alley from left to right on all fours, mid-stride, while two pedestrians in the background step back startled. Phone camera look: slightly overexposed sky, mild motion blur, digital noise, auto-exposure, ordinary afternoon light, not cinematic. Signs must be illegible or blurred, no readable text, no logos, no UI overlay, no watermarks.
Cut 4
4Street interview ①A local lady being interviewed in Shinsekai
📝 Image prompt (GPT Image 2.5)
Frame grab from a Japanese TV news street interview (man-on-the-street). A friendly Osaka woman in her early 60s with short permed reddish-brown hair, wearing a leopard-print blouse and a light cardigan, holding a shopping bag, stands on a busy old-fashioned downtown Osaka shopping street, medium close-up from chest up, positioned slightly right of center, looking just off-camera to the left toward an interviewer, mid-speech with an amused surprised expression, hand near her chest. The tip of a plain grey foam-covered handheld microphone without any logo enters from the lower left of the frame. Background: blurred colorful restaurant facades and lanterns, passersby, signs illegible. Natural afternoon daylight, broadcast ENG camera look, realistic skin texture, slightly compressed broadcast video quality, not cinematic. No text, no logos, no captions, no watermarks.
Cut 5
5Street interview ②A young man holding kushikatsu
📝 Image prompt (GPT Image 2.5)
Frame grab from a Japanese TV news street interview (man-on-the-street). A Japanese man in his early 20s with short black hair, wearing a grey hoodie under a dark jacket and holding a wooden skewer, stands on an old-fashioned downtown Osaka shopping street lined with small kushikatsu restaurants, medium close-up from chest up, positioned slightly left of center, looking just off-camera to the right toward an interviewer, mid-speech with a slightly stunned half-smile. The tip of a plain grey foam-covered handheld microphone without any logo enters from the lower right of the frame. Background: blurred red lanterns, restaurant facades, passersby, signs illegible. Natural afternoon daylight, broadcast ENG camera look, realistic skin texture, slightly compressed broadcast video quality, not cinematic. No text, no logos, no captions, no watermarks.
Cut 6
6Live field reportReporter in front of police tape, a patrol car and Tsutenkaku
📝 Image prompt (GPT Image 2.5)
Frame grab from a live Japanese TV news field report. A Japanese male field reporter in his mid 30s, short neat hair, dark navy suit with a plain dark windbreaker over it, holding a plain black handheld microphone without any logo, stands facing the camera, medium shot from waist up, slightly left of center, serious focused expression, mid-sentence. Behind him: a plain yellow barrier tape stretched across the street with no readable text, a Japanese police patrol car with its red rotating roof light glowing, a few uniformed police officers and zoo staff in green work uniforms, and in the background down the street the steel lattice Tsutenkaku tower in Osaka's Shinsekai district. Late afternoon overcast daylight, broadcast ENG camera look, realistic, slightly compressed broadcast video quality, not cinematic. No text, no logos, no captions, no watermarks.

Script

Video for each cut

Animated from the image and voice with Seedance 2.5 (Higgsfield MCP)

1StudioImage + voice → lip-sync to the voice
📝 Video prompt (Seedance 2.5)
Locked-off studio camera, Japanese evening TV news broadcast. The female news anchor at the desk reads the news straight to camera, speaking exactly the Japanese line from the audio reference with precise natural lip sync, calm professional delivery, small natural head movements and blinks, glances down at her script once briefly, hands resting on the papers. Background city screen subtly animated. No camera movement, no zoom. Quiet studio room tone only, no music. Realistic broadcast video, no text, no captions, no logos.
2Wide shot of the sceneImage + voice → the bear lumbers along
📝 Video prompt (Seedance 2.5)
Live news camera on a tripod with long telephoto lens, slight handheld-free micro shake and a slow small pan to follow the bear. The realistic golden honey-yellow bear walks slowly and heavily across the Shinsekai street on all fours with natural bear gait, shoulder muscles rolling, briefly sniffing the ground. Police officers keep their arms spread holding back the crowd; onlookers in the foreground raise smartphones and shift nervously. Tsutenkaku tower stays fixed in the background. The audio reference is an off-screen female news anchor voice-over; nobody on screen speaks or lip-syncs. Ambient sound: distant crowd murmur, faint police radio, street noise. Realistic broadcast video, no text, no captions, no logos.
3Viewer's smartphone videoImage + voice → shaky smartphone footage
📝 Video prompt (Seedance 2.5)
Vertical amateur smartphone footage, shaky handheld, the bystander filming steps back a little and jerkily re-frames to keep the bear in view, brief autofocus hunting. The realistic golden honey-yellow bear lumbers across the narrow kushikatsu alley from left to right on all fours with a natural heavy bear gait, then turns its head toward the camera for a moment. The two pedestrians in the background step back startled. The audio reference is the excited voice of the person holding the phone, off-camera; nobody on screen speaks. Ambient sound: alley chatter, a startled shout in the distance, phone-mic wind noise. Realistic phone video, no text, no UI, no logos.
4Street interview ①Image + voice → lip-sync to the voice
📝 Video prompt (Seedance 2.5)
Handheld TV news street interview, gentle natural camera sway. The Osaka woman in the leopard-print blouse speaks exactly the Kansai-dialect Japanese line from the audio reference to the interviewer just off-camera, with precise natural lip sync, animated expressive face, a surprised look then a laugh, hand patting her chest and a small wave of her hand. The interviewer's grey microphone stays in the lower left of frame. Passersby walk in the blurred background. Ambient sound: busy shopping street murmur. Realistic broadcast video, no text, no captions, no logos.
5Street interview ②Image + voice → lip-sync to the voice
📝 Video prompt (Seedance 2.5)
Handheld TV news street interview, gentle natural camera sway. The young man in the grey hoodie speaks exactly the Japanese line from the audio reference to the interviewer just off-camera, with precise natural lip sync, slightly stunned then a short embarrassed laugh, he gestures with the kushikatsu skewer to his side as if showing where the bear walked past. The interviewer's grey microphone stays in the lower right of frame. Passersby walk in the blurred background. Ambient sound: busy restaurant street murmur. Realistic broadcast video, no text, no captions, no logos.
6Live field reportImage + voice → lip-sync to the voice
📝 Video prompt (Seedance 2.5)
Live TV news field report, handheld camera on the cameraman's shoulder with slight natural sway. The male reporter holding the microphone speaks exactly the Japanese line from the audio reference straight to camera with precise natural lip sync, serious urgent but composed delivery, a short glance back over his shoulder toward the police line and then back to camera. Behind him the police car's red rotating light flashes, officers and zoo staff in green uniforms move around behind the yellow tape, Tsutenkaku tower in the background. Ambient sound: police radio chatter, distant siren, street noise. Realistic broadcast video, no text, no captions, no logos.

Script

Finishing with ffmpeg

Finish each cut first, then join all 7 at the end

①Finish each cut × 6 cuts

  1. TrimCut out only the part of the Seedance video we use
  2. Overlay captionsLay the station logo, LIVE, clock and headline on top as a transparent image
  3. Replace the audioSwap in the original ElevenLabs voice and nudge it to match the lips
Raw generated video
Raw generated video (no text)
+
On-screen textImage
Caption image (transparent background)
=
With captions
Cut with captions

Only cut 3 (smartphone video): place the vertical video in the center and fill the sides with a blurred background

②Make the closing disclaimer card 2 s

“※ This video is AI-generated fiction”

③Join the 7 clips in order

Cut 1 final
1
+
Cut 2 final
2
+
Cut 3 final
3
+
Cut 4 final
4
+
Cut 5 final
5
+
Cut 6 final
6
+
Cut 7 final
7

Hard cuts between shots (like real news). Finally, even out the loudness and export

→ Done: about 43 s, 1080p

bear-news-hook_v1_clean.mp4 / 43.15 s / 💎535.25 / ElevenLabs 367 chars

  • Station: Kamome TV
  • Zoo: Tennoji Zoo
  • Bear: realistic brown-bear build, pale golden fur, no clothes
  • Audio: ElevenLabs voices only (clean version with the doubled voice removed; the original with doubling is bear-news-hook_v1.mp4)
Bear reference image work/v1_backup/bear_ref.jpg 💎2.75
Prompt
Documentary wildlife reference photograph of a single real adult bear standing on all fours on an asphalt city street in daylight, full body side three-quarter view. The bear is a completely realistic animal with the anatomy of a brown bear (Ursus arctos): heavy shoulder hump, long snout, small rounded ears, dark nose, natural animal eyes, long claws. Its fur is an unusual soft pale golden honey-yellow color, thick and fluffy, slightly lighter on the face and darker golden on the back, natural fur texture with dust and clumps, realistic wildlife look. No clothing, no shirt, no accessories, no cartoon features, not a teddy bear, not anthropomorphic. Shot on a broadcast news ENG video camera, natural overcast afternoon light, neutral color, slightly compressed broadcast image quality, moderate depth of field. No text, no logos, no watermarks.
1

Cut 1: Studio / Anchor

images/c1.jpg (GPT Image 2.5 sunburst)💎2.75
clips/c1.mp4 (Seedance 2.5 / 9s)💎108

Line (Akari (female, calm) / eleven_v4 / 8.08s)
[calm, professional news anchor] News now. At around 2 p.m. today, a bear was spotted near Tsutenkaku Tower in Naniwa Ward, Osaka.

Image prompt
Frame grab from a Japanese evening television news broadcast. A calm Japanese female news anchor in her early 40s, short neat dark bob hair, navy blazer over a white blouse, sits behind a glossy news desk with a few paper script pages in front of her, medium shot from chest up, centered, looking straight into the camera, lips slightly parted mid-sentence, composed professional expression. Background: a modern generic news studio set with soft deep-green and warm amber panels and a large out-of-focus screen showing an abstract city at dusk. Even, flat broadcast studio lighting, realistic skin texture, shot on a studio broadcast camera, slightly compressed broadcast video quality, natural moderate depth of field. Absolutely no text, no captions, no logos, no channel bugs, no watermarks anywhere in the image.
Video prompt
Locked-off studio camera, Japanese evening TV news broadcast. The female news anchor at the desk reads the news straight to camera, speaking exactly the Japanese line from the audio reference with precise natural lip sync, calm professional delivery, small natural head movements and blinks, glances down at her script once briefly, hands resting on the papers. Background city screen subtly animated. No camera movement, no zoom. Quiet studio room tone only, no music. Realistic broadcast video, no text, no captions, no logos.
2

Cut 2: Wide live-camera shot (VO)

work/v1_backup/c2.jpg (GPT Image 2.5 sunburst)💎2.75
work/v1_backup/c2.mp4 (Seedance 2.5 / 7s)💎84

Line (Akari / eleven_v4 / 6.00s)
[calm, professional news anchor] According to police, the bear is believed to have escaped from the nearby Tennoji Zoo.

Image prompt
Frame grab from a live Japanese TV news helicopter-free ground camera, long telephoto lens from an elevated position. The Shinsekai shopping street in Osaka in the afternoon: a straight street lined with old-fashioned colorful restaurant facades, paper lanterns and oversized decorative signboards, and at the far end of the street the steel lattice tower Tsutenkaku rising against a hazy sky. In the foreground, backs of a crowd of onlookers holding up smartphones, two police officers in dark uniforms spreading their arms to hold people back. In the middle distance, in the center of the street, the same realistic golden honey-yellow bear from the reference image walks slowly across the pavement on all fours. All signboards are fictional and illegible from the distance, no readable words. Telephoto compression, slight heat haze, natural overcast daylight, broadcast ENG camera look, slightly compressed broadcast video quality, realistic, not cinematic. No text, no logos, no captions, no watermarks.
Video prompt
Live news camera on a tripod with long telephoto lens, slight handheld-free micro shake and a slow small pan to follow the bear. The realistic golden honey-yellow bear walks slowly and heavily across the Shinsekai street on all fours with natural bear gait, shoulder muscles rolling, briefly sniffing the ground. Police officers keep their arms spread holding back the crowd; onlookers in the foreground raise smartphones and shift nervously. Tsutenkaku tower stays fixed in the background. The audio reference is an off-screen female news anchor voice-over; nobody on screen speaks or lip-syncs. Ambient sound: distant crowd murmur, faint police radio, street noise. Realistic broadcast video, no text, no captions, no logos.
3

Cut 3: Viewer's smartphone video

work/v1_backup/c3.jpg (GPT Image 2.5 sunburst)💎2.75
work/v1_backup/c3.mp4 (Seedance 2.5 / 5s)💎60

Line (Haruta PVC (JA) / eleven_v4 / 3.84s)
[surprised] Wait, whoa—it's a bear! [excited] It's actually a bear! And it's yellow!

Image prompt
Vertical amateur smartphone video frame filmed by a bystander, handheld and slightly tilted, from about 10 meters away at eye level. A narrow old Osaka Shinsekai alley lined with small kushikatsu restaurants: red paper lanterns, noren curtains, plastic stools and a beer crate outside, menu boards with illegible scribbles. The same realistic golden honey-yellow bear from the reference image lumbers across the alley from left to right on all fours, mid-stride, while two pedestrians in the background step back startled. Phone camera look: slightly overexposed sky, mild motion blur, digital noise, auto-exposure, ordinary afternoon light, not cinematic. Signs must be illegible or blurred, no readable text, no logos, no UI overlay, no watermarks.
Video prompt
Vertical amateur smartphone footage, shaky handheld, the bystander filming steps back a little and jerkily re-frames to keep the bear in view, brief autofocus hunting. The realistic golden honey-yellow bear lumbers across the narrow kushikatsu alley from left to right on all fours with a natural heavy bear gait, then turns its head toward the camera for a moment. The two pedestrians in the background step back startled. The audio reference is the excited voice of the person holding the phone, off-camera; nobody on screen speaks. Ambient sound: alley chatter, a startled shout in the distance, phone-mic wind noise. Realistic phone video, no text, no UI, no logos.
4

Cut 4: Street interview 1 (local lady)

images/c4.jpg (GPT Image 2.5 sunburst)💎2.75
clips/c4.mp4 (Seedance 2.5 / 7s)💎84

Line (Aiko (Kansai, middle-aged female) / eleven_v4 / 6.08s)
[surprised] I was so surprised. It was yellow, so I thought it was a stuffed animal—[laughs] but it was the real thing!

Image prompt
Frame grab from a Japanese TV news street interview (man-on-the-street). A friendly Osaka woman in her early 60s with short permed reddish-brown hair, wearing a leopard-print blouse and a light cardigan, holding a shopping bag, stands on a busy old-fashioned downtown Osaka shopping street, medium close-up from chest up, positioned slightly right of center, looking just off-camera to the left toward an interviewer, mid-speech with an amused surprised expression, hand near her chest. The tip of a plain grey foam-covered handheld microphone without any logo enters from the lower left of the frame. Background: blurred colorful restaurant facades and lanterns, passersby, signs illegible. Natural afternoon daylight, broadcast ENG camera look, realistic skin texture, slightly compressed broadcast video quality, not cinematic. No text, no logos, no captions, no watermarks.
Video prompt
Handheld TV news street interview, gentle natural camera sway. The Osaka woman in the leopard-print blouse speaks exactly the Kansai-dialect Japanese line from the audio reference to the interviewer just off-camera, with precise natural lip sync, animated expressive face, a surprised look then a laugh, hand patting her chest and a small wave of her hand. The interviewer's grey microphone stays in the lower left of frame. Passersby walk in the blurred background. Ambient sound: busy shopping street murmur. Realistic broadcast video, no text, no captions, no logos.
5

Cut 5: Street interview 2 (young man)

images/c5.jpg (GPT Image 2.5 sunburst)💎2.75
clips/c5.mp4 (Seedance 2.5 / 6s)💎72

Line (Riku (Kansai, young male) / eleven_v4 / 5.28s)
I was eating kushikatsu, and it just walked right past me… [laughs] I did a double take.

Image prompt
Frame grab from a Japanese TV news street interview (man-on-the-street). A Japanese man in his early 20s with short black hair, wearing a grey hoodie under a dark jacket and holding a wooden skewer, stands on an old-fashioned downtown Osaka shopping street lined with small kushikatsu restaurants, medium close-up from chest up, positioned slightly left of center, looking just off-camera to the right toward an interviewer, mid-speech with a slightly stunned half-smile. The tip of a plain grey foam-covered handheld microphone without any logo enters from the lower right of the frame. Background: blurred red lanterns, restaurant facades, passersby, signs illegible. Natural afternoon daylight, broadcast ENG camera look, realistic skin texture, slightly compressed broadcast video quality, not cinematic. No text, no logos, no captions, no watermarks.
Video prompt
Handheld TV news street interview, gentle natural camera sway. The young man in the grey hoodie speaks exactly the Japanese line from the audio reference to the interviewer just off-camera, with precise natural lip sync, slightly stunned then a short embarrassed laugh, he gestures with the kushikatsu skewer to his side as if showing where the bear walked past. The interviewer's grey microphone stays in the lower right of frame. Passersby walk in the blurred background. Ambient sound: busy restaurant street murmur. Realistic broadcast video, no text, no captions, no logos.
6

Cut 6: Live field report

images/c6.jpg (GPT Image 2.5 sunburst)💎2.75
clips/c6.mp4 (Seedance 2.5 / 9s)💎108

Line (Masafumi (male, crisp) / eleven_v4 / 7.68s)
[serious, field reporter] Police and zoo staff are still searching for the bear. Residents in the area are advised to stay indoors.

Image prompt
Frame grab from a live Japanese TV news field report. A Japanese male field reporter in his mid 30s, short neat hair, dark navy suit with a plain dark windbreaker over it, holding a plain black handheld microphone without any logo, stands facing the camera, medium shot from waist up, slightly left of center, serious focused expression, mid-sentence. Behind him: a plain yellow barrier tape stretched across the street with no readable text, a Japanese police patrol car with its red rotating roof light glowing, a few uniformed police officers and zoo staff in green work uniforms, and in the background down the street the steel lattice Tsutenkaku tower in Osaka's Shinsekai district. Late afternoon overcast daylight, broadcast ENG camera look, realistic, slightly compressed broadcast video quality, not cinematic. No text, no logos, no captions, no watermarks.
Video prompt
Live TV news field report, handheld camera on the cameraman's shoulder with slight natural sway. The male reporter holding the microphone speaks exactly the Japanese line from the audio reference straight to camera with precise natural lip sync, serious urgent but composed delivery, a short glance back over his shoulder toward the police line and then back to camera. Behind him the police car's red rotating light flashes, officers and zoo staff in green uniforms move around behind the yellow tape, Tsutenkaku tower in the background. Ambient sound: police radio chatter, distant siren, street noise. Realistic broadcast video, no text, no captions, no logos.
7

Cut 7: Disclaimer card (2 s)

bear-news-hook_v2.mp4 / 42.96 s / added 💎152.25 (total 💎687.5) / ElevenLabs +72 chars

  • Station: Naniwa TV
  • Zoo: Takoyaki Zoo (cut 2 voice and captions)
  • Bear: more yellow, chubby round build, red vest (cuts 2 & 3 regenerated)
  • Audio: removed Seedance's audio and kept only the ElevenLabs voices (fixes the doubling). Outdoor cuts get a faint synthetic street ambience
Bear reference image images/bear_ref.jpg 💎2.75
Prompt
Documentary news-camera photograph of a single real adult bear standing on all fours on an asphalt city street in daylight, full body three-quarter view. The bear is a realistic living animal but with a chubby, round, cuddly build like a beloved storybook bear: plump round belly, rounded face with a short muzzle, small round ears, dark nose, gentle dark eyes. Its thick fluffy fur is a vivid warm golden yellow, clearly yellow like honey and sunflowers, not brown, evenly yellow all over. The bear is wearing a short bright red sleeveless vest (a cropped red knitted waistcoat) snugly over its upper body and chest, its yellow round belly showing below the vest. Realistic fur texture and real animal anatomy, photographed as a real event, not a costume, not a toy, not a cartoon. No honey pot, no other accessories. Shot on a broadcast news ENG video camera, natural overcast afternoon light, slightly compressed broadcast image quality, moderate depth of field. No text, no logos, no watermarks.
1

Cut 1: Studio / Anchor

images/c1.jpg (GPT Image 2.5 sunburst)💎2.75
clips/c1.mp4 (Seedance 2.5 / 9s)💎108

Line (Akari (female, calm) / eleven_v4 / 8.08s)
[calm, professional news anchor] News now. At around 2 p.m. today, a bear was spotted near Tsutenkaku Tower in Naniwa Ward, Osaka.

Image prompt
Frame grab from a Japanese evening television news broadcast. A calm Japanese female news anchor in her early 40s, short neat dark bob hair, navy blazer over a white blouse, sits behind a glossy news desk with a few paper script pages in front of her, medium shot from chest up, centered, looking straight into the camera, lips slightly parted mid-sentence, composed professional expression. Background: a modern generic news studio set with soft deep-green and warm amber panels and a large out-of-focus screen showing an abstract city at dusk. Even, flat broadcast studio lighting, realistic skin texture, shot on a studio broadcast camera, slightly compressed broadcast video quality, natural moderate depth of field. Absolutely no text, no captions, no logos, no channel bugs, no watermarks anywhere in the image.
Video prompt
Locked-off studio camera, Japanese evening TV news broadcast. The female news anchor at the desk reads the news straight to camera, speaking exactly the Japanese line from the audio reference with precise natural lip sync, calm professional delivery, small natural head movements and blinks, glances down at her script once briefly, hands resting on the papers. Background city screen subtly animated. No camera movement, no zoom. Quiet studio room tone only, no music. Realistic broadcast video, no text, no captions, no logos.
2

Cut 2: Wide live-camera shot (VO)

images/c2.jpg (GPT Image 2.5 sunburst)💎2.75
clips/c2.mp4 (Seedance 2.5 / 7s)💎84

Line (Akari / eleven_v4 / 5.76s)
[calm, professional news anchor] According to police, the bear is believed to have escaped from the nearby Takoyaki Zoo.

Image prompt
[v2: c2 v1Imageを編集] Edit the first image. Keep the entire scene exactly the same: the Shinsekai street, Tsutenkaku tower, the crowd with smartphones, the two police officers, camera angle, framing, lighting and broadcast news camera look. Only replace the brown bear in the middle of the street with the bear from the second image: a chubby round bear with vivid golden yellow fur wearing a short bright red sleeveless vest, same size and position and walking pose on all fours. Realistic, matching the lighting and telephoto compression of the scene. No text, no logos, no watermarks.
Video prompt
Live news camera on a tripod with long telephoto lens, slight handheld-free micro shake and a slow small pan to follow the bear. The chubby golden yellow bear wearing a short red sleeveless vest walks slowly and heavily across the Shinsekai street on all fours with a natural bear gait, belly swaying, briefly sniffing the ground. The red vest stays on and keeps its shape. Police officers keep their arms spread holding back the crowd; onlookers in the foreground raise smartphones and shift nervously. Tsutenkaku tower stays fixed in the background. The audio reference is an off-screen female news anchor voice-over; nobody on screen speaks or lip-syncs. Ambient sound: distant crowd murmur, faint police radio, street noise. Realistic broadcast video, no text, no captions, no logos.
3

Cut 3: Viewer's smartphone video

images/c3.jpg (GPT Image 2.5 sunburst)💎2.75
clips/c3.mp4 (Seedance 2.5 / 5s)💎60

Line (Haruta PVC (JA) / eleven_v4 / 3.84s)
[surprised] Wait, whoa—it's a bear! [excited] It's actually a bear! And it's yellow!

Image prompt
[v2: c3 v1Imageを編集] Edit the first image. Keep the entire scene exactly the same: the narrow kushikatsu alley, red lanterns, signs, beer crate, the two startled pedestrians, vertical smartphone framing, camera angle and amateur phone video look. Only replace the brown bear with the bear from the second image: a chubby round bear with vivid golden yellow fur wearing a short bright red sleeveless vest, same size, position and mid-stride walking pose on all fours, head turned slightly toward the camera. Realistic, matching the scene's lighting and phone camera noise. No text, no UI, no logos, no watermarks.
Video prompt
Vertical amateur smartphone footage, shaky handheld, the bystander filming steps back a little and jerkily re-frames to keep the bear in view, brief autofocus hunting. The chubby golden yellow bear wearing a short red sleeveless vest lumbers across the narrow kushikatsu alley from left to right on all fours with a natural heavy bear gait, then turns its head toward the camera for a moment. The red vest stays on and keeps its shape. The two pedestrians in the background step back startled. The audio reference is the excited voice of the person holding the phone, off-camera; nobody on screen speaks. Ambient sound: alley chatter, a startled shout in the distance, phone-mic wind noise. Realistic phone video, no text, no UI, no logos.
4

Cut 4: Street interview 1 (local lady)

images/c4.jpg (GPT Image 2.5 sunburst)💎2.75
clips/c4.mp4 (Seedance 2.5 / 7s)💎84

Line (Aiko (Kansai, middle-aged female) / eleven_v4 / 6.08s)
[surprised] I was so surprised. It was yellow, so I thought it was a stuffed animal—[laughs] but it was the real thing!

Image prompt
Frame grab from a Japanese TV news street interview (man-on-the-street). A friendly Osaka woman in her early 60s with short permed reddish-brown hair, wearing a leopard-print blouse and a light cardigan, holding a shopping bag, stands on a busy old-fashioned downtown Osaka shopping street, medium close-up from chest up, positioned slightly right of center, looking just off-camera to the left toward an interviewer, mid-speech with an amused surprised expression, hand near her chest. The tip of a plain grey foam-covered handheld microphone without any logo enters from the lower left of the frame. Background: blurred colorful restaurant facades and lanterns, passersby, signs illegible. Natural afternoon daylight, broadcast ENG camera look, realistic skin texture, slightly compressed broadcast video quality, not cinematic. No text, no logos, no captions, no watermarks.
Video prompt
Handheld TV news street interview, gentle natural camera sway. The Osaka woman in the leopard-print blouse speaks exactly the Kansai-dialect Japanese line from the audio reference to the interviewer just off-camera, with precise natural lip sync, animated expressive face, a surprised look then a laugh, hand patting her chest and a small wave of her hand. The interviewer's grey microphone stays in the lower left of frame. Passersby walk in the blurred background. Ambient sound: busy shopping street murmur. Realistic broadcast video, no text, no captions, no logos.
5

Cut 5: Street interview 2 (young man)

images/c5.jpg (GPT Image 2.5 sunburst)💎2.75
clips/c5.mp4 (Seedance 2.5 / 6s)💎72

Line (Riku (Kansai, young male) / eleven_v4 / 5.28s)
I was eating kushikatsu, and it just walked right past me… [laughs] I did a double take.

Image prompt
Frame grab from a Japanese TV news street interview (man-on-the-street). A Japanese man in his early 20s with short black hair, wearing a grey hoodie under a dark jacket and holding a wooden skewer, stands on an old-fashioned downtown Osaka shopping street lined with small kushikatsu restaurants, medium close-up from chest up, positioned slightly left of center, looking just off-camera to the right toward an interviewer, mid-speech with a slightly stunned half-smile. The tip of a plain grey foam-covered handheld microphone without any logo enters from the lower right of the frame. Background: blurred red lanterns, restaurant facades, passersby, signs illegible. Natural afternoon daylight, broadcast ENG camera look, realistic skin texture, slightly compressed broadcast video quality, not cinematic. No text, no logos, no captions, no watermarks.
Video prompt
Handheld TV news street interview, gentle natural camera sway. The young man in the grey hoodie speaks exactly the Japanese line from the audio reference to the interviewer just off-camera, with precise natural lip sync, slightly stunned then a short embarrassed laugh, he gestures with the kushikatsu skewer to his side as if showing where the bear walked past. The interviewer's grey microphone stays in the lower right of frame. Passersby walk in the blurred background. Ambient sound: busy restaurant street murmur. Realistic broadcast video, no text, no captions, no logos.
6

Cut 6: Live field report

images/c6.jpg (GPT Image 2.5 sunburst)💎2.75
clips/c6.mp4 (Seedance 2.5 / 9s)💎108

Line (Masafumi (male, crisp) / eleven_v4 / 7.68s)
[serious, field reporter] Police and zoo staff are still searching for the bear. Residents in the area are advised to stay indoors.

Image prompt
Frame grab from a live Japanese TV news field report. A Japanese male field reporter in his mid 30s, short neat hair, dark navy suit with a plain dark windbreaker over it, holding a plain black handheld microphone without any logo, stands facing the camera, medium shot from waist up, slightly left of center, serious focused expression, mid-sentence. Behind him: a plain yellow barrier tape stretched across the street with no readable text, a Japanese police patrol car with its red rotating roof light glowing, a few uniformed police officers and zoo staff in green work uniforms, and in the background down the street the steel lattice Tsutenkaku tower in Osaka's Shinsekai district. Late afternoon overcast daylight, broadcast ENG camera look, realistic, slightly compressed broadcast video quality, not cinematic. No text, no logos, no captions, no watermarks.
Video prompt
Live TV news field report, handheld camera on the cameraman's shoulder with slight natural sway. The male reporter holding the microphone speaks exactly the Japanese line from the audio reference straight to camera with precise natural lip sync, serious urgent but composed delivery, a short glance back over his shoulder toward the police line and then back to camera. Behind him the police car's red rotating light flashes, officers and zoo staff in green uniforms move around behind the yellow tape, Tsutenkaku tower in the background. Ambient sound: police radio chatter, distant siren, street noise. Realistic broadcast video, no text, no captions, no logos.
7

Cut 7: Disclaimer card (2 s)

Summary

What we made

10ImagesGPT Image 2.5
7Voice linesElevenLabs v4
8VideosSeedance 2.5
2Final videosffmpeg
~43secondsper video

Generation time & credits Higgsfield MCP

💎687.5Credits usedImages 27.5 + Videos 660≈ $30.46
⏱11m 29sActual waiting timeTotal wait, with jobs run in parallel
18Generation jobsImages 10 / Videos 8

Version 1

💎535.25 ≈ $23.72

  • 7 images: ⏱1m 26s (bear first → other 6 in parallel)
  • 6 videos: ⏱5m 07s (6 in parallel)

Version 2 (remake only)

💎152.25 ≈ $6.75

  • 3 images: ⏱1m 07s (bear first → 2 in parallel)
  • 2 videos: ⏱3m 49s (2 in parallel)

💴 1 credit = ¥7 ≈ $0.044 ($1 = ¥158) / 1 image 💎2.75 ≈ $0.12 / 1 s of video 💎12 ≈ $0.53

Pricing: GPT Image 2.5 (sunburst, high, 2K) = 💎2.75 per image / Seedance 2.5 (1080p) = 💎12 per second. Generation times are measured from each job's start to finish.

📋 Breakdown by job (18)
Ver.TypeItemModelCreditsUSDTime
v1ImageBear referenceGPT Image 2.5💎2.75$0.12⏱33 s
BothImageCut 1GPT Image 2.5💎2.75$0.12⏱35 s
v1ImageCut 2GPT Image 2.5💎2.75$0.12⏱46 s
v1ImageCut 3GPT Image 2.5💎2.75$0.12⏱52 s
BothImageCut 4GPT Image 2.5💎2.75$0.12⏱38 s
BothImageCut 5GPT Image 2.5💎2.75$0.12⏱37 s
BothImageCut 6GPT Image 2.5💎2.75$0.12⏱37 s
BothVideoCut 1 (9 s)Seedance 2.5💎108$4.78⏱4m 41s
v1VideoCut 2 (7 s)Seedance 2.5💎84$3.72⏱4m 49s
v1VideoCut 3 (5 s)Seedance 2.5💎60$2.66⏱4m 16s
BothVideoCut 4 (7 s)Seedance 2.5💎84$3.72⏱3m 45s
BothVideoCut 5 (6 s)Seedance 2.5💎72$3.19⏱2m 57s
BothVideoCut 6 (9 s)Seedance 2.5💎108$4.78⏱3m 31s
v2ImageBear referenceGPT Image 2.5💎2.75$0.12⏱32 s
v2ImageCut 2GPT Image 2.5💎2.75$0.12⏱35 s
v2ImageCut 3GPT Image 2.5💎2.75$0.12⏱34 s
v2VideoCut 2 (7 s)Seedance 2.5💎84$3.72⏱3m 49s
v2VideoCut 3 (5 s)Seedance 2.5💎60$2.66⏱3m 02s

Final videos

Version 1: Kamome TV / Tennoji Zoo
Version 2: Naniwa TV / Takoyaki Zoo / red vest

Outputs for each cut

Image GPT Image 2.5Voice ElevenLabs v4Video Seedance 2.5Finish ffmpeg
1Studio8.6 s
Image 💎2.75 ($0.12), ⏱35 s
Cut 1Image
Voice

“News now. At around 2 p.m. today, a bear was spotted near Tsutenkaku Tower in Naniwa Ward, Osaka.”

Video 💎108 ($4.78), ⏱4m 41s
With captions
2Wide shot of the scene6.4 s
Image 💎2.75 ($0.12), ⏱46 s
Cut 2Image
Voice

“According to police, the bear is believed to have escaped from the nearby Tennoji Zoo.”

Video 💎84 ($3.72), ⏱4m 49s
With captions
3Viewer's smartphone video4.6 s
Image 💎2.75 ($0.12), ⏱52 s
Cut 3Image
Voice

“Wait, whoa—it's a bear! It's actually a bear! And it's yellow!”

Video 💎60 ($2.66), ⏱4m 16s
With captions
4Street interview ①6.5 s
Image 💎2.75 ($0.12), ⏱38 s
Cut 4Image
Voice

“I was so surprised. It was yellow, so I thought it was a stuffed animal—but it was the real thing!”

Video 💎84 ($3.72), ⏱3m 45s
With captions
5Street interview ②6.0 s
Image 💎2.75 ($0.12), ⏱37 s
Cut 5Image
Voice

“I was eating kushikatsu, and it just walked right past me… I did a double take.”

Video 💎72 ($3.19), ⏱2m 57s
With captions
6Live field report8.9 s
Image 💎2.75 ($0.12), ⏱37 s
Cut 6Image
Voice

“Police and zoo staff are still searching for the bear. Residents in the area are advised to stay indoors.”

Video 💎108 ($4.78), ⏱3m 31s
With captions

Every generated image

Bear reference (v1)
Bear reference (v1)
Cut 1 (v1)
Cut 1 (v1)
Cut 2 (v1)
Cut 2 (v1)
Cut 3 (v1)
Cut 3 (v1)
Cut 4 (v1)
Cut 4 (v1)
Cut 5 (v1)
Cut 5 (v1)
Cut 6 (v1)
Cut 6 (v1)
Bear reference (v2)
Bear reference (v2)
Cut 2 (v2)
Cut 2 (v2)
Cut 3 (v2)
Cut 3 (v2)

Movie trailer

Haruta & the Bear

The story of a bear who escaped from the zoo and a boy named Haruta / ~30 s

💎253.75Credits usedImages 13.75 + Videos 240≈ $11.24
⏱6m 00sActual waiting timeTotal wait, with jobs run in parallel
9Generation jobsImages 5 / Videos 4

The two leads

Haruta
Haruta (an original AI-generated boy)💎2.75 ($0.12), ⏱37 s
Bear
The bear (referenced from the Version 2 bear image)

Outputs for each cut

Image GPT Image 2.5Video Seedance 2.5Voice ElevenLabs v4 (narrator: Koichi Yashiro / Haruta: Toru)
1A bear appears in Osaka
Image
A bear appears in Osaka
Video

Narration

“That day, a single bear wandered into the streets of Osaka.”

Reuses the news footage (Version 2, cut 2) with a film-style grade and letterbox

2The meeting
Image 💎2.75 ($0.12), ⏱42 s
The meeting
Video 💎72 ($3.19), ⏱3m 07s

Narration

“Everyone ran away… everyone but one boy.”

Haruta and the bear lock eyes in an alley at dusk

3Racing through Shinsekai at night
Image 💎2.75 ($0.12), ⏱45 s
Racing through Shinsekai at night
Video 💎48 ($2.13), ⏱3m 27s

(music only)

The two dash through neon and lanterns

4A promise in the rain
Image 💎2.75 ($0.12), ⏱42 s
A promise in the rain
Video 💎60 ($2.66), ⏱4m 38s

Haruta

“It's okay. I'll get you back home.”

Haruta shields the bear (lip-synced to his line)

5Rooftop at sunset
Image 💎2.75 ($0.12), ⏱43 s
Rooftop at sunset
Video 💎60 ($2.66), ⏱4m 09s

Narration

“Their small, grand adventure… begins.”

The two side by side on a rooftop overlooking Tsutenkaku

6Title
Image
Title
Video

Narration

“Haruta & the Bear. In theaters this summer.”

Title “Haruta & the Bear” + disclaimer

The music is the orchestral audio Seedance 2.5 generated with each clip, automatically ducked under the narration.

×
© 2026 Haruta — AI活用ガイド