8 steps across 3 phases — from finding paying local clients, to concepting a short film, to shipping your first edited draft, to inserting yourself into a fantasy scene. Ready-to-copy prompts, direct links, zero fluff.
3Phases
8Steps
100%Plug & Play
The problem: You know how to make AI Videos, but you don't have a pipeline of local clients willing to pay for them.
Open ChatGPT and paste the research prompt below, replacing the locality/city placeholder.
Let it pull the top 15 local jewellery stores (or any niche you target) with a website and product photos.
Review the tabular output: name, phone number, website, best outreach platform, and a ready script.
Start outreach with the highest-fit prospects first — the ones already running weak or outdated websites.
Ready-to-copy prompt
Hey Chatgpt, Assume you are a world class level of research analyst.
I want you to research through google my business and find me the top 15 local jewellery stores near me who have their website listed with some jewellery pics.
Further I want you to also find their phone numbers, website link, name, best platform and method with script to approach them and list it all to me in a tabular format. I live in [LOCALITY, CITY].
Swap "jewellery stores" for any local niche — real estate, salons, restaurants — to build a fresh prospect list in minutes.
The problem: Blank page syndrome. You need a 30-second brand video concept that's tight, emotional, and production-ready — fast.
Open the link above (or a normal ChatGPT chat if it doesn't load).
Paste the prompt below, rewritten for your specific brand/use-case.
Answer the top 5 questions it asks — each will come with copy-paste-ready options.
Receive your finished 30-second concept: a goosebumps-level narration with no on-screen dialogue.
Ready-to-copy prompt
Assume you are a world class level of creative director and a visual story teller. Who is extremely good at prompting as well. Help me craft a 30 seconds Brand Video Storyline for a brand.
I want you to come up with an extremely short and beautiful concept that can help my brand accomplish its goal. Make sure you do not generate stories which include minors or kids.
REMEMBER there should be a VoiceOver. The Voiceover should be in the form of a goosebumps-level {NARRATION} version. Rest visuals should consist of NO DIALOGUES between people talking to each other.
You are allowed to ask me only top 5 questions that can help you to craft the best story possible. Once you have answer to those 5 questions give me the Concept Directly.
You can ask Question like, What is my Brand name? How do I want the Voiceover to be like?, Style of video (3D, HyperRealistic etc). Make sure you give options to each question, so that i can copy and paste.
The tighter your 5 answers, the sharper the concept — write real brand details, not placeholders.
The problem: You have a script, but generation models need structured, consistent JSON prompts per scene — not prose.
In the same chat where you already discussed your script, paste the prompt below.
Let it work out how many videos/plots the scene requires, then generate a JSON prompt per plot.
Answer the top 3 clarifying questions it asks about the storyline.
Receive a tabular set of JSON prompts — one per scene — plus the complete voiceover at the end.
Feed each JSON prompt into VEO 3.1 to generate the visuals.
Ready-to-copy prompt
Assume you are a world class level of creative director, visual story teller & a Prompt Writer. You get desired output of a video that stuns the viewers with its quality and extreme attention to the details of video generation.
Study my entire script that WE HAVE ALREADY DISCUSSED ABOVE. First I want you to understand the scene and then plot how many videos we require. Then craft JSON Prompts for each plot under scene.
I will give it to VEO 3.1 to get stunning visual videos. Just make sure you maintain the same characters, actors, faces, fonts, music theme throughout the AI Video. color grade & tonality of the characters should be consistent. (Do not give me a separate master prompt or such for consistency include that inside all of the prompts)
You can Ask me TOP 3 questions that you might need answered in order to execute the storyline. Generate JSON Prompt for each scene in a tabular format and give it to me. At the end attach the complete VoiceOver.
Keep this in the same thread as Step 2 so the model retains full context of your concept and brand.
The problem: Generating scene-by-scene images that stay character-consistent across an entire storyboard is hard without a system.
Paste the context-setting prompt below to prime the model as a world-class storyteller/prompt designer.
When it asks for your Scene 1 JSON prompt, paste it in.
Go to ChatGPT and generate the frame from that JSON prompt.
Download the final image for that scene.
Paste your next scene's JSON prompt and repeat — the model keeps character DNA, camera metadata, and cinematic style consistent frame to frame.
Repeat until the entire storyboard is complete.
Ready-to-copy prompt
Now, Assume you are a world class level of creative story teller who very well knows designing prompts and getting the right output from image generation models like Gemini.
I want you to help me generate images that reveal the entire storyline frame by frame. So, I can create a Visual StoryBoard for my AI Video. I am going to give you JSON prompt for each scene, one after another. I want you to keep understanding the storyline and the give me character consistent frames accordingly.
Once you have understood it. I want you to ask me for a JSON Prompt of Scene 1. You can take my JSON Prompt convert it into the BEST Possible Image & Then. I will simply keep on pasting the next scenes and i want you to follow the same process, that you take the JSON Prompt and convert it into the next frame of the story
Important (How you will create the frames) — Every image you give should have:
- Consistent character DNA (all characters/products/environments involved)
- Consistent Cinematic camera metadata (lens, lighting, depth, grain)
- Output that matches my script's style of video generation (NOT AI-looking/slop)
If you are ready, Ask me for my first JSON Prompt?
Never start a new chat mid-storyboard — character consistency depends entirely on staying in the same thread.
The problem: You have static frames — now they need to become an animated, timeline-ready sequence.
Method 1 — Lyria by Google Gemini (Beginners, VoiceOver + Music)
Pros: free, copyright-free, fast. Cons: capped at 30 seconds, voiceover quality can drift.
Go to the Audio Engine inside Google Vids.
Click Music.
Copy and paste your entire voiceover script.
Hit Generate.
Download the track, or import it straight into your sequence. Deselect the music track if you only want the voiceover.
Ready-to-copy prompt (Lyria)
Assume you are a world class level of emotional music creator who is also very well versed with prompting and can write prompts for Google Gemini (Lyria) To help me generate the perfect music with voiceover narrative that can help my AI Video acheive it's goal. Here is the lyrics of the AI Video. {PASTE YOUR LYRICS, here}
Continuation prompt (for extending the track):
Continue in the same tone music & tonality. the lyrics attached below.
____________________
Method 2 — Suno (Technical creators, full song)
Pros: strong with local languages (Hindi, Tamil, Telugu, etc.), goosebump-level output. Cons: generates long audio (4-5 min) on the free tier.
Go to Suno and sign up / log in.
Click Create → Custom.
Paste your lyrics.
Add your style tags and hit Create.
Download and save your favorite generation.
Method 3 — Minimax (Intermediate, more control)
Pros: more control than Suno, can generate voiceovers too. Cons: weaker with local-language voiceovers.
Go to Minimax's audio engine and sign up / log in.
Click Create → Custom.
Paste your lyrics.
Add your style tags and hit Create.
Download and save your favorite generation.
We generally recommend generating voiceover + music under one app when possible — it saves an entire sync pass in editing.
The problem: Turning a folder of generated clips and a music track into a coherent first draft — without needing Premiere or DaVinci.
Note: Avoid DaVinci, Premiere Pro, Adobe Suite, or other advanced tools — keep it simple for the first draft.
Workflow steps
Cut unwanted elements out of the story.
Align the music and visuals to the beats.
Master the art of J-cuts / L-cuts — watch the tutorial linked below.
Add the desired effects, transitions, texts, and titles.
Export and share the video within your group before 6pm.
Style-match prompt:
Maintain the same style and font and alignation
instead of XXXXX, write there
______YYYYY_______
1920*1080
Animation-idea prompt:
Write me a prompt to animate this in a 6 sec cinematic title for my AI Video. Based on the photo suggest me top 3 ideas on how we can create the best cinematic title for my AI Video. Then ask me to select one idea, later create a detailed JSON prompt for it.
Ship the simple first draft before chasing the advanced homework — a finished basic edit beats an unfinished perfect one.
The problem: You want to see yourself starring in a cinematic wizarding-school style scene — full 30-second shot, scored and edited — without touching a camera or an editing suite.
Tools: Gemini (Nano Banana) for compositing, Google Vids' AI Video → Animate (powered by Veo 3) for turning stills into clips, Suno or Lyria for score, then CapCut or Google Vids again for the final edit/assembly.
IP note: Named characters, houses, crests, and film scores are copyrighted — models will often refuse exact-likeness prompts anyway. This workflow uses a generic "wizarding academy" aesthetic (own names/marks) so it generates cleanly and is safe to actually share. Swap in your own house colors/crest as you like.
Step 1 — Reference photo
Use a well-lit, front-facing photo of yourself, neutral expression, plain background — this is your character-DNA anchor.
Optional: a second photo in the rough pose you want for your final shot helps consistency.
Step 2 — Composite yourself into the world
Tool: Gemini only (this is your practice run before the real 4 beats)
Upload your photo plus a reference image of the setting/robe style you want.
Use the compositing prompt below and iterate 2-3 times until the face holds up close.
Take the person in image 1 and place them into a magical wizarding-school setting — stone castle corridor, candlelit, house robes, wand in hand. Keep their face, hair, and identity exactly consistent with image 1. Cinematic lighting, 35mm lens look, shallow depth of field, film grain. Output 1920x1080.
Once this test composite looks right, stay in this same Gemini chat — you'll reuse it for all 4 "FRAME" prompts in Step 3 below so the face stays locked.
Step 3 — Lock character DNA, build 4 beats
Stay in one continuous Gemini chat for all 4 "FRAME" prompts, so the face/outfit stay locked from one still image to the next. Each beat has two prompts, used in two different tools, in this order:
FRAME prompt → Gemini. Generates the still image for that beat. Approve or re-roll until the face is consistent, then download it.
ANIMATE prompt → Google Vids. Sign in to Google Vids → click "AI Video" → click "Animate" → upload the still image you just downloaded from Gemini as the starting frame → paste the ANIMATE prompt to turn it into a moving 6-8 second clip (this runs on Veo 3 behind the scenes).
Do this twice per beat (Gemini, then Google Vids) before moving to the next beat.
Beat 1 — Establishing shot (8s)
1a — Paste into Gemini to generate the still frame
{
"shot": "beat_1_establishing",
"subject": "The person from the reference photo, face and identity kept exactly consistent",
"wardrobe": "Dark academy robes with a subtle crest, deep green and bronze trim",
"setting": "Grand candlelit stone hall, floating candles, long wooden tables, vaulted ceiling",
"action": "Standing at the entrance archway, looking up in awe, wand held loosely at their side",
"camera": "Wide shot, low angle, 24mm lens, slight upward tilt",
"lighting": "Warm candlelight key, cool moonlight rim from tall windows, volumetric haze",
"grain": "Kodak Portra 400, 5% grain, cinematic color grade, teal-orange",
"style": "Hyper-realistic, NOT AI-looking, matches live-action fantasy film still"
}
1b — Google Vids: sign in → AI Video → Animate → upload image → paste this to animate it
{
"duration": "8s",
"camera_move": "Slow dolly-in from the archway toward the character",
"action": "Character steps forward into the hall, robes swaying, head turning slowly to take in the room",
"audio": "Distant echo of chatter, candle crackle, soft ambient orchestral swell begins",
"fps": 24,
"lens_look": "anamorphic, subtle flare from candlelight"
}
Beat 2 — Close-up resolve (6s)
2a — Paste into Gemini (same chat, so it stays consistent with Beat 1)
{
"shot": "beat_2_closeup",
"continuity": "SAME character, wardrobe, and lighting style as beat_1",
"action": "Close-up on face, a determined breath, gripping the wand tighter",
"camera": "Medium close-up, 85mm lens, shallow depth of field, background softly blurred hall",
"lighting": "Same warm/cool candlelight balance as beat_1",
"grain": "Match beat_1 exactly"
}
2b — Google Vids: AI Video → Animate → upload image → paste this
{
"shot": "beat_3_spellcast",
"continuity": "SAME character DNA and wardrobe as beats 1-2",
"action": "Character raises wand, tip glowing with golden-white light, small sparks arcing outward",
"camera": "Dynamic three-quarter angle, 35mm lens, slight low angle for power",
"lighting": "Glow from wand tip becomes the dominant light source, rim-lighting the character's face",
"grain": "Match previous beats",
"vfx": "Practical-looking spark/light particles, not overly digital or glossy"
}
3b — Google Vids: AI Video → Animate → upload image → paste this
{
"duration": "8s",
"camera_move": "Slight handheld sway, quick push toward the wand tip as it ignites",
"action": "Spell releases - a burst of light shoots upward, illuminating the ceiling briefly, candles flicker in response",
"audio": "Sharp magical whoosh, crackle of energy, orchestral hit on the release, hall gasps faintly in background",
"fps": 24,
"lens_look": "anamorphic flare bursts on the light release"
}
Beat 4 — Payoff, wide finish (8s)
4a — Paste into Gemini (same chat)
{
"shot": "beat_4_finish",
"continuity": "SAME character, wardrobe, hall setting as all prior beats",
"action": "Character lowers wand slightly, smiling, light from the spell fading around them, other students turning to look",
"camera": "Wide shot pulling back to reveal the full hall reacting",
"lighting": "Return to the beat_1 candlelight balance as the glow fades",
"grain": "Match previous beats"
}
4b — Google Vids: AI Video → Animate → upload image → paste this
{
"duration": "8s",
"camera_move": "Slow pull-back / crane up, revealing the full hall",
"action": "Glow fades to ambient candlelight, character stands a little taller, hall returns to murmuring chatter",
"audio": "Orchestral swell resolves to a warm final chord, ambient hall chatter returns, soft candle crackle",
"fps": 24
}
8s + 6s + 8s + 8s = 30 seconds exactly, across 4 Gemini stills and 4 Google Vids/Veo 3 clips. That's the whole scene, pre-planned.
Step 4 — Score the scene
Tool: Suno or Lyria (pick one — don't need both) — one continuous instrumental cue, matched to the emotional arc of your 4 beats, not four separate clips.
Instrumental orchestral cue, 30 seconds, cinematic fantasy score.
EMOTIONAL ARC (follow this exact structure):
0:00-0:08 - AWE: Soft, wide, wondrous. Solo french horn or harp arpeggios over sustained string pad. Slow tempo, spacious reverb, like walking into something magical for the first time.
0:08-0:14 - TENSION: Strings tighten, low cello/bass drone enters underneath. A single repeating pulse (like a heartbeat) builds quietly. Dynamics pull back - quieter, more focused, anticipatory. No percussion yet.
0:14-0:22 - RELEASE: Sudden orchestral hit on the downbeat - brass stab, choir swell, timpani accent. Tempo picks up, strings surge upward, energetic and bright. This is the "spell releasing" moment - big, magical, triumphant burst.
0:22-0:30 - TRIUMPH/RESOLVE: Full orchestra warms into a resolving major chord progression. French horns and strings carry the main theme home. Tempo eases back down. Ends on a warm, sustained final chord with a soft harp glissando trailing off.
INSTRUMENTATION: Full orchestra - strings, french horns, harp, choir (wordless "ahh" pads), light timpani/percussion only in the release section, no electronic/synth elements.
STYLE REFERENCE: Sweeping, wonder-filled fantasy adventure score - magical academy, first-day-of-wonder tone. Original composition, not mimicking any specific existing score.
MOOD KEYWORDS: wonder, anticipation, magic, triumph, warmth
TEMPO: Starts ~70 BPM, tightens to ~85 BPM at release, settles to ~75 BPM at resolve.
Avoid naming a specific composer or film score directly — generators often refuse it, and it keeps you clear of "sounds confusingly similar to" territory if you ever post the video publicly.
Step 5 — Align the cut to the music hit
Tool: CapCut or Google Vids — bring in your 4 Google Vids/Veo 3 clips + the Suno/Lyria track. The one cut that matters: Beat 2 → Beat 3. The orchestral hit at 0:14 needs to land within a few frames of the wand igniting, or the video reads as unsynced. Build the timeline backward from that moment.
Import all 4 clips + the music track; music sits on its own audio layer.
Scrub the music and mark the exact frame the brass/choir hit lands (~0:14).
Drag Beat 3 so its first frame of visible spark lines up with that marker, nudged frame-by-frame.
Work outward: Beat 2 butts directly against the marker, Beat 1 before it, Beat 4 starts the instant the light fades.
Crossfade (0.2-0.3s) for Beat 1→2 and Beat 3→4. Keep Beat 2→3 a hard cut — the punch sells the impact.
L-cut Beat 2's ambient audio to start ~0.5s before its visual cut in.
J-cut Beat 4's orchestral swell to bleed in slightly before its visual cut in.
Duck the ambient SFX under the music during the release beat so the brass hit isn't muddied.
Export at 1080p, 24fps to match your Veo generation settings.
Play the timeline at 0.5x speed around Beat 3 — the wand's first spark should appear within half a second of the orchestral hit. If it's off by more than ~3 frames at normal speed, nudge until it locks.