Pular para o conteúdo
← Back to Skalablog

Published article

AI Video Kaise Banaye: 5-Minute Mobile Workflow

ChatGPTGeminiOpenAI

AI video kaise banaye is a file-format question, not a prompt question. Text becomes still images, still images become animated clips, and clips become a timeline. The 2026 walkthrough examined here shows that chain across mobile apps, including the credit limit the video does not dwell on.

AI video kaise banaye on a mobile phone?

AI video kaise banaye on a phone: generate a story script, render each character as a still image, animate one short clip per scene, then join those clips in an editor. This article dissects a 16-minute Hindi walkthrough published on 16 September 2026 by the YouTube channel Creator Search 2.0, which demonstrates a "saas-bahu" drama format built from five consumer mobile apps.

The workflow needs no desktop machine and no paid plan. It does, however, have quirks the video only hints at, including a credit cap and a face-consistency trap that decides whether your characters look the same from scene to scene.

Independent evidence on AI video as a production method remains thin. The one large-scale result quoted in this article is a 2024 study reporting higher average quality scores for AI-assisted short-video production, and that result does not establish anything about monetization, reach, or revenue.

The demo channel used to prove the format, AI Katha, is shown in the video with more than 4 lakh subscribers and individual videos above 4 crore, 1.5 crore, and 1.5 crore views. Those totals might tempt you to treat this as a puzzle to solve, but they are the channel's numbers, not a promise about yours.

The five apps and what each one does

The described pipeline splits work across five apps, and each owns exactly one job. Substituting one app without checking its input and output format breaks the chain, because every step consumes the previous step's file.

The division matters more than the specific brands: a chat model writes text, an image model draws stills, a video model animates them with dialogue audio, an editor joins the clips, and the upload happens last.

AppRoleInput it consumesOutput it produces
ChatGPT (OpenAI)Writes the story and scene breakdownA master prompt plus one word, "start"10 story premises, then a scene-by-scene script with dialogue
Gemini (Google)Draws one still image per characterA character description copied from ChatGPTOne saved image per character in the phone gallery
Flow (Google Labs)Animates each sceneOne character image, its dialogue, and a video promptOne 6 to 10 second vertical clip per scene
VNJoins and trims the clipsThe downloaded clip files, opened in scene orderOne 1080p 30fps vertical video
YouTube Create (alternative)Same animation step, done in the YouTube appThe same character image plus dialogueA clip you can download from the app

Step 1: generate the story script in ChatGPT

The video opens ChatGPT, OpenAI's chatbot application, pastes a "master prompt" copied from the video description, and answers the single word "start". The bot returns ten short story premises in the Indian family-drama genre, and the creator picks idea number seven.

Typing a bare number such as "7" returns that story broken down scene by scene with dialogue. The transcript says the chat delivered 26 scenes for the chosen story. If none of the first ten premises fit, typing "and 10 stories" (the literal on-screen instruction) asks for another batch, which repeats indefinitely.

The master prompt itself is not reproduced in the transcript, only referenced as a link, so treat the exact wording as unavailable from this source. The structural lesson survives without it:

  1. Ask for premise options first, so you can reject weak ideas before spending credits on them.
  2. Expand one premise into numbered scenes second, so the script arrives in the order you will generate it.
  3. Keep the dialogue inside each scene entry, so the next app can copy it cleanly without you retyping anything.

The video presents this stage as effectively unlimited, which is accurate in the sense that the chat does not meter you. The constraint appears later, at the animation step, where each scene costs a generation.

Step 2: turn each character into a consistent image

Gemini, Google's AI assistant app, generates one still image per character, which is the step that decides visual consistency. The video copies the character description out of ChatGPT, switches to Gemini, opens the image-creation tool from the plus icon, pastes, and sends.

Images render in roughly two to five seconds. Each finished image is saved to the phone gallery through the save option under the image. In the demonstration there are three characters, the daughter-in-law, the mother-in-law, and a relative, and each one gets its own still.

The creator warns about a specific failure mode: if all characters are generated on the same Gemini page, the app may give every face the same features. Closing and reopening the app between characters, as shown, produces distinct faces. Check each render before saving rather than trusting the first batch.

This is also the cheapest place to fail. A bad face costs you one image render and about five seconds. The same mistake discovered after animation costs a video generation, which is the scarce resource in this workflow.

If you cannot install Gemini or ChatGPT, both are on the Play Store, as is VN. The transcript's one slip here is that it calls the editor KineMaster in a single sentence while the on-screen work is done in VN throughout; treat KineMaster as a naming error, not a second editor.

Step 3: build one video clip per scene in Flow

Flow, Google's AI filmmaking tool at labs.google/flow, animates each scene from a character image plus its dialogue and video prompt. The settings shown are aspect ratio for vertical Shorts, 720p resolution, 1x speed, and a clip length of 8 seconds, which the creator ties to dialogue length.

What the Flow settings mean for your output

The upload flow adds each character image via the plus icon, then the prompt field receives the scene dialogue followed by two spaces and then the video prompt text copied from ChatGPT. Long dialogue calls for a 10-second clip; short dialogue works at 6 to 8 seconds.

Generation takes roughly 5 to 10 seconds per scene. Finished clips download from the three-dot menu as 720p media, with a 1080p option also present. The same result can be produced in the YouTube Create app and downloaded from there, according to the transcript.

This is the credit-sensitive stage, and it is where the video's "unlimited" framing needs a caveat. The creator refers to a trick for continuing through Flow and does not demonstrate its mechanics on screen, so the sustainable pace is one clip per scene within whatever quota the account holds.

The consistency trap and the scene count

Face consistency across clips comes from reusing the same saved character images, not from retyping descriptions. Before generating each new scene, the creator returns to the ChatGPT tab, checks which characters appear in that scene, and uploads only those image files.

Scene two in the demo contains both the mother-in-law and the daughter-in-law, so both files are uploaded. Scene three repeats the pair. A scene with one character gets one file. Uploading an image the scene does not need wastes context; leaving out one it does need makes the model invent a second face mid-scene.

The scene count scales with the target runtime. The transcript's rule of thumb is that a 30-second video lands around 8 or 9 scenes, and a longer video simply means more clips generated and downloaded one at a time. There is no batch step in this workflow, which is why the promised five-minute turnaround is best read as the time spent on the steps themselves rather than the time you will spend waiting.

Step 4: merge and trim the clips in VN

VN, a free multi-track video editor, assembles the exported clips into one vertical video. A new project imports scene one, scene two, and scene three in order, the canvas ratio is switched to 9:16, and the timeline is trimmed.

The edit itself is minimal. A short tail of silence at the end of most generated clips, which the transcript calls the extra non-audio part, gets split off and deleted. This is the one edit that every scene needs; optional crossfade transitions between clips make the cuts less abrupt.

Export settings in the demonstration are 1080p at 30fps, with 720p available. The rendered file lands in the phone gallery ready for YouTube or Instagram. If you want a crossfade, tap the icon between two clips on the timeline and pick a short transition so the viewer does not notice the join.

The scoring study behind the format

A 2024 study in the journal Computers tested whether AI tools raise the quality of short-video production by tourism students. The AI-assisted group scored higher on average across the study's scoring rubric, which is evidence about production quality in a classroom setting, not a forecast about views or revenue.

The monetization claim in the source video rests on a subscriber count and view totals shown in a screen recording. That is one channel's result, not a rule, and it should not be read as a guarantee that the same format will monetize your channel.

Tool roles, compared

Each tool in this workflow owns one function, and mixing them up wastes the most time. The table below maps the role, the input it consumes, and the output it produces, based only on what the walkthrough demonstrates on screen.

ToolJobCost signalSwap for
ChatGPTScript and scene listFree tier is enough for textAny chat model that can hold numbered scenes
GeminiCharacter stillsFree tier, a few seconds per imageAny image generator with a save option
FlowClip generationCredit-limited, the bottleneckYouTube Create, same images and dialogue
VNTimeline assemblyFreeAny mobile editor with 9:16 and split
YouTube / InstagramPublishingChannel rules apply separatelyNothing in this chain

Mobile production in practice

Mobile production of this kind is increasingly used by faceless channel operators, and the appeal is obvious: one phone, no studio, no on-camera presence. The catch is that the entire visible result depends on the image step. If your character images look inconsistent, your video looks inconsistent no matter how clean the edit is.

That is why the recovered detail from the transcript about generating each character on a fresh Gemini session matters more than it sounds. It is not a tip, it is the mechanism that keeps the format usable across dozens of scenes.

A practical checklist before you commit a full script to Flow:

  1. Generate three or four test scenes and watch them back to back.
  2. Check that the same face returns in each test scene.
  3. Check that each clip has usable dialogue audio before you build the rest.
  4. Only then spend the rest of your credit allowance on the full scene list.

FAQ

What does "AI video kaise banaye" mean and what skills does it require?

The phrase is Hindi for "how to make an AI video" and covers any tool-driven video production workflow. This particular method needs no editing background, but it does require patience with app switching and a habit of checking generated output before saving it.

How long does one video take with this method?

The transcript claims a video can be finished in about five minutes on a phone, and that figure describes the mechanical steps, not the waiting time for image and clip generation. A finished short in the 30-second range needs roughly eight or nine scenes, each generated and downloaded separately, so expect the clock to run longer in practice.

Why do the characters' faces change between scenes?

Character drift happens when the image model is asked for every character in one conversation, which pushes it toward a single repeated face. The video's fix is to generate one character per fresh session and save each image to the gallery before starting the next one.

Can this workflow run without paying for anything?

Every app named in the walkthrough has a usable free tier, and the video presents the process as free on mobile. The Flow stage still consumes the account's generation credits, and the walkthrough never quantifies the free allowance, so treat "unlimited" as unverified.

Which app do I need first, and where do I get it?

Start with ChatGPT from the Play Store, because the master prompt and the scene list come from there. Then add Gemini for images, and VN for editing. Flow runs in the Chrome browser at labs.google/flow, so it is a website rather than an app.

What does the master prompt actually do?

It sets up the genre and the output shape. The transcript does not reproduce its wording, only links to it, but the behaviour it produces is visible: ten premises, then a numbered scene list with dialogue attached.

Do I need to record my own voice?

No. The dialogue is part of the prompt sent to Flow, and the generated clip carries audio. That is why the clip length is chosen to match dialogue length rather than your editing rhythm.

What should be checked before uploading the finished video?

Confirm that every scene uses the same saved character images, that no silent tail remains at the end of each clip, and that the exported file is vertical 9:16. Then upload through normal channel publishing, noting that platform monetization rules are separate from production questions.

Does the demo prove the format can be monetized?

No. The video shows a monetized channel as proof of concept, which is one example. The 2024 study supports a quality claim in a classroom setting, not a revenue claim, and neither source measures your channel.

What to take from this walkthrough

The durable part of this workflow is its segmentation: script first, stills second, clips third, assembly last. That order lets a weak step be fixed cheaply, because a bad face costs one image render instead of a regenerated scene.

The order is also what makes the format teachable. If you were shown only the finished video, you would have no way to tell which step produced a given artifact. The chain from ChatGPT to Gemini to Flow to VN is the actual lesson, and the specific prompt is replaceable.

Mobile production of this kind is increasingly used by faceless channel operators, but the video's monetization claims rest on one channel's subscriber count and view totals shown in a screen recording, which is an example rather than a general rule. Publish on your own judgment about what the platform and audience will accept.

If you have a long YouTube walkthrough of your own, Skala Blog turns the spoken steps into a structured written article. Paste the video URL, let it transcribe the audio, and edit the generated draft into your own article. Channels that already explain a workflow out loud, the way Dev Doido do canal do youtube does, can turn that explanation into a readable reference instead of leaving it buried in a 16-minute video.

Skala Blog

Source video

More on AI tools and workflows