CapCut text animation is built from three repeatable steps: choose a bold font, style the captions before animating, then add Notion. The instructor in the video builds custom movements with transform keyframes, opacity fades, and speed curves, then saves each Notion as a preset so the same captions can be reused on later projects.
How CapCut text animation works in three ordered steps
CapCut text animation works best in a fixed order. The instructor lays out three steps and follows them for the whole build instead of restyling captions one at a time afterward:
- Pick the font.
- Style every caption before animating.
- Animate last, using CapCut's pre-built motions plus a custom Notion layered on top.
The reason for styling first is mechanical. Font, weight, shadow and colour apply to the caption track, so changing them after animating means reopening every caption clip and repeating the same settings. Animating last keeps the keyframe work separate from the styling work. CapCut is ByteDance's video editing app for mobile and desktop.
Font choice splits in two. Body captions use a bold face; a highlighted power word can switch to a different font to stand out. The instructor names Poppins bold as the body font and drops character spacing to minus one, then adds a shadow with opacity around 30 percent.
Word-by-word captions need one text layer per word, which is the step most tutorials skip. The workflow is to build one animated word, turn it into a compound clip, duplicate it, reposition each copy, and retype the text. The video starts with easy motions and levels up through the build, and nothing is skippable because each Notion reuses the previous one's layer.
Which Notion does CapCut text animation build first?
Slide up is the first Notion built, and every other Notion in the video is a modification of it. The instructor avoids CapCut's built-in animation presets for this one and builds the movement manually with transform and opacity keyframes on a single text layer.
The build runs on a frame grid. A transform keyframe opens at the start of the layer, the playhead moves 15 frames forward, and a second transform keyframe is set. Returning to the first keyframe and dragging the text downward creates the travel; playback then slides the word up into place.
Opacity carries a second, shorter sequence. A blend keyframe drops to zero, two frames later it returns to 100, so the word fades in as it arrives rather than appearing at full strength the whole way.
Smoothing is what separates a finished Notion from a rough one. Opening the variable speed animation panel and applying cubic out to the transform keyframes removes the linear, mechanical feel, and the opacity keyframes get the same treatment.
Turning one animated word into reuse
A finished word becomes a reusable unit when it is wrapped as a compound clip, duplicated along the timeline, and retyped one word at a time. Each duplicate keeps the same keyframe timing, so the words arrive in sequence without rebuilding the Notion.
Spacing matters more than it looks. Duplicates have to sit at equal gaps on the timeline, otherwise the caption rhythm drifts. The instructor keeps a blue guide line in the centre of the canvas and aligns every word to it.
Saving the result as a preset is the step that pays off later. Right-clicking the clip and choosing save as a preset stores it in a presets folder, where it can be dragged onto any future timeline. Font, colour and size stay adjustable after insertion, so one preset covers multiple caption styles.
This is also where a creator library becomes useful. The instructor sells a pack of drag-and-drop presets called KeyBlocks through a link in the video description; the pack contains over 145 presets, and the preset workflow itself is available to anyone using CapCut without a purchase.
Slide down, left and right from the same layer
Slide down reverses the same setup: drag the first keyframe above the canvas instead of below it, then reapply cubic out to the transform keyframes. The instructor also pulls the first blue curve handle to make the word appear faster at the start of the move.
For a horizontal slide, the transform keyframes are reset rather than extended. Deactivate and reactivate the transform keyframe on the first frame to clear the stored position, set a fresh transform keyframe on the last frame, then move the word off to the side on the first frame. The horizontal slides in the video use the circle ease preset curve instead of cubic out, and the opacity keyframes are shifted slightly to the right so the fade only overlaps the part of the move that is visible on screen.
The direction is set by the position value. The instructor uses a horizontal offset of 1,495 pixels for the right-to-left entry, and a negative value of minus 1,495 to send the same Notion in from the left.
Those two numbers are specific to that project's canvas and text size. The practical rule is that the starting offset has to place the word fully off screen, so the value changes with resolution and font size.
Building the bounce and blur motions
Bounce up reuses the same layer with the keyframes reset, then layers rotation and scale changes across a short frame sequence. The instructor starts at scale 100 with a slight rotation, moves 10 frames to scale 84 with the rotation reversed, adds a five-frame step, then returns to scale 100 over three more frames.
The smoothing pass is more involved than cubic out. Selecting all transform keyframes and pressing shift plus W applies a bounce preset, then the curve handles on the second and third keyframes are dragged individually to shape the dip. In the walkthrough the third keyframe is deleted and the two remaining keyframes are re-smoothed before the dip is shaped. Editing the rotation keyframes the same way keeps the turn from snapping, and the instructor pulls the rotation curve handle to the left so the turn is fast at the start and slows into place.
A CapCut animation preset called inhale, set to 0.7 seconds, is added on top of the keyframed Notion the word scales into place while it bounces.
Blur slide and blur fade both apply blur, but the order of operations differs. For blur slide, a blur keyframe starts at full strength, moves 15 frames forward, and drops to zero. For blur fade, the compound clip gets CapCut's fade in preset at 0.8 seconds duration instead, and the blur is folded in alongside it.
Blur applied as a separate effect layer needs one extra nesting step. Selecting the effect layer and the compound clip together and creating a compound clip from both keeps the blur on the text and off the background, which matters if the background already has its own look. The video also shows a bonus combination: adding CapCut's pop up animation to a text layer, wrapping the text as a compound clip, then wrapping text and animation together for a bounce-plus-pop-up move.
Word highlight, glow blink and masked underline
Highlighting a single word uses the glow effect rather than a colour change. The instructor drags glow onto the timeline, drops its keyframe to zero at the start of the layer, raises it to 100 a few frames later, then returns it to zero three frames after that to create a blink.
Timing the blink is the fiddly part. The last glow keyframe is stretched so the highlight stays visible longer, and the middle curve handles are pushed to the right to speed the rise and fall. If the flash is too strong, lowering the value on the second keyframe softens it.
The glow effect only covers one word, so it is copy-pasted onto each additional word and re-timed against the caption. That is why the instructor builds each word as its own compound clip first.
An underline is drawn with a rectangle sticker from the shapes library, sized down, placed under the text, coloured, and then animated with a split mask. Mask keyframes move the split from left to right across a few frames, and the variable speed panel smooths the wipe.
Two settings finish the underline. Setting the rectangle to layer one puts it behind the letters, and turning off uniform scale lets width and height be adjusted separately, so the bar can be stretched to 500 in height or pulled narrow to match a thin font.
CapCut text animation FAQ
- Do you need CapCut animation presets to animate captions? No. The video builds slide up, slide down, horizontal slides and bounce with transform, scale, rotation, opacity and mask keyframes only. CapCut's built-in animations such as inhale, pop up and fade in are added on top of that keyframe work rather than replacing it.
- How many frames should a CapCut text Notion take? The builds in the video use short spans: 15 frames for the slide travel, 10 frames for the first bounce step, five and three frames for the remaining bounce moves, and two frames for an opacity return. Slower Notion needs more frames, but the timing stays in the tens of frames, not seconds.
- What does cubic out actually do? Cubic out is a speed curve inside CapCut's variable speed animation panel. It eases the Notion as it finishes, so a word decelerates into its final position instead of travelling at one constant rate. The instructor applies it to transform keyframes and to opacity keyframes on almost every Notion shown. The horizontal slides are the exception, where circle ease is used instead.
- Why convert captions into compound clips? A compound clip groups the keyframes, effects and styling of one animated word into a single unit. That grouping is what lets the instructor duplicate, retype and re-time words without rebuilding the Notion, and it keeps an effect like blur or glow bound to that text instead of the whole canvas.
- Can the finished motions be reused on another project? Yes. Right-clicking a finished clip and choosing save as a preset stores it in CapCut's presets folder, where it can be dragged onto a later timeline. Font, colour and size remain editable after insertion, so one preset works across different caption styles.
Frame values and curve choices at a glance
Every Notion in the walkthrough is a variation on the same handful of values. The table collects them so they can be copied without re-watching the video.
| Notion | Key values | Curve |
|---|---|---|
| Slide up | Transform keyframes 15 frames apart; opacity 0 to 100 over two frames | Cubic out |
| Slide down | Same keyframes, first position above the canvas | Cubic out, first handle pulled |
| Slide from right | Horizontal offset 1,495 px | Circle ease |
| Slide from left | Horizontal offset minus 1,495 | Circle ease |
| Bounce up | Scale 100, then 84 after 10 frames, five- and three-frame steps back to 100 | Bounce preset via shift + W, handles dragged |
| Blur slide | Blur keyframe at full strength, 15 frames to zero | Smoothed in variable speed panel |
| Blur fade | Blur keyframes plus CapCut fade in at 0.8 s | CapCut preset |
| Bounce plus inhale | Inhale animation at 0.7 s on top of keyframed scale | CapCut preset |
| Glow blink | Glow 0, 100, then 0 over three frames; last keyframe stretched | Middle handles pushed right |
| Masked underline | Split mask wipe across a few frames; height up to 500; layer 1 | Smoothed in variable speed panel |
Where these motions fit and what to check next
The motions in this video are construction techniques rather than finished caption styles. Each one is deliberately short, runs on a frame grid of around 15 frames for the main travel, and depends on a curve adjustment that CapCut does not apply by default, so the finished result comes from the smoothing pass as much as from the keyframes.
One caveat worth stating plainly: every timing and offset in the walkthrough is tied to that project's canvas, font size and clip length. The 1,495-pixel horizontal offset and the minus one character spacing will not produce the same result on a different resolution or a different typeface. Treat them as starting values to test, not fixed settings.
If the goal is a captions preset library, the useful habit from the video is the order of work: build one word properly, wrap it, duplicate it, then save it. Editing instructors who publish this kind of workflow, such as Gustavo dev doido in the Portuguese-language editing community, tend to emphasise the same sequencing point, because rebuilding a Notion per word is where time is lost.
After the seven motions, the next step the instructor points to is caption styling itself: ten viral caption styles broken down and rebuilt from scratch, covering fonts, colours and layout treatments that decide how the captions read before any Notion is applied. Notion gets the attention; the styling step is what makes it look intentional.
Turn a CapCut walkthrough into a written guide
CapCut text animation is the kind of knowledge that exists mostly inside screen recordings: a sequence of right-clicks, frame counts and curve handles that is easy to show and hard to describe. That is exactly the material a written guide handles well, because readers can skim the frame values, copy the settings and skip the parts they already know.
If you have a video where you walk through a workflow like this one, Skalablog turns it into a written article. Paste the YouTube URL, let the video be transcribed, and generate a structured draft you can review and publish.
Start at Skala Blog.
Fork this article
Start a new branch from the same video, shaped your way. You keep the credit; the original keeps the attribution.
A fork in another language is filed as a translation of this article, so the two pages point at each other. You can unlink it later from the editor.
0/240
You are creating
- Format
- For
- Language
- Source
- Your angle
You will be asked to sign in before it is generated.
Buy credits