시작하기
ClipCaster: Viral Clips No AI Voice

ClipCaster: Viral Clips No AI Voice

Turn any long video (podcast, stream, interview, lecture) into vertical clips ready for TikTok, Reels, or Shorts. Captions are the real transcript, synced word by word — no robotic AI voice. Auto-split the whole video into well-sized clips, or hand-pick and stitch together just the best moments, with optional hook/CTA text and a punch-in zoom at the end.
#비디오#분석#소셜 미디어
평점
더 많은 평가가 필요합니다
판매
0
사용 방법
다운로드

What it does

Give it a long-form video — a podcast, stream, interview, lecture, or any talking-head footage — and it hands back a folder of vertical (1080x1920), ready-to-post clips for TikTok, Instagram Reels, or YouTube Shorts.

Instead of an AI-written script read by a synthetic voice (which tends to sound robotic and hurts viewer trust), every clip is captioned with the real transcript, timed word by word — it reads naturally because it is natural. No AI voice is generated or used anywhere in the pipeline.

How it works
Transcribe — the source video is transcribed with word-level timestamps.
Decide what becomes a clip — two modes:
Auto-split: the whole video is divided into clips sized to fit its length (a short source gets ~3 minute clips, a long stream gets up to ~10 minute clips), trimming long silences between words for pacing — never deleting spoken words, so meaning is never altered.
Hand-picked: point it at the strongest few moments instead of the whole video. It can even stitch several non-contiguous moments from the source into one edited clip that skips the slow/rambling parts in between, like a real edit.
Render — burns in word-synced captions, an optional on-screen hook at the start and call-to-action at the end, and an optional "punch-in" zoom on the closing line for emphasis.
Example

Input: a single ~33 minute recording. Output: 6 vertical clips (~4.5-5 minutes each), captions synced word-by-word to what's actually said, each ready to upload as-is.

FAQ

Does it use an AI voice? No — captions are the real transcript, not a script read by a synthetic voice.

Does it post to TikTok/Instagram/YouTube for me? No — it hands back finished .mp4 files; uploading stays a manual step on your own accounts.

What languages does it support? Transcribes in whatever language is spoken in the source (Spanish by default, configurable). Captions render cleanly for Latin- and Cyrillic-script languages (Spanish, English, Portuguese, French, German, Russian, etc.). It does not translate — captions come out in the same language spoken in the video.

Does it check whether I have rights to the video? No — that's on you. Only use video you own or are licensed to use.

What do I need installed to run it? ffmpeg, Python 3.9+, and the faster-whisper / numpy packages. Internet access is needed the first time a given transcription model size is used (cached after that).

Is rendering instant? No — a clip built from many short cuts (auto-split on a talkative source can stitch 15-30 cuts into one clip) takes real time to render, since every cut is a separate seek+decode. Let a full-video batch run in the background rather than expecting it to finish immediately.


외부 API

huggingface.co