Turn Long Videos Into Shorts Locally: No Watermark
Want to go deeper than this article?
Free account unlocks the first chapter of all 25 courses — RAG, agents, MCP, voice AI, MLOps, real GitHub repos.
Go from reading about AI to building with AI 25 structured courses. Hands-on projects. Runs on your machine. Start free.
Short answer: use Anil-matcha/AI-Youtube-Shorts-Generator in --mode local, and export OPENAI_BASE_URL=http://localhost:11434/v1 so the highlight-ranking step hits Ollama instead of OpenAI. That one environment variable is the difference between "local mode" and actually local: the repo's local path already runs download, transcription, cutting and cropping on your machine, but ships requiring an OpenAI or Gemini key for the LLM that picks the clips.
No watermark, no per-clip credits, no minute cap. What you trade is polish: the vertical crop is an OpenCV Haar-cascade face tracker that follows the biggest face rather than the person speaking, and the final file comes out in the wrong codec unless you re-encode. Both are fixable in one line each, and both are covered below.
If you burned your Opus Clip free tier and came looking for a replacement, the honest framing is this: you will get 80% of the result for $0 and about twenty minutes of setup, and you will hand-fix the framing on multi-speaker footage.
The Short Answer, In Order
Five steps, four of them already local.
| Stage | What runs | Local by default? |
|---|---|---|
| Download | yt-dlp | Yes |
| Transcribe | faster-whisper (CPU or CUDA) | Yes |
| Rank highlights | OpenAI or Gemini client | No — this is the one you redirect |
| Cut | ffmpeg | Yes |
| Reframe to 9:16 | ffmpeg + OpenCV face tracking | Yes |
Everything on this page was verified by reading the current source of the repos involved on August 18, 2026. Where we could not run something end to end, we say so rather than reporting a number we did not measure.
Reading articles is good. Building is better.
Free account = the first chapter of all 25 courses, with a per-chapter AI tutor. No card.
What Is Dead and Why It Matters
Two of the three repos the internet recommends for this have not been touched in over a year. Checking that before you pip install -r requirements.txt will save you an afternoon.
| Repo | Stars | Last push | Verdict |
|---|---|---|---|
| Anil-matcha/AI-Youtube-Shorts-Generator | 4,620 | July 29, 2026 | Alive. The one to use. |
| RayVentura/ShortGPT | 7,839 | February 10, 2025 | Abandoned |
| ClipsAI/clipsai | 531 | January 17, 2024 | Abandoned |
Star counts are why people keep landing on ShortGPT — it has more of them than the maintained option. Stars are a lagging indicator of popularity, not a leading indicator of whether the dependency pins still resolve. Check the push date first, every time.
Supporting tools in the chain, all verified the same day: yt-dlp (185,191 stars, Unlicense, pushed August 17, 2026), m-bain/whisperX (23,616 stars, BSD-2-Clause, pushed July 13, 2026), and SYSTRAN/faster-whisper (24,965 stars, MIT). Note that faster-whisper's last push was November 19, 2025 — it is stable and widely embedded rather than dead, but it is not moving fast either.
Install the Local Mode
The install is two requirements files, and the second one is the one that matters.
git clone https://github.com/Anil-matcha/AI-Youtube-Shorts-Generator.git
cd AI-Youtube-Shorts-Generator
python3.10 -m venv venv
source venv/bin/activate
pip install -r requirements.txt
pip install -r requirements-local.txt # required for --mode local
requirements-local.txt pulls yt-dlp, faster-whisper, openai, google-genai and opencv-python. You also need ffmpeg on your PATH — the repo shells out to it directly and will not install it for you.
Then a .env in the project root:
LLM_PROVIDER=openai
OPENAI_API_KEY=ollama # any non-empty string
OPENAI_BASE_URL=http://localhost:11434/v1 # the line that makes it local
OPENAI_MODEL=<a model you have pulled>
LOCAL_WHISPER_MODEL=large-v3 # default is 'base' — too weak
LOCAL_WHISPER_DEVICE=cuda # auto | cpu | cuda
LOCAL_OUTPUT_DIR=output
Run it against a local file or a URL:
python main.py "/path/to/podcast.mp4" --mode local --num-clips 5 --aspect-ratio 9:16 --output-json result.json
Clips land at ./output/short_01.mp4, short_02.mp4, and so on. --output-json dumps the full transcript plus every candidate highlight with its score, hook line and reasoning — read that file before you read the videos, because it tells you whether the model understood the content or just chopped every ten minutes.
Point the LLM at Ollama
This works because of a two-line coincidence between the two projects, and we checked both.
The repo's local LLM backend builds its client like this — note there is no base_url argument:
client = OpenAI(api_key=require_openai_key())
response = client.chat.completions.create(
model=OPENAI_MODEL,
temperature=0.7,
messages=[{"role": "user", "content": prompt}],
)
And the OpenAI Python SDK, when base_url is not passed, resolves it from the environment before falling back to the hosted API:
elif base_url is None:
base_url = os.environ.get("OPENAI_BASE_URL")
So OPENAI_BASE_URL wins. Two gotchas that follow from the repo's own config module:
OPENAI_API_KEYcannot be empty.require_openai_key()raisesRuntimeError: OPENAI_API_KEY is not setbefore the request is ever built. Ollama ignores the value; the repo does not. Set it to anything.OPENAI_MODELdefaults togpt-4o-mini, which Ollama does not have. Set it to a tag you have pulled or the request 404s.
Sanity-check the endpoint before running the pipeline:
curl http://localhost:11434/v1/models
Honesty note: we verified this path by reading the current source of both projects, not by rendering a finished video. If it fails for you it will fail loudly at the highlight step with a connection error or a model-not-found from Ollama, not silently — and the transcript will already be cached, so a retry is cheap. New to running the server, see our Ollama guide.
Will a small model rank clips coherently?
This is the question everyone asks and it deserves a straight answer rather than a benchmark we did not run. The task is long-context structured extraction, not taste: read a transcript with timestamps, apply a fixed rubric, emit valid JSON with start_time and end_time floats. Small models fail at that in two specific ways — malformed JSON, and timestamps that do not exist in the transcript. Both are cheap to detect and expensive to ignore, because an out-of-range timestamp becomes a black or zero-length clip once ffmpeg gets hold of it.
So: use the largest instruct model your VRAM allows, keep the transcript chunked (the repo already does this), and validate the JSON against the transcript's duration before rendering. Fifteen lines of validation beats fifteen minutes of rendering garbage.
Run this on your own machine and stop paying every month
Pay once and keep it. No renewal, no per-token bill, and nothing you feed it ever leaves your hardware.
The Highlight Prompt
You can and should edit this — it is in shorts_generator/highlights.py, and it is the entire creative logic of the tool. Here is the ranked rubric it ships with, with the in-prompt examples trimmed:
Virality signals to prioritize (ranked by impact):
1. HOOK MOMENTS — statements that create immediate curiosity
2. EMOTIONAL PEAKS — genuine surprise, laughter, anger, vulnerability, excitement
3. OPINION BOMBS — strong, polarizing or counter-intuitive statements
4. REVELATION MOMENTS — surprising facts, stats, or confessions
5. CONFLICT/TENSION — disagreement, pushback, a problem confronted head-on
6. QUOTABLE ONE-LINERS — a sentence that works as a standalone quote card
7. STORY PEAKS — the climax or twist of an anecdote; the payoff moment
8. PRACTICAL VALUE — a concrete tip, hack, or insight
And the rules attached to it:
- Every highlight must open with a hook that lands within the first 3 seconds.
- Duration sweet spot: 45–90 seconds. Shorter (20–44s) only for a perfect standalone one-liner; longer (91–180s) only when a story arc needs the context.
- Never cut mid-sentence or mid-thought.
- Score 0–100 on viral potential, not general quality.
- Output only JSON:
{"highlights":[{"title","start_time","end_time","score","hook_sentence","virality_reason"}]}.
A separate classification pass first labels the video (podcast, interview, tutorial, lecture, commentary, debate, vlog, other) and its density (low/medium/high), and those get interpolated into the ranking prompt.
The chunking constants are the other thing worth tuning, and they are all in the same file:
| Constant | Default | Meaning |
|---|---|---|
LONG_VIDEO_THRESHOLD | 1800s | Videos longer than 30 minutes get chunked |
CHUNK_SIZE_SECONDS | 1200s | 20-minute chunks |
CHUNK_OVERLAP_SECONDS | 60s | Overlap so cross-boundary clips are not lost |
Overlapping candidates are then deduped: where two highlights overlap by more than 50%, the higher score wins. For a 90-minute podcast that means five LLM calls, not one — worth knowing when you are choosing a model, because five slow calls add up.
What the Crop Actually Does
It is a per-frame Haar cascade that follows the largest face, with no idea who is speaking. Knowing exactly this is the difference between using it well and being annoyed by it.
The reframer, in shorts_generator/local/clipper.py:
- Computes the largest crop rectangle that fits the source frame at the target ratio.
- Loads OpenCV's
haarcascade_frontalface_default.xml. - Per frame, runs
detectMultiScale(gray, scaleFactor=1.1, minNeighbors=5, minSize=(40, 40)). - Picks the largest detected face — the code comments this as "usually the speaker".
- Eases the crop centre toward it with a smoothing factor of 0.15 per frame.
- Falls back to the frame centre when nothing has ever been detected.
What follows from that design, in the order you will hit it:
- Profile shots break detection. Haar's frontal classifier does not fire on a turned head, so the window holds its last position until the face comes back. On an interview where one host is angled to the other, expect drift.
- "Largest face" is not "speaking face". On a two-shot the crop follows whoever sits closer to the camera for the whole clip. There is no audio-visual attribution anywhere in this path.
- The 0.15 smoothing is a deliberate lag. It stops the frame jittering between detections; it also means fast cuts between speakers arrive late. Raise it toward 1.0 for snappier tracking and more jitter, lower it for calmer, laggier motion.
- Single talking head works well. That is genuinely the case it was built for, and it does that case fine.
If your footage is multi-speaker and you need the right person on screen, the honest recommendation is to let this pass pick the clips and do the reframing by hand, or crop to a fixed region per speaker with plain ffmpeg.
Fix the Output Codec
The final clips are MPEG-4 Part 2, not H.264 — re-encode before you upload. This is not a bug report, it is a consequence of how the two stages are wired.
The cut stage encodes properly with x264:
ffmpeg -y -i source.mp4 -ss <start> -to <end> \
-c:v libx264 -preset fast -crf 20 -c:a aac -b:a 128k cut.mp4
But the reframe stage writes its output through OpenCV's VideoWriter with the mp4v fourcc, then muxes the audio back with -c:v copy — so the x264 encode is discarded and the shipped file carries the OpenCV codec. Platforms re-encode uploads anyway, so it works; you are just uploading a bigger file at worse quality per bit than you needed to.
One line fixes it:
for f in output/short_*.mp4; do
ffmpeg -y -i "$f" -c:v libx264 -crf 20 -preset medium -pix_fmt yuv420p \
-c:a aac -b:a 128k "${f%.mp4}_h264.mp4"
done
-pix_fmt yuv420p is not optional if you want the file to play everywhere.
The DIY Fallback Chain
Build this if the repo goes stale — it is four tools you already have and it has no single point of failure. Every stage is independently maintained, which is the whole point.
# 1. Download (your own content, or content you have rights to)
yt-dlp -f "bv*+ba/b" -o "source.%(ext)s" "<url>"
# 2. Word-level timestamps
pip install whisperx
whisperx source.mp4 --model large-v3 --output_format json --align_model WAV2VEC2_ASR_LARGE_LV60K_960H
# 3. Rank highlights with a local model, JSON out
ollama run <your-model> < prompt_with_transcript.txt > highlights.json
# 4. Cut and reframe
ffmpeg -i source.mp4 -ss 124.3 -to 187.6 \
-vf "crop=ih*9/16:ih:(iw-ih*9/16)/2:0,scale=1080:1920" \
-c:v libx264 -crf 20 -pix_fmt yuv420p -c:a aac -b:a 128k short_01.mp4
Why WhisperX rather than plain Whisper at step 2: it produces word-level timestamps via forced alignment, so your cut points land on word boundaries instead of Whisper's segment boundaries, which routinely drift by a second or more. A short that starts half a syllable late reads as broken. Full setup in our WhisperX guide; if you want the faster segment-level route instead, faster-whisper is the one the repo above uses internally, and our speech-to-text model comparison covers picking between them.
The crop filter above is a static centre crop — no face tracking, deliberately. It is predictable, it never drifts, and for a centred single speaker it beats the Haar tracker. Replace (iw-ih*9/16)/2 with a fixed x offset to lock onto an off-centre subject.
Captions are the last stage and a separate job: see local AI subtitles with Whisper for burning them in.
Honest Limitations
Things we did not measure, stated plainly rather than estimated.
- We did not time a 90-minute transcription on CPU versus a 3060. We do not have that hardware on the bench. What we can tell you from the config is the shape of the problem:
LOCAL_WHISPER_DEVICEresolves only tocpuorcuda, there is no Metal path because faster-whisper runs on CTranslate2, and the transcript is cached to.srtso you pay the cost once per source file. Time your own withtime python main.py ... --mode local. - We did not measure clip yield across content types. The repo returns
--num-clips(default 3) from a larger candidate pool, so "yield" is really "how many candidates scored well", and that is a property of your video. - We did not run the auto-crop against multi-speaker footage frame by frame. The behaviour described above is read from the source, and the source is unambiguous about picking the largest face. The visual result on your specific footage is still worth a five-minute check before you batch a hundred clips.
- Licence detail: the repo's README states MIT, but GitHub's licence detector does not classify the file, so the API reports none. Read
LICENSEyourself before commercial use. - Legal:
yt-dlpwill happily download anything. That does not make it yours. Clip your own long-form content, or content you have explicit permission to use.
Verdict
- The free local path is real, and it is one environment variable away from being genuinely free.
OPENAI_BASE_URL=http://localhost:11434/v1plus any placeholder API key moves the last remote call onto your own machine. - Check push dates, not stars. ShortGPT has 7,839 stars and has not shipped since February 2025. The 4,620-star option is the maintained one.
- Read
result.jsonbefore you watch the clips. Scores, hook lines and reasoning tell you immediately whether your local model understood the transcript. - Expect to fix framing on multi-speaker video. Largest-face tracking is not speaker tracking, and no amount of prompt tuning changes that.
- Re-encode to H.264 before uploading. One ffmpeg loop, smaller files, better quality.
What you actually save: a $20–$300/month subscription, a per-minute cap, and the requirement to upload unreleased footage to somebody else's server. What you actually spend: an evening of setup and some manual reframing. For most people cutting their own podcast, that trade is obvious. For an agency shipping fifty clips a week from multi-camera interviews, it is not.
If you want models that understand footage rather than cut it, local AI video analysis is the neighbouring piece.
Sources
- Anil-matcha/AI-Youtube-Shorts-Generator — README,
config.py,highlights.py,local/llm.py,local/clipper.py(read August 18, 2026) - openai/openai-python —
src/openai/_client.pyOPENAI_BASE_URLfallback (v3.2.0) - SYSTRAN/faster-whisper, m-bain/whisperX, yt-dlp/yt-dlp — repository metadata and star counts via the GitHub API
- ClipsAI/clipsai, RayVentura/ShortGPT — last-push dates via the GitHub API
FAQ
Go from reading about AI to building with AI
25 structured courses. Hands-on projects. Runs on your machine. Start free.
Liked this? 25 full AI courses are waiting.
From fundamentals to RAG, agents, MCP servers, voice AI, and production deployment with real GitHub repos. First chapter free, every course.
Build Real AI on Your Machine
RAG, agents, NLP, vision, and MLOps - chapters across 25 courses that take you from reading about AI to building AI.
Want the structured version?
Hands-on courses on local AI, from $8.99 a month. The first chapter of each is free.
Keep going
Comments (0)
No comments yet. Be the first to share your thoughts!