★ Reading this for free? Get 25 structured AI courses + per-chapter AI tutor — the first chapter of every course free, no card.Start free in 30 secondsOr own every course: $149 once
Video

Turn Long Videos Into Shorts Locally: No Watermark

October 4, 2026
14 min read
LocalAimaster Research Team

Want to go deeper than this article?

Free account unlocks the first chapter of all 25 courses — RAG, agents, MCP, voice AI, MLOps, real GitHub repos.

📚AI Learning Path

Go from reading about AI to building with AI 25 structured courses. Hands-on projects. Runs on your machine. Start free.

Start free
Or own it for life — Lifetime $149, pay once

Short answer: use Anil-matcha/AI-Youtube-Shorts-Generator in --mode local, and export OPENAI_BASE_URL=http://localhost:11434/v1 so the highlight-ranking step hits Ollama instead of OpenAI. That one environment variable is the difference between "local mode" and actually local: the repo's local path already runs download, transcription, cutting and cropping on your machine, but ships requiring an OpenAI or Gemini key for the LLM that picks the clips.

No watermark, no per-clip credits, no minute cap. What you trade is polish: the vertical crop is an OpenCV Haar-cascade face tracker that follows the biggest face rather than the person speaking, and the final file comes out in the wrong codec unless you re-encode. Both are fixable in one line each, and both are covered below.

If you burned your Opus Clip free tier and came looking for a replacement, the honest framing is this: you will get 80% of the result for $0 and about twenty minutes of setup, and you will hand-fix the framing on multi-speaker footage.


The Short Answer, In Order

Five steps, four of them already local.

StageWhat runsLocal by default?
Downloadyt-dlpYes
Transcribefaster-whisper (CPU or CUDA)Yes
Rank highlightsOpenAI or Gemini clientNo — this is the one you redirect
CutffmpegYes
Reframe to 9:16ffmpeg + OpenCV face trackingYes

Everything on this page was verified by reading the current source of the repos involved on August 18, 2026. Where we could not run something end to end, we say so rather than reporting a number we did not measure.


Reading articles is good. Building is better.

Free account = the first chapter of all 25 courses, with a per-chapter AI tutor. No card.

What Is Dead and Why It Matters

Two of the three repos the internet recommends for this have not been touched in over a year. Checking that before you pip install -r requirements.txt will save you an afternoon.

RepoStarsLast pushVerdict
Anil-matcha/AI-Youtube-Shorts-Generator4,620July 29, 2026Alive. The one to use.
RayVentura/ShortGPT7,839February 10, 2025Abandoned
ClipsAI/clipsai531January 17, 2024Abandoned

Star counts are why people keep landing on ShortGPT — it has more of them than the maintained option. Stars are a lagging indicator of popularity, not a leading indicator of whether the dependency pins still resolve. Check the push date first, every time.

Supporting tools in the chain, all verified the same day: yt-dlp (185,191 stars, Unlicense, pushed August 17, 2026), m-bain/whisperX (23,616 stars, BSD-2-Clause, pushed July 13, 2026), and SYSTRAN/faster-whisper (24,965 stars, MIT). Note that faster-whisper's last push was November 19, 2025 — it is stable and widely embedded rather than dead, but it is not moving fast either.


Install the Local Mode

The install is two requirements files, and the second one is the one that matters.

git clone https://github.com/Anil-matcha/AI-Youtube-Shorts-Generator.git
cd AI-Youtube-Shorts-Generator

python3.10 -m venv venv
source venv/bin/activate

pip install -r requirements.txt
pip install -r requirements-local.txt      # required for --mode local

requirements-local.txt pulls yt-dlp, faster-whisper, openai, google-genai and opencv-python. You also need ffmpeg on your PATH — the repo shells out to it directly and will not install it for you.

Then a .env in the project root:

LLM_PROVIDER=openai
OPENAI_API_KEY=ollama                       # any non-empty string
OPENAI_BASE_URL=http://localhost:11434/v1   # the line that makes it local
OPENAI_MODEL=<a model you have pulled>

LOCAL_WHISPER_MODEL=large-v3                # default is 'base' — too weak
LOCAL_WHISPER_DEVICE=cuda                   # auto | cpu | cuda
LOCAL_OUTPUT_DIR=output

Run it against a local file or a URL:

python main.py "/path/to/podcast.mp4" --mode local --num-clips 5 --aspect-ratio 9:16 --output-json result.json

Clips land at ./output/short_01.mp4, short_02.mp4, and so on. --output-json dumps the full transcript plus every candidate highlight with its score, hook line and reasoning — read that file before you read the videos, because it tells you whether the model understood the content or just chopped every ten minutes.


Point the LLM at Ollama

This works because of a two-line coincidence between the two projects, and we checked both.

The repo's local LLM backend builds its client like this — note there is no base_url argument:

client = OpenAI(api_key=require_openai_key())
response = client.chat.completions.create(
    model=OPENAI_MODEL,
    temperature=0.7,
    messages=[{"role": "user", "content": prompt}],
)

And the OpenAI Python SDK, when base_url is not passed, resolves it from the environment before falling back to the hosted API:

elif base_url is None:
    base_url = os.environ.get("OPENAI_BASE_URL")

So OPENAI_BASE_URL wins. Two gotchas that follow from the repo's own config module:

  1. OPENAI_API_KEY cannot be empty. require_openai_key() raises RuntimeError: OPENAI_API_KEY is not set before the request is ever built. Ollama ignores the value; the repo does not. Set it to anything.
  2. OPENAI_MODEL defaults to gpt-4o-mini, which Ollama does not have. Set it to a tag you have pulled or the request 404s.

Sanity-check the endpoint before running the pipeline:

curl http://localhost:11434/v1/models

Honesty note: we verified this path by reading the current source of both projects, not by rendering a finished video. If it fails for you it will fail loudly at the highlight step with a connection error or a model-not-found from Ollama, not silently — and the transcript will already be cached, so a retry is cheap. New to running the server, see our Ollama guide.

Will a small model rank clips coherently?

This is the question everyone asks and it deserves a straight answer rather than a benchmark we did not run. The task is long-context structured extraction, not taste: read a transcript with timestamps, apply a fixed rubric, emit valid JSON with start_time and end_time floats. Small models fail at that in two specific ways — malformed JSON, and timestamps that do not exist in the transcript. Both are cheap to detect and expensive to ignore, because an out-of-range timestamp becomes a black or zero-length clip once ffmpeg gets hold of it.

So: use the largest instruct model your VRAM allows, keep the transcript chunked (the repo already does this), and validate the JSON against the transcript's duration before rendering. Fifteen lines of validation beats fifteen minutes of rendering garbage.


Own it instead of renting it

Run this on your own machine and stop paying every month

Pay once and keep it. No renewal, no per-token bill, and nothing you feed it ever leaves your hardware.

The Highlight Prompt

You can and should edit this — it is in shorts_generator/highlights.py, and it is the entire creative logic of the tool. Here is the ranked rubric it ships with, with the in-prompt examples trimmed:

Virality signals to prioritize (ranked by impact):
1. HOOK MOMENTS — statements that create immediate curiosity
2. EMOTIONAL PEAKS — genuine surprise, laughter, anger, vulnerability, excitement
3. OPINION BOMBS — strong, polarizing or counter-intuitive statements
4. REVELATION MOMENTS — surprising facts, stats, or confessions
5. CONFLICT/TENSION — disagreement, pushback, a problem confronted head-on
6. QUOTABLE ONE-LINERS — a sentence that works as a standalone quote card
7. STORY PEAKS — the climax or twist of an anecdote; the payoff moment
8. PRACTICAL VALUE — a concrete tip, hack, or insight

And the rules attached to it:

  • Every highlight must open with a hook that lands within the first 3 seconds.
  • Duration sweet spot: 45–90 seconds. Shorter (20–44s) only for a perfect standalone one-liner; longer (91–180s) only when a story arc needs the context.
  • Never cut mid-sentence or mid-thought.
  • Score 0–100 on viral potential, not general quality.
  • Output only JSON: {"highlights":[{"title","start_time","end_time","score","hook_sentence","virality_reason"}]}.

A separate classification pass first labels the video (podcast, interview, tutorial, lecture, commentary, debate, vlog, other) and its density (low/medium/high), and those get interpolated into the ranking prompt.

The chunking constants are the other thing worth tuning, and they are all in the same file:

ConstantDefaultMeaning
LONG_VIDEO_THRESHOLD1800sVideos longer than 30 minutes get chunked
CHUNK_SIZE_SECONDS1200s20-minute chunks
CHUNK_OVERLAP_SECONDS60sOverlap so cross-boundary clips are not lost

Overlapping candidates are then deduped: where two highlights overlap by more than 50%, the higher score wins. For a 90-minute podcast that means five LLM calls, not one — worth knowing when you are choosing a model, because five slow calls add up.


What the Crop Actually Does

It is a per-frame Haar cascade that follows the largest face, with no idea who is speaking. Knowing exactly this is the difference between using it well and being annoyed by it.

The reframer, in shorts_generator/local/clipper.py:

  1. Computes the largest crop rectangle that fits the source frame at the target ratio.
  2. Loads OpenCV's haarcascade_frontalface_default.xml.
  3. Per frame, runs detectMultiScale(gray, scaleFactor=1.1, minNeighbors=5, minSize=(40, 40)).
  4. Picks the largest detected face — the code comments this as "usually the speaker".
  5. Eases the crop centre toward it with a smoothing factor of 0.15 per frame.
  6. Falls back to the frame centre when nothing has ever been detected.

What follows from that design, in the order you will hit it:

  • Profile shots break detection. Haar's frontal classifier does not fire on a turned head, so the window holds its last position until the face comes back. On an interview where one host is angled to the other, expect drift.
  • "Largest face" is not "speaking face". On a two-shot the crop follows whoever sits closer to the camera for the whole clip. There is no audio-visual attribution anywhere in this path.
  • The 0.15 smoothing is a deliberate lag. It stops the frame jittering between detections; it also means fast cuts between speakers arrive late. Raise it toward 1.0 for snappier tracking and more jitter, lower it for calmer, laggier motion.
  • Single talking head works well. That is genuinely the case it was built for, and it does that case fine.

If your footage is multi-speaker and you need the right person on screen, the honest recommendation is to let this pass pick the clips and do the reframing by hand, or crop to a fixed region per speaker with plain ffmpeg.


Fix the Output Codec

The final clips are MPEG-4 Part 2, not H.264 — re-encode before you upload. This is not a bug report, it is a consequence of how the two stages are wired.

The cut stage encodes properly with x264:

ffmpeg -y -i source.mp4 -ss <start> -to <end> \
  -c:v libx264 -preset fast -crf 20 -c:a aac -b:a 128k cut.mp4

But the reframe stage writes its output through OpenCV's VideoWriter with the mp4v fourcc, then muxes the audio back with -c:v copy — so the x264 encode is discarded and the shipped file carries the OpenCV codec. Platforms re-encode uploads anyway, so it works; you are just uploading a bigger file at worse quality per bit than you needed to.

One line fixes it:

for f in output/short_*.mp4; do
  ffmpeg -y -i "$f" -c:v libx264 -crf 20 -preset medium -pix_fmt yuv420p \
    -c:a aac -b:a 128k "${f%.mp4}_h264.mp4"
done

-pix_fmt yuv420p is not optional if you want the file to play everywhere.


The DIY Fallback Chain

Build this if the repo goes stale — it is four tools you already have and it has no single point of failure. Every stage is independently maintained, which is the whole point.

# 1. Download (your own content, or content you have rights to)
yt-dlp -f "bv*+ba/b" -o "source.%(ext)s" "<url>"

# 2. Word-level timestamps
pip install whisperx
whisperx source.mp4 --model large-v3 --output_format json --align_model WAV2VEC2_ASR_LARGE_LV60K_960H

# 3. Rank highlights with a local model, JSON out
ollama run <your-model> < prompt_with_transcript.txt > highlights.json

# 4. Cut and reframe
ffmpeg -i source.mp4 -ss 124.3 -to 187.6 \
  -vf "crop=ih*9/16:ih:(iw-ih*9/16)/2:0,scale=1080:1920" \
  -c:v libx264 -crf 20 -pix_fmt yuv420p -c:a aac -b:a 128k short_01.mp4

Why WhisperX rather than plain Whisper at step 2: it produces word-level timestamps via forced alignment, so your cut points land on word boundaries instead of Whisper's segment boundaries, which routinely drift by a second or more. A short that starts half a syllable late reads as broken. Full setup in our WhisperX guide; if you want the faster segment-level route instead, faster-whisper is the one the repo above uses internally, and our speech-to-text model comparison covers picking between them.

The crop filter above is a static centre crop — no face tracking, deliberately. It is predictable, it never drifts, and for a centred single speaker it beats the Haar tracker. Replace (iw-ih*9/16)/2 with a fixed x offset to lock onto an off-centre subject.

Captions are the last stage and a separate job: see local AI subtitles with Whisper for burning them in.


Honest Limitations

Things we did not measure, stated plainly rather than estimated.

  • We did not time a 90-minute transcription on CPU versus a 3060. We do not have that hardware on the bench. What we can tell you from the config is the shape of the problem: LOCAL_WHISPER_DEVICE resolves only to cpu or cuda, there is no Metal path because faster-whisper runs on CTranslate2, and the transcript is cached to .srt so you pay the cost once per source file. Time your own with time python main.py ... --mode local.
  • We did not measure clip yield across content types. The repo returns --num-clips (default 3) from a larger candidate pool, so "yield" is really "how many candidates scored well", and that is a property of your video.
  • We did not run the auto-crop against multi-speaker footage frame by frame. The behaviour described above is read from the source, and the source is unambiguous about picking the largest face. The visual result on your specific footage is still worth a five-minute check before you batch a hundred clips.
  • Licence detail: the repo's README states MIT, but GitHub's licence detector does not classify the file, so the API reports none. Read LICENSE yourself before commercial use.
  • Legal: yt-dlp will happily download anything. That does not make it yours. Clip your own long-form content, or content you have explicit permission to use.

Verdict

  1. The free local path is real, and it is one environment variable away from being genuinely free. OPENAI_BASE_URL=http://localhost:11434/v1 plus any placeholder API key moves the last remote call onto your own machine.
  2. Check push dates, not stars. ShortGPT has 7,839 stars and has not shipped since February 2025. The 4,620-star option is the maintained one.
  3. Read result.json before you watch the clips. Scores, hook lines and reasoning tell you immediately whether your local model understood the transcript.
  4. Expect to fix framing on multi-speaker video. Largest-face tracking is not speaker tracking, and no amount of prompt tuning changes that.
  5. Re-encode to H.264 before uploading. One ffmpeg loop, smaller files, better quality.

What you actually save: a $20–$300/month subscription, a per-minute cap, and the requirement to upload unreleased footage to somebody else's server. What you actually spend: an evening of setup and some manual reframing. For most people cutting their own podcast, that trade is obvious. For an agency shipping fifty clips a week from multi-camera interviews, it is not.

If you want models that understand footage rather than cut it, local AI video analysis is the neighbouring piece.


Sources


FAQ

🎯
AI Learning Path

Go from reading about AI to building with AI

25 structured courses. Hands-on projects. Runs on your machine. Start free.

Or own it for life — Lifetime $149 $599, pay once

Liked this? 25 full AI courses are waiting.

From fundamentals to RAG, agents, MCP servers, voice AI, and production deployment with real GitHub repos. First chapter free, every course.

Reading now
Join the discussion
TagsOpus Clip AlternativeVideo EditingWhisperWhisperXffmpegOllamaShorts

LocalAimaster Research Team

Local AI Master writes hands-on courses and hardware guides for running AI on machines you own. Content is checked against current releases and corrected when readers tell us it is wrong.

Build Real AI on Your Machine

RAG, agents, NLP, vision, and MLOps - chapters across 25 courses that take you from reading about AI to building AI.

Want the structured version?

Hands-on courses on local AI, from $8.99 a month. The first chapter of each is free.

AI Learning Path

Comments (0)

No comments yet. Be the first to share your thoughts!

Is there a free open-source alternative to Opus Clip?

Anil-matcha/AI-Youtube-Shorts-Generator (4,620 stars, MIT per its README, last pushed July 29, 2026) is the only actively maintained one we found that ships a genuine local mode. Its --mode local path uses yt-dlp for download, faster-whisper for transcription, an LLM for highlight ranking, and ffmpeg plus OpenCV face tracking for the vertical crop. Out of the box that LLM step still calls OpenAI or Gemini; the fix is to point the OpenAI client at Ollama with the OPENAI_BASE_URL environment variable. Two repos people commonly recommend are dead: ClipsAI (531 stars, last pushed January 17, 2024) and ShortGPT (7,839 stars, last pushed February 10, 2025).

Can I run the highlight ranking on Ollama instead of OpenAI?

Yes, without editing the repo. Its local LLM backend constructs the client as OpenAI(api_key=...) with no base_url argument, and the openai Python SDK falls back to os.environ.get("OPENAI_BASE_URL") when base_url is not passed — we confirmed both in source. So exporting OPENAI_BASE_URL=http://localhost:11434/v1 redirects the call to Ollama. You still need a non-empty OPENAI_API_KEY because the repo raises a RuntimeError on an empty one; any placeholder string works. Set OPENAI_MODEL to whatever you have pulled.

Which local model is good enough to pick viral clips?

The job is not creative — it is extracting structured JSON of start and end times from a long transcript against a scoring rubric, which is a long-context instruction-following task. That means the failure mode on small models is malformed JSON and hallucinated timestamps rather than boring clip choices. Start at the largest instruct model your VRAM allows and verify by checking every returned start_time and end_time against the transcript before rendering. Do not skip that check: a timestamp outside the video length turns into a zero-length or black clip at the ffmpeg stage.

Does the auto-crop keep the speaker centred on two-person footage?

Only by accident. The local reframer runs OpenCV's haarcascade_frontalface_default classifier per frame and picks the largest detected face, then eases the crop window toward it with a smoothing factor of 0.15. There is no audio-visual speaker attribution, so on a two-host podcast it follows whoever is closest to camera, not whoever is talking. Haar is also frontal-only — when a head turns in profile the detection drops and the window holds its last position. For talking-head footage with one person it is fine; for interview setups, expect to fix shots manually.

How long does transcribing a 90-minute podcast take?

It depends entirely on the device flag, and the repo defaults will surprise you. LOCAL_WHISPER_DEVICE defaults to auto and only ever resolves to cpu or cuda — there is no Metal path, because faster-whisper runs on CTranslate2, which has no Apple GPU backend. So on a Mac this is a CPU job regardless of your chip. On an NVIDIA card set LOCAL_WHISPER_DEVICE=cuda. The good news is you pay it once: the transcript is cached as an .srt next to the output and reused if it is newer than the source file.

Are the rendered clips actually watermark-free and platform-ready?

Watermark-free, yes — nothing in the local path draws an overlay. Platform-ready needs one extra step. The reframer writes its intermediate with OpenCV's mp4v fourcc and then muxes audio back with -c:v copy, so the final file is MPEG-4 Part 2 rather than H.264. Uploads generally still work because the platforms re-encode, but you get a larger file at lower quality per bit. Re-encode before uploading: ffmpeg -i short_01.mp4 -c:v libx264 -crf 20 -pix_fmt yuv420p -c:a aac -b:a 128k out.mp4.

Ready to Go Beyond Tutorials?

25 structured courses with hands-on chapters - build RAG chatbots, AI agents, and ML pipelines on your own hardware.

Bonus kit

Ollama Docker Templates

10 one-command Docker stacks for local models — get the Ollama backend behind this pipeline serving in minutes. Included with paid plans, or free after subscribing to both Local AI Master and Little AI Master on YouTube.

See Plans →

Was this helpful?

📅 Published: October 4, 2026🔄 Last Updated: October 4, 2026✓ Manually Reviewed
LM

Written by the Local AI Master Team

The team behind Local AI Master

We build Local AI Master around practical, testable local AI workflows: model selection, hardware planning, RAG systems, agents, and MLOps. The goal is to turn scattered tutorials into a structured learning path you can follow on your own hardware.

✓ Local AI Curriculum✓ Hands-On Projects✓ Open Source Contributor
📚
Free · no account required

Grab the AI Starter Kit — career roadmap, cheat sheet, setup guide

No spam. Unsubscribe with one click.

🎯
AI Learning Path

Go from reading about AI to building with AI

25 structured courses. Hands-on projects. Runs on your machine. Start free.

Or own it for life — Lifetime $149 $599, pay once
Free Tools & Calculators