Building a Local YouTube Transcription Pipeline on a 2017 Intel MacBook Pro

I wanted new recordings from a YouTube playlist transcribed automatically, saved locally, and uploaded to Google Drive. The important constraint was that speech recognition had to be entirely local: no hosted AI API, subscription, per-minute fee, or cloud inference fallback.

The result is a small transcription appliance running on fairway-frank, a 2017 Intel MacBook Pro with Omarchy Linux. Every Friday it checks the playlist, processes the newest recording, and uploads a timestamped UTF-8 transcript—all without requiring Codex or an interactive login to be running.

The machine

The available hardware shaped the design:

  • 2017 Apple MacBook Pro (MacBookPro14,1)
  • Intel Core i5-7360U, two physical cores and four threads
  • 16 GiB RAM and 30 GiB swap
  • Intel Iris Plus 640 graphics
  • Omarchy Linux on x86-64

This is modest hardware for modern speech recognition. GPU acceleration was not a practical target, so I designed the pipeline around reliable CPU inference, resumable processing, and a weekly workload where accuracy mattered more than speed.

The stack

The system uses a deliberately small collection of free and open-source tools:

  • yt-dlp for playlist metadata and source audio downloads
  • FFmpeg for loudness normalization, mono 16 kHz PCM conversion, and chunking
  • whisper.cpp for local Whisper inference
  • Whisper large-v3-turbo-q5_0 as the multilingual transcription model
  • Silero VAD v6.2.0 for local voice-activity detection
  • Python for orchestration, state tracking, retries, and transcript assembly
  • rclone for Google Drive uploads
  • systemd user services and timers for scheduling and supervision

No AI service receives the audio. Network access is used only to retrieve YouTube media, install tools or model weights, and upload completed text files to Drive.

Choosing a model for old Intel hardware

The recordings contain both English and Arabic, so an English-only model or a translation-first workflow was not acceptable. The goal was transcription in the language actually spoken.

I benchmarked a smaller multilingual Whisper model first. It was faster, but its mixed English-Arabic output was not reliable enough. I then tested the quantized multilingual large-v3-turbo-q5_0 model. On this dual-core CPU it ran at roughly three to four times the audio duration, but the improvement in multilingual output justified the extra time.

The effective command for each three-minute chunk is:

whisper-cli \
  --model ~/.local/share/youtube-transcriber/models/ggml-large-v3-turbo-q5_0.bin \
  --file CHUNK.wav \
  --language auto \
  --threads 4 \
  --no-gpu \
  --output-json \
  --vad \
  --vad-model ~/.local/share/youtube-transcriber/models/ggml-silero-v6.2.0.bin \
  --output-file CHUNK \
  --print-progress \
  --no-prints

Language detection is repeated for every chunk because a khutbah may move between Arabic and English. There is intentionally no --translate option.

How the workflow operates

At each run, the Python orchestration script:

  1. Reads the newest item from the configured YouTube playlist.
  2. Rejects live, upcoming, unavailable, or still-processing media as temporarily retryable.
  3. Downloads the best suitable original-language audio.
  4. Normalizes and converts it to mono 16 kHz WAV.
  5. Splits the recording into three-minute chunks.
  6. Transcribes each chunk locally with whisper.cpp and Silero VAD.
  7. Checkpoints every completed chunk as JSON.
  8. Merges the chunks into a readable UTF-8 transcript.
  9. Adds the title, publication date, source URL, video ID, engine details, and timestamps.
  10. Uploads the transcript with rclone and verifies the remote object.

The downloaded audio is removed only after upload verification. Metadata and transcription checkpoints remain available so a failed run can resume without repeating expensive inference.

Duplicate prevention and recovery

Each YouTube video receives its own state directory keyed by video ID. Separate markers record metadata, audio preparation, chunk transcription, merged output, and successful upload.

This separation matters. If Drive is temporarily unavailable, the next attempt reuses the completed transcript rather than spending hours transcribing it again. Existing Drive files are never deleted or overwritten automatically. A file with the expected name is inspected and handled conservatively.

A filesystem lock prevents overlapping runs. Scheduled attempts are bounded, with thirty minutes between retries, which gives a recently completed livestream time to finish YouTube’s post-processing without creating an infinite retry loop.

Scheduling and power management

A persistent systemd user timer runs the job every Friday at 1:58 p.m. America/Los_Angeles. Using the named IANA time zone means daylight-saving changes are handled automatically.

Systemd user lingering allows the timer to operate without an interactive desktop session. The service also takes a sleep-inhibitor lock while work is active. On AC power, the laptop is configured not to suspend when its lid is closed; on battery, normal lid-close suspension remains in place. A persistent timer catches a missed run after the next boot or wake.

The next-run status is visible with:

systemctl --user list-timers youtube-playlist-transcribe.timer

Logs are available through both the systemd journal and a persistent workflow log:

journalctl --user -u youtube-playlist-transcribe.service -f
less ~/.local/state/youtube-transcriber/workflow.log

Google Drive authentication

rclone initially worked through its shared Google OAuth client, but that client is being retired. For long-term reliability I created a personal Google OAuth Desktop app in a dedicated Cloud project, enabled the Google Drive API, placed the OAuth app in production mode, and reauthorized rclone with the target Google account.

The client credential and refresh token live only in rclone’s mode-0600 configuration. They are not stored in the orchestration source, documentation, or shell scripts. The downloaded credential JSON was removed after access and uploads were verified.

Challenges and solutions

Mixed Arabic and English accuracy

The smaller model was faster but not accurate enough. Benchmarking representative speech before committing to a model led to the quantized large-v3-turbo model, accepting slower inference for better multilingual output.

Long CPU-bound runs

On an older dual-core processor, a recording can take hours. Three-minute chunks, per-chunk checkpoints, and a sleep inhibitor make long runs recoverable and prevent routine interruptions from wasting completed work.

Livestream timing

A stream may end before the scheduled time but remain unavailable while YouTube processes it. Bounded delayed retries treat this as a temporary source condition instead of a permanent failure.

Avoiding repeated work

Download, transcription, assembly, and upload are separate persisted stages. An upload failure therefore retries only the network operation, while an interrupted transcription resumes at the first missing chunk.

Reliable unattended Drive access

The shared OAuth client was a future failure point. Migrating to a personal production-mode Desktop client removed that dependency and avoided the short refresh-token lifetime associated with testing-mode apps.

Verification

Commissioning included a complete video processed from download through local inference and verified Drive upload. The real Friday timer then independently processed the next khutbah and uploaded its transcript. The timer’s following occurrence, service state, Drive listing, and uploaded file hashes were all checked.

The result is intentionally unglamorous infrastructure: a repurposed Intel laptop, a handful of durable command-line tools, and enough checkpointing to make slow local inference dependable. It costs nothing per minute, preserves the original English and Arabic speech, and keeps the audio away from hosted transcription services.