⏳ This skill is pending AI review.
Scores will appear once the review pipeline completes.
youtube-fetcher
>-
Choose how to use this skill
You do not need every option. Choose the path your AI client supports. The stable page stays the same; versioned files are immutable.
1. Native installer
This listing has no registered native installer command. Use the complete package or source fallback below, depending on what your client supports.
Do not guess an installer command or replace an existing version without reviewing the diff.
2. Complete package recommended
Download the ZIP when available. It includes SKILL.md plus the references, security notes and version metadata.
No complete ProSkills package is published for this listing yet.3. Prompt-only
Copy the prompt above when the agent can read the stable page or when you want to adopt the workflow without installing a skill.
Need only the instruction file?
Download SKILL.md only if your client requires a single file. The complete ZIP is safer for a full installation because it preserves the references and release context.
No path installs or executes anything by itself. Your agent still needs access to the project files. Before updating, compare the installed version and review the diff.
// RATINGS
Not yet listed on ClawHub or SkillsMP
// README
Video Fetcher to Markdown
A video link in, a structured archival Markdown note out. Capture the transcript, creator metadata, description, chapters, actual language, and provenance in one Obsidian-ready file, without an API key. YouTube captions are read directly; Instagram, TikTok, X, Vimeo, Facebook and other sites are transcribed on your own machine with Whisper, with a contact sheet of frames for short videos.
npx skills add JimmySadek/video-fetcher-to-markdown
Read the v2.0.0 release notes for other video sites, local Whisper transcription, frames, and the login-wall fallback.
Formerly YouTube Fetcher to Markdown. Existing installs keep working and keep updating: the skill is still named
youtube-fetcher, and GitHub redirects the old address.npx skills updatebrings you the latest version.
An independent open-source tool, not affiliated with or endorsed by YouTube, Google, or any other video platform it reads.
What you get
Paste a YouTube link and receive a file such as:
~/yt_transcripts/2026-03-04_obsidian-the-king-of-learning-tools_[hSTy_BInQs8].md
---
title: "Obsidian: The King of Learning Tools (FULL GUIDE + SETUP)"
channel: "Odysseas"
url: "https://www.youtube.com/watch?v=hSTy_BInQs8"
video_id: "hSTy_BInQs8"
fetched: "2026-03-04"
source_project: "my-project"
language: "en"
caption_type: "manual"
duration: "36m 26s"
upload_date: "2024-04-24"
tags:
- yt-transcript
---
# Obsidian: The King of Learning Tools (FULL GUIDE + SETUP)
## Video Details
| Field | Value |
|----------|-------|
| URL | https://www.youtube.com/watch?v=hSTy_BInQs8 |
| Channel | Odysseas |
| Duration | 36m 26s |
| Uploaded | 2024-04-24 |
| Fetched | 2026-03-04 |
| Source | my-project |
| Language | en (manual) |
## Video Description
The creator's description, links, and chapter markers...
## Transcript
The complete caption text...
A video from another site gives the same kind of note, tagged media-transcript,
with platform, creator, transcription_engine and transcription_model in the
frontmatter, a Frames section that embeds the contact sheet with each tile's
time, and a Whisper transcript with timestamps:
~/yt_transcripts/2026-10-07_claude-motion-tips_[instagram-dehp8dpsimi].md
~/yt_transcripts/2026-10-07_claude-motion-tips_[instagram-dehp8dpsimi].frames.jpg
The YAML frontmatter makes a collection queryable through tools such as Dataview, while the Markdown remains portable to Logseq, other knowledge bases, and plain text workflows.
Why this exists
Most transcript extractors stop at raw caption text. An archival knowledge note also needs the source URL, creator, capture date, actual language, description, chapters, and a predictable filename. Video Fetcher to Markdown keeps that complete record in one local file. Short social videos often show the real content on screen (tool names, prompts, links) rather than saying it, so notes from those sites include frames as well as words.
Features
- Manual and auto-generated captions with optional timestamps
- Clickable timestamps and chapters that jump to the moment in the video
- Ordered language preferences, regional variants, automatic selection, and strict language matching
- Explicit YouTube translation, labeled with source language and machine-translation provenance
- Title, channel, duration, upload date, description, and chapters when available
- Safe YAML frontmatter and Markdown tables for dynamic metadata
- File protection for every format, safe replacement, and refreshes that update existing notes in place
- Obsidian-vault and custom-directory output
- Plain text, JSON, SRT, and WebVTT export
- Bounded network requests, useful errors, and optional metadata-free capture
- No API keys and no hosted service
- Other sites (new): Instagram, TikTok, X, Vimeo, Facebook and anything else
yt-dlpsupports, plus YouTube videos without captions and local files, transcribed on your machine with Whisper - Frames (new): a contact sheet of the video for reading on-screen text, by default for videos up to 3 minutes
- Login walls (new): a clear exit code and a browser fallback for sites such as Instagram; your browser login is used only when you ask
Installation
Install the skill
npx skills add JimmySadek/video-fetcher-to-markdown
Or clone the canonical repository:
git clone https://github.com/JimmySadek/video-fetcher-to-markdown.git
Install runtime dependencies
Python 3.8–3.14 is supported for captions. Python 3.10 or newer is recommended
for current optional yt-dlp releases. From the cloned or installed skill
directory, use an isolated environment so your system Python stays unchanged:
python3 -m venv .venv
.venv/bin/python -m pip install -r requirements.txt
.venv/bin/python scripts/fetch_transcript.py --check-deps
On Windows PowerShell:
py -m venv .venv
.venv\Scripts\python.exe -m pip install -r requirements.txt
.venv\Scripts\python.exe scripts\fetch_transcript.py --check-deps
Activate that environment before using the python3 examples below (source .venv/bin/activate on macOS/Linux), or use the full interpreter path each time.
An agent should also use that interpreter. If your skill installation is read-only,
create the environment in a writable location and pass the full path to
requirements.txt.
yt-dlp is optional for descriptions, chapters, duration, and upload dates:
.venv/bin/python -m pip install yt-dlp
# Windows: .venv\Scripts\python.exe -m pip install yt-dlp
Put its executable on PATH by activating the environment. Without it, oEmbed
still supplies title and channel when accessible. The script never installs
packages automatically. --no-metadata skips both metadata providers.
Other video sites need a few more command-line tools (ffmpeg, yt-dlp and a
Whisper tool); see Other video sites.
Usage
python3 scripts/fetch_transcript.py "https://youtu.be/VIDEO_ID"
An agent using the skill resolves scripts/fetch_transcript.py relative to its
installed SKILL.md; it does not depend on one fixed home-directory path.
Output location
The first configured option wins:
--outputfor one exact file--output-dirfor this runVIDEO_FETCHER_DIRfor a persistent directory (the olderYOUTUBE_FETCHER_DIRstill works)~/yt_transcripts/by default
# Save this note to an Obsidian vault
python3 scripts/fetch_transcript.py URL --output-dir ~/Notes/MyVault
# Set a persistent default
export VIDEO_FETCHER_DIR=~/Notes/MyVault
python3 scripts/fetch_transcript.py URL
# Save to one exact file
python3 scripts/fetch_transcript.py URL --output ~/Notes/video.md
Every format preserves an existing destination and exits with code 3, before
making a network request when the destination is already known. This is the same
in terminals and agent sessions; there is no hidden interactive prompt. --force
replaces the chosen file completely, including any annotations. A default
Markdown refresh reuses the existing note's path even if its title or capture date
has changed. An explicit --output is honored independently of other notes for
the same video, so distinct files can hold different languages or versions.
Writes use a temporary file beside the destination. Where the filesystem supports hard links, a new file appears only once its UTF-8 content is complete. Other filesystems use exclusive creation: they still refuse to open an existing file for writing, but a new file can be visible during the write. Handled write failures remove that par
// HOW IT'S BUILT
KEY FILES