--- name: reel description: Turn a Korean-language Instagram reel or post link into content an AI agent can read — caption, spoken words as text, and key frames as images. Use when the user shares an instagram.com/reel/, /reels/ or /p/ link, or asks to see, analyze or summarize an Instagram video (인스타 릴스, 릴스 분석, 릴스 내용). argument-hint: --- # Reel An Instagram link can't be opened without login, and many agents can't take video or audio directly. This skill converts a reel into things any agent can read: text (caption + transcript) and images (frames). The heavy lifting is done by `scripts/reel.py`; your job is to make sure it can run on this machine, run it, and read what it produces. Scope: Korean-language reels only. Speech recognition is fixed to Korean, so for a reel spoken in another language the transcript will be nonsense — ignore it, rely on the caption and frames, and tell the user the speech wasn't transcribed. ## 1. Check the environment (first run, or when something fails) If `config.json` exists next to this file and the Python it points to still works, skip to step 2. Otherwise, find out what this machine already has before installing anything: - **Python 3.10–3.12** in an isolated environment. Look beyond PATH: conda/miniconda, venvs, the `py` launcher, uv. On Windows, `python` may be a Microsoft Store stub — verify it really runs. - **ffmpeg** on PATH. - In that environment: **yt-dlp** and **faster-whisper** (`pip install yt-dlp faster-whisper`). Rules: - Never install into, upgrade, or modify an existing environment the user already uses. Create a dedicated one (e.g. conda env or venv named `reel`). - Pick install methods that fit this OS and what's already here (winget, brew, apt, conda, pip). - Before installing, tell the user exactly what will be installed, from where, and roughly how big, and wait for a yes. - When setup works, write `config.json` next to this file: `{"python": ""}`. ## 2. Extract Run: ``` /scripts/reel.py "" ``` The script prints the output folder when it finishes. Options: `--every ` (frame interval, default 1), `--model ` (Whisper size, default small), `--out `. - The first transcription downloads a Whisper model (small ≈ 500 MB). Tell the user before that happens. - If the download fails because Instagram requires login, don't guess at cookies. Explain the situation and ask the user how they want to provide login cookies (e.g. a `cookies.txt` export, passed with `--cookies `). ## 3. Read the result The output folder contains: - `caption.txt` — the full post text (the part under "more" included) - `meta.json` — author, date, duration, original URL - `transcript.txt` — spoken words with timestamps (empty if the reel has no speech) - `frames/` — JPG frames named by timestamp, near-duplicates already removed (video posts) - `images/` — the photos, in order (photo and carousel posts) - `media/` — the original downloaded files (only needed to re-run with different options) Read `caption.txt` first. The post text is often as important as the video: many creators put the full explanation, steps, links or code in the caption and keep the video short. Treat it as primary content, not metadata. Then read `meta.json` and `transcript.txt`. The transcript is raw machine speech recognition: it often mangles names, jargon, product names and English words inside the Korean speech (e.g. "클로드다데디" for "CLAUDE.md"). Correct these yourself from context, the caption and on-screen text — they are usually written by the creator and spelled right. Then look at the frames or images. Near-identical frames were already dropped, so each frame is a visual change worth seeing — view them all. If there are more than about 40, pick frames spread across the video and where the transcript suggests something is shown on screen. Pay attention to text on screen — tips reels often put the key point there. Then give the user a short account of what the reel says and shows, and ask what they want to do with it. Don't summarize, save, or act on the content beyond that unless asked.