⏳ This skill is pending AI review.
Scores will appear once the review pipeline completes.
agent-web-fetch
Read a public web page reliably and honestly. Use whenever the user gives you a URL to read, summarise, or extract from (articles, blog posts, WeChat 公众号 links, GitHub repos, X posts, JS-heavy pages). Escalates from a plain request to client rotation, structured-data extraction, alternates (AMP/RSS) and a headless browser only when blocked, and reports exactly why a page could not be read instead of bypassing logins, paywalls or CAPTCHAs.
Choose how to use this skill
You do not need every option. Choose the path your AI client supports. The stable page stays the same; versioned files are immutable.
1. Native installer
This listing has no registered native installer command. Use the complete package or source fallback below, depending on what your client supports.
Do not guess an installer command or replace an existing version without reviewing the diff.
2. Complete package recommended
Download the ZIP when available. It includes SKILL.md plus the references, security notes and version metadata.
No complete ProSkills package is published for this listing yet.3. Prompt-only
Copy the prompt above when the agent can read the stable page or when you want to adopt the workflow without installing a skill.
Need only the instruction file?
Download SKILL.md only if your client requires a single file. The complete ZIP is safer for a full installation because it preserves the references and release context.
No path installs or executes anything by itself. Your agent still needs access to the project files. Before updating, compare the installed version and review the diff.
// RATINGS
Not yet listed on ClawHub or SkillsMP
// README
agent-web-fetch
给 AI Agent 用的网页读取工具:先用最便宜的办法,被挡住了才升级;读不到时,明确告诉你为什么。
A web-reading toolkit for AI agents: cheapest first, escalate only when blocked, and always report why a page could not be read.
English · Python 库 · CLI (awf) · MCP Server (awf-mcp) · Agent Skill (skill/SKILL.md) · MIT
作者 Yomin Ma · GitHub @mrlong0129 · 姊妹项目:wechat-article-fetcher(专门读微信公众号文章) · 项目页
Quick start for agents
不用安装、不用配置,复制一行就能用(需要 uv;默认行为零配置):
# 1) 直接读一个网页(不安装)
uvx --from git+https://github.com/mrlong0129/agent-web-fetch awf fetch <url>
uvx --from git+https://github.com/mrlong0129/agent-web-fetch awf fetch <url> --format json # 结构化结果
# 2) Claude Code:一行注册 MCP server
claude mcp add agent-web-fetch -- uvx --from git+https://github.com/mrlong0129/agent-web-fetch awf-mcp
Cursor:写入 ~/.cursor/mcp.json(全局)或项目里的 .cursor/mcp.json:
{
"mcpServers": {
"agent-web-fetch": {
"command": "uvx",
"args": ["--from", "git+https://github.com/mrlong0129/agent-web-fetch", "awf-mcp"]
}
}
}
Agent Skill(不装包也能用,Agent 会照着手册用自己的工具读网页):
# Claude Code
mkdir -p ~/.claude/skills/agent-web-fetch && curl -fsSL https://raw.githubusercontent.com/mrlong0129/agent-web-fetch/main/skill/SKILL.md -o ~/.claude/skills/agent-web-fetch/SKILL.md
# Cursor
mkdir -p ~/.cursor/skills/agent-web-fetch && curl -fsSL https://raw.githubusercontent.com/mrlong0129/agent-web-fetch/main/skill/SKILL.md -o ~/.cursor/skills/agent-web-fetch/SKILL.md
给编码 Agent 的完整说明见 AGENTS.md。
输出长什么样
awf fetch https://mp.weixin.qq.com/s/<article-id>
---
title: "示例标题"
url: "https://mp.weixin.qq.com/s/<article-id>"
author: "示例作者"
published: "2026-10-05T23:59:51+08:00"
site_name: "示例公众号"
extraction_method: "wechat:content_noencode"
---
# 示例标题
正文(Markdown)……
读不到时,退出码为 2,stderr 里写明原因,例如 blocked_reason: login_required,以及每一步尝试了什么。
为什么要做这个
大多数给大模型用的"网页抓取"工具,逻辑都只有一步:发一个请求,拿到 HTML,转成文本。成功了很好;失败了,Agent 拿到的要么是一个报错,要么是一段"请完成验证"的页面文字,然后它会一本正经地"总结"这段验证页。
真实的网络不是这样的。同一个链接,换一个"访问者",拿到的可能完全是另一个页面:
- 微信公众号文章:从云服务器请求,经常返回"环境异常,完成验证后即可继续访问";用手机微信的身份请求,同一个链接返回完整文章。
- 很多新闻站:正文由 JavaScript 渲染,HTML 里只有一个空壳,但页面里藏着给搜索引擎看的 JSON-LD,里面就是全文。
- Next.js / Nuxt 站点:
<div id="__next"></div>是空的,正文在__NEXT_DATA__这段 JSON 里。 - YouTube:页面是重度 JS 应用,但视频标题和简介就在
ytInitialPlayerResponse里。 - 有的页面压根不该读:需要登录、付费墙、验证码。这时候 Agent 最需要的不是"再试一次",而是知道原因,好去问用户。
agent-web-fetch 把一个有经验的人读网页的思路,写成了一套 Agent 能直接调用的流程。
升级策略:先便宜,后昂贵
核心原则只有一句:一条路不通就换下一条,先用最便宜的办法,不行再上更重的;每一步都检查自己拿到的到底是不是正文。
┌──────────────────────────────────────────────────────────────────────┐
│ 0. 站点适配器 API GitHub README / X 公开嵌入接口 最便宜、最稳 │
│ 1. 普通 HTTP 请求 一套真实、自洽的浏览器请求头 │
│ 2. 换身份重试 桌面 Chrome → 手机 Safari → Android / 微信内置浏览器│
│ + 同站 Referer、Accept-Language、礼貌退避重试 │
│ 3. 每一步都做检测 Cloudflare?验证码?登录墙?付费墙?软 404?限流? │
│ 4. 多策略正文提取 trafilatura / readability / JSON-LD / JS 状态 / │
│ og: 元数据 / 站点专用规则,挑最丰富的那个 │
│ └ 正文太薄时 试 AMP 版本、RSS/Atom 订阅源 │
│ 5. 无头浏览器 Playwright(可选安装),真正把页面渲染出来 最贵 │
└──────────────────────────────────────────────────────────────────────┘
第 0 层:能用官方/公开接口,就别解析 HTML
如果一个网站本来就提供了不需要登录的公开接口,那它永远比抓网页更快、更稳、更礼貌。
- GitHub:仓库地址直接走
api.github.com拿 README 原文(Markdown),API 限额用完就退到raw.githubusercontent.com。 - X / Twitter:单条推文走嵌入推文用的公开接口
cdn.syndication.twimg.com,不行再走官方 oEmbed。个人主页、时间线、搜索需要登录,工具会直接报告login_required,不会去硬抓。
第 1 层:一次普通的请求,但要像个正常浏览器
很多工具默认的 User-Agent 是 python-requests/2.x,这本身就会被不少站点直接拒绝。我们发送的是一整套自洽的请求头:User-Agent、Accept、Accept-Language、sec-ch-ua 系列,彼此对得上,就像一个真的 Chrome。大部分公开页面在这一步就成功了,耗时几十到几百毫秒。
第 2 层:换一个"访问者"
被挡住时,最先要试的是换一个更合适的访问方式:
| 身份 | 适用场景 |
|---|---|
desktop_chrome | 默认,绝大多数网站 |
mobile_safari | 很多站点给手机版的是更轻、服务端渲染的页面 |
android_chrome | 同上,另一种常见移动客户端 |
wechat_ios | 微信内置浏览器。公众号文章就是为它设计的,所以微信适配器第一步就用它 |
从第二次尝试开始,会带上同站的 Referer(就像用户从网站首页点进来)。遇到 429 / 503 会读取 Retry-After,按指数退避礼貌地重试,并且对同一个站点的请求之间保持最小间隔。同一次抓取的多个尝试共享 Cookie,就像同一个浏览器会话。
第 3 层:每一步都检查"我拿到的到底是什么"
这是和普通抓取工具最大的区别。每次拿到响应,都会先判断它是不是正文,如果不是,给出一个机器可读的原因 blocked_reason:
blocked_reason | 含义 | 换身份有用吗 |
|---|---|---|
cloudflare_challenge | Cloudflare "Just a moment..." 等待页 | 可能(浏览器层可能通过) |
captcha | reCAPTCHA / hCaptcha / Turnstile / GeeTest / DataDome / 腾讯验证码 | 否:不解、不换身份绕,立即停止并报告 |
verification_wall | 环境验证页,例如微信"环境异常" | 可能 |
login_required | 需要登录才能看 | 否,立即停止 |
paywall | 付费墙(包括 JSON-LD 里 isAccessibleForFree=false) | 否,立即停止 |
rate_limited | 429 或"请求过于频繁" | 稍后再试 |
forbidden | 403 / 412 等,没有识别出具体的挑战页(例如风控) | 可能 |
not_found / soft_404 | 真 404,或 HTTP 200 但页面写着"找不到" | 否 |
content_removed | 内容已被删除 / 违规无法查看 | 否 |
js_required | 纯 JavaScript 空壳,需要浏览器渲染 | 交给浏览器层 |
robots_disallowed | robots.txt 禁止,且使用了 --robots strict | 否 |
两条细节:
- 关键词只在"页面很薄"时才算数。一篇讨论验证码的技术文章不会被误判成验证码页。
- Cloudflare 会往很多正常页面里注入
challenge-platform脚本,所以光有这个标记不算被挡,必须同时满足"挑战状态码"或"页面几乎没有内容"。
遇到验证码、登录墙、付费墙、404、已删除这类"换谁来都一样"的情况,梯子会立刻停下来,不浪费请求,也不去绕。
第 4 层:多策略提取,挑最丰富的,并记录谁赢了
拿到真正的页面之后,同时跑多种提取策略(它们都在本地运行,比再发一次网络请求便宜得多):
| 策略 | 擅长 |
|---|---|
trafilatura | 通用的正文提取(readability 类算法),去掉导航、页脚、广告 |
readability | Mozilla Readability 的 Python 移植,第二意见 |
json_ld | 新闻站给搜索引擎准备的 articleBody,常常就是全文 |
js_state | __NEXT_DATA__、__NUXT__、window.__INITIAL_STATE__、__APOLLO_STATE__、ytInitialPlayerResponse 等内嵌状态 |
meta | og: / twitter: 描述,最后的兜底(通常只是摘要) |
wechat:* 等 | 站点适配器自己的规则,例如公众号的 js_content / content_noencode / og:title |
按"文本长度 × 可信度权重"打分,选最高的。所有候选及其字数都写在 extraction_candidates 里,赢家写在 extraction_method 里,方便你调试,也方便 Agent 判断结果可不可靠。元数据(标题、作者、时间、站点名)则按字段合并:适配器 > JSON-LD > og 元数据 > trafilatura。
如果正文还是太薄,会尝试页面声明的 AMP 版本(<link rel="amphtml">),以及 RSS/Atom 订阅源里与当前链接匹配的那一条(订阅源本来就是给机器读的,又便宜又礼貌)。
第 5 层:真正打开一个浏览器
前面都不行、而且失败原因是"换个方式可能有用"的那一类时,才会启动 Playwright 无头浏览器,像普通访客一样把页面渲染出来,再走一遍检测和提取。它是可选依赖,没装就记录一条 skipped 然后跳过,不会报错。浏览器层不解验证码、不做交互式挑战;渲染后依然是验证码,就照实报告。
微信公众号:一个完整的例子
这个项目就是从读一篇公众号文章开始的。wechat 适配器把踩过的坑都写进去了:
- 第一个尝试就用 iPhone 微信内置浏览器的 UA,
Accept-Language: zh-CN。 - 识别"环境异常 / 完成验证后即可继续访问"的验证页 →
verification_wall;识别"该内容已被发布者删除""此内容因违规无法查看" →content_removed。 - 长文:正文在
#js_content里,图片地址在data-src里。 - 短图文(
item_show_type=10):#js_content是空的,全文在content_noencode这个 JS 变量里,og:title里通常也有一份完整的。 - 作者、公众号名、发布时间分别来自
author、nick_name、ori_create_time(没有时用create_time),时间统一输出为带+08:00的 ISO 格式。
安装
# 从 GitHub 安装(发布到 PyPI 后可直接 pip install agent-web-fetch)
pip install "git+https://github.com/mrlong0129/agent-web-fetch"
# 可选:无头浏览器层
pip install "agent-web-fetch[browser] @ git+https://github.com/mrlong0129/agent-web-fetch"
playwright install chromium # 或者系统里已有 Google Chrome 也可以,会自动尝试
# 不想装?用 uvx 直接跑
uvx --from git+https://github.com/mrlong0129/agent-web-fetch awf fetch https://example.com
需要 Python 3.10+。
使用
CLI
awf fetch <url> # 默认输出 Markdown(带 front matter)
awf fetch <url> --format json # 完整结构化结果
awf fetch <url> --format text # 纯文本
awf fetch <url> --browser # 直接用无头浏览器
awf fetch <url> --no-browser # 永远不用浏览器
awf fetch <url> --adapter none # 关闭站点适配器(auto | none | wechat | github | x)
awf fetch <url> --profile wechat_ios # 指定身份(可重复,按顺序尝试)
awf fetch <url> --robots strict # robots.txt 不允许就不抓(批量/爬取场景请用这个)
awf extract page.html --url <原链接> # 对本地 HTML 做检测 + 提取,不联网
awf adapters # 列出适配器
awf reasons # 列出所有 blocked_reason
退出码:0 成功,2 被挡或没有内容(原因打印在 stderr),1 其他错误。
JSON 输出字段:url、final_url、status、ok、`ti
// HOW IT'S BUILT
KEY FILES