⏳ This skill is pending AI review.

Scores will appear once the review pipeline completes.

v2.3.0

boss-zhipin-scraper

@eatmoreduck⭐ 1.5k stars

Scrape BOSS直聘 (job listing site) via Chrome CDP. Searches jobs by keyword/city/filters, fetches JD details, outputs structured JSON/CSV with plaintext salary, and can summarize scraped results into a job-market prompt. Use when user wants to search/analyze jobs on BOSS直聘 or zhipin.com.

Choose how to use this skill

You do not need every option. Choose the path your AI client supports. The stable page stays the same; versioned files are immutable.

1. Native installer

This listing has no registered native installer command. Use the complete package or source fallback below, depending on what your client supports.

Do not guess an installer command or replace an existing version without reviewing the diff.

2. Complete package recommended

Download the ZIP when available. It includes SKILL.md plus the references, security notes and version metadata.

No complete ProSkills package is published for this listing yet.

3. Prompt-only

Copy the prompt above when the agent can read the stable page or when you want to adopt the workflow without installing a skill.

Need only the instruction file?

Download SKILL.md only if your client requires a single file. The complete ZIP is safer for a full installation because it preserves the references and release context.

No path installs or executes anything by itself. Your agent still needs access to the project files. Before updating, compare the installed version and review the diff.

—/10

// RATINGS

⭐GitHub Stars
⭐⭐⭐⭐⭐ 1.5k on GitHubGitHub ↗

Very popular

🟢ProSkills Score
—
📍

Not yet listed on ClawHub or SkillsMP

// README

BOSS直聘爬虫 · 职位抓取工具 v2.3(Chrome/Edge CDP / 明文薪资)

🌐 English documentation: README.en.md

Python License Platform Version

一个轻量的 BOSS直聘爬虫(spider / crawler / scraper):通过 Chrome DevTools Protocol 连接本地已登录的 Chrome 或 Microsoft Edge,复用真实登录态调用 zhipin.com 搜索 API,绕过前端字体反爬,输出含明文薪资的职位数据(JSON / CSV),并生成薪资分布、技能词频和求职材料优化提示词。同时作为 Hermes Agent Skill 提供。

📌 一句话介绍:不用 Selenium/Playwright,直接通过 Chrome DevTools Protocol 连接本地已登录的 Chrome,复用真实登录态调搜索 API,输出含明文薪资的 JSON/CSV,并生成薪资分布、技能词频和求职材料优化提示词。

cover


⚠️ 免责声明

本项目仅供学习和技术研究参考,旨在探讨 Chrome DevTools Protocol、前端反爬机制与数据采集技术。请勿用于任何违反 BOSS直聘用户协议 或相关法律法规的用途,不得用于商业转售、恶意爬取或对目标网站造成负担的行为。使用本项目所产生的一切后果由使用者自行承担,作者不对任何滥用行为负责。


🚀 30 秒快速开始

# 1. 克隆 + 装依赖
git clone https://github.com/eatmoreduck/boss-zhipin-scraper.git
cd boss-zhipin-scraper
pip install -r requirements.txt          # 或 uv sync

# 2. 启动隔离 Chrome 并登录(只需一次,登录态持久保存)
python3 scripts/boss_cdp_raw.py --setup-chrome
# 也可以显式使用 Microsoft Edge:
# python3 scripts/boss_cdp_raw.py --setup-edge

# 3. 抓取 + 分析
python3 scripts/boss_cdp_raw.py --keyword "AI Agent" --city 上海 --pages 3 --analysis

# 支持全国城市(含三四五线),例如:
python3 scripts/boss_cdp_raw.py --keyword "前端" --city 赣州 --pages 3
# 查看支持的城市:--list-cities [关键词]
python3 scripts/boss_cdp_raw.py --list-cities 江

# 4. 抓取后生成聚合摘要 + 提示词(默认读取最新结果)
python3 scripts/job_summary.py

抓完直接拿到:薪资分布、经验要求、高频技能词、求职材料优化提示词。提示词只基于岗位数据,不读取本地简历文件,也不给岗位算个人匹配分。

✨ 特性

  • 明文薪资(API 模式,绕过字体反爬)
  • Boss 活跃状态独立字段(boss_active_status):列表兼容 bossOnline→「在线」,详情可得到「刚刚活跃」等更细状态
  • 跨轮新增标记(is_new):每条岗位带布尔字段标识「是否为本次新增」(终端同步以 🆕 显示),基准自动取同关键词同城市、日期早于本轮的最近一份结果;同一天多轮抓取折叠为一轮,--diff-base 可显式指定
  • JSON / CSV 双格式输出
  • 详情页 JD 抓取 + 技能分析
  • 抓取后聚合摘要 + 可复制提示词
  • 增量写入(异常退出不丢数据)
  • 一键环境检查 + 持久隔离 Chrome/Chromium CDP profile
  • 多维筛选(规模、融资、薪资、经验、学历、行业)
  • macOS + Linux + Windows 支持;macOS 会自动探测 Google Chrome 或 Chromium;Windows 已通过单元测试与基础 CLI 验证(GBK 控制台崩溃已修复),并支持 Chrome 或 Microsoft Edge CDP;真实抓取链路仍欢迎反馈
  • Selenium/Playwright 会启动完整的受控浏览器,体积大、指纹明显,容易触发 BOSS 的风控和验证码。
  • 本工具直接连接你已经登录的真实 Chrome(CDP),复用真实指纹和登录态,调用的也是页面内合法的搜索 API,返回的 salaryDesc 本就是明文——不需要解析被字体反爬加密的 DOM 薪资。
  • 因此比传统 DOM 抓取类爬虫更稳定,也更难被识别为自动化流量。

安装

方式 1:克隆到本地再安装(推荐)

由于 hermes skills install 的网络请求在某些环境下可能无法直接访问 GitHub,推荐先克隆仓库再本地安装:

# 1. 克隆仓库
git clone https://github.com/eatmoreduck/boss-zhipin-scraper.git
cd boss-zhipin-scraper

# 2. 复制到 Hermes skills 目录
mkdir -p ~/.hermes/skills/data-science/boss-zhipin-scraper/scripts
cp SKILL.md ~/.hermes/skills/data-science/boss-zhipin-scraper/
cp scripts/boss_cdp_raw.py ~/.hermes/skills/data-science/boss-zhipin-scraper/scripts/
cp scripts/job_summary.py ~/.hermes/skills/data-science/boss-zhipin-scraper/scripts/
mkdir -p ~/.hermes/skills/data-science/boss-zhipin-scraper/data
cp data/city_codes.json ~/.hermes/skills/data-science/boss-zhipin-scraper/data/

方式 2:curl 一键安装

不需要克隆整个仓库,直接下载必要文件:

mkdir -p ~/.hermes/skills/data-science/boss-zhipin-scraper/scripts && \
curl -sL https://raw.githubusercontent.com/eatmoreduck/boss-zhipin-scraper/master/SKILL.md \
  -o ~/.hermes/skills/data-science/boss-zhipin-scraper/SKILL.md && \
curl -sL https://raw.githubusercontent.com/eatmoreduck/boss-zhipin-scraper/master/scripts/boss_cdp_raw.py \
  -o ~/.hermes/skills/data-science/boss-zhipin-scraper/scripts/boss_cdp_raw.py && \
curl -sL https://raw.githubusercontent.com/eatmoreduck/boss-zhipin-scraper/master/scripts/job_summary.py \
  -o ~/.hermes/skills/data-science/boss-zhipin-scraper/scripts/job_summary.py && \
mkdir -p ~/.hermes/skills/data-science/boss-zhipin-scraper/data && \
curl -sL https://raw.githubusercontent.com/eatmoreduck/boss-zhipin-scraper/master/data/city_codes.json \
  -o ~/.hermes/skills/data-science/boss-zhipin-scraper/data/city_codes.json

方式 3:hermes skills install(需网络直连 GitHub)

hermes skills install https://raw.githubusercontent.com/eatmoreduck/boss-zhipin-scraper/master/SKILL.md --category data-science

注意:此方式依赖 hermes 进程能直接访问 GitHub,如果遇到超时或连接失败,请使用方式 1 或 2。

方式 4:skills.sh 一键安装(Claude Code 等 Agent Skills 兼容 agent)

npx skills add eatmoreduck/boss-zhipin-scraper

skills.sh 已收录本技能。任何支持 Agent Skills 格式的 agent(Claude Code、Codex、Gemini CLI、Cursor 等)都可以用这条命令安装,SKILL.md、脚本和城市码表随技能一起分发,按提示选择要安装到哪个 agent 即可。

方式 5:ClawHub 安装(OpenClaw / clawhub CLI)

本技能已上架 ClawHub(OpenClaw 生态的技能注册表),OpenClaw 用户可一条命令安装:

npx clawhub@latest install @eatmoreduck/boss-zhipin-scraper

验证安装

# 检查文件是否存在
ls ~/.hermes/skills/data-science/boss-zhipin-scraper/SKILL.md
ls ~/.hermes/skills/data-science/boss-zhipin-scraper/scripts/boss_cdp_raw.py
ls ~/.hermes/skills/data-science/boss-zhipin-scraper/scripts/job_summary.py
ls ~/.hermes/skills/data-science/boss-zhipin-scraper/data/city_codes.json

安装后直接在 Hermes 对话中说"帮我搜一下 BOSS直聘 上上海的 AI Agent 岗位"。

作为命令行工具使用

不想装成 Skill 也可以直接当 CLI 用:

# 1. 克隆 + 安装依赖
git clone https://github.com/eatmoreduck/boss-zhipin-scraper.git
cd boss-zhipin-scraper
pip install -r requirements.txt

# 2. 启动 Chrome CDP(也可改用 --setup-edge)
python3 scripts/boss_cdp_raw.py --setup-chrome
# macOS 会自动选择已安装的 Google Chrome 或 /Applications/Chromium.app
# 首次使用也不会复制主 Chrome 登录态;请在弹出的 BOSS 专用浏览器中登录 zhipin.com
# setup 会等待登录完成,并确认接口能返回明文薪资

# 3. 检查环境
python3 scripts/boss_cdp_raw.py --check

# 可选:真实浏览器/API smoke test(不写结果文件)
python3 scripts/boss_cdp_raw.py --smoke-test

# 4. 抓取
python3 scripts/boss_cdp_raw.py --keyword "AI Agent" --city 上海 --pages 3 --format csv --analysis

# 5. 抓取后摘要和提示词
python3 scripts/job_summary.py --top 15

参数

参数说明
--keyword搜索关键词(默认 "AI Agent")
--city城市(中文或 9 位代码,默认上海)。支持全国城市(一二三四五线全覆盖,共 300+ 个),运行时自动从 BOSS 同步最新城市码;码表见 data/city_codes.json,或用 --list-cities 查看。本地及在线码表均无法识别的城市名会报错退出,避免静默得到 0 条结果
--list-cities [关键词]打印支持的城市列表,可选关键词过滤,如 --list-cities 江
--pages页数(上限 10)
--formatjson / csv;csv 会同时导出列表和详情 CSV
--detail抓取详情页 JD(默认开启)
--no-detail不抓取详情页
--analysis分析报告
--merge FILE合并已有数据(按 job_id 去重)
--diff-base FILE指定 is_new 对比基准文件(默认自动取同关键词同城市、文件名日期早于本轮的最近一份结果;无基准时全体视为新增)
--allow-dom-fallbackAPI 无数据时允许降级 DOM 提取;默认关闭,薪资可能不可信
--check环境检查(CDP + 依赖 + 登录态)
--smoke-test用真实 Chrome/CDP 跑一次 BOSS 搜索 API smoke test,不写结果文件
--setup-chrome一键启动 Chrome CDP(持久隔离 profile)
--setup-edge一键启动 Microsoft Edge CDP(持久隔离 profile)
--browser配合 --setup-chrome 选择 chrome 或 edge(默认 chrome);--setup-edge 固定 Edge,显式传 --browser chrome 会以 --setup-edge 为准并提示
--copy-login-state手动导入主浏览器(--setup-chrome 取主 Chrome、--setup-edge 取主 Edge)的 Local State + Cookie 相关文件到隔离 profile(默认、首次启动、重复启动都不复制)
--reset-chrome-profile重建 BOSS 专用 Chrome profile,会清除此专用浏览器内的登录态
--no-wait-login--setup-chrome 启动后不等待登录完成
--login-timeout--setup-chrome 等待登录完成的秒数(默认 300)
--stop-chrome关闭 BOSS 专用 CDP Chrome(按隔离 profile 精准匹配,不碰主 Chrome)
--stop-edge关闭 BOSS 专用浏览器 CDP(与 --stop-chrome 共用隔离 profile)
--close-chrome抓取正常结束后自动关闭专用 Chrome(默认不关;异常退出不触发,保留登录态)
--output列表输出路径(默认 ~/.boss-zhipin-scraper/job-result/)
--detail-output详情输出路径(默认 ~/.boss-zhipin-scraper/job-result/)
--cdp-portCDP 端口(默认 9222)
--scale/--salary/--experience/--degree筛选条件

抓取后摘要与提示词

scripts/job_summary.py 只读取已抓取的 boss_jobs_*.json 和 boss_details_*.json,做简单聚合分析并生成一段可复制提示词。它不读取本地简历文件,不引入 PDF 依赖,也不给个人与岗位做分数判断。

# 读取默认结果目录下最新的 boss_jobs_*.json,并自动匹配同时间戳或最新详情文件
python3 scripts/job_summary.py

# 指定列表和详情文件

// HOW IT'S BUILT

KEY FILES

SKILL.mdREADME.md

// REPO STATS

1.5k stars