⏳ This skill is pending AI review.

Scores will appear once the review pipeline completes.

version unknown

cuawright-web

@microsoft⭐ 6.0k stars

Solve a user-specified web task code-as-action style by driving a local Playwright browser through one bash command at a time, saving screenshots and an action log into `final_runs/run_<id>/`, and visually verifying the result. Use when the user asks to automate a web task (search, filter, form-fill, multi-step flow, data extraction) and wants reusable scripts plus screenshot evidence rather than a one-shot answer.

Choose how to use this skill

You do not need every option. Choose the path your AI client supports. The stable page stays the same; versioned files are immutable.

1. Native installer

This listing has no registered native installer command. Use the complete package or source fallback below, depending on what your client supports.

Do not guess an installer command or replace an existing version without reviewing the diff.

2. Complete package recommended

Download the ZIP when available. It includes SKILL.md plus the references, security notes and version metadata.

No complete ProSkills package is published for this listing yet.

3. Prompt-only

Copy the prompt above when the agent can read the stable page or when you want to adopt the workflow without installing a skill.

Need only the instruction file?

Download SKILL.md only if your client requires a single file. The complete ZIP is safer for a full installation because it preserves the references and release context.

No path installs or executes anything by itself. Your agent still needs access to the project files. Before updating, compare the installed version and review the diff.

—/10

// RATINGS

⭐GitHub Stars
⭐⭐⭐⭐⭐ 6.0k on GitHubGitHub ↗

Very popular

🟢ProSkills Score
—
📍

Not yet listed on ClawHub or SkillsMP

// README

CUAWright

Webwright is now CUAWright. Alongside browser automation, you can now run desktop tasks in an Ubuntu VM. Use cuawright web for browser tasks and cuawright desktop for desktop tasks. Existing webwright commands and Python imports still work.

CUAWright gives coding models a terminal to automate browsers and desktop applications. Use Webwright to browse the web and build reusable Playwright scripts, or run your own desktop task and reproduce OSWorld-V2 benchmarks in an Ubuntu VM.

Start with the browser quick start or desktop quick start.

Already got your favorite agents, and wonder how to make Claude Code, Codex, Hermes more capable in browser tasks? Consider adding CUAWright browser plugin/skill!


📰 News

  • 2026-10-03 — Webwright → CUAWright: renamed the project to bring browser and desktop agents together. The Webwright browser runtime is now cuawright.webwright, and OSWorld desktop support lives at cuawright.desktop. Existing webwright commands and imports remain compatible. See desktop setup.
  • 2026-09-01 — Persistent step-by-step browsing and native run_command tool calls improve performance to 88.1% on Online-Mind2Web and 77.5% on Odysseys.
  • 2026-07-21 — Skill Factory: every solve leaves a script behind, distilled into reusable, verified, parameterized code skills that rerun standalone with no model (~40 s, zero tokens). On WebArena, reuse lifts held-out accuracy 55% → 70% (+15 pp). See the optional Skill Factory.
  • 2026-05-11 — Support Task2UI mode: Webwright completes the task and renders task results into an HTML-based web app you can easily view and reuse.
  • 2026-05-06 — Codex and Claude Code plugin manifests added; install via /plugin install cuawright@cuawright. Hermes Agent integration shipped; the same skills/cuawright-web/ folder now loads across Claude Code, Codex, and Hermes.
  • 2026-05-04 — Initial public release: ~1.5k LoC, OpenAI / Anthropic / OpenRouter backends, Playwright environment.

Most web agents today treat the browser session itself as the workspace: at each step the model receives the current page state and predicts a single next operation — a click, a type, a DOM selector, or a short tool call. Whatever the format, the agent is locked into predicting one web action at a time inside a predefined interaction loop. That harness was useful when LLMs were weaker. As models get stronger at writing and debugging code, the same harness becomes a bottleneck.

Webwright takes a different stance: separate the agent from the browser, and treat the browser as something the agent can launch, inspect, and discard while developing a program. The persistent artifact is not the browser session — it's the code and logs in the local workspace.

  • 🧱 Robust, reusable interaction with web environments — instead of fragile pixel-level actions, a coding agent with a terminal queries elements, waits for conditions, and handles dynamic behaviors like lazy loading or re-rendering. The resulting scripts can be rerun, adapted, and shared across tasks rather than rediscovered from scratch.
  • ⚡ Efficient composition of complex workflows — multi-step interactions like selecting a date or filling a form become a compact program. Loops, functions, and abstractions let the agent generalize across similar tasks (e.g. different dates) without re-predicting the same low-level sequences. Fewer interaction rounds, faster execution, less error accumulation on long horizons.
  • 🧪 Workspace-as-state, not browser-as-state — the agent can write exploratory scripts, spawn fresh browser sessions, and decide for itself when to capture screenshots and inspect failures, much like a human engineer iterating on an RPA script.
  • 🪄 Surprisingly effective despite being minimal — this stripped-down setup turns out to handle complex and especially long-horizon web tasks well (see Performance).

Most web agent frameworks bury the actual agent loop under layers of abstractions. Webwright takes the opposite stance:

  • 🪶 Readable core — browser agent loop, environments, models, and CLI are separate modules under src/cuawright/webwright/.
  • 🧩 Pluggable model backends — OpenAI, Anthropic, and OpenRouter, each ~150–200 lines.
  • 🔍 Zero hidden frameworks — just httpx, pydantic, playwright, and typer.
  • 🔁 Flat prompt → observe → execute script loop — readable end-to-end, easy to debug, easy to fork.
  • 🧪 Run-artifact first — every run writes trajectories and screenshots to disk for inspection.

If you want a minimal, easy-to-debug starting point for browser-using agents instead of another heavyweight platform, this is it.


How they differ at the architectural level:

Stagehand (Browserbase)agent-browser (Vercel)browser-useWebwright
ParadigmHybrid: code + NL primitives (act / extract / agent)CLI tool that another agent (Claude Code, Codex, etc.) callsAutonomous LLM agent loop over DOM/AX snapshotsCoding agent with a terminal; browser is just an environment it spawns
Action spacePlaywright code, or NL → LLM-translated PlaywrightDiscrete subcommands (open, click @e2, snapshot, eval)Indexed click/type actions selected by the LLMFree-form Python (writes Playwright scripts itself)
What is "state"?The browser sessionThe browser session (held by daemon across CLI calls)The browser sessionThe local workspace — code, screenshots, logs. Browser is disposable.
Loop shapeImperative; agent() does multi-step when neededOne CLI invocation per micro-stepobserve → predict next action → execute → repeatwrite code → execute → inspect screenshots → repair (code-as-action)

🎥 Demo

CUAWright demo

https://github.com/user-attachments/

// HOW IT'S BUILT

KEY FILES

skills/cuawright-web/SKILL.mdREADME.md

// REPO STATS

6.0k stars