⏳ This skill is pending AI review.

Scores will appear once the review pipeline completes.

version unknown

computer-use-action-picker

@mrmps⭐ 423 stars

Pick a browser or desktop agent's next action by choosing among the actions actually on screen instead of inventing one. Candidates come from the accessibility tree, the task goes in instructions, and a calibrated confidence decides whether to click or hand the step back to your own model. Use when building or debugging computer use, browser automation or a web agent, or on "it clicked the wrong thing" and "how do I stop it looping".

Choose how to use this skill

You do not need every option. Choose the path your AI client supports. The stable page stays the same; versioned files are immutable.

1. Native installer

This listing has no registered native installer command. Use the complete package or source fallback below, depending on what your client supports.

Do not guess an installer command or replace an existing version without reviewing the diff.

2. Complete package recommended

Download the ZIP when available. It includes SKILL.md plus the references, security notes and version metadata.

No complete ProSkills package is published for this listing yet.

3. Prompt-only

Copy the prompt above when the agent can read the stable page or when you want to adopt the workflow without installing a skill.

Need only the instruction file?

Download SKILL.md only if your client requires a single file. The complete ZIP is safer for a full installation because it preserves the references and release context.

No path installs or executes anything by itself. Your agent still needs access to the project files. Before updating, compare the installed version and review the diff.

—/10

// RATINGS

⭐GitHub Stars
⭐⭐⭐ 423 on GitHubGitHub ↗

Popular

🟢ProSkills Score
—
📍

Not yet listed on ClawHub or SkillsMP

// README

classifier.dev

Zero-shot text classification. Plain text in, a label and a calibrated confidence out. No key, no signup. Up to a thousand texts per request.

curl https://classifier.dev/spam,not+spam/Win+a+free+iPhone
spam

curl "https://classifier.dev/?labels=spam,not+spam&text=Win+a+free+iPhone"   # same call, query form
spam

Single Cloudflare Worker. No database, no framework, no build step beyond esbuild.

CLI

npm i -g classifier-dev
classify bug,feature,praise < feedback.txt

cli/ is a separate npm package (classifier-dev, bin classify): one dependency-free Node file, tests against a mock API (npm test), semver with its own CHANGELOG, released with npm run release patch|minor|major which tags cli-v<version> and lets .github/workflows/publish-cli.yml publish (needs an NPM_TOKEN repo secret). It talks to the API exactly like curl does.

Layout

src/index.ts    routing, validation, tiers, LLM fallback chain, analytics
src/query.ts    the GET query form, read and written with nuqs; the URL an error suggests
src/jev.ts      TypeSafe's Jev: packs inputs into requests, reads probabilities
src/limiter.ts  Durable Object: per-IP rate limiting
src/report.ts   digest — Analytics Engine SQL -> Resend, flags model fallbacks
src/alerts.ts   every 15 minutes; emails only when something is wrong
src/feedback.ts agent feedback, feedback.now protocol -> email
src/privacy.ts  keyed pseudonyms: nothing kept points back at a caller
src/cost.ts     per-request upstream spend, from the providers' own accounting
src/docs.ts     the site (GET / and GET /benchmark), plain text
src/home.ts     the same two documents rendered, for browsers only
src/ui.ts       the shared look: markdown in a terminal
cli/            the `classify` command, published to npm as classifier-dev
eval/           benchmarks; read eval/README.md before quoting a number
finish-dns.sh   one-shot DNS wiring, see below
wrangler.example.toml  the Worker config, minus the account-specific ids

The site

curl classifier.dev prints plain text, exactly as it always has. A browser sends Accept: text/html and gets the same document rendered — headings, bracketed links, copy buttons — from src/home.ts. Nothing is duplicated: the page is generated from DOCS and BENCHMARK at request time, so the text stays canonical and the two cannot drift. ?format=text opts out by hand, and both responses carry Vary: accept.

Agent feedback

Implements the feedback.now protocol (schema 1.1), so any agent that speaks it can report a problem without being told how:

GET  /.well-known/agent-feedback.json   what this host accepts
GET  /api/v1/policy                     categories, severities, limits
POST /api/v1/feedback                   full structured report
POST /api/v1/observations               lighter signal
POST /api/v1/feedback/{id}/attachments  more evidence, later
GET  /api/v1/receipts/{id}              did it land, and was it any good

Accepted submissions are emailed to REPORT_TO. Reports are kept in KV for 90 days. A repeat of the same domain + surface + category + title is stored and acknowledged as a duplicate but not emailed again, so one looping agent cannot empty itself into the inbox; the hourly budget is 100 per IP and the remainder comes back on every receipt.

Testimonials use signal.category: "testimonial". They require reporter.agent_type and reporter.agent_description, so praise arrives with enough context to understand what kind of agent benefited and how it works.

quality_score is a deterministic function of how complete the report is — an agent can read the rule and write a better one next time. Nothing here calls the classifier or Analytics Engine: this is where reports arrive saying those are broken, so it must work when they do not.

Deploy

Merging to main deploys. .github/workflows/deploy.yml typechecks, runs the Worker and CLI tests, runs wrangler deploy, and then asks the live service for /v1/health and one classification, so a deploy that uploads a broken Worker fails in CI rather than in somebody's terminal.

By hand, to try something before it is merged:

cp wrangler.example.toml wrangler.toml     # once, then fill in your own ids
npx wrangler deploy

wrangler.toml is gitignored and holds the two values that are specific to one Cloudflare account: account_id, and the STATS KV namespace id that npx wrangler kv namespace create STATS hands back. The tracked wrangler.example.toml carries everything else — crons, bindings, migrations — so the deployment shape is in the repository and only the identifiers are not.

CI has no wrangler.toml, so .github/render-wrangler.mjs writes one from the example and three repository secrets. That makes the example the deployed shape rather than a copy of it: change a binding in wrangler.toml alone and CI keeps deploying the old one.

Repository secrets the deploy needs:

CLOUDFLARE_API_TOKEN    dash.cloudflare.com > My Profile > API Tokens >
                        Create Token > "Edit Cloudflare Workers"
CLOUDFLARE_ACCOUNT_ID   the account_id from wrangler.toml
STATS_KV_ID             the STATS namespace id from wrangler.toml
REPORT_TO               where the daily digest goes

Secrets set with wrangler secret put live on the Worker, not in the script bundle, so a deploy leaves them alone and CI never needs to know them.

Secrets the Worker reads: TYPESAFE_API_KEY, AI_GATEWAY_API_KEY (Vercel's AI Gateway, which serves Jev on a free monthly credit; when set it is asked first and TypeSafe catches what it refuses), OPENROUTER_API_KEY, CONTEXT_API_KEY (context.dev, the chat's web search and page reads), RESEND_API_KEY, CF_ANALYTICS_TOKEN, REPORT_KEY, PRIVACY_SALT. Add one with npx wrangler secret put NAME; none of them are ever read from the repository. src/index.ts lists the rest in the Env interface.

Secrets are compared with secretEquals (src/secrets.ts), never ===: a plain comparison returns on the first wrong byte and tells a caller how much of a guess was right.

The model

Both tiers answer from TypeSafe's Jev, a decision model rather than a language model: it takes a state and typed questions and returns a calibrated probability per option, in ~150ms. That shape is why the API can do three things the LLM version could not.

A thousand inputs per request. State is an array of {id, text} and each input gets its own question, so the whole batch is one upstream call. The documented limit is 64k tokens per request; jev.ts packs to a conservative budget and runs the resulting requests eight at a time. Measured: 400 news headlines classified in 650ms end to end, and packing 100 items scored the same as sending them one at a time.

Confidence that means something. On 400 six-way emotion items, answers at

= 0.9 confidence were right 82% of the time and answers below 0.5 were right 29%. The previous model's logprob "confidence" put 87% of news items above 0.9 and was right on 68% of those. So tier: "smart" now means: re-ask the single-label answers below 0.7 of a fast reasoning model and replace them, marked escalated: true. Nothing else changes. Which model matters: on exactly the items Jev is unsure about, deepseek-v4-flash, qwen3.7-flash and mercury-2.5 were no better than Jev; gemini-3.8-flash took news topics from 87.5% to 90.0% and emotion from 61.8% to 63.7%, so that is the chain. A frontier model (claude-fable-5.1) gets 72.3% / 90.7% at ~3x the price; the numbers are on /benchmark if that trade ever looks worth it.

Multi-label in one pass. One yes/no question per label, labels at >= 0.7 returned most-likely-first with the full score map. F1 0.887 on the seven-case set against 0.799 for the sweep-and-verify LLM cascade it replaced, i

// HOW IT'S BUILT

KEY FILES

skills/computer-use-action-picker/SKILL.mdREADME.md

// REPO STATS

423 stars