⏳ This skill is pending AI review.
Scores will appear once the review pipeline completes.
computer-use-action-picker
Pick a browser or desktop agent's next action by choosing among the actions actually on screen instead of inventing one. Candidates come from the accessibility tree, the task goes in instructions, and a calibrated confidence decides whether to click or hand the step back to your own model. Use when building or debugging computer use, browser automation or a web agent, or on "it clicked the wrong thing" and "how do I stop it looping".
Choose how to use this skill
You do not need every option. Choose the path your AI client supports. The stable page stays the same; versioned files are immutable.
1. Native installer
This listing has no registered native installer command. Use the complete package or source fallback below, depending on what your client supports.
Do not guess an installer command or replace an existing version without reviewing the diff.
2. Complete package recommended
Download the ZIP when available. It includes SKILL.md plus the references, security notes and version metadata.
No complete ProSkills package is published for this listing yet.3. Prompt-only
Copy the prompt above when the agent can read the stable page or when you want to adopt the workflow without installing a skill.
Need only the instruction file?
Download SKILL.md only if your client requires a single file. The complete ZIP is safer for a full installation because it preserves the references and release context.
No path installs or executes anything by itself. Your agent still needs access to the project files. Before updating, compare the installed version and review the diff.
// RATINGS
Not yet listed on ClawHub or SkillsMP
// README
classifier.dev
Zero-shot text classification. Plain text in, a label and a calibrated confidence out. No key, no signup. Up to a thousand texts per request.
curl https://classifier.dev/spam,not+spam/Win+a+free+iPhone
spam
curl "https://classifier.dev/?labels=spam,not+spam&text=Win+a+free+iPhone" # same call, query form
spam
Single Cloudflare Worker. No database, no framework, no build step beyond esbuild.
CLI
npm i -g classifier-dev
classify bug,feature,praise < feedback.txt
cli/ is a separate npm package (classifier-dev, bin classify): one
dependency-free Node file, tests against a mock API (npm test), semver with
its own CHANGELOG, released with npm run release patch|minor|major which
tags cli-v<version> and lets .github/workflows/publish-cli.yml publish
(needs an NPM_TOKEN repo secret). It talks to the API exactly like curl does.
Layout
src/index.ts routing, validation, tiers, LLM fallback chain, analytics
src/query.ts the GET query form, read and written with nuqs; the URL an error suggests
src/jev.ts TypeSafe's Jev: packs inputs into requests, reads probabilities
src/limiter.ts Durable Object: per-IP rate limiting
src/report.ts digest — Analytics Engine SQL -> Resend, flags model fallbacks
src/alerts.ts every 15 minutes; emails only when something is wrong
src/feedback.ts agent feedback, feedback.now protocol -> email
src/privacy.ts keyed pseudonyms: nothing kept points back at a caller
src/cost.ts per-request upstream spend, from the providers' own accounting
src/docs.ts the site (GET / and GET /benchmark), plain text
src/home.ts the same two documents rendered, for browsers only
src/ui.ts the shared look: markdown in a terminal
cli/ the `classify` command, published to npm as classifier-dev
eval/ benchmarks; read eval/README.md before quoting a number
finish-dns.sh one-shot DNS wiring, see below
wrangler.example.toml the Worker config, minus the account-specific ids
The site
curl classifier.dev prints plain text, exactly as it always has. A browser
sends Accept: text/html and gets the same document rendered — headings,
bracketed links, copy buttons — from src/home.ts. Nothing is duplicated: the
page is generated from DOCS and BENCHMARK at request time, so the text
stays canonical and the two cannot drift. ?format=text opts out by hand, and
both responses carry Vary: accept.
Agent feedback
Implements the feedback.now protocol (schema 1.1), so any agent that speaks it can report a problem without being told how:
GET /.well-known/agent-feedback.json what this host accepts
GET /api/v1/policy categories, severities, limits
POST /api/v1/feedback full structured report
POST /api/v1/observations lighter signal
POST /api/v1/feedback/{id}/attachments more evidence, later
GET /api/v1/receipts/{id} did it land, and was it any good
Accepted submissions are emailed to REPORT_TO. Reports are kept in KV for 90
days. A repeat of the same domain + surface + category + title is stored and
acknowledged as a duplicate but not emailed again, so one looping agent cannot
empty itself into the inbox; the hourly budget is 100 per IP and the remainder
comes back on every receipt.
Testimonials use signal.category: "testimonial". They require
reporter.agent_type and reporter.agent_description, so praise arrives with
enough context to understand what kind of agent benefited and how it works.
quality_score is a deterministic function of how complete the report is — an
agent can read the rule and write a better one next time. Nothing here calls
the classifier or Analytics Engine: this is where reports arrive saying those
are broken, so it must work when they do not.
Deploy
Merging to main deploys. .github/workflows/deploy.yml typechecks, runs the
Worker and CLI tests, runs wrangler deploy, and then asks the live service for
/v1/health and one classification, so a deploy that uploads a broken Worker
fails in CI rather than in somebody's terminal.
By hand, to try something before it is merged:
cp wrangler.example.toml wrangler.toml # once, then fill in your own ids
npx wrangler deploy
wrangler.toml is gitignored and holds the two values that are specific to one
Cloudflare account: account_id, and the STATS KV namespace id that
npx wrangler kv namespace create STATS hands back. The tracked
wrangler.example.toml carries everything else — crons, bindings, migrations —
so the deployment shape is in the repository and only the identifiers are not.
CI has no wrangler.toml, so .github/render-wrangler.mjs writes one from the
example and three repository secrets. That makes the example the deployed shape
rather than a copy of it: change a binding in wrangler.toml alone and CI keeps
deploying the old one.
Repository secrets the deploy needs:
CLOUDFLARE_API_TOKEN dash.cloudflare.com > My Profile > API Tokens >
Create Token > "Edit Cloudflare Workers"
CLOUDFLARE_ACCOUNT_ID the account_id from wrangler.toml
STATS_KV_ID the STATS namespace id from wrangler.toml
REPORT_TO where the daily digest goes
Secrets set with wrangler secret put live on the Worker, not in the script
bundle, so a deploy leaves them alone and CI never needs to know them.
Secrets the Worker reads: TYPESAFE_API_KEY, AI_GATEWAY_API_KEY (Vercel's
AI Gateway, which serves Jev on a free monthly credit; when set it is asked
first and TypeSafe catches what it refuses), OPENROUTER_API_KEY,
CONTEXT_API_KEY (context.dev, the chat's web search and page reads),
RESEND_API_KEY, CF_ANALYTICS_TOKEN, REPORT_KEY, PRIVACY_SALT. Add one
with npx wrangler secret put NAME; none of them are ever read from the
repository. src/index.ts lists the rest in the Env interface.
Secrets are compared with secretEquals (src/secrets.ts), never ===: a
plain comparison returns on the first wrong byte and tells a caller how much
of a guess was right.
The model
Both tiers answer from TypeSafe's Jev, a decision model rather than a language model: it takes a state and typed questions and returns a calibrated probability per option, in ~150ms. That shape is why the API can do three things the LLM version could not.
A thousand inputs per request. State is an array of {id, text} and each
input gets its own question, so the whole batch is one upstream call. The
documented limit is 64k tokens per request; jev.ts packs to a conservative
budget and runs the resulting requests eight at a time. Measured: 400 news
headlines classified in 650ms end to end, and packing 100 items scored the
same as sending them one at a time.
Confidence that means something. On 400 six-way emotion items, answers at
= 0.9 confidence were right 82% of the time and answers below 0.5 were right 29%. The previous model's logprob "confidence" put 87% of news items above 0.9 and was right on 68% of those. So
tier: "smart"now means: re-ask the single-label answers below 0.7 of a fast reasoning model and replace them, markedescalated: true. Nothing else changes. Which model matters: on exactly the items Jev is unsure about, deepseek-v4-flash, qwen3.7-flash and mercury-2.5 were no better than Jev; gemini-3.8-flash took news topics from 87.5% to 90.0% and emotion from 61.8% to 63.7%, so that is the chain. A frontier model (claude-fable-5.1) gets 72.3% / 90.7% at ~3x the price; the numbers are on /benchmark if that trade ever looks worth it.
Multi-label in one pass. One yes/no question per label, labels at >= 0.7 returned most-likely-first with the full score map. F1 0.887 on the seven-case set against 0.799 for the sweep-and-verify LLM cascade it replaced, i
// HOW IT'S BUILT
KEY FILES