CLI
npx webforai converts a URL or a local HTML file to clean Markdown. It is
non-interactive by default: the Markdown goes to stdout and every log goes to stderr, so
the command composes with pipes, scripts and AI agents. The guided wizard is still there
behind -i.
# Markdown to stdout
npx webforai@latest https://example.com/article
# write to a file instead
npx webforai@latest https://example.com/article -o article.md
# convert a local HTML file
npx webforai@latest ./page.html
# machine-readable output
npx webforai@latest https://example.com --json | jq -r .markdownFlags
| flag | meaning |
|---|---|
-o, --output <path> | Write Markdown to a file instead of stdout. |
-l, --loader <loader> | fetch (default for URLs), playwright, or platform. Local files are detected automatically. |
-m, --mode <mode> | default or ai — ai strips links, tables and images down to plain text (local conversion only). |
--extractor <name> | Main-content extraction preset: auto (default: site adapters, then the learned kiwame classifier), kiwame, takumi (the heuristic, Readability-style extractor), minimal, or none for the whole page. The platform loader accepts all but kiwame. |
--frontmatter | Prepend YAML front matter built from the page metadata (title, author, canonical URL, …). |
--json | Print a JSON envelope instead of raw Markdown (below). |
--engine <engine> | Platform engine: auto (renders JavaScript only when the page needs it), fetch, browser, proxy-fetch, proxy-browser (proxy engines need a subscription). Implies -l platform. |
--region <region> | Platform proxy egress region: auto or jp (jp runs on the proxy tier, so it needs a subscription). Implies -l platform. |
--screenshot | Platform browser engines only: also capture a screenshot (expiring URL). |
--respect-robots | Platform loader only: honor the target site's robots.txt rules (off by default). A disallowed URL fails with robots_disallowed and is not billed. See robots.txt. |
--api-key <key> | Platform API key; defaults to $WEBFORAI_API_KEY. |
--platform-url <url> | Platform base URL; defaults to $WEBFORAI_PLATFORM_URL or the hosted instance. |
-i, --interactive | Guided prompt flow. |
-d, --debug | Verbose logs on stderr. |
Exit codes: 0 success, 1 runtime failure, 2 usage error. API failures include the
platform error code and retryAfter (seconds, when supplied)
on stderr; invalid_api_key, payment_required and spend_cap_reached name the platform
dashboard URL (your --platform-url when set), as
does the usage error for a missing API key. Explicit local
loaders cannot be combined with --engine or --region; this exits with code 2 before
making a request. Interactive mode applies the same flag validation and environment defaults.
Loaders
- fetch — plain HTTP fetch. Fast, no JavaScript execution. Follows
<meta http-equiv="refresh">redirect stubs (an HTTP 200 "Redirecting…" page) like ordinary redirects, up to 3 hops. - playwright — renders the page in local headless Chromium first. Install the browser once
with
npx playwright install chromium. - platform — the hosted webforai platform fetches and converts server-side. Useful for JavaScript-heavy pages without a local browser, bot-protected sites (proxy engines), region-pinned fetching, and screenshots.
export WEBFORAI_API_KEY=wfa_... # from https://platform.webforai.dev/dashboard
npx webforai@latest https://spa.example.com --engine auto
npx webforai@latest https://blocked.example.com --engine proxy-fetch --region jp # subscriptionBatch and crawl
Two subcommands run asynchronous platform jobs end to end: submit, wait (a live progress line on a terminal, quiet when piped), page through the results and write one Markdown file per page into a directory.
export WEBFORAI_API_KEY=wfa_...
# crawl a docs site (same-origin), plus llms.txt and llms-full.txt
npx webforai@latest crawl https://docs.example.com -o docs --limit 100 --include '^/docs' --llms-txt
# convert a list of URLs (arguments, or one per line with --file; "-" reads stdin)
npx webforai@latest batch https://a.example/post https://b.example/news -o pages
cat urls.txt | npx webforai@latest batch --file - -o pages --json| flag | meaning |
|---|---|
-o, --output <dir> | Required. Files mirror URL paths: /docs/intro → docs/intro.md, / → index.md; batch prefixes the host (a.example/post.md). A query string adds a short hash; collisions get _1, _2. |
--max-depth <n> | crawl: link depth from the seed, 0–5 (default 2). |
--limit <n> | crawl: maximum pages, 1–500 (default 50). |
--include <regex> / --exclude <regex> | crawl: pathname filters; repeat the flag for several patterns. |
--sitemap <mode> | crawl: skip (default), include (also queue the site's sitemap URLs) or only (seed + sitemap URLs, no link following). See crawl. |
--respect-robots / --no-respect-robots | crawl honors robots.txt by default; batch only with --respect-robots. |
--llms-txt | crawl: also write llms.txt (site title, summary and a - [title](url): description link per page, grouped by top-level path) and llms-full.txt (every page's Markdown, concatenated). |
-f, --file <path> | batch: read URLs from a file, one per line (# comments allowed); up to 100 URLs per job. |
--json | Print { jobId, status, credits, pages: [{ url, file, title, … }], failures: [{ url, code, message }], llmsTxt?, llmsFullTxt? } instead of the file list. |
--timeout <seconds> | Stop waiting after this long (default 1800). The job keeps running; results stay available for 7 days. |
--engine, --region, --extractor, --frontmatter, --api-key, --platform-url, -d | As above. |
stdout lists the written files, one per line (or the --json envelope); the progress and a
summary — including every failed page with its error code — go to stderr. Failed pages are not
billed and do not fail the command; it exits 1 when the job itself failed or no page was
converted. rate_limited responses are retried after Retry-After; too_many_jobs (3
concurrent jobs on the free tier, 20 with a subscription) is reported instead.
JSON output
--json prints a single envelope object:
{
"source": "https://example.com",
"loader": "platform",
"url": "https://example.com/",
"engine": "browser",
"region": "auto",
"markdown": "# …",
"metadata": { "title": "…" },
"credits": 2,
"output": "article.md"
}engine, region and credits appear only when the platform loader ran (screenshotUrl also with --screenshot, which adds 1 credit);
output only when -o also wrote a file. A warning appears when the result is probably
degraded — most commonly a client-rendered page fetched without JavaScript; rerun with
--engine auto or -l playwright when you see one.
For AI agents
The CLI ships an Agent Skill that teaches coding agents when and how to use it:
# print the SKILL.md
npx webforai@latest skill
# install it via Vercel's skills CLI (agent picker, symlinks, lockfile)
npx webforai@latest skill --install
# or install straight from the repository
npx skills add inaridiy/webforaiskill --install forwards --global, --agent <agents>, -y, --copy and --all to
npx skills add; --dir <path> writes the file directly without the skills CLI.
Interactive mode
npx webforai -i walks through source, loader (including the platform), mode and output
path with prompts.

