WebMCP Readiness Index · wri-v1

Methodology you can challenge

The index is designed to create testable hypotheses, not verdicts about companies. WRI v1 is now frozen and explicitly labeled uncalibrated: its raw records are preserved, its collection outage is quarantined, and its score limitations are published instead of silently rewriting history.

1 · Population

A reproducible popularity proxy

The target population is the first 100,000 domains in Tranco list GQJJK, dated 2026-08-31. Tranco is a research-oriented aggregate ranking. We call this “popular domains,” not measured traffic, visitors, or market share.

2 · Collection

One bounded, public homepage

One public homepage per ranked domain. HTTPS is attempted first; HTTP is a fallback. Redirects, response bytes, per-request time, and total time are capped.

Static response source only; no JavaScript execution, login, cookies, or screenshots. Page contents are analyzed in memory and discarded; the index retains only derived counts, scores, fetch metadata, and coverage state.

The research crawler checks robots.txt before fetching a homepage and records exclusions. A dedicated contact-bearing user agent is used. A root exclusion, 401/403 response, or temporary server error is not scored.

DNS and every redirect target are checked against private, reserved, local, and special-purpose address ranges before a request is made.

3 · Post-crawl integrity audit

100,000 scheduled is not 100,000 valid observations

The runner wrote 100,000 unique, contiguous rank/domain records. A post-run audit found two scanner-wide UPSTREAM_FAILURE windows—ranks 45,557–56,180 and 76,005–100,000—covering 34,620 records. Their speed and uninterrupted zero-success batches show a collection failure, not domain-specific evidence.

100,000

Scheduled ranks

65,380

Valid collection outcomes

34,620

Quarantined collection errors

Raw NDJSON remains immutable. The checked-in audit manifest can only remap rows inside those exact ranges when their original state and error code match. The public status is therefore audited partial, never “full corpus.”

4 · Frozen WRI v1 opportunity formula

Reproducible does not mean validated

55%

Workflow signal

action candidates × 13 + forms × 10 + controls × 2, capped at 100.

35%

Baseline friction

100 minus the source-only baseline actionability score.

10%

Implementation gap

100 when no static source hint appeared; 35 when one did.

opportunity = round(0.55 × workflow + 0.35 × friction + 0.10 × implementation gap)

Bands: 70–100 high leverage; 45–69 strong candidate; 20–44 foundation; 0–19 low interaction. The score says where a WebMCP investigation may be valuable. It does not evaluate business quality, security, accessibility conformance, or a production WebMCP implementation.

Known v1 validity limits: raw control volume can saturate workflow signal even with no named action; missing score dimensions were renormalized; absence of a static hint was treated as an implementation gap even though runtime tools may exist; and the model was not calibrated against an independently human-rated holdout corpus. Across the scored corpus, only 69 distinct integer opportunity scores exist. The audited UI now gives equal scores the same dense score rank and keeps Tranco popularity separate.

5 · What “after” means

No counterfactual theater

After-state visuals are illustrative design targets, never measured outcomes for third-party sites. A real after score requires a developer-provided or runtime-observed tool inventory plus paired journey runs with the same goal, state, fixtures, and assertions. Until that evidence exists, lift remains unknown.

6 · Access, corrections, and limitations

A homepage snapshot is not the whole product

  • Client-rendered, authenticated, regional, and consent-gated actions may be absent.
  • Source hints do not prove tools registered or executed at runtime.
  • Redirected domains may converge on the same final homepage.
  • Results change as sites, network paths, and the source list change.
  • WRI v1 classified 3,964 responses above its 640 KB cap as unreachable. The current quick scanner instead analyzes a bounded prefix and labels it partial; the frozen v1 rows are not retroactively rescored.
  • The coarse v1 robots bucket can include an explicit exclusion, access denial, or robots-fetch failure; it is not proof that a site owner deliberately blocked this research.

The server-rendered review slice contains 1,250 records for page performance and is not a representative sample. The explorer can load and verify the complete compressed 100,000-record attempt log on demand.

Site owners can request a correction, exclusion, or rescan by opening a private repository issue or emailing research@iswebmcp.com. Include the domain, snapshot list ID, and evidence.

← Return to the index