Abstract
Abstract
Source-based assessments can mislead when non-executing markup is mistaken for an active interface or equivalent HTML changes alter output metadata. We constructed 22 paired synthetic cases across inert markup, duplication, representation, accessible naming, intended sensitivity, and evidence boundaries. Each case declared an expected relation and its assumptions before execution. A deterministic local harness ran the analyzer without network collection, model calls, or human participants. Fifteen relations held and seven failed. Observed failures included hints inferred from commented or inert template scripts, score increases from commented metadata, case-sensitive confidence assignment, and full field-name credit for empty labels or unresolved references. The small white-box, purposively selected suite does not estimate real-world failure prevalence, predict agent task completion, or establish accessibility conformance.
22
paired synthetic cases
15 / 7
baseline relations passed / failed
22 / 22
same-suite post-fix relations passed
1 · Question and scope
Consistency of reported source evidence
isWebMCP extracts action candidates, source hints, category scores, and limitations from HTML. It does not execute the input application. This study therefore evaluates consistency of reported source evidence, not successful registration, authorization, task execution, end-user benefit, or WebMCP deployment readiness. The contribution is a reproducible set of counterexamples and preserved observations for one implementation; novelty and generality have not been established.
2 · Method
Twenty-two paired relations, six families
Metamorphic testing checks expected relationships between outputs for related inputs when a complete expected output is difficult to specify. The case catalog and protocol were written before the first execution, but after reviewing the scanner implementation and existing tests. This is white-box exploratory work, not preregistered confirmatory research.
| Family | Pairs | Pass / fail | Declared expectation |
|---|---|---|---|
| Inert markup | 6 | 2 / 4 | Comments, uninstantiated template content, and escaped code examples do not create active semantic evidence. |
| Duplication | 3 | 3 / 0 | Repeating existing signals can change raw counts without increasing semantic-diversity points. |
| Representation | 4 | 3 / 1 | Equivalent attribute order, tag case, action order, and entity encoding preserve normalized assessment. |
| Naming | 4 | 2 / 2 | Removing valid names lowers field-name points; empty labels and unresolved references receive no credit. |
| Intended sensitivity | 2 | 2 / 0 | A status region adds feedback evidence; an unencrypted entry hop lowers transport evidence. |
| Evidence boundaries | 3 | 3 / 0 | Registration stays a hint, truncated input retains uncertainty, and a requested goal cannot raise source-derived points. |
Inputs were owned HTML strings associated with a reserved.invalidURL. The analyzer was called directly; no URL was fetched. The harness fixed the analyzer date, computed UTF-8 byte lengths, and compared meaningful outputs while excluding random report IDs. It retained action confidence and risk, metric points, coverage, implementation classification, findings, and uncertainty. Every pair produced an analyzable result; there were no collection exceptions. The denominator is paired relations, not independently sampled applications, and many before-inputs are shared.
3 · Baseline findings
Seven declared relations failed
These are violations of the declared relations, not seven statistically independent root causes. Source inspection suggests three implementation-level explanations: raw-HTML metadata and script extraction bypassed inert-content filtering; action confidence used a case-sensitive tag-prefix check; and name detection credited attribute or relationship presence without establishing nonempty referenced text. Those explanations are grounded hypotheses, not a separate causal experiment across parsers.
- 1A registration-like script in an HTML comment changed implementation classification from no detected hint to a source hint; the aggregate score stayed 67.
- 2The same script in an uninstantiated template caused the same classification change, again with no score change.
- 3A commented JSON-LD block raised the score from 67 to 71 and counted as structured data.
- 4A commented title raised the score from 65 to 67 by contributing page-identity evidence.
- 5Changing the spelling of `button` to uppercase preserved the score of 46 but lowered action UI confidence from high to medium.
- 6Replacing a field’s only nonempty label with an empty associated label retained 55/55 field-name points and an aggregate score of 67.
- 7Replacing a valid `aria-labelledby` target with an unresolved reference retained 55/55 field-name points and an aggregate score of 49.
Preserved evidence boundaries
Duplicate buttons, status regions, and JSON-LD did not inflate tested score components. Ordinary label removal and restoration were detected. The tested inline-registration case gained no runtime quality or lift. Truncated input retained a prefix estimate, lowered confidence, and exposed a 0–100 full-page interval. A user-supplied goal did not raise the source score. These observations apply only to the tested cases.
4 · Separate intervention replay
The corrected version passed the same disclosed suite
After the baseline was recorded, the platform implementation was corrected without changing the case catalog, relation predicates, or protocol. A separately recorded run ofsource-actionability-v2.2satisfied all 22 relations; the seven formerly failing pairs passed and the 15 passing pairs remained passing. The source-hash comparison reports onlylib/scanner.tschanged among the eight tracked files.
This is targeted intervention evidence on the same disclosed cases, not held-out validation. It shows that these counterexamples were resolved under the recorded assumptions; it does not show generalization, calibrate scores against agent success, or support a quality certification.
5 · Limitations and next work
What these results do not establish
Selection and author bias
One product’s developers chose cases after reviewing its implementation; there was no random sample, blinded evaluation, or external relation review.
Construct validity
Score consistency is not accessible-name conformance, calibrated readiness, or agent task success.
Oracle assumptions
Relations cover simple fixtures; template activation, JavaScript strings, CSS visibility, malformed HTML, and alternative accessible names need more nuanced treatment.
Dependence
Cases share inputs, extraction logic, and relation families. Fractions describe this suite only; no population intervals or significance tests are justified.
Coverage
One analyzer version and local runtime; no browser execution, multilingual corpus, shadow DOM, authenticated state, network acquisition, or models.
Reproducibility
Hashes and local files identify artifacts, while the public page links to the protocol and results. The available evidence still needs independent review.
Next steps are independent review of the relations, validation of the simple DOM and naming assumptions against an isolated browser or standards-based oracle, held-out fixtures added before further implementation tuning, and comparison across multiple assessors with explicit decision rules. The one-time WRI v1 log was not accessed or modified; this synthetic dataset does not revise its interpretation.
Research artifacts
Inspect the protocol, data, and run records
The repository preserves the case catalog, protocol, runner, baseline and post-fix outputs, reports, and run manifests. The baseline remains unchanged after the fix. Results are local synthetic measurements, not field observations.
Generative-AI assistance and author responsibility
Generative AI was used substantially to inspect the implementation, propose and encode relations, implement the harness and provenance tools, interpret failures, suggest scanner corrections, and draft this manuscript. Measurements come from executed local software, not generated results. AI assistance is not independent review. Human authors must review the assumptions, code, sources, artifacts, and conclusions and take responsibility for any submitted version. No human-author sign-off or scholarly submission is claimed here.
Selected references
- T. Y. Chen, S. C. Cheung, and S. M. Yiu. Metamorphic Testing: A New Approach for Generating Next Test Cases. HKUST-CS98-01, 1998. Author-deposited copy.
- H. Liu, F.-C. Kuo, D. Towey, and T. Y. Chen. How Effectively Does Metamorphic Testing Alleviate the Oracle Problem? IEEE Transactions on Software Engineering 40(1), 4–22, 2014. DOI.
- WHATWG. HTML Living Standard: The template element. Specification.
- W3C. Accessible Name and Description Computation 1.2. Specification.
This is an initial positioning check, not a systematic literature review. A final novelty matrix and independent scholarly review would be needed before making broader research claims.