Dataset and method
A frozen audit of attempted captures—not a fresh availability test
The audit aggregated 169 retrieval attempts from an existing browser-surface matrix. It did not revisit or rescan any website. Eighty-one attempts produced a technically usable capture and 88 did not, for an observed technical yield of 47.9%. The recorded Wilson 95% interval was 40.5% to 55.4%.
The distinction matters because a website can enter a scoring evaluation only after enough public surface has been captured. A failed capture is missing evidence, not a low score, a high score or a negative website judgment. We therefore report the acquisition funnel before reporting any downstream model metric.
- 169 retrieval attempts in the frozen historical matrix
- 81 usable captures and 88 technical failures
- No domains or individual website scores are disclosed
- No failed row was converted into a model prediction
Decision matrix
The complete technical outcome distribution
Shares use all 169 attempts as the denominator. The largest failure class cannot be resolved more precisely from the retained historical evidence.
| Recorded outcome | Count | Share of attempts | What may be concluded |
|---|---|---|---|
| Successful capture | 81 | 47.9% | Enough public surface was retained for the Development evaluation |
| Navigation timeout, unresolved | 81 | 47.9% | The historical navigation exceeded its budget; no lower-level reason was retained |
| Client blocked | 3 | 1.8% | The client recorded a blocking response for these attempts |
| DNS unresolved | 3 | 1.8% | Name resolution did not complete for these attempts |
| Certificate error | 1 | 0.6% | TLS certificate handling prevented this capture |
Decision matrix
Yield was not evenly distributed across the historical sample
These are benchmark and collection cohorts, not claims about all human-built or AI-assisted websites. The imbalance is relevant because analysis of successful rows alone inherits the capture process.
| Historical segment | Attempted | Successful | Technical yield |
|---|---|---|---|
| Stable-human benchmark label | 86 | 53 | 61.6% (95% interval 51.1–71.2%) |
| Strong-AI benchmark label | 83 | 28 | 33.7% (95% interval 24.5–44.4%) |
| Existing collection cohort | 38 | 37 | 97.4% (95% interval 86.5–99.5%) |
| Expansion collection cohort | 131 | 44 | 33.6% (95% interval 26.1–42.0%) |
Self-review
Four rules for trustworthy website-scan reporting
A scan product should make the evidence pipeline auditable before presenting a precise-looking score.
Is capture yield shown before evaluation metrics?
- Strong signal
- Attempts, usable captures and failure classes are reported with denominators.
- Weak signal
- Only successfully scored pages are shown, hiding how much of the target set disappeared.
Are unresolved failures kept unresolved?
- Strong signal
- The report says only that navigation timed out under the recorded budget.
- Weak signal
- A timeout is relabelled as offline, blocked or defective without retained evidence.
Could technical success depend on the benchmark group?
- Strong signal
- Yield is compared by relevant cohort and the downstream limitation is stated.
- Weak signal
- Complete cases are treated as if they were a random sample of all attempts.
Is operational metadata separated from model evidence?
- Strong signal
- Hosting suffix and failure metadata stay outside the scoring feature set.
- Weak signal
- Infrastructure or retrieval success becomes a proxy for the benchmark label.
Interpretation
The main result is a selection-bias warning, not a claim about website quality
Technical yield differed by 27.9 percentage points between the two historical benchmark labels. That does not explain why captures failed, and it does not establish a general property of AI-assisted websites. It does mean that model metrics calculated only on the 81 successful captures are exposed to label-dependent selection bias.
For product decisions, the remedy is procedural: preserve richer failure reasons, monitor capture yield by cohort, retry under a documented policy and keep failed retrievals outside the score. The current VibeFootprint result should describe only the evidence actually observed.
Plain answers
Questions about the 169-attempt audit
Does 47.9% mean the current VibeFootprint scanner fails half the time?+
No. It is the yield of one frozen historical research collection generated on 13 August 2026 without a rescan. It is not a current production-service uptime measurement.
Were the 81 timeouts websites that were offline?+
That cannot be concluded. The retained artifact contains no lower-level reason, so the correct label is unresolved navigation timeout—not offline, unreachable or blocked.
Why publish a weak technical-yield result?+
Because downstream model metrics are easier to overstate when the acquisition funnel is hidden. Publishing the negative result makes the evidence boundary and future collection requirements explicit.