VVibeFootprintWebsite intelligence

VibeFootprint research data · n=100

Why 82.4% precision was not enough to pass our model gate

The historical blind run reached 82.4% precision and 85.7% recall across 99 successful cases. We still marked the gate as failed because none of those legacy captures contained explicit completeness evidence.

Format
Integrity gate data brief
For
Model reviewers, product teams and buyers evaluating evidence quality
Reading time
8 minutes

Published by VibeFootprint EditorialPublished · Last reviewed

Dataset and gate

A 100-case blind run with two separate release requirements

The independent historical confirmation manifest contained 100 blind cases. Ninety-nine completed technically and one returned an error, giving 99% technical coverage. The pre-existing numerical gate required at least 80% precision and at least 80% recall.

The integrity reconstruction also required explicit proof that each evaluated capture was complete under its capture contract. That second requirement is not cosmetic: a model can appear to meet a threshold while relying on partial or unverifiable inputs.

Decision matrix

The 99 technically successful cases produced this confusion matrix

Counts are reproduced from the frozen aggregate artifact. They describe agreement with the blind benchmark labels, not website quality or generated-code share.

OutcomeCountPlain-language meaningContribution
True positive42Positive benchmark case predicted positiveSupports precision and recall
False positive9Negative benchmark case predicted positiveReduces precision and specificity
True negative41Negative benchmark case predicted negativeSupports specificity and accuracy
False negative7Positive benchmark case predicted negativeReduces recall

Decision matrix

The numerical metric gate passed

Both primary thresholds exceeded their predefined minimum. The overall release gate still failed because numeric performance was only one required condition.

MetricObservedRequired thresholdThreshold result
Precision82.4%At least 80%Passed
Recall85.7%At least 80%Passed
Specificity82.0%Not a primary gateReported
Accuracy83.8%Not a primary gateReported
F184.0%Not a primary gateReported

Decision matrix

The evidence-integrity gate failed

The legacy result rows did not persist explicit capture-completeness state, so completeness could not be reconstructed after the fact.

Integrity checkObservedRequiredDecision
Explicitly complete rows0Completeness evidence for evaluated capturesFailed
Unverifiable legacy rows99No unverifiable evaluated rowsFailed
Metric thresholdsBoth passedPrecision and recall at least 80%Passed
Overall release gateNot passedMetrics and capture integrityRejected

Repeatable process

What the failed gate changed in the research process

The practical lesson is to design integrity evidence into capture and evaluation—not to reconstruct it after a promising number appears.

  1. 01

    Persist capture completeness

    Record whether every required surface and artifact completed under the declared contract.

  2. 02

    Separate technical, metric and integrity gates

    Require each gate explicitly and prevent one successful metric from overriding another failed condition.

  3. 03

    Hash immutable inputs and outputs

    Bind the manifest, raw results and evaluation artifact to stable hashes before interpretation.

  4. 04

    Publish the negative decision

    Report the observed metrics while preserving the fact that the confirmation was not accepted.

Interpretation

Good-looking metrics cannot repair missing provenance

Rejecting the run does not mean its 99 predictions never happened. It means the evidence package was insufficient for the stronger claim that the model had passed confirmation under the complete protocol. The appropriate status preserves both facts: the recorded metrics crossed 80%, and the release gate did not pass.

This is the standard VibeFootprint applies to public claims as well. A score should not imply more than the retained evidence can support, and missing provenance should lower the claim—not be filled with a confident assumption.

Plain answers

Questions about the failed confirmation gate

Did the model pass the 80/80 metric threshold?

Yes. The reconstruction recorded 82.4% precision and 85.7% recall. The overall release gate still failed because verified capture completeness was also mandatory.

Why not accept the metrics and add the missing evidence later?

Completeness evidence describes the inputs used for those exact predictions. It cannot be reliably recreated after the historical captures when the required state was not persisted.

Are these current VibeFootprint production metrics?

No. They belong to a frozen historical v0.4 confirmation reconstruction evaluated on 15 August 2026 and should not be used as a current production-performance claim.

Source notes

References used for this guide

We prefer first-party standards, primary documentation and a visible interpretation boundary. Links are provided for verification and deeper implementation work.

Public blind-confirmation aggregate (JSON)

A domain-free public extract containing technical coverage, confusion counts, metrics, the gate decision and the frozen source-artifact hash.

Confirmation evaluation script

The evaluation logic showing that capture completeness is required in addition to the numerical precision and recall thresholds.

VibeFootprint methodology

Defines the public-surface evidence boundary and the responsible interpretation of a Vibe-Footprint.

Apply the framework

Review a real public website.

See its pattern-similarity index, evidence breadth, separate security baseline and concrete findings.

Buy launch scan · €4.99