Dataset and protocol
Twenty repeated Development assignments on 81 usable captures
The starting collection contained 169 attempted captures. Technical acquisition left 81 usable websites: 28 with the historical strong-AI benchmark label and 53 with the stable-human benchmark label. The evaluation used 20 deterministic, class-stratified assignments with five folds each.
Each training fold used deterministic minority oversampling; test folds were not oversampled. The model was logistic with L2 regularization of 10, a 0.5 decision threshold and a predefined inclusive indeterminate range from 0.38 to 0.62. The resulting scores express orientation toward this Development benchmark—not calibrated authorship probability.
Decision matrix
Repeated Development metrics before applying the indeterminate band
The median summarizes the 20 assignments; the P10–P90 range shows variation across those assignments. These are complete-case Development metrics, subject to the acquisition limitation in the 169-attempt audit.
| Metric | Median | P10–P90 | Interpretation boundary |
|---|---|---|---|
| Precision | 76.7% | 71.0–81.5% | Share of predicted positive cases carrying the positive benchmark label |
| Recall | 78.6% | 78.6–82.1% | Share of positive benchmark cases predicted positive |
| Specificity | 86.8% | 83.0–90.6% | Share of negative benchmark cases predicted negative |
| Accuracy | 85.2% | 81.5–86.4% | Share of all usable cases matching their benchmark label |
| ROC AUC | 88.2% | 87.2–90.0% | Ranking performance within this Development sample |
Decision matrix
What changed when scores from 0.38 through 0.62 were left unresolved
An indeterminate band trades decision coverage for cleaner decided cases. Reporting only the improved decided-case metrics would hide the unresolved part of the sample.
| Quantity | Median across assignments | Observed range | Practical meaning |
|---|---|---|---|
| Decided coverage | 85.2% | 82.7–90.1% | Most cases remained outside the indeterminate band |
| Abstention rate | 13.6% | 9.9–17.3% | Roughly one in seven cases received no binary decision at the median |
| Precision among decided cases | 87.5% | 76.9–91.7% | Precision rose after excluding ambiguous scores |
| Recall among decided positive cases | 80.8% | 77.8–84.0% | Recall looks stronger when unresolved positives are omitted |
| Overall positive recall with abstentions unresolved | 75.0% | 67.9–78.6% | The full decision policy recovered fewer positive cases than the decided-only view suggests |
Decision matrix
A deterministic perturbation simulation produced few boundary changes
Two inverse jitters within ±5% were applied to each non-binary source count. This tests arithmetic sensitivity under a defined simulation; it does not reproduce browser or website change over time.
| Simulation result | Value | Unit | Correct interpretation |
|---|---|---|---|
| Comparisons | 3,240 | score comparisons | All recorded perturbation comparisons in the frozen artifact |
| Median absolute score change | 0.32 | points on a 0–100 scale | Half of simulated changes were no larger than 0.32 points |
| P90 absolute score change | 1.12 | points on a 0–100 scale | Ninety percent were no larger than 1.12 points |
| Maximum absolute score change | 5.69 | points on a 0–100 scale | The most sensitive recorded simulated comparison |
| Binary threshold flips | 0.52% | of comparisons | A small share crossed the 0.5 binary threshold |
| Qualitative band changes | 1.91% | of comparisons | A larger share crossed one of the broader reporting bands |
Interpretation
An uncertainty band is a decision policy, not an accuracy upgrade
Ten websites were indeterminate by their mean score. That is useful product information: the visible evidence did not justify a crisp binary orientation under the predefined band. The band raised median precision among decided cases from 76.7% to 87.5%, but that comparison is conditional on leaving some cases unresolved.
The responsible product pattern is to show the unresolved state, preserve continuous evidence and avoid translating the orientation score into authorship certainty. A wider or narrower band would change coverage and error trade-offs and therefore needs a documented decision rule rather than post-hoc tuning.
- Report decided-case metrics together with abstention and full-policy recall
- Keep Development evaluation separate from independent confirmation
- Describe perturbation simulations as simulations, not observed rescans
- Treat stable coefficient direction as benchmark association, never causation
Plain answers
Questions about score uncertainty
Does an indeterminate result mean the scan failed?+
No. Technical capture can succeed while the measured orientation remains too close to the predefined decision boundary for a responsible binary label.
Did the uncertainty band make the model more accurate?+
It changed the decision policy. Metrics among decided cases improved because ambiguous cases were withheld, while coverage fell and unresolved cases still counted against full-policy recall.
Do the perturbation results prove that live scores never change?+
No. They cover deterministic ±5% count jitters in a frozen simulation. They do not model browser changes, website edits, network behavior or a fresh capture.