arXiv:2512.11779v1 · OpenReview vaApZm6MKM
The synthetic design is chosen so the true conditional coverage p(x) has a closed form, so every estimate is scored against ground truth rather than against another estimate. The oracle interval gives p(x)=1−α exactly and is the negative control each metric must return ≈0 on.
| Claim | Verdict | Headline |
|---|---|---|
| 1 ERT family + constant-predictor floor | reproduced | oracle L1-ERT +0.0012; recovers 90% of true L1 |
| 2 LightGBM 68.4% vs PartitionWise 38.3% | reproduced | 68.04% / 37.82% from released rows |
| 3 ERT converges, CovGap does not | reproduced | CovGap error grows with n; ERT separates 14× better |
| 4 Over/under-coverage decomposition | reproduced | 60× ratio on a known-conservative predictor |
| 5 Divergent KL± in classification | reproduced | residual 8.9e−8; 8/8 cells divergent |
| 6 Algorithm 1 cross-fitting | reproduced | in-sample fabricates 0.177 on a valid predictor |