<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>Proceedings of Machine Learning Research</title>
    <description>Proceedings of the Fifteenth Symposium on Conformal and Probabilistic Prediction with Applications
  Held in M\&quot;olnlycke Health Care, M\&quot;olndal, Sweden on 02-04 September 2026

Published as Volume 329 by the Proceedings of Machine Learning Research on 30 August 2026.

Volume Edited by:
  Ernst Ahlberg
  Ulf Johansson
  Henrik Boström
  Alberto Carlevaro
  Johan Hallberg Szabadváry
  Lars Carlsson

Series Editors:
  Tegan Emerson
  Hoel Kervadec
  Neil D. Lawrence
</description>
    <link>https://proceedings.mlr.press/v329/</link>
    <atom:link href="https://proceedings.mlr.press/v329/feed.xml" rel="self" type="application/rss+xml"/>
    <pubDate>Sun, 30 Aug 2026 17:02:24 +0000</pubDate>
    <lastBuildDate>Sun, 30 Aug 2026 17:02:24 +0000</lastBuildDate>
    <generator>Jekyll v3.10.0</generator>
    
      <item>
        <title>What Can Conformal Risk Control Certify for PET/CT Tumour Segmentation?</title>
        <description>A conformal guarantee for segmentation is useful only when the controlled loss reflects the clinical decision. In PET/CT tumour imaging, missing a whole lesion can matter more than small boundary errors. We apply connected-component risk-controlling prediction sets (RCPS) to fixed, pretrained LesionTracer probability maps for 900 autoPET cases, asking whether one global threshold can certify a binary missed-lesion loss and a continuous voxel-level missed-tumour loss. At $\alpha$ = $\delta$ = 0.1, the binary loss cannot be certified: on lesion-positive calibration cases its empirical risk is 0.177 even at the smallest threshold we evaluated. The voxel loss behaves differently: an empirical-Bernstein bound certifies $\hat{\lambda} = 0.78$, giving P(Rvox($\hat{\lambda}$) $\leq$ 0.1) $\geq$ 0.9, once lesion-free cases contribute zero loss. This binary infeasibility holds for coverage requirements $\gamma$ from 0.5 to 0.9. We read the contrast as a statement about the prediction family: a global threshold cannot recover lesions the network never supported.</description>
        <pubDate>Sun, 30 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v329/zheng26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v329/zheng26a.html</guid>
        
        
      </item>
    
      <item>
        <title>When to Trust Simulated Data: Conformal Prediction for Selective Synthetic Data Integration</title>
        <description>Integrating simulated data into object detection training is a scalable strategy for alleviating data scarcity, yet simulated samples may exhibit geometric inaccuracies and domain artifacts that can degrade model performance when naively combined with real data. Rather than treating all synthetic samples as equally beneficial, we frame synthetic data inclusion as a reliability-aware selection problem and propose a conformal prediction-based framework to address it. The key insight is that a well-behaved synthetic sample should produce consistent object localizations across detectors trained on real and mixed data; violations of this consistency are captured as a nonconformity score, calibrated on a held-out set drawn from both real and synthetic data to derive a principled acceptance threshold. We evaluate on two benchmark datasets spanning autonomous driving and underwater sonar imaging, demonstrating consistent improvements in precision and recall across multiple model architectures. Explainability analysis further indicates that conformal selection yields more spatially coherent model attention, providing interpretable evidence that filtering acts on genuine domain artifacts rather than spurious correlations.</description>
        <pubDate>Sun, 30 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v329/yondem26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v329/yondem26a.html</guid>
        
        
      </item>
    
      <item>
        <title>Conformalized Time Series Anomaly Thresholding with Latent Space Features</title>
        <description>Time series anomaly detection is an important task applicable to a wide range of fields. Many methods have been proposed to address this task, but evaluation often suffers leakage from flawed thresholding methods. In this paper we explore leveraging latent space features in a conformal framework to automatically determine alarm thresholds without the risk of leakage and reducing false alarm rates. We show that our method is able to match or surpass the performance of a popular problematic thresholding method while using only information from the calibration set.</description>
        <pubDate>Sun, 30 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v329/xu26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v329/xu26a.html</guid>
        
        
      </item>
    
      <item>
        <title>Finite-Sample-Corrected Conformal Test Martingales for Wireless Link Adaptation</title>
        <description>Conformal test martingales (CTMs) bet on conformal p-values to accumulate evidence against exchangeability. With finite calibration, a threshold event has finite-grid uncertainty around its nominal probability. We use the Dvoretzky-Kiefer-Wolfowitz inequality to bound this probability and subtract the largest admissible positive excess expected gain. Under the stated continuous score null and a predictable betting schedule, the corrected single-score wealth is a nonnegative supermartingale on the calibration event. A simulated wireless replay illustrates how the corrected evidence process enters an online monitoring pipeline for alarms and retraining decisions.</description>
        <pubDate>Sun, 30 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v329/wu26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v329/wu26a.html</guid>
        
        
      </item>
    
      <item>
        <title>Aggregation in conformal e-classification</title>
        <description>Aggregating conformal predictors is a standard way of balancing their predictive and computational efficiency while retaining their validity, at least approximately. An important advantage of conformal e-predictors is that they are easier to aggregate without sacrificing their validity. This paper studies experimentally cross-conformal e-prediction, which is an existing method of aggregating conformal e-predictors, and its modifications that are conceptually simpler and more flexible.</description>
        <pubDate>Sun, 30 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v329/vovk26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v329/vovk26a.html</guid>
        
        
      </item>
    
      <item>
        <title>CRPS-Optimal Binning for Univariate Conformal Regression</title>
        <description>We propose a non-parametric, model-free method for conditional distribution estimation based on partitioning covariate-sorted observations into contiguous bins and using the within-bin empirical CDF as the predictive distribution. Bin boundaries are chosen to minimise the total leave-one-out Continuous Ranked Probability Score (LOO-CRPS), which admits a closed-form cost function with O(n2 log n) precomputation and O(n2) storage; the globally optimal K-partition is recovered by a dynamic programme in O(n2K) time. We select K by K-fold cross-validation of test CRPS, which yields a U-shaped criterion with a well-defined minimum. Having selected K* and fitted the full-data partition, we form a conformal prediction set based on CRPS as the nonconformity score, which carries a finite-sample marginal coverage guarantee at any prescribed level $\varepsilon$. The conformal prediction is transductive and data-efficient, as all observations are used for both partitioning and p-value calculation, with no need to reserve a hold-out set. On real-world small-data benchmarks against split-conformal competitors (Gaussian split conformal, CQR, CQR-QRF, CHR, and conformalized isotonic distributional regression), the full-n method produces narrower prediction intervals while maintaining near-nominal coverage; the advantage is not as clear-cut however in a matched-sample comparison restricted to the same training half as the competitors. Similarly, a comparison against QOOB, another transductive competitor, provides mixed results.</description>
        <pubDate>Sun, 30 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v329/toccaceli26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v329/toccaceli26a.html</guid>
        
        
      </item>
    
      <item>
        <title>Weighted Conformalized Quantile Regression for Active Learning</title>
        <description>A common approach in active learning is to improve models by acquiring samples for which the prediction is uncertain (Ren et al., 2021). Conformal quantile regression seems like a perfect match for estimating such prediction uncertainties, but any acquisition strategy used in the active learning setting would introduce selection and hence, violate the exchangeability assumption on which conformal prediction relies on. This paper addresses the challenge of maintaining valid uncertainty quantification in active learning for regression by utilizing Weighted Conformalized Quantile Regression with a Weighted CV+ approach. This restores predictive guarantees by re-weighting nonconformity scores based on the likelihood ratio between training and test distributions. Experimental results on the California Housing dataset demonstrate that uncertainty-based acquisition strategies accelerate error reduction (MSE), and that the weighting to maintain target coverage is required in order to keep the guarantees. Furthermore, the analysis identifies that there is a risk that greedy acquisition functions over-sample regions of irreducible noise, and demonstrates that a stochastic weighted strategy effectively mitigates this risk while balancing manifold exploration with predictive precision.</description>
        <pubDate>Sun, 30 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v329/stahl26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v329/stahl26a.html</guid>
        
        
      </item>
    
      <item>
        <title>Corrected Max-Rank for Finite-Sample Conformal Prediction</title>
        <description>Max-Rank is a dependence-aware family-wise error rate correction method used to construct hyperrectangular prediction regions in conformal multi-target regression. It ranks dimension-wise internal nonconformity scores and uses the largest dimension-wise rank as a joint nonconformity score, often achieving smaller regions than Bonferroni correction. In this paper, we revisit the finite-sample validity of Max-Rank and distinguish between the Max-Rank rule in the rank space and its implementation in the dimension-wise score space. The rank-space rule is valid. Its original score-space implementation, however, is incorrect and loses the finite-sample coverage guarantee. The implementation translates a rank threshold q into the qth dimension-wise calibration score. We show that this threshold is one order statistic too small and that the correct score threshold is the (q + 1)st dimension-wise calibration score, with the usual +$\infty$ convention when no large enough order statistic exists. We propose Corrected Max-Rank, a straightforward modification that uses the correct rank-to-score implementation. We prove that the corrected implementation satisfies the desired finite-sample coverage guarantee under exchangeability and dimension-separable internal scores. Experiments on synthetic multidimensional internal score distributions confirm that the original implementation can exceed the targeted error rate, especially for small calibration sets and a larger number of output dimensions. Corrected Max-Rank restores validity at the cost of slightly wider prediction regions.</description>
        <pubDate>Sun, 30 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v329/schlembach26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v329/schlembach26a.html</guid>
        
        
      </item>
    
      <item>
        <title>Conformal Poset Prediction for Label Ranking</title>
        <description>In supervised label ranking, an input X $\in$ X is associated with a ranking $\Pi$ $\in$ Sm of a fixed label set Y = {1,...,m} (Fürnkranz and Hüllermeier, 2010). Write a $\succ$$\pi$ b when $\pi$ ranks a before b, and identify $\pi$ with {(a,b): a $\succ$$\pi$ b}. A ranker typically returns one permutation $\hat{\sigma}(X)$. Such a point prediction suppresses uncertainty: a model may be confident that a precedes c while having little evidence for the comparison between a and b, adjacent in $\hat{\sigma}(X)$. A natural output is therefore a strict partial order, which asserts only selected pairwise preferences and leaves the remaining pairs incomparable1. Our goal is to predict an informative partial order R(x) whose asserted comparisons are simultaneously correct, that is, P{R(Xn+1) $\subseteq$ $\Pi$n+1} $\geq$ 1 - $\alpha$.</description>
        <pubDate>Sun, 30 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v329/sale26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v329/sale26a.html</guid>
        
        
      </item>
    
      <item>
        <title>Scaling-Score Conformal Prediction for Multi-Target Regression</title>
        <description>Multi-target regression requires a model to simultaneously predict several related outputs. Conformal prediction provides distribution-free, finite-sample marginal coverage guarantees, but extending these to joint multi-dimensional regions in a model-agnostic, sample-efficient manner remains challenging: max-aggregation ignores scale differences, copula-based methods are only asymptotically valid, rectangular methods typically split the calibration set, and quantile- or density-based methods require training a specialised model beyond a plain point predictor. We propose the scaling-score conformal method, which is model-agnostic (requires only component-wise absolute residuals), uses a single calibration set, and yields four nested output types: an outer rectangle (SCO) with valid joint coverage, the exact set R$\alpha$, a staircase (SC2) over-approximation of R$\alpha$, and an inner rectangle (SCI). A single hyperparameter $\gamma$ $\in$ (0,1) controls the base-rectangle quantile level independently of $\alpha$. We prove downward-closedness and a rectangular sandwich bound and derive a closed-form outer rectangle. Experiments on 29 real-world datasets confirm valid joint coverage; SC2 with $\gamma$ = 1 - $\alpha$ consistently achieves competitive volume relative to baselines, with the advantage growing with output dimension d.</description>
        <pubDate>Sun, 30 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v329/rousseau26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v329/rousseau26a.html</guid>
        
        
      </item>
    
      <item>
        <title>Probabilistic Object Detection with Conformal Prediction</title>
        <description>Conformal Prediction (CP) is a distribution-free method for constructing prediction sets with marginal finite-sample coverage guarantees, making it a suitable framework for reliable uncertainty quantification in safety-critical object detection. However, object detection introduces structured multi-output predictions, complicating the application of classical CP theory developed for single outputs. In addition, standard, unscaled CP produces fixed-width prediction intervals across inputs, leading to unnecessary width for low-uncertainty predictions. While scaled CP addresses this by adapting the interval width to an input-dependent uncertainty estimate, prior work has neither systematically compared unscaled and scaled CP for multi-class object detection, nor integrated CP with a complementary uncertainty quantification method in this setting. We fill this gap by: (i) applying CP coordinate-wise to bounding box corners with a Bonferroni correction for box-level guarantees; (ii) scaling the resulting intervals using per-prediction aleatoric uncertainty estimates derived from a probabilistic object detector trained with loss attenuation, evaluated in uncalibrated and two calibrated variants; (iii) extending to a two-step pipeline that constructs prediction sets for the class using Regularized Adaptive Prediction Sets (RAPS) and conditions the conformalized bounding boxes on the predicted class set. Across three autonomous driving datasets (KITTI, BDD, CODA), including a cross-domain setting under distribution shift, scaled CP consistently improves interval sharpness over unscaled CP, achieving up to 19% higher IoU and 39% lower interval scores, without sacrificing coverage. Class-wise calibration further improves coverage for both variants with a negligible effect on sharpness. Together, these improvements yield more actionable uncertainty estimates for real-time, real-world object detection. Code is available at https://github.com/mos-ks/OD-CP.</description>
        <pubDate>Sun, 30 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v329/ries26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v329/ries26a.html</guid>
        
        
      </item>
    
      <item>
        <title>TRACE-CRC: Trajectory-Adaptive Conformal Risk Control for Multi-Step Channel State Information Prediction</title>
        <description>Reliable prediction of time-varying channel state information (CSI) is essential for efficient wireless communication. Each CSI frame is a matrix-valued representation of the wireless channel response, and a sequence of CSI frames forms a temporal channel trajectory. Modern deep learning-based CSI predictors, however, often provide only point predictions and lack calibrated uncertainty estimates. This limitation is particularly problematic in multi-step CSI prediction, where the target is a sequence of future CSI matrices, and downstream decisions such as beamforming or scheduling may fail if any part of the predicted trajectory is unreliable. We propose trajectory-adaptive calibration and error profiling with conformal risk control (TRACE-CRC), a method for trajectory-aware uncertainty quantification in multi-step CSI prediction. TRACE-CRC constructs Frobenius-norm uncertainty balls around predicted CSI matrices and controls the risk that at least one future frame is uncovered. Instead of calibrating each future step independently, TRACE-CRC combines future-step-dependent error profiling, trajectory difficulty stratification, and learn-then-test (LTT) risk control. Empirically, TRACE-CRC achieves reliable trajectory-level coverage with substantially smaller uncertainty balls than conservative multi-step corrections, while avoiding the trajectory undercoverage of compact stepwise and adaptive conformal base-lines.</description>
        <pubDate>Sun, 30 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v329/rezaei26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v329/rezaei26a.html</guid>
        
        
      </item>
    
      <item>
        <title>Too Good to Conform: The Inverse Quality Paradox in Counterfactual Filtering</title>
        <description>As machine learning models enter high-stakes decision-making, explanations must be interpretable, reliable, and uncertainty-aware (Löfström et al., 2024, 2026). Counterfactuals (CFs) are a prominent explanation form, but their plausibility remains difficult to guarantee (Guidotti, 2024). Conformal prediction (Vovk et al., 2022) addresses this through generate-then-filter, which validates candidates post-hoc (Bates et al., 2023; Maalej et al., 2025; Adams et al., 2025), or filter-then-generate, which incorporates conformal constraints into generation (Altmeyer et al., 2024; Bilkhoo et al., 2025). We study the former. Although nonconformity-measure (NCM) design is actively studied in other conformal settings (Narteni et al., 2026), its role as a structural failure point in counterfactual filtering remains unexplored. We show that an NCM assigning greater nonconformity to candidates closer to xi can systematically reject the most proximal valid CFs — as generator quality (proximity to xi among candidates achieving the desired prediction change) improves, nonconformity increases, producing an inverse quality paradox. Our contributions are to: (i) formalise this failure for the reciprocal-distance score Aprox; (ii) show that it can reject reference-supported candidates while accepting candidates outside the data-supported region; and (iii) demonstrate that a score AR anchored to a fixed in-distribution reference set removes this dependence on xi and retains standard conformal validity under exchangeability.</description>
        <pubDate>Sun, 30 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v329/rabia-yapicioglu26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v329/rabia-yapicioglu26a.html</guid>
        
        
      </item>
    
      <item>
        <title>Conformal Prediction via Transported Beta Laws</title>
        <description>Split conformal prediction provides finite-sample marginal coverage under exchangeability, but this guarantee averages over the random calibration sample. We study instead the law of the calibration-conditional coverage induced by a realized conformal threshold. In the continuous i.i.d. setting this law is exactly Beta(k,n+1-k), so the usual marginal guarantee corresponds to its mean. We take this beta law as a finite-sample reference object and quantify departures from it using Wasserstein distances on [0,1]. The framework yields direct bounds on marginal coverage gaps and on bad-calibration probabilities, and separates different sources of non-i.i.d. behavior according to how they deform the beta reference: test-side shift acts through a transport map on the coverage scale, while calibration dependence changes the order-statistic law itself. We instantiate the framework in scale-shift, clustered, and stationary mixing settings, where the induced deformations can be characterized explicitly or through Berry–Esseen approximations. Simulations on dependent processes confirm that the first-order approximation tracks the empirical Wasserstein distance even at moderate sample sizes.</description>
        <pubDate>Sun, 30 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v329/r-ramos26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v329/r-ramos26a.html</guid>
        
        
      </item>
    
      <item>
        <title>Inductive Venn–Abers and related regressors</title>
        <description>Venn–Abers predictors are probabilistic predictors that enjoy appealing properties of validity, but their major limitation is that they are applicable only to the case of binary classification, with a recent extension to bounded regression. We generalize them to the case of unbounded regression, which requires adding an element of conformal prediction. In our simulation and empirical studies we investigate the predictive efficiency of point regressors derived from Venn–Abers regressors and argue that they somewhat improve the predictive efficiency of standard regressors for larger training sets.</description>
        <pubDate>Sun, 30 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v329/petej26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v329/petej26a.html</guid>
        
        
      </item>
    
      <item>
        <title>Continuous Assurance of Digital Assets via Conformal Prediction over Polytope-Defined Operating Regions</title>
        <description>Continuous assurance is a process for applying an agile, modular and iterative process to re-assure and verify changes in digital assets, such as simulation models, against a set of predefined requirements. In the recent years, several verification methods have been developed that can be used in the continuous assurance context to achieve statistically sufficient coverage of simulation space. Instead of exhaustively traversing the simulation search space to find a solution that falsifies the requirements, these methods partition the space into smaller and interpretable sub-regions. Statistical methods for computing prediction intervals for the subsequent sample provides a basis for determining whether simulation parameter sampled from these sub-regions are likely to satisfy or violate the requirement contract. In this work, we apply such techniques to test a converter controller model used in a grid-forming wind turbine, within the context of a continuous assurance cycle. Additionally, we extend the search space partitioning method from existing literature by creating regions of convex polytopes near class boundaries. These regions are linear in the optimization problem, with the aim to reduce the regions of uncertain areas during classification. We further demonstrate the partitioning method with an additional set of benchmarks including the mountain car and F-16 GCAS simulation model.</description>
        <pubDate>Sun, 30 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v329/park26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v329/park26a.html</guid>
        
        
      </item>
    
      <item>
        <title>Label-Free Distribution Shift Detection in Truck Mode Classification via Uncertainty-Gap Monitors</title>
        <description>Machine learning models in vehicle telematics face distribution shifts, but existing sequential monitors often require real-time ground-truth labels. We introduce the Uncertainty-Gap Monitor (UGM), a label-free empirical monitor based on the width of Venn-Abers probability intervals. In truck accelerometer data, UGM remained stable on benign shifts and produced large e-values that triggered alarms on the aggregated out-of-distribution fleet shifts.</description>
        <pubDate>Sun, 30 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v329/nurmatov26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v329/nurmatov26a.html</guid>
        
        
      </item>
    
      <item>
        <title>Exploring the Link Between Out-of-Distribution Detection and Conformal Prediction with Illustrations of Its Benefits</title>
        <description>Research on Out-Of-Distribution (OOD) detection focuses mainly on building scores that efficiently distinguish OOD data from In Distribution (ID) data. On the other hand, Conformal Prediction (CP) uses non-conformity scores to construct prediction sets with probabilistic coverage guarantees. In other words, the former designs scores, while the latter designs probabilistic guarantees based on scores. In this paper, we study how these two fields can benefit each other, and we formalize several aspects of this connection. First, for OOD detection, we show that in standard OOD benchmark settings, evaluation metrics can be affected by the validation dataset’s finite sample size. Extending the work of Bates et al. (2023), we define new conformal AUROC and conformal FPR@TPR95 metrics, which are corrections that provide probabilistic guarantees on the variability of the FPR involved in these metrics with respect to the validation datasets. We show the effect of these corrections on two reference OOD and anomaly detection benchmarks, OpenOOD (Yang et al., 2022), and ADBench (Han et al., 2022). Second, for CP, we study the use of OOD scores as non-conformity scores and show that they can improve the efficiency of the prediction sets obtained with CP in several settings.</description>
        <pubDate>Sun, 30 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v329/novello26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v329/novello26a.html</guid>
        
        
      </item>
    
      <item>
        <title>Distortion and Consistency in Conformal Prediction</title>
        <description>Conformal Prediction (CP) is a framework for reliable machine learning. Its aim is to convert a classification algorithm into a calibrated one with guaranteed validity properties. The core element (metaparameter) of a CP algorithm is a Non-Conformity Measure (CM) function that typically links the framework to an underlying algorithm. Typically, NCM represents an information distance between a data example from an object space and a bag of data examples. In most cases, it has a residual form: the difference between the true value of an example and the label predicted by the underlying algorithm. However, there is also an alternative principle for the NCM construction. Its core is a consistency function that is a function of a bag of data examples only and measures, in some sense, regularity in a bag. Consistency functions can be transformed into NCM functions in a standard way. Therefore, we suggest considering the consistency function as an intermediate chain between the underlying algorithm and the Conformity Measure. In the example of a simple nearest-neighbour approach, we demonstrate that it covers the capabilities of the standard method for defining conformity measures.</description>
        <pubDate>Sun, 30 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v329/nouretdinov26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v329/nouretdinov26a.html</guid>
        
        
      </item>
    
      <item>
        <title>Risk-Averse Evaluation of Conformal Scores through Conditional Value-at-Risk</title>
        <description>Conformal prediction provides distribution-free guarantees by constructing prediction sets that contain the true outcome with a predefined probability, under the mild assumption of exchangeability. While this marginal coverage guarantee is powerful, it remains an aggregate measure, by controlling the frequency of errors but being agnostic to their severity. In many high-stakes applications, however, not all errors are equally critical. Two conformal predictors achieving the same nominal coverage may exhibit different tail behaviors: one may incur into large deviations from the empirical quantile, while another may keep closer to it. However, standard evaluation metrics, such as empirical coverage and average set size, fail to capture this aspect. To address this limitation, we propose to evaluate conformal predictors through the lens of Conditional Value-at-Risk, a well-established measure of tail risk. This perspective enables a risk-averse selection of score functions, favoring those that control not only the frequency but also the magnitude of errors. In particular, we select the score function that minimizes this tail risk, thereby ensuring a tighter alignment between the worst errors observed on test data and the quantile estimated during calibration.</description>
        <pubDate>Sun, 30 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v329/narteni26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v329/narteni26a.html</guid>
        
        
      </item>
    
      <item>
        <title>Adaptive Conformal Prediction for 5G Link Adaptation under Distribution Shift</title>
        <description>In 5G link adaptation a modulation and coding scheme (MCS) is chosen under a 10% block-error-rate (BLER) constraint that channel non-stationarity undermines by breaking exchangeability. We drive the significance level $\alpha$t from a Conformal Test Martingale (CTM) used as a continuous control signal rather than a binary alarm: a windowed mixture martingale tracks non-exchangeability online and a bounded sigmoidal law maps its wealth to $\alpha$t. On real traces this restores BLER validity, matches error-counting baselines and—unlike them—returns to its nominal operating point once transient drift subsides.</description>
        <pubDate>Sun, 30 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v329/minotti26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v329/minotti26a.html</guid>
        
        
      </item>
    
      <item>
        <title>Towards Group-Conditional Coverage Guarantees via Group-Invariant Representations</title>
        <description>Split-conformal prediction is a solid framework that transforms any point predictor into a conformal set predictor with marginal coverage guarantees. In cases where observations are divided into different groups, these “average” guarantees may not be enough. Group-conditional coverage cannot be maintained, by default, using marginal calibration since different groups have different distributions of nonconformity scores. In this short paper, we propose a methodology for obtaining equalized group-conditional coverage via marginal calibration by learning group-invariant representations that disentangle the relationship between group-specific information and the nonconformity scores.</description>
        <pubDate>Sun, 30 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v329/melki26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v329/melki26a.html</guid>
        
        
      </item>
    
      <item>
        <title>Particle-Based Conformal Prediction for Contact-Aware Uncertainty Calibration in Stratified Configuration Spaces</title>
        <description>Reliable uncertainty representation is essential for deploying autonomous systems that interact with their environment, as robots must reason about how uncertainty arising from both stochasticity and model mismatch is impacted by contacts with obstacles (e.g., when navigating through a cluttered environment or inserting a part into an assembly). We propose Calibrated Particle-sets for Trans-dimensional Uncertainty Representation (CaPTURe), a geometry-aware, conformal prediction-based algorithm that generates probabilistically valid prediction regions of the unknown future system configuration using particle-based models of arbitrary fidelity. While calibrated uncertainty predictions are essential for safe and efficient planning, analytical or learned motion models are often inaccurate—due to limited data, simplifying assumptions, unmodeled effects, etc.—which can lead to unsafe executions or task failure. Additionally, when a robot contacts an obstacle, the distribution of its future configurations can become multimodal or disjoint, or lie along manifolds of lower intrinsic dimension than the space of possible robot configurations. Our method uses a calibration dataset of system transitions to locally calibrate motion uncertainty estimates, constructing regions guaranteed to contain the future robot configuration at a user-set probability. Our calibration procedure captures how motion uncertainty varies between contact-rich and contactless motions, leading to sufficient coverage in both cases. We evaluate our method on two simulated planning tasks: controlling a marble around a labyrinth and performing tight-tolerance peg-in-hole insertion with a manipulator. Compared to relevant baselines, CaPTURe achieves the user-specified coverage requirement both in and out of contact and achieves up to a 30% absolute improvement in task success rate over the best baseline. Project website: https://um-arm-lab.github.io/capture</description>
        <pubDate>Sun, 30 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v329/marques26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v329/marques26a.html</guid>
        
        
      </item>
    
      <item>
        <title>When Post-Hoc Calibration Hurts: Platt Scaling and Isotonic Regression on Modern Tabular Models</title>
        <description>Post-hoc calibration is routinely applied on the assumption that it cannot degrade a classifier. We test that over 3,150 runs per calibrator: 21 classifiers, 30 binary TabArena-v0.1 tasks, 5-fold cross-validation. Under log-loss, two of the most widely used methods are unreliable: isotonic regression has a mean relative change of +3.22% and improves log-loss in only 43.7% of runs, and Platt scaling improves it in 49.8%, showing no consistent directional advantage. Venn–Abers attains the largest mean reduction (-14.17%) and Beta calibration the highest improvement rate (67.1%). Mean AUC-ROC changes by under 1% throughout and mean accuracy is essentially unaffected except under Pearsonify, so changes in probabilistic quality are largely invisible in discrimination and classification metrics.</description>
        <pubDate>Sun, 30 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v329/manokhin26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v329/manokhin26a.html</guid>
        
        
      </item>
    
      <item>
        <title>Improving the Sensitivity of Gravitational Wave Detection with Weighted Conformal Prediction</title>
        <description>In the last decade, kilometre-scale interferometric gravitational-wave detectors have observed hundreds of compact binary mergers, the majority of which are binary black holes. However, the data are noise-dominated, and multiple independent search algorithms (pipelines) are used to enhance sensitivity and improve robustness. Rather than the standard approach of selecting the most significant pipeline output, we combine the outputs from all pipelines using a conformal prediction-based framework to provide statistically rigorous confidence estimates for candidate events. While combining pipelines improves sensitivity and ranking robustness, it requires a principled statistical framework that remains valid as data properties evolve across observing runs. A key challenge is distribution shifts between simulated datasets used for training and calibration and the real, unlabelled, observations used for testing, which can invalidate coverage guarantees and bias confidence estimates. In this work, we address this challenge by incorporating likelihood-ratio reweighting into our conformal prediction framework to account for covariate shift. Using mock datasets containing simulated signals, we demonstrate that weighted conformal prediction restores well-calibrated coverage under covariate shift and increases the confidence of events near the detection threshold, recovering true signals that would otherwise be missed.</description>
        <pubDate>Sun, 30 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v329/malz26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v329/malz26a.html</guid>
        
        
      </item>
    
      <item>
        <title>Uncertainty-aware Decision Support in Power Grid Demand Forecasting</title>
        <description>Effective management of modern power grids requires accurate demand forecasting and interpretable decision support for fluctuations in electricity consumption. This study extends the Calibrated Explanations method with online calibration support for uncertainty-aware probabilistic regression in offline, semi-online, and fully online settings. Since the application is a temporal forecasting problem, the proposed workflow should be understood as an empirically updated online decision-support procedure rather than as a new finite-sample conformal-validity theorem for arbitrary dependent time series. We apply Long Short-Term Memory (LSTM) networks and online ridge regression to an electricity-demand dataset to provide calibrated forecasts with quantified uncertainty. These forecasts provide threshold-risk information and support operator interpretation through Calibrated Explanations, including sensitivity analysis of key variables that affect energy demand. The proposed approach illustrates the practical use of online calibrated explanations in power-grid decision support.</description>
        <pubDate>Sun, 30 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v329/lofstrom26c.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v329/lofstrom26c.html</guid>
        
        
      </item>
    
      <item>
        <title>Conformal Repair on a Budget: An Empirical Protocol for Retraining Decisions</title>
        <description>Distribution shift can compromise the empirical reliability of conformal prediction intervals, but a coverage failure does not by itself imply that the predictive model must be refit. This paper studies budgeted conformal repair, an empirical maintenance rule for finite samples that selects the first sufficient action under a prespecified intervention-cost ordering, subject to empirical coverage and interval-width constraints. Given an existing linear predictor, conformal intervals, labelled shifted repair data, a target coverage level, and a width budget, the rule compares no intervention, residual-scale recalibration, conformal recalibration, intercept-only updating, and a linear refit on update data. In a controlled Gaussian linear-regression testbed, the resulting minimal repair map separates no-failure cases, failures repaired by low-cost recalibration or bias updating, failures for which a full linear refit is the minimal sufficient action, and failures that remain unrepairable within the width budget. Additional mixed-shift experiments show that compounding residual inflation with bias or slope shift can move otherwise repairable failures into the unrepairable-within-budget category. Many reliability failures are repairable without refitting the predictor. High noise-scale shifts are often width-limited rather than refit-limited, while large pure slope shifts provide the clearest case where refitting becomes necessary. The contribution is a concrete empirical decision rule for distinguishing lower-intervention repair, refitting on update data, and no acceptable repair under a stated width budget.</description>
        <pubDate>Sun, 30 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v329/lofstrom26b.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v329/lofstrom26b.html</guid>
        
        
      </item>
    
      <item>
        <title>Guarded Explanations: Conformal-Style Filtering for Distribution-Aware Rule Conditions</title>
        <description>In high-stakes decision support, post-hoc explanation methods are often used to estimate feature influence by perturbing inputs. However, a fundamental limitation of this approach is that candidate perturbations frequently fall outside the support of the calibration data, yielding unsupported and potentially misleading explanations. We introduce guarded explanations, an extension of the Calibrated Explanations framework. Guarded explanations apply a conformal-anomaly-detection-inspired filter to candidate perturbations, retaining a representative candidate only when its empirical guard score meets a user-chosen threshold $\varepsilon$. To facilitate this, we replace the standard binary discretiser with a multi-bin discretiser, generating interval-valued conditions alongside one-sided thresholds. Emitted perturbations are then scored using the standard Calibrated Explanations uncertainty-aware backend. In a synthetic constraint setting, we demonstrate that guarded factual explanations reduce the fraction of emitted rules whose representative values violate a known domain constraint from approximately 2.5% to near 0.0% at $\varepsilon$ = 0.2. Crucially, we isolate this effect from discretisation granularity alone by comparing against an unpruned multi-bin (no-guard) baseline. We frame guarded explanations as a representative-level plausibility filter, where the guard demonstrably reduces calibration-unsupported representative candidates in emitted rules. The guard score is a conformal-style empirical rank, not a standard conformal p-value, and $\varepsilon$ has no finite-sample error-rate guarantee for constructed perturbations.</description>
        <pubDate>Sun, 30 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v329/lofstrom26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v329/lofstrom26a.html</guid>
        
        
      </item>
    
      <item>
        <title>No Free Lunch in Few-Step Flow Matching for Conformal Prediction</title>
        <description>Conformal prediction (CP) turns a black-box predictor into prediction sets with finite-sample marginal coverage (Vovk et al., 2005). When p(y | x) is multimodal, a single interval either overcovers the gaps between modes or misses them; probabilistic conformal prediction (PCP; Wang et al., 2023) instead draws K samples from a conditional generative model and returns the union of balls of calibrated radius around them, tracing the support rather than summarizing it. Flow matching (FM; Lipman et al., 2023; Tong et al., 2024) is a natural generator, and recent work builds conformal procedures directly on flows (Li and Boström, 2025; Fang et al., 2025, 2026; Lee et al., 2025). All of this work assumes an exact sampler. FM sampling instead solves an ODE numerically, one evaluation of the learned velocity field per solver step, so its cost is the number of function evaluations (NFE), which the FM literature invests heavily in cutting (Liu et al., 2023; Song et al., 2023). Does that endanger the guarantee? No: for any fixed sampler, calibration and test scores stay exchangeable, so PCP retains coverage at any NFE, down to a single Euler step (Proposition 1, proved in Appendix B). However, there is no free lunch. (i) Coarse sampling is safe but not free. Set volume grows as the budget shrinks, by up to 1.51$\times$ at a single step, and across our four datasets the cost is ordered by $\hat{q}$/$\sigma$, the calibrated radius over the sampler’s conditional spread (Section 2), computable at full budget before any sweep. Where the ratio is large the K balls already coincide and the set is one ball around a point prediction, so few-step sampling is free exactly where sample-based CP was buying nothing over an interval method. (ii) Budget mismatch silently breaks coverage. Exchangeability requires the same sampler at calibration and deployment, an easily violated and previously unexamined condition: calibrating at 50 steps and deploying at 1 drops coverage from 90% to 48%, and to 11% on our hardest task, with no error signal. The rule: sample cheaply if you like, but calibrate with the sampler you deploy.</description>
        <pubDate>Sun, 30 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v329/li26b.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v329/li26b.html</guid>
        
        
      </item>
    
      <item>
        <title>A Deployment Audit of Release-Side Risk in Conformal Triage under Prevalence Shift</title>
        <description>Conformal triage converts predictive scores into deployment actions that either release a case, flag it for urgent attention, or defer it to human review. Under an observed change in target-event prevalence, however, marginal coverage and human-review rate can miss whether patients who experience the target event are released without review. To address this gap, we introduce a leakage-aware deployment audit for release-side conformal triage. It first assigns target subjects to three non-overlapping roles: prevalence correction, conformal calibration, and held-out release-side evaluation. This separation then lets the audit evaluate release directly: how many event-positive patients are cleared without review, whether the pilot has enough event labels for calibration, and how the release-review trade-off shifts. Applying this audit to a retrospective non-small-cell lung cancer (NSCLC) target cohort shows why lower review can be misleading: after prevalence correction, the pooled conformal branch lowers review by releasing more patients, some of whom are event-positive. Within the audit, the classwise branch acts as a scarcity diagnostic: the pilot has too few event labels to support a low-review release rule.</description>
        <pubDate>Sun, 30 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v329/li26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v329/li26a.html</guid>
        
        
      </item>
    
      <item>
        <title>Adaptive Conformal Prediction for Full-Scale Plant pH Forecasting under Process Drift</title>
        <description>Soft sensors in full-scale wastewater treatment face operating drift; overconfident point forecasts can mask compliance risk or force overly conservative chemical dosing. We evaluate conformal prediction intervals for 10-minute-ahead reactor pH forecasting on historical data from a semiconductor wastewater plant: a residual MLP trained on one source unit is evaluated offline across four anonymized units over multiple operating months. Shift diagnostics relative to a source reference month show strong covariate separability (domain-classifier AUC) and large increases in relative MAE (RMAE), indicating model-error-scale drift under real operation. Comparing split, weighted, and online adaptive (ACI) conformal methods and an EnbPI-style rolling residual interval at nominal 90% coverage, we find static and covariate-weighted intervals undercover shifted regimes while online methods recover near-nominal coverage at moderate width cost.</description>
        <pubDate>Sun, 30 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v329/lee26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v329/lee26a.html</guid>
        
        
      </item>
    
      <item>
        <title>Pixelwise Split Conformal Prediction for Global Temperature Emulation via Leave-One-Out Ensemble Distillation</title>
        <description>Neural emulators of climate model output rarely come with statistically principled uncertainty estimates. We study how split conformal prediction (SCP) can be repurposed from a post-hoc uncertainty wrapper into a training-time spatial prior for ensemble distillation. Given nine CMIP6 Earth System Models, we train nine leave-one-out (LOO) UNet++ teachers and calibrate each on its held-out target member, producing per-pixel calibration quantiles with the standard marginal split-CP validity interpretation at each spatial location and for each teacher, under exchangeability. From these calibration residual statistics, we derive calibration-informed center weights that score, for each pixel, how consistently each teacher tracks the bias-corrected ensemble consensus, and pixel-loss weights that re-weight the distillation objective so that the student is pushed harder on historically difficult regions. On global monthly near-surface air temperature emulation across a 192$\times$288 grid, the student achieves a test MAE of 0.9625 K versus 0.9782 K for a strong UNet++ baseline (-1.61%), decreases the error of 96.4% of grid cells, and reaches empirical pixelwise coverage of 90.6%. This student coverage is reported as an empirical calibration check rather than as a formal conformal guarantee for the distilled predictor. Our source code is available at https://github.com/bunyaminkorkut/COPA-Global-Temperature-Emulation.</description>
        <pubDate>Sun, 30 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v329/korkut26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v329/korkut26a.html</guid>
        
        
      </item>
    
      <item>
        <title>Conformal Score Prediction in Figure Skating Competitions</title>
        <description>Predicting figure-skating scores is a challenging modelling problem because the International Skating Union (ISU) scoring system combines objective, rule-based evaluation of technical elements with subjective human judgement of artistic and performance quality. This study examines whether the final score of a skater can be estimated using only technical-element data, combining reliability (validity) and information (precision) of the prediction intervals produced by the Conformal Prediction framework for regression. In general, the results indicate that machine learning models can effectively learn scoring patterns in figure skating.</description>
        <pubDate>Sun, 30 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v329/kisel26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v329/kisel26a.html</guid>
        
        
      </item>
    
      <item>
        <title>Target-Aligned Full Conformal Bayes under Continuous Label Shift</title>
        <description>Standard Full Conformal Bayes relies on exchangeability between the source observations and the target point. Under continuous label shift, weighting by the oracle label-density ratio transports rank aggregation to the target label distribution and restores target-domain coverage. However, the posterior predictive that defines the ranks may still reflect source-domain geometry, which can reduce interval efficiency. We propose Target-Aligned Full Conformal Bayes (TA-FCB), which keeps the weighted rank aggregation and additionally tilts each candidate-augmented posterior predictive before the scores are computed. For Bayesian ridge regression, candidate augmentation is a rank-one posterior update, and under an exponential label tilt the predictive tilt is a variance-scaled mean shift. Consequently, the procedure admits a closed-form implementation without candidate-wise refitting. The unrestricted oracle construction has finite-sample target-domain coverage under permutation-symmetric scoring, and the experiments use a shared finite-grid implementation. On two molecular property benchmarks, weighting recovers the coverage lost to controlled label shift, and candidate-wise tilting shortens intervals relative to weighting alone at comparable coverage. Estimated-ratio variants are evaluated separately and are not covered by the finite-sample guarantee.</description>
        <pubDate>Sun, 30 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v329/kim26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v329/kim26a.html</guid>
        
        
      </item>
    
      <item>
        <title>MuLaConf: Multi-label Conformal Prediction</title>
        <description>This paper presents MuLaConf, a Scikit-Learn-compatible Python package that implements Inductive (or Split) Conformal Prediction for multi-label classification tasks. Conformal Prediction is a rigorous machine learning framework for constructing prediction regions with coverage guarantees under the minimal assumption of data exchangeability. In multi-label classification, the Powerset Scoring approach explicitly computes p-values for all possible label combinations. However, it often yields un informatively large prediction sets. MuLaConf addresses this by defining nonconformity measures in the error vector space. This allows users to utilize either the Mahalanobis distance to capture label correlations or the standard Euclidean Norm. Furthermore, the package incorporates structural penalties, based on Hamming distance and label-set cardinality, to significantly reduce prediction set sizes while strictly maintaining theoretical coverage. To overcome the exponential computational complexity of the powerset space, MuLaConf leverages PyTorch-accelerated tensor operations and a memory management mechanism. Finally, its modular architecture features lazy evaluation, enabling the on-the-fly updating of distance measures and penalty weights without the need to retrain the underlying machine learning models. Together, these features make MuLaConf a practical tool for reliable uncertainty quantification in multi-label settings and a foundation for future research in Conformal Prediction.</description>
        <pubDate>Sun, 30 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v329/katsios26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v329/katsios26a.html</guid>
        
        
      </item>
    
      <item>
        <title>Generating Uncertainty Sets Using Conformal Prediction for Robust Optimization in Minimum Cost Flow</title>
        <description>Uncertainty presents a significant challenge in optimization, particularly in problems where unpredictable costs influence decision-making. Robust Optimization (RO) addresses this issue by modeling uncertainty through uncertainty sets, which ensure solutions hold under worst-case scenarios, with the model’s success depending directly on the accuracy of these sets. This research explores the application of conformal prediction to construct uncertainty sets for RO, an area not yet extensively investigated. The study specifically tests split and full conformal prediction within a robust optimization minimum cost flow problem. These methods are compared against traditional techniques, including interval-based and normal-based ellipsoidal uncertainty sets. Experiments were conducted across various network structures and under different error distributions, encompassing normal, skewed, and bimodal cases. The results indicate that conformal prediction uncertainty sets perform comparably to traditional methods, exhibiting slight variations in solution quality and coverage across different scenarios. While the conformal prediction uncertainty sets generally demonstrated tighter cost variability, their median performance did not consistently surpass that of traditional approaches. Notably, the full conformal models exhibited inconsistencies in coverage, a finding that warrants further investigation into the best method for generating an uncertainty set via full conformal prediction. The findings also highlight that network characteristics, particularly structural properties and inherent cost distributions, may influence the effectiveness of these uncertainty sets.</description>
        <pubDate>Sun, 30 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v329/johnson26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v329/johnson26a.html</guid>
        
        
      </item>
    
      <item>
        <title>Conformal Selection of Counterfactual Explanations: From Minimal to Robust Recourse</title>
        <description>Counterfactual explanations are commonly generated by identifying minimally perturbed instances that change a model prediction. However, minimal perturbations do not necessarily correspond to realistic, representative, or robust forms of recourse. In particular, counterfactuals located close to the decision boundary may satisfy the desired prediction outcome while still exhibiting atypical feature configurations or relying on unusual model reasoning patterns. This paper proposes a conformal framework for selecting counterfactual explanations based on their conformity to the target-class distribution. Rather than selecting the closest generated counterfactual, the proposed approach selects the candidate with highest conformal p-value according to conformal anomaly detection. Conformity is evaluated both in feature space and in SHAP space, enabling robustness to be assessed not only in terms of geometric similarity, but also in terms of explanatory consistency. Experiments on twelve benchmark datasets demonstrate that substantially more conforming counterfactual explanations can often be obtained with only modest increases in perturbation magnitude. In particular, the results show that counterfactuals with similar feature-space proximity may nevertheless differ substantially in SHAP-space conformity, suggesting that minimally perturbed counterfactuals may rely on atypical explanatory patterns. Qualitative examples further illustrate how conformity-based selection can produce explanations that appear more representative of genuine target-class instances. The proposed framework further highlights how conformity-based reasoning can extend beyond predictive uncertainty estimation and provide principled tools for assessing the quality and consistency of AI explanations.</description>
        <pubDate>Sun, 30 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v329/johansson26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v329/johansson26a.html</guid>
        
        
      </item>
    
      <item>
        <title>Lightweight Online Multivariate Conformal Prediction</title>
        <description>In the context of Advanced Driver Assistance Systems (ADAS), autonomous planning modules require rigorous uncertainty quantification alongside trajectory forecasting to ensure safe and proactive decision-making. While Conformal Prediction (CP) provides statistically valid confidence guarantees, its extension to multivariate spaces remains under-explored and, when implemented, often results in significant computational overhead that limits real-time applicability. To address this, we propose a Lightweight Online Multivariate Conformal Prediction (L-OMCP) framework designed to operate as an efficient top layer for existing prediction modules. Our approach constructs 2D confidence regions by introducing a joint non-conformity measure based on the empirical aspect ratio of spatial prediction errors, maintaining the lightweight computational overhead necessary for online deployment. To handle non-stationary driving dynamics, the framework utilizes an Exponential Moving Average (EMA) as its online update mechanism integrated with Adaptive Conformal Inference (ACI). This coupled formulation is projected into a bounded angular phase space to ensure robust stabilization and mitigate the impact of transient outliers or numerical instabilities. Comprehensive evaluations on synthetic scenarios and the real-world Renault dataset demonstrate that the proposed framework maintains the target coverage. Compared to both independent 1D baselines and state of the art multivariate methods, our approach generates tighter, context-aware bounding regions, providing a solution for real-time uncertainty quantification in ADAS.</description>
        <pubDate>Sun, 30 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v329/hourani26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v329/hourani26a.html</guid>
        
        
      </item>
    
      <item>
        <title>Online monotone density estimation and log-optimal calibration</title>
        <description>We study the problem of online monotone density estimation, where density estimators must be constructed in a predictable manner from sequentially observed data. We propose two online estimators: an online analogue of the classical Grenander estimator, and an expert aggregation estimator inspired by exponential weighting methods from the online learning literature. In the well-specified stochastic setting, where the underlying density is monotone, we show that the expected cumulative log-likelihood gap between the online estimators and the true density admits an O(n1/3) bound. We further establish a $\sqrt{}$n log n pathwise regret bound for the expert aggregation estimator relative to the best offline monotone estimator chosen in hindsight, under minimal regularity assumptions on the observed sequence. As an application of independent interest, we show that the problem of constructing log-optimal p-to-e calibrators for sequential hypothesis testing can be formulated as an online monotone density estimation problem. We adapt the proposed estimators to build empirically adaptive p-to-e calibrators and establish their optimality. Numerical experiments illustrate the theoretical results.</description>
        <pubDate>Sun, 30 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v329/hore26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v329/hore26a.html</guid>
        
        
      </item>
    
      <item>
        <title>Conformal Anomaly Detection in Python: Moving Beyond Heuristic Thresholds with nonconform</title>
        <description>Most anomaly detection systems output scores rather than calibrated decisions, leaving practitioners to choose thresholds heuristically and without clear statistical interpretation. Conformal anomaly detection addresses this limitation by converting anomaly scores into calibrated p-values that are valid under the statistical assumption of data exchangeability, with a growing literature extending this idea beyond that setting. We present nonconform, a Python package for applying conformal anomaly detection within existing machine-learning workflows, and use it as the basis for an implementation-grounded introduction to the field. The package integrates with scikit-learn, PyOD, and custom anomaly detectors, and provides a unified interface for calibration, p-value generation, and false discovery rate control. It supports several conformalization strategies, ranging from simple split-conformal calibration to more data-efficient and shift-aware extensions. Through a progression from foundational concepts to advanced conformalization strategies, complemented by code examples, the paper connects the statistical ideas behind conformal anomaly detection to their practical use in nonconform. Empirical results demonstrate that the implemented methods enable statistically principled anomaly detection. Together, the package and exposition aim to make core conformal anomaly detection workflows more accessible and reproducible in experimental and production-oriented settings.</description>
        <pubDate>Sun, 30 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v329/hennhofer26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v329/hennhofer26a.html</guid>
        
        
      </item>
    
      <item>
        <title>Beyond the Predicted Class: Calibrated Explanations for Real Multiclass Decisions*</title>
        <description>Calibrated Explanations provide uncertainty-aware feature importance and probabilistic predictions for binary classification, but existing multiclass extensions explain only the most probable class. This limitation reduces interpretability, especially when prediction confidence is low. We extend Calibrated Explanations to full multiclass support using a one-vs-rest strategy, generating calibrated probability intervals and class-specific feature importance for all classes. Experiments on four datasets show lower repeated-run variance than calibrated SHAP and LIME, while uncalibrated SHAP and LIME achieved lower variance in the robustness evaluation. A user study with 22 experts found that the proposed explanation format received the highest mean trust score and strongest overall preference ranking, while also influencing participants’ decisions.</description>
        <pubDate>Sun, 30 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v329/hanna26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v329/hanna26a.html</guid>
        
        
      </item>
    
      <item>
        <title>Venn Predictive Decision-Making</title>
        <description>Venn predictors generate multiprobability predictions, which are sets of probability distributions with provable calibration guarantees, quantifying epistemic ambiguity (Knightian uncertainty) rather than providing a single probability estimate. This paper introduces Venn predictive decision-making systems (Venn-PDMSs), a framework that integrates multiprobability predictions with decision-theoretic principles to guide decision-making under uncertainty. As Venn predictors explicitly account for epistemic ambiguity, the resulting Venn-PDMSs possess the capability to “know what they do not know,” thereby facilitating safe human-in-the-loop automation in numerous scenarios. We formalise utility- and regret-based decision criteria, parametrised by an optimism index that allows for interpolation between pessimistic and optimistic subjective stances towards uncertainty. We establish theoretical guarantees, including invariance under affine transformations of the utility function and convergence to Bayes-optimal decisions in certain cases. Through synthetic examples and straightforward applications to mushroom classification and financial credit assessment, we demonstrate that Venn-PDMSs effectively balance epistemic ambiguity and risk (which is probabilistically quantifiable). We also highlight the limitations of aggregating multiprobabilities into point estimates and emphasise the importance of explicitly incorporating ambiguity-aware decision criteria in high-stakes predictive decision-making.</description>
        <pubDate>Sun, 30 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v329/hallberg-szabadvary26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v329/hallberg-szabadvary26a.html</guid>
        
        
      </item>
    
      <item>
        <title>Martingale Detection of Treatment Effects in Wound Care</title>
        <description>In clinical research, the ability to act on accumulating evidence motivates methods that support valid sequential decision-making. We apply a martingale-based framework for detecting treatment effects in a one-arm clinical trial in wound care, where outcomes are continuously compared against a reference distribution. We then extend this framework to individualized treatment evaluation by constructing patient-specific counterfactual trajectories from a real-world reference cohort using Gaussian kernel weighting. Because the underlying null distribution is not fully known, treatment effects are quantified through a martingale-inspired divergence measure, and statistical significance is assessed using an empirical null distribution generated from 5,000 resampled reference cohorts.</description>
        <pubDate>Sun, 30 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v329/gustafsson26c.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v329/gustafsson26c.html</guid>
        
        
      </item>
    
      <item>
        <title>Trustworthy AI in Wound Care Documentation</title>
        <description>Efficient and trustworthy support for clinical documentation is becoming increasingly important as reporting requirements expand and EHR systems continue to increase in complexity. In wound care, where highly detailed structured assessments are a routine part of practice, there is a particular need for methods that assist clinicians through reliable recommendations. AI-based tools offer promising opportunities to reduce the documentation burden, but their value ultimately depends on the trustworthiness of the guidance they provide, as their recommendations feed directly into patient care and outcomes. In this work, we build on a unified uncertainty-quantification framework that combines Conformal Prediction with Venn–Abers predictors and extend it by integrating Shapley-value explanations into the same workflow. The resulting system produces prediction sets that narrow the label space for structured wound-care documentation while providing statistical coverage guarantees, calibrated inclusion probabilities, and interpretable explanations for the associated confidence. In practice, this kind of system could reduce documentation time and give clinicians a straightforward view of both the model’s uncertainty and the reasons behind its recommendations.</description>
        <pubDate>Sun, 30 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v329/gustafsson26b.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v329/gustafsson26b.html</guid>
        
        
      </item>
    
      <item>
        <title>Well-calibrated probabilities for prediction sets</title>
        <description>Conformal prediction is a model-agnostic, distribution-free statistical framework that outputs prediction sets such that, for each prediction, the probability that the true label falls outside the set is less than a user-specified tolerance, $\varepsilon$. Notably, this guarantee holds conditionally on the calibration dataset and thus only prior to observing a new object. We present a method that assigns prediction probabilities to any set prediction. These probabilities approximate the true conditional probability that the label lies within the set, and converge to it with increasing data. Venn-Abers predictors are (multi)-probabilistic predictors combining Venn-predictors and isotonic regression to output well-calibrated probabilities for binary prediction problems. We propose applying Venn-Abers to the constructed prediction sets from conformal prediction and, thereby, assigning a well-calibrated probability of the true label being in that set. We further introduce a novel self-consistent scoring rule for Venn-Abers predictors, which treats a predicted probability as being as likely as the event it postulates.</description>
        <pubDate>Sun, 30 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v329/gustafsson26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v329/gustafsson26a.html</guid>
        
        
      </item>
    
      <item>
        <title>Performance Estimation in Hybrid Interpretable Models with Venn Predictors</title>
        <description>The design of hybrid interpretable models has recently emerged as a promising paradigm for the development of explainable artificial intelligence. This modeling framework relies on a collaborative scheme in which black-box and interpretable models are paired and cooperate to produce a final prediction. Intuitively, it assumes that there are some regions of the feature space where a black-box classifier can be replaced by an interpretable model without losing predictive performance. For hybrid interpretable models to be useful in practice, users should have statistical guarantees on their performance for a given transparency level (i.e., the fraction of samples delegated to the interpretable classifier). In this work, we propose the use of Venn prediction to construct reliable accuracy estimators for hybrid interpretable models in the absence of ground truth labels. In particular, we derive estimators for the marginal accuracy (i.e., for the hybrid model as a single predictor) and the component-conditional accuracy (i.e., for each of the internal components of the hybrid model). We demonstrate the benefits of our estimation methodology, consistently outperforming alternative approaches for six different datasets, with the conditional Venn predictor providing the most reliable component-conditional estimates.</description>
        <pubDate>Sun, 30 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v329/garcia-galindo26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v329/garcia-galindo26a.html</guid>
        
        
      </item>
    
      <item>
        <title>Multivariate Change-Point Detection Using Feature-Based Inductive Conformal Martingales</title>
        <description>Change-point detection in sequential data can be formulated as the problem of identifying violations of exchangeability as observations arrive over time. In this work, we study feature-based change-point detection within the framework of Inductive Conformal Martingales, which are well known for providing theoretical validity guarantees on the probability of false alarms. We construct nonconformity measures that quantify how unusual the features of a test instance are. In particular, our proposed nonconformity measure is based on an ellipsoidal k-nearest neighbour density estimator combined with Mahalanobis distance. We evaluate the proposed method on four real-world datasets in which an artificial change-point is inserted, as well as on an air-quality dataset where the goal is to estimate possible sensor recalibration points. We also compare our approach with two competing nonparametric multivariate change-detection methods with false-alarm control mechanisms. The results illustrate the potential of feature-based conformal martingales for practical sequential change-point detection in multivariate data.</description>
        <pubDate>Sun, 30 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v329/eliades26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v329/eliades26a.html</guid>
        
        
      </item>
    
      <item>
        <title>When Is Conformal Coverage Free? Switching Thresholds for Predict-then-Optimize</title>
        <description>Machine learning increasingly drives operational decisions (dispatching power plants, routing vehicles, allocating medical supplies), yet its forecasts are uncertain. Conformal prediction turns a forecast into a calibrated uncertainty set with distribution-free guarantees, and a robust optimizer can then hedge its decision against that set. Before adopting this, a practitioner wants to know: will maintaining the uncertainty set actually change the decision the system makes, or will it leave the decision untouched and merely add an audit trail? We answer with a single quantity, the switching threshold: the point at which calibrated uncertainty begins to alter the chosen action. While the uncertainty stays below this threshold, coverage is free, and the system gains calibrated, auditable uncertainty without changing what it does. We show how to read this threshold from the structure of the decision problem, characterize the fluctuations of the online quantile that decide which side of it a problem falls on, and reduce these to a single pre-deployment safety margin that labels a problem as free, borderline, or costly. The label is set by the decision problem’s structure rather than the choice of predictor. Across real benchmarks in energy, routing, public health, and logistics, dispatching power under real market prices is free (coverage never changes which generators run and adds no cost), whereas routing on a dense city map changes the route about half the time. The certificate stays reliable at city scale, on road networks with tens of thousands of streets, and even when the predictor is a vision model reading images. The result is a simple test for whether adding conformal coverage will cost a decision system anything.</description>
        <pubDate>Sun, 30 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v329/dronavajjala26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v329/dronavajjala26a.html</guid>
        
        
      </item>
    
      <item>
        <title>Detecting Adversarial Attacks in Image Sequences with Conformal Test Martingales</title>
        <description>Deep neural networks used for image classification remain vulnerable to adversarial examples (AEs), especially under white-box attacks where the attacker can exploit information about the target model. Detecting whether a model is under attack is therefore an important step toward mitigating adversarial manipulation. Most existing AE detectors focus on instance-level or batch-level decisions, i.e., determining whether a single image or a fixed batch contains AEs. In this paper, we study a sequence-level detection setting using conformal test martingales. We formulate several nonconformity scores based on image and target-model quantities, convert them into conformal p-values, and monitor the resulting sequences with conformal test martingale algorithms. The resulting detection procedure is evaluated on ImageNet-100 image sequences with a pretrained ResNet-50 target model and three representative white-box attacks. The experiments show that conformal test martingales can detect attacks even when the monitoring segment contains a relatively small number of adversarial examples, but that detection performance depends strongly on both the nonconformity score and the martingale algorithm. Empirically, no single nonconformity score and algorithm pair works well for all attacks. These findings suggest that applying conformal test martingales to detect attacks may benefit from considering multiple image and target-model quantities and designing nonconformity scores according to the observed attack-induced changes.</description>
        <pubDate>Sun, 30 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v329/dong26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v329/dong26a.html</guid>
        
        
      </item>
    
      <item>
        <title>Robust Model Predictive Control via Conformal Prediction</title>
        <description>Model predictive control (Rawlings et al., 2017) is an established and widely used method in control theory. More specifically, it is a feedback control strategy that uses a dynamic model of the system under control to predict its future behavior over a finite time horizon. At each time step, it solves an optimization problem to determine the best control action, based on the predicted system behavior. In real applications, the prediction of the future system behavior is commonly afflicted with uncertainty. In this work, we therefore consider the use of conformal prediction (Vovk et al., 2005) to increase safety and make control more robust. Broadly speaking, the idea is to conformalize the predicted system behavior, replacing precise trajectories by confidence bands that cover the true behavior with high probability. Safe control inputs can then be selected in a more “cautious” way by optimizing worst-case scenarios.</description>
        <pubDate>Sun, 30 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v329/czaja26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v329/czaja26a.html</guid>
        
        
      </item>
    
      <item>
        <title>ConfBatch: Conformal Prediction for Adaptive Batch Sizing</title>
        <description>We study whether conformal uncertainty can guide adaptive batch size selection during neural network training. We propose ConfBatch that periodically computes conformal prediction sets on held-out data, using their size as an uncertainty score to adjust batch size. On CIFAR-10, the opposite-direction ConfBatch variants achieve the highest accuracies, while standard ConfBatch has the lowest wall time among adaptive batch-size methods.</description>
        <pubDate>Sun, 30 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v329/chatzipapadopoulou26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v329/chatzipapadopoulou26a.html</guid>
        
        
      </item>
    
      <item>
        <title>A Gibbs-Boltzmann Approach to Conformal Prediction under Distribution Shift</title>
        <description>Building on the foundation of standard Conformal Prediction and its known vulnerabilities to covariate shift, we formalize our Thermodynamic DRO framework. While recent methods such as Aolaritei et al. (2025) address score shift via topological metrics like the Lévy-Prokhorov distance, our approach introduces a statistical mechanics perspective. By applying a physics-inspired thermodynamic base measure directly to 1D non-conformity scores within a KL-divergence ambiguity set, we yield a smooth, closed-form partition function that expands conformal thresholds to absorb out-of-distribution errors with minimal computational overhead.</description>
        <pubDate>Sun, 30 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v329/channapura-sudhakara26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v329/channapura-sudhakara26a.html</guid>
        
        
      </item>
    
      <item>
        <title>Fair Conformal Prediction for Individuals and Subgroups in Natural and Adversarial Settings: Theoretical Guarantees and Impossibility Results</title>
        <description>Conformal prediction offers a principled, model-agnostic framework for constructing prediction sets with formal coverage guarantees under the assumption of exchangeability. In many practical applications, however, coverage alone is insufficient, and additional properties of the resulting prediction sets are often required. For instance, recent work has explored how to ensure CP fair treatment of both individuals and population subgroups and robustness to adversarial perturbations. In this paper, we study whether it is possible to construct conformal prediction sets that preserve coverage guarantees while simultaneously ensuring fair treatment of individuals, in the sense of counterfactual fairness by means of Equal Set Size, and of population subgroups, in the sense of Equalized Coverage and Equalized Average Set Size, in both natural and adversarial settings. In the adversarial setting, we consider an attacker whose goal is to induce unfair behavior toward either individual subjects or population subgroups, and we develop the first attack and defense methods for this scenario. We analyze both the case in which the sensitive attribute is available at test time and the more realistic setting in which it is not. For the natural setting, we propose conformal prediction methods that are provably efficient and satisfy fairness guarantees, and we also establish corresponding impossibility results. For the adversarial setting, we show that, under a realistic threat model in which the adversarial strategy is explicitly specified, the guarantees achieved in the natural setting can be recovered. Finally, experiments on real-world classification datasets involving fairness-sensitive tasks demonstrate the effectiveness and practical relevance of our approach, while also highlighting the limitations of existing methods.</description>
        <pubDate>Sun, 30 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v329/carlevaro26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v329/carlevaro26a.html</guid>
        
        
      </item>
    
      <item>
        <title>Explaining Conformal Prediction: Diagnosing Reliability through Feature-Conditioned Outcomes and p-Value Margins</title>
        <description>Conformal prediction produces prediction sets with finite-sample coverage guarantees, and its practical utility for decision-making grows when the reliability of its predictions can be explained at the feature level. Yet, the intersection of conformal prediction and explainability remains relatively underexplored, leaving practitioners without tools to interpret how input features shape prediction reliability. We introduce a framework for explaining the reliability of conformal classifiers at the feature level using two tools. First, Feature-Conditioned Outcome Distribution (FCOD) plots visualize how correct, incorrect, and ambiguous conformal prediction set outcomes vary across feature ranges, identifying regions of the feature space associated with different empirical behaviors in terms of correctness and uncertainty of the predictions. Second, we propose two metrics derived from class-conditional conformal p-values: the label-free evidence margin, which provides a signed contrast between competing classes without the true label, and the label-dependent correctness margin, which compares the conformal p-value of the true class with that of its strongest alternative. They are used as SHAP targets for local feature attribution and aggregated global analysis, enabling feature-level explanations of relative conformal evidence between competing classes and support for the true class.</description>
        <pubDate>Sun, 30 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v329/caparrini26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v329/caparrini26a.html</guid>
        
        
      </item>
    
      <item>
        <title>Skew-adaptive conformal prediction</title>
        <description>To account for predictive uncertainty in regression that may vary across the feature space not only in scale but also in asymmetry, we develop a skew-adaptive extension of split conformal prediction. The construction starts from an asymmetric interval family centered at a point prediction and uses the gauge approach to deduce the conformity score induced by this family. The inverse hyperbolic sine transform of signed scaled residuals provides the training target for an additional predictive model, whose role is to learn how predictive uncertainty should tilt across the feature space. The resulting procedure preserves the finite-sample marginal validity of split conformal prediction under exchangeability, while producing intervals that adapt to both local scale and local skewness. We also develop a calibration-sample-based estimator for comparing the expected relative future width of the skew-adaptive and classical scaled-score intervals. Experiments on a variety of datasets indicate gains in prediction interval efficiency over the scaled-score construction and con-formalized quantile regression, and also show that the proposed estimator closely matches the corresponding average width ratio observed on the test sample.</description>
        <pubDate>Sun, 30 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v329/c-marques-f-26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v329/c-marques-f-26a.html</guid>
        
        
      </item>
    
      <item>
        <title>Efficient Multi-Label Conformal Classification</title>
        <description>In multi-label classification, the label space consists of the power set of a given set of class labels. Generating conformal classifiers for such tasks may be challenging, as the label space grows exponentially with the number of class labels, potentially resulting in a very high computational cost. Moreover, prediction sets containing label sets may be difficult to interpret; for example, the inclusion of a superset does not entail the inclusion of any of its subsets. A novel approach to multi-label conformal classification is proposed that addresses these problems by assuming an object-specific scoring function over the class labels, defining an ordering according to which labels are included in the prediction set. Results from an empirical investigation on six multi-label classification datasets, using both multiple single-target models and single multiple-target models, confirm that the theoretically guaranteed coverage is achieved and show that the choice of scoring function may have a substantial impact on predictive efficiency.</description>
        <pubDate>Sun, 30 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v329/bostrom26c.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v329/bostrom26c.html</guid>
        
        
      </item>
    
      <item>
        <title>Detecting Label Shift for Large Language Models with Conformal Test Martingales</title>
        <description>We study the problem of detecting label shift in deployed Large Language Models (LLMs), using model outputs, in the form of class labels, confidence estimates, and justifications, as proxies for true labels. We propose the use of conformal test martingales, which can detect shifts in both mean and dispersion across multiple streams while controlling the false alarm rate. We evaluate the approach on three binary and three multiclass text classification datasets using LLMs ranging from 3B to 32B parameters. We consider three martingale algorithms: Simple Jumper, Sleeper/Stayer, and Sleeper/Drifter, as well as their composite. The observed false alarm rates are consistently below the nominal level. Detection-time rankings are stable across datasets, models, and signals, with Sleeper/Drifter and the composite martingale outperforming the others, while Simple Jumper is least competitive. The composite provides only marginal gains over Sleeper/Drifter alone. Larger models tend to enable earlier detection of label shift, although the effect is uneven and non-monotonic. Among the signals, confidence is the most reliable proxy for binary datasets, whereas combining all three signals yields the best performance on two multiclass datasets. Justification is highly dataset-dependent: it fails for several datasets and models, yet is the only effective signal for the smallest model in one case.</description>
        <pubDate>Sun, 30 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v329/bostrom26b.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v329/bostrom26b.html</guid>
        
        
      </item>
    
      <item>
        <title>Mondrian Conformal Regressors with Overlapping Categories</title>
        <description>Normalized conformal regressors have two potential limitations: i) they may generate prediction intervals that are several times larger than the largest observed absolute residual in the calibration set, and ii) the interval sizes may exhibit high dispersion even when using non-informative difficulty estimators. Mondrian conformal regressors provide a remedy by forming categories through equal-sized binning of the difficulty estimates and applying standard conformal regressors within each category. A drawback of this approach is that the mapping from difficulty estimates to prediction intervals may be very coarse, in particular for smaller calibration sets. Moreover, the calibration set size of each category cannot be fully controlled, potentially leading to overly conservative prediction intervals. We introduce a generalization of conformal predictors that allows overlapping Mondrian categories, and present an instantiation of this class for conformal regression. The proposed approach results in a smoother distribution of interval sizes while ensuring that each Mondrian category contains exactly the desired number of examples. A large-scale experimental evaluation on 33 datasets shows that the proposed approach consistently achieves superior average rankings with respect to prediction interval size compared with standard, normalized, and non-overlapping Mondrian conformal regressors over a range of difficulty estimators and significance levels.</description>
        <pubDate>Sun, 30 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v329/bostrom26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v329/bostrom26a.html</guid>
        
        
      </item>
    
      <item>
        <title>Geometry-Aware Conformal Decoding for Motor BCI Systems</title>
        <description>Brain-computer interface (BCI) motor decoders need accurate kinematic commands and reliable uncertainty. In closed loop, velocity uncertainty can gate or scale commands or give users feedback. Prior non-invasive BCI work used conformal prediction to defer uncertain discrete exoskeleton commands (Eliades and Papadopoulos, 2019); invasive work remains limited to one fixed-width global interval around 1D reach direction or position (Wei et al., 2024). We instead compare complete decoder-geometry systems, not geometries around a common predictor, for continuous 2D velocity decoding.</description>
        <pubDate>Sun, 30 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v329/bonini26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v329/bonini26a.html</guid>
        
        
      </item>
    
      <item>
        <title>Conformal prediction as a solution to distribution shifts in medical image analysis</title>
        <description>Before medical AI systems can be widely deployed in clinics, there is a need for uncertainty quantification and improved handling of distribution shifts. Conformal prediction (CP) is an uncertainty quantification method that is highly useful but requires datasets to exhibit exchangeability, something that is broken by distribution shifts. In this work, we employ CP on real-world pathology image data to investigate distribution shifts’ impact on the validity of CP. We find that the main driver for loss of exchangeability is disagreements between pathologists, rather than medical differences between patients.</description>
        <pubDate>Sun, 30 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v329/boman26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v329/boman26a.html</guid>
        
        
      </item>
    
      <item>
        <title>Conformal Calibration for Multi-Modal Regression with Missing Modalities</title>
        <description>Prediction intervals for multi-modal regression with tabular variables, text, images, or other input sources are difficult to calibrate when those sources disagree or one is missing. A single global quantile averages these regimes together instead of calibrating to the modality pattern observed at test time. We address this through a modality-aware conformal calibration layer. The layer trains or reuses one predictor per modality, computes a disagreement score from their predictions, and uses that score in split conformal calibration under a strict split protocol. We use the score in two complementary ways. First, a continuous disagreement-scaled method reallocates interval width across examples while preserving the usual marginal split-conformal guarantee. Second, a Mondrian (stratified) method calibrates within groups defined by disagreement or modality availability fixed before calibration, giving group guarantees under joint exchangeability of the calibration and test examples. Across four multi-modal datasets, the disagreement-scaled layer matches or improves the marginal conformal baseline in 59 of 60 paired runs for interval continuous ranked probability score (CRPS) and in 52 of 60 for interval width, while keeping empirical coverage near the 95% target. In stress tests with missing modalities, mask-matched recalibration recovers up to 19.5 percentage points of coverage in the hardest fixed-mask regime. The result is a simple, model-agnostic reliability layer for multi-modal regression systems. A project page is available at https://unco3892.github.io/modality-aware-conformal.</description>
        <pubDate>Sun, 30 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v329/azizi26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v329/azizi26a.html</guid>
        
        
      </item>
    
      <item>
        <title>Attune, Don’t Prune: A Conformal Framework for Fact Preservation in News Content Attunement</title>
        <description>We propose a news content attunement system that rewrites articles according to a reader-defined graphic sensitivity setting, or abstains if rewriting is not possible. Conformal importance selection provides a finite-sample article-level recall guarantee for important source sentences, while conformal risk control calibrates a bounded aggregate attunement loss balancing preservation, residual graphic content, and utility. On a held-out dataset, we preserve 86% of essential facts whilst neutralising $\approx$97% of graphic content.</description>
        <pubDate>Sun, 30 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v329/ashby26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v329/ashby26a.html</guid>
        
        
      </item>
    
      <item>
        <title>Turning Feature Attributions into Sufficient Explanations Using Conformal Prediction</title>
        <description>Feature attribution methods are widely used for explaining image-based predictions, as they provide feature-level insights that can be intuitively visualized. However, such explanations often vary in their robustness and may fail to reflect the reasoning of the underlying black-box model faithfully. To address these limitations, we propose a novel conformal prediction-based approach that enables users to assign confidence levels to the important explanation regions. The proposed approach employs any feature-attribution method to identify a subset of salient features sufficient to preserve the model’s prediction, regardless of the information carried by the excluded features, without demanding access to ground-truth explanations for calibration. Four conformity functions are proposed to quantify the extent to which explanations conform to the model’s predictions. The approach is empirically evaluated using five explainers across six image datasets. The empirical results demonstrate that FastSHAP consistently outperforms the competing methods in terms of both fidelity and informational efficiency, the latter measured by the size of the explanation regions. Furthermore, the results reveal that conformity measures based on super-pixels are more effective than their pixel-wise counterparts.</description>
        <pubDate>Sun, 30 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v329/alkhatib26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v329/alkhatib26a.html</guid>
        
        
      </item>
    
      <item>
        <title>Reliable Environmental Planning at HS2 Construction Sites</title>
        <description>High Speed 2 (HS2) is a major UK railway infrastructure project with substantial economic and transport benefits. However, its scale and operational complexity have potentially created environmental impacts. Thus, forecasting air quality and noise levels may support more informed site management. Nevertheless, existing Machine Learning models do not reveal the uncertainty of their predictions, limiting their reliability in such high-stakes environments. Therefore, we propose a state-conditioned conformal prediction (CP) framework that uses recent operational-state information to produce reliable and adaptive prediction intervals. Experiments on real-world HS2 air-quality and noise monitors show that state-conditioned calibration reduced mean interval width by up to 40.38%. </description>
        <pubDate>Sun, 30 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v329/ahluwalia26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v329/ahluwalia26a.html</guid>
        
        
      </item>
    
      <item>
        <title>Preface</title>
        <description>This preface introduces the proceedings of the Fifteenth Symposium on Conformal and Probabilistic Prediction with Applications (COPA 2026), highlighting the key themes and contributions presented at the conference.</description>
        <pubDate>Sun, 30 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v329/ahlberg26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v329/ahlberg26a.html</guid>
        
        
      </item>
    
  </channel>
</rss>
