[edit]
PRISM: Calibrated Bayesian Fusion and Auditable Attribution for Reliable LLM Event Prediction from Text
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:2595-2617, 2026.
Abstract
We study event prediction from news text when intermediate drivers are not expert-annotated. Directly prompting large language models ({LLMs}) for forecast probabilities can yield miscalibrated outputs and explanations that are hard to audit. We propose the Probabilistic Reliability and Interpretability System, PRISM, which separates semantic extraction from probabilistic decision making. PRISM uses {LLMs} only to extract interpretable factors, treats repeated extractions as noisy measurements of latent states, and performs {Bayesian} inference to produce posterior predictive probabilities that propagate extraction and measurement uncertainty. For empirical reliability under temporal dependence, we add low-capacity post-hoc calibration and time-adaptive conformal prediction and report calibration and coverage diagnostics. PRISM also reports factor-contrast association summaries with uncertainty and explicit interpretation boundaries. On a {UCDP}–{GDELT} conflict benchmark, PRISM achieves AUROC 0.821 and {ECE} 0.058 on the test set; at 90% target coverage, it attains empirical coverage 0.908 with average set size 1.38.