ALIGN: Adversarial Learning for Generalizable Speech Neuroprosthesis

Zhanqi Zhang, Shun Li, Bernardo L. Sabatini, Mikio Christian Aoi, Gal Mishne
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:8006-8025, 2026.

Abstract

Intracortical brain-computer interfaces (BCIs) can decode speech from neural activity with high accuracy when trained on data pooled across recording sessions. In realistic deployment, however, models must generalize to new sessions without labeled data, and performance often degrades due to cross-session nonstationarities (e.g., electrode shifts, neural turnover, and changes in user strategy). In this paper, we propose {ALIGN}, a session-invariant learning framework based on multi-domain adversarial neural networks for semi-supervised cross-session adaptation. {ALIGN} trains a feature encoder jointly with a phoneme classifier and a domain classifier operating on the latent representation. Through adversarial optimization, the encoder is encouraged to preserve task-relevant information while suppressing session-specific cues. We evaluate {ALIGN} on intracortical speech decoding and find that it generalizes consistently better to previously unseen sessions, improving both phoneme error rate and word error rate relative to baselines. These results indicate that adversarial domain alignment is an effective approach for mitigating session-level distribution shift and enabling robust longitudinal BCI decoding.

Cite this Paper


BibTeX
@InProceedings{pmlr-v337-zhang26a, title = {{ALIGN}: Adversarial Learning for Generalizable Speech Neuroprosthesis}, author = {Zhang, Zhanqi and Li, Shun and Sabatini, Bernardo L. and Aoi, Mikio Christian and Mishne, Gal}, booktitle = {Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence}, pages = {8006--8025}, year = {2026}, editor = {Perković, Emilija and Malinsky, Daniel}, volume = {337}, series = {Proceedings of Machine Learning Research}, month = {17--21 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v337/main/assets/zhang26a/zhang26a.pdf}, url = {https://proceedings.mlr.press/v337/zhang26a.html}, abstract = {Intracortical brain-computer interfaces (BCIs) can decode speech from neural activity with high accuracy when trained on data pooled across recording sessions. In realistic deployment, however, models must generalize to new sessions without labeled data, and performance often degrades due to cross-session nonstationarities (e.g., electrode shifts, neural turnover, and changes in user strategy). In this paper, we propose {ALIGN}, a session-invariant learning framework based on multi-domain adversarial neural networks for semi-supervised cross-session adaptation. {ALIGN} trains a feature encoder jointly with a phoneme classifier and a domain classifier operating on the latent representation. Through adversarial optimization, the encoder is encouraged to preserve task-relevant information while suppressing session-specific cues. We evaluate {ALIGN} on intracortical speech decoding and find that it generalizes consistently better to previously unseen sessions, improving both phoneme error rate and word error rate relative to baselines. These results indicate that adversarial domain alignment is an effective approach for mitigating session-level distribution shift and enabling robust longitudinal BCI decoding.} }
Endnote
%0 Conference Paper %T ALIGN: Adversarial Learning for Generalizable Speech Neuroprosthesis %A Zhanqi Zhang %A Shun Li %A Bernardo L. Sabatini %A Mikio Christian Aoi %A Gal Mishne %B Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2026 %E Emilija Perković %E Daniel Malinsky %F pmlr-v337-zhang26a %I PMLR %P 8006--8025 %U https://proceedings.mlr.press/v337/zhang26a.html %V 337 %X Intracortical brain-computer interfaces (BCIs) can decode speech from neural activity with high accuracy when trained on data pooled across recording sessions. In realistic deployment, however, models must generalize to new sessions without labeled data, and performance often degrades due to cross-session nonstationarities (e.g., electrode shifts, neural turnover, and changes in user strategy). In this paper, we propose {ALIGN}, a session-invariant learning framework based on multi-domain adversarial neural networks for semi-supervised cross-session adaptation. {ALIGN} trains a feature encoder jointly with a phoneme classifier and a domain classifier operating on the latent representation. Through adversarial optimization, the encoder is encouraged to preserve task-relevant information while suppressing session-specific cues. We evaluate {ALIGN} on intracortical speech decoding and find that it generalizes consistently better to previously unseen sessions, improving both phoneme error rate and word error rate relative to baselines. These results indicate that adversarial domain alignment is an effective approach for mitigating session-level distribution shift and enabling robust longitudinal BCI decoding.
APA
Zhang, Z., Li, S., Sabatini, B.L., Aoi, M.C. & Mishne, G.. (2026). ALIGN: Adversarial Learning for Generalizable Speech Neuroprosthesis. Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research 337:8006-8025 Available from https://proceedings.mlr.press/v337/zhang26a.html.

Related Material