<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>Proceedings of Machine Learning Research</title>
    <description>Proceedings of the 2nd Conference on Topology, Algebra, and Geometry in Data Science(TAG-DS 2026)
  Held in Boston, Massachusetts, USA on 18-20 August 2026

Published as Volume 334 by the Proceedings of Machine Learning Research on 16 August 2026.

Volume Edited by:
  Eddie Berman
  Guillermo Bernárdez
  Samantha Chen
  Alex Cloninger
  Timothy Doster
  Tegan Emerson
  J. Elisenda Grigsby
  Henry Kvinge
  Hannah Lawrence
  Tim Marrinan
  Audun Myers
  Mathilde Papillon
  Behrooz Tahmasebi
  Lev Telyatnikov
  Robin Walters
  Melanie Weber
  YuQing Xie
  Eric Yeats

Series Editors:
  Neil D. Lawrence
  Hoel Kervadec
  Tegan Emerson
</description>
    <link>https://proceedings.mlr.press/v334/</link>
    <atom:link href="https://proceedings.mlr.press/v334/feed.xml" rel="self" type="application/rss+xml"/>
    <pubDate>Sun, 16 Aug 2026 14:52:37 +0000</pubDate>
    <lastBuildDate>Sun, 16 Aug 2026 14:52:37 +0000</lastBuildDate>
    <generator>Jekyll v3.10.0</generator>
    
      <item>
        <title>Dimensionless Controls of Plasticity Under Alternating Tasks: From Evolutionary Biology to Continual Learning</title>
        <description>Plasticity under changing environments is central to both evolutionary biology and continual learning. Motivated by recent work on genotype–phenotype maps, we study a minimal deep-learning analogue where a network is trained alternately on two Boolean label sets, and ask which biological controls of plasticity survive the translation to gradient descent. Reinterpreting four proposed biological factors as quantities of training dynamics, we find the system reduces to two dimensionless controls: the task disagreement $r$, the fraction of disagreeing labels, and the reach $\eta T$, the product of learning rate and switching period. We derive two bounds on plasticity: $r$ alone fixes an extremal geometric floor on the utopia distance, while $r$ and $\eta T$ jointly bound forgetting. Across 9720 trajectories, an ANOVA confirms that $r$, $\eta$, and $T$ dominate, while the effect of neutral-set size (emphasized in the biological setting) is negligible. The optimal reach itself follows an approximate inverse power law, yielding a heuristic that sets the optimal reach $\eta T^*$ from the task disagreement alone. The analogy that survives is therefore dynamical rather than geometric, and our setting enables a view of plasticity through the lens of other driven systems in physics and engineering.</description>
        <pubDate>Sun, 16 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v334/skriloff26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v334/skriloff26a.html</guid>
        
        
      </item>
    
      <item>
        <title>Scalable Graph Coreset Selection via Greedy Sampling</title>
        <description>Sampling representative nodes from large graphs is fundamental to graph signal processing and network analysis, yet existing methods require access to the full graph Laplacian, making them impractical at scale. We propose a simple and effective column-selective graph sampling algorithm based on a minimum inner product greedy selection rule. At each iteration, the algorithm accesses only a small random subset of Laplacian columns, requiring no eigendecomposition or global graph traversal, making it well-suited for large-scale graphs where the full Laplacian cannot be stored in memory. We analyze the algorithm under the stochastic block model and show that, when the degree distribution is balanced across nodes, the algorithm achieves sampling proportional to cluster size, and that the resulting mean estimate is controlled for band-limited graph signals in the Paley-Wiener space, with the error decaying as inter-cluster connectivity weakens. Numerical experiments on both synthetic and real-world data validate the effectiveness of the proposed method.</description>
        <pubDate>Sun, 16 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v334/shen26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v334/shen26a.html</guid>
        
        
      </item>
    
      <item>
        <title>Topological Simplification in Predictive Coding Networks</title>
        <description>We study the topology of learned representations in predictive coding networks (PCNs), a neuro-inspired bidirectional architecture, using a quantitative layer-wise persistent homology analysis. We train well-performing PCNs on a synthetic classification dataset ($\geq 99.9$% test accuracy) and on MNIST ($\geq 95$% test accuracy), and measure how topological features change across layers for different architectures and activation functions. We find that smaller PCNs collapse connected components across layers earlier than larger models (Spearman $\rho \in [0.72, 0.79]$ across activations), with model size measured as the sum of hidden-layer widths. We also observe a strong negative correlation ($\rho = -0.58$) between the depth at which simplification occurs and reconstruction error; i.e., architectures that simplify later reconstruct better. Finally, a seed-level bootstrap comparison across architectures and activations shows that PCNs consistently collapse connected components later than matched MLPs, with an average difference of $3.6$ layers. These results suggest that persistent homology offers a useful quantitative lens on the compression–reconstruction tradeoff in PCNs, and that both model capacity and the recurrent, bidirectional dynamics of predictive coding inference shape when this tradeoff is resolved across layers.</description>
        <pubDate>Sun, 16 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v334/shaw26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v334/shaw26a.html</guid>
        
        
      </item>
    
      <item>
        <title>Probing the effect of data representation on transformer learning</title>
        <description>The choice of data representation can affect what a model learns, motivating the need for systematic methods to study this effect. Permutations and their statistics provide a controlled setting in which the underlying object is fixed and only its input representation varies. In this paper, we train small transformers on multiple representations of the same permutations and analyze them using linear probes, attention head ablations, and representational similarity measures. We find that convergence across input representations is task-dependent: models trained to compute the length of the longest increasing subsequence from different representations converge on similar internal computations, while models trained to compute other statistics show more varied behavior, including representation-specific algorithms. Beyond these specific findings, this work provides a controlled testbed for future interpretability research studying when models trained on different input representations learn the same mechanism.</description>
        <pubDate>Sun, 16 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v334/scullen26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v334/scullen26a.html</guid>
        
        
      </item>
    
      <item>
        <title>Finsler Geometry, Graph Neural Networks, and You</title>
        <description>Graph neural network architectures based on the graph Laplacian approximate the Laplace-Beltrami operator, thus limiting their application to isotropic operators. As a nonlinear alternative to the Laplace-Beltrami operator, we consider estimates of the Finsler Laplacian on point clouds sampled from a manifold. We prove that these discrete estimates converge to the true operator on the manifold as the number of point samples grows. Moreover, we show that this operator can be expressed as a graph neural network layer, which we use to define a family of Finslerian graph neural networks constrained to express Finsler geometry. We show that Finslerian graph neural networks recover the geometry underlying nonlinear diffusion equations in practice.</description>
        <pubDate>Sun, 16 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v334/roddenberry26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v334/roddenberry26a.html</guid>
        
        
      </item>
    
      <item>
        <title>The Provable Unsupervised Learning Rule Known as Batch Normalization</title>
        <description>Batch normalization (BN) accelerates training and improves generalization, yet why it works remains unsettled. Prior explanations focus on the loss landscape but cannot explain why BN still helps in architectures that already optimize well. We offer a complementary geometric account: BN is an unsupervised learning rule. For networks with ReLU-like activations, BN’s centering forces every neuron’s hyperplane through the mini-batch mean, anchoring the network’s spline partition to the data, independently of the labels or loss. We prove that this anchoring makes random hyperplanes at initialization separate low-dimensional data clusters with high probability and that BN’s scaling refines partition boundaries near the data manifold. BN thus acts as a label-free mechanism for geometric adaptation and smart initialization.</description>
        <pubDate>Sun, 16 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v334/riedi26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v334/riedi26a.html</guid>
        
        
      </item>
    
      <item>
        <title>Reassessing Muon for Matrix Factorization</title>
        <description>Muon has recently emerged as a strong optimizer for large-scale deep learning, where it reshapes gradient updates through approx- imate orthogonalization and has been reported to outperform Adam and AdamW on large language model training. Its empirical success has motivated a growing theoretical liter- ature that interprets Muon as steepest descent under the spectral norm. Yet it remains unclear which of Muon{’}s advantages stem from its update rule itself and which are arti- facts of the scale, architecture, and data of modern deep networks. In this work we isolate the optimizer from these confounders by studying Muon on a simple, well-understood, and spectrally structured problem: low-rank matrix factorization. Through a controlled and systematically tuned comparison against adaptive baselines, we find that Muon does not consistently outperform AdamW in this setting, and that several previously reported advan- tages are sensitive to hyperparameter choices. Our results give a more nuanced picture of when spectrum-aware orthogonalization helps, and argue for evaluating modern optimizers on controlled problems in addition to end-to-end benchmarks.</description>
        <pubDate>Sun, 16 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v334/parviz26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v334/parviz26a.html</guid>
        
        
      </item>
    
      <item>
        <title>Look Before You Lift: Visual and Quantitative Diagnostics for Topological Deep Learning</title>
        <description>Topological deep learning (TDL) methods rely on lifting raw data into higher-order discrete domains such as simplicial complexes, cell complexes, and hypergraphs. In practice, this lifting step is often treated as a black box: practitioners select a lifting and then tune architectures, with limited visibility into whether the induced higher-order connectivity is meaningful for the downstream task. To address this missing diagnostic layer, we propose a visualization technique called TopoExplorer that leverages the strictly augmented Hasse graph form of topological datasets for exploratory data analysis. For the first time, practitioners can easily visualize the incidence- and adjacency-based neighborhoods that define the lifted dataset, as well as read off key graph metrics that describe its structural and feature landscape. Via an extensive set of experiments across many datasets and liftings, we show that several of these metrics correlate with downstream model performance, suggesting they can help inform TDL preprocessing design. Our perspective reframes the TDL workflow from *lift-train* to *lift-look-design-train*, enabling more principled, interpretable, and efficient model development. TopoExplorer is hosted  at topoexplorer.pagekite.me, and its source code is available at github.com/geometric-intelligence/topoexplorer.</description>
        <pubDate>Sun, 16 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v334/papillon26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v334/papillon26a.html</guid>
        
        
      </item>
    
      <item>
        <title>Tracking Representation Dynamics in Large Language Models with Persistent Homology</title>
        <description>Large language models are commonly aligned through supervised fine-tuning, yet little is known about how their internal representations evolve during this process. We study alignment dynamics using persistent homology by tracking the topology of activation spaces throughout fine-tuning. Across four transformer language models ranging from 1B to 7B parameters and three alignment objectives corresponding to helpful, harmless, and mixed training data, we find that the majority of topological reorganization occurs during the earliest stages of training. A dense checkpoint analysis reveals a transient peak in topological activity followed by rapid stabilization. We further show that different alignment objectives induce distinguishable topological trajectories, while instruction-tuned and pretrained models exhibit qualitatively different patterns of evolution. Our results suggest that persistent homology provides a complementary perspective on alignment, revealing representation-level changes that are not apparent from behavioral metrics alone.</description>
        <pubDate>Sun, 16 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v334/malhotra26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v334/malhotra26a.html</guid>
        
        
      </item>
    
      <item>
        <title>Even with AI, Bijection Discovery is Still Hard: The Opportunities and Challenges of OpenEvolve for Novel Bijection Construction</title>
        <description>Evolutionary program synthesis systems such as AlphaEvolve, OpenEvolve, and ShinkaEvolve offer a new approach to AI-assisted mathematical discovery. These systems use large language models (LLMs) to generate candidate solutions to a problem as human readable code, which is then evolved to improve beyond single-shot outputs. While existing mathematical applications have mostly focused on problems of establishing bounds (e.g., sphere packing), the program synthesis approach is well suited to any problem where the solution takes the form of an explicit construction. With this in mind, in this paper we explore the use of OpenEvolve for combinatorial bijection discovery. We describe the results of applying OpenEvolve to three bijection construction problems involving Dyck paths, two of which are known and one of which is open. We find that while systems like OpenEvolve show promise as a valuable tool for combinatorialists, the problem of finding novel, research-level bijections remains a challenging task for current frontier systems, reinforcing the need for human mathematicians in the loop.</description>
        <pubDate>Sun, 16 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v334/jenne26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v334/jenne26a.html</guid>
        
        
      </item>
    
      <item>
        <title>A Weisfeiler–Leman Characterization of Global-Attention Graph Transformers for Mixed-Integer Linear Programs</title>
        <description>Graph foundation models (GFMs) equipped with global attention are increasingly used to learn representations of mixed-integer linear programs (MILPs), with the stated aim of capturing structural information beyond the locality of standard graph neural networks. We study the expressive power of these architectures through the lens of graph isomorphism testing and ask which MILP instances they map to identical representations. We prove that a broad class of hierarchical graph transformers combining global linear attention, edge-weighted cross-attention, and bipartite message passing is bounded by the one-dimensional Weisfeiler{–}Leman (1-WL) test: for any parameter setting, any pair of 1-WL-equivalent MILP graphs receives an identical graph embedding. The result follows from a compositional analysis in which each architectural component is shown to be a symmetric multiset function and therefore to preserve 1-WL equivalence. We validate the characterization across ten architecturally diverse graph encoders, including Graphormer-, GraphGPS-, Set-Transformer-, and Gasse-style models. Across model capacities, graph scales, and pooling operators, all tested encoders map 1-WL-equivalent non-isomorphic graph pairs to numerically identical embeddings. We then analyze the consequences of this representation equivalence for downstream prediction: graph invariants that vary within a 1-WL equivalence class are not recoverable from the resulting representations. Finally, we localize the source of expressiveness beyond 1-WL to the input encoding rather than the attention mechanism. Random-walk positional encodings separate the constructed pairs, while additional constructions characterize the limits of this remedy. Together, these results provide a theoretical and empirical characterization of the expressive power of global-attention graph foundation models, together with an encoder-agnostic diagnostic for detecting 1-WL-induced representation equivalence.</description>
        <pubDate>Sun, 16 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v334/jahin26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v334/jahin26a.html</guid>
        
        
      </item>
    
      <item>
        <title>Scalably computing metric magnitude</title>
        <description>Applications of metric magnitude often rely on numerically exact results in order to exploit a connection with information theory. We examine various approaches for scaling the dense linear algebra involved and identify hierarchical low-rank solvers as a preferred approach, with a clear path to scales of $10^5$ points on a single powerful workstation, and larger scales for clusters and/or supercomputers using our containerized C++/MPI pipeline.</description>
        <pubDate>Sun, 16 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v334/huntsman26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v334/huntsman26a.html</guid>
        
        
      </item>
    
      <item>
        <title>Local Manifold Explanations with Tangent Space Regression</title>
        <description>Low-dimensional manifold learning is used to embed and visualize high-dimensional data, revealing its underlying geometry. However, identifying which features drive local variation along the manifold remains difficult. Many post-hoc explanation methods target explaining extrinsic embedding coordinates rather than intrinsic manifold structure, or provide only global explanations. In this work, we introduce Local Tangent Space Regression Explanations (LTSREx)), a method to explain the local structure of a manifold in terms of interpretable features by performing sparse linear regression in the tangent space of the manifold at each point, coupled with Tikhonov denoising via the connection Laplacian to ensure that explanations are consistent and vary smoothly across nearby points. We show that our method produces meaningful local explanations on synthetic data, rotated MNIST digits, and two single-cell gene expression datasets. Our code is available at https://github.com/he-jesse/LTSREx.</description>
        <pubDate>Sun, 16 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v334/he26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v334/he26a.html</guid>
        
        
      </item>
    
      <item>
        <title>Freeze, Diffuse, Decode: Task-Aware Adaptation of Transformer Embeddings for Antimicrobial Peptide Design</title>
        <description>Pretrained transformers provide general-purpose molecular embeddings for downstream tasks. While these embeddings provide task-agnostic structural patterns, they lack task-specific alignment, limiting downstream performance. Here, we introduce Freeze, Diffuse, Decode (FDD), a diffusion-based framework that adapts pretrained transformer embeddings to downstream tasks by building a task supervised diffusion geometry over the frozen embeddings, without any backbone training. Applied to antimicrobial peptide design, FDD yields low-dimensional, predictive, and interpretable representations that support property prediction, retrieval, and latent-space interpolation.</description>
        <pubDate>Sun, 16 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v334/gawade26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v334/gawade26a.html</guid>
        
        
      </item>
    
      <item>
        <title>GrAE ScaLE: Grassmannian AutoEncoder for Scalable Linearly-invariant Embeddings</title>
        <description>Supervised learning and generation on subspace-valued data remain difficult in practice, since existing techniques for embeddings of points on the Grassmann manifold can be costly and difficult to scale. In this work, we introduce the Grassmannian Autoencoder (GrAE), an autoencoder architecture designed for subspace-valued data, together with a variational extension, the Grassmannian Variational Autoencoder (GrVAE). These architectures produce informative, robust, low-dimensional latent representations with improved computational scalability. We demonstrate our approaches across three novel subspace datasets designed to help disambiguate failure modes and support benchmarking for future subspace-learning architectures. Our GrAE architectures{’} compact latent representations offer an efficient and expressive alternative for supervised learning on the Grassmann manifold and a practical step toward scalable subspace learning.</description>
        <pubDate>Sun, 16 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v334/emerson26b.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v334/emerson26b.html</guid>
        
        
      </item>
    
      <item>
        <title>Disentanglement by Prediction: A Twin Autoencoder with a Factored Latent Traveler</title>
        <description>We recast disentangled representation learning as latent-space prediction with a factored predictor. Existing approaches—$\beta$-VAE, FactorVAE, $\beta$-TCVAE—cast disentanglement as a statistical property of a marginal latent distribution: they identify what factors describe an image but cannot predict how the latent should move when a factor changes. Drawing on the joint-embedding predictive architecture (JEPA) framework, in which predicting in latent space yields structured representations, we ask: what is the right structure for that predictor when the world is governed by independent generative factors? We propose the Factored-Traversal Twin Autoencoder (FT-TAE): a JEPA-inspired twin-encoder model in which a discrete factor predictor (Gumbel-Softmax) identifies what changed between two views and a learnable state-embedding table $\mathbf{E}\!\in\!\mathbb{R}^{K\times S\times d}$ composes the latent transition as a sum of per-factor displacements. Each displacement is the difference between two table lookups, so unchanged factors contribute exactly zero by construction and the additive form is order-independent without any explicit regularizer. The system is trained end-to-end with a single reconstruction loss—disentanglement emerges as a byproduct of learning accurate factored transitions, not from a marginal-statistics objective. We evaluate on 3D Shapes and MPI3D, measuring factor alignment, traversal accuracy, and compositionality.</description>
        <pubDate>Sun, 16 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v334/emerson26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v334/emerson26a.html</guid>
        
        
      </item>
    
      <item>
        <title>Topology of a Smile: Persistent Homology in Dental Imaging</title>
        <description>CBCT (Cone Beam Computed Tomography) scans provide detailed three-dimensional images, widely used in dentistry for diagnostic and treatment planning tasks. While invaluable, analyzing and documenting these scans is labor-intensive, prompting efforts to automate key steps like the classification and segmentation of anatomical structures to identify tooth types and associated pathologies.  In this article, we propose an approach to automation that leverages persistent homology, a framework from topological data analysis that studies the shape of data by identifying features like connected components, holes, and voids across multiple scales. Persistent homology, together with a support vector machine, allows us to classify teeth in a CBCT scan and to perform diagnostics.  Our method advances the state of the art, reaching average accuracy scores of 97.67% for tooth-labeling and 96.77% for diagnostic tasks, outperforming a CNN trained on the same data with accuracy of 70.27% and 86.67%, respectively.</description>
        <pubDate>Sun, 16 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v334/dahlmeier26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v334/dahlmeier26a.html</guid>
        
        
      </item>
    
      <item>
        <title>Encoding the Euler Characteristic Transform</title>
        <description>The Euler Characteristic Curve (ECC) records the Euler characteristic of a linearly embedded cell complex as a function of filtration height in a given direction, and the Euler Characteristic Transform (ECT) is the injective shape descriptor obtained by collecting ECCs over many directions. How the ECT is encoded for a neural network is itself an inductive bias, conventionally fixed by discretizing each ECC. We introduce a continuous encoding: for each direction and each vertex it records the net Euler-characteristic change attributed to that vertex, producing a per-direction token sequence that a small transformer maps to a feature vector. We separate the resulting pipeline into two stages on orthogonal axes: an ECC encoder that acts within each direction, mapping its curve to a fixed-length vector, and an ECT representation that acts across directions, aggregating the per-direction vectors into one. We study six ECT representation architectures spanning a range of inductive biases, from a structure-agnostic feedforward baseline to convolutional and complex-valued models that preserve equivariance under planar rotations. Across six classification benchmarks covering point clouds, graphs, cubical complexes, and meshes, the continuous encoding improves accuracy on all six datasets, and control experiments attribute the gain to the tokenization itself rather than to the added transformer capacity. The representation architecture matters less than the encoding, and the payoff from its inductive biases depends on the encoding: a feedforward network performs best under continuous encoding but is less robust under discretization than convolutional architectures.</description>
        <pubDate>Sun, 16 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v334/blaser26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v334/blaser26a.html</guid>
        
        
      </item>
    
      <item>
        <title>Loss Landscape Geometry of Partial Differential Equation Emulators: Or, Symmetry Learning via Gradient Alignment</title>
        <description>We study how neural emulators of partial differential equation solution operators learn physical symmetries from data by introducing a hat-matrix diagnostic that quantifies the alignment of parameter updates between symmetry related training examples.  The diagnostic is a metric-weighted overlap of loss        gradients evaluated across group orbits, giving a proximal influence function for symmetry-related examples. Our measurements of gradient alignment  across both translations and rotations for models trained as autoregressive fluid-flow emulators suggest that equivariance arises when training dynamics            propagate gradients coherently throughout symmetry orbits. This finding is based on an empirical correspondence between equivariance error and cross-orbit influence. Both our UNet and ViT architectures exhibit approximate translation equivariance, yet their gradient alignment profiles differ by uniform         versus periodically concentrated influence over the orbit. On a Navier-Stokes dataset, pronounced dihedral equivariance error coincides with suppressed cross-influence, identifying the failure as due to decoupled learning of symmetry group elements.  Our diagnostic is architecture-agnostic and isolates the     mechanism of learning symmetries from data, extending beyond forward-pass equivariance tests by directly assessing whether learning dynamics share information across physically equivalent configurations. We identify symmetry generalization performance with symmetry-compatible gradient transport, implying       that evaluation of scientific machine learning models requires dynamical probes of loss landscape geometry in addition to predictive accuracy.</description>
        <pubDate>Sun, 16 Aug 2026 00:00:00 +0000</pubDate>
        <link>https://proceedings.mlr.press/v334/amarel26a.html</link>
        <guid isPermaLink="true">https://proceedings.mlr.press/v334/amarel26a.html</guid>
        
        
      </item>
    
  </channel>
</rss>
