Cardiff University | Prifysgol Caerdydd ORCA
Online Research @ Cardiff 
WelshClear Cookie - decide language by browser settings

Dynamic Fractal Mamba: A neural renormalization group flow for scale-invariant sequence modeling

Fang, Shenglei, Sun, Xianfang ORCID: https://orcid.org/0000-0002-6114-0766 and Zhou, You ORCID: https://orcid.org/0000-0002-1743-1291 2026. Dynamic Fractal Mamba: A neural renormalization group flow for scale-invariant sequence modeling. Presented at: ICML 2026, Seoul, South Korea, 6-11 July 2026. Proceedings of the 43rd International Conference on Machine Learning. Proceedings of Machine Learning Research. , vol.306 pp. 29318-29346.

[thumbnail of 1016_Dynamic_Fractal_Mamba_A_N.pdf] PDF - Published Version
Available under License Creative Commons Attribution.

Download (2MB)

Abstract

Sequence models typically operate at a fixed temporal or spatial scale and struggle to generalize to substantially longer horizons or higher resolutions without retraining. Existing hierarchical architectures expand receptive fields but rely on scale-specific parameters and lack mechanisms to enforce consistent dynamics across scales. We propose Dynamic Fractal Mamba (DF-Mamba), a recursive state-space model that applies a single shared operator across multiple scales. By sharing parameters across recursion depths and exponentially scaling the effective time step, DF-Mamba achieves an exponentially expanding receptive field while preserving linear computational complexity. A learned content-aware coarse-graining module aggregates representations across scales. Auxiliary reconstruction and cross-scale consistency objectives stabilize recursive training. We evaluate DF-Mamba on long-range time-series forecasting, spatial transcriptomics, and computational pathology. Across all tasks, DF-Mamba consistently outperforms Transformers and flat Mamba baselines while using fewer parameters and maintaining linear-time scalability. Importantly, models trained on short sequences or low-resolution inputs generalize in a zero-shot manner to substantially larger temporal and spatial scales unseen during training. These results demonstrate that recursive parameter sharing provides an effective inductive bias for learning scale-consistent and efficient sequence representations.

Item Type: Conference or Workshop Item - published (Paper)
Status: Published
Schools: Schools > Medicine
Schools > Computer Science & Informatics
ISSN: 1938-7228
Related URLs:
Date of First Compliant Deposit: 21 July 2026
Last Modified: 07 Oct 2026 11:24
URI: https://orca.cardiff.ac.uk/id/eprint/188359

Actions (repository staff only)

Edit Item Edit Item

Downloads

Downloads per month over past year

View more statistics