When Attention Beats Fourier: Multi-Scale Transformers for PDE Solving on Irregular Domains
Researchers from arXiv have introduced the Multi-Scale Attention Transformer (MSAT), a new deep learning architecture designed to solve partial differential equations (PDEs) on irregular domains. The study addresses the critical problem of architecture selection, determining when transformer-based models with learned attention outperform traditional Fourier-domain neural operators. MSAT encodes spatiotemporal solution histories as token sequences and utilizes a composite supervised objective with optional physics-informed regularization. In comprehensive empirical evaluations against nine baselines, including PINNs, FNO, and Mamba-NO, MSAT achieved state-of-the-art generalization on complex geometry problems. Notably, it demonstrated a 3.7x improvement over FNO on the Heat2D-CG benchmark and significantly faster inference times compared to Mamba-NO. Ablation studies revealed that while physics priors reduce error in diffusion-dominated problems, they may degrade performance in chaotic flow regimes. The paper also provides theoretical approximation error bounds based on domain boundary complexity, offering a principled rule for selecting appropriate architectures for specific PDE solving tasks.
Wire timeline
When Attention Beats Fourier: Multi-Scale Transformers for PDE Solving on Irregular Domains
Researchers from arXiv have introduced the Multi-Scale Attention Transformer (MSAT), a new deep learning architecture designed to solve partial differential equations (PDEs) on irregular domains. The study addresses the critical problem of architecture selection, determining when transformer-based models with learned attention outperform traditional Fourier-domain neural operators. MSAT encodes spatiotemporal solution histories as token sequences and utilizes a composite supervised objective with optional physics-informed regularization. In comprehensive empirical evaluations against nine baselines, including PINNs, FNO, and Mamba-NO, MSAT achieved state-of-the-art generalization on complex geometry problems. Notably, it demonstrated a 3.7x improvement over FNO on the Heat2D-CG benchmark and significantly faster inference times compared to Mamba-NO. Ablation studies revealed that while physics priors reduce error in diffusion-dominated problems, they may degrade performance in chaotic flow regimes. The paper also provides theoretical approximation error bounds based on domain boundary complexity, offering a principled rule for selecting appropriate architectures for specific PDE solving tasks.
cs.AI updates on arXiv.org