2024/10/30 by Johann Brehmer, Brehmer, Johann, Sönke Behrends +5 · 5 voices · 20 citations
Computer Science · Materials Science · Mathematics · Physics and Astronomy · #Advanced Graph Neural Networks #Cartography #Geography #Machine Learning in Materials Science #Mathematics #Model Reduction and Neural Networks #Scale (ratio)
paper · pdf · doi:10.48550/arxiv.2410.23179
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2024/10/30 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Given large datasets and sufficient compute, is it beneficial to design neural architectures for the structure and symmetries of each problem? Or is it more efficient to learn them from data? We study empirically how equivariant and non-equivariant networks scale with compute and training samples. Focusing on a benchmark problem of rigid-body interactions and on general-purpose transformer architectures, we perform a series of experiments, varying the model size, training steps, and dataset size. We find evidence for three conclusions. First, equivariance improves data efficiency, but training non-equivariant models with data augmentation can close this gap given sufficient epochs. Second, scaling with compute follows a power law, with equivariant models outperforming non-equivariant ones at each tested compute budget. Finally, the optimal allocation of a compute budget onto model size and training duration differs between equivariant and non-equivariant models.