vix.ing · top · new · best · stats · spec

Crown, Frame, Reverse: Layer-Wise Scaling Variants for LLM Pre-Training

2025/09/08 by Andrei Baroian, Baroian, Andrei, Kasper Notebomer +1 · 1 citation
Engineering · Medicine · #Artificial Intelligence (cs.AI) #Cardiac Valve Diseases and Treatments #Computation and Language (cs.CL) #FOS: Computer and information sciences #Orthopaedic implants and arthroplasty #Tunneling and Rock Mechanics

paper · pdf · doi:10.48550/arxiv.2509.06518

openalex publication_date 2025/09/08 · openalex created_date 2025/10/11 · openalex updated_date 2026/07/28

Abstract

Transformer-based language models traditionally use uniform (isotropic) layer sizes, yet they ignore the diverse functional roles that different depths can play and their computational capacity needs. Building on Layer-Wise Scaling (LWS) and pruning literature, we introduce three new LWS variants - Framed, Reverse, and Crown - that redistribute FFN widths and attention heads via two or three-point linear interpolation in the pre-training stage. We present the first systematic ablation of LWS and its variants, on a fixed budget of 180M parameters, trained on 5B tokens. All models converge to similar losses and achieve better performance compared to an equal-cost isotropic baseline, without a substantial decrease in training throughput. This work represents an initial step into the design space of layer-wise architectures for pre-training, but future work should scale experiments to orders of magnitude more tokens and parameters to fully assess their potential.

Citations

Cited by

Related