Query-Key Normalization for Transformers
2020/10/08 by Alex Henry, Prudhvi Raj Dachapally, Henry, Alex +5 · 79 citations
Biochemistry, Genetics and Molecular Biology · Computer Science · #Biomedical Text Mining and Ontologies #Natural Language Processing Techniques #Topic Modeling #cs.AI #cs.CL #cs.LG
paper · pdf · doi:10.48550/arxiv.2010.04245
8 pages, 2 figures, accepted at Findings of EMNLP 2020
arxiv created 2020/10/08 · arxiv updated 2020/10/12
Abstract
Low-resource language translation is a challenging but socially valuable NLP task. Building on recent work adapting the Transformer's normalization to this setting, we propose QKNorm, a normalization technique that modifies the attention mechanism to make the softmax function less prone to arbitrary saturation without sacrificing expressivity. Specifically, we apply ℓ2 normalization along the head dimension of each query and key matrix prior to multiplying them and then scale up by a learnable parameter instead of dividing by the square root of the embedding dimension. We show improvements averaging 0.928 BLEU over state-of-the-art bilingual benchmarks for 5 low-resource translation pairs from the TED Talks corpus and IWSLT'15.
Cited by
- Completed Hyperparameter Transfer across Modules, Width, Depth, Batch and Duration
- Raven: High-Recall Sequence Modeling with Sparse Memory Routing
- Next-Embedding Prediction Makes Strong Vision Learners
- DVGT: Driving Visual Geometry Transformer
- VFMF: World Modeling by Forecasting Vision Foundation Model Features
- Long-LRM++: Preserving Fine Details in Feed-Forward Wide-Coverage Reconstruction
- PCMind-2.1-Kaiyuan-2B Technical Report
- Multi-view Pyramid Transformer: Look Coarser to See Broader
- Controlling changes to attention logits
- Selective Rotary Position Embedding
- Mixture of States: Routing Token-Level Dynamics for Multimodal Generation
- Learning to Focus: Focal Attention for Selective and Scalable Transformers
- Arcee Trinity Large Technical Report
- Generative Modeling via Drifting
- DynaRend: Learning 3D Dynamics via Masked Future Rendering for Robotic Manipulation
- Key and Value Weights Are Probably All You Need: On the Necessity of the Query, Key, Value weight Triplet in Self-Attention Transformers
- SeeDNorm: Self-Rescaled Dynamic Normalization
- LongCat-Video Technical Report
- Streaming Generation for Music Accompaniment
- Scalable Vision-Language-Action Model Pretraining for Robotic Manipulation with Real-Life Human Activity Videos
- VFM-VAE: Vision Foundation Models Can Be Good Tokenizers for Latent Diffusion Models
- Latent Diffusion Model without Variational Autoencoder
- Generative World Modelling for Humanoids: 1X World Model Challenge Technical Report
- Exact Causal Attention with 10% Fewer Operations
- UniVid: The Open-Source Unified Video Model
- Beyond Outliers: A Study of Optimizers Under Quantization
- ZUNA1.1: A more flexible EEG foundation model for Denoising and Super-resolution
- ELF: Embedded Language Flows
- Tenma: Robust Cross-Embodiment Robot Manipulation with Diffusion Transformer
- Open-sci-ref-0.01: open and reproducible reference baselines for language model and dataset comparison
- FLOWER: Democratizing Generalist Robot Policies with Efficient Vision-Language-Action Flow Policies
- WISCA: A Lightweight Model Transition Method to Improve LLM Training via Weight Scaling
- Efficient Patent Searching Using Graph Transformers
- Whisfusion: Parallel ASR Decoding with Masked Diffusion
- Matrix-Driven Identification and Reconstruction of LLM Weight Homology
- H3R: Hybrid Multi-view Correspondence for Generalizable 3D Reconstruction
- Motif 2.6B Technical Report
- iLRM: An Iterative Large 3D Reconstruction Model
- Training Transformers with Enforced Lipschitz Constants
- GR-3 Technical Report
- Mellum2 Technical Report
- Enigma: An Efficient Model for Deciphering Regulatory Genomics
- DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD
- The Less You Depend, The More You Learn: Synthesizing Novel Views from Sparse, Unposed Images with Minimal 3D Knowledge
- NeoBabel: A Multilingual Open Tower for Visual Generation
- Lost in Latent Space: An Empirical Study of Latent Diffusion Models for Physics Emulation
- IntFold: A Controllable Foundation Model for General and Specialized Biomolecular Structure Prediction
- Characterization and Mitigation of Training Instabilities in Microscaling Formats
- CausalPFN: Amortized Causal Effect Estimation via In-Context Learning
- CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech
- DeepVerse: 4D Autoregressive Video Generation as a World Model
- D-AR: Diffusion via Autoregressive Models
- Test-Time Training Done Right
- Learning in Compact Spaces with Approximately Normalized Transformer
- RenderFormer: Transformer-based Neural Rendering of Triangle Meshes with Global Illumination
- RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding
- Taming Transformer Without Using Learning Rate Warmup
- Towards Fully FP8 GEMM LLM Training at Scale
- Absolute Coordinates Make Motion Generation Easy
- P-MUSE: Prompt-MIDI-Optional Model for Unified Instrumental Music Synthesis and Editing
- Applications of Modular Co-Design for De Novo 3D Molecule Generation
- Revealing Language Model Trajectories via Kullback-Leibler Divergence
- Dense Local Dependencies Induce Attention-Logit Explosion and Training Instability During Long-Sequence Transformer Training
- One QK Channel, Many Sources: Guarding Low-Precision Attention Collapse
- Training nGPT
- Fast Text-to-Audio Generation with Adversarial Post-Training
- Towards Foundation Models for Experimental Readout Systems Combining Discrete and Continuous Data
- PluRel: Synthetic Data unlocks Scaling Laws for Relational Foundation Models
- RayZer: A Self-supervised Large View Synthesis Model
- Fractional Rotation, Full Potential? Investigating Performance and Convergence of Partial RoPE
- Pion: A Spectrum-Preserving Optimizer via Orthogonal Equivalence Transformation
- Hallucination in World Models is Predictable and Preventable
- WEAVER, Better, Faster, Longer: An Effective World Model for Robotic Manipulation
- Mean Mode Screaming: Mean--Variance Split Residuals for 1000-Layer Diffusion Transformers
- Screening Is Enough
- Hybrid Linear Attention Done Right: Efficient Distillation and Effective Architectures for Extremely Long Contexts
- Surface Keypoint Representation for Multi-Object and Articulated Human-Object Interaction Generation
- LoopMTP: A looped transformer guided by latent multi-token prediction
- Foundation Models for Time Series: A Survey
Related