RoFormer: Enhanced Transformer with Rotary Position Embedding
2021/04/20 by Jianlin Su, Su, Jianlin, Yu Lu +8 · 792 citations
Computer Science · #Advanced Text Analysis Techniques #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2104.09864
openalex publication_date 2021/04/20 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Position encoding recently has shown effective in the transformer architecture. It enables valuable supervision for dependency modeling between elements at different positions of the sequence. In this paper, we first investigate various methods to integrate positional information into the learning process of transformer-based language models. Then, we propose a novel method named Rotary Position Embedding(RoPE) to effectively leverage the positional information. Specifically, the proposed RoPE encodes the absolute position with a rotation matrix and meanwhile incorporates the explicit relative position dependency in self-attention formulation. Notably, RoPE enables valuable properties, including the flexibility of sequence length, decaying inter-token dependency with increasing relative distances, and the capability of equipping the linear self-attention with relative position encoding. Finally, we evaluate the enhanced transformer with rotary position embedding, also called RoFormer, on various long text classification benchmark datasets. Our experiments show that it consistently overcomes its alternatives. Furthermore, we provide a theoretical analysis to explain some experimental results. RoFormer is already integrated into Huggingface: \urlhttps://huggingface.co/docs/transformers/modeldoc/roformer.
Citations
Cited by
- Self-Attention Dynamics with Rotary Position Embeddings: Twisted States and Explicit Consensus Rates on the Sphere
- Qwen-Music Technical Report
- HY-Motion 1.0: Scaling Flow Matching Models for Text-To-Motion Generation
- Diversity or Precision? A Deep Dive into Next Token Prediction
- Argus: Token Aware Distributed LLM Inference Optimization
- Chessformer: A Unified Architecture for Chess Modeling
- WeDLM: Reconciling Diffusion Language Models with Standard Causal Attention for Fast Inference
- SPECTRE: Spectral Pre-training Embeddings with Cylindrical Temporal Rotary Position Encoding for Fine-Grained sEMG-Based Movement Decoding
- Qwen-Audio-3.0-TTS: Freely Controllable and Highly Robust Speech Synthesis with Multi-Stage Training Paradigm
- On the Existence and Behavior of Secondary Attention Sinks
- RayRoPE: Projective Ray Positional Encoding for Multi-view Attention
- Selecting Language Models for Social Science: Start Small, Start Open, and Validate
- The Gate Always Closes: On Injecting Auxiliary Signals into Frozen Vision-Language Models
- seqLens: Optimizing Language Models for Genomic Predictions
- StreamIndex: Memory-Bounded Compressed Sparse Attention via Streaming Top-k
- A satellite foundation model for improved wealth monitoring
- Doc-to-LoRA: Learning to Instantly Internalize Contexts
- Accelerating Language Model Workflows with Prompt Choreography
- Scale Weight Decay and Train Better
- STEER: Steerable Dyadic Head Avatars
- Disentangling Semantic Attention from Structural Bias in the Attention Manifold
- KAP: Bridging the Knowledge Selection-Runtime Consumption Gap in LLM Systems
- CameraAnything: Refilming Videos with Arbitrary Camera Control
- CAPT: A Multi-task Continuous Autoregressive Transformer enabling Cross-dataset and Cross-species Transfer for Calcium Population Dynamics
- The Spatial Blindspot of Vision-Language Models
- OmniMate: Open-Ended Real-Time Streaming Audio-Visual Generation for Interactive Avatars
- Bekko Embedding: Parameter-Efficient Multilingual Retrieval with Ultra-Compact Encoders
- Raven: High-Recall Sequence Modeling with Sparse Memory Routing
- Phase Structure in Rotary Attention: A Spectral Framework for Semantic Continuity and Execution-Boundary Governance
- Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model
- Ranked by Position: Order Sensitivity as an Exploitable Attack Surface in LLM Listwise Recommenders
- Out-of-Length Scene Text Recognition: A Two-Axis Diagnosis and a Training-Free Fix
- T2LDM++: A Self-Conditioned Representation Guided Diffusion Model for Realistic Text-to-LiDAR Scene Generation
- Temporal Context Reinstatement Drives Episodic-Like Order Memory in Long-Context Language Models
- Self-Distillation of Hidden Layers for Self-Supervised Representation Learning
- ZipMap: Linear-Time Stateful 3D Reconstruction via Test-Time Training
- CLIP Is Shortsighted: Paying Attention Beyond the First Sentence
- PixelGen: Improving Pixel Diffusion with Perceptual Supervision
- Deep Delta Learning
- DPAR: Dynamic Patchification for Efficient Autoregressive Visual Generation
- StreamAvatar: Streaming Diffusion Models for Real-Time Interactive Human Avatars
- Gamayun's Path to Multilingual Mastery: Cost-Efficient Training of a 1.5B-Parameter LLM
- AstraNav-World: World Model for Foresight Control and Consistency
- Towards Long-window Anchoring in Vision-Language Model Distillation
- UltraShape 1.0: High-Fidelity 3D Shape Generation via Scalable Geometric Refinement
- Learning from Next-Frame Prediction: Autoregressive Video Modeling Encodes Effective Representations
- Benchmarking and Enhancing VLM for Compressed Image Understanding
- Beyond Weight Adaptation: Feature-Space Domain Injection for Cross-Modal Ship Re-Identification
- Architectural Trade-offs in Small Language Models Under Compute Constraints
- DreaMontage: Arbitrary Frame-Guided One-Shot Video Generation
- Generating the Past, Present and Future from a Motion-Blurred Image
- Pushing the Frontier of Audiovisual Perception with Large-Scale Multimodal Correspondence Learning
- LoGoPlanner: Localization Grounded Navigation Policy with Metric-aware Visual Geometry
- StoryMem: Multi-shot Long Video Storytelling with Memory
- D2Pruner: Debiased Importance and Structural Diversity for MLLM Token Pruning
- ReasonCD: A Multimodal Reasoning Large Model for Implicit Change-of-Interest Semantic Mining
- DIVER-1: Scaling Intracranial EEG Foundation Models for Transferable Representations
- DeepGESI: A Non-Intrusive Objective Evaluation Model for Predicting Speech Intelligibility in Hearing-Impaired Listeners
- OmniEgoCap: Camera-Agnostic Sequence-Level Egocentric Motion Reconstruction
- Alternative positional encoding functions for neural transformers
- Memorize-and-Generate: Towards Long-Term Consistency in Real-Time Video Generation
- SmartSight: Mitigating Hallucination in Video-LLMs Without Compromising Video Understanding via Temporal Attention Collapse
- Uni-Neur2Img: Unified Neural Signal-Guided Image Generation, Editing, and Stylization via Diffusion Transformers
- Layout-Aware Text Editing for Efficient Transformation of Academic PDFs to Markdown
- Diffusion Forcing for Multi-Agent Interaction Sequence Modeling
- Mitigating Forgetting in Low Rank Adaptation
- Physics of Language Models: Part 4.1, Architecture Design and the Magic of Canon Layers
- XLM: A Python package for non-autoregressive language models
- Next-Embedding Prediction Makes Strong Vision Learners
- In-Context Algebra
- Hearing to Translate: The Effectiveness of Speech Modality Integration into LLMs
- How Smoothing is N-simplicial Attention?
- CTkvr: KV Cache Retrieval for Long-Context LLMs via Centroid then Token Indexing
- T5Gemma 2: Seeing, Reading, and Understanding Longer
- Native and Compact Structured Latents for 3D Generation
- Segmental Attention Decoding With Long Form Acoustic Encodings
- WorldPlay: Towards Long-Term Geometric Consistency for Real-Time Interactive World Modeling
- Dual-objective Language Models: Training Efficiency Without Overfitting
- SS4D: Native 4D Generative Model via Structured Spacetime Latents
- TorchTraceAP: A New Benchmark Dataset for Detecting Performance Anti-Patterns in Computer Vision Models
- Context Representation via Action-Free Transformer encoder-decoder for Meta Reinforcement Learning
- FacEDiT: Unified Talking Face Editing and Generation via Facial Motion Infilling
- EXAONE Path 2.5: Pathology Foundation Model with Multi-Omics Alignment
- Let's (not) just put things in Context: Test-Time Training for Long-Context LLMs
- The Devil is in Attention Sharing: Improving Complex Non-rigid Image Editing Faithfulness via Attention Synergy
- BlossomRec: Block-level Fused Sparse Attention Mechanism for Sequential Recommendations
- LitePT: Lighter Yet Stronger Point Transformer
- ReFusion: A Diffusion Large Language Model with Parallel Autoregressive Decoding
- Improving Recursive Transformers with Mixture of LoRAs
- MiniLingua: A Small Open-Source LLM for European Languages
- RecTok: Reconstruction Distillation along Rectified Flow
- From Small to Large: Generalization Bounds for Transformers on Variable-Size Inputs
- Unlocking Generalization in Polyp Segmentation with DINO Self-Attention "keys"
- Cross-Modal Representational Knowledge Distillation for Enhanced Spike-Informed LFP Modeling
- StreamingAssistant: Efficient Visual Token Pruning for Accelerating Online Video Understanding
- CurvaDion: Curvature-Adaptive Distributed Orthonormalization
- Boosting Monocular Metric Depth Estimation via Bokeh Rendering
- BaRISTA: Brain Scale Informed Spatiotemporal Representation of Human Intracranial Neural Activity
- EditMGT: Unleashing Potentials of Masked Generative Transformers in Image Editing
- JoyAvatar: Real-time and Infinite Audio-Driven Avatar Generation with Autoregressive Diffusion
- Sliced ReLU attention: Quasi-linear contextual expressivity via sorting
- REST: Diffusion-based Real-time End-to-end Streaming Talking Head Generation via ID-Context Caching and Asynchronous Streaming Distillation
- Seeing to Act, Prompting to Specify: A Bayesian Factorization of Vision Language Action Policy
- Reframing Music-Driven 2D Dance Pose Generation as Multi-Channel Image Generation
- Mining Legal Arguments to Study Judicial Formalism
- AutoRefiner: Improving Autoregressive Video Diffusion Models via Reflective Refinement Over the Stochastic Sampling Path
- AlcheMinT: Fine-grained Temporal Control for Multi-Reference Consistent Video Generation
- Network and Compiler Optimizations for Efficient Linear Algebra Kernels in Private Transformer Inference
- SceneMaker: Open-set 3D Scene Generation with Decoupled De-occlusion and Pose Estimation Model
- Bidirectional Normalizing Flow: From Data to Noise and Back
- OmniView: An All-Seeing Diffusion Model for 3D and 4D View Synthesis
- Asynchronous Reasoning: Training-Free Interactive Thinking LLMs
- SpaceDrive: Infusing Spatial Awareness into VLM-based Autonomous Driving
- Cross-modal Retrieval Models for Stripped Binary Analysis
- EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs
- StereoSpace: Depth-Free Synthesis of Stereo Geometry via End-to-End Diffusion in a Canonical Space
- HiF-VLA: Hindsight, Insight and Foresight through Motion Representation for Vision-Language-Action Models
- Mixture of Lookup Key-Value Experts
- Circuits, Features, and Heuristics in Molecular Transformers
- Toward Closed-loop Molecular Discovery via Language Model, Property Alignment and Strategic Search
- GimbalDiffusion: Gravity-Aware Camera Control for Video Generation
- Learning Unmasking Policies for Diffusion Language Models
- Revisiting the Scaling Properties of Downstream Metrics in Large Language Model Training
- InfiniteVL: Synergizing Linear and Sparse Attention for Highly-Efficient, Unlimited-Input Vision-Language Models
- ContextDrag: Precise Drag-Based Image Editing via Context-Preserving Token Injection and Position-Aligned Attention
- The Unseen Bias: How Norm Discrepancy in Pre-Norm MLLMs Leads to Visual Information Loss
- EgoX: Egocentric Video Generation from a Single Exocentric Video
- A scalable and real-time neural decoder for topological quantum codes
- Meta Lattice: Model Space Redesign for Cost-Effective Industry-Scale Ads Recommendations
- Short-Context Dominance: How Much Local Context Natural Language Actually Needs?
- Unveiling Latent Knowledge in Chemistry Language Models through Sparse Autoencoders
- One Layer Is Enough: Adapting Pretrained Visual Encoders for Image Generation
- Mary, the Cheeseburger-Eating Vegetarian: Do LLMs Recognize Incoherence in Narratives?
- ViSA: 3D-Aware Video Shading for Real-Time Upper-Body Avatar Creation
- PCMind-2.1-Kaiyuan-2B Technical Report
- Flash Multi-Head Feed-Forward Network
- Beyond Real: Imaginary Extension of Rotary Position Embeddings for Long-Context LLMs
- Unified Video Editing with Temporal Reasoner
- Generalization of Long-Range Machine Learning Potentials in Complex Chemical Spaces
- ContextAnyone: Context-Aware Diffusion for Character-Consistent Text-to-Video Generation
- Materium: An Autoregressive Approach for Material Generation
- A Hetero-Associative Sequential Memory Model Utilizing Neuromorphic Signals: Validated on a Mobile Manipulator
- JoPano: Unified Panorama Generation via Joint Modeling
- JT-DA: Enhancing Data Analysis with Tool-Integrated Table Reasoning Large Language Models
- Large Language Model-Based Generation of Discharge Summaries
- Unleashing the Intrinsic Visual Representation Capability of Multimodal Large Language Models
- SCAIL: Towards Studio-Grade Character Animation via In-Context Learning of 3D-Consistent Pose Representations
- TinyMyo: a Tiny Foundation Model for Flexible EMG Signal Processing at the Edge
- NEAT: Neighborhood-Guided, Efficient, Autoregressive Set Transformer for 3D Molecular Generation
- ShaRP: SHAllow-LayeR Pruning for Video Large Language Models Acceleration
- Deep Forcing: Training-Free Long Video Generation with Deep Sink and Participative Compression
- BulletTime: Decoupled Control of Time and Camera Pose for Video Generation
- HiPPO: Exploring A Novel Hierarchical Pronunciation Assessment Approach for Spoken Languages
- GeoPE:A Unified Geometric Positional Embedding for Structured Tensors
- EMMA: Efficient Multimodal Understanding, Generation, and Editing with a Unified Architecture
- Generative Recursive Reasoning
- Parabolic Position Encoding: Vision-Centric, Principled, Extrapolatable, General
- Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models
- Reward Forcing: Efficient Streaming Video Generation with Rewarded Distribution Matching Distillation
- Denoise to Track: Harnessing Video Diffusion Priors for Robust Correspondence
- Self-Paced and Self-Corrective Masked Prediction for Movie Trailer Generation
- UltraImage: Rethinking Resolution Extrapolation in Image Diffusion Transformers
- Jina-VLM: Small Multilingual Vision Language Model
- Reconstructing KV Caches with Cross-layer Fusion For Enhanced Transformers
- Context-Aware Hierarchical Learning: A Two-Step Paradigm towards Safer LLMs
- Overcoming State Inertia: Minimally Invasive Temporal Alignment for Evolving Contexts
- CSMapping: Scalable Crowdsourced Semantic Mapping and Topology Inference for Autonomous Driving
- A Preliminary Study on the Promises and Challenges of Native Top-k Sparse Attention
- RELIC: Interactive Video World Model with Long-Horizon Memory
- UniQL: Unified Quantization and Low-rank Compression for Adaptive Edge LLMs
- CAMEO: Correspondence-Attention Alignment for Multi-View Diffusion Models
- MultiShotMaster: A Controllable Multi-Shot Video Generation Framework
- Video4Spatial: Towards Visuospatial Intelligence with Context-Guided Video Generation
- AutoBrep: Autoregressive B-Rep Generation with Unified Topology and Geometry
- In-Context Sync-LoRA for Portrait Video Editing
- MSPT: Efficient Large-Scale Physical Modeling via Parallelized Multi-Scale Attention
- Graph VQ-Transformer (GVT): Fast and Accurate Molecular Generation via High-Fidelity Discrete Latents
- When Refusals Fail: Unstable Safety Mechanisms in Long-Context LLM Agents
- Spatiotemporal Pyramid Flow Matching for Climate Emulation
- TUNA: Taming Unified Visual Representations for Native Unified Multimodal Models
- Improved Mean Flows: On the Challenges of Fastforward Generative Models
- Four Over Six: More Accurate NVFP4 Quantization with Adaptive Block Scaling
- Rectifying LLM Thought from Lens of Optimization
- Scaling and context steer LLMs along the same computational path as the human brain
- Efficient Training of Diffusion Mixture-of-Experts Models: A Practical Recipe
- ResDiT: Evoking the Intrinsic Resolution Scalability in Diffusion Transformers
- Know Thyself by Knowing Others: Learning Neuron Identity from Population Context
- DreamingComics: A Story Visualization Pipeline via Subject and Layout Customized Generation using Video Models
- Cosine-Similarity Methods for Efficient Training and Sampling in High-Dimensional Latent Spaces
- G-KV: Decoding-Time KV Cache Eviction with Global Attention
- BioArc: Discovering Optimal Neural Architectures for Biological Foundation Models
- AnyTalker: Scaling Multi-Person Talking Video Generation with Interactivity Refinement
- LUMOS: Large User MOdels for User Behavior Prediction
- Captain Safari: A World Engine with Pose-Aligned 3D Memory
- Ovis-Image Technical Report
- Markovian Scale Prediction: A New Era of Visual Autoregressive Generation
- Every Token Counts: Generalizing 16M Ultra-Long Context in Large Language Models
- Flowing Backwards: Improving Normalizing Flows via Reverse Representation Alignment
- The Collapse of Patches
- DriveVGGT: Visual Geometry Transformer for Autonomous Driving
- Controlling changes to attention logits
- Evaluation of Large Language Models for Numeric Anomaly Detection in Power Systems
- Generating Separated Singing Vocals Using a Diffusion Model Conditioned on Music Mixtures
- HTTM: Head-wise Temporal Token Merging for Faster VGGT
- Softmax Transformers are Turing-Complete
- Dynamical Properties of Tokens in Self-Attention and Effects of Positional Encoding
- Building a Foundation Model for Trajectory from Scratch
- Adam Simplified: Bias Correction Debunked
- Object-Centric Vision Token Pruning for Vision Language Models
- Uplifting Table Tennis: A Robust, Real-World Application for 3D Trajectory and Spin Estimation
- UltraViCo: Breaking Extrapolation Limits in Video Diffusion Transformers
- MambaEye: A Size-Agnostic Visual Encoder with Causal Sequential Processing
- LumiTex: Towards High-Fidelity PBR Texture Generation with Illumination Context
- GigaWorld-0: World Models as Data Engine to Empower Embodied AI
- ReDirector: Creating Any-Length Video Retakes with Rotary Camera Encoding
- Zero-Knowledge Proof Based Verifiable Inference of Models
- Predicting partially observable dynamical systems via diffusion models with a multiscale inference scheme
- HunyuanOCR Technical Report
- DeCo: Frequency-Decoupled Pixel Diffusion for End-to-End Image Generation
- MapFormer: Self-Supervised Learning of Cognitive Maps with Input-Dependent Positional Embeddings
- SimDiff: Simpler Yet Better Diffusion Model for Time Series Point Forecasting
- SWAN: Sparse Winnowed Attention for Reduced Inference Memory via Decompression-Free KV-Cache Compression
- STCDiT: Spatio-Temporally Consistent Diffusion Transformer for High-Quality Video Super-Resolution
- LATTICE: Democratize High-Fidelity 3D Generation at Scale
- Phase-Aligned RoPE for Mixed-Resolution Diffusion Transformer
- CycleChemist: A Dual-Pronged Machine Learning Framework for Organic Photovoltaic Discovery
- InstructAudio: Unified speech and music generation with natural language instruction
- NAF: Zero-Shot Feature Upsampling via Neighborhood Attention Filtering
- MammothModa2: A Unified AR-Diffusion Framework for Multimodal Understanding and Generation
- Muskie: Multi-view Masked Image Modeling for 3D Vision Pre-training
- AdaPerceiver: Transformers with Adaptive Width, Depth, and Tokens
- UltraFlux: Data-Model Co-Design for High-quality Native 4K Text-to-Image Generation across Diverse Aspect Ratios
- PrefixGPT: Prefix Adder Optimization by a Generative Pre-trained Transformer
- Generative Adversarial Post-Training Mitigates Reward Hacking in Live Human-AI Music Interaction
- Blu-WERP (Web Extraction and Refinement Pipeline): A Scalable Pipeline for Preprocessing Large Language Model Datasets
- A cross-species neural foundation model for end-to-end speech decoding
- Unmasking Airborne Threats: Guided-Transformers for Portable Aerosol Mass Spectrometry
- Selective Rotary Position Embedding
- DeepCoT: Deep Continual Transformers for Real-Time Inference on Data Streams
- Scaling Self-Supervised and Cross-Modal Pretraining for Volumetric CT Transformers
- RoSA: Enhancing Parameter-Efficient Fine-Tuning via RoPE-aware Selective Adaptation in Large Language Models
- MolSight: Optical Chemical Structure Recognition with SMILES Pretraining, Multi-Granularity Learning and Reinforcement Learning
- Predicting one-year clinical instability and mortality in heart failure patients using sequence modeling
- Walrus: A Cross-Domain Foundation Model for Continuum Dynamics
- NaTex: Seamless Texture Generation as Latent Color Diffusion
- SUNAC: Source-aware Unified Neural Audio Codec
- Recurrent Autoregressive Diffusion: Global Memory Meets Local Attention
- Generative Photographic Control for Scene-Consistent Video Cinematic Editing
- RoMa v2: Harder Better Faster Denser Feature Matching
- Hierarchical Token Prepending: Enhancing Information Flow in Decoder-based LLM Embeddings
- LiteCache: A Query Similarity-Driven, GPU-Centric KVCache Subsystem for Efficient LLM Inference
- Segment Anything Across Shots: A Method and Benchmark
- OPFormer: Object Pose Estimation leveraging foundation model with geometric encoding
- LOBERT: Generative AI Foundation Model for Limit Order Book Messages
- On-Device Fine-Tuning via Backprop-Free Zeroth-Order Optimization
- Φeat: Physically Grounded Material Feature Representation
- PAS: A Training-Free Stabilizer for Temporal Encoding in Video LLMs
- Know Your Limits: Entropy Estimation Modeling for Compression and Generalization
- A3: Attention-Aware Accurate KV Cache Fusion for Fast Large Language Model Serving
- SiDGen: Structure-informed Diffusion for Generative modeling of Ligands for Proteins
- Making Every Head Count: Sparse Attention Without the Speed-Performance Trade-off
- DoPE: Denoising Rotary Position Embedding
- Branching Flows: Discrete, Continuous, and Manifold Flow Matching with Splits and Deletions
- Leveraging unlabelled data for generalizable neural population decoding
- Clifford Algebraic Rotor Embeddings : Maybe embeddings should start to CARE
- Galactification: painting galaxies onto dark matter only simulations using a transformer-based model
- Do traveling waves make good positional encodings?
- A Unified Geometric Field Theory Framework for Transformers: From Manifold Embeddings to Kernel Modulation
- Gate-level boolean evolutionary geometric attention neural networks
- Laytrol: Preserving Pretrained Knowledge in Layout Control for Multimodal Diffusion Transformers
- CellARC: Measuring Intelligence with Cellular Automata
- A Circular Argument : Does RoPE need to be Equivariant for Vision?
- oboro: Text-to-Image Synthesis on Limited Data using Flow-based Diffusion Transformer with MMH Attention
- Rethinking Generative Image Pretraining: How Far Are We From Scaling Up Next-Pixel Prediction?
- StreamKV: Streaming Video Question-Answering with Segment-based KV Cache Retrieval and Compression
- P3-LLM: An Integrated NPU-PIM Accelerator for LLM Inference Using Hybrid Numerical Formats
- Learning to Focus: Focal Attention for Selective and Scalable Transformers
- Diagnose Like A REAL Pathologist: An Uncertainty-Focused Approach for Trustworthy Multi-Resolution Multiple Instance Learning
- Vocabulary In-Context Learning in Transformers: Benefits of Positional Encoding
- EcoSpa: Efficient Transformer Training with Coupled Sparsity
- MambaOVSR: Multiscale Fusion with Global Motion Modeling for Chinese Opera Video Super-Resolution
- VLDrive: Vision-Augmented Lightweight MLLMs for Efficient Language-grounded Autonomous Driving
- Guardian-regularized Safe Offline Reinforcement Learning for Smart Weaning of Mechanical Circulatory Devices
- Make It Long, Keep It Fast: End-to-End 10k-Sequence Modeling at Billion Scale on Douyin
- Next-Latent Prediction Transformers Learn Compact World Models
- BiPETE: A Bi-Positional Embedding Transformer Encoder for Risk Assessment of Alcohol and Substance Use Disorder with Electronic Health Records
- Deep Progressive Training: scaling up depth capacity of zero/one-layer models
- BudgetMem: Learning Selective Memory Policies for Cost-Efficient Long-Context Processing in Language Models
- Motif 2 12.7B technical report
- InfinityStar: Unified Spacetime AutoRegressive Modeling for Visual Generation
- MoSa: Motion Generation with Scalable Autoregressive Modeling
- PETRA: Pretrained Evolutionary Transformer for SARS-CoV-2 Mutation Prediction
- SyMuPe: Affective and Controllable Symbolic Music Performance
- Enhancing composition-based materials property prediction by cross-modal knowledge transfer
- Generative Sequential Recommendation via Hierarchical Behavior Modeling
- PLUTO-4: Frontier Pathology Foundation Models
- Orion-MSP: Multi-Scale Sparse Attention for Tabular In-Context Learning
- Improving DF-Conformer Using Hydra For High-Fidelity Generative Speech Enhancement on Discrete Codec Token
- Using Span Queries to Optimize for Cache and Attention Locality
- KV Cache Transform Coding for Compact Storage in LLM Inference
- On the Emergence of Induction Heads for In-Context Learning
- RefVTON: person-to-person Try on with Additional Unpaired Visual Reference
- OMEGA: Optimized Multimodal Position Encoding Index Derivation with Global Adaptive Scaling for Vision-Language Models
- FlashEVA: Accelerating LLM inference via Efficient Attention
- Tree Training: Accelerating Agentic LLMs Training via Shared Prefix Reuse
- BiSparse-AAS: Bilinear Sparse Attention and Adaptive Spans Framework for Scalable and Efficient Text Summarization
- A Sensing Whole Brain Zebrafish Foundation Model for Neuron Dynamics and Behavior
- Consciousness-ECG Transformer for Conscious State Estimation System with Real-Time Monitoring
- LongCat-Flash-Omni Technical Report
- LLMs Process Lists With General Filter Heads
- Running VLAs at Real-time Speed
- Encoder-Decoder or Decoder-Only? Revisiting Encoder-Decoder Large Language Model
- Emu3.5: Native Multimodal Models are World Learners
- Angular Steering: Behavior Control via Rotation in Activation Space
- Nirvana: A Specialized Generalist Model With Task-Aware Memory Mechanism
- VFXMaster: Unlocking Dynamic Visual Effect Generation via In-Context Learning
- Efficient Vocal Source Separation Through Windowed Sink Attention
- ScaleDiff: Higher-Resolution Image Synthesis via Efficient and Model-Agnostic Diffusion
- Controlling Contrastive Self-Supervised Learning with Knowledge-Driven Multiple Hypothesis: Application to Beat Tracking
- Cache Merging as a Convergent Replicated State for Multi-Agent Latent Reasoning
- ClockRoPE: Random Fourier Rotations for Temporal Routine Modeling
- CineWeaver: Training-Free Reference-Controllable Multi-Shot Long Video Generation for Cinematic Storytelling
- Visko Orbis 1.0: A Live Model for Real-Time Interactive Long Video Generation
- Rad-JEPA 3D: Radiology Joint-Embedding Predictive Model for 3D Computed Tomography
- Data Fusion and Contrastive Alignment for Unconstrained IR Molecular Structure Elucidation
- MetaKoopman: Bayesian Meta-Learning of Koopman Operators for Modeling Structured Dynamics under Distribution Shifts
- RedKnot: Efficient Long-Context LLM Serving with Head-Aware KV Reuse and SegPagedAttention
- Keyless Attention: Value-Space Routing and Value-Only Caching for Efficient Transformers
- Journey Operators for Structured Multi-Axis Composition
- Ripple: Real-Time Streaming Audio-Video Generation With Cross-Modal Recurrent Memory
- Symbol-Equivariant Recurrent Reasoning Models
- Massive Spikes in LLMs are Bias Vectors: Mechanistic Uncovering and Spike-Free Quantization
- The Scaling Properties of Implicit Deductive Reasoning in Transformers
- HRM-Text: Efficient Pretraining Beyond Scaling
- Zero-shot de novo peptide sequencing with open posttranslational modification discovery
- Lattice Deduction Transformers
- Guess-Verify-Refine: Data-Aware Top-K for Sparse-Attention Decoding on Blackwell via Temporal Correlation
- Arcee Trinity Large Technical Report
- Understanding and Optimizing Attention-Based Sparse Matching for Diverse Local Features
- Generative Modeling via Drifting
- Linguistically Informed Evaluation of Multilingual ASR for African Languages
- Ministral 3
- What Drives Success in Physical Planning with Joint-Embedding Predictive World Models?
- BSFA: Leveraging the Subspace Dichotomy to Accelerate Neural Network Training
- Beyond Length: Quantifying Long-Range Information for Long-Context LLM Pretraining Data
- DRIP: Dynamic patch Reduction via Interpretable Pooling
- Uniform Discrete Diffusion with Metric Path for Video Generation
- Group Relative Attention Guidance for Image Editing
- UltraImageGen: Efficient Ultra-High-Resolution Image Generation with Hierarchical Local Attention
- SALS: Sparse Attention in Latent Space for KV cache Compression
- DynaRend: Learning 3D Dynamics via Masked Future Rendering for Robotic Manipulation
- EddyFormer: Accelerated Neural Simulations of Three-Dimensional Turbulence at Scale
- Language-Conditioned Representations and Mixture-of-Experts Policy for Robust Multi-Task Robotic Manipulation
- MeCeFO: Enhancing LLM Training Robustness via Fault-Tolerant Optimization
- Key and Value Weights Are Probably All You Need: On the Necessity of the Query, Key, Value weight Triplet in Decoder-Only Transformers
- LightGlueStick: a Fast and Robust Glue for Joint Point-Line Matching
- BitSkip: An Empirical Analysis of Quantization and Early Exit Composition
- PixelRefer: A Unified Framework for Spatio-Temporal Object Referring with Arbitrary Granularity
- Revisiting Multimodal Positional Encoding in Vision-Language Models
- A Survey on LLM Mid-Training
- Simple Denoising Diffusion Language Models
- Scaling Up Occupancy-centric Driving Scene Generation: Dataset and Method
- Positional Preservation Embedding for Multimodal Large Language Models
- Batch Speculative Decoding Done Right
- SeeDNorm: Self-Rescaled Dynamic Normalization
- LooGLE v2: Are LLMs Ready for Real World Long Dependency Challenges?
- Moving Beyond Diffusion: Hierarchy-to-Hierarchy Autoregression for fMRI-to-Image Reconstruction
- LUNA: Efficient and Topology-Agnostic Foundation Model for EEG Signal Analysis
- LongCat-Video Technical Report
- Streaming Generation for Music Accompaniment
- Mitigating Coordinate Prediction Bias from Positional Encoding Failures
- Agentic Reinforcement Learning for Real-World Code Repair
- OmniVinci: Enhancing Architecture and Data for Omni-Modal Understanding LLM
- Transformer Based Linear Attention with Optimized GPU Kernel Implementation
- BachVid: Training-Free Video Generation with Consistent Background and Character
- StylePitcher: Generating Style-Following and Expressive Pitch Curves for Versatile Singing Tasks
- FlexIO: Flexible Single- and Multi-Channel Speech Separation and Enhancement
- Chronos-2: From Univariate to Universal Forecasting
- Sparser Block-Sparse Attention via Token Permutation
- Smule Renaissance Small: Efficient General-Purpose Vocal Restoration
- Blockwise Flow Matching: Improving Flow Matching Models For Efficient High-Quality Generation
- Attention Sinks in Diffusion Language Models
- Stateful KV Cache Management for LLMs: Balancing Space, Time, Accuracy, and Positional Fidelity
- Video-As-Prompt: Unified Semantic Control for Video Generation
- Alleviating Forgetfulness of Linear Attention by Hybrid Sparse Attention and Contextualized Learnable Token Eviction
- A Scalable, Causal, and Energy Efficient Framework for Neural Decoding with Spiking Neural Networks
- EmbodiedBrain: Expanding Performance Boundaries of Task Planning for Embodied Intelligence
- Positional Encoding Field
- DyPE: Dynamic Position Extrapolation for Ultra High Resolution Diffusion
- The Impact of Negated Text on Hallucination with Large Language Models
- Rotate Both Ways: Time-and-Order RoPE for Generative Recommendation
- Forging GEMs: Advancing Greek NLP through Quality-Based Corpus Curation
- Seed3D 1.0: From Images to High-Fidelity Simulation-Ready 3D Assets
- Memo: Training Memory-Efficient Embodied Agents with Reinforcement Learning
- What is the Best Sequence Length for BABYLM?
- GigaBrain-0: A World Model-Powered Vision-Language-Action Model
- Stream: Scaling up Mechanistic Interpretability to Long Context in LLMs via Sparse Attention
- Loopholing Discrete Diffusion: Deterministic Bypass of the Sampling Wall
- CAGE: Curvature-Aware Gradient Estimation For Accurate Quantization-Aware Training
- Investigating LLM Capabilities on Long Context Comprehension for Medical Question Answering
- EMA-SAM: Exponential Moving-average for SAM-based PTMC Segmentation
- Unifying and Enhancing Graph Transformers via a Hierarchical Mask Framework
- MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training
- MoTVLA: A Vision-Language-Action Model with Unified Fast-Slow Reasoning
- Latent-Augmented Discrete Diffusion Models
- Demystifying Transition Matching: When and Why It Can Beat Flow Matching
- Glyph: Scaling Context Windows via Visual-Text Compression
- ReXMoE: Reusing Experts with Minimal Overhead in Mixture-of-Experts
- StreamingThinker: Large Language Models Can Think While Reading
- JT-Safe: Intrinsically Enhancing the Safety and Trustworthiness of LLMs
- Reasoning Distillation and Structural Alignment for Improved Code Generation
- EventFormer: A Node-graph Hierarchical Attention Transformer for Action-centric Video Event Prediction
- Learning to play: A Multimodal Agent for 3D Game-Play
- MuonBP: Faster Muon via Block-Periodic Orthogonalization
- UniGTE: Unified Graph-Text Encoding for Zero-Shot Generalization across Graph Tasks and Domains
- All You Need is One: Capsule Prompt Tuning with a Single Vector
- Improving Model Representation and Reducing KV Cache via Skip Connections with First Value Heads
- BiMax: Bidirectional MaxSim Score for Document-Level Alignment
- Extending Audio Context for Long-Form Understanding in Large Audio-Language Models
- StretchySnake: Flexible SSM Training Unlocks Action Recognition Across Spatio-Temporal Scales
- Attention Is All You Need for KV Cache in Diffusion LLMs
- Predicting Task Performance with Context-aware Scaling Laws
- From Loop Nests to Silicon: Mapping AI Workloads onto AMD NPUs with MLIR-AIR
- Adapting Self-Supervised Representations as a Latent Space for Efficient Generation
- STANCE: Motion Coherent Video Generation Via Sparse-to-Dense Anchored Encoding
- Understanding the Ability of LLMs to Handle Character-Level Perturbation
- MatchAttention: Matching the Relative Positions for High-Resolution Cross-View Matching
- NOSA: Native and Offloadable Sparse Attention
- True Self-Supervised Novel View Synthesis is Transferable
- Reciprocal Space Attention for Learning Long-Range Interactions
- UniCalli: A Unified Diffusion Framework for Column-Level Generation and Recognition of Chinese Calligraphy
- Shortcutting Pre-trained Flow Matching Diffusion Models is Almost Free Lunch
- Chinese ModernBERT with Whole-Word Masking
- Litespark Technical Report: High-Throughput, Energy-Efficient LLM Training Framework
- KVCOMM: Online Cross-context KV-cache Communication for Efficient LLM-based Multi-agent Systems
- What If : Understanding Motion Through Sparse Interactions
- FlashVSR: Towards Real-Time Diffusion-Based Streaming Video Super-Resolution
- Fine-grained Analysis of Brain-LLM Alignment through Input Attribution
- Chimera: State Space Models Beyond Sequences
- APCE: Adaptive Progressive Context Expansion for Long Context Processing
- SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models
- DiSTAR: Diffusion over a Scalable Token Autoregressive Representation for Speech Generation
- Catch Your Breath: Adaptive Computation for Self-Paced Sequence Production
- Inpainting the Neural Picture: Inferring Unrecorded Brain Area Dynamics from Multi-Animal Datasets
- DiT360: High-Fidelity Panoramic Image Generation via Hybrid Training
- VideoNSA: Native Sparse Attention Scales Video Understanding
- ContextGen: Contextual Layout Anchoring for Identity-Consistent Multi-Instance Generation
- ShishuLM: Lightweight Language Model with Hybrid Decoder-MLP Architecture and Paired Weight Sharing
- High-Fidelity Speech Enhancement via Discrete Audio Tokens
- AdaViewPlanner: Adapting Video Diffusion Models for Viewpoint Planning in 4D Scenes
- Head-wise Adaptive Rotary Positional Encoding for Fine-Grained Image Generation
- SoundReactor: Frame-level Online Video-to-Audio Generation
- ProteinAE: Protein Diffusion Autoencoders for Structure Encoding
- UltraLLaDA: Scaling the Context Length to 128K for Diffusion Large Language Models
- Translution: Unifying Self-attention and Convolution for Adaptive and Relative Modeling
- DreamX-World 1.0: A General-Purpose Interactive World Model
- Accelerating Attention with Basis Decomposition
- DELTA: Dynamic Layer-Aware Token Attention for Efficient Long-Context Reasoning
- Design Principles for Sequence Models via Coefficient Dynamics
- RoDyn: Taming Interactive Robot-Dynamic 2.5D World Model for Robotic Manipulation
- Value-State Gated Attention for Mitigating Extreme-Token Phenomena in Transformers
- DiTSinger: Scaling Singing Voice Synthesis with Diffusion Transformer and Implicit Alignment
- MaP: A Unified Framework for Reliable Evaluation of Pre-training Dynamics
- Efficient Autoregressive Inference for Transformer Probabilistic Models
- KORMo: Korean Open Reasoning Model for Everyone
- Production-Grade Local LLM Inference on Apple Silicon: A Comparative Study of MLX, MLC-LLM, Ollama, llama.cpp, and PyTorch MPS
- Graph Diffusion Transformers are In-Context Molecular Designers
- Scaling Laws for Code: A More Data-Hungry Regime
- Evaluating Small Vision-Language Models on Distance-Dependent Traffic Perception
- UniMMVSR: A Unified Multi-Modal Framework for Cascaded Video Super-Resolution
- VideoNorms: Benchmarking Cultural Awareness of Video Language Models
- SkipSR: Faster Super Resolution with Token Skipping
- VideoCanvas: Unified Video Completion from Arbitrary Spatiotemporal Patches via In-Context Conditioning
- Post-Norm can Resharpen Attention
- D-CoDe: Scaling Image-Pretrained VLMs to Video via Dynamic Compression and Question Decomposition
- When Thoughts Meet Facts: Reusable Reasoning for Long-Context LMs
- Haystack Engineering: Context Engineering for Heterogeneous and Agentic Long-Context Evaluation
- GenPilot: A Multi-Agent System for Test-Time Prompt Optimization in Image Generation
- ReSSFormer: A Recursive Sparse Structured Transformer for Scalable and Long-Context Reasoning
- Efficient numeracy in language models through single-token number embeddings
- JAI-1: A Thai-Centric Large Language Model
- AWM: Accurate Weight-Matrix Fingerprint for Large Language Models
- Reusing Overtrained Language Models Saturates Scaling
- SIGMA-GEN: Structure and Identity Guided Multi-subject Assembly for Image Generation
- ATOM: A Pretrained Neural Operator for Multitask Molecular Dynamics
- ShapeGen4D: Towards High Quality 4D Shape Generation from Videos
- Pack and Force Your Memory: Long-form and Consistent Video Generation
- The Role of Feature Interactions in Graph-based Tabular Deep Learning
- HRTFformer: A Spatially-Aware Transformer for Individual HRTF Upsampling in Immersive Audio Rendering
- Compressed Convolutional Attention: Efficient Attention in a Compressed Latent Space
- MT-DAO: Multi-Timescale Distributed Adaptive Optimizers with Local Updates
- On the Limitations and Capabilities of Position Embeddings for Length Generalization
- Bridging Text and Video Generation: A Survey
- UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models
- ChronoEdit: Towards Temporal Reasoning for Image Editing and World Simulation
- Scaling Sequence-to-Sequence Generative Neural Rendering
- Optimal Scaling Needs Optimal Norm
- Allocation of Parameters in Transformers
- Exploring the Hierarchical Reasoning Model for Small Natural-Image Classification Without Augmentation
- Self-Speculative Masked Diffusions
- Improving GUI Grounding with Explicit Position-to-Coordinate Mapping
- SALSA-V: Shortcut-Augmented Long-form Synchronized Audio from Videos
- Evaluation Framework for Highlight Explanations of Context Utilisation in Language Models
- Coevolutionary Continuous Discrete Diffusion: Make Your Diffusion Language Model a Latent Reasoner
- Audio Driven Real-Time Facial Animation for Social Telepresence
- GeoGraph: Geometric and Graph-based Ensemble Descriptors for Intrinsically Disordered Proteins
- Erased, But Not Forgotten: Erased Rectified Flow Transformers Still Remain Unsafe Under Concept Attack
- Composer: A Search Framework for Hybrid Neural Architecture Design
- Arbitrary Generative Video Interpolation
- SAGE-Music: Low-Latency Symbolic Music Generation via Attribute-Specialized Key-Value Head Sharing
- AReUReDi: Annealed Rectified Updates for Refining Discrete Flows with Multi-Objective Guidance
- Delayed Attention Training Improves Length Generalization in Transformer--RNN Hybrids
- Thoughtbubbles: an Unsupervised Method for Parallel Thinking in Latent Space
- Query-Kontext: An Unified Multimodal Model for Image Generation and Editing
- Source Separation for A Cappella Music
- HilbertA: Hilbert Attention for Image Generation with Diffusion Models
- Refine Drugs, Don't Complete Them: Uniform-Source Discrete Flows for Fragment-Based Drug Discovery
- SeedPrints: Fingerprints Can Even Tell Which Seed Your Large Language Model Was Trained From
- Fading to Grow: Growing Preference Ratios via Preference Fading Discrete Diffusion for Recommendation
- Kairos: Towards Adaptive and Generalizable Time Series Foundation Models
- LaTo: Landmark-tokenized Diffusion Transformer for Fine-grained Human Face Editing
- Boundary-to-Region Supervision for Offline Safe Reinforcement Learning
- Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation
- GenVarFormer: Predicting gene expression from long-range mutations in cancer
- Spontaneous High-Order Generalization in Neural Theory-of-Mind Networks
- Rolling Forcing: Autoregressive Long Video Diffusion in Real Time
- An empirical study on the limitation of Transformers in program trace generation
- Efficient Hyperparameter Tuning via Trajectory Invariance Principle
- LVT: Large-Scale Scene Reconstruction via Local View Transformers
- Double Descent as a Lens for Sample Efficiency in Autoregressive vs. Discrete Diffusion Models
- SemShareKV: Efficient KVCache Sharing for Semantically Similar Prompts via Token-Level LSH Matching
- VSSFlow: Unifying Video-conditioned Sound and Speech Generation via Joint Learning
- Enabling Physical AI through Biological Principles
- DINOReg: Strong Point Cloud Registration with Vision Foundation Model
- UI-UG: A Unified MLLM for UI Understanding and Generation
- Hyperspherical Latents Improve Continuous-Token Autoregressive Generation
- AuON: A Linear-time Alternative to Orthogonal Momentum Updates
- Rethinking JEPA: Compute-Efficient Video SSL with Frozen Teachers
- UniVid: The Open-Source Unified Video Model
- TR2-D2: Tree Search Guided Trajectory-Aware Fine-Tuning for Discrete Diffusion
- Ultra-Fast Language Generation via Discrete Diffusion Divergence Instruct
- Bacterial proteome foundation model enhances functional prediction from enzymes to ecological interactions
- Short window attention enables long-term memorization
- Muon: Training and Trade-offs with Latent Attention and MoE
- Scalable GANs with Transformers
- Training Agents Inside of Scalable World Models
- SANA-Video: Efficient Video Generation with Block Linear Diffusion Transformer
- Pretraining with hierarchical memories: separating long-tail and common knowledge
- Which course? Discourse! Teaching Discourse and Generation in the Era of LLMs
- Disentangling Score Content and Performance Style for Joint Piano Rendering and Transcription
- FraudTransformer: Time-Aware GPT for Transaction Fraud Detection
- Internal Planning in Language Models: Characterizing Horizon and Branch Awareness
- HunyuanImage 3.0 Technical Report
- CORRECT: COndensed eRror RECognition via knowledge Transfer in multi-agent systems
- Beyond Outliers: A Study of Optimizers Under Quantization
- Benchmarking DINOv3 for Multi-Task Stroke Analysis on Non-Contrast CT
- AttAnchor: Guiding Cross-Modal Token Alignment in VLMs with Attention Anchors
- Effective Quantization of Muon Optimizer States
- Look Back to Reason Forward: Revisitable Memory for Long-Context LLM Agents
- GeLoc3r: Enhancing Relative Camera Pose Regression with Geometric Consistency Regularization
- IIET: Efficient Numerical Transformer via Implicit Iterative Euler Method
- Partial Parameter Updates for Efficient Distributed Training
- A model of errors in transformers
- Stochastic activations
- Aurora: Towards Universal Generative Multimodal Time Series Forecasting
- Wavelet-Induced Rotary Encodings: RoPE Meets Graphs
- MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing
- From Bias to Balance: Exploring and Mitigating Spatial Bias in LVLMs
- Syncphony: Synchronized Audio-to-Video Generation with Diffusion Transformers
- DiTraj: training-free trajectory control for video diffusion transformer
- ChaosNexus: A Foundation Model for Universal Chaotic System Forecasting with Multi-scale Representations
- SynerGen: Contextualized Generative Recommender for Unified Search and Recommendation
- Rethinking RoPE Scaling in Quantized LLM: Theory, Outlier, and Channel-Band Analysis with Weight Rescaling
- DeLiVR: Differential Spatiotemporal Lie Bias for Efficient Video Deraining
- Compute-Optimal Quantization-Aware Training
- Unsupervised Speech Enhancement using Data-defined Priors
- Learning Inter-Atomic Potentials without Explicit Equivariance
- X-Streamer: Unified Human World Modeling with Audiovisual Interaction
- Nova: Real-Time Agentic Vision-Language Model Serving with Adaptive Cross-Stage Parallelization
- LayerNorm Induces Recency Bias in Transformer Decoders
- TF-Restormer: Complex Spectral Prediction for Speech Restoration
- Learning to Summarize by Learning to Quiz: Adversarial Agentic Collaboration for Long Document Summarization
- Learning Greens Operators through Hierarchical Neural Networks Inspired by the Fast Multipole Method
- CoSupFormer : A Contrastive Supervised learning approach for EEG signal Classification
- Mamba Modulation: On the Length Generalization of Mamba
- Multimodal Language Models with Modality-Specific Experts for Financial Forecasting from Interleaved Sequences of Text and Time Series
- GyRot: Leveraging Hidden Synergy between Rotation and Fine-grained Group Quantization for Low-bit LLM Inference
- MUL-T: Decoding Spatial Cellular Architecture in Multiplexed Tissue Images
- THGFM: Dual-Branch Temporal Heterogeneous Graph Fusion Model
- Now You Have My Healthy Attention: A U-DiT for Brain-MRI Inpainting
- Edge Prediction for Roof Wireframe Reconstruction with Transformers
- The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence
- SEGA: Spectral-Energy Guided Attention for Resolution Extrapolation in Diffusion Transformers
- ELF: Embedded Language Flows
- How Can We Synthesize High-Quality Pretraining Data? A Systematic Study of Prompt Design, Generator Model, and Source Data
- Self-Supervised Flow Matching for Scalable Multi-Modal Synthesis
- DAGE: Dual-Stream Architecture for Efficient and Fine-Grained Geometry Estimation
- SleepLM: Natural-Language Intelligence for Human Sleep
- ILRe: Intermediate Layer Retrieval for Context Compression in Causal Language Models
- Reading the unreadable: creating a dataset of 19th century English newspapers using image-to-text language models
- PoRe: Position-Reweighted Visual Token Pruning for Vision Language Models
- False Friends Are Not Foes: Investigating Vocabulary Overlap in Multilingual Language Models
- An overview of neural architectures for self-supervised audio representation learning from masked spectrograms
- Reading Images Like Texts: Sequential Image Understanding in Vision-Language Models
- StereoFoley: Object-Aware Stereo Audio Generation from Video
- Turk-LettuceDetect: A Hallucination Detection Models for Turkish RAG Applications
- LAWCAT: Efficient Distillation from Quadratic to Linear Attention with Convolution across Tokens for Long Context Modeling
- Qwen3-Omni Technical Report
- MAST: Multi-Agent Spatial Transformer for Learning to Collaborate
- Scalable Multi Agent Diffusion Policies for Coverage Control
- JCo-MVTON: Jointly Controllable Multi-Modal Diffusion Transformer for Mask-Free Virtual Try-on
- GWM: Towards Scalable Gaussian World Models for Robotic Manipulation
- ViTCAE: ViT-based Class-conditioned Autoencoder
- DISCO: Disentangled Communication Steering for Large Language Models
- Causality-Induced Positional Encoding for Transformer-Based Representation Learning of Non-Sequential Features
- Pico: A Modular Framework for Hypothesis-Driven Small Language Model Research
- ENSAM: an efficient foundation model for interactive segmentation of 3D medical images
- Lynx: Towards High-Fidelity Personalized Video Generation
- Deep Learning Empowered Super-Resolution: A Comprehensive Survey and Future Prospects
- The Rhythm In Anything: Audio-Prompted Drums Generation with Masked Language Modeling
- SolarCrossFormer: Improving day-ahead Solar Irradiance Forecasting by Integrating Satellite Imagery and Ground Sensors
- Language Modeling with Learned Meta-Tokens
- DyWPE: Signal-Aware Dynamic Wavelet Positional Encoding for Time Series Transformers
- Back to Ear: Perceptually Driven High Fidelity Music Reconstruction
- Patent Language Model Pretraining with ModernBERT
- Synthetic bootstrapped pretraining
- Canary-1B-v2 & Parakeet-TDT-0.6B-v3: Efficient and High-Performance Models for Multilingual ASR and AST
- Condition Weaving Meets Expert Modulation: Towards Universal and Controllable Image Generation
- ST-LINK: Spatially-Aware Large Language Models for Spatio-Temporal Forecasting
- Re-purposing SAM into Efficient Visual Projectors for MLLM-Based Referring Image Segmentation
- ProTDyn: a foundation Protein language model for Thermodynamics and Dynamics generation
- DanceEditor: Towards Iterative Editable Music-driven Dance Generation with Open-Vocabulary Descriptions
- PosBridge: Multi-View Positional Embedding Transplant for Identity-Aware Image Editing
- From Next Token Prediction to (STRIPS) World Models -- Preliminary Results
- Multi-Model Synthetic Training for Mission-Critical Small Language Models
- Investigating ReLoRA: Effects on the Learning Dynamics of Small Language Models
- MFAF: An EVA02-Based Multi-scale Frequency Attention Fusion Method for Cross-View Geo-Localization
- Positional Encoding via Token-Aware Phase Attention
- Enhancing Video Large Language Models with Structured Multi-Video Collaborative Reasoning
- LazyDrag: Enabling Stable Drag-Based Editing on Multi-Modal Diffusion Transformers via Explicit Correspondence
- Dynamic Relational Priming Improves Transformer in Multivariate Time Series
- EgoMem: Lifelong Memory Agent for Full-duplex Omnimodal Models
- Preservation of Language Understanding Capabilities in Speech-aware Large Language Models
- Context-Aware Language Models for Forecasting Market Impact from Sequences of Financial News
- SSG-Dit: A Spatial Signal Guided Framework for Controllable Video Generation
- Length-Aware Rotary Position Embedding for Text-Speech Alignment
- Predictable Compression Failures: Why Language Models Actually Hallucinate
- LayerLock: Non-collapsing Representation Learning with Progressive Freezing
- Opening the Black Box: Interpretable LLMs via Semantic Resonance Architecture
- DiFlow-TTS: Discrete Flow Matching with Factorized Speech Tokens for Low-Latency Zero-Shot Text-To-Speech
- CoPE: A Lightweight Complex Positional Encoding
- ENSI: Efficient Non-Interactive Secure Inference for Large Language Models
- DATE: Dynamic Absolute Time Enhancement for Long Video Understanding
- Vejde: A Framework for Inductive Deep Reinforcement Learning Based on Factor Graph Color Refinement
- ViRanker: A BGE-M3 & Blockwise Parallel Transformer Cross-Encoder for Vietnamese Reranking
- When FinTech Meets Privacy: Securing Financial LLMs with Differential Private Fine-Tuning
- Streaming Sequence-to-Sequence Learning with Delayed Streams Modeling
- Customizing the Inductive Biases of Softmax Attention using Structured Matrices
- A Survey of Long-Document Retrieval in the PLM and LLM Era
- LSMTCR: A Scalable Multi-Architecture Model for Epitope-Specific T Cell Receptor de novo Design
- Improving Machine Learning-Based Robot Self-Collision Checking with Input Positional Encoding
- TaDiCodec: Text-aware Diffusion Speech Tokenizer for Speech Language Modeling
- Causal Attention with Lookahead Keys
- NOWJ@COLIEE 2025: A Multi-stage Framework Integrating Embedding Models and Large Language Models for Legal Retrieval and Entailment
- Video Parallel Scaling: Aggregating Diverse Frame Subsets for VideoLLMs
- CausNVS: Autoregressive Multi-view Diffusion for Flexible 3D Novel View Synthesis
- F1: A Vision-Language-Action Model Bridging Understanding and Generation to Actions
- mmBERT: A Modern Multilingual Encoder with Annealed Language Learning
- WindFM: An Open-Source Foundation Model for Zero-Shot Wind Power Forecasting
- Home-made Diffusion Model from Scratch to Hatch
- LatinX: Aligning a Multilingual TTS Model with Direct Preference Optimization
- Crosscoding Through Time: Tracking Emergence & Consolidation Of Linguistic Representations Throughout LLM Pretraining
- HoPE: Hyperbolic Rotary Positional Encoding for Stable Long-Range Dependency Modeling in Large Language Models
- Elucidating the Design Space of Decay in Linear Attention
- COGITAO: A Visual Reasoning Framework To Study Compositionality & Generalization
- SpikingBrain: Spiking Brain-inspired Large Models
- SAC-MIL: Spatial-Aware Correlated Multiple Instance Learning for Histopathology Whole Slide Image Classification
- LatPhon: Lightweight Multilingual G2P for Romance Languages and English
- Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture
- MOSAIC: Multi-Subject Personalized Generation via Correspondence-Aware Alignment and Disentanglement
- LLMs for LLMs: A Structured Prompting Methodology for Long Legal Documents
- Do LLM Modules Generalize? A Study on Motion Generation for Autonomous Driving
- Meta-Pretraining for Zero-Shot Cross-Lingual Named Entity Recognition in Low-Resource Philippine Languages
- Dynamic Sparse Attention on Mobile SoCs
- Flavors of Moonshine: Tiny Specialized ASR Models for Edge Devices
- Kwai Keye-VL 1.5 Technical Report
- CommonKV: Compressing KV Cache with Cross-layer Parameter Sharing
- Towards More Diverse and Challenging Pre-training for Point Cloud Learning: Self-Supervised Cross Reconstruction with Decoupled Views
- LongCat-Flash Technical Report
- Imputing Missing Long-Term Spatiotemporal Multivariate Atmospheric Data with CNN-Transformer Machine Learning
- Entropy-based Coarse and Compressed Semantic Speech Representation Learning
- ParaThinker: Native Parallel Thinking as a New Paradigm to Scale LLM Test-time Compute
- Learn from What We HAVE: History-Aware VErifier that Reasons about Past Interactions Online
- QZhou-Embedding Technical Report
- ECHO: Ego-Centric modeling of Human-Object interactions
- ELV-Halluc: Benchmarking Semantic Aggregation Hallucinations in Long Video Understanding
- WaveLLDM: Design and Development of a Lightweight Latent Diffusion Model for Speech Enhancement and Restoration
- CineScale: Free Lunch in High-Resolution Cinematic Visual Generation
- Mixture of Contexts for Long Video Generation
- Waver: Wave Your Way to Lifelike Video Generation
- StreamMem: Query-Agnostic KV Cache Memory for Streaming Video Understanding
- Provable Benefits of In-Tool Learning for Large Language Models
- Evaluating Compositional Generalisation in VLMs and Diffusion Models
- CAMÕES: A Comprehensive Automatic Speech Recognition Benchmark for European Portuguese
- OmniHuman-1.5: Instilling an Active Mind in Avatars via Cognitive Simulation
- FastMesh: Efficient Artistic Mesh Generation via Component Decoupling
- Beyond flattening: a geometrically principled positional encoding for vision transformers with Weierstrass elliptic functions
- CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis
- Enhancing compact convolutional transformers with super attention
- Optimal Sparsity of Mixture-of-Experts Language Models for Reasoning Tasks
- Integral Transformer: Denoising Attention, Not Too Much Not Too Little
- Position Bias Mitigates Position Bias:Mitigate Position Bias Through Inter-Position Knowledge Distillation
- TOAST: Fast and scalable auto-partitioning based on principled static analysis
- PENGUIN: Enhancing Transformer with Periodic-Nested Group Attention for Long-term Time Series Forecasting
- OmniTry: Virtual Try-On Anything without Masks
- DPad: Efficient Diffusion Language Models with Suffix Dropout
- Maximum Score Routing For Mixture-of-Experts
- 4DNeX: Feed-Forward 4D Generative Modeling Made Easy
- EgoTwin: Dreaming Body and View in First Person
- Matrix-game 2.0: An open-source real-time and streaming interactive world model
- Next Visual Granularity Generation
- TBGRecall: A Generative Retrieval Model for E-commerce Recommendation Scenarios
- Ovis2.5 Technical Report
- ToonComposer: Streamlining Cartoon Production with Generative Post-Keyframing
- NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at Scale
- FuXi-β: Towards a Lightweight and Fast Large-Scale Generative Recommendation Model
- XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization
- Amazon Nova AI Challenge -- Trusted AI: Advancing secure, AI-assisted software development
- VisCodex: Unified Multimodal Code Generation via Merging Vision and Coding Models
- Physical Autoregressive Model for Robotic Manipulation without Action Pretraining
- Learning Spatial Decay for Vision Transformers
- Verify Distributed Deep Learning Model Implementation Refinement with Iterative Relation Inference
- Animate-X++: Universal Character Image Animation with Dynamic Backgrounds
- Shadow in the Cache: Unveiling and Mitigating Privacy Risks of KV-cache in LLM Inference
- RelayFormer: A Unified Local-Global Attention Framework for Scalable Image and Video Manipulation Localization
- OpenCUA: Open Foundations for Computer-Use Agents
- RealisMotion: Decomposed Human Motion Control and Video Generation in the World Space
- GeoVLA: Empowering 3D Representations in Vision-Language-Action Models
- MuGa-VTON: Multi-Garment Virtual Try-On via Diffusion Transformers with Prompt Customization
- DiffractGPT: Atomic Structure Determination from X-ray Diffraction Patterns using Generative Pre-trained Transformer
- Mem4D: Decoupling Static and Dynamic Memory for Dynamic Scene Reconstruction
- LaVieID: Local Autoregressive Diffusion Transformers for Identity-Preserving Video Creation
- Exploring Multimodal Diffusion Transformers for Enhanced Prompt-based Image Editing
- UniFlow: Unifying Speech Front-End Tasks via Continuous Generative Modeling
- Stand-In: A Lightweight and Plug-and-Play Identity Control for Video Generation
- OMGSR: You Only Need One Mid-timestep Guidance for Real-World Image Super-Resolution
- Consistent and Controllable Image Animation with Motion Linear Diffusion Transformers
- AR-GRPO: Training Autoregressive Image Generation Models via Reinforcement Learning
- Many-Turn Jailbreaking
- Whisfusion: Parallel ASR Decoding with Masked Diffusion
- Zero-Shot Cellular Trajectory Map Matching
- gpt-oss-120b & gpt-oss-20b Model Card
- RayDer: Scalable Self-Supervised Novel View Synthesis from Real-World Video
- Multimodal learning with next-token prediction for large multimodal models
- Benchmarking Pretrained Molecular Embedding Models For Molecular Representation Learning
- MeanAudio: Fast and Faithful Text-to-Audio Generation with Mean Flows
- Matrix-Driven Identification and Reconstruction of LLM Weight Homology
- Estimating Musical Surprisal from Audio in Autoregressive Diffusion Model Noise Spaces
- PoseGen: In-Context LoRA Finetuning for Pose-Controllable Long Human Video Generation
- Voost: A Unified and Scalable Diffusion Transformer for Bidirectional Virtual Try-On and Try-Off
- Sculptor: Empowering LLMs with Cognitive Agency via Active Context Management
- Live Music Models
- LayerT2V: Interactive Multi-Object Trajectory Layering for Video Generation
- CodonMoE: DNA Language Models for mRNA Analyses
- MahaTTS: A Unified Framework for Multilingual Text-to-Speech Synthesis
- Agent Lightning: Train ANY AI Agents with Reinforcement Learning
- Hidden Dynamics of Massive Activations in Transformer Training
- READ: Real-time and Efficient Asynchronous Diffusion for Audio-driven Talking Head Generation
- SecoustiCodec: Cross-Modal Aligned Streaming Single-Codecbook Speech Codec
- VITRIX-CLIPIN: Enhancing Fine-Grained Visual Understanding in CLIP via Instruction Editing Data and Long Captions
- LaMPE: Length-aware Multi-grained Positional Encoding for Adaptive Long-context Scaling Without Training
- Learning Dynamics of Meta-Learning in Small Model Pretraining
- What are you sinking? A geometric approach on attention sink
- HGTS-Former: Hierarchical HyperGraph Transformer for Multivariate Time Series Analysis
- A Glimpse to Compress: Dynamic Visual Token Pruning for Large Vision-Language Models
- FluidFormer: Transformer with Continuous Convolution for Particle-based Fluid Simulation
- Frequency-Constrained Learning for Long-Term Forecasting
- Training Dynamics of the Cooldown Stage in Warmup-Stable-Decay Learning Rate Scheduler
- Kronos: A Foundation Model for the Language of Financial Markets
- PiKV: KV Cache Management System for Mixture of Experts
- EdgeInfinite-Instruct: Bridging SFT-Based Optimization and NPU-Level Efficiency for Edge Devices
- Beyond Gloss: A Hand-Centric Framework for Gloss-Free Sign Language Translation
- Unveiling Super Experts in Mixture-of-Experts Large Language Models
- PixNerd: Pixel Neural Field Diffusion
- VMatcher: State-Space Semi-Dense Local Feature Matching
- BS-1-to-N: Diffusion-Based Environment-Aware Cross-BS Channel Knowledge Map Generation for Cell-Free Networks
- On the Expressiveness of Softmax Attention: A Recurrent Neural Network Perspective
- NeedleChain: Measuring Intact Long-Context Reasoning Capability of Large Language Models
- Context-aware Rotary Position Embedding
- X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again
- PanoSplatt3R: Leveraging Perspective Pretraining for Generalized Unposed Wide-Baseline Panorama Reconstruction
- VN-MTEB: Vietnamese Massive Text Embedding Benchmark
Related