FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
2022/05/27 by Tri Dao, Dao, Tri, Daniel Y. Fu +7 · 8 voices · 1,192 citations
Computer Science · #Advanced Neural Network Applications #Algorithm #Artificial intelligence #Computer science #Domain Adaptation and Few-Shot Learning #FLOPS #Language model #Machine Learning and Data Classification #Memory bandwidth #Parallel computing #Perplexity #Speedup #cs.LG
paper · pdf · doi:10.48550/arxiv.2205.14135
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2022/05/27 · openalex created_date 2022/06/13 · arxiv created 2022/06/23 · arxiv updated 2022/06/24 · openalex updated_date 2026/08/05
Abstract
Transformers are slow and memory-hungry on long sequences, since the time and memory complexity of self-attention are quadratic in sequence length. Approximate attention methods have attempted to address this problem by trading off model quality to reduce the compute complexity, but often do not achieve wall-clock speedup. We argue that a missing principle is making attention algorithms IO-aware -- accounting for reads and writes between levels of GPU memory. We propose FlashAttention, an IO-aware exact attention algorithm that uses tiling to reduce the number of memory reads/writes between GPU high bandwidth memory (HBM) and GPU on-chip SRAM. We analyze the IO complexity of FlashAttention, showing that it requires fewer HBM accesses than standard attention, and is optimal for a range of SRAM sizes. We also extend FlashAttention to block-sparse attention, yielding an approximate attention algorithm that is faster than any existing approximate attention method. FlashAttention trains Transformers faster than existing baselines: 15% end-to-end wall-clock speedup on BERT-large (seq. length 512) compared to the MLPerf 1.1 training speed record, 3× speedup on GPT-2 (seq. length 1K), and 2.4× speedup on long-range arena (seq. length 1K-4K). FlashAttention and block-sparse FlashAttention enable longer context in Transformers, yielding higher quality models (0.7 better perplexity on GPT-2 and 6.4 points of lift on long-document classification) and entirely new capabilities: the first Transformers to achieve better-than-chance performance on the Path-X challenge (seq. length 16K, 61.4% accuracy) and Path-256 (seq. length 64K, 63.1% accuracy).
Cited by
- DeCoRAG: Cognitive Decoupling and Semantic-Aware Cropping for Complex Document Understanding
- Moving Beyond Diversity: Visual Token Pruning as Subspace Reconstruction for Efficient VLMs
- SAFE-Pruner: Semantic Attention-Guided Future-Aware Token Pruning for Efficient Vision-Language-Action Manipulation
- RIS-Kernel: A Model-Agnostic Architecture for Long-Context LLM Inference via Sparse Attention
- IQ-JEPA: A Joint-Embedding Predictive Architecture with a Hermitian Vision Transformer for Sound Speed and Attenuation Estimation from Ultrasound IQ Data
- Flash EQ-Linear: Accelerating Equivariant Linear Layers via Group-wise Discrete Fourier Transform
- Eddy-VL 1.9B: Structural Pruning and Layered Distillation for Edge-Deployable Multimodal Embedding
- Trainable Log-linear Sparse Attention for Efficient Diffusion Transformers
- PReM: Learning What to Preserve and When to Refresh for Context Compression
- HijackKV: New Threat in Position-Independent KV Cache Reuse
- HeadCast: Casting Attention Heads for Efficient Autoregressive Video Generation
- Pixel-Space Diffusion Transformers
- ELSAA: Efficient Low-Rank and Sparse Attention Approximation for Training Transformers
- MoA-Structured Decode Attention DNF Derivation, KV-Cache Accumulation, GQA/MQA, and OpenACC Kernel
- Leaky Language Models: Stealing Architecture and Inference Optimizations via Per-Token Timing
- Timeripple: Accelerating vDiTs by Understanding the Spatio-Temporal Correlations in Latent Space
- GQLA: Group-Query Latent Attention for Hardware-Adaptive Large Language Model Decoding
- Matching Ranks Over Probability Yields Truly Deep Safety Alignment
- Harness Engineering for LLM-Driven GPU Kernel Generation
- Patch Policy: Efficient Embodied Control via Dense Visual Representations
- The Capability Convergence Hypothesis: Capability from Access Structure, Not Scale
- Vectorizing the Trie: Efficient Constrained Decoding for LLM-based Generative Retrieval on Accelerators
- Codec-Gauge: Learning Compression-Friendly Gauges for Transformer KV Caches
- Hardware Mechanisms to Dynamically Throttle AI Performance
- Sobek: Streaming Equivariant Tensor Product Convolutions
- WHALE: A Scalable Unified Model for Recommendation with Wukong-HSTU Architecture
- RoVE: Rotary Value Embeddings Attention for Relative Position-dependent Value Pathways
- Do Value Vectors in Deep Layers Need Context from the Residual Stream?
- Points as Tori: Fast Pointwise Signed Distance for Point Clouds
- Constraint-Driven Model Optimization: An Industry Framework for Selecting Compression and Acceleration Techniques in Modern Machine Learning Systems
- Long-Context Fine-Tuning with Limited VRAM
- Spectral Heat Flow for Conservative Token Condensation in Vision-Language Models
- RhinoVLA Technical Report
- DualKV: Shared-Prompt Flash Attention for Efficient RL Training with Large Rollouts and Long Contexts
- DSTAR: Accelerating Diffusion Transformers via Spatial and Temporal Redundancy Reduction
- FVAttn: Adaptive Sparse Attention with Runtime Load Balancing for Video Generation
- An MLIR-Based Compilation Method for Large Language Models
- RePlan: Reasoning-guided Region Planning for Complex Instruction-based Image Editing
- Attention to Detail: Evaluating Energy, Performance, and Accuracy Trade-offs Across vLLM Configurations
- HeteroMosaic: Exposing and Exploiting Heterogeneous Execution Opportunities for Energy-Efficient Edge LLM Inference
- ACID: Adaptive Caching for vIDeo generation
- Bifocal Attention: Harmonizing Geometric and Spectral Positional Embeddings for Algorithmic Generalization
- FlashDecoder: Real-Time Latent-to-Pixel Streaming Decoder with Transformers
- Reflex: Real-Time VLA Control through Streaming Inference
- Native Multi-Dimensional Subquadratic Operators via Input Dependent Long Convolutions
- High-accuracy Low-Bit KV-Cache Quantization via Local Distribution Restoration
- KernelBench-Verified: Do LLM-Generated Kernels Actually Beat PyTorch?
- Efficient and Training-Free Single-Image Diffusion Models
- The Devil is in the Spectrum: Mitigating Representation Collapse in LLMs via Topologically Regularized Side-Path
- Do Transformers Need Three Projections? Systematic Study of QKV Variants
- InferenceBench: A Benchmark for Open-Ended LLM Inference Optimization by AI Agents
- Multi-Stream LLMs: Unblocking Language Models with Parallel Streams of Thoughts, Inputs and Outputs
- JAXBench: Benchmarking Autonomous TPU Kernel Optimization
- Dropping the Anchor: Statistical Context Summarization for Distributed Systems via Pulsar Attention
- Attention to Mamba: A Recipe for Cross-Architecture Distillation
- PoM: A Linear-Time Replacement for Attention with the Polynomial Mixer
- M2RNN: Non-Linear RNNs with Matrix-Valued States for Scalable Language Modeling
- Flash-KMeans: Fast and Memory-Efficient Exact K-Means
- AutoKernel: Autonomous GPU Kernel Optimization via Iterative Agent-Driven Search
- Speculative Speculative Decoding
- Deep models of protein evolution in time generate realistic evolutionary trajectories and functional proteins
- Do LLMs Benefit From Their Own Words?
- Spelling Bee Embeddings for Language Modeling
- Even GPT-5.2 Can't Count to Five: The Case for Zero-Error Horizons in Trustworthy LLMs
- mHC: Manifold-Constrained Hyper-Connections
- Dynamic Large Concept Models: Latent Reasoning in an Adaptive Semantic Space
- PHOTON: Hierarchical Autoregressive Modeling for Lightspeed and Memory-Efficient Language Generation
- Extending the Context of Pretrained LLMs by Dropping Their Positional Embeddings
- Jasmine: A Simple, Performant and Scalable JAX-based World Modeling Codebase
- Kimi Linear: An Expressive, Efficient Attention Architecture
- Expected Attention: KV Cache Compression by Estimating Attention from Future Queries Distribution
- Fast and Simplex: 2-Simplicial Attention in Triton
- JAFAR: Jack up Any Feature at Any Resolution
- Protriever: End-to-End Differentiable Protein Homology Search for Fitness Prediction
- QiMeng: Fully Automated Hardware and Software Design for Processor Chip
- Leveraging Natural Language Processing to Unravel the Mystery of Life: A Review of NLP Approaches in Genomics, Transcriptomics, and Proteomics
- AbsenceBench: Language Models Can't Tell What's Missing
- Log-Linear Attention
- Flash Invariant Point Attention
- ModernBERT or DeBERTaV3? Examining Architecture and Data Influence on Transformer Encoder Models Performance
- 70% Size, 100% Accuracy: Lossless LLM Compression for Efficient GPU Inference via Dynamic-Length Float (DFloat11)
- TransMamba: A Sequence-Level Hybrid Transformer-Mamba Language Model
- ZSMerge: Zero-Shot KV Cache Compression for Memory-Efficient Long-Context LLMs
- SPECTRE: An FFT-Based Efficient Drop-In Replacement to Self-Attention for Long Contexts
- InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU
- Extending the RANGE of Graph Neural Networks: Relaying Attention Nodes for Global Encoding
- Block Diffusion: Interpolating Between Autoregressive and Diffusion Language Models
- Scaling Laws in Patchification: An Image Is Worth 50,176 Tokens And More
- Fast3R: Towards 3D Reconstruction of 1000+ Images in One Forward Pass
- Accelerating Time Series Foundation Models with Speculative Decoding
- MXAttention: Data-Free Optimal Scaling and Pre-Normalization Quantization for MXFP4 Attention
- AKG kernel Agent: A Multi-Agent Framework for Cross-Platform Kernel Synthesis
- Deep learning for pedestrians: backpropagation in Transformers
- Trust Region Masking for Long-Horizon LLM Reinforcement Learning
- Breaking the Memory Wall: Exact Analytical Differentiation via Tiled Operator-Space Evolution
- Viability and Performance of a Private LLM Server for SMBs: A Benchmark Analysis of Qwen3-30B on Consumer-Grade Hardware
- Fleet: Hierarchical Task-based Abstraction for Megakernels on Multi-Die GPUs
- Latent Multi-Head Attention for Small Language Models
- Chessformer: A Unified Architecture for Chess Modeling
- Bridging Global Intent with Local Details: A Hierarchical Representation Approach for Semantic Validation in Text-to-SQL
- WeDLM: Reconciling Diffusion Language Models with Standard Causal Attention for Fast Inference
- Learning When Not to Attend Globally
- Role-Based Fault Tolerance System for LLM RL Post-Training
- DiRL: An Efficient Post-Training Framework for Diffusion Language Models
- MatKV: Trading Compute for Flash Storage in LLM Inference
- DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention
- RayRoPE: Projective Ray Positional Encoding for Multi-view Attention
- Compressing Observation History into Agent Memory: Distilling Transformers into Recurrent Transformers
- Omni-Prune: Query-Aware Unified Token Pruning for Efficient Omnimodal Large Language Models
- MobileWan: Closing the Quality Gap for Mobile Video Diffusion
- LiteTopK: Exploiting the Curse of Dimensionality for a Fused Indexer-TopK Kernel in Long-Context Sparse Attention
- Express Language Modeling
- StreamIndex: Memory-Bounded Compressed Sparse Attention via Streaming Top-k
- Full event interpretation with machine-learning-based particle-flow reconstruction in the CMS detector
- Accelerating Language Model Workflows with Prompt Choreography
- Anchor Attention, Small Cache: Code Generation with Large Language Models
- Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification
- KAP: Bridging the Knowledge Selection-Runtime Consumption Gap in LLM Systems
- X-Stage: An Overlooked Pipeline Stage for Communication-Computation Overlap in DiT Inference
- WHTMix: Efficient Stereo Depth Estimation via Walsh-Hadamard Token Mixing
- Salient Knowledge Pathways: Sparse Cross-Modal Routing for Efficient Knowledge-Intensive Multimodal Question Answering
- CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention
- Intrinsic and Triangulation-Agnostic Attention: A Simple and Powerful Approach for Learning on Meshes
- FusionML: Prefill, Not Decode - Mechanism and Boundaries of CPU+GPU Co-Execution on Unified-Memory Apple Silicon
- SeedFold: Scaling Biomolecular Structure Prediction
- Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs
- Agent Team Work Zone: An Automated, Persistent Workspace for Long-Lived Claude Code Agent Teams
- Visible-Light Imaging Diagnosis of Neutral Particle Emission Tomography in the Tokamak Divertor: An Efficient Transformer-based Surrogate Model
- Bumblebee: Interleaved Mixed-Layer Building Blocks for Large-Scale Recommendation Systems
- xMIx: High-Performance Serving-Time Platform for Mechanistic Interpretability Apps
- Kernel Forge: An Agent Harness for LLM-based Generation and Optimization of CUDA Kernels
- FlashEvaluator: Expanding Search Space with Parallel Sequence-Level Evaluation
- SkipOPU: An FPGA-based Overlay Processor for Large Language Models with Dynamically Allocated Computation
- Plain Transformers are Surprisingly Powerful Link Predictors
- DPAR: Dynamic Patchification for Efficient Autoregressive Visual Generation
- EasyOmnimatte: Taming Pretrained Inpainting Diffusion Models for End-to-End Video Layered Decomposition
- Scene-VLM: Multimodal Video Scene Segmentation via Vision-Language Models
- Analyzing the Mechanism of Attention Collapse in VGGT from a Dynamics Perspective
- Mesh-Attention: A New Communication-Efficient Distributed Attention with Improved Data Locality
- FlashVLM: Text-Guided Visual Token Selection for Large Multimodal Models
- Designing Spatial Architectures for Sparse Attention: STAR Accelerator via Cross-Stage Tiling
- ActionFlow: A Pipelined Action Acceleration for Vision Language Models on Edge
- Predictive-LoRA: A Proactive and Fragmentation-Aware Serverless Inference System for LLMs
- Memory as Resonance: A Biomimetic Architecture for Infinite Context Memory on Ergodic Phonetic Manifolds
- RAPID-LLM: Resilience-Aware Performance analysis of Infrastructure for Distributed LLM Training and Inference
- CascadeInfer: Length-Aware Scheduling of LLM Serving with Low Latency and Load Balancing
- In-Context Audio Control of Video Diffusion Transformers
- A Multi-agent Text2SQL Framework using Small Language Models and Execution Feedback
- Optimal Software Pipelining and Warp Specialization for Tensor Core GPUs
- KV Admission: Learning What to Write for Efficient Long-Context Inference
- LoPA: Scaling dLLM Inference via Lookahead Parallel Decoding
- MRG-R1: Reinforcement Learning for Clinically Aligned Medical Report Generation
- INTELLECT-3: Technical Report
- From Theory to Throughput: CUDA-Optimized APML for Large-Batch 3D Learning
- Dynamic Rebatching for Efficient Early-Exit Inference with DREX
- End-to-End Training for Autoregressive Video Diffusion via Self-Resampling
- LLMQ: Efficient Lower-Precision Pretraining for Consumer GPUs
- Dual-Density Inference for Efficient Language Model Reasoning
- Mixture of Attention Schemes (MoAS): Learning to Route Between MHA, GQA, and MQA
- Towards Nepali-language LLMs: Efficient GPT training with a Nepali BPE tokenizer
- TEMP: A Memory Efficient Physical-aware Tensor Partition-Mapping Framework on Wafer-scale Chips
- Ladder Up, Memory Down: Low-Cost Fine-Tuning With Side Nets
- BLURR: A Boosted Low-Resource Inference for Vision-Language-Action Models
- A Unified Sparse Attention via Multi-Granularity Compression
- VajraV1 -- The most accurate Real Time Object Detector of the YOLO family
- LitePT: Lighter Yet Stronger Point Transformer
- Improving Recursive Transformers with Mixture of LoRAs
- Soul: Breathe Life into Digital Human for High-fidelity Long-term Multimodal Animation
- Transform Trained Transformer: Accelerating Naive 4K Video Generation Over 10×
- NRR-Core: Non-Resolution Reasoning as a Computational Framework for Contextual Identity and Ambiguity Preservation
- SneakPeek: Future-Guided Instructional Streaming Video Generation
- Investigating Data Pruning for Pretraining Biological Foundation Models at Scale
- ProServe: Unified Multi-Priority Request Scheduling for LLM Serving
- Cross-Modal Representational Knowledge Distillation for Enhanced Spike-Informed LFP Modeling
- Exact Flow Linear Attention: Exact Solution from Continuous-Time Dynamics
- StreamingAssistant: Efficient Visual Token Pruning for Accelerating Online Video Understanding
- TA-KAND: Two-stage Attention Triple Enhancement and U-KAN based Diffusion For Few-shot Knowledge Graph Completion
- BOOST: BOttleneck-Optimized Scalable Training Framework for Low-Rank Large Language Models
- BLASST: Dynamic BLocked Attention Sparsity via Softmax Thresholding
- xGR: Efficient Generative Recommendation Serving at Scale
- Adaptive Soft Rolling KV Freeze with Entropy-Guided Recovery: Sublinear Memory Growth for Efficient LLM Inference
- Mining Legal Arguments to Study Judicial Formalism
- Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap
- FUSER: Feed-Forward MUltiview 3D Registration Transformer and SE(3)N Diffusion Refinement
- RACAM: Enhancing DRAM with Reuse-Aware Computation and Automated Mapping for ML Inference
- TritonForge: Profiling-Guided Framework for Automated Triton Kernel Optimization
- Should AI Become an Intergenerational Civil Right?
- LaMoSys3.5D: Enabling 3.5D-IC-Based Large Language Model Inference Serving Systems via Hardware/Software Co-Design
- Towards Lossless Ultimate Vision Token Compression for VLMs
- Chopper: A Multi-Level GPU Characterization Tool & Derived Insights Into LLM Training Inefficiency
- HybridToken-VLM: Hybrid Token Compression for Vision-Language Models
- MobileFineTuner: A Mobile-Native Framework for On-Device LLM Fine-Tuning in Real-World Embedded AI Applications
- GatedFWA: Linear Flash Windowed Attention with Gated Associative Memory
- A Mathematical Theory of Top-k Sparse Attention via Total Variation Distance
- Flash Multi-Head Feed-Forward Network
- Persian-Phi: Efficient Cross-Lingual Adaptation of Compact LLMs via Curriculum Learning
- DCO: Dynamic Cache Orchestration for LLM Accelerators through Predictive Management
- Materium: An Autoregressive Approach for Material Generation
- Block Sparse Flash Attention
- RLAX: Large-Scale, Distributed Reinforcement Learning for Large Language Models on TPUs
- ReCAD: Reinforcement Learning Enhanced Parametric CAD Model Generation with Vision-Language Models
- Hierarchical geometric deep learning enables scalable analysis of molecular dynamics
- EgoEdit: Dataset, Real-Time Streaming Model, and Benchmark for Egocentric Video Editing
- ShaRP: SHAllow-LayeR Pruning for Video Large Language Models Acceleration
- StreamFlow: Theory, Algorithm, and Implementation for High-Efficiency Rectified Flow Generation
- Parabolic Position Encoding: Vision-Centric, Principled, Extrapolatable, General
- State Space Models for Bioacoustics: A Comparative Evaluation with Transformers
- Reconstructing KV Caches with Cross-layer Fusion For Enhanced Transformers
- A Preliminary Study on the Promises and Challenges of Native Top-k Sparse Attention
- Agentic Operator Generation for ML ASICs
- PSA: Pyramid Sparse Attention for Efficient Video Understanding and Generation
- AutoBrep: Autoregressive B-Rep Generation with Unified Topology and Geometry
- GeoZero: Incentivizing Reasoning from Scratch on Geospatial Scenes
- Fairy2i: Training Complex LLMs from Real LLMs with All Parameters in \± 1, ± i\
- When Refusals Fail: Unstable Safety Mechanisms in Long-Context LLM Agents
- KV Pareto: Systems-Level Optimization of KV Cache and Model Compression for Long Context Inference
- SVRG and Beyond via Posterior Correction
- A Systematic Characterization of LLM Inference on GPUs
- FlashVGGT: Efficient and Scalable Visual Geometry Transformers with Compressed Descriptor Attention
- VLASH: Real-Time VLAs via Future-State-Aware Asynchronous Inference
- Accelerating Streaming Video Large Language Models via Hierarchical Token Compression
- Feed-Forward 3D Gaussian Splatting Compression with Long-Context Modeling
- SIMPLE: Disaggregating Sampling from GPU Inference into a Decision Plane for Faster Distributed LLM Serving
- InvarDiff: Cross-Scale Invariance Caching for Accelerated Diffusion Models
- Serving Heterogeneous LoRA Adapters in Distributed LLM Inference Systems
- Energy-Efficient Vision Transformer Inference for Edge-AI Deployment
- GSPN-2: Efficient Parallel Sequence Modeling
- Ovis-Image Technical Report
- Behavior-Equivalent Token: Single-Token Replacement for Long Prompts in LLMs
- InstanceV: Instance-Level Video Generation
- OOCO: Latency-disaggregated Architecture for Online-Offline Co-locate LLM Serving
- PIBNet: a Physics-Inspired Boundary Network for Multiple Scattering Simulations
- IntAttention: A Fully Integer Attention Pipeline for Efficient Edge Inference
- Frequency-Aware Token Reduction for Efficient Vision Transformer
- Subjective Depth and Timescale Transformers: Learning Where and When to Compute
- MeanFlow Transformers with Representation Autoencoders
- Generating Separated Singing Vocals Using a Diffusion Model Conditioned on Music Mixtures
- DOPD: A Dynamic PD-Disaggregation Architecture for Maximizing Goodput in LLM Inference Serving
- LocateAnything3D: Vision-Language 3D Detection with Chain-of-Sight
- Adam Simplified: Bias Correction Debunked
- FREE: Uncertainty-Aware Autoregression for Parallel Diffusion Transformers
- Scaling LLM Speculative Decoding: Non-Autoregressive Forecasting in Large-Batch Scenarios
- UltraViCo: Breaking Extrapolation Limits in Video Diffusion Transformers
- QiMeng-Kernel: Macro-Thinking Micro-Coding Paradigm for LLM-Based High-Performance GPU Kernel Generation
- HKRAG: Holistic Knowledge Retrieval-Augmented Generation Over Visually-Rich Documents
- Tell Model Where to Look: Mitigating Hallucinations in MLLMs by Vision-Guided Attention
- Terminal Velocity Matching
- MapFormer: Self-Supervised Learning of Cognitive Maps with Input-Dependent Positional Embeddings
- EventSTU: Event-Guided Efficient Spatio-Temporal Understanding for Video Large Language Models
- Nemotron-Flash: Towards Latency-Optimal Hybrid Small Language Models
- Parallel Vision Token Scheduling for Fast and Accurate Multimodal LMMs Inference
- VLM in a flash: I/O-Efficient Sparsification of Vision-Language Model via Neuron Chunking
- NEZHA: A Zero-sacrifice and Hyperspeed Decoding Architecture for Generative Recommendations
- Kitty: Accurate and Efficient 2-bit KV Cache Quantization with Dynamic Channel-wise Precision Boost
- Deterministic Inference across Tensor Parallel Sizes That Eliminates Training-Inference Mismatch
- Selective Rotary Position Embedding
- DeepCoT: Deep Continual Transformers for Real-Time Inference on Data Streams
- MUCH: A Multilingual Claim Hallucination Benchmark
- Energy Scaling Laws for Diffusion Models: Quantifying Compute and Carbon Emissions in Image Generation
- Optimizing PyTorch Inference with LLM-Based Multi-Agent Systems
- Revisiting Multimodal KV Cache Compression: A Frequency-Domain-Guided Outlier-KV-Aware Approach
- Neo: Real-Time On-Device 3D Gaussian Splatting with Reuse-and-Update Sorting Acceleration
- Hemlet: A Heterogeneous Compute-in-Memory Chiplet Architecture for Vision Transformers with Group-Level Parallelism
- Reasoning in Diffusion Large Language Models is Concentrated in Dynamic Confusion Zones
- A Comprehensive Study on Visual Token Redundancy for Discrete Diffusion-based Multimodal Large Language Models
- OmniZip: Audio-Guided Dynamic Token Compression for Fast Omnimodal Large Language Models
- From Projection to Prediction: Beyond Logits for Scalable Language Models
- MergeDNA: Context-aware Genome Modeling with Dynamic Tokenization through Token Merging
- TabFlash: Efficient Table Understanding with Progressive Question Conditioning and Token Focusing
- MACKO: Sparse Matrix-Vector Multiplication for Low Sparsity
- EchoLSTM: A Self-Reflective Recurrent Network for Stabilizing Long-Range Memory
- Equivalence Checking of ML GPU Kernels
- Evaluating Latent Generative Paradigms for High-Fidelity 3D Shape Completion from a Single Depth Image
- Striking the Right Balance between Compute and Copy: Improving LLM Inferencing Under Speculative Decoding
- Arcee: Differentiable Recurrent State Chain for Generative Vision Modeling with Mamba SSMs
- DocSLM: A Small Vision-Language Model for Long Multimodal Document Understanding
- LEMUR: Large scale End-to-end MUltimodal Recommendation
- Seeing the Forest and the Trees: Query-Aware Tokenizer for Long-Video Multimodal Language Models
- Optimizing Mixture of Block Attention
- Faster Algorithms for Structured Matrix Multiplication via Flip Graph Search
- Speculative Decoding in Decentralized LLM Inference: Turning Communication Latency into Computation Throughput
- Retrieval as a Decision: Training-Free Adaptive Gating for Efficient RAG
- Making Every Head Count: Sparse Attention Without the Speed-Performance Trade-off
- Branching Flows: Discrete, Continuous, and Manifold Flow Matching with Splits and Deletions
- Think-at-Hard: Selective Latent Iterations to Improve Reasoning Language Models
- Moebius: 0.2B Lightweight Image Inpainting Framework with 10B-Level Performance
- DPRM: A Dual Implicit Process Reward Model in Multi-Hop Question Answering
- Generalizable Insights for Graph Transformers in Theory and Practice
- Sharp Eyes and Memory for VideoLLMs: Information-Aware Visual Token Pruning for Efficient and Reliable VideoLLM Reasoning
- Taking the Road Less Scheduled with Adaptive Polyak Steps
- A Circular Argument : Does RoPE need to be Equivariant for Vision?
- StreamDiffusionV2: A Streaming System for Dynamic and Interactive Video Generation
- TNT: Improving Chunkwise Training for Test-Time Memorization
- Learning to Focus: Focal Attention for Selective and Scalable Transformers
- You Had One Job: Per-Task Quantization Using LLMs' Hidden Representations
- PRAGMA: A Profiling-Reasoned Multi-Agent Framework for Automatic Kernel Optimization
- We Can Hear You with mmWave Radar! An End-to-End Eavesdropping System
- Optimizing Long-context LLM Serving via Fine-grained Sequence Parallelism
- MoSKA: Mixture of Shared KV Attention for Efficient Long-Sequence LLM Inference
- Make It Long, Keep It Fast: End-to-End 10K Long User Behavior Sequence Modeling for Billion-Scale Douyin Recommendation
- Lethe: Layer- and Time-Adaptive KV Cache Pruning for Reasoning-Intensive LLM Serving
- Hilbert-Guided Sparse Local Attention
- OckBench: Measuring the Efficiency of LLM Reasoning
- Rethinking Metrics and Diffusion Architecture for 3D Point Cloud Generation
- LoPT: Lossless Parallel Tokenization Acceleration for Long Context Inference of Large Language Model
- BudgetMem: Learning Selective Memory Policies for Cost-Efficient Long-Context Processing in Language Models
- DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing
- FuseFlow: A Fusion-Centric Compilation Framework for Sparse Deep Learning on Streaming Dataflow
- Cambrian-S: Towards Spatial Supersensing in Video
- PICNIC: Silicon Photonic Interconnected Chiplets with Computational Network and In-memory Computing for LLM Inference Acceleration
- PETRA: Pretrained Evolutionary Transformer for SARS-CoV-2 Mutation Prediction
- Temporal Zoom Networks: Distance Regression and Continuous Depth for Efficient Action Localization
- NAPS: Attention-Based Fusion of Heterogeneous Physiological Signals
- SnapStream: Efficient Long Sequence Decoding on Dataflow Accelerators
- Divide, Cache, Conquer: Dichotomic Prompting for Efficient Multi-Label LLM-Based Classification
- From Models to Operators: Rethinking Autoscaling Granularity for Large Generative Models
- Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live
- Optimizing Attention on GPUs by Exploiting GPU Architectural NUMA Effects
- Flashlight: PyTorch Compiler Extensions to Accelerate Attention Variants
- Assessing LLM Reasoning Steps via Principal Knowledge Grounding
- GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding
- FlashEVA: Accelerating LLM inference via Efficient Attention
- Reasoning Planning for Language Models
- SpecAttn: Speculating Sparse Attention
- H-FA: A Hybrid Floating-Point and Logarithmic Approach to Hardware Accelerated FlashAttention
- Running VLAs at Real-time Speed
- OneTrans: Unified Feature Interaction and Sequence Modeling with One Transformer in Industrial Recommender
- Gaperon: A Peppered English-French Generative Language Model Suite
- PureKV: Plug-and-Play KV Cache Optimization with Spatial-Temporal Sparse Attention for Vision-Language Large Models
- LLMET: Enabling Cross-Layer Evaluation of Emerging M3D Memories for Energy-Efficient LLM Serving
- PSG: Pair-Space Generation for Efficient Generative Reranking
- NELSSA: A GPU-PNM Heterogeneous System for Mixed-Length LLM Serving via Length-based Request Placement
- From Tokens to Watt-hours: Analytical Energy Estimation for LLM Inference on Modern GPUs
- Steering Instruction Hierarchies at Inference Time
- Mixture-of-Depths Attention
- Don't Just Chase "Highlighted Tokens" in MLLMs: Revisiting Visual Holistic Context Retention
- A Compositional Theory of Causally Masked Transformers
- CuTe Layout Representation and Algebra
- ThriftAttention: Selective Mixed Precision for Long-Context FP4 Attention
- Zero-shot de novo peptide sequencing with open posttranslational modification discovery
- Simplified Sparse Attention via Gist Tokens
- VisionSelector: End-to-End Learnable Visual Token Compression for Efficient Multimodal LLMs
- Fine-Tuning GPT-5 for GPU Kernel Generation
- HySparse: A Hybrid Sparse Attention Architecture with Oracle Token Selection and KV Cache Sharing
- Group Relative Attention Guidance for Image Editing
- Improving LLM Reasoning via Dependency-Aware Query Decomposition and Logic-Parallel Content Expansion
- Long-Context Modeling with Dynamic Hierarchical Sparse Attention for On-Device LLMs
- UltraImageGen: Efficient Ultra-High-Resolution Image Generation with Hierarchical Local Attention
- APTBench: Benchmarking Agentic Potential of Base LLMs During Pre-Training
- ETC: training-free diffusion models acceleration with Error-aware Trend Consistency
- Beyond Line-Level Filtering for the Pretraining Corpora of LLMs
- Pie: A Programmable Serving System for Emerging LLM Applications
- Key and Value Weights Are Probably All You Need: On the Necessity of the Query, Key, Value weight Triplet in Self-Attention Transformers
- A Survey on Efficient Vision-Language-Action Models
- Knocking-Heads Attention
- Can Language Models Compose Skills In-Context?
- Rethinking Inference Placement for Deep Learning across Edge and Cloud Platforms: A Multi-Objective Optimization Perspective and Future Directions
- Massive Memorization with Hundreds of Trillions of Parameters for Sequential Transducer Generative Recommenders
- Beyond Reasoning Gains: Mitigating General Capabilities Forgetting in Large Reasoning Models
- Transformer Based Linear Attention with Optimized GPU Kernel Implementation
- A Unified Model for Multi-Task Drone Routing in Post-Disaster Road Assessment
- Sparser Block-Sparse Attention via Token Permutation
- Correlation Dimension of Auto-Regressive Large Language Models
- Mitra: Mixed Synthetic Priors for Enhancing Tabular Foundation Models
- CudaForge: An Agent Framework with Hardware Feedback for CUDA Kernel Optimization
- REx86: A Local Large Language Model for Assisting in x86 Assembly Reverse Engineering
- HoloCine: Holistic Generation of Cinematic Multi-Shot Long Video Narratives
- Alleviating Forgetfulness of Linear Attention by Hybrid Sparse Attention and Contextualized Learnable Token Eviction
- HA-RAG: Hotness-Aware RAG Acceleration via Mixed Precision and Data Placement
- PPMStereo: Pick-and-Play Memory Construction for Consistent Dynamic Stereo Matching
- Rotate Both Ways: Time-and-Order RoPE for Generative Recommendation
- Seed3D 1.0: From Images to High-Fidelity Simulation-Ready 3D Assets
- A Survey on Cache Methods in Diffusion Models: Toward Efficient Multi-Modal Generation
- Stream: Scaling up Mechanistic Interpretability to Long Context in LLMs via Sparse Attention
- Every Attention Matters: An Efficient Hybrid Architecture for Long-Context Reasoning
- Reasoning Language Model Inference Serving Unveiled: An Empirical Study
- EdgeReasoning: Characterizing Reasoning LLM Deployment on Edge GPUs
- ssToken: Self-modulated and Semantic-aware Token Selection for LLM Fine-tuning
- Unifying and Enhancing Graph Transformers via a Hierarchical Mask Framework
- MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training
- Accelerating Vision Transformers with Adaptive Patch Sizes
- Unbiased Gradient Low-Rank Projection
- This is Going to Sound Crazy, But What If We Used Large Language Models to Boost Automatic Database Tuning Algorithms By Leveraging Prior History? We Will Find Better Configurations More Quickly Than Retraining From Scratch!
- Enhanced Motion Forecasting with Plug-and-Play Multimodal Large Language Models
- Efficient Long-context Language Model Training by Core Attention Disaggregation
- Efficient Vision-Language-Action Models for Embodied Manipulation: A Systematic Survey
- Finding Manifolds With Bilinear Autoencoders
- All You Need is One: Capsule Prompt Tuning with a Single Vector
- Improving Model Representation and Reducing KV Cache via Skip Connections with First Value Heads
- Predicting Task Performance with Context-aware Scaling Laws
- To Infinity and Beyond: Tool-Use Unlocks Length Generalization in State Space Models
- xLLM Technical Report
- First Attentions Last: Better Exploiting First Attentions for Efficient Transformer Training
- State-Space Models for Tabular Prior-Data Fitted Networks
- End-to-End Multi-Modal Diffusion Mamba
- BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Disaggregated LLM Serving in AI Infrastructure
- GatePro: Parameter-Free Expert Selection Optimization for Mixture-of-Experts Models
- Taming the Fragility of KV Cache Eviction in LLM Inference
- Trace Anything: Representing Any Video in 4D via Trajectory Fields
- Attention Illuminates LLM Reasoning: The Preplan-and-Anchor Rhythm Enables Fine-Grained Policy Optimization
- Chinese ModernBERT with Whole-Word Masking
- Litespark Technical Report: High-Throughput, Energy-Efficient LLM Training Framework
- FlashVSR: Towards Real-Time Diffusion-Based Streaming Video Super-Resolution
- APCE: Adaptive Progressive Context Expansion for Long Context Processing
- PAINT: Parallel-in-time Neural Twins for Dynamical System Reconstruction
- Task-Aware Reduction for Scalable LLM-Database Systems
- HoMer: Addressing Heterogeneities by Modeling Sequential and Set-wise Contexts for CTR Prediction
- Protenix-Mini+: efficient structure prediction model with scalable pairformer
- KVComm: Enabling Efficient LLM Communication through Selective KV Sharing
- DCP: Addressing Input Dynamism In Long-Context Training via Dynamic Context Parallelism
- PVDetector: Detecting Prompt Injection Attacks on Purpose-Specific LLM Agents through Policy-Violation Concept Analysis
- GPU-Tile-Sim: A Tile-Centric GPU Simulation Framework for LLM Hardware-Software Co-Design
- AutoMegaKernel: A Statically-Checked Agent Harness for Self-Retargeting Megakernel Synthesis
- AURA: Action-Gated Memory for Robot Policies at Constant VRAM
- Sparse Query Attention (SQA): A Computationally Efficient Attention Mechanism with Query Heads Reduction
- Attention Once Is All You Need: Efficient Streaming Inference with Stateful Transformers
- Accelerating Attention with Basis Decomposition
- On the Quantization Robustness of Diffusion Language Models in Coding Benchmarks
- SOL-ExecBench: Speed-of-Light Benchmarking for Real-World GPU Kernels Against Hardware Limits
- An Efficient Heterogeneous Co-Design for Fine-Tuning on a Single GPU
- PolyBlocks: A Compiler Infrastructure for AI Chips and Programming Frameworks
- Token Is All You Price
- MRMR: A Realistic and Expert-Level Multidisciplinary Benchmark for Reasoning-Intensive Multimodal Retrieval
- KORMo: Korean Open Reasoning Model for Everyone
- Neptune: Advanced ML Operator Fusion for Locality and Parallelism on GPUs
- SViM3D: Stable Video Material Diffusion for Single Image 3D Generation
- SPAD: Specialized Prefill and Decode Hardware for Disaggregated LLM Inference
- Beyond Sunk Costs: Boosting LLM Pre-training Efficiency via Orthogonal Growth of Mixture-of-Experts
- From Tokens to Layers: Redefining Stall-Free Scheduling for MoE Serving with Layered Prefill
- Haystack Engineering: Context Engineering for Heterogeneous and Agentic Long-Context Evaluation
- HTMformer: Hybrid Time and Multivariate Transformer for Time Series Forecasting
- Syncopate: Efficient Multi-GPU AI Kernels via Automatic Chunk-Centric Compute-Communication Overlap
- Executable Counterfactuals: Improving LLMs' Causal Reasoning Through Code
- Vectorized FlashAttention with Low-cost Exponential Computation in RISC-V Vector Processors
- Reusing Overtrained Language Models Saturates Scaling
- ConCuR: Conciseness Makes State-of-the-Art Kernel Generation
- AsyncSpade: Efficient Test-Time Scaling with Asynchronous Sparse Decoding
- Relational Transformer: Toward Zero-Shot Foundation Models for Relational Data
- VecInfer: Efficient LLM Inference with Low-Bit KV Cache via Outlier-Suppressed Vector Quantization
- lm-Meter: Unveiling Runtime Inference Latency for On-Device Language Models
- Shaken or Stirred? An Analysis of MetaFormer's Token Mixing for Medical Imaging
- The Anatomy of a Triton Attention Kernel
- Luth: Efficient French Specialization for Small Language Models and Cross-Lingual Transfer
- Staircase Streaming for Low-Latency Multi-Agent Inference
- SwiReasoning: Switch-Thinking in Latent and Explicit for Pareto-Superior Reasoning LLMs
- Transformers Discover Molecular Structure Without Graph Priors
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Hybrid Architectures for Language Models: Systematic Analysis and Design Insights
- Pushing on Multilingual Reasoning Models with Language-Mixed Chain-of-Thought
- Exact Causal Attention with 10% Fewer Operations
- RACE Attention: A Strictly Linear-Time Attention Layer for Training on Outrageously Large Contexts
- OneFlow: Concurrent Mixed-Modal and Interleaved Generation with Edit Flows
- Malice in Agentland: Down the Rabbit Hole of Backdoors in the AI Supply Chain
- Gather-Scatter Mamba: Accelerating Propagation with Efficient State Space Model
- Automated Structured Radiology Report Generation with Rich Clinical Context
- Efficient Multi-modal Large Language Models via Progressive Consistency Distillation
- Local Linear Attention: An Optimal Interpolation of Linear and Softmax Attention For Test-Time Regression
- Thoughtbubbles: an Unsupervised Method for Parallel Thinking in Latent Space
- LoRAFusion: Efficient LoRA Fine-Tuning for LLMs
- TASP: Topology-aware Sequence Parallelism
- Rearchitecting Datacenter Lifecycle for AI: A TCO-Driven Framework
- SlimPack: Fine-Grained Asymmetric Packing for Balanced and Efficient Variable-Length LLM Training
- Collaborative Compression for Large-Scale MoE Deployment on Edge
- TTT3R: 3D Reconstruction as Test-Time Training
- Context-Driven Performance Modeling for Causal Inference Operators on Neural Processing Units
- A Scalable Distributed Framework for Multimodal GigaVoxel Image Registration
- LVT: Large-Scale Scene Reconstruction via Local View Transformers
- Attention Surgery: An Efficient Recipe to Linearize Your Video Diffusion Transformer
- HAPT: Heterogeneity-Aware Automated Parallel Training on Heterogeneous Clusters
- ProxyAttn: Guided Sparse Attention via Representative Heads
- SparseServe: Unlocking Parallelism for Dynamic Sparse Attention in Long-Context LLM Serving
- Uni-X: Mitigating Modality Conflict with a Two-End-Separated Architecture for Unified Multimodal Models
- Hyperspherical Latents Improve Continuous-Token Autoregressive Generation
- Muon: Training and Trade-offs with Latent Attention and MoE
- Optimizing Privacy-Preserving Primitives to Support LLM-Scale Applications
- Brain Harmony: A Multimodal Foundation Model Unifying Morphology and Function into 1D Tokens
- MACE: A Hybrid LLM Serving System with Colocated SLO-aware Continuous Retraining Alignment
- SLA: Beyond Sparsity in Diffusion Transformers via Fine-Tunable Sparse-Linear Attention
- HiDe: Rethinking The Zoom-IN method in High Resolution MLLMs via Hierarchical Decoupling
- A Predictive and Synergistic Two-Layer Scheduling Framework for LLM Serving
- OracleGS: Grounding Generative Priors for Sparse-View Gaussian Splatting
- Where to Add PDE Diffusion in Transformers
- From Long to Lean: Performance-aware and Adaptive Chain-of-Thought Compression via Multi-round Refinement
- Reinforcement Learning-Guided Chain-of-Draft for Token-Efficient Code Generation
- Zeppelin: Balancing Variable-length Workloads in Data Parallel Large Model Training
- ChaosNexus: A Foundation Model for Universal Chaotic System Forecasting with Multi-scale Representations
- A Unifying Framework for Parallelizing Sequential Models with Linear Dynamical Systems
- Compute-Optimal Quantization-Aware Training
- OjaKV: Context-Aware Online Low-Rank KV Cache Compression with Oja's Rule
- InfiniPipe: Elastic Pipeline Parallelism for Efficient Variable-Length Long-Context LLM Training
- AIBA: Attention-based Instrument Band Alignment for Text-to-Audio Diffusion
- Real-Time Object Detection Meets DINOv3
- TyphoonMLA: A Mixed Naive-Absorb MLA Kernel For Shared Prefix
- RollPacker: Mitigating Long-Tail Rollouts for Fast, Synchronous RL Post-Training
- ARMesh: Autoregressive Mesh Generation via Next-Level-of-Detail Prediction
- Fine-Tuning LLMs to Analyze Multiple Dimensions of Code Review: A Maximum Entropy Regulated Long Chain-of-Thought Approach
- AMLA: MUL by ADD in FlashAttention Rescaling
- WEST: LLM based Speech Toolkit for Speech Understanding, Generation, and Interaction
- Enhancing Linear Attention with Residual Learning
- BurstEngine: an Efficient Distributed Framework for Training Transformers on Extremely Long Sequences of over 1M Tokens
- Gyges: Dynamic Cross-Instance Parallelism Transformation for Efficient LLM Inference
- RoboSSM: Scalable In-context Imitation Learning via State-Space Models
- Are We Scaling the Right Thing? A System Perspective on Test-Time Scaling
- scPlantAnnotate: an accurate and robust transformer-based model for plant cell type annotation
- Recall Before You Rank: Similarity-Guided Top-K Reuse for Efficient Long-Context Attention
- Back from the Future: Key-Value Cache Management by Counter-Causal Surprise
- SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer
- ReGenVC: End-to-End Real-Time Generative Video Coding at Ultra-Low Bitrate
- ROCS: Request-Oriented Compute Sharing for Efficient Large-Scale Recommendation
- Beyond KV Reconstruction: Functional Reconstruction for MLA Draft Models in Speculative Decoding
- FORGE: Fused On-Register Gradient Elimination for Memory-Efficient LLM Training
- Recursive transformers for semiconductor thermo-mechanical reliability
- KernelGenBench: A Multi-Source and Multi-Chip Benchmark for LLM-based Kernel Generation
- Multi-Head Attention Residuals
- Defusing the Trigger: Tail-Risk-Informed Attention Rebalancing for LLM Backdoor Mitigation
- Training Compute-Optimal Protein Language Models
- Gener anno : A Genomic Foundation Model for Metagenomic Annotation
- DAGE: Dual-Stream Architecture for Efficient and Fine-Grained Geometry Estimation
- Use What You Know: Causal Foundation Models with Partial Graphs
- Transformers and genome language models
- A Unified Transformer Architecture for Low-Latency and Scalable Wireless Signal Processing
- Protein Language Models: Is Scaling Necessary?
- CR-Net: Scaling Parameter-Efficient Training with Cross-Layer Low-Rank Structure
- LLM-based Agents Suffer from Hallucinations: A Survey of Taxonomy, Methods, and Directions
- Memory in Large Language Models: Mechanisms, Evaluation and Evolution
- An overview of neural architectures for self-supervised audio representation learning from masked spectrograms
- Reading Images Like Texts: Sequential Image Understanding in Vision-Language Models
- Visual Instruction Pretraining for Domain-Specific Foundation Models
- Probabilistic Token Alignment for Large Language Model Fusion
- STCast: Adaptive Boundary Alignment for Global and Regional Weather Forecasting
- FG-Attn: Leveraging Fine-Grained Sparse Attention in Video Diffusion Models
- Text-Scene: A Scene-to-Language Parsing Framework for 3D Scene Understanding
- Pico: A Modular Framework for Hypothesis-Driven Small Language Model Research
- HERO: Hierarchical Extrapolation and Refresh for Efficient World Models
- Efficient Pre-Training of LLMs via Topology-Aware Communication Alignment on More Than 9600 GPUs
- SuperGen: An Efficient Ultra-high-resolution Video Generation System with Sketching and Tiling
- Analyzing and Mitigating Surface Bias in Code Evaluation Metrics
- MedFact-R1: Towards Factual Medical Reasoning via Pseudo-Label Augmentation
- Evaluating the Limitations of Local LLMs in Solving Complex Programming Challenges
- LEAP: LLM Inference on Scalable PIM-NoC Architecture with Balanced Dataflow and Fine-Grained Parallelism
- SALT4Decompile: Inferring Source-level Abstract Logic Tree for LLM-Based Binary Decompilation
- Patent Language Model Pretraining with ModernBERT
- Value-Guided KV Compression for LLMs via Approximated CUR Decomposition
- TinySR: Pruning Diffusion for Real-World Image Super-Resolution
- Jim137/qkan: v0.1.0
- DiCache: Let Diffusion Model Determine Its Own Cache
- BWCache: Accelerating Video Diffusion Transformers through Block-Wise Caching
- FLAME: A Serving System Optimized for Large-Scale Generative Recommendation with Efficiency
- Re-purposing SAM into Efficient Visual Projectors for MLLM-Based Referring Image Segmentation
- Multi-Model Synthetic Training for Mission-Critical Small Language Models
- xOffense: An AI-driven autonomous penetration testing framework with offensive knowledge-enhanced LLMs and multi agent systems
- HPIM: Heterogeneous Processing-In-Memory-based Accelerator for Large Language Models Inference
- Reversible Deep Equilibrium Models
- Positional Encoding via Token-Aware Phase Attention
- ScaleDoc: Scaling LLM-based Predicates over Large Document Collections
- CoViPAL: Layer-wise Contextualized Visual Token Pruning for Large Vision-Language Models
- Large Language Models Imitate Logical Reasoning, but at what Cost?
- FastMTP: Accelerating LLM Inference with Enhanced Multi-Token Prediction
- ProtHyena: A fast and efficient foundation protein language model at single amino acid Resolution
- MindVL: Towards Efficient and Effective Training of Multimodal Large Language Models on Ascend NPUs
- SpeCa: Accelerating Diffusion Transformers with Speculative Feature Caching
- Genome-Factory: A Library for Tuning, Deploying, and Interpreting Genomic Foundation Models
- Multipole Semantic Attention: A Fast Approximation of Softmax Attention for Pretraining
- Characterizing the Efficiency of Distributed Training: A Power, Performance, and Thermal Perspective
- Compute Only 16 Tokens in One Timestep: Accelerating Diffusion Transformers with Cluster-Driven Feature Caching
- ButterflyQuant: Ultra-low-bit LLM Quantization through Learnable Orthogonal Butterfly Transforms
- Conditioning on PDE Parameters to Generalise Deep Learning Emulation of Stochastic and Chaotic Dynamics
- DiTReducio: A Training-Free Acceleration for DiT-Based TTS via Progressive Calibration
- Large Language Models in Document Intelligence: A Comprehensive Survey, Recent Advances, Challenges, and Future Trends
- Efficient Transformer-Based Piano Transcription With Sparse Attention Mechanisms
- DeepAries: Adaptive Rebalancing Interval Selection for Enhanced Portfolio Selection
- ViRanker: A BGE-M3 & Blockwise Parallel Transformer Cross-Encoder for Vietnamese Reranking
- LLM Architecture, Scaling Laws, and Economics: A Quick Summary
- Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism
- EvolKV: Evolutionary KV Cache Compression for LLM Inference
- A Survey of Long-Document Retrieval in the PLM and LLM Era
- Causal Attention with Lookahead Keys
- Block-Sparse Global Attention for Efficient Multi-View Geometry Transformers
- Dato: A Task-Based Programming Model for Dataflow Accelerators
- LatinX: Aligning a Multilingual TTS Model with Direct Preference Optimization
- Sticker-TTS: Learn to Utilize Historical Experience with a Sticker-driven Test-Time Scaling Framework
- Set Block Decoding is a Language Model Inference Accelerator
- SAC-MIL: Spatial-Aware Correlated Multiple Instance Learning for Histopathology Whole Slide Image Classification
- Differentiable Entropy Regularization: A Complexity-Aware Approach for Neural Optimization
- Scalable hybrid quantum Monte Carlo simulation of U(1) gauge field coupled to fermions on GPU
- Batch Query Processing and Optimization for Agentic Workflows
- AudioRWKV: Efficient and Stable Bidirectional RWKV for Audio Pattern Recognition
- Improving Large Vision and Language Models by Learning from a Panel of Peers
- Latent Space Single-Pixel Imaging Under Low-Sampling Conditions
- LobRA: Multi-tenant Fine-tuning over Heterogeneous Data
- GPT-OSS-20B: A Comprehensive Deployment-Centric Analysis of OpenAI's Open-Weight Mixture of Experts Model
- DSDE: Dynamic Speculative Decoding with KLD Stability for Real-World Serving
- Efficient Large Language Models with Zero-Shot Adjustable Acceleration
- Any-Order Flexible Length Masked Diffusion
- COMET: A Framework for Modeling Compound Operation Dataflows with Explicit Collectives
- Backdoor Samples Detection Based on Perturbation Discrepancy Consistency in Pre-trained Language Models
- Chunked TabPFN: Exact Training-Free In-Context Learning for Long-Context Tabular Data
- Balanced Actor Initialization: Stable RLHF Training of Distillation-Based Reasoning Models
- Accelerating Mixture-of-Experts Inference by Hiding Offloading Latency with Speculative Decoding
- Machine-learning based particle-flow algorithm in CMS
- Mixture of Contexts for Long Video Generation
- Provable Benefits of In-Tool Learning for Large Language Models
- Boltz-1 Democratizing Biomolecular Interaction Modeling
- OLMoASR: Open Models and Data for Training Robust Speech Recognition Models
- SwizzlePerf: Hardware-Aware LLMs for GPU Kernel Performance Optimization
- LMDeploy Accelerates Mixed-Precision LLM Inference with TurboMind
- Hits to Higgs: Reconstruction-Free Higgs Classification from Raw LHC Detector Data Using Higgsformers
- Conflict-Aware Soft Prompting for Retrieval-Augmented Generation
- Enhancing compact convolutional transformers with super attention
- ClusterFusion: Expanding Operator Fusion Scope for LLM Inference via Cluster-Level Collective Primitive
- Rethinking Caching for LLM Serving Systems: Beyond Traditional Heuristics
- Revisiting associative recall in modern recurrent models
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- Exploring Scaling Laws of CTR Model for Online Performance Improvement
- CARFT: Boosting LLM Reasoning via Contrastive Learning with Annotated Chain-of-Thought-based Reinforced Fine-Tuning
- Position Bias Mitigates Position Bias:Mitigate Position Bias Through Inter-Position Knowledge Distillation
- Lorentz-Equivariance without Limitations
- H2EAL: Hybrid-Bonding Architecture with Hybrid Sparse Attention for Efficient Long-Context LLM Inference
- MoEcho: Exploiting Side-Channel Attacks to Compromise User Privacy in Mixture-of-Experts LLMs
- Beyond Turing: Memory-Amortized Inference as a Foundation for Cognitive Computation
- Toward Better EHR Reasoning in LLMs: Reinforcement Learning with Expert Attention Guidance
- NovoMolGen: Rethinking Molecular Language Model Pretraining
- Prune2Drive: A Plug-and-Play Framework for Accelerating Vision-Language Models in Autonomous Driving
- Compact Attention: Exploiting Structured Spatio-Temporal Sparsity for Fast Video Generation
- FLARE: Fast Low-rank Attention Routing Engine
- CarelessWhisper: Turning Whisper into a Causal Streaming Model
- The Rise of Generative AI for Metal-Organic Framework Design and Synthesis
- E-CaTCH: Event-Centric Cross-Modal Attention with Temporal Consistency and Class-Imbalance Handling for Misinformation Detection
- Novel Parasitic Dual-Scale Modeling for Efficient and Accurate Multilingual Speech Translation
- Dataset Construction for Training LLM to Learn Analog Circuit Knowledge
- A learning-driven automatic planning framework for proton PBS treatments of H&N cancers
- BLADE: Block-Sparse Attention Meets Step Distillation for Efficient Video Generation
- Verify Distributed Deep Learning Model Implementation Refinement with Iterative Relation Inference
- Faster and Memory-Efficient Training of Sequential Recommendation Models for Large Catalogs
- READER: Retrieval-Assisted Drafter for Efficient LLM Inference
- Profiling Large Language Model Inference on Apple Silicon: A Quantization Perspective
- Retrospective Sparse Attention for Efficient Long-Context Generation
- Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions
- Exploring Multimodal Diffusion Transformers for Enhanced Prompt-based Image Editing
- CATP: Contextually Adaptive Token Pruning for Efficient and Enhanced Multimodal In-Context Learning
- GLiClass: Generalist Lightweight Model for Sequence Classification Tasks
- MobileViCLIP: An Efficient Video-Text Model for Mobile Devices
- DySK-Attn: A Framework for Efficient, Real-Time Knowledge Updating in Large Language Models via Dynamic Sparse Knowledge Attention
- The Fused Kernel Library: A C++ API to Develop Highly-Efficient GPU Libraries
- Multi-level Advantage Credit Assignment for Cooperative Multi-Agent Reinforcement Learning
- gpt-oss-120b & gpt-oss-20b Model Card
- Scalable Swin Transformer network for brain tumor segmentation from incomplete MRI modalities
- DeepFold-PLM: accelerating protein structure prediction via efficient homology search using protein language models
- Fewer Denoising Steps or Cheaper Per-Step Inference: Towards Compute-Optimal Diffusion Model Deployment
- AdaptInfer: Adaptive Token Pruning for Vision-Language Model Inference with Dynamical Text Guidance
- Crisp Attention: Regularizing Transformers via Structured Sparsity
- KnapFormer: An Online Load Balancer for Efficient Diffusion Transformers Training
- Estimating Musical Surprisal from Audio in Autoregressive Diffusion Model Noise Spaces
- Making Prompts First-Class Citizens for Adaptive LLM Pipelines
- iFairy: the First 2-bit Complex LLM with All Parameters in \±1, ± i\
- Why Attention Fails: A Taxonomy of Faults in Attention-Based Neural Networks
- HierarchicalPrune: Position-Aware Compression for Large-Scale Diffusion Models
- Share Your Attention: Transformer Weight Sharing via Matrix-based Dictionary Learning
- Efficient Inter-Task Attention for Multitask Transformer Models
- VLMQ: Efficient Post-Training Quantization for Large Vision-Language Models via Hessian Augmentation
- Macro-from-Micro Planning for High-Quality and Parallelized Autoregressive Long Video Generation
- Long Story Generation via Knowledge Graph and Literary Theory
- Following Route Instructions using Large Vision-Language Models: A Comparison between Low-level and Panoramic Action Spaces
- VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo
- Dynaword: From One-shot to Continuously Developed Datasets
- Forecasting When to Forecast: Accelerating Diffusion Models with Confidence-Gated Taylor
- Free-MoRef: Instantly Multiplexing Context Perception Capabilities of Video-MLLMs within Single Inference
- Trainable Dynamic Mask Sparse Attention
- A Glimpse to Compress: Dynamic Visual Token Pruning for Large Vision-Language Models
- FluidFormer: Transformer with Continuous Convolution for Particle-based Fluid Simulation
- Can3Tok: Canonical 3D Tokenization and Latent Modeling of Scene-Level 3D Gaussians
- SiLQ: Simple Large Language Model Quantization-Aware Training
- PiKV: KV Cache Management System for Mixture of Experts
- FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models
- Optimal Scheduling Algorithms for LLM Inference: Theory and Practice
- HiPrune: Training-Free Visual Token Pruning via Hierarchical Attention in Vision-Language Models
- Sortblock: Similarity-Aware Feature Reuse for Diffusion Model
- Representation Shift: Unifying Token Compression with FlashAttention
- Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts
- Systematic Evaluation of Optimization Techniques for Long-Context Language Models
- Calibrated Language Models and How to Find Them with Label Smoothing
- Self-Foveate: Enhancing Diversity and Difficulty of Synthesized Instructions from Unsupervised Text via Multi-Level Foveation
- BAR Conjecture: the Feasibility of Inference Budget-Constrained LLM Services with Authenticity and Reasoning
- Your Spending Needs Attention: Modeling Financial Habits with Transformers
- trAIce3D: A Prompt-Driven Transformer Based U-Net for Semantic Segmentation of Microglial Cells from Large-Scale 3D Microscopy Images
- TriangleMix: Accelerating Prefilling via Decoding-time Contribution Sparsity
- Model-free Speculative Decoding for Transformer-based ASR with Token Map Drafting
- Forecasting LLM Inference Performance via Hardware-Agnostic Analytical Modeling
- LeMix: Unified Scheduling for LLM Training and Inference on Multi-GPU Systems
- EgoPrune: Efficient Token Pruning for Egomotion Video Reasoning in Embodied Agent
- TransPrune: Token Transition Pruning for Efficient Large Vision-Language Model
- Sustainable AI Training via Hardware-Software Co-Design on NVIDIA, AMD, and Emerging GPU Architectures
- Clustering by Attention: Leveraging Prior Fitted Transformers for Data Partitioning
- EcoTransformer: Attention without Multiplication
- A Survey of Token Compression for Efficient Multimodal Large Language Models
- StackTrans: From Large Language Model to Large Pushdown Automata Model
- Rethinking LLM Inference Bottlenecks: Insights from Latent Attention and Mixture-of-Experts
- HydraMamba: Multi-Head State Space Model for Global Point Cloud Learning
- Modality Agnostic Efficient Long Range Encoder
- Iwin Transformer: Hierarchical Vision Transformer using Interleaved Windows
- Linear Memory SE(2) Invariant Attention
- CTA-Pipelining: A Latency-Oriented Spatial Scaling Method for Multi-GPU Systems
- Relax: Composable Abstractions for End-to-End Dynamic Machine Learning
- The Role of Feedback Alignment in Self-Distillation
- Eywa: Provenance-Grounded Long-Term Memory for AI Agents
- Evaluating CUDA Tile for AI Workloads on Hopper and Blackwell GPUs
- Kerncap: Automated Kernel Extraction and Isolation for AMD GPUs
- SAW-INT4: System-Aware 4-Bit KV-Cache Quantization for Real-World LLM Serving
- PRAGMA: Revolut Foundation Model
- Docopilot: Improving Multimodal Models for Document-Level Understanding
- K-Search: LLM Kernel Generation via Co-Evolving Intrinsic World Model
- Retrieval-Aware Distillation for Transformer-SSM Hybrids
- Characterizing WebGPU Dispatch Overhead for LLM Inference Across Four GPU Vendors, Three Backends, and Three Browsers
- Enigma: An Efficient Model for Deciphering Regulatory Genomics
- Photonic Fabric Platform for AI Accelerators
- A Comprehensive Review of Transformer-based language models for Protein Sequence Analysis and Design
- Hyperbolic Deep Learning for Foundation Models: A Survey
- DistrAttention: An Efficient and Flexible Self-Attention Mechanism on Modern GPUs
- BrownoutServe: SLO-Aware Inference Serving under Bursty Workloads for MoE-based LLMs
- BucketServe: Bucket-Based Dynamic Batching for Smart and Efficient LLM Inference Serving
- Pixel-Resolved Long-Context Learning for Turbulence at Exascale: Resolving Small-scale Eddies Toward the Viscous Limit
- Custom Algorithm-based Fault Tolerance for Attention Layers in Transformers
- VGGT-Long: Chunk it, Loop it, Align it -- Pushing VGGT's Limits on Kilometer-scale Long RGB Sequences
- VideoITG: Multimodal Video Understanding with Instructed Temporal Grounding
- Characterizing State Space Model (SSM) and SSM-Transformer Hybrid Language Model Performance with Long Context Length
- GitChameleon 2.0: Evaluating AI Code Generation Against Python Library Version Incompatibilities
- On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention
- Scalable Non-Equivariant 3D Molecule Generation via Rotational Alignment
- Kevin: Multi-Turn RL for Generating CUDA Kernels
- ELK: Exploring the Efficiency of Inter-core Connected AI Chips with Deep Learning Compiler Techniques
- FastFLUX: Pruning FLUX with Block-wise Replacement and Sandwich Training
- Breaking the ICE: Exploring promises and challenges of benchmarks for Inference Carbon & Energy estimation for LLMs
- Plug-and-play linear attention with provable guarantees for training-free image restoration
- TTrace: Lightweight Error Checking and Diagnosis for Distributed Training
- MagCache: Fast Video Generation with Magnitude-Aware Cache
- KV-Latent: Dimensional-level KV Cache Reduction with Frequency-aware Rotary Positional Embedding
- LaCache: Ladder-Shaped KV Caching for Efficient Long-Context Modeling of Large Language Models
- Mixture-of-Recursions: Learning Dynamic Recursive Depths for Adaptive Token-Level Computation
- Past-Future Scheduler for LLM Serving under SLA Guarantees
- Brevity is the soul of sustainability: Characterizing LLM response lengths
- Uncovering Causal Relation Shifts in Event Sequences under Out-of-Domain Interventions
- ElasticMM: Efficient Multimodal LLMs Serving with Elastic Multimodal Parallelism
- Inter2Former: Dynamic Hybrid Attention for Efficient High-Precision Interactive
- Compute Requirements for Algorithmic Innovation in Frontier AI Models
- Simplifying Traffic Anomaly Detection with Video Foundation Models
- Optimizing Sequential Multi-Step Tasks with Parallel LLM Agents
- Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA)
- Efficient Deployment of Vision-Language Models on Mobile Devices: A Case Study on OnePlus 13R
- Upsample What Matters: Region-Adaptive Latent Sampling for Accelerated Diffusion Transformers
- Scalable Spatiotemporal Inference with Biased Scan Attention Transformer Neural Processes
- Quantifying Mix Network Privacy Erosion with Generative Models
- BlindSight: Harnessing Sparsity for Efficient Vision-Language Models
- Compactor: Calibrated Query-Agnostic KV Cache Compression with Approximate Leverage Scores
- Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs
- Bridging the Plausibility-Validity Gap by Fine-Tuning a Reasoning-Enhanced LLM for Chemical Synthesis and Discovery
- Intra-DP: A High Performance Collaborative Inference System for Mobile Edge Computing
- TransX: Scaling Transformer-based Recommendation via Behavioral and Serving Stream Crossings
- ETT: Expanding the Long Context Understanding Capability of LLMs at Test-Time
- WikiChat: Stopping the Hallucination of Large Language Model Chatbots by Few-Shot Grounding on Wikipedia
- Improving BERT for Symbolic Music Understanding Using Token Denoising and Pianoroll Prediction
- AXLearn: Modular, Hardware-Agnostic Large Model Training
- Scaling Context Requires Rethinking Attention
- RAT: Bridging RNN Efficiency and Attention Accuracy via Chunk-based Sequence Modeling
- High-Level Big Integer Arithmetic in Futhark for GPUs
- WitCert: Sound Runtime Risk Observability and Gating for KV-Cache Quantization
- OMAR-RQ: Open Music Audio Representation Model Trained with Multi-Feature Masked Token Prediction
- Analysis and Optimized CXL-Attached Memory Allocation for Long-Context LLM Fine-Tuning
- HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference
- System-performance and cost modeling of Large Language Model training and inference
- Resolving Turbulent Magnetohydrodynamics: A Hybrid Operator-Diffusion Framework
- SD-Acc: Accelerating Stable Diffusion through Phase-aware Sampling and Hardware Co-Optimizations
- Stem: Rethinking Causal Information Flow in Sparse Attention
- La RoSA: Enhancing LLM Efficiency via Layerwise Rotated Sparse Activation
- ZeCO: Zero Communication Overhead Sequence Parallelism for Linear Attention
- MambAttention: Mamba with Multi-Head Attention for Generalizable Single-Channel Speech Enhancement
- VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator
- Recurrent Memory-Augmented Transformers with Chunked Attention for Long-Context Language Modeling
- HelixPipe: Efficient Distributed Training of Long Sequence Transformers with Attention Parallel Pipeline Parallelism
- CGEarthEye:A High-Resolution Remote Sensing Vision Foundation Model Based on the Jilin-1 Satellite Constellation
- Improve Underwater Object Detection through YOLOv12 Architecture and Physics-informed Augmentation
- A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation
- Agent.xpu: Efficient Scheduling of Agentic LLM Workloads on Heterogeneous SoC
- VMoBA: Mixture-of-Block Attention for Video Diffusion Models
- Pipelined Decoder for Efficient Context-Aware Text Generation
- Not All Explanations for Deep Learning Phenomena Are Equally Valuable
- Token Activation Map to Visually Explain Multimodal LLMs
- Masked Gated Linear Unit
- Ovis-U1 Technical Report
- RL4CO: an Extensive Reinforcement Learning for Combinatorial Optimization Benchmark
- A Systematic Study of Compositional Syntactic Transformer Language Models
- STR-Match: Matching SpatioTemporal Relevance Score for Training-Free Video Editing
- BEST-Route: Adaptive LLM Routing with Test-Time Optimal Compute
- DriveMRP: Enhancing Vision-Language Models with Synthetic Motion Data for Motion Risk Prediction
- Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models
- MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive Operators
- GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling
- Multi-View Contrastive Learning for Robust Domain Adaptation in Medical Time Series Analysis
- A Survey of LLM Inference Systems
- BitVLA: 1-bit Vision-Language-Action Models for Robotics Manipulation
- FOCUS: Internal MLLM Representations for Efficient Fine-Grained Visual Question Answering
- MoE-GPS: Guidlines for Prediction Strategy for Dynamic Expert Duplication in MoE Load Balancing
- MoQAE: Mixed-Precision Quantization for Long-Context LLM Inference via Mixture of Quantization-Aware Experts
- Fine-grained Token Allocation Via Operation Pruning for Efficient MLLMs
- Radial Attention: O(nlog n) Sparse Attention with Energy Decay for Long Video Generation
- RecLLM-R1: A Two-Stage Training Paradigm with Reinforcement Learning and Chain-of-Thought v1
- MegaFold: Efficient Training of Next-Generation 3D Attention Protein Models on Cross-Platform GPUs
- MiniCPM4: Ultra-Efficient LLMs on End Devices
- Comparative Analysis of Lion and AdamW Optimizers for Cross-Encoder Reranking with MiniLM, GTE, and ModernBERT
- Mechanistic Interpretability in the Presence of Architectural Obfuscation
- VeriLocc: End-to-End Cross-Architecture Register Allocation via LLM
- Foundation Model Empowered Synesthesia of Machines (SoM): AI-native Intelligent Multi-Modal Sensing-Communication Integration
- Optimizing MoE Routers: Design, Implementation, and Evaluation in Transformer Models
- FinCoT: Grounding Chain-of-Thought in Expert Financial Reasoning
- LazyEviction: Lagged KV Eviction with Attention Pattern Observation for Efficient Long Reasoning
- LBMamba: Locally Bi-directional Mamba
- REIS: A High-Performance and Energy-Efficient Retrieval System with In-Storage Processing
- When Does Divide and Conquer Work for Long Context LLM? A Noise Decomposition Framework
- T-SHRED: Symbolic Regression for Regularization and Model Discovery with Transformer Shallow Recurrent Decoders
- Mondrian: Transformer Operators via Domain Decomposition
- eLLM: Elastic Memory Management Framework for Efficient LLM Serving
- Zero-Shot Reinforcement Learning Under Partial Observability
- Scaling Intelligence: Designing Data Centers for Next-Gen Language Models
- Earth Observation Foundation Model PhilEO: Pretraining on the MajorTOM and FastTOM Datasets
- Paged Attention Meets FlexAttention: Unlocking Long-Context Efficiency in Deployed Inference
- Efficient Serving of LLM Applications with Probabilistic Demand Modeling
- Think Clearly: Improving Reasoning via Redundant Token Pruning
- GLU Attention Improve Transformer
- AlphaEvolve: A coding agent for scientific and algorithmic discovery
- Discrete Diffusion in Large Language and Multimodal Models: A Survey
- PRISM2: Unlocking Multi-Modal General Pathology AI with Clinical Dialogue
- DDiT: Dynamic Resource Allocation for Diffusion Transformer Model Serving
- StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns
- Scaling Algorithm Distillation for Continuous Control with Mamba
- Universal Jailbreak Suffixes Are Strong Attention Hijackers
- GTA: Grouped-head latenT Attention
- Guiding Cross-Modal Representations with MLLM Priors via Preference Alignment
- Bridging the Digital Divide: Small Language Models as a Pathway for Physics and Photonics Education in Underdeveloped Regions
- LIFELONG SOTOPIA: Evaluating Social Intelligence of Language Agents Over Lifelong Social Interactions
- Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling
- QiMeng-Attention: SOTA Attention Operator is generated by SOTA Attention Algorithm
- Two Heads Are Better than One: Simulating Large Transformers with Small Ones
- Lag-Relative Sparse Attention In Long Context Training
- CentroidKV: Efficient Long-Context LLM Inference via KV Cache Clustering
- TruncQuant: Truncation-Ready Quantization for DNNs with Flexible Weight Bit Precision
- A Survey of Foundation Models for IoT: Taxonomy and Criteria-Based Analysis
- Semantic Scheduling for LLM Inference
- The Impact of Partial Computations on the Red-Blue Pebble Game
- ConTextTab: A Semantics-Aware Tabular In-Context Learner
- SWE-Factory: Your Automated Factory for Issue Resolution Training Data and Evaluation Benchmarks
- Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs
- NoLoCo: No-all-reduce Low Communication Training Method for Large Models
- Don't Pay Attention
- PyLO: Towards Accessible Learned Optimizers in PyTorch
- RARL: Improving Medical VLM Reasoning and Generalization with Reinforcement Learning and LoRA under Data and Hardware Constraints
- Training with Confidence: Catching Silent Errors in Deep Learning Training with Automated Proactive Checks
- Accelerating Sparse Transformer Inference on GPU
- Cartridges: Lightweight and general-purpose long context representations via self-study
- S2GO: Streaming Sparse Gaussian Occupancy Prediction
- Diagonal Batching Unlocks Parallelism in Recurrent Memory Transformers for Long Contexts
- Astraea: A Token-wise Acceleration Framework for Video Diffusion Transformers
- FPSAttention: Training-Aware FP8 and Sparsity Co-Design for Fast Video Diffusion
- TabFlex: Scaling Tabular Learning to Millions with Linear Attention
- FlashMoE: Fast Distributed MoE in a Single Kernel
- Mixture-of-Experts Meets In-Context Reinforcement Learning
- Kinetics: Rethinking Test-Time Scaling Laws
- Sentinel: SOTA model to protect against prompt injections
- Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment
- Video, How Do Your Tokens Merge?
- CAD-Llama: Leveraging Large Language Models for Computer-Aided Design Parametric 3D Model Generation
- Rethinking LLM Advancement: Compute-Dependent and Independent Paths to Progress
- MFSeg: Efficient Multi-frame 3D Semantic Segmentation
- AhaKV: Adaptive Holistic Attention-Driven KV Cache Eviction for Efficient Inference of Large Language Models
- TokAlign: Efficient Vocabulary Adaptation via Token Alignment
- FullDiT2: Efficient In-Context Conditioning for Video Diffusion Transformers
- Homogeneous Keys, Heterogeneous Values: Exploiting Local KV Cache Asymmetry for Long-Context LLMs
- HMAR: Efficient Hierarchical Masked Auto-Regressive Image Generation
- Vision Remember: Recovering Visual Information in Efficient LVLM with Vision Feature Resampling
- Parallel CPU-GPU Execution for LLM Inference on Constrained GPUs
- QKV Projections Require a Fraction of Their Memory
- Contrast & Compress: Learning Lightweight Embeddings for Short Trajectories
- TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models
- HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference
- Rethinking Dynamic Networks and Heterogeneous Computing with Automatic Parallelization
- Chipmunk: Training-Free Acceleration of Diffusion Transformers with Dynamic Column-Sparse Deltas
- Comba: Improving Bilinear RNNs with Closed-loop Control
- SynPO: Synergizing Descriptiveness and Preference Optimization for Video Detailed Captioning
- Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective
- Foresight: Adaptive Layer Reuse for Accelerated and High-Quality Text-to-Video Generation
- Probabilistic Forecasting for Building Energy Systems using Time-Series Foundation Models
- SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling
- 50 Years of Automated Face Recognition
- Large Language Model Meets Constraint Propagation
- AnchorAttention: Difference-Aware Sparse Attention with Stripe Granularity
- Fortune: Formula-Driven Reinforcement Learning for Symbolic Table Reasoning in Language Models
- AC-ODM: Actor--Critic Online Data Mixing for Sample-Efficient LLM Pretraining
- RetroInfer: A Vector-Storage Approach for Scalable Long-Context LLM Inference
- Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism
- Speeding up Model Loading with fastsafetensors
- MemAscend: System Memory Optimization for SSD-Offloaded LLM Fine-Tuning
- Test-Time Training Done Right
- LoLA: Low-Rank Linear Attention With Sparse Caching
- RiverMamba: A State Space Model for Global River Discharge and Flood Forecasting
- Advancing Expert Specialization for Better MoE
- Curse of High Dimensionality Issue in Transformer for Long-context Modeling
- Learning in Compact Spaces with Approximately Normalized Transformer
- Towards Scalable Language-Image Pre-training for 3D Medical Imaging
- EFIM: Efficient Serving of LLMs for Infilling Tasks with Improved KV Cache Reuse
- Balanced Token Pruning: Accelerating Vision Language Models Beyond Local Optimization
- VScan: Rethinking Visual Token Reduction for Efficient Large Vision-Language Models
- RePaViT: Scalable Vision Transformer Acceleration via Structural Reparameterization on Feedforward Network Layers
- Tevatron 2.0: Unified Document Retrieval Toolkit across Scale, Language, and Modality
- Geometric Hyena Networks for Large-scale Equivariant Learning
- FlashFormer: Whole-Model Kernels for Efficient Low-Batch Inference
- The quest for the GRAph Level autoEncoder (GRALE)
- HAD: Hybrid Architecture Distillation Outperforms Teacher in Genomic Sequence Modeling
- SpecExtend: A Drop-in Enhancement for Speculative Decoding of Long Sequences
- SageAttention2++: A More Efficient Implementation of SageAttention2
- Continuous-Time Attention: PDE-Guided Mechanisms for Long-Sequence Transformers
- In Search of Adam's Secret Sauce
- FlashDLM: Accelerating Diffusion Language Model Inference via Efficient KV Caching and Guided Diffusion
- HoliTom: Holistic Token Merging for Fast Video Large Language Models
- Visual Product Graph: Bridging Visual Products And Composite Images For End-to-End Style Recommendations
- Hardware-Efficient Attention for Fast Decoding
- efunc: An Efficient Function Representation without Neural Networks
- Lorentz Local Canonicalization: How to Make Any Network Lorentz-Equivariant
- Large Language Models as Autonomous Spacecraft Operators in Kerbal Space Program
- Skrull: Towards Efficient Long Context Fine-tuning through Dynamic Data Scheduling
- SCFormer: Structured Channel-wise Transformer with Cumulative Historical State for Multivariate Time Series Forecasting
- Small Language Models: Architectures, Techniques, Evaluation, Problems and Future Adaptation
- FlowCut: Rethinking Redundancy via Information Flow for Efficient Vision-Language Models
- REARANK: Reasoning Re-ranking Agent via Reinforcement Learning
- Understanding Transformer from the Perspective of Associative Memory
- Jodi: Unification of Visual Generation and Understanding via Joint Modeling
- 100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability?
- Shifting AI Efficiency From Model-Centric to Data-Centric Compression
- Estimating Online Influence Needs Causal Modeling! Counterfactual Analysis of Social Media Engagement
- SRDiffusion: Accelerate Video Diffusion Inference via Sketching-Rendering Cooperation
- Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling
- ToDRE: Effective Visual Token Pruning via Token Diversity and Task Relevance
- Equivariant Flow Matching for Point Cloud Assembly
- MonarchAttention: Zero-Shot Conversion to Fast, Hardware-Aware Structured Attention
- MTGR: Industrial-Scale Generative Recommendation Framework in Meituan
- Linear Multi-Timescale Retention as a Memory-Efficient Vision-Language Bridge
- FAR: Function-preserving Attention Replacement for IMC-friendly Inference
- FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators
- VORTA: Efficient Video Diffusion via Routing Sparse Attention
- Lookahead Q-Cache: Achieving More Consistent KV Cache Eviction via Pseudo Query
- SpectraLDS: Provable Distillation for Linear Dynamical Systems
- Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization
- The Real Barrier to LLM Agent Usability is Agentic ROI
- Learning What to Remember: Test-Time Training via Context Distillation
- BehaveGPT: A Foundation Model for Large-scale User Behavior Modeling
- Less Context, Same Performance: A RAG Framework for Resource-Efficient LLM-Based Clinical NLP
- Cooperative Coevolution for Resource-Constrained Agentic LLM Post-Training
- Fast and Accurate Quotation Attribution in Literary Texts
- Longer Context, Deeper Thinking: Uncovering the Role of Long-Context Ability in Reasoning
- Training-Free Efficient Video Generation via Dynamic Token Carving
- Mamba with Hierarchical Memory: Solving Representation Bottleneck in Long Sequence Modeling
- PaTH Attention: Position Encoding via Accumulating Householder Transformations
- Efficient All-Pairs Correlation Volume Sampling for Optical Flow Estimation
- MARché: Fast Masked Autoregressive Image Generation with Cache-Aware Attention
- CASTILLO: Characterizing Response Length Distributions of Large Language Models
- Understanding Differential Transformer Unchains Pretrained Self-Attentions
- Native Segmentation Vision Transformers
- After Retrieval, Before Generation: Enhancing the Trustworthiness of Large Language Models in RAG
- DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling
- Revealing Language Model Trajectories via Kullback-Leibler Divergence
- SUS backprop: linear backpropagation algorithm for long inputs in transformers
- TELLER: Non-intrusive Cross-Layer Root-Cause Analysis for LLM Inference
- SwarmDiff: Swarm Robotic Trajectory Planning in Cluttered Environments via Diffusion Transformer
- One QK Channel, Many Sources: Guarding Low-Precision Attention Collapse
- RAG or Fine-tuning? A Comparative Study on LCMs-based Code Completion in Industry
- Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM
- LiveVLM: Efficient Online Video Understanding via Streaming-Oriented KV Cache and Retrieval
- MaxPoolBERT: Enhancing BERT Classification via Layer- and Token-Wise Aggregation
- UNet with Self-Adaptive Mamba-Like Attention and Causal-Resonance Learning for Medical Image Segmentation
- 3DZip: Spatial-Aware Feature Diversity-Guided Token Compression for 3D Question Answering
- Balanced and Elastic End-to-end Training of Dynamic LLMs
- ServerlessLoRA: Minimizing Latency and Cost in Serverless Inference for LoRA-Based LLMs
- Low-Cost FlashAttention with Fused Exponential and Multiplication Hardware Operators
- Video Compression Commander: Plug-and-Play Inference Acceleration for Video Large Language Models
- Polar Sparsity: High Throughput Batched LLM Inferencing with Scalable Contextual Sparsity
- ModRWKV: Transformer Multimodality in Linear Time
- FLASH-D: FlashAttention with Hidden Softmax Division
- P2HCT: Plug-and-Play Hierarchical C2F Transformer for Multi-Scale Feature Fusion
- HydraInfer: Hybrid Disaggregated Scheduling for Multimodal Large Language Model Serving
- Sense and Sensitivity: Examining the Influence of Semantic Recall on Long Context Code Understanding
- An Empirical Study of Many-to-Many Summarization with Large Language Models
- Optimizing Anytime Reasoning via Budget Relative Policy Optimization
- Sandwich: Joint Configuration Search and Hot-Switching for Efficient CPU LLM Serving
- VSA: Faster Video Diffusion with Trainable Sparse Attention
- Celty: SpMspV GPU Kernel and SIMT Co-Design for Efficient Dual-Sparse LLM Inference
- A Token is Worth over 1,000 Tokens: Efficient Knowledge Distillation through Low-Rank Clone
- PSC: Extending Context Window of Large Language Models via Phase Shift Calibration
- Bishop: Sparsified Bundling Spiking Transformers on Heterogeneous Cores with Error-Constrained Pruning
- Fast RoPE Attention: Combining the Polynomial Method and Fast Fourier Transform
- Zellige: Moldable Sequence Placement for Mixed Image-Video DiT Training
- AutoMedEval: Harnessing Language Models for Automatic Medical Capability Evaluation
- Efficiently Building a Domain-Specific Large Language Model from Scratch: A Case Study of a Classical Chinese Large Language Model
- Arrow: Adaptive Scheduling Mechanisms for Disaggregated LLM Inference Architecture
- GeoMaNO: Geometric Mamba Neural Operator for Partial Differential Equations
- DraftAttention: Fast Video Diffusion via Low-Resolution Attention Guidance
- Riemannian Attention Mechanisms for Transformers: A Theoretical Framework and Architecture Design
- GRACE: Generative Recommender Acceleration Engine for Real-Time Ads Retrieval
- FlashBias: Fast Computation of Attention with Bias
- Stop When Memory Suffices: Evidence-Conditioned Progressive Execution for LLM Agents
- MegaScale-MoE: Large-Scale Communication-Efficient Training of Mixture-of-Experts Models in Production
- METHOD: Modular Efficient Transformer for Health Outcome Discovery
- AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction
- Accurate KV Cache Quantization with Outlier Tokens Tracing
- QVGen: Pushing the Limit of Quantized Video Generative Models
- Efficient Attention via Pre-Scoring: Prioritizing Informative Keys in Transformers
- SageAttention3: Microscaling FP4 Attention for Inference and An Exploration of 8-Bit Training
- Parallel Scaling Law for Language Models
- MathCoder-VL: Bridging Vision and Code for Enhanced Multimodal Mathematical Reasoning
- Insights into DeepSeek-V3: Scaling Challenges and Reflections on Hardware for AI Architectures
- Aquarius: A Family of Industry-Level Video Generation Models for Marketing Scenarios
- Scaling Context, Not Parameters: Training a Compact 7B Language Model for Efficient Long-Context Processing
- Nova: An End-to-End MLIR Compiler for Deep Learning
- Motif-Mamba: network motif improved mamba for long-range sequence modeling
- HERALD: High-Throughput Block Diffusion LLM Serving via CPU-GPU Cooperative KV Cache Retrieval
- AgentCompile: An LLM-Guided Compiler for Direct CUDA Inference
- HSRAI: Permutation-Preserving Address Interleaving with Hierarchical Balance Metrics
- FlashMLA-ETAP: Efficient Transpose Attention Pipeline for Accelerating MLA Inference on NVIDIA H20 GPUs
- Putting It All into Context: Simplifying Agents with LCLMs
- Fused3S: Fast Sparse Attention on Tensor Cores
- Is MixIT Really Unsuitable for Correlated Sources? Exploring MixIT for Unsupervised Pre-training in Music Source Separation
- OLinear: A Linear Model for Time Series Forecasting in Orthogonally Transformed Domain
- QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration
- Attention Is Not All You Need: The Importance of Feedforward Networks in Transformer Models
- Challenging GPU Dominance: When CPUs Outperform for On-Device LLM Inference
- AROpt: An Optimization Method for Autoregressive Time Series Forecasting
- Small Clips, Big Gains: Learning Long-Range Refocused Temporal Information for Video Super-Resolution
- Argus: Metric Panoramic 3D Reconstruction for Indoor Scenes
- Phantora: Maximizing Code Reuse in Simulation-based Machine Learning System Performance Estimation
- PipeSpec: Breaking Stage Dependencies in Hierarchical LLM Decoding
- CodeSSM: Towards State Space Models for Code Understanding
- Time is Not Compute: Scaling Laws for Wall-Clock Constrained Training on Consumer GPUs
- Sustainability Is Not Linear: Quantifying Performance, Energy, and Privacy Trade-offs in On-Device Intelligence
- Reroute, Don't Remove: Recoverable Visual Token Routing for Vision-Language Models
- Kwai Keye-VL-2.0 Technical Report
- Six Open Questions in Machine-Learned Interatomic Potential Foundation Models
- Linear-Time and Constant-Memory Text Embeddings Based on Recurrent Language Models
- FlashAttention-4: Algorithm and Kernel Pipelining Co-Design for Asymmetric Hardware Scaling
- PackFlow: Generative Molecular Crystal Structure Prediction via Reinforcement Learning Alignment
- Output-Aware Rotation for INT2 KV-Cache Quantization
- Logic-Gated Time-Shared Feedforward Networks for Alternating Finite Automata: Exact Simulation and Learnability
- Mixture-of-Parallelisms: Towards Memory-Efficient Training Stack for Mixture-of-Experts Models
- KernelSight-LM: A Kernel-Level LLM Inference Simulator
- Hardware Acceleration for Neural Networks: A Comprehensive Survey
- Epiphany-Aware KV Cache Eviction Without the Attention Matrix
- BluTrain: A C++/CUDA Framework for AI Systems
- RoPE-Aware Bit Allocation for KV-Cache Quantization
- A 35B Hybrid-Attention Mixture-of-Experts Model on a 6GB 2011 GPU: Hand-Written 4-bit CUDA Inference for Fermi
- Scalable Physics-Inspired Transformers for Spin Glasses
- ARGUS: Production-Scale Tracing and Performance Diagnosis for over 10,000-GPU Clusters
- Mojo: A Promising Tool for Scalable Financial AI Efficiency
- MiniMax Sparse Attention
- SPADE: Split-and-Delay Embeddings for Autoregressive High-Granularity Calorimeter Simulation
- GRPO Does Not Close the Multi-Agent Coordination Gap
- Scaling On-Device GPU Inference for Large Generative Models
- Vision Mamba in Remote Sensing: A Comprehensive Survey of Techniques, Applications and Outlook
- Helios: Real Real-Time Long Video Generation Model
- Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing
- Ada-MK: Adaptive MegaKernel Optimization via Automated DAG-based Search for LLM Inference
- Tessera: Unlocking Heterogeneous GPUs through Kernel-Granularity Disaggregation
- SparseX: Efficient Segment-Level KV Cache Sharing for Interleaved LLM Serving
- Trading Complexity for Expressivity Through Structured Generalized Linear Token Mixing
- RTP-LLM: High-Performance Alibaba LLM Inference Engine
- Structured-Sparse Attention for Entity Tracking with Subquadratic Sequence Complexity
- FlashInfer-Bench: Building the Virtuous Cycle for AI-driven LLM Systems
- GPU Performance Portability needs Autotuning
- KernelBrain: Coarse-to-Fine, Budget-Aware Search for Agentic GPU Kernel Optimization
- Separating Intelligence from Inference: A Standard for Edge-Native AI Computing
- KV-Fold: One-Step KV-Cache Recurrence for Long-Context Inference
- DisagMoE: Computation-Communication overlapped MoE Training via Disaggregated AF-Pipe Parallelism
- Spark3R: Asymmetric Token Reduction Makes Fast Feed-Forward 3D Reconstruction
- When Quantization Is Free: An int4 KV Cache That Outruns fp16 on Apple Silicon
- Softpick: No Attention Sink, No Massive Activations with Rectified Softmax
- Blockbuster, Part 1: Block-level AI Operator Fusion
- Combatting Dimensional Collapse in LLM Pre-Training Data via Diversified File Selection
- Ascendra: Dynamic Request Prioritization for Efficient LLM Serving
- Reviving Any-Subset Autoregressive Models with Principled Parallel Sampling and Speculative Decoding
- Hetu v2: A General and Scalable Deep Learning System with Hierarchical and Heterogeneous Single Program Multiple Data Annotations
- CellxPert: Inference-Time MCMC Steering of a Multi-Omics Single-Cell Foundation Model for In-Silico Perturbation
- A Comparative Study on Positional Encoding for Time-frequency Domain Dual-path Transformer-based Source Separation Models
- Boosting LLM Serving through Spatial-Temporal GPU Resource Sharing
- The State-Prediction Separation Hypothesis
- QuasiMoTTo: Quasi-Monte Carlo Test-Time Scaling
- Tool Attention Is All You Need: Dynamic Tool Gating and Lazy Schema Loading for Eliminating the MCP/Tools Tax in Scalable Agentic Workflows
- Mean Mode Screaming: Mean--Variance Split Residuals for 1000-Layer Diffusion Transformers
- Screening Is Enough
- Facts as First Class Objects: Knowledge Objects for Persistent LLM Memory
- AI+HW 2035: Shaping the Next Decade
- Deep Sequence Modeling with Quantum Dynamics: Language as a Wave Function
- Interleaved Head Attention
- Kunlun: Establishing Scaling Laws for Massive-Scale Recommendation Systems through Unified Architecture Design
- Optimizing Agentic Workflows using Meta-tools
- semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage
- R-Sparse: Rank-Aware Activation Sparsity for Efficient LLM Inference
- Neural Garbage Collection: Learning to Forget while Learning to Reason
- CIMple: Standard-cell SRAM-based CIM with LUT-based split softmax for attention acceleration
- Small Models, Big Tasks: An Exploratory Empirical Study on Small Language Models for Function Calling
- Ragged Paged Attention: A High-Performance and Flexible LLM Inference Kernel for TPU
- A Method for the Architecture of a Medical Vertical Large Language Model Based on Deepseek R1
- EcoServe: Enabling Cost-effective LLM Serving with Proactive Intra- and Inter-Instance Orchestration
- The Long-Horizon Task Mirage? Diagnosing Where and Why Agentic Systems Break
- SWE-AGILE: A Software Agent Framework for Efficiently Managing Dynamic Reasoning Context
- WaveTune: Wave-aware Bilinear Modeling for Efficient GPU Kernel Auto-tuning
- Nexus: Same Pretraining Loss, Better Downstream Generalization via Common Minima
- Blink: CPU-Free LLM Inference by Delegating the Serving Stack to GPU and SmartNIC
- Sawtooth Wavefront Reordering: Enhanced CuTile FlashAttention on NVIDIA GB10
- VibeTensor: System Software for Deep Learning, Fully Generated by AI Agents
- PipeWeave: Synergizing Analytical and Learning Models for Unified GPU Performance Prediction
- Threshold Differential Attention for Sink-Free, Ultra-Sparse, and Non-Dispersive Language Modeling
- DNATokenizer: A GPU-First Byte-to-Identifier Tokenizer for High-Throughput DNA Language Models
- CUDA MPC: A GPU-Native Solver for Model Predictive Control
- Attention Needs to Focus: A Unified Perspective on Attention Allocation
- SPADE: An Input-Adaptive Sparse Attention Engine for Fast Video Diffusion Models Inference
- Embedding Empirical Distributions for Computing Optimal Transport Maps
- An Empirical Study on Prompt Compression for Large Language Models
- TileLang: A Composable Tiled Programming Model for AI Systems
- L3: DIMM-PIM Integrated Architecture and Coordination for Scalable Long-Context LLM Inference
- JITServe: SLO-aware LLM Serving with Imprecise Request Information
- EvtGraph: Event-Adaptive Compression for Sparse Temporal Graph Learning in Multimodal Time Series
- Architectural Implications of Agentic AI Workflows
- When does training on downscaled images yield the same gradients?
- Deltoris: Enabling Real-time VLA Inference in Embodied AI via Bit-level Sparsity and Speculative Inference
- Training-Free Hashing-Based Attention via Binary Principal Components
- Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms
- SHIELD: A Segmented Hierarchical Memory Architecture for Energy-Efficient LLM Inference on Edge NPUs
- RooflineBench: A Benchmarking Framework for On-Device LLMs via Roofline Analysis
- PARD: Accelerating LLM Inference with Low-Cost PARallel Draft Model Adaptation
- Generalized Neighborhood Attention: Multi-dimensional Sparse Attention at the Speed of Light
- Hexcute: A Tile-based Programming Language with Automatic Layout and Task-Mapping Synthesis
- High-performance training and inference for deep equivariant interatomic potentials
- Large models for machinery fault diagnosis: Current advances and future directions
- llm-jp-modernbert: A ModernBERT Model Trained on a Large-Scale Japanese Corpus with Long Context Length
- SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference
- Understanding the Skill Gap in Recurrent Language Models: The Role of the Gather-and-Aggregate Mechanism
- COBRA: Algorithm-Architecture Co-optimized Binary Transformer Accelerator for Edge Inference
- GeneMamba: An Efficient and Effective Foundation Model on Single Cell Data
- TeLLMe: An Energy-Efficient Ternary LLM Accelerator for Prefilling and Decoding on Edge FPGAs
- KeyDiff: Key Similarity-Based KV Cache Eviction for Long-Context LLM Inference in Resource-Constrained Environments
- Efficient Pretraining Length Scaling
- Hardware-based Heterogeneous Memory Management for Large Language Model Inference
- gLLM: Global Balanced Pipeline Parallelism System for Distributed LLM Serving with Token Throttling
- Think2SQL: Reinforce LLM Reasoning Capabilities for Text2SQL
- LLMSched: Uncertainty-Aware Workload Scheduling for Compound LLM Applications
- SlimPipe: Memory-Thrifty and Efficient Pipeline Parallelism for Long-Context LLM Training
- Learning to Attribute with Attention
- An All-Atom Generative Model for Designing Protein Complexes
- Antidistillation Sampling
- Frame Context Packing and Drift Prevention in Next-Frame-Prediction Video Diffusion Models
- Collaborative Learning of On-Device Small Model and Cloud-Based Large Model: Advances and Future Directions
- MOM: Memory-Efficient Offloaded Mini-Sequence Inference for Long Context Language Models
- A Review of YOLOv12: Attention-Based Enhancements vs. Previous Versions
- TacoDepth: Towards Efficient Radar-Camera Depth Estimation with One-stage Fusion
- Shared Disk KV Cache Management for Efficient Multi-Instance Inference in RAG-Powered LLMs
- Window Token Concatenation for Efficient Visual Large Language Models
- TurboFNO: High-Performance Fourier Neural Operator with Fused FFT-GEMM-iFFT on GPU
- Characterizing and Optimizing LLM Inference Workloads on CPU-GPU Coupled Architectures
- TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure
- PLoRA: An NDP-Enhanced Pooled-Memory System for Cost-Efficient Multi-LoRA Serving
- The Impossibility Triangle of Long-Context Modeling
- RIG-RoPE: Relation- and Instance-Gated Rotary Positional Encoding with Duration-Aware Temporal Coordinates
- QAMA: Scalable Quantum Annealing Multi-Head Attention Operator for Deep Learning
- VEXP: A Low-Cost RISC-V ISA Extension for Accelerated Softmax Computation in Transformers
- Summarization of Multimodal Presentations with Vision-Language Models: Study of the Effect of Modalities and Structure
- OVERLORD: Ultimate Scaling of DataLoader for Multi-Source Large Foundation Model Training
- Training LLMs on HPC Systems: Best Practices from the OpenGPT-X Project
- Towards Quantifying Commonsense Reasoning with Mechanistic Insights
- AlayaDB: The Data Foundation for Efficient and Effective Long-context LLM Inference
- VDocRAG: Retrieval-Augmented Generation over Visually-Rich Documents
- Efficient LLM Serving on Hybrid Real-time and Best-effort Requests
- FVQ: A Large-Scale Dataset and an LMM-based Method for Face Video Quality Assessment
- Position: Beyond Euclidean -- Foundation Models Should Embrace Non-Euclidean Geometries
- Particle Hit Clustering and Identification Using Point Set Transformers in Liquid Argon Time Projection Chambers
- PACT: Pruning and Clustering-Based Token Reduction for Faster Visual Language Models
- Token Level Routing Inference System for Edge Devices
- Pangu Ultra: Pushing the Limits of Dense Large Language Models on Ascend NPUs
- Kimi-VL Technical Report
- CHIME: A Compressive Framework for Holistic Interest Modeling
- Beyond Moore's Law: Harnessing the Redshift of Generative AI with Effective Hardware-Software Co-Design
- Crafting Query-Aware Selective Attention for Single Image Super-Resolution
- Distilling Textual Priors from LLM to Efficient Image Fusion
- Mosaic: Composite Projection Pruning for Resource-efficient LLMs
- Accelerating LLM Inference Throughput via Asynchronous KV Cache Prefetching
- SPIRe: Boosting LLM Inference Throughput with Speculative Decoding
- PaMi-VDPO: Mitigating Video Hallucinations by Prompt-Aware Multi-Instance Video Preference Learning
- One-Minute Video Generation with Test-Time Training
- Text-to-image personalization [wikipedia]
Discussions
- New AI/LLM Breakthrough - FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness [lemmy, 63 points, 4 comments]
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness [hn, 3 points, 1 comments]
- Flash Attention [hn, 2 points, 0 comments]
- Flash Attention accelerates GPT2 3x [hn, 2 points, 1 comments]
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness [hn, 1 points, 1 comments]
- 9/ Here is the paper https://buff.ly/3PdL8gp, I recommend you give it a read! [bsky, 1 points, 0 comments]
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness https://arxiv.org/pdf/2205.14135.pdf [bsky, 1 points, 0 comments]
- This article explores AI advancements and their industry transformation, stressing the importance of ethical use and responsible integration. https://arxiv.org/pdf/2205.14135.pdf [bsky, 0 points, 0 comments]
Related