FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
2022/05/27 by Tri Dao, Dao, Tri, Daniel Y. Fu +7 · 8 voices · 684 citations
Computer Science · #Advanced Neural Network Applications #Domain Adaptation and Few-Shot Learning #Machine Learning and Data Classification #cs.LG
paper · pdf · doi:10.48550/arxiv.2205.14135
openalex publication_date 2022/05/27 · openalex created_date 2022/06/13 · openalex updated_date 2026/07/28
Abstract
Machine learning's explosive growth has produced a proliferation of accelerator architectures, yet no framework maps a workload's computational character to the hardware that serves it best. We build that framework: decomposing ML workloads — transformers, CNNs, GNNs, diffusion and state-space models — into core compute primitives characterized by arithmetic intensity, memory-access pattern, and parallelism, and evaluating how GPUs, TPUs, systolic arrays, dataflow processors, wafer-scale engines, and neuromorphic chips serve them. Using hierarchical roofline and utilization modelling, we show the binding constraint has shifted from peak FLOPS to memory bandwidth and data-movement energy — which quantization only sharpens — and consolidate eight persistent, unsolved problems. These motivate the central contribution: the Fused Memory-Compute Tile (FMCT, codename Habanero), an SRAM-first, HBM-free accelerator combining a three-mode (GEMM / bandwidth / fused) phase-adaptive tile switched by compiler directive, a non-linear unit fused into the systolic datapath, compute-in-memory as a selectable mode rather than the whole architecture, and an inline KV-cache compression engine. A validated component-level energy model shows ~4-7× batch-1 efficiency gains over HBM GPUs. We position FMCT against its closest SRAM-first, no-HBM analogs — d-Matrix Corsair and Tenstorrent — and OpenAI's HBM-retaining Jalapeño ASIC, closing with principles for memory-centric, workload-adaptive accelerators.
Cited by
- DeCoRAG: Cognitive Decoupling and Semantic-Aware Cropping for Complex Document Understanding
- Moving Beyond Diversity: Visual Token Pruning as Subspace Reconstruction for Efficient VLMs
- SAFE-Pruner: Semantic Attention-Guided Future-Aware Token Pruning for Efficient Vision-Language-Action Manipulation
- RIS-Kernel: A Model-Agnostic Architecture for Long-Context LLM Inference via Sparse Attention
- IQ-JEPA: A Joint-Embedding Predictive Architecture with a Hermitian Vision Transformer for Sound Speed and Attenuation Estimation from Ultrasound IQ Data
- Flash EQ-Linear: Accelerating Equivariant Linear Layers via Group-wise Discrete Fourier Transform
- Eddy-VL 1.9B: Structural Pruning and Layered Distillation for Edge-Deployable Multimodal Embedding
- Trainable Log-linear Sparse Attention for Efficient Diffusion Transformers
- PReM: Learning What to Preserve and When to Refresh for Context Compression
- HijackKV: New Threat in Position-Independent KV Cache Reuse
- HeadCast: Casting Attention Heads for Efficient Autoregressive Video Generation
- Pixel-Space Diffusion Transformers
- ELSAA: Efficient Low-Rank and Sparse Attention Approximation for Training Transformers
- MoA-Structured Decode Attention DNF Derivation, KV-Cache Accumulation, GQA/MQA, and OpenACC Kernel
- Leaky Language Models: Stealing Architecture and Inference Optimizations via Per-Token Timing
- Timeripple: Accelerating vDiTs by Understanding the Spatio-Temporal Correlations in Latent Space
- GQLA: Group-Query Latent Attention for Hardware-Adaptive Large Language Model Decoding
- Matching Ranks Over Probability Yields Truly Deep Safety Alignment
- Harness Engineering for LLM-Driven GPU Kernel Generation
- Patch Policy: Efficient Embodied Control via Dense Visual Representations
- Capability from Access Structure, Not Scale: Lower Bounds and Pre-Registered Tests for Hybrid Sequence Models
- Vectorizing the Trie: Efficient Constrained Decoding for LLM-based Generative Retrieval on Accelerators
- Codec-Gauge: Learning Compression-Friendly Gauges for Transformer KV Caches
- Hardware Mechanisms to Dynamically Throttle AI Performance
- Sobek: Streaming Equivariant Tensor Product Convolutions
- WHALE: A Scalable Unified Model for Recommendation with Wukong-HSTU Architecture
- RoVE: Rotary Value Embeddings Attention for Relative Position-dependent Value Pathways
- Do Value Vectors in Deep Layers Need Context from the Residual Stream?
- Points as Tori: Fast Pointwise Signed Distance for Point Clouds
- Constraint-Driven Model Optimization: An Industry Framework for Selecting Compression and Acceleration Techniques in Modern Machine Learning Systems
- Long-Context Fine-Tuning with Limited VRAM
- Spectral Heat Flow for Conservative Token Condensation in Vision-Language Models
- RhinoVLA Technical Report
- DualKV: Shared-Prompt Flash Attention for Efficient RL Training with Large Rollouts and Long Contexts
- DSTAR: Accelerating Diffusion Transformers via Spatial and Temporal Redundancy Reduction
- FVAttn: Adaptive Sparse Attention with Runtime Load Balancing for Video Generation
- An MLIR-Based Compilation Method for Large Language Models
- RePlan: Reasoning-guided Region Planning for Complex Instruction-based Image Editing
- Attention to Detail: Evaluating Energy, Performance, and Accuracy Trade-offs Across vLLM Configurations
- HeteroMosaic: Exposing and Exploiting Heterogeneous Execution Opportunities for Energy-Efficient Edge LLM Inference
- ACID: Adaptive Caching for vIDeo generation
- Bifocal Attention: Harmonizing Geometric and Spectral Positional Embeddings for Algorithmic Generalization
- FlashDecoder: Real-Time Latent-to-Pixel Streaming Decoder with Transformers
- Reflex: Real-Time VLA Control through Streaming Inference
- Native Multi-Dimensional Subquadratic Operators via Input Dependent Long Convolutions
- High-accuracy Low-Bit KV-Cache Quantization via Local Distribution Restoration
- KernelBench-Verified: Do LLM-Generated Kernels Actually Beat PyTorch?
- Efficient and Training-Free Single-Image Diffusion Models
- The Devil is in the Spectrum: Mitigating Representation Collapse in LLMs via Topologically Regularized Side-Path
- Do Transformers Need Three Projections? Systematic Study of QKV Variants
- InferenceBench: A Benchmark for Open-Ended LLM Inference Optimization by AI Agents
- Multi-Stream LLMs: Unblocking Language Models with Parallel Streams of Thoughts, Inputs and Outputs
- JAXBench: Benchmarking Autonomous TPU Kernel Optimization
- Dropping the Anchor: Statistical Context Summarization for Distributed Systems via Pulsar Attention
- Attention to Mamba: A Recipe for Cross-Architecture Distillation
- PoM: A Linear-Time Replacement for Attention with the Polynomial Mixer
- M2RNN: Non-Linear RNNs with Matrix-Valued States for Scalable Language Modeling
- Flash-KMeans: Fast and Memory-Efficient Exact K-Means
- AutoKernel: Autonomous GPU Kernel Optimization via Iterative Agent-Driven Search
- Speculative Speculative Decoding
- Deep models of protein evolution in time generate realistic evolutionary trajectories and functional proteins
- Do LLMs Benefit From Their Own Words?
- Spelling Bee Embeddings for Language Modeling
- Even GPT-5.2 Can't Count to Five: The Case for Zero-Error Horizons in Trustworthy LLMs
- mHC: Manifold-Constrained Hyper-Connections
- Dynamic Large Concept Models: Latent Reasoning in an Adaptive Semantic Space
- PHOTON: Hierarchical Autoregressive Modeling for Lightspeed and Memory-Efficient Language Generation
- Extending the Context of Pretrained LLMs by Dropping Their Positional Embeddings
- Jasmine: A Simple, Performant and Scalable JAX-based World Modeling Codebase
- Kimi Linear: An Expressive, Efficient Attention Architecture
- Expected Attention: KV Cache Compression by Estimating Attention from Future Queries Distribution
- Fast and Simplex: 2-Simplicial Attention in Triton
- JAFAR: Jack up Any Feature at Any Resolution
- Protriever: End-to-End Differentiable Protein Homology Search for Fitness Prediction
- QiMeng: Fully Automated Hardware and Software Design for Processor Chip
- Leveraging Natural Language Processing to Unravel the Mystery of Life: A Review of NLP Approaches in Genomics, Transcriptomics, and Proteomics
- AbsenceBench: Language Models Can't Tell What's Missing
- Log-Linear Attention
- Flash Invariant Point Attention
- ModernBERT or DeBERTaV3? Examining Architecture and Data Influence on Transformer Encoder Models Performance
- 70% Size, 100% Accuracy: Lossless LLM Compression for Efficient GPU Inference via Dynamic-Length Float (DFloat11)
- TransMamba: A Sequence-Level Hybrid Transformer-Mamba Language Model
- ZSMerge: Zero-Shot KV Cache Compression for Memory-Efficient Long-Context LLMs
- SPECTRE: An FFT-Based Efficient Drop-In Replacement to Self-Attention for Long Contexts
- InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU
- Extending the RANGE of Graph Neural Networks: Relaying Attention Nodes for Global Encoding
- Block Diffusion: Interpolating Between Autoregressive and Diffusion Language Models
- Scaling Laws in Patchification: An Image Is Worth 50,176 Tokens And More
- Fast3R: Towards 3D Reconstruction of 1000+ Images in One Forward Pass
- Accelerating Time Series Foundation Models with Speculative Decoding
- MXAttention: Data-Free Optimal Scaling and Pre-Normalization Quantization for MXFP4 Attention
- AKG kernel Agent: A Multi-Agent Framework for Cross-Platform Kernel Synthesis
- Deep learning for pedestrians: backpropagation in Transformers
- Trust Region Masking for Long-Horizon LLM Reinforcement Learning
- Breaking the Memory Wall: Exact Analytical Differentiation via Tiled Operator-Space Evolution
- Viability and Performance of a Private LLM Server for SMBs: A Benchmark Analysis of Qwen3-30B on Consumer-Grade Hardware
- Fleet: Hierarchical Task-based Abstraction for Megakernels on Multi-Die GPUs
- Chessformer: A Unified Architecture for Chess Modeling
- Bridging Global Intent with Local Details: A Hierarchical Representation Approach for Semantic Validation in Text-to-SQL
- WeDLM: Reconciling Diffusion Language Models with Standard Causal Attention for Fast Inference
- Learning When Not to Attend Globally
- Role-Based Fault Tolerance System for LLM RL Post-Training
- DiRL: An Efficient Post-Training Framework for Diffusion Language Models
- MatKV: Trading Compute for Flash Storage in LLM Inference
- DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention
- RayRoPE: Projective Ray Positional Encoding for Multi-view Attention
- Compressing Observation History into Agent Memory: Distilling Transformers into Recurrent Transformers
- Omni-Prune: Query-Aware Unified Token Pruning for Efficient Omnimodal Large Language Models
- MobileWan: Closing the Quality Gap for Mobile Video Diffusion
- LiteTopK: Exploiting the Curse of Dimensionality for a Fused Indexer-TopK Kernel in Long-Context Sparse Attention
- Express Language Modeling
- StreamIndex: Memory-Bounded Compressed Sparse Attention via Streaming Top-k
- Full event interpretation with machine-learning-based particle-flow reconstruction in the CMS detector
- Accelerating Language Model Workflows with Prompt Choreography
- Anchor Attention, Small Cache: Code Generation with Large Language Models
- Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification
- KAP: Bridging the Knowledge Selection-Runtime Consumption Gap in LLM Systems
- X-Stage: An Overlooked Pipeline Stage for Communication-Computation Overlap in DiT Inference
- WHTMix: Efficient Stereo Depth Estimation via Walsh-Hadamard Token Mixing
- Salient Knowledge Pathways: Sparse Cross-Modal Routing for Efficient Knowledge-Intensive Multimodal Question Answering
- CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention
- Intrinsic and Triangulation-Agnostic Attention: A Simple and Powerful Approach for Learning on Meshes
- FusionML: Prefill, Not Decode - Mechanism and Boundaries of CPU+GPU Co-Execution on Unified-Memory Apple Silicon
- SeedFold: Scaling Biomolecular Structure Prediction
- Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs
- Agent Team Work Zone: An Automated, Persistent Workspace for Long-Lived Claude Code Agent Teams
- Visible-Light Imaging Diagnosis of Neutral Particle Emission Tomography in the Tokamak Divertor: An Efficient Transformer-based Surrogate Model
- Bumblebee: Interleaved Mixed-Layer Building Blocks for Large-Scale Recommendation Systems
- xMIx: High-Performance Serving-Time Platform for Mechanistic Interpretability Apps
- Kernel Forge: An Agent Harness for LLM-based Generation and Optimization of CUDA Kernels
- FlashEvaluator: Expanding Search Space with Parallel Sequence-Level Evaluation
- SkipOPU: An FPGA-based Overlay Processor for Large Language Models with Dynamically Allocated Computation
- Plain Transformers are Surprisingly Powerful Link Predictors
- DPAR: Dynamic Patchification for Efficient Autoregressive Visual Generation
- EasyOmnimatte: Taming Pretrained Inpainting Diffusion Models for End-to-End Video Layered Decomposition
- Scene-VLM: Multimodal Video Scene Segmentation via Vision-Language Models
- Analyzing the Mechanism of Attention Collapse in VGGT from a Dynamics Perspective
- Mesh-Attention: A New Communication-Efficient Distributed Attention with Improved Data Locality
- FlashVLM: Text-Guided Visual Token Selection for Large Multimodal Models
- Designing Spatial Architectures for Sparse Attention: STAR Accelerator via Cross-Stage Tiling
- ActionFlow: A Pipelined Action Acceleration for Vision Language Models on Edge
- Predictive-LoRA: A Proactive and Fragmentation-Aware Serverless Inference System for LLMs
- Memory as Resonance: A Biomimetic Architecture for Infinite Context Memory on Ergodic Phonetic Manifolds
- RAPID-LLM: Resilience-Aware Performance analysis of Infrastructure for Distributed LLM Training and Inference
- CascadeInfer: Length-Aware Scheduling of LLM Serving with Low Latency and Load Balancing
- In-Context Audio Control of Video Diffusion Transformers
- A Multi-agent Text2SQL Framework using Small Language Models and Execution Feedback
- Optimal Software Pipelining and Warp Specialization for Tensor Core GPUs
- KV Admission: Learning What to Write for Efficient Long-Context Inference
- LoPA: Scaling dLLM Inference via Lookahead Parallel Decoding
- MRG-R1: Reinforcement Learning for Clinically Aligned Medical Report Generation
- INTELLECT-3: Technical Report
- From Theory to Throughput: CUDA-Optimized APML for Large-Batch 3D Learning
- Dynamic Rebatching for Efficient Early-Exit Inference with DREX
- End-to-End Training for Autoregressive Video Diffusion via Self-Resampling
- LLMQ: Efficient Lower-Precision Pretraining for Consumer GPUs
- Dual-Density Inference for Efficient Language Model Reasoning
- Mixture of Attention Schemes (MoAS): Learning to Route Between MHA, GQA, and MQA
- Towards Nepali-language LLMs: Efficient GPT training with a Nepali BPE tokenizer
- TEMP: A Memory Efficient Physical-aware Tensor Partition-Mapping Framework on Wafer-scale Chips
- Ladder Up, Memory Down: Low-Cost Fine-Tuning With Side Nets
- BLURR: A Boosted Low-Resource Inference for Vision-Language-Action Models
- A Unified Sparse Attention via Multi-Granularity Compression
- VajraV1 -- The most accurate Real Time Object Detector of the YOLO family
- LitePT: Lighter Yet Stronger Point Transformer
- Improving Recursive Transformers with Mixture of LoRAs
- Soul: Breathe Life into Digital Human for High-fidelity Long-term Multimodal Animation
- Transform Trained Transformer: Accelerating Naive 4K Video Generation Over 10×
- Non-Resolution Reasoning (NRR): A Computational Framework for Contextual Identity and Ambiguity Preservation
- SneakPeek: Future-Guided Instructional Streaming Video Generation
- Investigating Data Pruning for Pretraining Biological Foundation Models at Scale
- ProServe: Unified Multi-Priority Request Scheduling for LLM Serving
- Cross-Modal Representational Knowledge Distillation for Enhanced Spike-Informed LFP Modeling
- Exact Flow Linear Attention: Exact Solution from Continuous-Time Dynamics
- StreamingAssistant: Efficient Visual Token Pruning for Accelerating Online Video Understanding
- TA-KAND: Two-stage Attention Triple Enhancement and U-KAN based Diffusion For Few-shot Knowledge Graph Completion
- BOOST: BOttleneck-Optimized Scalable Training Framework for Low-Rank Large Language Models
- BLASST: Dynamic BLocked Attention Sparsity via Softmax Thresholding
- xGR: Efficient Generative Recommendation Serving at Scale
- Adaptive Soft Rolling KV Freeze with Entropy-Guided Recovery: Sublinear Memory Growth for Efficient LLM Inference
- Mining Legal Arguments to Study Judicial Formalism
- Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap
- FUSER: Feed-Forward MUltiview 3D Registration Transformer and SE(3)N Diffusion Refinement
- RACAM: Enhancing DRAM with Reuse-Aware Computation and Automated Mapping for ML Inference
- TritonForge: Profiling-Guided Framework for Automated Triton Kernel Optimization
- Should AI Become an Intergenerational Civil Right?
- LaMoSys3.5D: Enabling 3.5D-IC-Based Large Language Model Inference Serving Systems via Hardware/Software Co-Design
- Towards Lossless Ultimate Vision Token Compression for VLMs
- Chopper: A Multi-Level GPU Characterization Tool & Derived Insights Into LLM Training Inefficiency
- HybridToken-VLM: Hybrid Token Compression for Vision-Language Models
- MobileFineTuner: A Unified End-to-End Framework for Fine-Tuning LLMs on Mobile Phones
- GatedFWA: Linear Flash Windowed Attention with Gated Associative Memory
- A Mathematical Theory of Top-k Sparse Attention via Total Variation Distance
- Flash Multi-Head Feed-Forward Network
- Persian-Phi: Efficient Cross-Lingual Adaptation of Compact LLMs via Curriculum Learning
- DCO: Dynamic Cache Orchestration for LLM Accelerators through Predictive Management
- Materium: An Autoregressive Approach for Material Generation
- Block Sparse Flash Attention
- RLAX: Large-Scale, Distributed Reinforcement Learning for Large Language Models on TPUs
- ReCAD: Reinforcement Learning Enhanced Parametric CAD Model Generation with Vision-Language Models
- Hierarchical geometric deep learning enables scalable analysis of molecular dynamics
- EgoEdit: Dataset, Real-Time Streaming Model, and Benchmark for Egocentric Video Editing
- ShaRP: SHAllow-LayeR Pruning for Video Large Language Models Acceleration
- StreamFlow: Theory, Algorithm, and Implementation for High-Efficiency Rectified Flow Generation
- Parabolic Position Encoding: Vision-Centric, Principled, Extrapolatable, General
- State Space Models for Bioacoustics: A comparative Evaluation with Transformers
- Reconstructing KV Caches with Cross-layer Fusion For Enhanced Transformers
- A Preliminary Study on the Promises and Challenges of Native Top-k Sparse Attention
- Agentic Operator Generation for ML ASICs
- PSA: Pyramid Sparse Attention for Efficient Video Understanding and Generation
- AutoBrep: Autoregressive B-Rep Generation with Unified Topology and Geometry
- GeoZero: Incentivizing Reasoning from Scratch on Geospatial Scenes
- Fairy2i: Training Complex LLMs from Real LLMs with All Parameters in \± 1, ± i\
- When Refusals Fail: Unstable Safety Mechanisms in Long-Context LLM Agents
- KV Pareto: Systems-Level Optimization of KV Cache and Model Compression for Long Context Inference
- SVRG and Beyond via Posterior Correction
- A Systematic Characterization of LLM Inference on GPUs
- FlashVGGT: Efficient and Scalable Visual Geometry Transformers with Compressed Descriptor Attention
- VLASH: Real-Time VLAs via Future-State-Aware Asynchronous Inference
- Accelerating Streaming Video Large Language Models via Hierarchical Token Compression
- Feed-Forward 3D Gaussian Splatting Compression with Long-Context Modeling
- SIMPLE: Disaggregating Sampling from GPU Inference into a Decision Plane for Faster Distributed LLM Serving
- InvarDiff: Cross-Scale Invariance Caching for Accelerated Diffusion Models
- Serving Heterogeneous LoRA Adapters in Distributed LLM Inference Systems
- Energy-Efficient Vision Transformer Inference for Edge-AI Deployment
- GSPN-2: Efficient Parallel Sequence Modeling
- Ovis-Image Technical Report
- Behavior-Equivalent Token: Single-Token Replacement for Long Prompts in LLMs
- InstanceV: Instance-Level Video Generation
- OOCO: Latency-disaggregated Architecture for Online-Offline Co-locate LLM Serving
- PIBNet: a Physics-Inspired Boundary Network for Multiple Scattering Simulations
- IntAttention: A Fully Integer Attention Pipeline for Efficient Edge Inference
- Frequency-Aware Token Reduction for Efficient Vision Transformer
- Subjective Depth and Timescale Transformers: Learning Where and When to Compute
- MeanFlow Transformers with Representation Autoencoders
- Generating Separated Singing Vocals Using a Diffusion Model Conditioned on Music Mixtures
- DOPD: A Dynamic PD-Disaggregation Architecture for Maximizing Goodput in LLM Inference Serving
- LocateAnything3D: Vision-Language 3D Detection with Chain-of-Sight
- Adam Simplified: Bias Correction Debunked
- FREE: Uncertainty-Aware Autoregression for Parallel Diffusion Transformers
- Scaling LLM Speculative Decoding: Non-Autoregressive Forecasting in Large-Batch Scenarios
- UltraViCo: Breaking Extrapolation Limits in Video Diffusion Transformers
- QiMeng-Kernel: Macro-Thinking Micro-Coding Paradigm for LLM-Based High-Performance GPU Kernel Generation
- HKRAG: Holistic Knowledge Retrieval-Augmented Generation Over Visually-Rich Documents
- Tell Model Where to Look: Mitigating Hallucinations in MLLMs by Vision-Guided Attention
- Terminal Velocity Matching
- MapFormer: Self-Supervised Learning of Cognitive Maps with Input-Dependent Positional Embeddings
- EventSTU: Event-Guided Efficient Spatio-Temporal Understanding for Video Large Language Models
- Nemotron-Flash: Towards Latency-Optimal Hybrid Small Language Models
- Parallel Vision Token Scheduling for Fast and Accurate Multimodal LMMs Inference
- VLM in a flash: I/O-Efficient Sparsification of Vision-Language Model via Neuron Chunking
- NEZHA: A Zero-sacrifice and Hyperspeed Decoding Architecture for Generative Recommendations
- Kitty: Accurate and Efficient 2-bit KV Cache Quantization with Dynamic Channel-wise Precision Boost
- Deterministic Inference across Tensor Parallel Sizes That Eliminates Training-Inference Mismatch
- Selective Rotary Position Embedding
- DeepCoT: Deep Continual Transformers for Real-Time Inference on Data Streams
- MUCH: A Multilingual Claim Hallucination Benchmark
- Energy Scaling Laws for Diffusion Models: Quantifying Compute and Carbon Emissions in Image Generation
- Optimizing PyTorch Inference with LLM-Based Multi-Agent Systems
- Revisiting Multimodal KV Cache Compression: A Frequency-Domain-Guided Outlier-KV-Aware Approach
- Neo: Real-Time On-Device 3D Gaussian Splatting with Reuse-and-Update Sorting Acceleration
- Hemlet: A Heterogeneous Compute-in-Memory Chiplet Architecture for Vision Transformers with Group-Level Parallelism
- Reasoning in Diffusion Large Language Models is Concentrated in Dynamic Confusion Zones
- A Comprehensive Study on Visual Token Redundancy for Discrete Diffusion-based Multimodal Large Language Models
- OmniZip: Audio-Guided Dynamic Token Compression for Fast Omnimodal Large Language Models
- From Projection to Prediction: Beyond Logits for Scalable Language Models
- MergeDNA: Context-aware Genome Modeling with Dynamic Tokenization through Token Merging
- TabFlash: Efficient Table Understanding with Progressive Question Conditioning and Token Focusing
- MACKO: Sparse Matrix-Vector Multiplication for Low Sparsity
- EchoLSTM: A Self-Reflective Recurrent Network for Stabilizing Long-Range Memory
- Equivalence Checking of ML GPU Kernels
- Evaluating Latent Generative Paradigms for High-Fidelity 3D Shape Completion from a Single Depth Image
- Striking the Right Balance between Compute and Copy: Improving LLM Inferencing Under Speculative Decoding
- Arcee: Differentiable Recurrent State Chain for Generative Vision Modeling with Mamba SSMs
- DocSLM: A Small Vision-Language Model for Long Multimodal Document Understanding
- LEMUR: Large scale End-to-end MUltimodal Recommendation
- Seeing the Forest and the Trees: Query-Aware Tokenizer for Long-Video Multimodal Language Models
- Optimizing Mixture of Block Attention
- Faster Algorithms for Structured Matrix Multiplication via Flip Graph Search
- Speculative Decoding in Decentralized LLM Inference: Turning Communication Latency into Computation Throughput
- Retrieval as a Decision: Training-Free Adaptive Gating for Efficient RAG
- Making Every Head Count: Sparse Attention Without the Speed-Performance Trade-off
- Branching Flows: Discrete, Continuous, and Manifold Flow Matching with Splits and Deletions
- Think-at-Hard: Selective Latent Iterations to Improve Reasoning Language Models
- Moebius: 0.2B Lightweight Image Inpainting Framework with 10B-Level Performance
- DPRM: A Dual Implicit Process Reward Model in Multi-Hop Question Answering
- Generalizable Insights for Graph Transformers in Theory and Practice
- Sharp Eyes and Memory for VideoLLMs: Information-Aware Visual Token Pruning for Efficient and Reliable VideoLLM Reasoning
- Schedulers for Schedule-free: Theoretically inspired hyperparameters
- A Circular Argument : Does RoPE need to be Equivariant for Vision?
- StreamDiffusionV2: A Streaming System for Dynamic and Interactive Video Generation
- TNT: Improving Chunkwise Training for Test-Time Memorization
- Learning to Focus: Focal Attention for Selective and Scalable Transformers
- You Had One Job: Per-Task Quantization Using LLMs' Hidden Representations
- PRAGMA: A Profiling-Reasoned Multi-Agent Framework for Automatic Kernel Optimization
- We Can Hear You with mmWave Radar! An End-to-End Eavesdropping System
- Optimizing Long-context LLM Serving via Fine-grained Sequence Parallelism
- MoSKA: Mixture of Shared KV Attention for Efficient Long-Sequence LLM Inference
- Make It Long, Keep It Fast: End-to-End 10k-Sequence Modeling at Billion Scale on Douyin
- Lethe: Layer- and Time-Adaptive KV Cache Pruning for Reasoning-Intensive LLM Serving
- Hilbert-Guided Sparse Local Attention
- OckBench: Measuring the Efficiency of LLM Reasoning
- Rethinking Metrics and Diffusion Architecture for 3D Point Cloud Generation
- LoPT: Lossless Parallel Tokenization Acceleration for Long Context Inference of Large Language Model
- BudgetMem: Learning Selective Memory Policies for Cost-Efficient Long-Context Processing in Language Models
- DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing
- FuseFlow: A Fusion-Centric Compilation Framework for Sparse Deep Learning on Streaming Dataflow
- Cambrian-S: Towards Spatial Supersensing in Video
- PICNIC: Silicon Photonic Interconnected Chiplets with Computational Network and In-memory Computing for LLM Inference Acceleration
- PETRA: Pretrained Evolutionary Transformer for SARS-CoV-2 Mutation Prediction
- Temporal Zoom Networks: Distance Regression and Continuous Depth for Efficient Action Localization
- NAPS: Attention-Based Fusion of Heterogeneous Physiological Signals
- SnapStream: Efficient Long Sequence Decoding on Dataflow Accelerators
- Divide, Cache, Conquer: Dichotomic Prompting for Efficient Multi-Label LLM-Based Classification
- From Models to Operators: Rethinking Autoscaling Granularity for Large Generative Models
- Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live
- Optimizing Attention on GPUs by Exploiting GPU Architectural NUMA Effects
- Flashlight: PyTorch Compiler Extensions to Accelerate Attention Variants
- Assessing LLM Reasoning Steps via Principal Knowledge Grounding
- GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding
- FlashEVA: Accelerating LLM inference via Efficient Attention
- Reasoning Planning for Language Models
- SpecAttn: Speculating Sparse Attention
- H-FA: A Hybrid Floating-Point and Logarithmic Approach to Hardware Accelerated FlashAttention
- Running VLAs at Real-time Speed
- OneTrans: Unified Feature Interaction and Sequence Modeling with One Transformer in Industrial Recommender
- Gaperon: A Peppered English-French Generative Language Model Suite
- PureKV: Plug-and-Play KV Cache Optimization with Spatial-Temporal Sparse Attention for Vision-Language Large Models
- LLMET: Enabling Cross-Layer Evaluation of Emerging M3D Memories for Energy-Efficient LLM Serving
- PSG: Pair-Space Generation for Efficient Generative Reranking
- NELSSA: A GPU-PNM Heterogeneous System for Mixed-Length LLM Serving via Length-based Request Placement
- From Tokens to Watt-hours: Analytical Energy Estimation for LLM Inference on Modern GPUs
- Steering Instruction Hierarchies at Inference Time
- Mixture-of-Depths Attention
- Don't Just Chase "Highlighted Tokens" in MLLMs: Revisiting Visual Holistic Context Retention
- A Compositional Theory of Causally Masked Transformers
- CuTe Layout Representation and Algebra
- ThriftAttention: Selective Mixed Precision for Long-Context FP4 Attention
- Zero-shot de novo peptide sequencing with open posttranslational modification discovery
- Simplified Sparse Attention via Gist Tokens
- VisionSelector: End-to-End Learnable Visual Token Compression for Efficient Multimodal LLMs
- Fine-Tuning GPT-5 for GPU Kernel Generation
- HySparse: A Hybrid Sparse Attention Architecture with Oracle Token Selection and KV Cache Sharing
- Group Relative Attention Guidance for Image Editing
- Improving LLM Reasoning via Dependency-Aware Query Decomposition and Logic-Parallel Content Expansion
- Long-Context Modeling with Dynamic Hierarchical Sparse Attention for On-Device LLMs
- UltraImageGen: Efficient Ultra-High-Resolution Image Generation with Hierarchical Local Attention
- APTBench: Benchmarking Agentic Potential of Base LLMs During Pre-Training
- ETC: training-free diffusion models acceleration with Error-aware Trend Consistency
- Beyond Line-Level Filtering for the Pretraining Corpora of LLMs
- Pie: A Programmable Serving System for Emerging LLM Applications
- Key and Value Weights Are Probably All You Need: On the Necessity of the Query, Key, Value weight Triplet in Decoder-Only Transformers
- A Survey on Efficient Vision-Language-Action Models
- Knocking-Heads Attention
- Can Language Models Compose Skills In-Context?
- Rethinking Inference Placement for Deep Learning across Edge and Cloud Platforms: A Multi-Objective Optimization Perspective and Future Directions
- Massive Memorization with Hundreds of Trillions of Parameters for Sequential Transducer Generative Recommenders
- Beyond Reasoning Gains: Mitigating General Capabilities Forgetting in Large Reasoning Models
- Transformer Based Linear Attention with Optimized GPU Kernel Implementation
- A Unified Model for Multi-Task Drone Routing in Post-Disaster Road Assessment
- Sparser Block-Sparse Attention via Token Permutation
- Correlation Dimension of Auto-Regressive Large Language Models
- Mitra: Mixed Synthetic Priors for Enhancing Tabular Foundation Models
- CudaForge: An Agent Framework with Hardware Feedback for CUDA Kernel Optimization
- REx86: A Local Large Language Model for Assisting in x86 Assembly Reverse Engineering
- HoloCine: Holistic Generation of Cinematic Multi-Shot Long Video Narratives
- Alleviating Forgetfulness of Linear Attention by Hybrid Sparse Attention and Contextualized Learnable Token Eviction
- HA-RAG: Hotness-Aware RAG Acceleration via Mixed Precision and Data Placement
- PPMStereo: Pick-and-Play Memory Construction for Consistent Dynamic Stereo Matching
- Rotate Both Ways: Time-and-Order RoPE for Generative Recommendation
- Seed3D 1.0: From Images to High-Fidelity Simulation-Ready 3D Assets
- A Survey on Cache Methods in Diffusion Models: Toward Efficient Multi-Modal Generation
- Stream: Scaling up Mechanistic Interpretability to Long Context in LLMs via Sparse Attention
- Every Attention Matters: An Efficient Hybrid Architecture for Long-Context Reasoning
- Reasoning Language Model Inference Serving Unveiled: An Empirical Study
- EdgeReasoning: Characterizing Reasoning LLM Deployment on Edge GPUs
- ssToken: Self-modulated and Semantic-aware Token Selection for LLM Fine-tuning
- Unifying and Enhancing Graph Transformers via a Hierarchical Mask Framework
- MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training
- Accelerating Vision Transformers with Adaptive Patch Sizes
- Unbiased Gradient Low-Rank Projection
- This is Going to Sound Crazy, But What If We Used Large Language Models to Boost Automatic Database Tuning Algorithms By Leveraging Prior History? We Will Find Better Configurations More Quickly Than Retraining From Scratch!
- Enhanced Motion Forecasting with Plug-and-Play Multimodal Large Language Models
- Efficient Long-context Language Model Training by Core Attention Disaggregation
- Efficient Vision-Language-Action Models for Embodied Manipulation: A Systematic Survey
- Finding Manifolds With Bilinear Autoencoders
- All You Need is One: Capsule Prompt Tuning with a Single Vector
- Improving Model Representation and Reducing KV Cache via Skip Connections with First Value Heads
- Predicting Task Performance with Context-aware Scaling Laws
- To Infinity and Beyond: Tool-Use Unlocks Length Generalization in State Space Models
- xLLM Technical Report
- First Attentions Last: Better Exploiting First Attentions for Efficient Transformer Training
- State-Space Models for Tabular Prior-Data Fitted Networks
- End-to-End Multi-Modal Diffusion Mamba
- BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Disaggregated LLM Serving in AI Infrastructure
- GatePro: Parameter-Free Expert Selection Optimization for Mixture-of-Experts Models
- Taming the Fragility of KV Cache Eviction in LLM Inference
- Trace Anything: Representing Any Video in 4D via Trajectory Fields
- Attention Illuminates LLM Reasoning: The Preplan-and-Anchor Rhythm Enables Fine-Grained Policy Optimization
- Chinese ModernBERT with Whole-Word Masking
- Litespark Technical Report: High-Throughput, Energy-Efficient LLM Training Framework
- FlashVSR: Towards Real-Time Diffusion-Based Streaming Video Super-Resolution
- APCE: Adaptive Progressive Context Expansion for Long Context Processing
- PAINT: Parallel-in-time Neural Twins for Dynamical System Reconstruction
- Task-Aware Reduction for Scalable LLM-Database Systems
- HoMer: Addressing Heterogeneities by Modeling Sequential and Set-wise Contexts for CTR Prediction
- Protenix-Mini+: efficient structure prediction model with scalable pairformer
- KVComm: Enabling Efficient LLM Communication through Selective KV Sharing
- DCP: Addressing Input Dynamism In Long-Context Training via Dynamic Context Parallelism
- PVDetector: Detecting Prompt Injection Attacks on Purpose-Specific LLM Agents through Policy-Violation Concept Analysis
- GPU-Tile-Sim: A Tile-Centric GPU Simulation Framework for LLM Hardware-Software Co-Design
- AutoMegaKernel: A Statically-Checked Agent Harness for Self-Retargeting Megakernel Synthesis
- AURA: Action-Gated Memory for Robot Policies at Constant VRAM
- Sparse Query Attention (SQA): A Computationally Efficient Attention Mechanism with Query Heads Reduction
- Attention Once Is All You Need: Efficient Streaming Inference with Stateful Transformers
- Accelerating Attention with Basis Decomposition
- On the Quantization Robustness of Diffusion Language Models in Coding Benchmarks
- SOL-ExecBench: Speed-of-Light Benchmarking for Real-World GPU Kernels Against Hardware Limits
- An Efficient Heterogeneous Co-Design for Fine-Tuning on a Single GPU
- PolyBlocks: A Compiler Infrastructure for AI Chips and Programming Frameworks
- Token Is All You Price
- MRMR: A Realistic and Expert-Level Multidisciplinary Benchmark for Reasoning-Intensive Multimodal Retrieval
- KORMo: Korean Open Reasoning Model for Everyone
- Neptune: Advanced ML Operator Fusion for Locality and Parallelism on GPUs
- SViM3D: Stable Video Material Diffusion for Single Image 3D Generation
- SPAD: Specialized Prefill and Decode Hardware for Disaggregated LLM Inference
- Beyond Sunk Costs: Boosting LLM Pre-training Efficiency via Orthogonal Growth of Mixture-of-Experts
- From Tokens to Layers: Redefining Stall-Free Scheduling for MoE Serving with Layered Prefill
- Haystack Engineering: Context Engineering for Heterogeneous and Agentic Long-Context Evaluation
- HTMformer: Hybrid Time and Multivariate Transformer for Time Series Forecasting
- Syncopate: Efficient Multi-GPU AI Kernels via Automatic Chunk-Centric Compute-Communication Overlap
- Executable Counterfactuals: Improving LLMs' Causal Reasoning Through Code
- Vectorized FlashAttention with Low-cost Exponential Computation in RISC-V Vector Processors
- Reusing Overtrained Language Models Saturates Scaling
- ConCuR: Conciseness Makes State-of-the-Art Kernel Generation
- AsyncSpade: Efficient Test-Time Scaling with Asynchronous Sparse Decoding
- Relational Transformer: Toward Zero-Shot Foundation Models for Relational Data
- VecInfer: Efficient LLM Inference with Low-Bit KV Cache via Outlier-Suppressed Vector Quantization
- lm-Meter: Unveiling Runtime Inference Latency for On-Device Language Models
- Shaken or Stirred? An Analysis of MetaFormer's Token Mixing for Medical Imaging
- The Anatomy of a Triton Attention Kernel
- Luth: Efficient French Specialization for Small Language Models and Cross-Lingual Transfer
- Staircase Streaming for Low-Latency Multi-Agent Inference
- SwiReasoning: Switch-Thinking in Latent and Explicit for Pareto-Superior Reasoning LLMs
- Transformers Discover Molecular Structure Without Graph Priors
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Hybrid Architectures for Language Models: Systematic Analysis and Design Insights
- Pushing on Multilingual Reasoning Models with Language-Mixed Chain-of-Thought
- Exact Causal Attention with 10% Fewer Operations
- RACE Attention: A Strictly Linear-Time Attention Layer for Training on Outrageously Large Contexts
- OneFlow: Concurrent Mixed-Modal and Interleaved Generation with Edit Flows
- Malice in Agentland: Down the Rabbit Hole of Backdoors in the AI Supply Chain
- Gather-Scatter Mamba: Accelerating Propagation with Efficient State Space Model
- Automated Structured Radiology Report Generation with Rich Clinical Context
- Efficient Multi-modal Large Language Models via Progressive Consistency Distillation
- Local Linear Attention: An Optimal Interpolation of Linear and Softmax Attention For Test-Time Regression
- Thoughtbubbles: an Unsupervised Method for Parallel Thinking in Latent Space
- LoRAFusion: Efficient LoRA Fine-Tuning for LLMs
- TASP: Topology-aware Sequence Parallelism
- Rearchitecting Datacenter Lifecycle for AI: A TCO-Driven Framework
- SlimPack: Fine-Grained Asymmetric Packing for Balanced and Efficient Variable-Length LLM Training
- Collaborative Compression for Large-Scale MoE Deployment on Edge
- TTT3R: 3D Reconstruction as Test-Time Training
- Context-Driven Performance Modeling for Causal Inference Operators on Neural Processing Units
- A Scalable Distributed Framework for Multimodal GigaVoxel Image Registration
- LVT: Large-Scale Scene Reconstruction via Local View Transformers
- Attention Surgery: An Efficient Recipe to Linearize Your Video Diffusion Transformer
- HAPT: Heterogeneity-Aware Automated Parallel Training on Heterogeneous Clusters
- ProxyAttn: Guided Sparse Attention via Representative Heads
- SparseServe: Unlocking Parallelism for Dynamic Sparse Attention in Long-Context LLM Serving
- Uni-X: Mitigating Modality Conflict with a Two-End-Separated Architecture for Unified Multimodal Models
- Hyperspherical Latents Improve Continuous-Token Autoregressive Generation
- Muon: Training and Trade-offs with Latent Attention and MoE
- Optimizing Privacy-Preserving Primitives to Support LLM-Scale Applications
- Brain Harmony: A Multimodal Foundation Model Unifying Morphology and Function into 1D Tokens
- MACE: A Hybrid LLM Serving System with Colocated SLO-aware Continuous Retraining Alignment
- SLA: Beyond Sparsity in Diffusion Transformers via Fine-Tunable Sparse-Linear Attention
- HiDe: Rethinking The Zoom-IN method in High Resolution MLLMs via Hierarchical Decoupling
- A Predictive and Synergistic Two-Layer Scheduling Framework for LLM Serving
- OracleGS: Grounding Generative Priors for Sparse-View Gaussian Splatting
- PDE-Transformer: A Continuous Dynamical Systems Approach to Sequence Modeling
- From Long to Lean: Performance-aware and Adaptive Chain-of-Thought Compression via Multi-round Refinement
- Reinforcement Learning-Guided Chain-of-Draft for Token-Efficient Code Generation
- Zeppelin: Balancing Variable-length Workloads in Data Parallel Large Model Training
- ChaosNexus: A Foundation Model for Universal Chaotic System Forecasting with Multi-scale Representations
- A Unifying Framework for Parallelizing Sequential Models with Linear Dynamical Systems
- Compute-Optimal Quantization-Aware Training
- OjaKV: Context-Aware Online Low-Rank KV Cache Compression with Oja's Rule
- Data-Centric Elastic Pipeline Parallelism for Efficient Long-Context LLM Training
- AIBA: Attention-based Instrument Band Alignment for Text-to-Audio Diffusion
- Real-Time Object Detection Meets DINOv3
- TyphoonMLA: A Mixed Naive-Absorb MLA Kernel For Shared Prefix
- RollPacker: Mitigating Long-Tail Rollouts for Fast, Synchronous RL Post-Training
- ARMesh: Autoregressive Mesh Generation via Next-Level-of-Detail Prediction
- Fine-Tuning LLMs to Analyze Multiple Dimensions of Code Review: A Maximum Entropy Regulated Long Chain-of-Thought Approach
- AMLA: MUL by ADD in FlashAttention Rescaling
- WEST: LLM based Speech Toolkit for Speech Understanding, Generation, and Interaction
- Enhancing Linear Attention with Residual Learning
- BurstEngine: an Efficient Distributed Framework for Training Transformers on Extremely Long Sequences of over 1M Tokens
- Gyges: Dynamic Cross-Instance Parallelism Transformation for Efficient LLM Inference
- RoboSSM: Scalable In-context Imitation Learning via State-Space Models
- Are We Scaling the Right Thing? A System Perspective on Test-Time Scaling
- scPlantAnnotate: an accurate and robust transformer-based model for plant cell type annotation
- Recall Before You Rank: Similarity-Guided Top-K Reuse for Efficient Long-Context Attention
- Back from the Future: Key-Value Cache Management by Counter-Causal Surprise
- SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer
- ReGenVC: End-to-End Real-Time Generative Video Coding at Ultra-Low Bitrate
- ROCS: Request-Oriented Compute Sharing for Efficient Large-Scale Recommendation
- Beyond KV Reconstruction: Functional Reconstruction for MLA Draft Models in Speculative Decoding
- FORGE: Fused On-Register Gradient Elimination for Memory-Efficient LLM Training
- Recursive transformers for semiconductor thermo-mechanical reliability
- KernelGenBench: A Multi-Source and Multi-Chip Benchmark for LLM-based Kernel Generation
- Multi-Head Attention Residuals
- Defusing the Trigger: Tail-Risk-Informed Attention Rebalancing for LLM Backdoor Mitigation
- Training Compute-Optimal Protein Language Models
- Gener <i>anno</i> : A Genomic Foundation Model for Metagenomic Annotation
- DAGE: Dual-Stream Architecture for Efficient and Fine-Grained Geometry Estimation
- Use What You Know: Causal Foundation Models with Partial Graphs
- Transformers and genome language models
- A Unified Transformer Architecture for Low-Latency and Scalable Wireless Signal Processing
- Protein Language Models: Is Scaling Necessary?
- CR-Net: Scaling Parameter-Efficient Training with Cross-Layer Low-Rank Structure
- LLM-based Agents Suffer from Hallucinations: A Survey of Taxonomy, Methods, and Directions
- Memory in Large Language Models: Mechanisms, Evaluation and Evolution
- An overview of neural architectures for self-supervised audio representation learning from masked spectrograms
- Reading Images Like Texts: Sequential Image Understanding in Vision-Language Models
- Visual Instruction Pretraining for Domain-Specific Foundation Models
- Probabilistic Token Alignment for Large Language Model Fusion
- STCast: Adaptive Boundary Alignment for Global and Regional Weather Forecasting
- FG-Attn: Leveraging Fine-Grained Sparse Attention in Video Diffusion Models
- Text-Scene: A Scene-to-Language Parsing Framework for 3D Scene Understanding
- Pico: A Modular Framework for Hypothesis-Driven Small Language Model Research
- HERO: Hierarchical Extrapolation and Refresh for Efficient World Models
- Efficient Pre-Training of LLMs via Topology-Aware Communication Alignment on More Than 9600 GPUs
- SuperGen: An Efficient Ultra-high-resolution Video Generation System with Sketching and Tiling
- Analyzing and Mitigating Surface Bias in Code Evaluation Metrics
- MedFact-R1: Towards Factual Medical Reasoning via Pseudo-Label Augmentation
- Evaluating the Limitations of Local LLMs in Solving Complex Programming Challenges
- LEAP: LLM Inference on Scalable PIM-NoC Architecture with Balanced Dataflow and Fine-Grained Parallelism
- SALT4Decompile: Inferring Source-level Abstract Logic Tree for LLM-Based Binary Decompilation
- Patent Language Model Pretraining with ModernBERT
- Value-Guided KV Compression for LLMs via Approximated CUR Decomposition
- TinySR: Pruning Diffusion for Real-World Image Super-Resolution
- Jim137/qkan: v0.1.0
- DiCache: Let Diffusion Model Determine Its Own Cache
- BWCache: Accelerating Video Diffusion Transformers through Block-Wise Caching
- FLAME: A Serving System Optimized for Large-Scale Generative Recommendation with Efficiency
- Re-purposing SAM into Efficient Visual Projectors for MLLM-Based Referring Image Segmentation
- Multi-Model Synthetic Training for Mission-Critical Small Language Models
- xOffense: An AI-driven autonomous penetration testing framework with offensive knowledge-enhanced LLMs and multi agent systems
- HPIM: Heterogeneous Processing-In-Memory-based Accelerator for Large Language Models Inference
- Reversible Deep Equilibrium Models
- Positional Encoding via Token-Aware Phase Attention
- ScaleDoc: Scaling LLM-based Predicates over Large Document Collections
- CoViPAL: Layer-wise Contextualized Visual Token Pruning for Large Vision-Language Models
- Large Language Models Imitate Logical Reasoning, but at what Cost?
- FastMTP: Accelerating LLM Inference with Enhanced Multi-Token Prediction
- ProtHyena: A fast and efficient foundation protein language model at single amino acid Resolution
- MindVL: Towards Efficient and Effective Training of Multimodal Large Language Models on Ascend NPUs
- SpeCa: Accelerating Diffusion Transformers with Speculative Feature Caching
- Genome-Factory: A Library for Tuning, Deploying, and Interpreting Genomic Foundation Models
- Multipole Semantic Attention: A Fast Approximation of Softmax Attention for Pretraining
- Characterizing the Efficiency of Distributed Training: A Power, Performance, and Thermal Perspective
- Compute Only 16 Tokens in One Timestep: Accelerating Diffusion Transformers with Cluster-Driven Feature Caching
- ButterflyQuant: Ultra-low-bit LLM Quantization through Learnable Orthogonal Butterfly Transforms
- Conditioning on PDE Parameters to Generalise Deep Learning Emulation of Stochastic and Chaotic Dynamics
- DiTReducio: A Training-Free Acceleration for DiT-Based TTS via Progressive Calibration
- Large Language Models in Document Intelligence: A Comprehensive Survey, Recent Advances, Challenges, and Future Trends
- Efficient Transformer-Based Piano Transcription With Sparse Attention Mechanisms
- DeepAries: Adaptive Rebalancing Interval Selection for Enhanced Portfolio Selection
- ViRanker: A BGE-M3 & Blockwise Parallel Transformer Cross-Encoder for Vietnamese Reranking
- LLM Architecture, Scaling Laws, and Economics: A Quick Summary
- Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism
- EvolKV: Evolutionary KV Cache Compression for LLM Inference
- A Survey of Long-Document Retrieval in the PLM and LLM Era
- Causal Attention with Lookahead Keys
- Block-Sparse Global Attention for Efficient Multi-View Geometry Transformers
- Dato: A Task-Based Programming Model for Dataflow Accelerators
- LatinX: Aligning a Multilingual TTS Model with Direct Preference Optimization
- Sticker-TTS: Learn to Utilize Historical Experience with a Sticker-driven Test-Time Scaling Framework
- Set Block Decoding is a Language Model Inference Accelerator
- SAC-MIL: Spatial-Aware Correlated Multiple Instance Learning for Histopathology Whole Slide Image Classification
- Differentiable Entropy Regularization: A Complexity-Aware Approach for Neural Optimization
- Scalable hybrid quantum Monte Carlo simulation of U(1) gauge field coupled to fermions on GPU
- Batch Query Processing and Optimization for Agentic Workflows
- AudioRWKV: Efficient and Stable Bidirectional RWKV for Audio Pattern Recognition
- Improving Large Vision and Language Models by Learning from a Panel of Peers
- Latent Space Single-Pixel Imaging Under Low-Sampling Conditions
- LobRA: Multi-tenant Fine-tuning over Heterogeneous Data
- GPT-OSS-20B: A Comprehensive Deployment-Centric Analysis of OpenAI's Open-Weight Mixture of Experts Model
- DSDE: Dynamic Speculative Decoding with KLD Stability for Real-World Serving
- Efficient Large Language Models with Zero-Shot Adjustable Acceleration
- Any-Order Flexible Length Masked Diffusion
- COMET: A Framework for Modeling Compound Operation Dataflows with Explicit Collectives
- Backdoor Samples Detection Based on Perturbation Discrepancy Consistency in Pre-trained Language Models
- Chunked TabPFN: Exact Training-Free In-Context Learning for Long-Context Tabular Data
- Balanced Actor Initialization: Stable RLHF Training of Distillation-Based Reasoning Models
- Accelerating Mixture-of-Experts Inference by Hiding Offloading Latency with Speculative Decoding
- Machine-learning based particle-flow algorithm in CMS
- Mixture of Contexts for Long Video Generation
- Provable Benefits of In-Tool Learning for Large Language Models
- Boltz-1 Democratizing Biomolecular Interaction Modeling
- OLMoASR: Open Models and Data for Training Robust Speech Recognition Models
- SwizzlePerf: Hardware-Aware LLMs for GPU Kernel Performance Optimization
- LMDeploy Accelerates Mixed-Precision LLM Inference with TurboMind
- Hits to Higgs: Reconstruction-Free Higgs Classification from Raw LHC Detector Data Using Higgsformers
- Conflict-Aware Soft Prompting for Retrieval-Augmented Generation
- Enhancing compact convolutional transformers with super attention
- ClusterFusion: Expanding Operator Fusion Scope for LLM Inference via Cluster-Level Collective Primitive
- Rethinking Caching for LLM Serving Systems: Beyond Traditional Heuristics
- Revisiting associative recall in modern recurrent models
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- Exploring Scaling Laws of CTR Model for Online Performance Improvement
- CARFT: Boosting LLM Reasoning via Contrastive Learning with Annotated Chain-of-Thought-based Reinforced Fine-Tuning
- Position Bias Mitigates Position Bias:Mitigate Position Bias Through Inter-Position Knowledge Distillation
- Lorentz-Equivariance without Limitations
- H2EAL: Hybrid-Bonding Architecture with Hybrid Sparse Attention for Efficient Long-Context LLM Inference
- MoEcho: Exploiting Side-Channel Attacks to Compromise User Privacy in Mixture-of-Experts LLMs
- Beyond Turing: Memory-Amortized Inference as a Foundation for Cognitive Computation
- Toward Better EHR Reasoning in LLMs: Reinforcement Learning with Expert Attention Guidance
- NovoMolGen: Rethinking Molecular Language Model Pretraining
- Prune2Drive: A Plug-and-Play Framework for Accelerating Vision-Language Models in Autonomous Driving
- Compact Attention: Exploiting Structured Spatio-Temporal Sparsity for Fast Video Generation
- FLARE: Fast Low-rank Attention Routing Engine
- CarelessWhisper: Turning Whisper into a Causal Streaming Model
- The Rise of Generative AI for Metal-Organic Framework Design and Synthesis
- E-CaTCH: Event-Centric Cross-Modal Attention with Temporal Consistency and Class-Imbalance Handling for Misinformation Detection
- Novel Parasitic Dual-Scale Modeling for Efficient and Accurate Multilingual Speech Translation
- Dataset Construction for Training LLM to Learn Analog Circuit Knowledge
- A learning-driven automatic planning framework for proton PBS treatments of H&N cancers
- BLADE: Block-Sparse Attention Meets Step Distillation for Efficient Video Generation
- Verify Distributed Deep Learning Model Implementation Refinement with Iterative Relation Inference
- Faster and Memory-Efficient Training of Sequential Recommendation Models for Large Catalogs
- READER: Retrieval-Assisted Drafter for Efficient LLM Inference
- Profiling Large Language Model Inference on Apple Silicon: A Quantization Perspective
- Retrospective Sparse Attention for Efficient Long-Context Generation
- Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions
- Exploring Multimodal Diffusion Transformers for Enhanced Prompt-based Image Editing
- CATP: Contextually Adaptive Token Pruning for Efficient and Enhanced Multimodal In-Context Learning
- GLiClass: Generalist Lightweight Model for Sequence Classification Tasks
- MobileViCLIP: An Efficient Video-Text Model for Mobile Devices
- DySK-Attn: A Framework for Efficient, Real-Time Knowledge Updating in Large Language Models via Dynamic Sparse Knowledge Attention
- The Fused Kernel Library: A C++ API to Develop Highly-Efficient GPU Libraries
- Multi-level Advantage Credit Assignment for Cooperative Multi-Agent Reinforcement Learning
- gpt-oss-120b & gpt-oss-20b Model Card
- Scalable Swin Transformer network for brain tumor segmentation from incomplete MRI modalities
- DeepFold-PLM: accelerating protein structure prediction via efficient homology search using protein language models
- Fewer Denoising Steps or Cheaper Per-Step Inference: Towards Compute-Optimal Diffusion Model Deployment
- AdaptInfer: Adaptive Token Pruning for Vision-Language Model Inference with Dynamical Text Guidance
- Crisp Attention: Regularizing Transformers via Structured Sparsity
- KnapFormer: An Online Load Balancer for Efficient Diffusion Transformers Training
- Estimating Musical Surprisal from Audio in Autoregressive Diffusion Model Noise Spaces
- Making Prompts First-Class Citizens for Adaptive LLM Pipelines
- iFairy: the First 2-bit Complex LLM with All Parameters in \±1, ± i\
- Why Attention Fails: A Taxonomy of Faults in Attention-Based Neural Networks
- HierarchicalPrune: Position-Aware Compression for Large-Scale Diffusion Models
- Share Your Attention: Transformer Weight Sharing via Matrix-based Dictionary Learning
- Efficient Inter-Task Attention for Multitask Transformer Models
- VLMQ: Efficient Post-Training Quantization for Large Vision-Language Models via Hessian Augmentation
- Macro-from-Micro Planning for High-Quality and Parallelized Autoregressive Long Video Generation
- Long Story Generation via Knowledge Graph and Literary Theory
- Following Route Instructions using Large Vision-Language Models: A Comparison between Low-level and Panoramic Action Spaces
- VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo
- Dynaword: From One-shot to Continuously Developed Datasets
- Forecasting When to Forecast: Accelerating Diffusion Models with Confidence-Gated Taylor
- Free-MoRef: Instantly Multiplexing Context Perception Capabilities of Video-MLLMs within Single Inference
- Trainable Dynamic Mask Sparse Attention
- A Glimpse to Compress: Dynamic Visual Token Pruning for Large Vision-Language Models
- FluidFormer: Transformer with Continuous Convolution for Particle-based Fluid Simulation
- Can3Tok: Canonical 3D Tokenization and Latent Modeling of Scene-Level 3D Gaussians
- PiKV: KV Cache Management System for Mixture of Experts
- FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models
- Optimal Scheduling Algorithms for LLM Inference: Theory and Practice
- HiPrune: Training-Free Visual Token Pruning via Hierarchical Attention in Vision-Language Models
- Sortblock: Similarity-Aware Feature Reuse for Diffusion Model
- Representation Shift: Unifying Token Compression with FlashAttention
- Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts
- Systematic Evaluation of Optimization Techniques for Long-Context Language Models
- Calibrated Language Models and How to Find Them with Label Smoothing
- Self-Foveate: Enhancing Diversity and Difficulty of Synthesized Instructions from Unsupervised Text via Multi-Level Foveation
- BAR Conjecture: the Feasibility of Inference Budget-Constrained LLM Services with Authenticity and Reasoning
- Your Spending Needs Attention: Modeling Financial Habits with Transformers
- trAIce3D: A Prompt-Driven Transformer Based U-Net for Semantic Segmentation of Microglial Cells from Large-Scale 3D Microscopy Images
- Text-to-image personalization [wikipedia]
Discussions
- New AI/LLM Breakthrough - FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness [lemmy, 63 points, 4 comments]
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness [hn, 3 points, 1 comments]
- Flash Attention [hn, 2 points, 0 comments]
- Flash Attention accelerates GPT2 3x [hn, 2 points, 1 comments]
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness [hn, 1 points, 1 comments]
- 9/ Here is the paper https://buff.ly/3PdL8gp, I recommend you give it a read! [bsky, 1 points, 0 comments]
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness https://arxiv.org/pdf/2205.14135.pdf [bsky, 1 points, 0 comments]
- This article explores AI advancements and their industry transformation, stressing the importance of ethical use and responsible integration. https://arxiv.org/pdf/2205.14135.pdf [bsky, 0 points, 0 comments]
Related