Mamba: Linear-Time Sequence Modeling with Selective State Spaces
2023/12/01 by Albert Gu, Tri Dao, Gu, Albert +1 · 13 voices · 885 citations
Computer Science · #Machine Learning and Algorithms #Neural Networks and Applications #Topic Modeling #cs.AI #cs.LG
paper · pdf · doi:10.48550/arxiv.2312.00752
openalex publication_date 2023/12/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/29
Abstract
Foundation models, now powering most of the exciting applications in deep learning, are almost universally based on the Transformer architecture and its core attention module. Many subquadratic-time architectures such as linear attention, gated convolution and recurrent models, and structured state space models (SSMs) have been developed to address Transformers' computational inefficiency on long sequences, but they have not performed as well as attention on important modalities such as language. We identify that a key weakness of such models is their inability to perform content-based reasoning, and make several improvements. First, simply letting the SSM parameters be functions of the input addresses their weakness with discrete modalities, allowing the model to selectively propagate or forget information along the sequence length dimension depending on the current token. Second, even though this change prevents the use of efficient convolutions, we design a hardware-aware parallel algorithm in recurrent mode. We integrate these selective SSMs into a simplified end-to-end neural network architecture without attention or even MLP blocks (Mamba). Mamba enjoys fast inference (5× higher throughput than Transformers) and linear scaling in sequence length, and its performance improves on real data up to million-length sequences. As a general sequence model backbone, Mamba achieves state-of-the-art performance across several modalities such as language, audio, and genomics. On language modeling, our Mamba-3B model outperforms Transformers of the same size and matches Transformers twice its size, both in pretraining and downstream evaluation.
Cited by
- LiMuon: Light and Fast Muon Optimizer for Large Models
- The Lift Spectrum: How Measurement-to-Space Adaptivity Shapes Robustness in Image-Free Single-Pixel Sensing
- CODA: Cascaded Online Discontinuity-Aware Alignment for Real-Time Image-Based Score Following
- Relaxed activation analysis of dataflow networks - A clock calculus for machine learning and real-time scheduling
- DM3D: Dynamic Mamba via Offset-Guided Feature Resampling for Point Cloud Understanding
- Pretraining Recurrent Networks without Recurrence
- Safe In-Context Reinforcement Learning
- IR275K: A Benchmark for Infrared Multi-Frame Super-Resolution Toward Efficient Remote Sensing
- Indexing: the Beginning and the End
- DCVC-MB: Neural B-Frame Video Compression using State Space Models
- The Second LoViF 2026 Challenge on Real-World All-in-One Image Restoration: Methods and Results
- CHERRY: Compressed Hierarchical Experts with Recurrent Representational Yield
- The Orthogonalized Read Is a Removable Training Scaffold for Recurrent Memory
- Physics-Aware Complex-Valued State Space Model with Scattering-Prior Feature Modulation for PolSAR Image Classification
- Schrödinger Bridge Mamba for One-Step Speech Enhancement
- Cumsum-Composable Phase Transport for Low-Cost Streaming Keyword Spotting
- SpikingMOT: A Spike-Driven Multi-Object Tracker
- Frequency-Hierarchical Active k-Space Sampling for Diagnostic MRI
- User-Centric Modeling of Transactional Sequences with Explainable State Space Models
- WorldPack: Dynamic Frame Compression for Long-context Video World Modeling
- SR-TTT Does Not Learn Retrieval: A Correction and Mechanistic Post-Mortem of Surprisal-Aware Residual Test-Time Training
- SPECTRA: State-Space Exogenous Context and Temporal-Frequency Resolution Architecture for Probabilistic Energy Forecasting
- CAPS: A Cascaded Reconstruction Model to Power Saving in Hearables Using Sub-Nyquist Sampling with Bandwidth Extension
- How Far Can Wearable-Compatible Signals Go? A Controlled Decomposition of Non-EEG Sleep Staging
- Causal dictionary learning reveals and validates transcription-factor binding features in genomic language models
- Stochastic Thermodynamics for Autoregressive Generative Models: A Non-Markovian Perspective
- Robust Betatron-Tune Measurement from Schottky Spectra: Complementary Classical and Deep-Learning Paradigms
- Team RAS in 11th ABAW Competition: Multimodal Ambivalence Recognition Approach
- Norm or Direction? Decoding Vision Mambas for High-Resolution Vision
- CS-ETS: Chaos-Inspired Samba-Based EMG-To-Speech Synthesis with Nonlinear Chaotic Losses
- Countercurrent Multiplier Networks: A Renal-Inspired Iterative Operator with Provably Bounded Fixed-Point Dynamics
- Selective State-Space Adaptation and Retrieval for Language Model Reasoning
- Mambalaya: Einsum-Based Fusion Optimizations on State-Space Models
- Adaptive Mamba Neural Operators
- The Aura in the Machine: Genealogy and the Status of the Work of Art in the Generative Era
- MambaLSTM: A Spatio-Temporal Framework for Enhanced Traffic Accident Risk Prediction
- The Art of Not Forgetting
- Capability from Access Structure, Not Scale: Lower Bounds and Pre-Registered Tests for Hybrid Sequence Models
- Kernelized Linear Attention: Breaking the Capacity Wall with Symmetric Cones
- The Truncation Blind Spot: How Decoding Strategies Systematically Exclude Human-Like Token Choices
- LLM-Guided Task-Semantic Field Factorization for Industrial Process Forecasting
- HyBDM: Multi-Scale Hybrid Experts for Time Series Forecasting with Bidirectional Dependency Modeling
- A Better Start for Language Models: Domain-Conditional Position Offsets
- RTL-Sequencer: Towards Scalable RTL Timing Prediction with the Sequence-based Paradigm
- SPEED: One-Step Pixel Diffusion for High-quality Video Frame Interpolation
- VideoSEMA: a scalable and efficient Mamba-like attention for video understanding
- ReCal3R: Reliability-Calibrated Learning Rates for Streaming 3D Reconstruction
- Language as a Wave Phenomenon: Semantic Phase Locking and Interference in Neural Networks
- A Modern Multimodal Assistant on a 6 GB 2011 GPU: Stage-Validated, All-GPU CUDA Inference for Fermi
- Robust Filter Attention: Self-Attention as Precision-Weighted State Estimation
- Sharp Stability Threshold and Certification for Designing Stable Residual Architectures
- A Minimal Interpretable Architecture for Zero-Shot Reconstruction of Dynamical Systems
- Can Tokens Compete? Token Representations against Supervised CNN Backbones for BirdCLEF+ 2026
- Autoregressive Guidance of Deep Spatially Selective Filters using Bayesian Tracking for Efficient Extraction of Moving Speakers
- xHC: Expanded Hyper-Connections
- Campaign Diagrams: Visualizing the March Through the Phases of a Workload
- Weakly-Supervised RGB-D Salient Object Detection via SAM-driven Pseudo Annotation and State Space Interaction-based Diffusion
- SEMA: a Scalable and Efficient Mamba like Attention via Token Localization and Averaging
- Teacher-Guided Causal Interventions for Image Denoising: Orthogonal Content-Noise Disentanglement in Vision Transformers
- Structured Latent Space Modeling over Multi-Scale Temporal Patches for Multivariate Time Series Forecasting
- Native Multi-Dimensional Subquadratic Operators via Input Dependent Long Convolutions
- AdaSurvMamba: Dynamic Fusion and Semantic Scanning for Multimodal Survival Analysis
- Spectral-LSH: Sub-Quadratic Prompt Compression via Krylov-Projected Locality-Sensitive Hashing
- The Devil is in the Spectrum: Mitigating Representation Collapse in LLMs via Topologically Regularized Side-Path
- Do Transformers Need Three Projections? Systematic Study of QKV Variants
- Stress-Testing EEG Foundation Models for Clinical Decoding: Dataset Identity and Targeted Negative Controls
- JAXBench: Benchmarking Autonomous TPU Kernel Optimization
- Attention to Mamba: A Recipe for Cross-Architecture Distillation
- Effective Distillation to Hybrid xLSTM Architectures
- Mamba-3: Improved Sequence Modeling using State Space Principles
- M2RNN: Non-Linear RNNs with Matrix-Valued States for Scalable Language Modeling
- Lost in Backpropagation: The LM Head is a Gradient Bottleneck
- Speculative Speculative Decoding
- Position: A Dynamical Systems Perspective is Needed to Advance Time Series Modeling
- Self-Attention at Constant Cost per Token via Symmetry-Aware Taylor Approximation
- Remapping and navigation of an embedding space via error minimization: a fundamental organizational principle of cognition in natural and artificial systems
- Innovative tooth segmentation using hierarchical features and bidirectional sequence modeling
- Kinaema: a recurrent sequence model for memory and pose in motion
- Kimi Linear: An Expressive, Efficient Attention Architecture
- Expected Attention: KV Cache Compression by Estimating Attention from Future Queries Distribution
- Transformers are Inherently Succinct
- Generalized Orders of Magnitude for Scalable, Parallel, High-Dynamic-Range Computation
- AI for scientific discovery is a social problem
- FakeParts: a New Family of AI-Generated DeepFakes
- Universal Learning of Nonlinear Dynamics
- Fast weight programming and linear transformers: from machine learning to neurobiology
- AlphaGo Moment for Model Architecture Discovery
- Dynamic Chunking for End-to-End Hierarchical Sequence Modeling
- Beyond Sensor Data: Foundation Models of Behavioral Data from Wearables Improve Health Predictions
- Fast and Simplex: 2-Simplicial Attention in Triton
- Log-Linear Attention
- Mercury: Ultra-Fast Language Models Based on Diffusion
- Leveraging Natural Language Processing to Unravel the Mystery of Life: A Review of NLP Approaches in Genomics, Transcriptomics, and Proteomics
- Zebra-Llama: Towards Extremely Efficient Hybrid Models
- Blending Complementary Memory Systems in Hybrid Quadratic-Linear Transformers
- True Zero-Shot Inference of Dynamical Systems Preserving Long-Term Statistics
- Flash Invariant Point Attention
- TransMamba: A Sequence-Level Hybrid Transformer-Mamba Language Model
- Transformers without Normalization
- SPECTRE: An FFT-Based Efficient Drop-In Replacement to Self-Attention for Long Contexts
- Idiosyncrasies in Large Language Models
- Scaling Laws in Patchification: An Image Is Worth 50,176 Tokens And More
- MambaGlue: Fast and Robust Local Feature Matching With Mamba
- Context-Selective State Space Models: Feedback is All You Need
- Random Controlled Differential Equations
- Foundation models for electrocardiogram interpretation: clinical implications
- Hierarchical Decision Mamba Meets Agentic AI: A Novel Approach for RAN Slicing in 6G
- ECG-RAMBA: Zero-Shot ECG Generalization by Morphology-Rhythm Disentanglement and Long-Range Modeling
- Text-Routed Sparse Mixture-of-Experts Model with Explanation and Temporal Alignment for Multi-Modal Sentiment Analysis
- MAPLE: Efficient and Diverse Multi-Alpha Generation for Portfolio Construction
- Reservoir Computing inspired Matrix Multiplication-free Language Model
- Breaking the Memory Wall: Exact Analytical Differentiation via Tiled Operator-Space Evolution
- TYTAN: Taylor-series based Non-Linear Activation Engine for Deep Learning Accelerators
- Why Are Linear RNNs More Parallelizable?
- What Matters in Deep Learning for Time Series Forecasting?
- Visual Autoregressive Modelling for Monocular Depth Estimation
- Lessons from Neuroscience for AI: How integrating Actions, Compositional Structure and Episodic Memory could enable Safe, Interpretable and Human-Like AI
- Learning When Not to Attend Globally
- MEGA-PCC: A Mamba-based Efficient Approach for Joint Geometry and Attribute Point Cloud Compression
- DeMoGen: Towards Decompositional Human Motion Generation with Energy-Based Diffusion Models
- Masking Teacher and Reinforcing Student for Distilling Vision-Language Models
- Transformer Reconstructed with Dynamic Value Attention
- PointCHR: Point Cloud Analysis via Curvature-Aware Hyperbolic Rectification
- Compressing Observation History into Agent Memory: Distilling Transformers into Recurrent Transformers
- UniSkip-Mamba: A Frequency-Aware State Space Model for Audio-Visual Temporal Forgery Localization
- A Controlled Visual-Backbone Benchmark for Multimodal Short-Term Solar Irradiance Forecasting
- Multi-Modal Object Re-Identification with Prompt-S6 and Semantic-Aware Knowledge Guidance
- Direction-adaptive Mamba: Spatial-Frequency Dual-Domain Collaborative Learning for PolSAR Image Classification
- QuantFlow: A Federated Mamba-Based Post-Transformer Foundation Model for Time-Series Forecasting
- StreamIndex: Memory-Bounded Compressed Sparse Attention via Streaming Top-k
- Scale Weight Decay and Train Better
- No Free Lunch in Flow Surrogates under Time-Varying Boundary Conditions: A Two-Regime Study
- LCMamNet: A Lightweight Cross-scale Mamba Network for Infrared Small Target Detection
- WHTMix: Efficient Stereo Depth Estimation via Walsh-Hadamard Token Mixing
- Memory for Large Language Models
- Raven: High-Recall Sequence Modeling with Sparse Memory Routing
- Disentangling Multi-View Scanning in Mamba for Network Traffic Anomaly Detection
- Hierarchical Grading in Large Language Models
- DAMamba-UNet3D: A Parameter-Efficient Mamba State Space U-Net with Dynamic Adaptive Scan for 3D Medical Image Segmentation
- PointCSP: Cross-Sample Semantic Propagation and Stability Preservation in Self-Supervised Point Cloud Learning
- VEMamba: Efficient Isotropic Reconstruction of Volume Electron Microscopy with Axial-Lateral Consistent Mamba
- ZipMap: Linear-Time Stateful 3D Reconstruction via Test-Time Training
- Reverso: Efficient Time Series Foundation Models for Zero-shot Forecasting
- CellMamba: Adaptive Mamba for Accurate and Efficient Cell Detection
- Self-attention vector output similarities reveal how machines pay attention
- Causal-HM: Restoring Physical Generative Logic in Multimodal Anomaly Detection via Hierarchical Modulation
- RAPTOR: Real-Time High-Resolution UAV Video Prediction with Efficient Video Attention
- Spatiotemporal-Untrammelled Mixture of Experts for Multi-Person Motion Prediction
- Self-supervised Multiplex Consensus Mamba for General Image Fusion
- Learning to Sense for Driving: Joint Optics-Sensor-Model Co-Design for Semantic Segmentation
- Distilling to Hybrid Attention Models via KL-Guided Layer Selection
- Recurrent Off-Policy Deep Reinforcement Learning Doesn't Have to be Slow
- milliMamba: Specular-Aware Human Pose Estimation via Dual mmWave Radar with Multi-Frame Mamba Fusion
- HGAN-SDEs: Learning Neural Stochastic Differential Equations with Hermite-Guided Adversarial Training
- JDPNet: A Network Based on Joint Degradation Processing for Underwater Image Enhancement
- Generative Krylov Subspace Representations for Scalable Quantum Eigensolvers
- Mamba-Based Modality Disentanglement Network for Multi-Contrast MRI Reconstruction
- Efficient Jailbreak Mitigation Using Semantic Linear Classification in a Multi-Staged Pipeline
- Efficient Vision Mamba for MRI Super-Resolution via Hybrid Selective Scanning
- DeltaMIL: Gated Memory Integration for Efficient and Discriminative Whole Slide Image Analysis
- Lag Operator SSMs: A Geometric Framework for Structured State Space Modeling
- Rectification Reimagined: A Unified Mamba Model for Image Correction and Rectangling with Prompts
- From Shortcut to Induction Head: How Data Diversity Shapes Algorithm Selection in Transformers
- When Does Learning Renormalize? Sufficient Conditions for Power Law Spectral Dynamics
- MACE-Dance: Motion-Appearance Cascaded Experts for Music-Driven Dance Video Generation
- MambaMIL+: Modeling Long-Term Contextual Patterns for Gigapixel Whole Slide Image
- A Systematic Reproducibility Study of BSARec for Sequential Recommendation
- WDFFU-Mamba: A Wavelet-guided Dual-attention Feature Fusion Mamba for Breast Tumor Segmentation in Ultrasound Images
- Physics of Language Models: Part 4.1, Architecture Design and the Magic of Canon Layers
- KineST: A Kinematics-guided Spatiotemporal State Space Model for Human Motion Tracking from Sparse Signals
- KOSS: Kalman-Optimal Selective State Spaces for Long-Term Sequence Modeling
- CPMamba: Selective State Space Models for MIMO Channel Prediction in High-Mobility Environments
- LaverNet: Lightweight All-in-one Video Restoration via Selective Propagation
- Prefix Sums via Kronecker Products
- The Evolution of Reranking Models in Information Retrieval: From Heuristic Methods to Large Language Models
- AIE4ML: An End-to-End Framework for Compiling Neural Networks for the Next Generation of AMD AI Engines
- Keep the Core: Adversarial Priors for Significance-Preserving Brain MRI Segmentation
- Characterizing Mamba's Selective Memory using Auto-Encoders
- OccSTeP: Benchmarking 4D Occupancy Spatio-Temporal Persistence
- MMMamba: A Versatile Cross-Modal In Context Fusion Framework for Pan-Sharpening and Zero-Shot Image Enhancement
- How Many Heads Make an SSM? A Unified Framework for Attention and State Space Models
- LADY: Linear Attention for Autonomous Driving Efficiency without Transformers
- BarcodeMamba+: Advancing State-Space Models for Fungal Biodiversity Research
- PSMamba: Progressive Self-supervised Vision Mamba for Plant Disease Recognition
- FLAME: Flow Enhanced Legendre Memory Models for General Time Series Forecasting
- Kinetic-Mamba: Mamba-Assisted Predictions of Stiff Chemical Kinetics
- BlossomRec: Block-level Fused Sparse Attention Mechanism for Sequential Recommendations
- Grab-3D: Detecting AI-Generated Videos from 3D Geometric Temporal Consistency
- CoRA: A Collaborative Robust Architecture with Hybrid Fusion for Efficient Perception
- Sliding Window Recurrences for Sequence Models
- COBRA: Catastrophic Bit-flip Reliability Analysis of State-Space Models
- Exact Flow Linear Attention: Exact Solution from Continuous-Time Dynamics
- Referring Change Detection in Remote Sensing Imagery
- WATOS: Efficient LLM Training Strategies and Architecture Co-exploration for Wafer-scale Chip
- SCIR: A Self-Correcting Iterative Refinement Framework for Enhanced Information Extraction Based on Schema
- BLASST: Dynamic BLocked Attention Sparsity via Softmax Thresholding
- All-in-One ASR: Unifying Encoder-Decoder Models of CTC, Attention, and Transducer in Dual-Mode ASR
- TSkel-Mamba: Temporal Dynamic Modeling via State Space Model for Human Skeleton-based Action Recognition
- An Anatomy of Vision-Language-Action Models: From Modules to Milestones and Challenges
- Hybrid Transformer-Mamba for Weakly Supervised Volumetric Medical Segmentation
- GDKVM: Echocardiography Video Segmentation via Spatiotemporal Key-Value Memory with Gated Delta Rule
- Provably Learning from Modern Language Models via Low Logit Rank
- StateSpace-SSL: Linear-Time Self-supervised Learning for Plant Disease Detection
- Understanding temperature tuning in energy-based models
- InfiniteVL: Synergizing Linear and Sparse Attention for Highly-Efficient, Unlimited-Input Vision-Language Models
- Fourier-RWKV: A Multi-State Perception Network for Efficient Image Dehazing
- GatedFWA: Linear Flash Windowed Attention with Gated Associative Memory
- UltrasODM: A Dual Stream Optical Flow Mamba Network for 3D Freehand Ultrasound Reconstruction
- How Far are Modern Trackers from UAV-Anti-UAV? A Million-Scale Benchmark and New Baseline
- MultiAPI Spoof: A Multi-API Dataset and Local-Attention Network for Speech Anti-spoofing Detection
- Multi-view Pyramid Transformer: Look Coarser to See Broader
- PlantBiMoE: A Bidirectional Foundation Model with SparseMoE for Plant Genomes
- Dual Refinement Cycle Learning: Unsupervised Text Classification of Mamba and Community Detection on Text Attributed Graph
- Adaptive Normalization Mamba with Multi Scale Trend Decomposition and Patch MoE Encoding
- Quantifying Memory Use in Reinforcement Learning with Temporal Range
- ProAgent: Harnessing On-Demand Sensory Contexts for Proactive LLM Agent Systems
- Always Keep Your Promises: A Model-Agnostic Attribution Algorithm for Neural Networks
- On Memory: A comparison of memory mechanisms in world models
- TextMamba: Scene Text Detector with Mamba
- When Gender is Hard to See: Multi-Attribute Support for Long-Range Recognition
- Rethinking Training Dynamics in Scale-wise Autoregressive Generation
- FacePhys: State of the Heart Learning
- QL-LSTM: A Parameter-Efficient LSTM for Stable Long-Sequence Modeling
- TinyMyo: a Tiny Foundation Model for Flexible EMG Signal Processing at the Edge
- Evolutionary System 2 Reasoning: An Empirical Proof
- What Happens When: Learning Temporal Orders of Events in Videos
- Can Multi-Modal LLMs Provide Live Step-by-Step Task Guidance?
- Binary-30K: A Heterogeneous Dataset for Deep Learning in Binary Analysis and Malware Detection
- Recurrent Neural Networks with Linear Structures for Electricity Price Forecasting
- Disentangling Progress in Medical Image Registration: Beyond Trend-Driven Architectures towards Domain-Specific Strategies
- Long-Horizon Model-Based Offline Reinforcement Learning Without Explicit Conservatism
- State Space Models for Bioacoustics: A comparative Evaluation with Transformers
- Traffic Image Restoration under Adverse Weather via Frequency-Aware Mamba
- LSRS: Latent Scale Rejection Sampling for Visual Autoregressive Modeling
- Overcoming State Inertia: Minimally Invasive Temporal Alignment for Evolving Contexts
- AaPE: Aliasing-aware Patch Embedding for Self-Supervised Audio Representation Learning
- DF-Mamba: Deformable State Space Modeling for 3D Hand Pose Estimation in Interactions
- SeeU: Seeing the Unseen World via 4D Dynamics-aware Generation
- Video4Spatial: Towards Visuospatial Intelligence with Context-Guided Video Generation
- Instant Video Models: Universal Adapters for Stabilizing Image-Based Networks
- PanFoMa: A Lightweight Foundation Model and Benchmark for Pan-Cancer
- Leveraging Large-Scale Pretrained Spatial-Spectral Priors for General Zero-Shot Pansharpening
- Deep Learning-Based Joint Uplink-Downlink CSI Acquisition for Next-Generation Upper Mid-Band Systems
- Unrolled Networks are Conditional Probability Flows in MRI Reconstruction
- Beyond SFT: Reinforcement Learning for Safer Large Reasoning Models with Better Reasoning Ability
- ViT3: Unlocking Test-Time Training in Vision
- Parallel Delayed Memory Units for Enhanced Temporal Modeling in Biomedical and Bioacoustic Signal Analysis
- Toward Content-based Indexing and Retrieval of Head and Neck CT with Abscess Segmentation
- MasHeNe: A Benchmark for Head and Neck CT Mass Segmentation using Window-Enhanced Mamba with Frequency-Domain Integration
- PointNet4D: A Lightweight 4D Point Cloud Video Backbone for Online and Offline Perception in Robotic Applications
- Panda: Self-distillation of Reusable Sensor-level Representations for High Energy Physics
- TBT-Former: Learning Temporal Boundary Distributions for Action Localization
- MindFuse: Towards GenAI Explainability in Marketing Strategy Co-Creation
- Upper Approximation Bounds for Neural Oscillators
- Sleep Apnea Detection on a Wireless Multimodal Wearable Device Without Oxygen Flow Using a Mamba-based Deep Learning Approach
- MambaScope: Coarse-to-Fine Scoping for Efficient Vision Mamba
- SelfAI: Building a Self-Training AI System with LLM Agents
- MM-DETR: An Efficient Multimodal Detection Transformer with Mamba-Driven Dual-Granularity Fusion and Frequency-Aware Modality Adapters
- SMamDiff: Spatial Mamba for Stochastic Human Motion Prediction
- BioArc: Discovering Optimal Neural Architectures for Biological Foundation Models
- ReactionMamba: Generating Short &Long Human Reaction Sequences
- Distributed Dynamic Associative Memory via Online Convex Optimization
- GSPN-2: Efficient Parallel Sequence Modeling
- PerfMamba: Performance Analysis and Pruning of Selective State Space Models
- Intelligent Neural Networks: From Layered Architectures to Graph-Organized Intelligence
- Every Token Counts: Generalizing 16M Ultra-Long Context in Large Language Models
- Adaptive Dueling Double Deep Q-networks in Uniswap V3 Replication and Extension with Mamba
- Harmonic-Percussive Disentangled Neural Audio Codec for Bandwidth Extension
- MMA: A Momentum Mamba Architecture for Human Activity Recognition with Inertial Sensors
- MobileI2V: Fast and High-Resolution Image-to-Video on Mobile Devices
- PathMamba: A Hybrid Mamba-Transformer for Topologically Coherent Road Segmentation in Satellite Imagery
- DeepRFTv2: Kernel-level Learning for Image Deblurring
- FlowerDance: MeanFlow for Efficient and Refined 3D Dance Generation
- Gated KalmaNet: A Fading Memory Layer Through Test-Time Ridge Regression
- Towards a Foundation Model for Partial Differential Equations Across Physics Domains
- TaCo: Capturing Spatio-Temporal Semantic Consistency in Remote Sensing Change Detection
- DAPointMamba: Domain Adaptive Point Mamba for Point Cloud Completion
- A Compact Hybrid Convolution--Frequency State Space Network for Learned Image Compression
- MambaEye: A Size-Agnostic Visual Encoder with Causal Sequential Processing
- Towards Edge General Intelligence: Knowledge Distillation for Mobile Agentic AI
- ChessMamba: Structure-Aware Interleaving of State Spaces for Change Detection in Remote Sensing Images
- MapFormer: Self-Supervised Learning of Cognitive Maps with Input-Dependent Positional Embeddings
- MambaRefine-YOLO: A Dual-Modality Small Object Detector for UAV Imagery
- Nemotron-Flash: Towards Latency-Optimal Hybrid Small Language Models
- FVAR: Visual Autoregressive Modeling via Next Focus Prediction
- MambaX: Image Super-Resolution with State Predictive Control
- SAMBA: Toward a Long-Context EEG Foundation Model via Spatial Embedding and Differential Mamba
- HiFi-MambaV2: Hierarchical Shared-Routed MoE for High-Fidelity MRI Reconstruction
- RNN as Linear Transformer: A Closer Investigation into Representational Potentials of Visual Mamba Models
- DiM-TS: Bridge the Gap between Selective State Space Models and Time Series for Generative Modeling
- Coherent Multi-Agent Trajectory Forecasting in Team Sports with CausalTraj
- Compact neural networks for astronomy with optimal transport bias correction
- MambaTAD: When State-Space Models Meet Long-Range Temporal Action Detection
- Equivalence of Context and Parameter Updates in Modern Transformer Blocks
- Selective Rotary Position Embedding
- Feature Partitioning and Semantic Equalization for Intrinsic Robustness in Semantic Communication under Packet Loss
- UAM: A Unified Attention-Mamba Backbone of Multimodal Framework for Tumor Cell Classification
- Nemotron Elastic: Towards Efficient Many-in-One Reasoning LLMs
- Goal-Directed Search Outperforms Goal-Agnostic Memory Compression in Long-Context Memory Tasks
- Predicting one-year clinical instability and mortality in heart failure patients using sequence modeling
- TimeViper: A Hybrid Mamba-Transformer Vision-Language Model for Efficient Long Video Understanding
- Simba: Towards High-Fidelity and Geometrically-Consistent Point Cloud Completion via Transformation Diffusion
- Recurrent Autoregressive Diffusion: Global Memory Meets Local Attention
- CausalMamba: Interpretable State Space Modeling for Temporal Rumor Causality
- MamTiff-CAD: Multi-Scale Latent Diffusion with Mamba+ for Complex Parametric Sequence
- AccelOpt: A Self-Improving LLM Agentic System for AI Accelerator Kernel Optimization
- MambaTrack3D: A State Space Model Framework for LiDAR-Based Object Tracking under High Temporal Variation
- Context Cascade Compression: Exploring the Upper Limits of Text Compression
- EfficientSAM3: Progressive Hierarchical Distillation for Video Concept Segmentation from SAM1, 2, and 3
- Addressing the gravitational collapse of a massless scalar field with Physics-Informed Neural Networks
- Quality-Controlled Multimodal Emotion Recognition in Conversations with Identity-Based Transfer Learning and MAMBA Fusion
- Parameter Aware Mamba Model for Multi-task Dense Prediction
- Accelerating Automatic Differentiation of Direct Form Digital Filters
- Compute-in-Memory Implementation of State Space Models for Event Sequence Processing
- MergeDNA: Context-aware Genome Modeling with Dynamic Tokenization through Token Merging
- Naga: Vedic Encoding for Deep State Space Models
- Uni-Hand: Universal Hand Motion Forecasting in Egocentric Views
- WinMamba: Multi-Scale Shifted Windows in State Space Model for 3D Object Detection
- ExplicitLM: Decoupling Knowledge from Parameters via Explicit Memory Banks
- MMRINet: Efficient Mamba-Based Segmentation with Dual-Path Refinement for Low-Resource MRI Analysis
- Through-Foliage Surface-Temperature Reconstruction for Early Wildfire Detection
- DensePercept-NCSSD: Vision Mamba towards Real-time Dense Visual Perception with Non-Causal State Space Duality
- A Disease-Aware Dual-Stage Framework for Chest X-ray Report Generation
- BudgetLeak: Membership Inference Attacks on RAG Systems via the Generation Budget Side Channel
- Arcee: Differentiable Recurrent State Chain for Generative Vision Modeling with Mamba SSMs
- Rethinking Progression of Memory State in Robotic Manipulation: An Object-Centric Perspective
- VitalBench: A Rigorous Multi-Center Benchmark for Long-Term Vital Sign Prediction in Intraoperative Care
- Reinforcing Trustworthiness in Multimodal Emotional Support Systems
- Viper-F1: Fast and Fine-Grained Multimodal Understanding with Cross-Modal State-Space Modulation
- Multitask GLocal OBIA-Mamba for Sentinel-2 Landcover Mapping
- STAGE: A Symbolic Tensor grAph GEnerator for distributed AI system co-design
- AdaptViG: Adaptive Vision GNN with Exponential Decay Gating
- Do Blind Spots Matter for Word-Referent Mapping? A Computational Study with Infant Egocentric Video
- Koopman Invariants as Drivers of Emergent Time-Series Clustering in Joint-Embedding Predictive Architectures
- Making Every Head Count: Sparse Attention Without the Speed-Performance Trade-off
- MPCM-Net: Multi-scale network integrates partial attention convolution with Mamba for ground-based cloud image segmentation
- CMI-MTL: Cross-Mamba interaction based multi-task learning for medical visual question answering
- SasMamba: A Lightweight Structure-Aware Stride State Space Model for 3D Human Pose Estimation
- SSMRadNet : A Sample-wise State-Space Framework for Efficient and Ultra-Light Radar Segmentation and Object Detection
- MVSMamba: Multi-View Stereo with State Space Model
- UniLION: Towards Unified Autonomous Driving Model with Linear Group RNNs
- Towards General Auditory Intelligence: Large Multimodal Models for Machine Listening and Speaking
- Leveraging unlabelled data for generalizable neural population decoding
- Hybrid Quantum-Classical Selective State Space Artificial Intelligence
- ReIDMamba: Learning Discriminative Features with Visual State Space Model for Person Re-Identification
- CloudMamba: Grouped Selective State Spaces for Point Cloud Analysis
- Fast Multi-Organ Fine Segmentation in CT Images with Hierarchical Sparse Sampling and Residual Transformer
- TNT: Improving Chunkwise Training for Test-Time Memorization
- 4DSTR: Advancing Generative 4D Gaussians with Spatial-Temporal Rectification for High-Quality and Consistent 4D Generation
- CenterMamba-SAM: Center-Prioritized Scanning and Temporal Prototypes for Brain Lesion Segmentation
- Rethinking Crystal Symmetry Prediction: A Decoupled Perspective
- Learning to Focus: Focal Attention for Selective and Scalable Transformers
- Resource Efficient Sleep Staging via Multi-Level Masking and Prompt Learning
- MRT: Learning Compact Representations with Mixed RWKV-Transformer for Extreme Image Compression
- MirrorMamba: Towards Scalable and Robust Mirror Detection in Videos
- Spatial-Frequency Enhanced Mamba for Multi-Modal Image Fusion
- Attention and Compression is all you need for Controllably Efficient Language Models
- A Risk-Neutral Neural Operator for Arbitrage-Free SPX-VIX Term Structures
- MambaOVSR: Multiscale Fusion with Global Motion Modeling for Chinese Opera Video Super-Resolution
- Next-Latent Prediction Transformers Learn Compact World Models
- TYrPPG: Uncomplicated and Enhanced Learning Capability rPPG for Remote Heart Rate Estimation
- Towards Frequency-Adaptive Learning for SAR Despeckling
- GroupKAN: Efficient Kolmogorov-Arnold Networks via Grouped Spline Modeling
- LoPT: Lossless Parallel Tokenization Acceleration for Long Context Inference of Large Language Model
- 3D Gaussian Point Encoders
- Cambrian-S: Towards Spatial Supersensing in Video
- Reusing Pre-Training Data at Test Time is a Compute Multiplier
- Forget BIT, It is All about TOKEN: Towards Semantic Information Theory for LLMs
- Enhancing Multimodal Protein Function Prediction Through Dual-Branch Dynamic Selection with Reconstructive Pre-Training
- Enhancing Medical Image Segmentation via Heat Conduction Equation
- Improving DF-Conformer Using Hydra For High-Fidelity Generative Speech Enhancement on Discrete Codec Token
- Apriel-H1: Towards Efficient Enterprise Reasoning Models
- Machine Learning for RNA Secondary Structure Prediction: a review of current methods and challenges
- M2S2L: Mamba-based Multi-Scale Spatial-temporal Learning for Video Anomaly Detection
- Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live
- MM-UNet: Morph Mamba U-shaped Convolutional Networks for Retinal Vessel Segmentation
- Sample-Efficient Language Modeling with Linear Attention and Lightweight Enhancements
- Hyper Hawkes Processes: Interpretable Models of Marked Temporal Point Processes
- The Hidden Power of Normalization Layers in Neural Networks: Exponential Capacity Control
- FlashEVA: Accelerating LLM inference via Efficient Attention
- Tree Training: Accelerating Agentic LLMs Training via Shared Prefix Reuse
- Region-Aware Reconstruction Strategy for Pre-training fMRI Foundation Model
- MambaNetLK: Enhancing Colonoscopy Point Cloud Registration with Mamba
- Versatile and Efficient Medical Image Super-Resolution Via Frequency-Gated Mamba
- Higher-order Linear Attention
- AFM-Net: Advanced Fusing Hierarchical CNN Visual Priors with Global Sequence Modeling for Remote Sensing Image Scene Classification
- Advancing AI Challenges for the United States Department of the Air Force
- Understanding and Enhancing Mamba-Transformer Hybrids for Memory Recall and Language Modeling
- Cross-Platform Evaluation of Reasoning Capabilities in Foundation Models
- Context Engineering 2.0: The Context of Context Engineering
- Hybrid Quantum-Classical Recurrent Neural Networks
- MossNet: Mixture of State-Space Experts is a Multi-Head Attention
- Nirvana: A Specialized Generalist Model With Task-Aware Memory Mechanism
- Predicate Renaming via Large Language Models
- TempoPFN: Synthetic Pre-training of Linear RNNs for Zero-shot Time Series Forecasting
- Semantic-Aware Temporal Adaptation for UAV Anti-UAV Tracking
- The Art of Not Forgetting A Local Learning Architecture for Continual Learning
- CMT-RAG: Complementary Memory Traces for Multi-turn Multi-hop RAG
- Phoneme- vs. Character-Level Targets and Selective State-Space Models for Intracortical Brain-to-Text
- Metis: Memory Foundation Model
- FMMVCC: Fuzzy Mamba-based Multi-View Contrastive Clustering for Univariate Time Series
- Benchmarking DNA large language models on quadruplexes
- Journey Operators for Structured Multi-Axis Composition
- Reactive Transformer (RxT) -- Stateful Real-Time Processing for Event-Driven Reactive Language Models
- Latent-IM: Latent Interaction Management for Speech LLMs
- CommFuse: Hiding Tail Latency via Communication Decomposition and Fusion for Distributed LLM Training
- Integrative neurocybernetic modeling in the era of large-scale neuroscience
- Simplified Sparse Attention via Gist Tokens
- Low-latency Event-based Object Detection with Spatially-Sparse Linear Attention
- Enhancing Compositional Reasoning in CLIP via Reconstruction and Alignment of Text Descriptions
- HySparse: A Hybrid Sparse Attention Architecture with Oracle Token Selection and KV Cache Sharing
- FrugalPrompt: Reducing Contextual Overhead in Large Language Models via Token Attribution
- The Bayesian Geometry of Transformer Attention
- TV-Rec: Time-Variant Convolutional Filter for Sequential Recommendation
- State Space and Self-Attention Collaborative Network with Feature Aggregation for DOA Estimation
- MaGNet: A Mamba Dual-Hypergraph Network for Stock Prediction via Temporal-Causal and Global Relational Learning
- Larger Hausdorff Dimension in Scanning Pattern Facilitates Mamba-Based Methods in Low-Light Image Enhancement
- Sequences of Logits Reveal the Low Rank Structure of Language Models
- Language Model Behavioral Phases are Consistent Across Architecture, Training Data, and Scale
- The Kinetics of Reasoning: How Chain-of-Thought Shapes Learning in Transformers?
- Long-Context Modeling with Dynamic Hierarchical Sparse Attention for On-Device LLMs
- DeshadowMamba: Deshadowing as 1D Sequential Similarity
- Causal Convolutional Neural Networks as Finite Impulse Response Filters
- Beyond Line-Level Filtering for the Pretraining Corpora of LLMs
- Sequence Modeling with Spectral Mean Flows
- A Survey on Efficient Vision-Language-Action Models
- HieraMamba: Video Temporal Grounding via Hierarchical Anchor-Mamba Pooling
- Hankel Singular Value Regularization for Highly Compressible State Space Models
- GTR-Mamba: Geometry-to-Tangent Routing for Hyperbolic POI Recommendation
- Beyond Semantics: How Temporal Biases Shape Retrieval in Transformer and State-Space Models
- Scalable Neural Decoders for Practical Real-Time Quantum Error Correction
- Energy-Efficient Domain-Specific Artificial Intelligence Models and Agents: Pathways and Paradigms
- Massive Memorization with Hundreds of Trillions of Parameters for Sequential Transducer Generative Recommenders
- ATLAS: Adaptive Transfer Scaling Laws for Multilingual Pretraining, Finetuning, and Decoding the Curse of Multilinguality
- Transformer Based Linear Attention with Optimized GPU Kernel Implementation
- Structure-Aware Fusion with Progressive Injection for Multimodal Molecular Representation Learning
- Correlation Dimension of Auto-Regressive Large Language Models
- RadioMapMotion: A Dataset and Baseline for Proactive Spatio-Temporal Radio Environment Prediction
- WaveSeg: Enhancing Segmentation Precision via High-Frequency Prior and Mamba-Driven Spectrum Decomposition
- ParaRNN: Unlocking Parallel Training of Nonlinear RNNs for Large Language Models
- LLM-Integrated Bayesian State Space Models for Multimodal Time-Series Forecasting
- Alleviating Forgetfulness of Linear Attention by Hybrid Sparse Attention and Contextualized Learnable Token Eviction
- PointMapPolicy: Structured Point Cloud Processing for Multi-Modal Imitation Learning
- Deep Learning Based Domain Adaptation Methods in Remote Sensing: A Comprehensive Survey
- Context-level Language Modeling by Learning Predictive Context Embeddings
- Mixing Importance with Diversity: Joint Optimization for KV Cache Compression in Large Vision-Language Models
- From Masks to Worlds: A Hitchhiker's Guide to World Models
- What is the Best Sequence Length for BABYLM?
- Tibetan Language and AI: A Comprehensive Survey of Resources, Methods and Challenges
- Beyond Hearing: Learning Task-Agnostic ExG Representations from Earphones via Physiology-Informed Tokenization
- Auditory Attention Decoding from Ear-EEG Signals: A Dataset with Dynamic Attention Switching and Rigorous Cross-Validation
- PRGCN: A Graph Memory Network for Cross-Sequence Pattern Reuse in 3D Human Pose Estimation
- Every Attention Matters: An Efficient Hybrid Architecture for Long-Context Reasoning
- UltraGen: High-Resolution Video Generation with Hierarchical Attention
- MLMA: Towards Multilingual ASR With Mamba-based Architectures
- Simple and Efficient Heterogeneous Temporal Graph Neural Network
- OmniNWM: Omniscient Driving Navigation World Models
- Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs
- LIME: Link-based user-item Interaction Modeling with decoupled xor attention for Efficient test time scaling
- Condition-Invariant fMRI Decoding of Speech Intelligibility with Deep State Space Model
- Δt-Mamba3D: A Time-Aware Spatio-Temporal State-Space Model for Breast Cancer Risk Prediction
- Cortical-SSM: A Deep State Space Model for EEG and ECoG Motor Imagery Decoding
- Unbiased Gradient Low-Rank Projection
- A novel water quality prediction model based on BiMKANsDformer
- CausalMamba: Scalable Conditional State Space Models for Neural Causal Inference
- Understanding and Improving Length Generalization in Hierarchical Sparse Attention Models
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- Efficient Vision-Language-Action Models for Embodied Manipulation: A Systematic Survey
- Saber: An Efficient Sampling with Adaptive Acceleration and Backtracking Enhanced Remasking for Diffusion Language Model
- Evaluating protein binding interfaces with PUMBA
- End-to-end Listen, Look, Speak and Act
- Neuronal Group Communication for Efficient Neural representation
- EMRRG: Efficient Fine-Tuning Pre-trained X-ray Mamba Networks for Radiology Report Generation
- WaMaIR: Image Restoration via Multiscale Wavelet Convolutions and Mamba-based Channel Modeling with Texture Enhancement
- EdgeNavMamba: Mamba Optimized Object Detection for Energy Efficient Edge Devices
- PRISM: Probabilistic Runtime Insights and Scalable Performance Modeling for Large-Scale Distributed Training
- Still Competitive: Revisiting Recurrent Models for Irregular Time Series Prediction
- Rethinking Convergence in Deep Learning: The Predictive-Corrective Paradigm for Anatomy-Informed Brain MRI Segmentation
- StretchySnake: Flexible SSM Training Unlocks Action Recognition Across Spatio-Temporal Scales
- VM-BeautyNet: A Synergistic Ensemble of Vision Transformer and Mamba for Facial Beauty Prediction
- To Infinity and Beyond: Tool-Use Unlocks Length Generalization in State Space Models
- State-Space Models for Tabular Prior-Data Fitted Networks
- A Deep State-Space Model Compression Method using Upper Bound on Output Error
- Vision Mamba for Permeability Prediction of Porous Media
- ASecond-Order SpikingSSM for Wearables
- Document Intelligence in the Era of Large Language Models: A Survey
- End-to-End Multi-Modal Diffusion Mamba
- Prompt-based Adaptation in Large-scale Vision Models: A Survey
- DMTrack: Deformable State-Space Modeling for UAV Multi-Object Tracking with Kalman Fusion and Uncertainty-Aware Association
- CymbaDiff: Structured Spatial Diffusion for Sketch-based 3D Semantic Urban Scene Generation
- One Dimensional CNN ECG Mamba for Multilabel Abnormality Classification in 12 Lead ECG
- Learning Human Motion with Temporally Conditional Mamba
- Fine-grained Analysis of Brain-LLM Alignment through Input Attribution
- State Space Prompting via Gathering and Spreading Spatio-Temporal Information for Video Understanding
- Chimera: State Space Models Beyond Sequences
- Mamba Can Learn Low-Dimensional Targets In-Context via Test-Time Feature Learning
- Deconstructing Attention: Investigating Design Principles for Effective Language Modeling
- Situat3DChange: Situated 3D Change Understanding Dataset for Multimodal Large Language Model
- Robust Photoplethysmography Signal Denoising via Mamba Networks
- High-Resolution Spatiotemporal Modeling with Global-Local State Space Models for Video-Based Human Pose Estimation
- ShishuLM: Lightweight Language Model with Hybrid Decoder-MLP Architecture and Paired Weight Sharing
- xLSTM Scaling Laws: Competitive Performance with Linear Time-Complexity
- Protenix-Mini+: efficient structure prediction model with scalable pairformer
- MSF-Mamba: Motion-aware State Fusion Mamba for Efficient Micro-Gesture Recognition
- SaFiRe: Saccade-Fixation Reiteration with Mamba for Referring Image Segmentation
- Artificial intelligence as a surrogate brain: Bridging neural dynamical models and data
- Exploring Resolution-Wise Shared Attention in Hybrid Mamba-U-Nets for Improved Cross-Corpus Speech Enhancement
- AURA: Action-Gated Memory for Robot Policies at Constant VRAM
- Forget Attention: Importance-Aware Attention Is All You Need
- Sparse Query Attention (SQA): A Computationally Efficient Attention Mechanism with Query Heads Reduction
- Attention Once Is All You Need: Efficient Streaming Inference with Stateful Transformers
- DELTA: Dynamic Layer-Aware Token Attention for Efficient Long-Context Reasoning
- Titans Revisited: A Lightweight Reimplementation and Critical Analysis of a Test-Time Memory Model
- Online Video Depth Anything: Temporally-Consistent Depth Prediction with Low Memory Consumption
- MambaH-Fit: Rethinking Hyper-surface Fitting-based Point Cloud Normal Estimation via State Space Modelling
- Minkowski-MambaNet: A Point Cloud Framework with Selective State Space Models for Forest Biomass Quantification
- Which Heads Matter for Reasoning? RL-Guided KV Cache Compression
- gLSTM: Mitigating Over-Squashing by Increasing Storage Capacity
- MetaDefense: Defending Finetuning-based Jailbreak Attack Before and During Generation
- Deep Neural Networks Inspired by Differential Equations
- Language Models Do Not Embed Numbers Continuously
- Long-Tailed Recognition via Information-Preservable Two-Stage Learning
- LinVideo: A Post-Training Framework towards O(n) Attention in Efficient Video Generation
- Knowledge-Aware Mamba for Joint Change Detection and Classification from MODIS Times Series
- Accuracy, Memory Efficiency and Generalization: A Comparative Study on Liquid Neural Networks and Recurrent Neural Networks
- When Thoughts Meet Facts: Reusable Reasoning for Long-Context LMs
- Artificial Hippocampus Networks for Efficient Long-Context Modeling
- ELMUR: External Layer Memory with Update/Rewrite for Long-Horizon RL Problems
- Native Hybrid Attention for Efficient Sequence Modeling
- Revisiting Node Affinity Prediction in Temporal Graphs
- End-to-End Test-Time Training for Long Context
- BOTANIC-0: a series of foundation models for plant genomic data
- DeRainMamba: A Frequency-Aware State Space Model with Detail Enhancement for Image Deraining
- A Comparative Analysis of Contextual Representation Flow in State-Space and Transformer Architectures
- The Markovian Thinker: Architecture-Agnostic Linear Scaling of Reasoning
- CLAQS: Compact Learnable All-Quantum Token Mixer with Shared-ansatz for Text Classification
- Microstructure sensitive recurrent neural network surrogate model of crystal plasticity
- The Anatomy of a Triton Attention Kernel
- Untangling Component Imbalance in Hybrid Linear Attention Conversion Methods
- WaveSP-Net: Learnable Wavelet-Domain Sparse Prompt Tuning for Speech Deepfake Detection
- Rivaling Transformers: Multi-Scale Structured State-Space Mixtures for Agentic 6G O-RAN
- On Structured State-Space Duality
- REN: Anatomically-Informed Mixture-of-Experts for Interstitial Lung Disease Diagnosis
- Speak, Edit, Repeat: High-Fidelity Voice Editing and Zero-Shot TTS with Cross-Attentive Mamba
- Noise or Signal? Deconstructing Contradictions and An Adaptive Remedy for Reversible Normalization in Time Series Forecasting
- Pack and Force Your Memory: Long-form and Consistent Video Generation
- SpikingMamba: Towards Energy-Efficient Large Language Models via Knowledge Distillation from Mamba
- Poolformer: Recurrent Networks with Pooling for Long-Sequence Modeling
- Compressed Convolutional Attention: Efficient Attention in a Compressed Latent Space
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Representation Potentials of Foundation Models for Multimodal Alignment: A Survey
- Benchmarking M-LTSF: Frequency and Noise-Based Evaluation of Multivariate Long Time Series Forecasting Models
- How Different from the Past? Spatio-Temporal Time Series Forecasting with Self-Supervised Deviation Learning
- Hybrid Architectures for Language Models: Systematic Analysis and Design Insights
- SSM-CGM: Interpretable State-Space Forecasting Model of Continuous Glucose Monitoring for Personalized Diabetes Management
- Wave-PDE Nets: Trainable Wave-Equation Layers as an Alternative to Attention
- Detecting Invariant Manifolds in ReLU-Based RNNs
- Understanding Transformers for Time Series: Rank Structure, Flow-of-ranks, and Compressibility
- MambaCAFU: Hybrid Multi-Scale and Multi-Attention Model with Mamba-Based Fusion for Medical Image Segmentation
- How We Won BraTS-SSA 2025: Brain Tumor Segmentation in the Sub-Saharan African Population Using Segmentation-Aware Data Augmentation and Model Ensembling
- Flip Distribution Alignment VAE for Multi-Phase MRI Synthesis
- Flamed-TTS: Flow Matching Attention-Free Models for Efficient Generating and Dynamic Pacing Zero-shot Text-to-Speech
- RAxSS: Retrieval-Augmented Sparse Sampling for Explainable Variable-Length Medical Time Series Classification
- Brain-Language Model Alignment: Insights into the Platonic Hypothesis and Intermediate-Layer Advantage
- Linear RNNs for autoregressive generation of long music samples
- Uncovering the Computational Ingredients of Human-Like Representations in LLMs
- Gather-Scatter Mamba: Accelerating Propagation with Efficient State Space Model
- MASH: Modeling Abstention via Selective Help-Seeking
- Memory Determines Learning Direction: A Theory of Gradient-Based Optimization in State Space Models
- Can Mamba Learn In Context with Outliers? A Theoretical Generalization Analysis
- Retrieval-Augmented Framework for LLM-Based Clinical Decision Support
- Local Linear Attention: An Optimal Interpolation of Linear and Softmax Attention For Test-Time Regression
- Delayed Attention Training Improves Length Generalization in Transformer--RNN Hybrids
- Recursive Self-Aggregation Unlocks Deep Thinking in Large Language Models
- Learning to See Before Seeing: Demystifying LLM Visual Priors from Language Pre-training
- PRISM: Progressive Rain removal with Integrated State-space Modeling
- Bringing Emerging Architectures to Sequence Labeling in NLP
- TTT3R: 3D Reconstruction as Test-Time Training
- Context-Driven Performance Modeling for Causal Inference Operators on Neural Processing Units
- Benchmarking ECG Foundational Models: A Reality Check Across Clinical Tasks
- From Ambiguity to Verdict: A Semiotic-Grounded Multi-Perspective Agent for LLM Logical Reasoning
- Wavelet-Assisted Mamba for Satellite-Derived Sea Surface Temperature Super-Resolution
- Similarity-Aware Selective State-Space Modeling for Semantic Correspondence
- BALR-SAM: Boundary-Aware Low-Rank Adaptation of SAM for Resource-Efficient Medical Image Segmentation
- BERT Bi-modal self-supervised learning for crop classification using Sentinel-2 and Planetscope
- Alternatives To Next Token Prediction In Text Generation -- A Survey
- HyMaTE: A Hybrid Mamba and Transformer Model for EHR Representation Learning
- ResFormer: All-Time Reservoir Memory for Long Sequence Classification
- Sim-DETR: Unlock DETR for Temporal Sentence Grounding
- AdaPtis: Reducing Pipeline Bubbles with Adaptive Pipeline Parallelism on Heterogeneous Models
- MSD-KMamba: Bidirectional Spatial-Aware Multi-Modal 3D Brain Segmentation via Multi-scale Self-Distilled Fusion Strategy
- EfficientMIL: Efficient Linear-Complexity MIL Method for WSI Classification
- VAMamba: An Efficient Visual Adaptive Mamba for Image Restoration
- MemMamba: Rethinking Memory Patterns in State Space Model
- Balanced Diffusion-Guided Fusion for Multimodal Remote Sensing Classification
- PDE-Transformer: A Continuous Dynamical Systems Approach to Sequence Modeling
- PonderLM-2: Pretraining LLM with Latent Thoughts in Continuous Space
- Look Back to Reason Forward: Revisitable Memory for Long-Context LLM Agents
- Multimodal Slice Interaction Network Enhanced by Transfer Learning for Precise Segmentation of Internal Gross Tumor Volume in Lung Cancer PET/CT Imaging
- TRUST: Test-Time Refinement using Uncertainty-Guided SSM Traverses
- Orochi: Versatile Biomedical Image Processor
- StateX: Enhancing RNN Recall via Post-training State Expansion
- JointDiff: Bridging Continuous and Discrete in Multi-Agent Trajectory Generation
- IIET: Efficient Numerical Transformer via Implicit Iterative Euler Method
- A model of errors in transformers
- Johnson-Lindenstrauss Lemma Guided Network for Efficient 3D Medical Segmentation
- Structured Sparse Transition Matrices to Enable State Tracking in State-Space Models
- Self-Supervised Point Cloud Completion based on Multi-View Augmentations of Single Partial Point Cloud
- Statistical Advantage of Softmax Attention: Insights from Single-Location Regression
- IPDnet2: an efficient and improved inter-channel phase difference estimation network for sound source localization
- StableDub: Taming Diffusion Prior for Generalized and Efficient Visual Dubbing
- A Unifying Framework for Parallelizing Sequential Models with Linear Dynamical Systems
- MedVSR: Medical Video Super-Resolution with Cross State-Space Propagation
- SlideMamba: Entropy-Based Adaptive Fusion of GNN and Mamba for Enhanced Representation Learning in Digital Pathology
- TF-Restormer: Complex Spectral Prediction for Speech Restoration
- Aligning Inductive Bias for Data-Efficient Generalization in State Space Models
- CoSupFormer : A Contrastive Supervised learning approach for EEG signal Classification
- A HyperGraphMamba-Based Multichannel Adaptive Model for ncRNA Classification
- U-Mamba2-SSL for Semi-Supervised Tooth and Pulp Segmentation in CBCT
- SpecMamba: Accelerating Mamba Inference on FPGA with Speculative Decoding
- SAGE:State-Aware Guided End-to-End Policy for Multi-Stage Sequential Tasks via Hidden Markov Decision Process
- Enhancing Linear Attention with Residual Learning
- RoboSSM: Scalable In-context Imitation Learning via State-Space Models
- Mamba Modulation: On the Length Generalization of Mamba
- ExPe: Exact Positional Encodings for Generative Transformer Models with Extrapolating Capabilities
- Subtract or Replay? Exact Deletion from Language-Model Memory
- FootprintNet: State-Transition-Guided Dynamic Footprint Learning for Multi-temporal Remote Sensing Change Detection
- AuricularWorld: Hierarchical Action-Guided World Modeling for Fine-Grained Auricular Structure Segmentation from CT Scans
- Recursive transformers for semiconductor thermo-mechanical reliability
- Retrieval from Within: An Intrinsic Capability of Attention-Based Models
- On the Importance of Multistability for Horizon Generalization in Reinforcement Learning
- Improving the Performance and Learning Stability of Parallelizable RNNs Designed for Ultra-Low Power Applications
- The Topological Trouble With Transformers
- SeedPolicy: Horizon Scaling via Self-Evolving Diffusion Policy for Robot Manipulation
- LoGeR: Long-Context Geometric Reconstruction with Hybrid Memory
- Transformers and genome language models
- Beyond computational equivalence: the behavioral inference principle for machine consciousness
- scGMB: A scRNA‐seq Cell Classification Method Combining GCN and Mamba
- AI-CNet3D: An Anatomically-Informed Cross-Attention Network with Multi-Task Consistency Fine-tuning for 3D Glaucoma Classification
- MsFIN: Multi-scale Feature Interaction Network for Traffic Accident Anticipation
- LEAF-Mamba: Local Emphatic and Adaptive Fusion State Space Model for RGB-D Salient Object Detection
- LCMF: Lightweight Cross-Modality Mambaformer for Embodied Robotics VQA
- Characterizing the Behavior of Training Mamba-based State Space Models on GPUs
- An overview of neural architectures for self-supervised audio representation learning from masked spectrograms
- Diff-GNSS: Diffusion-based Pseudorange Error Estimation
- Finding Outliers in a Haystack: Anomaly Detection for Large Pointcloud Scenes
- SISMA: Semantic Face Image Synthesis with Mamba
- Achilles' Heel of Mamba: Essential difficulties of the Mamba architecture demonstrated by synthetic data
- DA-Mamba: Dialogue-aware selective state-space model for multimodal engagement estimation
- LAWCAT: Efficient Distillation from Quadratic to Linear Attention with Convolution across Tokens for Long Context Modeling
- SynergyNet: Fusing Generative Priors and State-Space Models for Facial Beauty Prediction
- TSGym: Design Choices for Deep Multivariate Time-Series Forecasting
- EvoBrain: Dynamic Multi-Channel EEG Graph Modeling for Time-Evolving Brain Networks
- History-Aware Visuomotor Policy Learning via Point Tracking
- Surgical-MambaLLM: Mamba2-enhanced Multimodal Large Language Model for VQLA in Robotic Surgery
- Sequential Token Merging: Revisiting Hidden States
- DistillMatch: Leveraging Knowledge Distillation from Vision Foundation Model for Multimodal Image Matching
- UniGist: Towards General and Hardware-aligned Sequence-level Long Context Compression
- Interpretable Modeling of Articulatory Temporal Dynamics from real-time MRI for Phoneme Recognition
- Bridging Graph and State-Space Modeling for Intensive Care Unit Length of Stay Prediction
- FiLM-Nav: Efficient and Generalizable Navigation via VLM Fine-tuning
- SAM: A Mamba-2 State-Space Audio-Language Model
- Estimating Clinical Lab Test Result Trajectories from PPG using Physiological Foundation Model and Patient-Aware State Space Model -- a UNIPHY+ Approach
- Blind Room Impulse Response Identification via Reverberant Speech Spectrum Reconstruction
- HybridMamba: A Dual-domain Mamba for 3D Medical Image Segmentation
- VLHSA: Vision-Language Hierarchical Semantic Alignment for Jigsaw Puzzle Solving with Eroded Gaps
- Pre-Manipulation Alignment Prediction with Parallel Deep State-Space and Transformer Models
- TENET: An Efficient Sparsity-Aware LUT-Centric Architecture for Ternary LLM Inference On Edge
- State Space Models over Directed Graphs
- UM-Depth : Uncertainty Masked Self-Supervised Monocular Depth Estimation with Visual Odometry
- Cross-Distribution Diffusion Priors-Driven Iterative Reconstruction for Sparse-View CT
- FoundDiff: Foundational Diffusion Model for Generalizable Low-Dose CT Denoising
- The CCF AATC 2025 Speech Restoration Challenge: A Retrospective
- Accelerating Long-Term Molecular Dynamics with Physics-Informed Time-Series Forecasting
- TFANet: Three-Stage Image-Text Feature Alignment Network for Robust Referring Image Segmentation
- Match Chat: Real Time Generative AI and Generative Computing for Tennis
- CECT-Mamba: a Hierarchical Contrast-enhanced-aware Model for Pancreatic Tumor Subtyping from Multi-phase CECT
- U-Mamba2: Scaling State Space Models for Dental Anatomy Segmentation in CBCT
- Predictability Enables Parallelization of Nonlinear State Space Models
- Neuromorphic Intelligence
- SAM-TTT: Segment Anything Model via Reverse Parameter Configuration and Test-Time Training for Camouflaged Object Detection
- Detecting Multilevel Manipulation from Limit Order Book via Cascaded Contrastive Representation Learning
- Joint-octamamba:an octa joint segmentation network based on feature enhanced mamba
- MixANT: Observation-dependent Memory Propagation for Stochastic Dense Action Anticipation
- Mamba Outpaces Reformer in Stock Prediction with Sentiments from Top Ten LLMs
- MEMBOT: Memory-Based Robot in Intermittent POMDP
- Real-time reinforcement learning for turbulent state-dependent control in a bluff-body wake
- Genome-Factory: A Library for Tuning, Deploying, and Interpreting Genomic Foundation Models
- Long Context Automated Essay Scoring with Language Models
- ARMA Block: A CNN-Based Autoregressive and Moving Average Module for Long-Term Time Series Forecasting
- Event Camera Guided Visual Media Restoration & 3D Reconstruction: A Survey
- OnlineHOI: Towards Online Human-Object Interaction Generation and Perception
- DyKen-Hyena: Dynamic Kernel Generation via Cross-Modal Attention for Multimodal Intent Recognition
- Towards Understanding Visual Grounding in Visual Language Models
- Purge-Gate: Backpropagation-Free Test-Time Adaptation for Point Clouds Classification via Token Purging
- FlexiD-Fuse: Flexible number of inputs multi-modal medical image fusion based on diffusion model
- FS-Diff: Semantic guidance and clarity-aware simultaneous multimodal image fusion and super-resolution
- AWM-Fuse: Multi-Modality Image Fusion for Adverse Weather via Global and Local Text Perception
- First-order State Space Model for Lightweight Image Super-resolution
- Hyperspectral Mamba for Hyperspectral Object Tracking
- Customizing the Inductive Biases of Softmax Attention using Structured Matrices
- Multimodal Contrastive Pretraining of CBCT and IOS for Enhanced Tooth Segmentation
- DEPFusion: Dual-Domain Enhancement and Priority-Guided Mamba Fusion for UAV Multispectral Object Detection
- TransMPC: Transformer-based Explicit MPC with Variable Prediction Horizon
- Astra: A Multi-Agent System for GPU Kernel Performance Optimization
- Testing chatbots on the creation of encoders for audio conditioned image generation
- Electricity Demand and Grid Impacts of AI Data Centers: Challenges and Prospects
- Scaling Transformer-Based Novel View Synthesis Models with Token Disentanglement and Synthetic Data
- Beyond Memorization: Extending Reasoning Depth with Recurrence, Memory and Test-Time Compute Scaling
- PUUMA (Placental patch and whole-Uterus dual-branch U-Mamba-based Architecture): Functional MRI Prediction of Gestational Age at Birth and Preterm Risk
- Effectively obtaining acoustic, visual and textual data from videos
- Hyperbolic Large Language Models
- TreeGPT: Pure TreeFFN Encoder-Decoder Architecture for Structured Reasoning Without Attention Mechanisms
- PLaMo 2 Technical Report
- Exploring Non-Local Spatial-Angular Correlations with a Hybrid Mamba-Transformer Framework for Light Field Super-Resolution
- Elucidating the Design Space of Decay in Linear Attention
- ACE-RL: Adaptive Constraint-Enhanced Reward for Long-form Generation Reinforcement Learning
- CD-Mamba: Cloud detection with long-range spatial dependency modeling
- MambaLite-Micro: Memory-Optimized Mamba Inference on MCUs
- VCMamba: Bridging Convolutions with Multi-Directional Mamba for Efficient Visual Representation
- Echo State Networks as State-Space Models: A Systems Perspective
- Rethinking the long-range dependency in Mamba/SSM and transformer models
- SAC-MIL: Spatial-Aware Correlated Multiple Instance Learning for Histopathology Whole Slide Image Classification
- Human Motion Video Generation: A Survey
- Focus Through Motion: RGB-Event Collaborative Token Sparsification for Efficient Object Detection
- The Personality Illusion: Revealing Dissociation Between Self-Reports & Behavior in LLMs
- Time-Scaling State-Space Models for Dense Video Captioning
- RTGMFF: Enhanced fMRI-based Brain Disorder Diagnosis via ROI-driven Text Generation and Multimodal Feature Fusion
- AR-KAN: Autoregressive-Weight-Enhanced Kolmogorov-Arnold Network for Time Series Forecasting
- S2M2ECG: Spatio-temporal bi-directional State Space Model Enabled Multi-branch Mamba for ECG
- Agentic AI Empowered Multi-UAV Trajectory Optimization in Low-Altitude Economy Networks
- Multi-level SSL Feature Gating for Audio Deepfake Detection
- TRACE: Unlocking Effective CXL Bandwidth via Lossless Compression and Precision Scaling
- ESTM: An Enhanced Dual-Branch Spectral-Temporal Mamba for Anomalous Sound Detection
- Comprehensive Analysis and Exclusion Hypothesis of α-Approximation Method for Discretizing Analog Systems
- AudioRWKV: Efficient and Stable Bidirectional RWKV for Audio Pattern Recognition
- DSGC-Net: A Dual-Stream Graph Convolutional Network for Crowd Counting via Feature Correlation Mining
- Mamba-CNN: A Hybrid Architecture for Efficient and Accurate Facial Beauty Prediction
- Learn to Jump: Adaptive Random Walks for Long-Range Propagation through Graph Hierarchies
- A Unified Voxel Diffusion Module for Point Cloud 3D Object Detection
- SpectMamba: Integrating Frequency and State Space Models for Enhanced Medical Image Detection
- Self-Exploring Language Models for Explainable Link Forecasting on Temporal Graphs via Reinforcement Learning
- SCOUT: Toward Sub-Quadratic Attention via Segment Compression for Optimized Utility in Transformers
- DarkVRAI: Capture-Condition Conditioning and Burst-Order Selective Scan for Low-light RAW Video Denoising
- CSFMamba: Cross State Fusion Mamba Operator for Multimodal Remote Sensing Image Classification
- IndiaWeatherBench: A Dataset and Benchmark for Data-Driven Regional Weather Forecasting over India
- MV-SSM: Multi-View State Space Modeling for 3D Human Pose Estimation
- COMET: A Framework for Modeling Compound Operation Dataflows with Explicit Collectives
- Quantum-Optimized Selective State Space Model for Efficient Time Series Prediction
- Scaling Legal AI: Benchmarking Mamba and Transformers for Statutory Classification and Case Law Retrieval
- SincQDR-VAD: A Noise-Robust Voice Activity Detection Framework Leveraging Learnable Filters and Ranking-Aware Optimization
- Uncovering the Spectral Bias in Diagonal State Space Models
- Enhancing Mamba Decoder with Bidirectional Interaction in Multi-Task Dense Prediction
- HiddenObject: Modality-Agnostic Fusion for Multimodal Hidden Object Detection
- ASCMamba: Multimodal Time-Frequency Mamba for Acoustic Scene Classification
- Deep Equilibrium Convolutional Sparse Coding for Hyperspectral Image Denoising
- Task-Generalized Adaptive Cross-Domain Learning for Multimodal Image Fusion
- Atrial Fibrillation Prediction Using a Lightweight Temporal Convolutional and Selective State Space Architecture
- Autoregressive Universal Video Segmentation Model
- The Ramon Llull's Thinking Machine for Automated Ideation
- Boosting Micro-Expression Analysis via Prior-Guided Video-Level Regression
- Hierarchical Spatio-temporal Segmentation Network for Ejection Fraction Estimation in Echocardiography Videos
- Revisiting associative recall in modern recurrent models
- Parallelizing MCMC Across the Sequence Length
- The Computational Complexity of Satisfiability in State Space Models
- Frozen in Time: Parameter-Efficient Time Series Transformers via Reservoir-Induced Feature Expansion and Fixed Random Dynamics
- Comprehensively stratifying MCIs into distinct risk subtypes based on brain imaging genetics fusion learning
- Survey of Vision-Language-Action Models for Embodied Manipulation
- Hydra: A Modular Architecture for Efficient Long-Context Reasoning
- Beyond Individuals: Collective Predictive Coding for Memory, Attention, and the Emergence of Language
- Mamba2 Meets Silence: Robust Vocal Source Separation for Sparse Regions
- DIME-Net: A Dual-Illumination Adaptive Enhancement Network Based on Retinex and Mixture-of-Experts
- Learning to See Through Flare
- Diversity-enhanced Collaborative Mamba for Semi-supervised Medical Image Segmentation
- Towards Efficient Vision State Space Models via Token Merging
- Prediction of Hospital Associated Infections During Continuous Hospital Stays
- Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey
- Point upsampling networks for single-photon sensing
- HRS: Hybrid Representation Framework with Scheduling Awareness for Time Series Forecasting in Crowdsourced Cloud-Edge Platforms
- Score-informed Neural Operator for Enhancing Ordering-based Causal Discovery
- The impact of tokenizer selection in genomic language models
- Towards High-Resolution Industrial Image Anomaly Detection
- STM3: Mixture of Multiscale Mamba for Long-Term Spatio-Temporal Time-Series Prediction
- Cost-Aware Contrastive Routing for LLMs
- SRMA-Mamba: Spatial Reverse Mamba Attention Network for Pathological Liver Segmentation in MRI Volumes
- The Cultural Gene of Large Language Models: A Study on the Impact of Cross-Corpus Training on Model Values and Biases
- LocoMamba: Vision-Driven Locomotion via End-to-End Deep Reinforcement Learning with Mamba
- ENA: Efficient N-dimensional Attention
- HPD: Hybrid Projection Decomposition for Robust State Space Models on Analog CIM Hardware
- TrajSV: A Trajectory-based Model for Sports Video Representations and Applications
- Multi-State Tracker: Enhancing Efficient Object Tracking via Multi-State Specialization and Interaction
- Denoise-then-Retrieve: Text-Conditioned Video Denoising for Video Moment Retrieval
- Human-in-Context: Unified Cross-Domain 3D Human Motion Modeling via In-Context Learning
- Advances in Speech Separation: Techniques, Challenges, and Future Trends
- Privacy-enhancing Sclera Segmentation Benchmarking Competition: SSBC 2025
- GCRPNet: Graph-Enhanced Contextual and Regional Perception Network for Salient Object Detection in Optical Remote Sensing Images
- EgoMusic-driven Human Dance Motion Estimation with Skeleton Mamba
- Trajectory-aware Shifted State Space Models for Online Video Super-Resolution
- eMamba: Efficient Acceleration Framework for Mamba Models in Edge Computing
- Leveraging OS-Level Primitives for Robotic Action Management
- Deep Learning for Crack Detection: A Review of Learning Paradigms, Generalizability, and Datasets
- rETF-semiSL: Semi-Supervised Learning for Neural Collapse in Temporal Data
- Enhancing Monocular 3D Hand Reconstruction with Learned Texture Priors
- Learning Spatial Decay for Vision Transformers
- Animate-X++: Universal Character Image Animation with Dynamic Backgrounds
- FM4NPP: A Scaling Foundation Model for Nuclear and Particle Physics
- A Signer-Invariant Conformer and Multi-Scale Fusion Transformer for Continuous Sign Language Recognition
- Fake-Mamba: Real-Time Speech Deepfake Detection Using Bidirectional Mamba as Self-Attention's Alternative
- Text-conditioned State Space Model For Domain-generalized Change Detection Visual Question Answering
- Hierarchical Visual Prompt Learning for Continual Video Instance Segmentation
- DeepFleet: Multi-Agent Foundation Models for Mobile Robots
- ME-TST+: Micro-expression Analysis via Temporal State Transition with ROI Relationship Awareness
- Stochastic dynamics learning with state-space systems
- Deep Space Weather Model: Long-Range Solar Flare Prediction from Multi-Wavelength Images
- MambaTrans: Multimodal Fusion Image Translation via Large Language Model Priors for Downstream Visual Tasks
- Parity Requires Unified Input Dependence and Negative Eigenvalues in SSMs
- HSA-Net: Hierarchical and Structure-Aware Framework for Efficient and Scalable Molecular Language Modeling
- Keyword Mamba: Spoken Keyword Spotting with State Space Models
- Consistent and Controllable Image Animation with Motion Linear Diffusion Transformers
- ASM-UNet: Adaptive Scan Mamba Integrating Group Commonalities and Individual Variations for Fine-Grained Segmentation
- Unsupervised Real-World Super-Resolution via Rectified Flow Degradation Modelling
- Similarity Matters: A Novel Depth-guided Network for Image Restoration and A New Dataset
- RORPCap: Retrieval-based Objects and Relations Prompt for Image Captioning
- BrainATCL: Adaptive Temporal Brain Connectivity Learning for Functional Link Prediction and Age Estimation
- Generative AI for Intent-Driven Network Management in 6G: A Case Study on Hierarchical Learning Approach
- Hypergraph Neural Network with State Space Models for Node Classification
- Lightweight Quad Bayer HybridEVS Demosaicing via State Space Augmented Cross-Attention
- Recurrent Deep Differentiable Logic Gate Networks
- ME3-BEV: Mamba-Enhanced Deep Reinforcement Learning for End-to-End Autonomous Driving with BEV-Perception
- MOSEv2: A More Challenging Dataset for Video Object Segmentation in Complex Scenes
- PRvL: Quantifying the Capabilities and Risks of Large Language Models for PII Redaction
- HiSTM: Hierarchical Spatiotemporal Mamba for Cellular Traffic Forecasting
- SONAR-LLM: Autoregressive Transformer that Thinks in Sentence Embeddings and Speaks in Tokens
- FlowState: Sampling Rate Invariant Time Series Forecasting
- HiFi-Mamba: Dual-Stream W-Laplacian Enhanced Mamba for High-Fidelity MRI Reconstruction
- S2M-Former: Spiking Symmetric Mixing Branchformer for Brain Auditory Attention Detection
- Generative Artificial Intelligence in Medical Imaging: Foundations, Progress, and Clinical Translation
- MedMambaLite: Hardware-Aware Mamba for Medical Image Classification
- Log2Sig: Frequency-Aware Insider Threat Detection via Multivariate Behavioral Signal Decomposition
- MambaITD: An Efficient Cross-Modal Mamba Network for Insider Threat Detection
- RetinexDual: Retinex-based Dual Nature Approach for Generalized Ultra-High-Definition Image Restoration
- Position: The Current AI Conference Model is Unsustainable! Diagnosing the Crisis of Centralized AI Conference
- Share Your Attention: Transformer Weight Sharing via Matrix-based Dictionary Learning
- DDTracking: A Deep Generative Framework for Diffusion MRI Tractography with Streamline Local-Global Spatiotemporal Modeling
- PRISM: Lightweight Multivariate Time-Series Classification through Symmetric Multi-Resolution Convolutional Layers
- Small transformer architectures for task switching
- BadTime: An Effective Backdoor Attack on Multivariate Long-Term Time Series Forecasting
- Uncertainty-Aware Spatial Color Correlation for Low-Light Image Enhancement
- TNet: Terrace Convolutional Decoder Network for Remote Sensing Image Semantic Segmentation
- Prototype-Driven Structure Synergy Network for Remote Sensing Images Segmentation
- CodonMoE: DNA Language Models for mRNA Analyses
- Veila: Panoramic LiDAR Generation from a Monocular RGB Image
- Minimal Convolutional RNNs Accelerate Spatiotemporal Learning
- Revisiting Heat Flux Analysis of Tungsten Monoblock Divertor on EAST using Physics-Informed Neural Network
- Beyond Illumination: Fine-Grained Detail Preservation in Extreme Dark Image Restoration
- Rethinking Selectivity in State Space Models: A Minimal Predictive Sufficiency Approach
- SSFMamba: Symmetry-driven Spatial-Frequency Feature Fusion for 3D Medical Image Segmentation
- ClinicalFMamba: Advancing Clinical Assessment using Mamba-based Multimodal Neuroimaging Fusion
- Mamba-X: An End-to-End Vision Mamba Accelerator for Edge Computing Devices
- JustDense: Just using Dense instead of Sequence Mixer for Time Series analysis
- ByteGen: A Tokenizer-Free Generative Model for Orderbook Events in Byte Space
- Content-Aware Mamba for Learned Image Compression
- After the Party: Navigating the Mapping From Color to Ambient Lighting
- Trainable Dynamic Mask Sparse Attention
- DeflareMamba: Hierarchical Vision Mamba for Contextually Consistent Lens Flare Removal
- GlaBoost: A multimodal Structured Framework for Glaucoma Risk Stratification
- Pulse Shape Discrimination Algorithms: Survey and Benchmark
- Self-Navigated Residual Mamba for Universal Industrial Anomaly Detection
- DMTrack: Spatio-Temporal Multimodal Tracking via Dual-Adapter
- RoadMamba: A Dual Branch Visual State Space Model for Road Surface Classification
- PIF-Net: Ill-Posed Prior Guided Multispectral and Hyperspectral Image Fusion via Invertible Mamba and Fusion-Aware LoRA
- Mamba for Wireless Communications and Networking: Principles and Opportunities
- EdgeInfinite-Instruct: Bridging SFT-Based Optimization and NPU-Level Efficiency for Edge Devices
- Multimodal Referring Segmentation: A Survey
- Guided Depth Map Super-Resolution via Multi-Scale Fusion U-shaped Mamba Network
- Stable at Any Speed: Speed-Driven Multi-Object Tracking with Learnable Kalman Filtering
- UIS-Mamba: Exploring Mamba for Underwater Instance Segmentation via Dynamic Tree Scan and Hidden State Weaken
- MVHybrid: Improving Spatial Transcriptomics Prediction with Hybrid State Space-Vision Transformer Backbone in Pathology Vision Foundation Models
- TriP-LLM: A Tri-Branch Patch-wise Large Language Model Framework for Time-Series Anomaly Detection
- Mamba-based Efficient Spatio-Frequency Motion Perception for Video Camouflaged Object Detection
- MamV2XCalib: V2X-based Target-less Infrastructure Camera Calibration with State Space Model
- Merging Memory and Space: A State Space Neural Operator
- VMatcher: State-Space Semi-Dense Local Feature Matching
- Hybrid CNN-Mamba Enhancement Network for Robust Multimodal Sentiment Analysis
- Quantifying surprise in clinical care: Detecting highly informative events in electronic health records with foundation models
Discussions
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces [hn, 130 points, 37 comments]
- MoE-Mamba: Efficient Selective State Space Models with Mixture of Experts [lemmy, 19 points, 0 comments]
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces [lemmy, 8 points, 1 comments]
- Mamba model for sequence modeling, offering linear-time processing & higher efficiency compared to Transformers.Paper by Albert Gu & Tri Dao for insights on selective state spaces in deep learning. #A [bsky, 4 points, 0 comments]
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces [hn, 3 points, 0 comments]
- The paper "Mamba: Linear-Time Sequence Modeling with Selective State Spaces" presents a new model that enhances efficiency via selective state spaces, allowing faster inference and linear scaling for [bsky, 2 points, 0 comments]
- arxiv.org/abs/2312.00752 github.com/havenhq/mamb... www.linkedin.com/posts/datapi... [bsky, 2 points, 0 comments]
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces [lobsters, 2 points, 0 comments]
- Was catching up on a massive (year+) backlog of papers and came across Mamba: arxiv.org/abs/2312.00752 Seems like a promising approach - I really like the idea of being hardware-aware and running sca [bsky, 1 points, 0 comments]
- 選択的状態空間モデルだなんて,やだこれ萌えちゃう。 https://arxiv.org/abs/2312.00752 [bsky, 0 points, 0 comments]
- Late to the party... was looking through this paper about giant snakes and realize I need to follow up on the use of chinchillas, hyenas, hippopotamuses, and various camelids in state-of-the-art AI h [bsky, 0 points, 0 comments]
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces https://lobste.rs/s/ntv2lz #ai [bsky, 0 points, 0 comments]
- Cool new #AI stuff: 1. “Mamba: Linear-Time Sequence Modeling With Selective State Spaces”, Albert Gu & Tri Dao (arxiv.org/abs/2312.00752). On HN: news.ycombinator.com/item?id=3852... 2. “Mixtral” 8 [bsky, 0 points, 1 comments]
Related