Neural Discrete Representation Learning
2017/11/02 by Aaron van den Oord, Aäron van den Oord, Oord, Aaron van den +4 · 2 voices · 1,207 citations
Computer Science · #FOS: Computer and information sciences #Machine Learning (cs.LG) #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing #cs.LG
paper · pdf · doi:10.48550/arxiv.1711.00937
openalex publication_date 2017/11/02 · arxiv published 2017/11/02 · arxiv created 2018/05/30 · arxiv updated 2018/05/31 · openalex created_date 2019/07/30 · openalex updated_date 2026/07/28
Abstract
Learning useful representations without supervision remains a key challenge in machine learning. In this paper, we propose a simple yet powerful generative model that learns such discrete representations. Our model, the Vector Quantised-Variational AutoEncoder (VQ-VAE), differs from VAEs in two key ways: the encoder network outputs discrete, rather than continuous, codes; and the prior is learnt rather than static. In order to learn a discrete latent representation, we incorporate ideas from vector quantisation (VQ). Using the VQ method allows the model to circumvent issues of "posterior collapse" -- where the latents are ignored when they are paired with a powerful autoregressive decoder -- typically observed in the VAE framework. Pairing these representations with an autoregressive prior, the model can generate high quality images, videos, and speech as well as doing high quality speaker conversion and unsupervised learning of phonemes, providing further evidence of the utility of the learnt representations.
Cited by
- Marrying Autoregressive Transformer and Diffusion with Multi-Reference Autoregression
- ThinkGen: Generalized Thinking for Visual Generation
- Quantum Generative Models for Computational Fluid Dynamics: A First Exploration of Latent Space Learning in Lattice Boltzmann Simulations
- Visual Autoregressive Modelling for Monocular Depth Estimation
- CogRec: Structure-Cognitive Fast-and-Slow Reasoning for Generative Recommendation
- EEGForceFusion: Joint Tokenised-Continuous Representation Learning for Subject-Independent Grasp Force Decoding
- Envision: Embodied Visual Planning via Goal-Imagery Video Diffusion
- CoDS: Collaborative Perception via Digital Semantic Communication
- Pose-Guided Residual Refinement for Interpretable Text-to-Motion Generation and Editing
- Learning Reusable Hybrid Motion Priors for Humanoid Locomotion from Motion Imitation
- Codebook Capacity Governs Perceptual Quality Across Resolutions in Hierarchical Discrete Video Compression
- mmSimPrior: Learning Simulation Priors for Data-Efficient and Generalizable Real-World Radar-based Human Motion Reconstruction
- TreeAdapter: Hierarchical Taxonomy-Guided Adapter Composition for Fine-Grained Species Image Generation
- Where Quality Breaks in Compressed Short-Text Generation: Staged Bottleneck Localization
- UniGen-AR: Unifying Visual Generation with Auto-Regressive Modeling
- PriSAR: 3D Geometric-Prior-Guided Diffusion for Parameter-Controlled SAR Image Generation
- Argus-Unified: Towards A Compact and Economical Unified Model for Image Understanding and Generation
- Tokens are All You Need: Dual-purpose Semantic IDs for Achieving LLM-Level I/O Efficiency in recommendation systems
- A Diagnostic Gap Framework for Evaluating Reconstruction Fidelity in Weakly Supervised Mammography
- MIME: Multimodal Interactive Motion Encoder
- ForgeDrive: Bidirectional Cross-Conditioning for Unified Visual-Action Generation in Autonomous Driving
- A General Deep Learning Framework for Wireless Resource Allocation under Discrete Constraints
- The Equalizer: Introducing Shape-Gain Decomposition in Neural Audio Codecs
- GQ-VAE: A gated quantized VAE for learning variable length tokens
- Fast Inference of Visual Autoregressive Model with Adjacency-Adaptive Dynamical Draft Trees
- LibContinual: A Comprehensive Library towards Realistic Continual Learning
- Residual Prior Diffusion: A Probabilistic Framework Integrating Coarse Latent Priors with Diffusion Models
- Videos are Sample-Efficient Supervisions: Behavior Cloning from Videos via Latent Representations
- Multi-Attribute guided Thermal Face Image Translation based on Latent Diffusion Model
- FluencyVE: Marrying Temporal-Aware Mamba with Bypass Attention for Video Editing
- SegMo: Segment-aligned Text to 3D Human Motion Generation
- Branch Learning in MRI: More Data, More Models, More Training
- LiDARDraft: Generating LiDAR Point Cloud from Versatile Inputs
- Aliasing-Free Neural Audio Synthesis
- Generative Latent Coding for Ultra-Low Bitrate Image Compression
- VA-π: Variational Policy Alignment for Pixel-Aware Autoregressive Generation
- Towards Closed-Loop Embodied Empathy Evolution: Probing LLM-Centric Lifelong Empathic Motion Generation in Unseen Scenarios
- Emotion-Director: Bridging Affective Shortcut in Emotion-Oriented Image Generation
- Memorize-and-Generate: Towards Long-Term Consistency in Real-Time Video Generation
- Social Comparison without Explicit Inference of Others' Reward Values: A Constructive Approach Using a Probabilistic Generative Model
- MACE-Dance: Motion-Appearance Cascaded Experts for Music-Driven Dance Video Generation
- Efficient Optimization of Hierarchical Identifiers for Generative Recommendation
- Learning Semantic Atomic Skills for Multi-Task Robotic Manipulation
- Diffusion Forcing for Multi-Agent Interaction Sequence Modeling
- Reasoning Palette: Modulating Reasoning via Latent Contextualization for Controllable Exploration for (V)LMs
- Next-Embedding Prediction Makes Strong Vision Learners
- SFTok: Bridging the Performance Gap in Discrete Tokenizers
- KineST: A Kinematics-guided Spatiotemporal State Space Model for Human Motion Tracking from Sparse Signals
- Prime and Reach: Synthesising Body Motion for Gaze-Primed Object Reach
- BEST-STD2.0: Balanced and Efficient Speech Tokenizer for Spoken Term Detection
- InfoTok: Adaptive Discrete Video Tokenizer via Information-Theoretic Compression
- Vision-Language-Action Models for Autonomous Driving: Past, Present, and Future
- Evaluation of Generative Models for Emotional 3D Animation Generation in VR
- OccSTeP: Benchmarking 4D Occupancy Spatio-Temporal Persistence
- Audio-Visual Cross-Modal Compression for Generative Face Video Coding
- HD-Prot: A Protein Language Model for Joint Sequence-Structure Modeling with Continuous Structure Tokens
- Spherical Leech Quantization for Visual Tokenization and Generation
- OmniGen: Unified Multimodal Sensor Generation for Autonomous Driving
- Smudged Fingerprints: A Systematic Evaluation of the Robustness of AI Image Fingerprints
- ViewMask-1-to-3: Multi-View Consistent Image Generation via Multimodal Discrete Diffusion Models
- RecTok: Reconstruction Distillation along Rectified Flow
- DisCo-Speech: Controllable Zero-Shot Speech Generation with A Disentangled Speech Codec
- Towards Unified Co-Speech Gesture Generation via Hierarchical Implicit Periodicity Learning
- CoRe3D: Collaborative Reasoning as a Foundation for 3D Intelligence
- GenieDrive: Towards Physics-Aware Driving World Model with 4D Occupancy Guided Video Generation
- Skillful Subseasonal-to-Seasonal Forecasting of Extreme Events with a Multi-Sphere Coupled Probabilistic Model
- Ultra-Low Bitrate Perceptual Image Compression with Shallow Encoder
- EditMGT: Unleashing Potentials of Masked Generative Transformers in Image Editing
- BAgger: Backwards Aggregation for Mitigating Drift in Autoregressive Video Diffusion Models
- ARC-AGI Without Pretraining
- Autoregressive Video Autoencoder with Decoupled Temporal and Spatial Context
- An Anatomy of Vision-Language-Action Models: From Modules to Milestones and Challenges
- WholeBodyVLA: Towards Unified Latent VLA for Whole-Body Loco-Manipulation Control
- Closing the Train-Test Gap in World Models for Gradient-Based Planning
- A Unified Generative-Predictive Framework for Deterministic Inverse Design
- Training One Model to Master Cross-Level Agentic Actions via Reinforcement Learning
- UniLS: End-to-End Audio-Driven Avatars for Unified Listening and Speaking
- Masked Generative Policy for Robotic Control
- Fast-ARDiff: An Entropy-informed Acceleration Framework for Continuous Space Autoregressive Generation
- Learning to Control Physically-simulated 3D Characters via Generating and Mimicking 2D Motions
- Error-Resilient Semantic Communication for Speech Transmission over Packet-Loss Networks
- Chain-of-Image Generation: Toward Monitorable and Controllable Image Generation
- Voxify3D: Pixel Art Meets Volumetric Rendering
- Demo: Generative AI helps Radiotherapy Planning with User Preference
- Training-Free Vector Quantization via Gaussian VAEs
- MeshRipple: Structured Autoregressive Generation of Artist-Meshes
- M-STAR: Multi-Scale Spatiotemporal Autoregression for Human Mobility Modeling
- SUCCESS-GS: Survey of Compactness and Compression for Efficient Static and Dynamic Gaussian Splatting
- Self-Supervised Learning on Molecular Graphs: A Systematic Investigation of Masking Design
- See Once, Then Act: Vision-Language-Action Model with Task Learning from One-Shot Video Demonstrations
- Foresight Prediction Enhanced Live-Streaming Recommendation
- Rethinking Training Dynamics in Scale-wise Autoregressive Generation
- World Models That Know When They Don't Know - Controllable Video Generation with Calibrated Uncertainty
- EmoStyle: Emotion-Driven Image Stylization
- Geometry-Aware Single-Image 4D Synthesis via Dense Trajectory Generation
- Efficient Generative Transformer Operators For Million-Point PDEs
- FASTer: Toward Efficient Autoregressive Vision Language Action Modeling via Neural Action Tokenization
- SparseWorld-TC: Trajectory-Conditioned Sparse Occupancy World Model
- Autoregressive Image Generation Needs Only a Few Lines of Cached Tokens
- Order Matters: 3D Shape Generation from Sequential VR Sketches
- Denoise to Track: Harnessing Video Diffusion Priors for Robust Correspondence
- TARDis: Time Attenuated Representation Disentanglement for Incomplete Multi-Modal Tumor Segmentation and Classification
- COOPER: A Unified Model for Cooperative Perception and Reasoning in Spatial Intelligence
- Controllable Long-term Motion Generation with Extended Joint Targets
- DeRA: Decoupled Representation Alignment for Video Tokenization
- LSRS: Latent Scale Rejection Sampling for Visual Autoregressive Modeling
- FloodDiffusion: Tailored Diffusion Forcing for Streaming Motion Generation
- AutoBrep: Autoregressive B-Rep Generation with Unified Topology and Geometry
- Mitigating Intra- and Inter-modal Forgetting in Continual Learning of Unified Multimodal Models
- U4D: Uncertainty-Aware 4D World Modeling from LiDAR Sequences
- Graph VQ-Transformer (GVT): Fast and Accurate Molecular Generation via High-Fidelity Discrete Latents
- Contrastive Deep Learning for Variant Detection in Wastewater Genomic Sequencing
- TUNA: Taming Unified Visual Representations for Native Unified Multimodal Models
- Real-World Robot Control by Deep Active Inference With a Temporally Hierarchical World Model
- Deconstructing Generative Diversity: An Information Bottleneck Analysis of Discrete Latent Generative Models
- BrepGPT: Autoregressive B-rep Generation with Voronoi Half-Patch
- Mofasa: A Step Change in Metal-Organic Framework Generation
- Two-Dimensional Quantization for Geometry-Aware Audio Coding
- Lotus-2: Advancing Geometric Dense Prediction with Powerful Image Generative Model
- Neural Discrete Representation Learning for Sparse-View CBCT Reconstruction: From Algorithm Design to Prospective Multicenter Clinical Evaluation
- Smol-GS: Compact Representations for Abstract 3D Gaussian Splatting
- ReactionMamba: Generating Short & Long Human Reaction Sequences
- Visual Generation Tuning
- Quantized-Tinyllava: a new multimodal foundation model enables efficient split learning
- VQRAE: Representation Quantization Autoencoders for Multimodal Understanding, Generation and Reconstruction
- REVEAL: Reasoning-Enhanced Forensic Evidence Analysis for Explainable AI-Generated Image Detection
- Markovian Scale Prediction: A New Era of Visual Autoregressive Generation
- PURE Codec: Progressive Unfolding of Residual Entropy for Speech Codec Learning
- AIA: Rethinking Architecture Decoupling Strategy In Unified Multimodal Model
- Harmonic-Percussive Disentangled Neural Audio Codec for Bandwidth Extension
- DiverseVAR: Balancing Diversity and Quality of Next-Scale Visual Autoregressive Models
- Infinite-Story: A Training-Free Consistent Text-to-Image Generation
- RAVQ-HoloNet: Rate-Adaptive Vector-Quantized Hologram Compression
- Progress by Pieces: Test-Time Scaling for Autoregressive Image Generation
- Operationalizing Quantized Disentanglement
- DINO-Tok: Adapting DINO for Visual Tokenizers
- SKEL-CF: Coarse-to-Fine Biomechanical Skeleton and Surface Mesh Recovery
- Text-guided Controllable Diffusion for Realistic Camouflage Images Generation
- PRADA: Probability-Ratio-Based Attribution and Detection of Autoregressive-Generated Images
- AI/ML based Joint Source and Channel Coding for HARQ-ACK Payload
- Temporal-Visual Semantic Alignment: A Unified Architecture for Transferring Spatial Priors from Vision Models to Zero-Shot Temporal Tasks
- SSA: Sparse Sparse Attention by Aligning Full and Sparse Attention Outputs in Feature Space
- fMRI-LM: Towards a Universal Foundation Model for Language-Aligned fMRI Understanding
- Learning Massively Multitask World Models for Continuous Control
- Multiscale Vector-Quantized Variational Autoencoder for Endoscopic Image Synthesis
- Robust Long-term Test-Time Adaptation for 3D Human Pose Estimation through Motion Discretization
- FVAR: Visual Autoregressive Modeling via Next Focus Prediction
- DELTA: Language Diffusion-based EEG-to-Text Architecture
- TRIDENT: A Trimodal Cascade Generative Framework for Drug and RNA-Conditioned Cellular Morphology Synthesis
- Spanning Tree Autoregressive Visual Generation
- METIS: Multi-Source Egocentric Training for Integrated Dexterous Vision-Language-Action Model
- LAOF: Robust Latent Action Learning with Optical Flow Constraints
- Progressive Supernet Training for Efficient Visual Autoregressive Modeling
- Flow and Depth Assisted Video Prediction with Latent Transformer
- Mem-MLP: Real-Time 3D Human Motion Generation from Sparse Inputs
- LiSTAR: Ray-Centric World Models for 4D LiDAR Sequences in Autonomous Driving
- AMS-KV: Adaptive KV Caching in Multi-Scale Visual Autoregressive Transformers
- MamTiff-CAD: Multi-Scale Latent Diffusion with Mamba+ for Complex Parametric Sequence
- CoordAR: One-Reference 6D Pose Estimation of Novel Objects via Autoregressive Coordinate Map Generation
- Taming Generative Synthetic Data for X-ray Prohibited Item Detection
- Learning Human-Like RL Agents Through Trajectory Optimization With Action Quantization
- WiCo-MG: Wireless Channel Foundation Model for Multipath Generation via Synesthesia of Machines
- GeoSceneGraph: Geometric Scene Graph Diffusion Model for Text-guided 3D Indoor Scene Synthesis
- B-Rep Distance Functions (BR-DF): How to Represent a B-Rep Model by Volumetric Distance Functions?
- Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech
- StreamingTalker: Audio-driven 3D Facial Animation with Autoregressive Diffusion Model
- Semantic Context Matters: Improving Conditioning for Autoregressive Models
- Self-Supervised Compression and Artifact Correction for Streaming Underwater Imaging Sonar
- ActVAR: Activating Mixtures of Weights and Tokens for Efficient Visual Autoregressive Generation
- MergeDNA: Context-aware Genome Modeling with Dynamic Tokenization through Token Merging
- InterMoE: Individual-Specific 3D Human Interaction Generation via Dynamic Temporal-Selective MoE
- Generalized Denoising Diffusion Codebook Models (gDDCM): Tokenizing images using a pre-trained diffusion model
- DAP: A Discrete-token Autoregressive Planner for Autonomous Driving
- ExplicitLM: Decoupling Knowledge from Parameters via Explicit Memory Banks
- Through-Foliage Surface-Temperature Reconstruction for Early Wildfire Detection
- DINO-Detect: A Simple yet Effective Framework for Blur-Robust AI-Generated Image Detection
- VLA-R: Vision-Language Action Retrieval toward Open-World End-to-End Autonomous Driving
- DEMIST: Decoupled Multi-stream latent diffusion for Quantitative Myelin Map Synthesis
- Seg-VAR: Image Segmentation with Visual Autoregressive Modeling
- LiDAR-GS++:Improving LiDAR Gaussian Reconstruction via Diffusion Priors
- Evaluating Latent Generative Paradigms for High-Fidelity 3D Shape Completion from a Single Depth Image
- MixAR: Mixture Autoregressive Image Generation
- ReCast: Reliability-aware Codebook Assisted Lightweight Time Series Forecasting
- Improved Masked Image Generation with Knowledge-Augmented Token Representations
- Protein Structure Tokenization via Geometric Byte Pair Encoding
- Point Cloud Quantization through Multimodal Prompting for 3D Understanding
- A Style is Worth One Code: Unlocking Code-to-Style Image Generation with Discrete Style Space
- Stroke Modeling Enables Vectorized Character Generation with Large Vectorized Glyph Model
- Semantic Communication with Hopfield Memories
- Towards Leveraging Sequential Structure in Animal Vocalizations
- Optimizing Input of Denoising Score Matching is Biased Towards Higher Score Norm
- Learning Binary Autoencoder-Based Codes with Progressive Training
- TransactionGPT
- Large Sign Language Models: Toward 3D American Sign Language Translation
- Embodied Cognition Augmented End2End Autonomous Driving
- Generative AI Meets 6G and Beyond: Diffusion Models for Semantic Communications
- Retrospective motion correction in MRI using disentangled embeddings
- Re-coding for Uncertainties: Edge-awareness Semantic Concordance for Resilient Event-RGB Segmentation
- From IDs to Semantics: A Generative Framework for Cross-Domain Recommendation with Adaptive Semantic Tokenization
- MonoCLUE : Object-Aware Clustering Enhances Monocular 3D Object Detection
- SONIC: Supersizing Motion Tracking for Natural Humanoid Whole-Body Control
- ViPRA: Video Prediction for Robot Actions
- Rethinking Generative Image Pretraining: How Far Are We From Scaling Up Next-Pixel Prediction?
- VAEVQ: Enhancing Discrete Visual Tokenization through Variational Modeling
- CAST-LUT: Tokenizer-Guided HSV Look-Up Tables for Purple Flare Removal
- PuzLM: Solving Jigsaw Puzzles with Sequence-to-Sequence Language Models
- Enhancing Multimodal Misinformation Detection by Replaying the Whole Story from Image Modality Perspective
- Multimodal Diffusion Forcing for Forceful Manipulation
- InfinityStar: Unified Spacetime AutoRegressive Modeling for Visual Generation
- MoSa: Motion Generation with Scalable Autoregressive Modeling
- PhysCorr: Dual-Reward DPO for Physics-Constrained Text-to-Video Generation with Automated Preference Selection
- Speech-Based Prioritization for Schizophrenia Intervention
- Unified Diffusion VLA: Vision-Language-Action Model via Joint Discrete Denoising Diffusion Process
- XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations
- Wave-Particle (Continuous-Discrete) Dualistic Visual Tokenization for Unified Understanding and Generation
- iFlyBot-VLA Technical Report
- MalDataGen: A Modular Framework for Synthetic Tabular Data Generation in Malware Detection
- Co-Evolving Latent Action World Models
- Modeling strategies for speech enhancement in the latent space of a neural audio codec
- Sketch2PoseNet: Efficient and Generalized Sketch to 3D Human Pose Prediction
- Alpamayo-R1: Bridging Reasoning and Action Prediction for Generalizable Autonomous Driving in the Long Tail
- Cloning Deterministic Worlds: The Critical Role of Latent Geometry in Long-Horizon World Models
- Hawk: Leveraging Spatial Context for Faster Autoregressive Text-to-Image Generation
- Product-Quantised Image Representation for High-Quality Image Synthesis
- Input-Aware Sparse Attention for Real-Time Co-Speech Video Generation
- Multi-scale Autoregressive Models are Laplacian, Discrete, and Latent Diffusion Models in Disguise
- Explainable Disentanglement on Discrete Speech Representations for Noise-Robust ASR
- Uniform Discrete Diffusion with Metric Path for Video Generation
- MuseTok: Symbolic Music Tokenization for Generation and Semantic Understanding
- TokenAR: Multiple Subject Generation via Autoregressive Token-level enhancement
- See the Speaker: Crafting High-Resolution Talking Faces from Speech with Prior Guidance and Region Refinement
- Unlocking Out-of-Distribution Generalization in Dynamics through Physics-Guided Augmentation
- MC-SJD : Maximal Coupling Speculative Jacobi Decoding for Autoregressive Visual Generation Acceleration
- OneCast: Structured Decomposition and Modular Generation for Cross-Domain Time Series Forecasting
- A Survey on Efficient Vision-Language-Action Models
- FARMER: Flow AutoRegressive Transformer over Pixels
- More Than Generation: Unifying Generation and Depth Estimation via Text-to-Image Diffusion Models
- SoulX-Podcast: Towards Realistic Long-form Podcasts with Dialectal and Paralinguistic Diversity
- Low-Resource Audio Codec (LRAC): 2025 Challenge Description
- Autoregressive Styled Text Image Generation, but Make it Reliable
- Quantizing Space and Time: Fusing Time Series and Images for Earth Observation
- Nested AutoRegressive Models
- VALA: Learning Latent Anchors for Training-Free and Temporally Consistent
- Switchable Token-Specific Codebook Quantization For Face Image Compression
- Optimize Any Topology: A Foundation Model for Shape- and Resolution-Free Structural Topology Optimization
- NetBurst: Event-Centric Forecasting of Bursty, Intermittent Time Series
- Moving Beyond Diffusion: Hierarchy-to-Hierarchy Autoregression for fMRI-to-Image Reconstruction
- The Lossy Horizon: Error-Bounded Predictive Coding for Lossy Text Compression (Episode I)
- BADiff: Bandwidth Adaptive Diffusion Model
- Pctx: Tokenizing Personalized Context for Generative Recommendation
- SpecTokenizer: A Lightweight Streaming Codec in the Compressed Spectrum Domain
- Distilled Decoding 2: One-step Sampling of Image Auto-regressive Models with Conditional Score Distillation
- Video Prediction of Dynamic Physical Simulations With Pixel-Space Spatiotemporal Transformers
- ARGenSeg: Image Segmentation with Autoregressive Image Generation Model
- Speaking Clearly: A Simplified Whisper-Based Codec for Low-Bitrate Speech Coding
- EditInfinity: Image Editing with Binary-Quantized Generative Models
- SpikeFit: Towards Optimal Deployment of Spiking Networks on Neuromorphic Hardware
- Diffusion Autoencoders with Perceivers for Long, Irregular and Multimodal Astronomical Sequences
- Better Tokens for Better 3D: Advancing Vision-Language Modeling in 3D Medical Imaging
- Exploring Conditions for Diffusion models in Robotic Control
- Adaptive Distribution-aware Quantization for Mixed-Precision Neural Networks
- KnowMol: Advancing Molecular Large Language Models with Multi-Level Chemical Knowledge
- NeSyPr: Neurosymbolic Proceduralization For Efficient Embodied Reasoning
- Multi-Rate Task-Oriented Communication for Multi-Edge Cooperative Inference
- Modeling Turn-Taking with Semantically Informed Gestures
- SITS-DECO: A Generative Decoder Is All You Need For Multitask Satellite Image Time Series Modelling
- GPTFace: Generative Pre-training of Facial-Linguistic Transformer by Span Masking and Weakly Correlated Text-image Data
- UniGenBench++: A Unified Semantic Evaluation Benchmark for Text-to-Image Generation
- Accelerating Vision Transformers with Adaptive Patch Sizes
- MEG-GPT: A transformer-based foundation model for magnetoencephalography data
- OmniCast: A Masked Latent Diffusion Model for Weather Forecasting Across Time Scales
- AION-1: Omnimodal Foundation Model for Astronomical Sciences
- Conveying Meaning through Gestures: An Investigation into Semantic Co-Speech Gesture Generation
- SAC: Neural Speech Codec with Semantic-Acoustic Dual-Stream Quantization
- OmniMotion: Multimodal Motion Generation with Continuous Masked Autoregression
- ScaleWeaver: Weaving Efficient Controllable T2I Generation with Multi-Scale Reference Attention
- LightQANet: Quantized and Adaptive Feature Learning for Low-Light Image Enhancement
- QDepth-VLA: Quantized Depth Prediction as Auxiliary Supervision for Vision-Language-Action Models
- SkyDreamer: Interpretable End-to-End Vision-Based Drone Racing with Model-Based Reinforcement Learning
- Adapting Self-Supervised Representations as a Latent Space for Efficient Generation
- GemiRec: Interest Quantization and Generation for Multi-Interest Recommendation
- DCMIL: A Progressive Representation Learning of Whole Slide Images for Cancer Prognosis Analysis
- Synergistic Integration and Discrepancy Resolution of Contextualized Knowledge for Personalized Recommendation
- Vector Quantization in the Brain: Grid-like Codes in World Models
- Simplicial Embeddings Improve Sample Efficiency in Actor-Critic Agents
- CanvasMAR: Improving Masked Autoregressive Video Prediction With Canvas
- EEGChaT: A Transformer-Based Modular Channel Selector for SEEG Analysis
- Group-Wise Optimization for Self-Extensible Codebooks in Vector Quantized Models
- Adaptive Visual Conditioning for Semantic Consistency in Diffusion-Based Story Continuation
- Universal Image Restoration Pre-training via Masked Degradation Classification
- NExT-OMNI: Towards Any-to-Any Omnimodal Foundation Models with Discrete Flow Matching
- T(R,O) Grasp: Efficient Graph Diffusion of Robot-Object Spatial Transformation for Cross-Embodiment Dexterous Grasping
- A Hierarchical Quantized Tokenization Framework for Task-Adaptive Graph Representation Learning
- BIGFix: Bidirectional Image Generation with Token Fixing
- Your VAR Model is Secretly an Efficient and Explainable Generative Classifier
- Diffusion Transformers with Representation Autoencoders
- Diffusion Models for Reinforcement Learning: Foundations, Taxonomy, and Development
- Pre to Post-Treatment Glioblastoma MRI Prediction using a Latent Diffusion Model
- Test-Time Anchoring for Discrete Diffusion Posterior Sampling
- HiMaCon: Discovering Hierarchical Manipulation Concepts from Unlabeled Multi-Modal Data
- Perspective-aware 3D Gaussian Inpainting with Multi-view Consistency
- SoundReactor: Frame-level Online Video-to-Audio Generation
- VividAnimator: An End-to-End Audio and Pose-driven Half-Body Human Animation Framework
- Are Video Models Emerging as Zero-Shot Learners and Reasoners in Medical Imaging?
- MelTok: 2D Tokenization for Single-Codebook Audio Compression
- Benchmarking Language Modeling for Lossless Compression of Full-Fidelity Audio
- Cross-Sensor Touch Generation
- Prototype-based Self-Supervised Multimodal Learning for PPG and Accelerometry Signals
- Precoder Design in Multi-User FDD Systems with VQ-VAE and GNN
- Mono4DEditor: Text-Driven 4D Scene Editing from Monocular Video via Point-Level Localization of Language-Embedded Gaussians
- RoDyn: Taming Interactive Robot-Dynamic 2.5D World Model for Robotic Manipulation
- Neural Codecs as Biosignal Tokenizers
- Discrete Facial Encoding: : A Framework for Data-driven Facial Display Discovery
- A Theoretically-Grounded Codebook for Digital Semantic Communications
- Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications
- Growing Visual Generative Capacity for Pre-Trained MLLMs
- Heptapod: Language Modeling on Visual Signals
- VUGEN: Visual Understanding priors for GENeration
- A Bridge from Audio to Video: Phoneme-Viseme Alignment Allows Every Face to Speak Multiple Languages
- We Can Hide More Bits: The Unused Watermarking Capacity in Theory and in Practice
- Discrete Diffusion Models with MLLMs for Unified Medical Multimodal Generation
- \bfD3QE: Learning Discrete Distribution Discrepancy-aware Quantization Error for Autoregressive-Generated Image Detection
- BlockGPT: Spatio-Temporal Modelling of Rainfall via Frame-Level Autoregression
- PointNSP: Autoregressive 3D Point Cloud Generation with Next-Scale Level-of-Detail Prediction
- Efficient Conditional Generation on Scale-based Visual Autoregressive Models
- Traj-Transformer: Diffusion Models with Transformer for GPS Trajectory Generation
- Information-Theoretic Policy Pre-Training with Empowerment
- DYMO-Hair: Generalizable Volumetric Dynamics Modeling for Robot Hair Manipulation
- VChain: Chain-of-Visual-Thought for Reasoning in Video Generation
- SSDD: Single-Step Diffusion Decoder for Efficient Image Tokenization
- Provable Speech Attributes Conversion via Latent Independence
- ReactDiff: Fundamental Multiple Appropriate Facial Reaction Diffusion Model
- Language Model Based Text-to-Audio Generation: Anti-Causally Aligned Collaborative Residual Transformers
- REAR: Rethinking Visual Autoregressive Models via Generator-Tokenizer Consistency Regularization
- CodeFormer++: Blind Face Restoration Using Deformable Registration and Deep Metric Learning
- Bridging Text and Video Generation: A Survey
- GenAR: Next-Scale Autoregressive Generation for Spatial Gene Expression Prediction
- MulVuln: Enhancing Pre-trained LMs with Shared and Language-Specific Knowledge for Multilingual Vulnerability Detection
- MASC: Boosting Autoregressive Image Generation with a Manifold-Aligned Semantic Clustering
- Désentrelacement Fréquentiel Doux pour les Codecs Audio Neuronaux
- Soft Disentanglement in Frequency Bands for Neural Audio Codecs
- TIT-Score: Evaluating Long-Prompt Based Text-to-Image Alignment via Text-to-Image-to-Text Consistency
- Flip Distribution Alignment VAE for Multi-Phase MRI Synthesis
- Rate-Adaptive Semantic Communication via Multi-Stage Vector Quantization
- Purrception: Variational Flow Matching for Vector-Quantized Image Generation
- Eliciting Chain-of-Thought Reasoning for Time Series Analysis using Reinforcement Learning
- SoftCFG: Uncertainty-guided Stable Guidance for Visual Autoregressive Model
- Ultra-Efficient Decoding for End-to-End Neural Compression and Reconstruction
- Flow Autoencoders are Effective Protein Tokenizers
- Baseline Systems For The 2025 Low-Resource Audio Codec Challenge
- Learning Energy-based Variational Latent Prior for VAEs
- DiVeQ: Differentiable Vector Quantization Using the Reparameterization Trick
- Go with Your Gut: Scaling Confidence for Autoregressive Image Generation
- MARS: Sound Generation via Multi-Channel Autoregression on Spectrograms
- PUREVQ-GAN: Defending Data Poisoning Attacks through Vector-Quantized Bottlenecks
- LieHMR: Autoregressive Human Mesh Recovery with SO(3) Diffusion
- LaTo: Landmark-tokenized Diffusion Transformer for Fine-grained Human Face Editing
- MoReFlow: Motion Retargeting Learning through Unsupervised Flow Matching
- Aligning Visual Foundation Encoders to Tokenizers for Diffusion Models
- Score-based Membership Inference on Diffusion Models
- Beyond Softmax: A Natural Parameterization for Categorical Random Variables
- Discrete Variational Autoencoding via Policy Search
- Uni-X: Mitigating Modality Conflict with a Two-End-Separated Architecture for Unified Multimodal Models
- Hyperspherical Latents Improve Continuous-Token Autoregressive Generation
- SynthCloner: Synthesizer Preset Conversion via Factorized Codec with ADSR Envelope Control
- Cycle Diffusion Model for Counterfactual Image Generation
- DyMoDreamer: World Modeling with Dynamic Modulation
- SenSE: Semantic-Aware High-Fidelity Universal Speech Enhancement
- Reinforcement Learning with Inverse Rewards for World Model Post-training
- GSID: Generative Semantic Indexing for E-Commerce Product Understanding
- Texture Vector-Quantization and Reconstruction Aware Prediction for Generative Super-Resolution
- ResAD++: Towards Class Agnostic Anomaly Detection via Residual Feature Learning
- HieraTok: Multi-Scale Visual Tokenizer Improves Image Reconstruction and Generation
- AudioMoG: Guiding Audio Generation with Mixture-of-Guidance
- MotionVerse: A Unified Multimodal Framework for Motion Comprehension, Generation and Editing
- DiffInk: Glyph- and Style-Aware Latent Diffusion Transformer for Text to Online Handwriting Generation
- Internal Planning in Language Models: Characterizing Horizon and Branch Awareness
- An Efficient Transfer Learning Method Based on Adapter with Local Attributes for Speech Emotion Recognition
- Entering the Era of Discrete Diffusion Models: A Benchmark for Schrödinger Bridges and Entropic Optimal Transport
- Geometry-Aware Losses for Structure-Preserving Text-to-Sign Language Generation
- ARSS: Taming Decoder-only Autoregressive Visual Generation for View Synthesis From Single View
- Soft-Di[M]O: Improving One-Step Discrete Image Generation with Soft Embeddings
- Group Critical-token Policy Optimization for Autoregressive Image Generation
- One Prompt Fits All: Universal Graph Adaptation for Pretrained Models
- StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs
- Rate-Distortion Optimized Communication for Collaborative Perception
- Developing Vision-Language-Action Model from Egocentric Videos
- AUV: Teaching Audio Universal Vector Quantization with Single Nested Codebook
- Pushing Toward the Simplex Vertices: A Simple Remedy for Code Collapse in Smoothed Vector Quantization
- Residual Vector Quantization For Communication-Efficient Multi-Agent Perception
- CAD-Tokenizer: Towards Text-based CAD Prototyping via Modality-Specific Tokenization
- From Samples to Scenarios: A New Paradigm for Probabilistic Forecasting
- AJAHR: Amputated Joint Aware 3D Human Mesh Recovery
- MUGEN: A Unified Framework for Efficient Motion Understanding and Generation
- Hierarchical Latent Reasoning for LLM-based Recommendation
- Collaborative feature aggregation for face super-resolution and robust re-identification
- OneShot: Index-in-Ranking with Neural Scoring for Large-Scale Retrieval
- Shared Semantic Codebook Distillation for Unpaired Cross-Modal Medical Classification
- Temporal Straightening for Latent Planning
- OSF: On Pre-training and Scaling of Sleep Foundation Models
- Learning a Sampling-Free Variational DNN Plugin from Tiny Training Sets to Refine OOD Segmentation With Uncertainty Estimation
- EEG-FM-Bench: A Comprehensive Benchmark for the Systematic Evaluation of EEG Foundation Models
- Adversarially-Refined VQ-GAN with Dense Motion Tokenization for Spatio-Temporal Heatmaps
- Enhancing Noise Robustness for Neural Speech Codecs through Resource-Efficient Progressive Quantization Perturbation Simulation
- DiSSECT: Structuring Transfer-Ready Medical Image Representations through Discrete Self-Supervision
- COLT: Enhancing Video Large Language Models with Continual Tool Usage
- Online Adaptation via Dual-Stage Alignment and Self-Supervision for Fast-Calibration Brain-Computer Interfaces
- Understanding-in-Generation: Reinforcing Generative Capability of Unified Model via Infusing Understanding into Generation
- Improving Test-Time Performance of RVQ-based Neural Codecs
- Learning Dexterous Manipulation with Quantized Hand State
- Drum-to-Vocal Percussion Sound Conversion and Its Evaluation Methodology
- VCE: Safe Autoregressive Image Generation via Visual Contrast Exploitation
- DA-Font: Few-Shot Font Generation via Dual-Attention Hybrid Integration
- MANZANO: A Simple and Scalable Unified Multimodal Model with a Hybrid Vision Tokenizer
- Inverse Optimization Latent Variable Models for Learning Costs Applied to Route Problems
- Mental Accounts for Actions: EWA-Inspired Attention in Decision Transformers
- Purely Semantic Indexing for LLM-based Generative Recommendation and Retrieval
- SAMPO:Scale-wise Autoregression with Motion PrOmpt for generative world models
- Attention Schema-based Attention Control (ASAC): A Cognitive-Inspired Approach for Attention Management in Transformers
- T2I-ReasonBench: Benchmarking Reasoning-Informed Text-to-Image Generation
- Llama-Mimi: Exploring the Limits of Flattened Speech Language Modeling
- Generative AI Meets Wireless Sensing: Towards Wireless Foundation Model
- Back to Ear: Perceptually Driven High Fidelity Music Reconstruction
- OpenViGA: Video Generation for Automotive Driving Scenes by Streamlining and Fine-Tuning Open Source Models with Public Data
- AToken: A Unified Tokenizer for Vision
- MoSA: Motion-Coherent Human Video Generation via Structure-Appearance Decoupling
- AnyAccomp: Generalizable Accompaniment Generation via Quantized Melodic Bottleneck
- ProTDyn: a foundation Protein language model for Thermodynamics and Dynamics generation
- Improving 3D Gaussian Splatting Compression by Scene-Adaptive Lattice Vector Quantization
- SPGen: Spherical Projection as Consistent and Flexible Representation for Single Image 3D Shape Generation
- VQT-Light:Lightweight HDR Illumination Map Prediction with Richer Texture.pdf
- Image Tokenizer Needs Post-Training
- Lost in Embeddings: Information Loss in Vision-Language Models
- VQL: An End-to-End Context-Aware Vector Quantization Attention for Ultra-Long User Behavior Modeling
- CoachMe: Decoding Sport Elements with a Reference-Based Coaching Instruction Generation Model
- PoolingVQ: A VQVAE Variant for Reducing Audio Redundancy and Boosting Multi-Modal Fusion in Music Emotion Analysis
- FuseCodec: Semantic-Contextual Fusion and Supervision for Neural Codecs
- OpenHA: A Series of Open-Source Hierarchical Agentic Models in Minecraft
- InfGen: A Resolution-Agnostic Paradigm for Scalable Image Synthesis
- A Discrepancy-Based Perspective on Dataset Condensation
- Scalable Training for Vector-Quantized Networks with 100% Codebook Utilization
- DiFlow-TTS: Discrete Flow Matching with Factorized Speech Tokens for Low-Latency Zero-Shot Text-To-Speech
- Finite Scalar Quantization Enables Redundant and Transmission-Robust Neural Audio Compression at Low Bit-rates
- A novel method and dataset for depth-guided image deblurring from smartphone Lidar
- DeCodec: Rethinking Audio Codecs as Universal Disentangled Representation Learners
- CoDiCodec: Unifying Continuous and Discrete Compressed Representations of Audio
- Integrating Anatomical Priors into a Causal Diffusion Model
- Learning Turbulent Flows with Generative Models: Super-resolution, Forecasting, and Sparse Flow Reconstruction
- Tokenizing Loops of Antibodies
- LatentVoiceGrad: Nonparallel Voice Conversion with Latent Diffusion/Flow-Matching Models
- Bitrate-Controlled Diffusion for Disentangling Motion and Content in Video
- TA-VLA: Elucidating the Design Space of Torque-aware Vision-Language-Action Models
- TaDiCodec: Text-aware Diffusion Speech Tokenizer for Speech Language Modeling
- Reconstruction Alignment Improves Unified Multimodal Models
- Learning Decomposed Contextual Token Representations from Pretrained and Collaborative Signals for Generative Recommendation
- UniSearch: Rethinking Search System with a Unified Generative Architecture
- Continuous Audio Language Models
- Compression Beyond Pixels: Semantic Compression with Multimodal Foundation Models
- LatinX: Aligning a Multilingual TTS Model with Direct Preference Optimization
- Exploring Autoregressive Vision Foundation Models for Image Compression
- Missing Fine Details in Images: Last Seen in High Frequencies
- Say More with Less: Variable-Frame-Rate Speech Tokenization via Adaptive Clustering and Implicit Duration Coding
- Phonological Representation Learning for Isolated Signs Improves Out-of-Vocabulary Generalization
- Human Motion Video Generation: A Survey
- Skywork UniPic 2.0: Building Kontext Model with Online RL for Unified Multimodal Model
- OneSearch: A Preliminary Exploration of the Unified End-to-End Generative Framework for E-commerce Search
- RecBase: Generative Foundation Model Pretraining for Zero-Shot Recommendation
- SynBT: High-quality Tumor Synthesis for Breast Tumor Segmentation by 3D Diffusion Model
- Heavy-Tailed Class-Conditional Priors for Long-Tailed Generative Modeling
- Spectrogram Patch Codec: A 2D Block-Quantized VQ-VAE and HiFi-GAN for Neural Speech Coding
- PromptFlare: Prompt-Generalized Defense via Cross-Attention Decoy in Diffusion-Based Inpainting
- Discrete Noise Inversion for Next-scale Autoregressive Text-based Image Editing
- Distillation of a tractable model from the VQ-VAE
- Hierarchical Motion Captioning Utilizing External Text Data Source
- GPSToken: Gaussian Parameterized Spatially-adaptive Tokenization for Image Representation and Generation
- Entropy-based Coarse and Compressed Semantic Speech Representation Learning
- Vevo2: A Unified and Controllable Framework for Speech and Singing Voice Generation
- Generative AI for Industrial Contour Detection: A Language-Guided Vision System
- Physics Informed Generative Models for Magnetic Field Images
- Embracing Aleatoric Uncertainty: Generating Diverse 3D Human Motion
- FORGE: Foundational Optimization Representations from Graph Embeddings
- Disentangling Latent Embeddings with Sparse Linear Concept Subspaces (SLiCS)
- Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies
- Quantum latent distributions in deep generative models
- Controllable Skin Synthesis via Lesion-Focused Vector Autoregression Model
- Fast 3D Diffusion for Scalable Granular Media Synthesis
- MRExtrap: Longitudinal Aging of Brain MRIs using Linear Modeling in Latent Space
- CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis
- Interpretable by AI Mother Tongue: Native Symbolic Reasoning in Neural Models
- PCR-CA: Parallel Codebook Representations with Contrastive Alignment for Multiple-Category App Recommendation
- Image-Conditioned 3D Gaussian Splat Quantization
- Visual Autoregressive Modeling for Instruction-Guided Image Editing
- Robust Residual Finite Scalar Quantization for Neural Compression
- Long-Context Speech Synthesis with Context-Aware Memory
- SATURN: Autoregressive Image Generation Guided by Scene Graphs
- Taming Transformer for Emotion-Controllable Talking Face Generation
- Making Pose Representations More Expressive and Disentangled via Residual Vector Quantization
- A Survey on Video Anomaly Detection via Deep Learning: Human, Vehicle, and Environment
- Democratizing News Recommenders: Modeling Multiple Perspectives for News Candidate Generation with VQ-VAE
- EmoSLLM: Parameter-Efficient Adaptation of LLMs for Speech Emotion Recognition
- Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey
- Learning local and global prototypes with optimal transport for unsupervised anomaly detection and localization
- Learn Faster and Remember More: Balancing Exploration and Exploitation for Continual Test-time Adaptation
- EgoTwin: Dreaming Body and View in First Person
- Scalable RF Simulation in Generative 4D Worlds
- VimoRAG: Video-based Retrieval-augmented 3D Motion Generation for Motion Language Models
- QuickMerge++: Fast Token Merging with Autoregressive Prior
- MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts
- Representation Quantization for Collaborative Filtering Augmentation
- OpenBEATs: A Fully Open-Source General-Purpose Audio Encoder
- Object Fidelity Diffusion for Remote Sensing Image Generation
- Ultra-High-Definition Reference-Based Landmark Image Super-Resolution with Generative Diffusion Prior
- Dissecting Generalized Category Discovery: Multiplex Consensus under Self-Deconstruction
- Exploiting Discriminative Codebook Prior for Autoregressive Image Generation
- Energy-Based Models for Predicting Mutational Effects on Proteins
- DAS: Dual-Aligned Semantic IDs Empowered Industrial Recommender System
- OneVAE: Joint Discrete and Continuous Optimization Helps Discrete Video VAE Train Better
- EEGDM: EEG Representation Learning via Generative Diffusion Model
- PaCo-FR: Patch-Pixel Aligned End-to-End Codebook Learning for Facial Representation Pre-training
- Edge General Intelligence Through World Models and Agentic AI: Fundamentals, Solutions, and Challenges
- Masquerade: Learning from In-the-wild Human Videos using Data-Editing
- How Safe Will I Be Given What I Saw? Calibrated Prediction of Safety Chances for Image-Controlled Autonomy
- Turbo-VAED: Fast and Stable Transfer of Video-VAEs to Mobile Devices
- Spatial-Temporal Multi-Scale Quantization for Flexible Motion Generation
- Closing the Performance Gap in Generative Recommenders with Collaborative Tokenization and Efficient Modeling
- VQ-VAE Based Digital Semantic Communication with Importance-Aware OFDM Transmission
- SPARC: Soft Probabilistic Adaptive multi-interest Retrieval Model via Codebooks for recommender system
- Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model
- Vision Generalist Model: A Survey
- Grouped Speculative Decoding for Autoregressive Image Generation
- LaVieID: Local Autoregressive Diffusion Transformers for Identity-Preserving Video Creation
- CoT-Pose: Chain-of-Thought Reasoning for 3D Pose Generation from Abstract Prompts
- MolmoAct: Action Reasoning Models that can Reason in Space
- DCoAR: Deep Concept Injection into Unified Autoregressive Models for Personalized Text-to-Image Generation
- AR-GRPO: Training Autoregressive Image Generation Models via Reinforcement Learning
- BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models
- Bounding Distributional Shifts in World Modeling through Novelty Detection
- Graph is a Natural Regularization: Revisiting Vector Quantization for Graph Representation Learning
- Bifrost-1: Bridging Multimodal LLMs and Diffusion Models with Patch-level CLIP Latents
- WeTok: Powerful Discrete Tokenization for High-Fidelity Visual Reconstruction
- Negative Binomial Variational Autoencoders for Overdispersed Latent Modeling
- Smoothing Slot Attention Iterations and Recurrences
- SONAR-LLM: Autoregressive Transformer that Thinks in Sentence Embeddings and Speaks in Tokens
- X-MoGen: Unified Motion Generation across Humans and Animals
- Automatic Image Colorization with Convolutional Neural Networks and Generative Adversarial Networks
- CRAM: Large-scale Video Continual Learning with Bootstrapped Compression
- Steering One-Step Diffusion Model with Fidelity-Rich Decoder for Fast Image Compression
- Live Music Models
- HiD-VAE: Interpretable Generative Recommendation via Hierarchical and Disentangled Semantic IDs
- UniTalker: Conversational Speech-Visual Synthesis
- Skeleton Motion Words for Unsupervised Skeleton-Based Temporal Action Segmentation
- Cross-Domain Image Synthesis: Generating H&E from Multiplex Biomarker Imaging
- MahaTTS: A Unified Framework for Multilingual Text-to-Speech Synthesis
- LiDARCrafter: Dynamic 4D World Modeling from LiDAR Sequences
- La La LiDAR: Large-Scale Layout Generation from LiDAR Data
- Veila: Panoramic LiDAR Generation from a Monocular RGB Image
- CloudBreaker: Breaking the Cloud Covers of Sentinel-2 Images using Multi-Stage Trained Conditional Flow Matching on Sentinel-1
- Zero-Variance Gradients for Variational Autoencoders
- GL-LCM: Global-Local Latent Consistency Models for Fast High-Resolution Bone Suppression in Chest X-Ray Images
- CIVQLLIE: Causal Intervention with Vector Quantization for Low-Light Image Enhancement
- SecoustiCodec: Cross-Modal Aligned Streaming Single-Codecbook Speech Codec
- FinWorld: An All-in-One Open-Source Platform for End-to-End Financial AI Research and Deployment
- Content-Aware Mamba for Learned Image Compression
- Remembering without (representational) memory: a neuro-computational study on regaining categoricity and compositionality from minimal traces
- A Spatio-temporal Continuous Network for Stochastic 3D Human Motion Prediction
- CLASS: Contrastive Learning via Action Sequence Supervision for Robot Manipulation
- Kronos: A Foundation Model for the Language of Financial Markets
- NS-Net: Decoupling CLIP Semantic Information through NULL-Space for Generalizable AI-Generated Image Detection
- Unraveling Hidden Representations: A Multi-Modal Layer Analysis for Better Synthetic Content Forensics
- Video Color Grading via Look-Up Table Generation
- Boosting Vision Semantic Density with Anatomy Normality Modeling for Medical Vision-language Pre-training
- ECGTwin: Personalized ECG Generation Using Controllable Diffusion Model
- VQ-DeepISC: Vector Quantized-Enabled Digital Semantic Communication with Channel Adaptive Image Transmission
- Dream, Lift, Animate: From Single Images to Animatable Gaussian Avatars
- Cooperative Perception: A Resource-Efficient Framework for Multi-Drone 3D Scene Reconstruction Using Federated Diffusion and NeRF
- XSpecMesh: Quality-Preserving Auto-Regressive Mesh Generation Acceleration via Multi-Head Speculative Decoding
- Adaptively Distilled ControlNet: Accelerated Training and Superior Sampling for Medical Image Synthesis
- UniLiP: Adapting CLIP for Unified Multimodal Understanding, Generation and Editing
- EB-gMCR: Energy-Based Generative Modeling for Signal Unmixing and Multivariate Curve Resolution
- Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space
- DiSC-Med: Diffusion-based Semantic Communications for Robust Medical Image Transmission
- A Practical Investigation of Spatially-Controlled Image Generation with Transformers
- EMP: Executable Motion Prior for Humanoid Robot Standing Upper-body Motion Imitation
- GestureHYDRA: Semantic Co-speech Gesture Synthesis via Hybrid Modality Diffusion Transformer and Cascaded-Synchronized Retrieval-Augmented Generation
- Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
- Generative Recommendation with Semantic IDs: A Practitioner's Handbook
- Discrete Tokenization for Multimodal LLMs: A Comprehensive Survey
- DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation
- Neural Autoregressive Modeling of Brain Aging
- Reconstructing 4D Spatial Intelligence: A Survey
- JAM: A Tiny Flow-based Song Generator with Fine-grained Controllability and Aesthetic Alignment
- Implicit Counterfactual Learning for Audio-Visual Segmentation
- Frequency-Aware Autoregressive Modeling for Efficient High-Resolution Image Synthesis
- Learning to Gridize: Segment Physical World by Wireless Communication Channel
- MagicAnime: A Hierarchically Annotated, Multimodal and Multitasking Dataset with Benchmarks for Cartoon Animation Generation
- TADT-CSA: Temporal Advantage Decision Transformer with Contrastive State Abstraction for Generative Recommendation
- KB-DMGen: Knowledge-Based Global Guidance and Dynamic Pose Masking for Human Image Generation
- A Survey on Generative Model Unlearning: Fundamentals, Taxonomy, Evaluation, and Future Direction
- FineMotion: A Dataset and Benchmark with both Spatial and Temporal Annotation for Fine-grained Motion Generation and Editing
- SCALAR: Scale-wise Controllable Visual Autoregressive Learning
- GABRIL: Gaze-Based Regularization for Mitigating Causal Confusion in Imitation Learning
- Back to the Features: DINO as a Foundation for Video World Models
- SemGes: Semantics-aware Co-Speech Gesture Generation using Semantic Coherence and Relevance Learning
- Reconstruct or Generate: Exploring the Spectrum of Generative Modeling for Cardiac MRI
- Orbis: Overcoming Challenges of Long-Horizon Prediction in Driving World Models
- A Survey of Multimodal Hallucination Evaluation and Detection
- TTS-VAR: A Test-Time Scaling Framework for Visual Auto-Regressive Generation
- Exploiting Gaussian Agnostic Representation Learning with Diffusion Priors for Enhanced Infrared Small Target Detection
- Enhancing Scene Transition Awareness in Video Generation via Post-Training
- Music-Aligned Holistic 3D Dance Generation via Hierarchical Motion Modeling
- Quantizing Text-attributed Graphs for Semantic-Structural Integration
- SimPersona: Learning Discrete Buyer Personas from Raw Clickstreams for Grounded E-Commerce Agents
- Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning
- Generative Video Compression with One-Dimensional Latent Representation
- Advances in Feed-Forward 3D Reconstruction and View Synthesis: A Survey
- Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion
- Learning Deformable Body Interactions With Adaptive Spatial Tokenization
- Generative AI-Driven High-Fidelity Human Motion Simulation
- Learning Deblurring Texture Prior from Unpaired Data with Diffusion Model
- From Points to Spheres: A Geometric Reinterpretation of Variational Autoencoders
- Toward Long-Tailed Online Anomaly Detection through Class-Agnostic Concepts
- TTMBA: Towards Text To Multiple Sources Binaural Audio Generation
- Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling
- Towards channel foundation models (CFMs): Motivations, methodologies and opportunities
- Generative Multi-Target Cross-Domain Recommendation
- Autoregressive Speech Enhancement via Acoustic Tokens
- Task-Specific Audio Coding for Machines: Machine-Learned Latent Features Are Codes for That Machine
- Resurrect Mask AutoRegressive Modeling for Efficient and Scalable Image Generation
- AU-Blendshape for Fine-grained Stylized 3D Facial Expression Manipulation
- A Survey of Deep Learning for Geometry Problem Solving
- Compositional Discrete Latent Code for High Fidelity, Productive Diffusion Models
- Locally Masked Convolution for Autoregressive Models
- crank: An Open-Source Software for Nonparallel Voice Conversion Based on Vector-Quantized Variational Autoencoder
- NWT: Towards natural audio-to-video generation with representation learning
- Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation
- Autoregressive Semantic Visual Reconstruction Helps VLMs Understand Better
- Eccentric Regularization: Minimizing Hyperspherical Energy without explicit projection
- Learning Sampling in Financial Statement Audits using Vector Quantised Autoencoder Neural Networks
- Online Learned Continual Compression with Adaptive Quantization Modules
- Learning source-aware representations of music in a discrete latent space
- Parrot: Data-Driven Behavioral Priors for Reinforcement Learning
- Manifolds for Unsupervised Visual Anomaly Detection
- Transformer VQ-VAE for Unsupervised Unit Discovery and Speech Synthesis: ZeroSpeech 2020 Challenge
- Diffusion-Based Imaginative Coordination for Bimanual Manipulation
- Computational principles of intelligence: learning and reasoning with neural networks
- EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing
- Alleviating Textual Reliance in Medical Language-guided Segmentation via Prototype-driven Semantic Approximation
- Teach Me Sign: Stepwise Prompting LLM for Sign Language Production
- Beyond Task-Specific Reasoning: A Unified Conditional Generative Framework for Abstract Visual Reasoning
- EEG Foundation Models: A Critical Review of Current Progress and Future Directions
- CodeBrain: Bridging Decoupled Tokenizer and Multi-Scale Architecture for EEG Foundation Model
- Quantize-then-Rectify: Efficient VQ-VAE Training
- Chat with AI: The Surprising Turn of Real-time Video Communication from Human to AI
- Latent Diffusion Models with Masked AutoEncoders
- IGD: Instructional Graphic Design with Multimodal Layer Generation
- MCGA: Mixture of Codebooks Hyperspectral Reconstruction via Grayscale-Aware Attention
- Voice Conversion Based Speaker Normalization for Acoustic Unit Discovery
- Implicit Rank-Minimizing Autoencoder
- A Preliminary Study of a Two-Stage Paradigm for Preserving Speaker Identity in Dysarthric Voice Conversion
- Inter2Former: Dynamic Hybrid Attention for Efficient High-Precision Interactive
- Intention-Conditioned Flow Occupancy Models
- AlphaVAE: Unified End-to-End RGBA Image Reconstruction and Generation with Alpha-Aware Representation Learning
- Improving Lossless Compression Rates via Monte Carlo Bits-Back Coding
- I2-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting
- SnapMoGen: Human Motion Generation from Expressive Texts
- Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Image Generation
- SemAlignVC: Enhancing zero-shot timbre conversion using semantic alignment
- MGVQ: Could VQ-VAE Beat VAE? A Generalizable Tokenizer with Multi-group Quantization
- D-CNN and VQ-VAE Autoencoders for Compression and Denoising of Industrial X-ray Computed Tomography Images
- Behave Your Motion: Habit-preserved Cross-category Animal Motion Transfer
- RIGEL: Real-time Optical Anomaly Diagnosis with Stateful In-Network Inference based on Distributed On-switch GNNs
- A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality
- Vision-Language-Vision Auto-Encoder: Scalable Knowledge Distillation from Diffusion Models
- Generating Multi-Table Time Series EHR from Latent Space with Minimal Preprocessing
- STARS: A Unified Framework for Singing Transcription, Alignment, and Refined Style Annotation
- FIFA: Unified Faithfulness Evaluation Framework for Text-to-Video and Video-to-Text Generation
- PaletteID: Prototype-Composed Semantic Identifiers for Multimodal CTR Prediction
- FairDiffuseVQVAE: Sampling-Time Fairness in Tabular Diffusion via Conditional Refinement of Vector-Quantized Latents
- Goal-Oriented Skill Abstraction for Offline Multi-Task Reinforcement Learning
- LangMamba: A Language-driven Mamba Framework for Low-dose CT Denoising with Vision-language Models
- Bridging Sequential Deep Operator Network and Video Diffusion: Residual Refinement of Spatio-Temporal PDE Solutions
- Omni-Video: Democratizing Unified Video Understanding and Generation
- Text-Guided Token Communication for Wireless Image Transmission
- VisualSpeaker: Visually-Guided 3D Avatar Lip Synthesis
- Trainable Class Prototypes for Few-Shot Learning
- Taming Data Challenges in ML-based Security Tasks Using Generative AI
- Latent Actions from Factorized Transition Effects under Agent Ambiguity
- MoRAE: Flow-Friendly Self-Supervised Latents for Text-to-Motion Generation
- Motion Generation: A Survey of Generative Approaches and Benchmarks
- Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens
- DC-AR: Efficient Masked Autoregressive Image Generation with Deep Compression Hybrid Tokenizer
- PointGAC: Geometric-Aware Codebook for Masked Point Cloud Modeling
- LLMs as Architects and Critics for Multi-Source Opinion Summarization
- Do Music Foundation Models Embed Pitch in Helical Structure?
- FastS2S-VC: Streaming Non-Autoregressive Sequence-to-Sequence Voice Conversion
- ICAS: Detecting Training Data from Autoregressive Image Generative Models
- Critique of World Model
- Testing for Typicality with Respect to an Ensemble of Learned Distributions
- Learning Disentangled Phone and Speaker Representations in a Semi-Supervised VQ-VAE Paradigm
- GALA: Generative Aligned Learning for Adaptive Multimodal Representation in the Taobao Shangou Recommender System
- MambaVideo for Discrete Video Tokenization with Channel-Split Quantization
- A Training-Free Style-Personalization via SVD-Based Feature Decomposition
- Foveation for Segmentation of Ultra-High Resolution Images
- Tractable Representation Learning with Probabilistic Circuits
- Accurate and Efficient World Modeling with Masked Latent Transformers
- Pronunciation-Lexicon Free Training for Phoneme-based Crosslingual ASR via Joint Stochastic Approximation
- SMCLM: Semantically Meaningful Causal Language Modeling for Autoregressive Paraphrase Generation
- Reconstructing hadronically decaying tau leptons with a jet foundation model
- Global Variational Inference Enhanced Robust Domain Adaptation
- LangScene-X: Reconstruct Generalizable 3D Language-Embedded Scenes with TriMap Video Diffusion
- Mesh Silksong: Auto-Regressive Mesh Generation as Weaving Silk
- Hita: Holistic Tokenizer for Autoregressive Image Generation
- LATTE: Latent Trajectory Embedding for Diffusion-Generated Image Detection
- Energy-Based Transformers are Scalable Learners and Thinkers
- Communication-Computation Trade-Off in Resource-Constrained Edge Inference
- Enhancing Multi-Exposure High Dynamic Range Imaging with Overlapped Codebook for Improved Representation Learning
- Exploring Classical Piano Performance Generation with Expressive Music Variational AutoEncoder
- Differentiable Product Quantization for End-to-End Embedding Compression
- Frequency Domain-Based Diffusion Model for Unpaired Image Dehazing
- Autoregressive Image Generation with Linear Complexity: A Spatial-Aware Decay Perspective
- Is Visual in-Context Learning for Compositional Medical Tasks within Reach?
- BEV-VAE: Multi-view Image Generation with Spatial Consistency for Autonomous Driving
- Effectiveness of self-supervised pre-training for speech recognition
- ARIG: Autoregressive Interactive Head Generation for Real-time Conversations
- Flexible Language Modeling in Continuous Space with Transformer-based Autoregressive Flows
- Neural Spectral Band Generation for Audio Coding
- A Survey: Learning Embodied Intelligence from Physical Simulators and World Models
- MAMBO: High-Resolution Generative Approach for Mammography Images
- Epona: Autoregressive Diffusion World Model for Autonomous Driving
- MotionGPT3: Human Motion as a Second Modality
- Unified Multimodal Understanding via Byte-Pair Visual Encoding
- CycleVAR: Repurposing Autoregressive Model for Unsupervised One-Step Image Translation
- BridgeShape: Latent Diffusion Schrödinger Bridge for 3D Shape Completion
- Curious Causality-Seeking Agents Learn Meta Causal World
- Learning Causal State Representations of Partially Observable Environments
- Squeeze3D: Your 3D Generation Model is Secretly an Extreme Neural Compressor
- Concept Pinpoint Eraser for Text-to-image Diffusion Models via Residual Attention Gate
- EgoM2P: Egocentric Multimodal Multitask Pretraining
- Unsupervised Audiovisual Synthesis via Exemplar Autoencoders
- D2C: Diffusion-Denoising Models for Few-shot Conditional Generation
- Hierarchical Characterization of Brain Dynamics via State Space-based Vector Quantization
- 3D Shape Generation: A Survey
- Irregular Convolutional Auto-Encoder on Point Clouds
- Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy
- StableCodec: Taming One-Step Diffusion for Extreme Image Compression
- CodecSlime: Temporal Redundancy Compression of Neural Speech Codec via Dynamic Frame Rate
- Discovering Dialog Structure Graph for Open-Domain Dialog Generation
- Capturing User Interests from Data Streams for Continual Sequential Recommendation
- Instella-T2I: Pushing the Limits of 1D Discrete Latent Space Image Generation
- Exploring Image Generation via Mutually Exclusive Probability Spaces and Local Correlation Hypothesis
- Bridging Compositional and Distributional Semantics: A Survey on Latent Semantic Geometry via AutoEncoder
- EAR: Erasing Concepts from Unified Autoregressive Models
- OLALa: Online Learned Adaptive Lattice Codes for Heterogeneous Federated Learning
- UniCode2: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation
- Diffuse Everything: Multimodal Diffusion Models on Arbitrary State Spaces
- SoK: Can Synthetic Images Replace Real Data? A Survey of Utility and Privacy of Synthetic Image Generation
- Cross-Layer Discrete Concept Discovery for Interpreting Language Models
- Kling-Foley: Multimodal Diffusion Transformer for High-Quality Video-to-Audio Generation
- Style Transfer: A Decade Survey
- LeVo: High-Quality Song Generation with Multi-Preference Alignment
- Discrete Representations Strengthen Vision Transformer Robustness
- Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion
- From Virtual Games to Real-World Play
- Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations
- CommVQ: Commutative Vector Quantization for KV Cache Compression
- DuetGen: Music Driven Two-Person Dance Generation via Hierarchical Masked Modeling
- Normality Prior Guided Multi-Semantic Fusion Network for Unsupervised Image Anomaly Detection
- Transformer World Model for Sample Efficient Multi-Agent Reinforcement Learning
- Auto-Regressively Generating Multi-View Consistent Images
- Exponential Tilting of Generative Models: Improving Sample Quality by Training and Sampling from Latent Energy
- Highly Compressed Tokenizer Can Generate Without Training
- VIVAT: Virtuous Improving VAE Training through Artifact Mitigation
- Make It Efficient: Dynamic Sparse Attention for Autoregressive Image Generation
- IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech
- MiCo: Multiple Instance Learning with Context-Aware Clustering for Whole Slide Image Analysis
- Variational Variance: Simple, Reliable, Calibrated Heteroscedastic Noise Variance Parameterization
- PlanMoGPT: Flow-Enhanced Progressive Planning for Text to Motion Synthesis
- SIGNER: Temporally Grounded Sign Language Generation via Time-Resolved Conditioning
- Learning golf swing signatures from a single wrist-worn inertial sensor
- Deep generative models as the probability transformation functions
- Visual-Instructed Degradation Diffusion for All-in-One Image Restoration
- Single-step Diffusion for Image Compression at Ultra-Low Bitrates
- A Tutorial on VAEs: From Bayes' Rule to Lossless Compression
- Watermarking Autoregressive Image Generation
- CRIA: A Cross-View Interaction and Instance-Adapted Pre-training Framework for Generalizable EEG Representations
- Towards Bitrate-Efficient and Noise-Robust Speech Coding with Variable Bitrate RVQ
- Advanced Sign Language Video Generation with Compressed and Quantized Multi-Condition Tokenization
- Aligning Text, Images, and 3D Structure Token-by-Token
- Privacy-Preserving Chest X-ray Classification in Latent Space with Homomorphically Encrypted Neural Inference
- Enhancing Vector Quantization with Distributional Matching: A Theoretical and Empirical Study
- HOIDiNi: Human-Object Interaction through Diffusion Noise Optimization
- Exploring Disentanglement with Multilingual and Monolingual VQ-VAE
- Construction of an Organ Shape Atlas Using a Hierarchical Mesh Variational Autoencoder
- SyncTalk++: High-Fidelity and Efficient Synchronized Talking Heads Synthesis Using Gaussian Splatting
- VQ-DRAW: A Sequential Discrete VAE
- Discrete JEPA: Learning Discrete Token Representations without Reconstruction
- Causally Steered Diffusion for Automated Video Counterfactual Generation
- AMPLIFY: Actionless Motion Priors for Robot Learning from Videos
- Leveraging External Factors in Household-Level Electrical Consumption Forecasting using Hypernetworks
- Learned transform compression with optimized entropy encoding
- Latent Anomaly Detection: Masked VQ-GAN for Unsupervised Segmentation in Medical CBCT
- SeqPE: Transformer with Sequential Position Encoding
- COME: Adding Scene-Centric Forecasting Control to Occupancy World Model
- Discrete Diffusion in Large Language and Multimodal Models: A Survey
- DicFace: Dirichlet-Constrained Variational Codebook Learning for Temporally Coherent Video Face Restoration
- Interpretable and Reliable Detection of AI-Generated Images via Grounded Reasoning in MLLMs
- CrossGen: Learning and Generating Cross Fields for Quad Meshing
- Hierarchical Group-wise Ranking Framework for Recommendation Models
- Enhancing Large Language Models for Mobility Analytics with Semantic Location Tokenization
- LaTtE-Flow: Layerwise Timestep-Expert Flow-based Transformer
- GGBall: Graph Generative Model on Poincaré Ball
- BG-HOP: A Bimanual Generative Hand-Object Prior
- RL from Physical Feedback: Aligning Large Motion Models with Humanoid Control
- Good Noise Makes Good Edits: A Training-Free Diffusion-Based Video Editing with Image and Text Prompts
- Recent Advances in Autoencoder-Based Representation Learning
- Exploring the Effectiveness of Deep Features from Domain-Specific Foundation Models in Retinal Image Synthesis
- Efficient Text-Attributed Graph Learning through Selective Annotation and Graph Alignment
- Efficient Multi-Camera Tokenization with Triplanes for End-to-End Driving
- Task-Driven Discrete Representation Learning
- Subgoal-Guided Policy Heuristic Search with Learned Subgoals
- SpectralAR: Spectral Autoregressive Visual Generation
- Unsupervised Deformable Image Registration with Structural Nonparametric Smoothing
- DanceChat: Large Language Model-Guided Music-to-Dance Generation
- A look at adversarial attacks on radio waveforms from discrete latent space
- Prompt-Guided Latent Diffusion with Predictive Class Conditioning for 3D Prostate MRI Generation
- Discond-VAE: Disentangling Continuous Factors from the Discrete
- DGAE: Diffusion-Guided Autoencoder for Efficient Latent Representation Learning
- Towards AI-Native Fronthaul: Neural Compression for NextG Cloud RAN
- HeartcareGPT: A Unified Multimodal ECG Suite for Dual Signal-Image Modeling and Understanding
- Projectable Models: One-Shot Generation of Small Specialized Transformers from Large Ones
- Bridging the Modality Gap: Softly Discretizing Audio Representation for LLM-based Automatic Speech Recognition
- RecGPT: A Foundation Model for Sequential Recommendation
- CLAM: Continuous Latent Action Models for Robot Learning from Unlabeled Demonstrations
- A Factorial Mixture Prior for Compositional Deep Generative Models
- (LiFT) Lightweight Fitness Transformer: A language-vision model for Remote Monitoring of Physical Training
- AE-OT-GAN: Training GANs from data specific latent distribution
- FontAdapter: Instant Font Adaptation in Visual Text Generation
- MOGO: Residual Quantized Hierarchical Causal Transformer for High-Quality and Real-Time 3D Human Motion Generation
- Aligning Latent Spaces with Flow Priors
- Multi-scale Image Super Resolution with a Single Auto-Regressive Model
- Towards Sequence Modeling Alignment between Tokenizer and Autoregressive Model
- TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation
- X-Driver: Explainable Autonomous Driving with Vision-Language Models
- UniMate: A Unified Model for Mechanical Metamaterial Generation, Property Prediction, and Condition Confirmation
- Improving AI-generated music with user-guided training
- Bringing Interpretability to Neural Audio Codecs
- Beyond Token-Level Cross-Entropy: Fréchet Distributional Post-Training for Autoregressive Image Generation
- Distillation-Enabled Knowledge Alignment Protocol for Semantic Communication in AI Agent Networks
- Variable-rate discrete representation learning
- A Comparison of Discrete Latent Variable Models for Speech Representation Learning
- Autoencoding Under Normalization Constraints
- AI-Generated Fall Data: Assessing LLMs and Diffusion Model for Wearable Fall Detection
- ALFEE: Adaptive Large Foundation Model for EEG Representation
- Unsupervised Acoustic Unit Representation Learning for Voice Conversion using WaveNet Auto-encoders
- Occupancy World Model for Robots
- A Universal Music Translation Network
- Semi-supervised Grasp Detection by Representation Learning in a Vector Quantized Latent Space
- Learning Physical Concepts in Cyber-Physical Systems: A Case Study
- Unsupervised Paraphrasing by Simulated Annealing
- GRILL: Gradient Signal Restoration in Ill-Conditioned Layers to Enhance Adversarial Attacks on Autoencoders
- From Pixels to Polygons: A Survey of Deep Learning Approaches for Medical Image-to-Mesh Reconstruction
- Fixed-Length Dense Fingerprint Representation with Alignment and Robust Enhancement
- STAR: Learning Diverse Robot Skill Abstractions through Rotation-Augmented Vector Quantization
- Video-Text Pre-training with Learned Regions
- Illiterate DALL-E Learns to Compose
- HMAR: Efficient Hierarchical Masked Auto-Regressive Image Generation
- Recent Advancements in Self-Supervised Paradigms for Visual Feature Representation
- Assessing the Completeness of Traffic Scenario Categories for Automated Highway Driving Functions via Cluster-based Analysis
- Efficient Tactile Perception with Soft Electrical Impedance Tomography and Pre-trained Transformer
- FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens
- InterMamba: Efficient Human-Human Interaction Generation with Adaptive Spatio-Temporal Mamba
- A platform for multimodal in vivo pooled genetic screens reveals regulators of liver function
- FreqPolicy: Frequency Autoregressive Visuomotor Policy with Continuous Tokens
- Self-supervised Latent Space Optimization with Nebula Variational Coding
- Generative Next POI Recommendation with Semantic ID
- Ultra-High-Resolution Image Synthesis: Data, Method and Evaluation
- Tomographic Foundation Model -- FORCE: Flow-Oriented Reconstruction Conditioning Engine
- Benchmarking Neural Speech Codec Intelligibility with SITool
- On-device Streaming Discrete Speech Units
- ShapeLLM-Omni: A Native Multimodal LLM for 3D Generation and Understanding
- Unraveling Spatio-Temporal Foundation Models via the Pipeline Lens: A Comprehensive Review
- Automatic Stage Lighting Control: Is it a Rule-Driven Process or Generative Task?
- Enhancing Interpretable Image Classification Through LLM Agents and Conditional Concept Bottleneck Models
- Speaking Beyond Language: A Large-Scale Multimodal Dataset for Learning Nonverbal Cues from Video-Grounded Dialogues
- Efficient And Scalable Neural Residual Waveform Coding With Collaborative Quantization
- VQVAE Unsupervised Unit Discovery and Multi-scale Code2Spec Inverter for Zerospeech Challenge 2019
- Counterfactual Activation Editing for Post-hoc Prosody and Mispronunciation Correction in TTS Models
- Uncertainty-Aware Prototype Semantic Decoupling for Text-Based Person Search in Full Images
- Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching
- Autoregressive Images Watermarking through Lexical Biasing: An Approach Resistant to Regeneration Attack
- Synthetic weather radar using hybrid quantum-classical machine learning
- Humanoid World Models: Open World Foundation Models for Humanoid Robotics
- MagiCodec: Simple Masked Gaussian-Injected Codec for High-Fidelity Reconstruction and Generation
- Probabilistic Forecasting for Building Energy Systems using Time-Series Foundation Models
- Interpretable agent communication from scratch (with a generic visual processor emerging on the side)
- Concept-Centric Token Interpretation for Vector-Quantized Generative Models
- DLM-One: Diffusion Language Models for One-Step Sequence Generation
- On Designing Diffusion Autoencoders for Efficient Generation and Representation Learning
- DiffWave: A Versatile Diffusion Model for Audio Synthesis
- Are Any-to-Any Models More Consistent Across Modality Transfers Than Specialists?
- A Multi-Level Causal Intervention Framework for Mechanistic Interpretability in Variational Autoencoders
- SwitchCodec: A High-Fidelity Nerual Audio Codec With Sparse Quantization
- DS-Codec: Dual-Stage Training with Mirror-to-NonMirror Architecture Switching for Speech Codec
- SignBot: Learning Human-to-Humanoid Sign Language Interaction
- MAGREF: Masked Guidance for Any-Reference Video Generation with Subject Disentanglement
- D-AR: Diffusion via Autoregressive Models
- Semantics-Aware Human Motion Generation from Audio Instructions
- LAFR: Efficient Diffusion-based Blind Face Restoration via Latent Codebook Alignment Adapter
- Variational Inference for Data-Efficient Model Learning in POMDPs
- Fine-Tuning Next-Scale Visual Autoregressive Models with Group Relative Policy Optimization
- MMGT: Motion Mask Guided Two-Stage Network for Co-Speech Gesture Video Generation
- EAD: An EEG Adapter for Automated Classification
- Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation
- A Short Note on Analyzing Sequence Complexity in Trajectory Prediction Benchmarks
- Compositional Fine-Grained Low-Shot Learning
- Normalizing Flows are Capable Models for Continuous Control
- MGE-LDM: Joint Latent Diffusion for Simultaneous Music Generation and Source Extraction
- Lane-Wise Highway Anomaly Detection
- Deep Learning-Based CSI Feedback for Wi-Fi Systems With Temporal Correlation
- ACE-Step: A Step Towards Music Generation Foundation Model
- Latent Reasoning via Sentence Embedding Prediction
- Towards Scalable Language-Image Pre-training for 3D Medical Imaging
- OSPO: Object-Centric Self-Improving Preference Optimization for Text-to-Image Generation
- PacTure: Efficient PBR Texture Generation on Packed Views with Visual Autoregressive Models
- AudioTurbo: Fast Text-to-Audio Generation with Rectified Diffusion
- EF-VI: Enhancing End-Frame Injection for Video Inbetweening
- Improving Brain-to-Image Reconstruction via Fine-Grained Text Bridging
- CVC: Contrastive Learning for Non-parallel Voice Conversion
- Revisiting Multi-Agent World Modeling from a Diffusion-Inspired Perspective
- Incorporating Real-world Noisy Speech in Neural-network-based Speech Enhancement Systems
- Exploring Fluent Query Reformulations with Text-to-Text Transformers and Reinforcement Learning
- MetaSlot: Break Through the Fixed Number of Slots in Object-Centric Learning
- OmniResponse: Online Multimodal Conversational Response Generation in Dyadic Interactions
- What Do Latent Action Models Actually Learn?
- DetailFlow: 1D Coarse-to-Fine Autoregressive Image Generation via Next-Detail Prediction
- MoPFormer: Motion-Primitive Transformer for Wearable-Sensor Activity Recognition
- LeDiFlow: Learned Distribution-guided Flow Matching to Accelerate Image Generation
- Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech
- Discrete Markov Bridge
- Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression
- LlamaSeg: Image Segmentation via Autoregressive Mask Generation
- StyleAR: Customizing Multimodal Autoregressive Model for Style-Aligned Text-to-Image Generation
- Deep Sequence Learning for Video Anticipation: From Discrete and Deterministic to Continuous and Stochastic
- Information-theoretic Generalization Analysis for VQ-VAEs: A Role of Latent Variables
- Unified Multimodal Understanding and Generation Models: Advances, Challenges, and Opportunities
- BrainStratify: Coarse-to-Fine Disentanglement of Intracranial Neural Dynamics
- Learning the Imaging Landmarks: Unsupervised Key point Detection in Lung Ultrasound Videos
- Advancing Limited-Angle CT Reconstruction Through Diffusion-Based Sinogram Completion
- FlowSE: Efficient and High-Quality Speech Enhancement via Flow Matching
- Knowledge-Aligned Counterfactual-Enhancement Diffusion Perception for Unsupervised Cross-Domain Visual Emotion Recognition
- Direct Noisy Speech Modeling for Noisy-to-Noisy Voice Conversion
- AniCrafter: Customizing Realistic Human-Centric Animation via Avatar-Background Conditioning in Video Diffusion Models
- DiEmo-TTS: Disentangled Emotion Representations via Self-Supervised Distillation for Cross-Speaker Emotion Transfer in Text-to-Speech
- LAFITE: Towards Language-Free Training for Text-to-Image Generation
- WorldEval: World Model as Real-World Robot Policies Evaluator
- Plug-and-Play Context Feature Reuse for Efficient Masked Generation
- Unpaired Modality-Agnostic Generative Recommendation
- PIGPVAE: Physics-Informed Gaussian Process Variational Autoencoders
- ASTRA: Communication-Efficient Acceleration for Multi-Device Transformer Inference
- Joint-stochastic-approximation Autoencoders with Application to Semi-supervised Learning
- BiomechGPT: Extending Motion-Language Models to Clinical Motion Understanding
- Distinctive Feature Codec: An Adaptive Efficient Speech Representation for Depression Detection
- v1: Learning to Point Visual Tokens for Multimodal Grounded Reasoning
- Image Tokens Matter: Mitigating Hallucination in Discrete Tokenizer-based Large Vision-Language Models via Latent Editing
- Bridging Supervised Learning and Reinforcement Learning in Math Reasoning
- RestoreVAR: Visual Autoregressive Generation for All-in-One Image Restoration
- Imagine Beyond! Distributionally Robust Auto-Encoding for State Space Coverage in Online Reinforcement Learning
- FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous Driving
- Discrete-Valued Neural Communication
- Sparse Diffusion Autoencoder for Test-time Adapting Prediction of Complex Systems
- UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information
- Slot-MLLM: Object-Centric Visual Tokenization for Multimodal LLM
- Maximizing Mutual Information for Tacotron
- MEGADance: Mixture-of-Experts Architecture for Genre-Aware 3D Dance Generation
- TokBench: Evaluating Your Visual Tokenizer before Visual Generation
- High-Fidelity Functional Ultrasound Reconstruction via A Visual Auto-Regressive Framework
- Learning Adaptive and Temporally Causal Video Tokenization in a 1D Latent Space
- Mesh-RFT: Enhancing Mesh Generation via Fine-grained Reinforcement Fine-Tuning
- Generative Latent Coding for Ultra-Low Bitrate Image and Video Compression
- CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning
- MARché: Fast Masked Autoregressive Image Generation with Cache-Aware Attention
- CLIPDraw: Exploring Text-to-Drawing Synthesis through Language-Image Encoders
- From Prediction to Perfection: Introducing Refinement to Autoregressive Image Generation
- Learning to Tessellate: Point Cloud Generation via Recursive Spectral Partitioning
- ALTo: Adaptive-Length Tokenizer for Autoregressive Mask Generation
- FPQVAR: Floating Point Quantization for Visual Autoregressive Model with FPGA Hardware Co-design
- Differentiable K-means for Fully-optimized Discrete Token-based ASR
- ChemMLLM: Chemical Multimodal Large Language Model
- OBoW: Online Bag-of-Visual-Words Generation for Self-Supervised Learning
- AvatarShield: Visual Reinforcement Learning for Human-Centric Synthetic Video Detection
- Learning Interpretable Representations Leads to Semantically Faithful EEG-to-Text Generation
- EASY: Emotion-aware Speaker Anonymization via Factorized Distillation
- Nonparallel Voice Conversion with Augmented Classifier Star Generative Adversarial Networks
- Intentional Gesture: Deliver Your Intentions with Gestures for Speech
- Deep Encoder-Decoder Models for Unsupervised Learning of Controllable Speech Synthesis
- The Multi-speaker Multi-style Voice Cloning Challenge 2021
- Segmentation-Variant Codebooks for Preservation of Paralinguistic and Prosodic Information
- Discrete Audio Representations for Automated Audio Captioning
- Interspatial Attention for Efficient 4D Human Video Generation
- StreamTalk: Streaming Co-Speech Gesture Generation with Key-Pose Anchoring
- Large Language Models Implicitly Learn to See and Hear Just By Reading
- Rethinking Video Token Compression with a Global Codebook: Learning Once, Compressing Everywhere
- Byte Pair Encoding for Efficient Time Series Forecasting
- Plane Geometry Problem Solving with Multi-modal Reasoning: A Survey
- MSDformer: Multi-scale Discrete Transformer For Time Series Generation
- Impact of Frame Rates on Speech Tokenizer: A Case Study on Mandarin and English
- UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation
- RLVR-World: Training World Models with Reinforcement Learning
- A Semantic Information-based Hierarchical Speech Enhancement Method Using Factorized Codec and Diffusion Model
- Sat2Sound: A Unified Framework for Zero-Shot Soundscape Mapping
- Latent Programmer: Discrete Latent Codes for Program Synthesis
- VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation
- VesselGPT: Autoregressive Modeling of Vascular Geometry
- Denoising Diffusion Probabilistic Model for Point Cloud Compression at Low Bit-Rates
- Efficient Speech Language Modeling via Energy Distance in Continuous Latent Space
- Universal Semantic Disentangled Privacy-preserving Speech Representation Learning
- DualCodec: A Low-Frame-Rate, Semantically-Enhanced Neural Audio Codec for Speech Generation
- OZSpeech: One-step Zero-shot Speech Synthesis with Learned-Prior-Conditioned Flow Matching
- GANCompress: GAN-Enhanced Neural Image Compression with Binary Spherical Quantization
- Understanding Complexity in VideoQA via Visual Program Generation
- Touch2Shape: Touch-Conditioned 3D Diffusion for Shape Exploration and Reconstruction
- Speech-to-speech Translation between Untranscribed Unknown Languages
- MVAR: Visual Autoregressive Modeling with Scale and Spatial Markovian Conditioning
- ChromFound: Towards A Universal Foundation Model for Single-Cell Chromatin Accessibility Data
- FreqSelect: Frequency-Aware fMRI-to-Image Reconstruction
- Training Latent Diffusion Models with Interacting Particle Algorithms
- Hyperbolic Residual Quantization: Discrete Representations for Data with Latent Hierarchies
- Context-Aware Autoregressive Models for Multi-Conditional Image Generation
- Learning High-Order Relationships with Hypergraph Attention-based Spatio-Temporal Aggregation for Brain Disease Analysis
- Evidential Softmax for Sparse Multimodal Distributions in Deep Generative Models
- Patient-Specific Autoregressive Models for Organ Motion Prediction in Radiotherapy
- Learning Speaker Embedding from Text-to-Speech
- Unpaired Image-to-Image Translation via Latent Energy Transport
- MARRS: Masked Autoregressive Unit-based Reaction Synthesis
- Channel-Agnostic Semantic Compression for Bandwidth-Limited Visual Communication
- Prot2Text-V2: Protein Function Prediction with Multimodal Contrastive Alignment
- Self-supervised perception for tactile skin covered dexterous hands
- EA-3DGS: Efficient and Adaptive 3D Gaussians with Highly Enhanced Quality for outdoor scenes
- ToDMA: Large Model-Driven Massive Token Communications for Semantic Multiple Access
- Observational causality by states and interaction type for scientific discovery
- Towards Robust and Controllable Text-to-Motion via Masked Autoregressive Diffusion
- TACO: Rethinking Semantic Communications with Task Adaptation and Context Embedding
- Dyadic Mamba: Long-term Dyadic Human Motion Synthesis
- UICopilot: Automating UI Synthesis via Hierarchical Code Generation from Webpage Designs
- An Introduction to Discrete Variational Autoencoders
- Multi-Token Prediction Needs Registers
- FlowDreamer: A RGB-D World Model with Flow-based Motion Representations for Robot Manipulation
- Robust Representation Learning via Perceptual Similarity Metrics
- Text-driven Motion Generation: Overview, Challenges and Directions
- Reinforcement Learning meets Masked Video Modeling : Trajectory-Guided Adaptive Token Selection
- TT-DF: A Large-Scale Diffusion-Based Dataset and Benchmark for Human Body Forgery Detection
- M3G: Multi-Granular Gesture Generator for Audio-Driven Full-Body Human Motion Synthesis
- EventDiff: A Unified and Efficient Diffusion Model Framework for Event-based Video Frame Interpolation
- AI and Generative AI Transforming Disaster Management: A Survey of Damage Assessment and Response Techniques
- Using Seismic Statistical Features and VQ-VAE to Improve Spatiotemporal Seismicity Predictability
- Towards Foundation Models for Experimental Readout Systems Combining Discrete and Continuous Data
- Feature Quantization Improves GAN Training
- Many-to-Many Voice Conversion using Cycle-Consistent Variational Autoencoder with Multiple Decoders
- MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder
- Metrics that matter: Evaluating image quality metrics for medical image generation
- Continuous Visual Autoregressive Generation via Score Maximization
- H3DP: Triply-Hierarchical Diffusion Policy for Visuomotor Learning
- Entropy optimized semi-supervised decomposed vector-quantized variational autoencoder model based on transfer learning for multiclass text classification and generation
- Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning
- Multi-band Frequency Reconstruction for Neural Psychoacoustic Coding
- Joint Low-level and High-level Textual Representation Learning with Multiple Masking Strategies
- A Temporal Variational Model for Story Generation
- ProFashion: Prototype-guided Fashion Video Generation with Multiple Reference Images
- A Short Overview of Multi-Modal Wi-Fi Sensing
- Video-Enhanced Offline Reinforcement Learning: A Model-Based Approach
- Accurate and Efficient Multivariate Time Series Forecasting via Offline Clustering
- Noise-Consistent Siamese-Diffusion for Medical Image Synthesis and Segmentation
- Asymmetric Generative Recommendation via Kronecker Residual Bridge and Multi-Faceted Hierarchical Quantization
- Prompt Codebooks: Discrete Compositional Optimization for Language Model Instruction Refinement
- tvGP-VAE: Tensor-variate Gaussian Process Prior Variational Autoencoder
- UniVLA: Learning to Act Anywhere with Task-centric Latent Actions
- The Moon's Many Faces: A Single Unified Transformer for Multimodal Lunar Reconstruction
- Latent Preference Coding: Aligning Large Language Models via Discrete Latent Codes
- ReactDance: Hierarchical Representation for High-Fidelity and Coherent Long-Form Reactive Dance Generation
- Light Forcing: Accelerating Autoregressive Video Diffusion via Sparse Attention
- A Survey on Generative Diffusion Models
- FlowDubber: Movie Dubbing with LLM-based Semantic-aware Learning and Flow Matching based Voice Enhancing
- FreePCA: Integrating Consistency Information across Long-short Frames in Training-free Long Video Generation via Principal Component Analysis
- Fast Decoding in Sequence Models using Discrete Latent Variables
- VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling
- BitDance: Scaling Autoregressive Generative Models with Binary Tokens
- Voice Cloning: Comprehensive Survey
- Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing
- Adaptive Protein Tokenization
- Tailor Made Embeddings for Quantum Machine Learning
- SLAM: A Unified Encoder for Speech and Language Modeling via Speech-Text Joint Pre-Training
- Hierarchical Latent Action Model
- Unsupervised Disentanglement of Linear-Encoded Facial Semantics
- SOM-VQ: Topology-Aware Tokenization for Interactive Generative Models
- M6: A Chinese Multimodal Pretrainer
- VLANeXt: Recipes for Building Strong VLA Models
- Cryo-SWAN: the Multi-Scale Wavelet-decomposition-inspired Autoencoder Network for molecular density representation of molecular volumes
- End-to-End Autoregressive Image Generation with 1D Semantic Tokenizer
- Neuro-Symbolic ODE Discovery with Latent Grammar Flow
- The Design Space of Tri-Modal Masked Diffusion Models
- Representation Forcing for Bottleneck-Free Unified Multimodal Models
- Semantic Motion Anchors: Bridging Motion and Meaning in Co-Speech Gestures
- GPIC: A Giant Permissive Image Corpus for Visual Generation
- Atom-level Protein Representation Learning Improves Protein Structure Prediction
- Learning Latent Action World Models In The Wild
- VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction
- SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture
- Yeti: A compact protein structure tokenizer for reconstruction and multi-modal generation
- Large Language Models are Universal Reasoners for Visual Generation
- CMT: A Cascade MAR with Topology Predictor for Multimodal Conditional CAD Generation
- Efficient Listener: Dyadic Facial Motion Synthesis via Action Diffusion
- Gated Memory Policy: In-Context Memorization and Adaptation
- TesserAct: Learning 4D Embodied World Models
- Towards Robust Multimodal Physiological Foundation Models: Handling Arbitrary Missing Modalities
- ELSA3D: Elastic Semantic Anchoring for Unified 3D Understanding and Generation
- Learning Action Priors for Cross-embodiment Robot Manipulation
- PiD: Fast and High-Resolution Latent Decoding with Pixel Diffusion
- Representation Learning on a Random Lattice
- EarthMapper: Visual Autoregressive Models for Controllable Bidirectional Satellite-Map Translation
- REED-VAE: RE-Encode Decode Training for Iterative Image Editing with Diffusion Models
- Preserving Seasonal and Trend Information: A Variational Autoencoder-Latent Space Arithmetic Based Approach for Non-stationary Learning
- Spatial PixelCNN: Generating Images from Patches
- ArrowFlow: Hierarchical Machine Learning in the Space of Permutations
- CalM: A Self-Supervised Foundation Model for Population Dynamics in Calcium Imaging Data
- Test-time Generalization for Physics through Neural Operator Splitting
- A Survey on Semantic Communication for Vision: Categories, Frameworks, Enabling Techniques, and Applications
- VQ-HPS
- DIVE: Inverting Conditional Diffusion Models for Discriminative Tasks
- Variational Autoencoder with Implicit Optimal Priors
- A hierarchical approach to imitation learning for manipulation tasks requiring time varying forces
- DigitCode: Symbolic Tokenization of Hand Motion by Anatomical Units
- UniGD: A Unified Generative-Discriminative Framework for Industrial Retrieval
- ATLAS: Learning to Recommend Across Unseen Domains
- Neural Bayes: A Generic Parameterization Method for Unsupervised Representation Learning
- T2VAttack: Adversarial Attack on Text-to-Video Diffusion Models
- SSD-Poser: Avatar Pose Estimation with State Space Duality from Sparse Observations
- Disentangle Identity, Cooperate Emotion: Correlation-Aware Emotional Talking Portrait Generation
- Token-Shuffle: Towards High-Resolution Image Generation with Autoregressive Models
- Enhancing Variational Autoencoders with Smooth Robust Latent Encoding
- Fast Autoregressive Models for Continuous Latent Generation
- DisMix: Order-Aware Mixup for Medical Imaging via Disentangling Ordinal and Non-Ordinal Features
- Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes
- RUTA: Principled Visual Token Allocation via Rate-Utility Optimization
- Teaching Foundation Models to Read mmWave: Pose-Guided Kinematic Representation for Human Behavior Understanding
- Brain2Word: Decoding Brain Activity for Language Generation
- Zero-shot Voice Conversion via Self-supervised Prosody Representation Learning
- Latent Denoising Improves Visual Alignment in Large Multimodal Models
- RiboSphere: Learning Unified and Efficient Representations of RNA Structures
- Distilling Specialized Orders for Visual Generation
- Hyper-Transforming Latent Diffusion Models
- Weakly Supervised Disentangled Representation for Goal-conditioned Reinforcement Learning
- Progressive Tandem Learning for Pattern Recognition with Deep Spiking Neural Networks
- Pre-training Generative Recommender with Multi-Identifier Item Tokenization
- Learning from Few Samples: A Survey
- Autoencoding sensory substitution
- Learning Compositional Transferability of Time Series for Source-Free Domain Adaptation
- Unifying Image Counterfactuals and Feature Attributions with Latent-Space Adversarial Attacks
- Generative Multimodal Pretraining with Discrete Diffusion Timestep Tokens
- UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding
- SUDO: Enhancing Text-to-Image Diffusion Models with Self-Supervised Direct Preference Optimization
- Adversarial Locomotion and Motion Imitation for Humanoid Policy Learning
- Towards Explainable Fake Image Detection with Multi-Modal Large Language Models
- A survey on Variational Autoencoders from a GreenAI perspective
- Transformers predicting the future. Applying attention in next-frame and time series forecasting
- Proceedings of the 2020 Joint Conference on AI Music Creativity
- A Brief Overview of Unsupervised Neural Speech Representation Learning
- Self-Supervised VQ-VAE for One-Shot Music Style Transfer
- Preventing Posterior Collapse with delta-VAEs
- Learning Hierarchical Discrete Linguistic Units from Visually-Grounded Speech
- IB-DRR: Incremental Learning with Information-Back Discrete Representation Replay
- Unified Signal Compression Using a GAN with Iterative Latent Representation Optimization
- Neuro-Symbolic Representations for Video Captioning: A Case for Leveraging Inductive Biases for Vision and Language
- A synthetic dataset of French electric load curves with temperature conditioning
- Image Editing with Diffusion Models: A Survey
- DiTaiListener: Controllable High Fidelity Listener Video Generation with Diffusion
- Personalized Text-to-Image Generation with Auto-Regressive Models
- Towards Learning to Complete Anything in Lidar
- ESC-MVQ: End-to-End Semantic Communication With Multi-Codebook Vector Quantization
- How Do I Do That? Synthesizing 3D Hand Motion and Contacts for Everyday Interactions
- Support is All You Need for Certified VAE Training
- Reasoning Errors Have a Region and a Direction in the Residual-Stream Trajectory of LLMs
- Temporal Bridges for Spatial Resolution: Enhancing Climate Data Super-Resolution with Bidirectional Alignment
- Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model
- MotionMAR: Multi-scale Auto-Regressive Human Motion Reconstruction from Sparse Observations
- TerraMind: Large-Scale Generative Multimodality for Earth Observation
- Efficient Reasoning Models: A Survey
- Elucidating the Design Space of Multimodal Protein Language Models
- Scalable Transceiver Design for Multi-User Communication in FDD Massive MIMO Systems via Deep Learning
- QAMA: Scalable Quantum Annealing Multi-Head Attention Operator for Deep Learning
- REPA-E: Unlocking VAE for End-to-End Tuning with Latent Diffusion Transformers
- FedRecon: Missing Modality Reconstruction in Heterogeneous Distributed Environments
- OctGPT: Octree-based Multiscale Autoregressive Models for 3D Shape Generation
- Investigating the Role of Bilateral Symmetry for Inpainting Brain MRI
- Exploration into Translation-Equivariant Image Quantization
- GigaTok: Scaling Visual Tokenizers to 3 Billion Parameters for Autoregressive Image Generation
- On The Landscape of Spoken Language Models: A Comprehensive Survey
- Millions of States: Designing a Scalable MoE Architecture with RWKV-7 Meta-learner
- MineWorld: a Real-Time and Open-Source Interactive World Model on Minecraft
- MotionDreamer: One-to-Many Motion Synthesis with Localized Generative Masked Transformer
- Vector Quantized-Elites: Unsupervised and Problem-Agnostic Quality-Diversity Optimization
- Scaling Laws for Native Multimodal Models
- Synthetic CT Generation from Time-of-Flight Non-Attenutaion-Corrected PET for Whole-Body PET Attenuation Correction
- Psychoacoustic Calibration of Loss Functions for Efficient End-to-End Neural Audio Coding
- Multi-Sense Embeddings for Language Models and Knowledge Distillation
- GARF: Learning Generalizable 3D Reassembly for Real-World Fractures
- Low-Rate Semantic Communication with Codebook-based Conditional Generative Models
- Bridging the Gap between Continuous and Informative Discrete Representations by Random Product Quantization
Discussions
Related