Mo, Shentong
- DiT-3D: Exploring Plain Diffusion Transformers for 3D Shape Generation
2023/07/04 by Shentong Mo, Enze Xie, Mo, Shentong +11 · 24 citations
Engineering · Computer Science · #3D Shape Modeling and Analysis #Generative Adversarial Networks and Image Synthesis #Advanced Vision and Imaging
- Context Autoencoder for Self-Supervised Representation Learning
2022/02/07 by Chen, Xiaokang, Ding, Mingyu, Wang, Xiaodi +7 · 16 citations
#Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
- Tree of Uncertain Thoughts Reasoning for Large Language Models
2023/09/14 by Shentong Mo, Mo, Shentong, Xin Miao +2 · 1 voice · 3 citations
Computer Science · #Artificial Intelligence (cs.AI) #Bayesian Modeling and Causal Inference #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling #cs.AI #cs.CL #cs.LG
- Unveiling the Power of Audio-Visual Early Fusion Transformers with Dense Interactions through Masked Modeling
2023/12/02 by Mo, Shentong, Morgado, Pedro · 9 citations
#Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimedia (cs.MM) #Sound (cs.SD)
- Audio-Visual Grouping Network for Sound Localization from Mixtures
2023/03/29 by Shentong Mo, Mo, Shentong, Yapeng Tian +1 · 7 citations
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #Digital Media Forensic Detection #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimedia (cs.MM) #Music and Audio Processing #Speech and Audio Processing
- A Closer Look at Weakly-Supervised Audio-Visual Source Localization
2022/08/30 by Shentong Mo, Mo, Shentong, Pedro Morgado +1 · 6 citations
Computer Science · #Speech and Audio Processing #Music and Audio Processing #Video Analysis and Summarization
- Audio-Synchronized Visual Animation
2024/03/08 by Lin Zhang, Zhang, Lin, Shentong Mo +5 · 7 citations
Arts and Humanities · Engineering · Social Sciences · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Motion and Animation #Multimedia Communication and Technology #Subtitles and Audiovisual Media
- Localizing Visual Sounds the Easy Way
2022/03/17 by Mo, Shentong, Morgado, Pedro · 4 citations
#Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Multimedia (cs.MM) #Sound (cs.SD) #electronic engineering #information engineering
- AV-SAM: Segment Anything Model Meets Audio-Visual Localization and Segmentation
2023/05/03 by Mo, Shentong, Tian, Yapeng · 5 citations
#Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Multimedia (cs.MM) #Sound (cs.SD) #electronic engineering #information engineering
- Weakly-Supervised Audio-Visual Segmentation
2023/11/25 by Mo, Shentong, Raj, Bhiksha · 5 citations
#Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Multimedia (cs.MM) #Sound (cs.SD) #electronic engineering #information engineering
- DiffComplete: Diffusion-based Generative 3D Shape Completion
2023/06/28 by Ruihang Chu, Enze Xie, Chu, Ruihang +11 · 4 citations
Computer Science · Engineering · #3D Shape Modeling and Analysis #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Human Pose and Action Recognition
- High-Modality Multimodal Transformer: Quantifying Modality & Interaction Heterogeneity for High-Modality Representation Learning
2022/03/02 by Paul Pu Liang, Liang, Paul Pu, Yiwei Lyu +13 · 3 citations
Computer Science · #Speech and dialogue systems #Multimodal Machine Learning Applications #Music and Audio Processing
- Text-to-Audio Generation Synchronized with Videos
2024/03/08 by Mo, Shentong, Shi, Jing, Tian, Yapeng · 4 citations
#Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Multimedia (cs.MM) #Sound (cs.SD) #electronic engineering #information engineering
- Audio-Visual Class-Incremental Learning
2023/08/21 by Pian, Weiguo, Mo, Shentong, Guo, Yunhui +1 · 3 citations
#Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
- Connecting Joint-Embedding Predictive Architecture with Contrastive Self-supervised Learning
2024/10/25 by Mo, Shentong, Tong, Shengbang · 4 citations
#Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Image and Video Processing (eess.IV) #Machine Learning (cs.LG) #Signal Processing (eess.SP) #electronic engineering #information engineering
- Class-Incremental Grouping Network for Continual Audio-Visual Learning
2023/09/11 by Shentong Mo, Weiguo Pian, Mo, Shentong +3 · 2 citations
Computer Science · Neuroscience · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Hearing Loss and Rehabilitation #Machine Learning (cs.LG) #Multimedia (cs.MM) #Music and Audio Processing #Speech and Audio Processing
- Fast Training of Diffusion Transformer with Extreme Masking for 3D Point Clouds Generation
2023/12/12 by Mo, Shentong, Xie, Enze, Wu, Yue +3 · 2 citations
#Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Continual Audio-Visual Sound Separation
2024/11/05 by Weiguo Pian, Yiyang Nan, Pian, Weiguo +9 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Image and Signal Denoising Methods #Machine Learning (cs.LG) #Multimedia (cs.MM) #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
- MultiMed: Massively Multimodal and Multitask Medical Understanding
2024/08/22 by Mo, Shentong, Liang, Paul Pu · 2 citations
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimedia (cs.MM)
- Multi-modal Self-supervised Pre-training for Regulatory Genome Across Cell Types
2021/10/11 by Shentong Mo, Mo, Shentong, Xi Fu +15 · 1 citation
Biochemistry, Genetics and Molecular Biology · #Artificial Intelligence (cs.AI) #FOS: Biological sciences #FOS: Computer and information sciences #Genomics (q-bio.GN) #Genomics and Chromatin Dynamics #Genomics and Phylogenetic Studies #Machine Learning (cs.LG) #RNA and protein synthesis mechanisms
- Scaling Diffusion Mamba with Bidirectional SSMs for Efficient Image and Video Generation
2024/05/24 by Shentong Mo, Mo, Shentong, Yapeng Tian +1 · 2 citations
Computer Science · #Advanced Image and Video Retrieval Techniques #Advanced Vision and Imaging #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Face recognition and analysis #Machine Learning (cs.LG)
- Siamese Prototypical Contrastive Learning
2022/08/18 by Mo, Shentong, Sun, Zhun, Li, Chao · 1 citation
#Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Rethinking Prototypical Contrastive Learning through Alignment, Uniformity and Correlation
2022/10/18 by Shentong Mo, Zhun Sun, Mo, Shentong +3 · 1 citation
Computer Science · Medicine · #Artificial Intelligence (cs.AI) #COVID-19 diagnosis using AI #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications
- Variantional autoencoder with decremental information bottleneck for disentanglement
2023/03/22 by Jiantao Wu, Shentong Mo, Wu, Jiantao +10 · 1 citation
Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Digital Media Forensic Detection #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Machine Learning (cs.LG)
- Aligning Audio-Visual Joint Representations with an Agentic Workflow
2024/10/30 by Shentong Mo, Mo, Shentong, Yibing Song +1 · 2 citations
Computer Science · Engineering · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Human Motion and Animation #Machine Learning (cs.LG) #Multimedia (cs.MM) #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
- LSPT: Long-term Spatial Prompt Tuning for Visual Representation Learning
2024/02/27 by Shentong Mo, Yansen Wang, Mo, Shentong +5 · 1 citation
Computer Science · #Advanced Image and Video Retrieval Techniques #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications
- DiffGAP: A Lightweight Diffusion Module in Contrastive Space for Bridging Cross-Model Gap
2025/03/15 by Mo, Shentong, Chen, Zehua, Bao, Fan +1 · 2 citations
#Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
- Semantic Grouping Network for Audio Source Separation
2024/07/04 by Mo, Shentong, Tian, Yapeng · 1 citation
#Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Multimedia (cs.MM) #Sound (cs.SD) #electronic engineering #information engineering
- Rethinking Positive Pairs in Contrastive Learning
2024/10/23 by Wu, Jiantao, Atito, Sara, Feng, Zhenhua +3 · 1 citation
#Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- MA-AVT: Modality Alignment for Parameter-Efficient Audio-Visual Transformers
2024/06/07 by Tanvir Mahmud, Mahmud, Tanvir, Shentong Mo +5 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
- Modality-Inconsistent Continual Learning of Multimodal Large Language Models
2024/12/17 by Weiguo Pian, Pian, Weiguo, Shijian Deng +8 · 1 citation
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- Multi-scale Multi-instance Visual Sound Localization and Segmentation
2024/08/31 by Mo, Shentong, Wang, Haofan · 1 citation
#Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Multimedia (cs.MM) #Sound (cs.SD) #electronic engineering #information engineering
- The Dynamic Duo of Collaborative Masking and Target for Advanced Masked Autoencoder Learning
2024/12/23 by Shentong Mo, Mo, Shentong · 1 citation
Computer Science · #Neural Networks and Applications #Image Processing and 3D Reconstruction
- IoT-LM: Large Multisensory Language Models for the Internet of Things
2024/07/13 by Shentong Mo, Mo, Shentong, Russ R. Salakhutdinov +5 · 1 citation
Agricultural and Biological Sciences · Social Sciences · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Food Supply Chain Traceability #Machine Learning (cs.LG) #Multimedia (cs.MM) #Public Relations and Crisis Communication