vix.ing · top · new · best · stats · spec

Mo, Shentong

  1. DiT-3D: Exploring Plain Diffusion Transformers for 3D Shape Generation
    2023/07/04 by Shentong Mo, Enze Xie, Mo, Shentong +11 · 24 citations
    Engineering · Computer Science · #3D Shape Modeling and Analysis #Generative Adversarial Networks and Image Synthesis #Advanced Vision and Imaging
  2. Context Autoencoder for Self-Supervised Representation Learning
    2022/02/07 by Chen, Xiaokang, Ding, Mingyu, Wang, Xiaodi +7 · 16 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  3. Tree of Uncertain Thoughts Reasoning for Large Language Models
    2023/09/14 by Shentong Mo, Mo, Shentong, Xin Miao +2 · 1 voice · 3 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Bayesian Modeling and Causal Inference #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling #cs.AI #cs.CL #cs.LG
  4. Unveiling the Power of Audio-Visual Early Fusion Transformers with Dense Interactions through Masked Modeling
    2023/12/02 by Mo, Shentong, Morgado, Pedro · 9 citations
    #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimedia (cs.MM) #Sound (cs.SD)
  5. Audio-Visual Grouping Network for Sound Localization from Mixtures
    2023/03/29 by Shentong Mo, Mo, Shentong, Yapeng Tian +1 · 7 citations
    Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #Digital Media Forensic Detection #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimedia (cs.MM) #Music and Audio Processing #Speech and Audio Processing
  6. A Closer Look at Weakly-Supervised Audio-Visual Source Localization
    2022/08/30 by Shentong Mo, Mo, Shentong, Pedro Morgado +1 · 6 citations
    Computer Science · #Speech and Audio Processing #Music and Audio Processing #Video Analysis and Summarization
  7. Audio-Synchronized Visual Animation
    2024/03/08 by Lin Zhang, Zhang, Lin, Shentong Mo +5 · 7 citations
    Arts and Humanities · Engineering · Social Sciences · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Motion and Animation #Multimedia Communication and Technology #Subtitles and Audiovisual Media
  8. Localizing Visual Sounds the Easy Way
    2022/03/17 by Mo, Shentong, Morgado, Pedro · 4 citations
    #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Multimedia (cs.MM) #Sound (cs.SD) #electronic engineering #information engineering
  9. AV-SAM: Segment Anything Model Meets Audio-Visual Localization and Segmentation
    2023/05/03 by Mo, Shentong, Tian, Yapeng · 5 citations
    #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Multimedia (cs.MM) #Sound (cs.SD) #electronic engineering #information engineering
  10. Weakly-Supervised Audio-Visual Segmentation
    2023/11/25 by Mo, Shentong, Raj, Bhiksha · 5 citations
    #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Multimedia (cs.MM) #Sound (cs.SD) #electronic engineering #information engineering
  11. DiffComplete: Diffusion-based Generative 3D Shape Completion
    2023/06/28 by Ruihang Chu, Enze Xie, Chu, Ruihang +11 · 4 citations
    Computer Science · Engineering · #3D Shape Modeling and Analysis #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Human Pose and Action Recognition
  12. High-Modality Multimodal Transformer: Quantifying Modality & Interaction Heterogeneity for High-Modality Representation Learning
    2022/03/02 by Paul Pu Liang, Liang, Paul Pu, Yiwei Lyu +13 · 3 citations
    Computer Science · #Speech and dialogue systems #Multimodal Machine Learning Applications #Music and Audio Processing
  13. Text-to-Audio Generation Synchronized with Videos
    2024/03/08 by Mo, Shentong, Shi, Jing, Tian, Yapeng · 4 citations
    #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Multimedia (cs.MM) #Sound (cs.SD) #electronic engineering #information engineering
  14. Audio-Visual Class-Incremental Learning
    2023/08/21 by Pian, Weiguo, Mo, Shentong, Guo, Yunhui +1 · 3 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  15. Connecting Joint-Embedding Predictive Architecture with Contrastive Self-supervised Learning
    2024/10/25 by Mo, Shentong, Tong, Shengbang · 4 citations
    #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Image and Video Processing (eess.IV) #Machine Learning (cs.LG) #Signal Processing (eess.SP) #electronic engineering #information engineering
  16. Class-Incremental Grouping Network for Continual Audio-Visual Learning
    2023/09/11 by Shentong Mo, Weiguo Pian, Mo, Shentong +3 · 2 citations
    Computer Science · Neuroscience · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Hearing Loss and Rehabilitation #Machine Learning (cs.LG) #Multimedia (cs.MM) #Music and Audio Processing #Speech and Audio Processing
  17. Fast Training of Diffusion Transformer with Extreme Masking for 3D Point Clouds Generation
    2023/12/12 by Mo, Shentong, Xie, Enze, Wu, Yue +3 · 2 citations
    #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  18. Continual Audio-Visual Sound Separation
    2024/11/05 by Weiguo Pian, Yiyang Nan, Pian, Weiguo +9 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Image and Signal Denoising Methods #Machine Learning (cs.LG) #Multimedia (cs.MM) #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
  19. MultiMed: Massively Multimodal and Multitask Medical Understanding
    2024/08/22 by Mo, Shentong, Liang, Paul Pu · 2 citations
    #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimedia (cs.MM)
  20. Multi-modal Self-supervised Pre-training for Regulatory Genome Across Cell Types
    2021/10/11 by Shentong Mo, Mo, Shentong, Xi Fu +15 · 1 citation
    Biochemistry, Genetics and Molecular Biology · #Artificial Intelligence (cs.AI) #FOS: Biological sciences #FOS: Computer and information sciences #Genomics (q-bio.GN) #Genomics and Chromatin Dynamics #Genomics and Phylogenetic Studies #Machine Learning (cs.LG) #RNA and protein synthesis mechanisms
  21. Scaling Diffusion Mamba with Bidirectional SSMs for Efficient Image and Video Generation
    2024/05/24 by Shentong Mo, Mo, Shentong, Yapeng Tian +1 · 2 citations
    Computer Science · #Advanced Image and Video Retrieval Techniques #Advanced Vision and Imaging #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Face recognition and analysis #Machine Learning (cs.LG)
  22. Siamese Prototypical Contrastive Learning
    2022/08/18 by Mo, Shentong, Sun, Zhun, Li, Chao · 1 citation
    #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  23. Rethinking Prototypical Contrastive Learning through Alignment, Uniformity and Correlation
    2022/10/18 by Shentong Mo, Zhun Sun, Mo, Shentong +3 · 1 citation
    Computer Science · Medicine · #Artificial Intelligence (cs.AI) #COVID-19 diagnosis using AI #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications
  24. Variantional autoencoder with decremental information bottleneck for disentanglement
    2023/03/22 by Jiantao Wu, Shentong Mo, Wu, Jiantao +10 · 1 citation
    Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Digital Media Forensic Detection #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Machine Learning (cs.LG)
  25. Aligning Audio-Visual Joint Representations with an Agentic Workflow
    2024/10/30 by Shentong Mo, Mo, Shentong, Yibing Song +1 · 2 citations
    Computer Science · Engineering · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Human Motion and Animation #Machine Learning (cs.LG) #Multimedia (cs.MM) #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
  26. LSPT: Long-term Spatial Prompt Tuning for Visual Representation Learning
    2024/02/27 by Shentong Mo, Yansen Wang, Mo, Shentong +5 · 1 citation
    Computer Science · #Advanced Image and Video Retrieval Techniques #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications
  27. DiffGAP: A Lightweight Diffusion Module in Contrastive Space for Bridging Cross-Model Gap
    2025/03/15 by Mo, Shentong, Chen, Zehua, Bao, Fan +1 · 2 citations
    #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
  28. Semantic Grouping Network for Audio Source Separation
    2024/07/04 by Mo, Shentong, Tian, Yapeng · 1 citation
    #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Multimedia (cs.MM) #Sound (cs.SD) #electronic engineering #information engineering
  29. Rethinking Positive Pairs in Contrastive Learning
    2024/10/23 by Wu, Jiantao, Atito, Sara, Feng, Zhenhua +3 · 1 citation
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  30. MA-AVT: Modality Alignment for Parameter-Efficient Audio-Visual Transformers
    2024/06/07 by Tanvir Mahmud, Mahmud, Tanvir, Shentong Mo +5 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
  31. Modality-Inconsistent Continual Learning of Multimodal Large Language Models
    2024/12/17 by Weiguo Pian, Pian, Weiguo, Shijian Deng +8 · 1 citation
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  32. Multi-scale Multi-instance Visual Sound Localization and Segmentation
    2024/08/31 by Mo, Shentong, Wang, Haofan · 1 citation
    #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Multimedia (cs.MM) #Sound (cs.SD) #electronic engineering #information engineering
  33. The Dynamic Duo of Collaborative Masking and Target for Advanced Masked Autoencoder Learning
    2024/12/23 by Shentong Mo, Mo, Shentong · 1 citation
    Computer Science · #Neural Networks and Applications #Image Processing and 3D Reconstruction
  34. IoT-LM: Large Multisensory Language Models for the Internet of Things
    2024/07/13 by Shentong Mo, Mo, Shentong, Russ R. Salakhutdinov +5 · 1 citation
    Agricultural and Biological Sciences · Social Sciences · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Food Supply Chain Traceability #Machine Learning (cs.LG) #Multimedia (cs.MM) #Public Relations and Crisis Communication