Tseng, Shao-Yen
- LDM3D: Latent Diffusion Model for 3D
2023/05/18 by Gabriela Ben Melech Stan, Diana Wofk, Stan, Gabriela Ben Melech +19 · 1 voice · 12 citations
Computer Science · #cs.CV
- LVLM-Interpret: An Interpretability Tool for Large Vision-Language Models
2024/04/03 by Gabriela Ben Melech Stan, Stan, Gabriela Ben Melech, Estelle Aflalo +17 · 22 citations
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Natural Language Processing Techniques #Semantic Web and Ontologies
- VL-InterpreT: An Interactive Visualization Tool for Interpreting Vision-Language Transformers
2022/03/30 by Estelle Aflalo, Meng Du, Aflalo, Estelle +11 · 7 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling
- Automatic prediction of suicidal risk in military couples using\n multimodal interaction cues from couples conversations
2019/11/26 by Sandeep Nallan Chakravarthula, Chakravarthula, Sandeep Nallan, Md Nasir +15 · 2 citations
Psychology · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Mental Health via Writing #electronic engineering #information engineering
- LDM3D-VR: Latent Diffusion Model for 3D VR
2023/11/06 by Gabriela Ben Melech Stan, Stan, Gabriela Ben Melech, Diana Wofk +11 · 3 citations
Computer Science · #Advanced Image and Video Retrieval Techniques #Advanced Vision and Imaging #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Video Analysis and Summarization
- LLaVA-Gemma: Accelerating Multimodal Foundation Models with a Compact Language Model
2024/03/29 by Musashi Hinck, Hinck, Musashi, Matthew Olson +7 · 2 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Semantic Web and Ontologies #Speech and dialogue systems
- MuMUR : Multilingual Multimodal Universal Retrieval
2022/08/24 by Madasu, Avinash, Aflalo, Estelle, Stan, Gabriela Ben Melech +4 · 1 citation
#Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
- Debiasing Large Vision-Language Models by Ablating Protected Attribute Representations
2024/10/17 by Neale Ratzlaff, Ratzlaff, Neale, Matthew Olson +9 · 2 citations
Computer Science · #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling
- FastRM: An efficient and automatic explainability framework for multimodal generative models
2024/12/02 by Gabriela Ben-Melech Stan, Stan, Gabriela Ben-Melech, Estelle Aflalo +13 · 2 citations
Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Topic Modeling
- Probing the Representational Power of Sparse Autoencoders in Vision Models
2025/08/15 by Olson, Matthew Lyle, Hinck, Musashi, Ratzlaff, Neale +4 · 3 citations
#Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability
2024/12/19 by Estelle Aflalo, Gabriela Ben Melech Stan, Aflalo, Estelle +13 · 1 citation
Computer Science · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling
- ManagerTower: Aggregating the Insights of Uni-Modal Experts for Vision-Language Representation Learning
2023/05/31 by Xu Xiao, Bei Li, Xu, Xiao +15 · 1 citation
Computer Science · #Multimodal Machine Learning Applications #Advanced Image and Video Retrieval Techniques #Domain Adaptation and Few-Shot Learning