vix.ing · top · new · best · stats · spec

Maaz, Muhammad

  1. Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
    2023/06/08 by Muhammad Maaz, Hanoona Rasheed, Maaz, Muhammad +5 · 1 voice · 201 citations
    Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Pose and Action Recognition #Multimodal Machine Learning Applications #Topic Modeling #cs.CV
  2. MaPLe: Multi-modal Prompt Learning
    2022/10/06 by Muhammad Uzair Khattak, Khattak, Muhammad Uzair, Hanoona Rasheed +7 · 134 citations
    Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Natural Language Processing Techniques
  3. GLaMM: Pixel Grounding Large Multimodal Model
    2023/11/06 by Hanoona Rasheed, Muhammad Maaz, Rasheed, Hanoona +16 · 112 citations
    Arts and Humanities · Computer Science · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Speech and dialogue systems #Subtitles and Audiovisual Media
  4. UNETR++: Delving into Efficient and Accurate 3D Medical Image Segmentation
    2022/12/08 by Abdelrahman Shaker, Shaker, Abdelrahman, Muhammad Maaz +9 · 31 citations
    Computer Science · Medicine · #Advanced Neural Network Applications #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Medical Image Segmentation Techniques #Radiomics and Machine Learning in Medical Imaging
  5. Fine-tuned CLIP Models are Efficient Video Learners
    2022/12/06 by Hanoona Rasheed, Rasheed, Hanoona, Muhammad Uzair Khattak +7 · 29 citations
    Biochemistry, Genetics and Molecular Biology · Computer Science · #Artificial Intelligence (cs.AI) #Cancer-related molecular mechanisms research #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Multimodal Machine Learning Applications
  6. VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding
    2024/06/13 by Muhammad Maaz, Maaz, Muhammad, Hanoona Rasheed +5 · 34 citations
    Computer Science · #Advanced Vision and Imaging #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Human Pose and Action Recognition
  7. EdgeNeXt: Efficiently Amalgamated CNN-Transformer Architecture for Mobile Vision Applications
    2022/06/21 by Muhammad Maaz, Maaz, Muhammad, Abdelrahman Shaker +11 · 16 citations
    Computer Science · Medicine · #Advanced Neural Network Applications #COVID-19 diagnosis using AI #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences
  8. Class-agnostic Object Detection with Multi-modal Transformer
    2021/11/22 by Maaz, Muhammad, Rasheed, Hanoona, Khan, Salman +3 · 9 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  9. Bridging the Gap between Object and Image-level Representations for Open-Vocabulary Detection
    2022/07/07 by Hanoona Rasheed, Muhammad Maaz, Rasheed, Hanoona +7 · 10 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling
  10. PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding
    2025/04/17 by Jang Hyun Cho, Cho, Jang Hyun, Andrea Madotto +54 · 30 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Machine Learning (cs.LG) #Multimodal Machine Learning Applications
  11. SwiftFormer: Efficient Additive Attention for Transformer-based Real-time Mobile Vision Applications
    2023/03/27 by Shaker, Abdelrahman, Maaz, Muhammad, Rasheed, Hanoona +3 · 10 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  12. PALO: A Polyglot Large Multimodal Model for 5B People
    2024/02/22 by Muhammad Maaz, Hanoona Rasheed, Maaz, Muhammad +15 · 9 citations
    Social Sciences · #Human Mobility and Location-Based Analysis
  13. PG-Video-LLaVA: Pixel Grounding Large Video-Language Models
    2023/11/22 by Shehan Munasinghe, R. Thushara, Munasinghe, Shehan +11 · 5 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling
  14. VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos
    2025/06/05 by Hanoona Rasheed, Abdelrahman Shaker, Rasheed, Hanoona +10 · 8 citations
    Computer Science · Social Sciences · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Intelligent Tutoring Systems and Adaptive Learning #Mathematics Education and Teaching Techniques #Multimodal Machine Learning Applications
  15. Mobile-VideoGPT: Fast and Accurate Model for Mobile Video Understanding
    2025/03/27 by Abdelrahman Shaker, Shaker, Abdelrahman, Muhammad Maaz +9 · 3 citations
    Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Pose and Action Recognition #Multimodal Machine Learning Applications #Video Analysis and Summarization