vix.ing · top · new · best · stats · spec

Yang, Antoine

  1. Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
    2025/07/07 by Gheorghe Comanici, Eric Bieber, Comanici, Gheorghe +6844 · 8 voices · 1363 citations
    #cs.CL #cs.AI
  2. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
    2024/03/08 by Gemini Robotics Team, Petko Georgiev, Gemini Team +2277 · 4 voices · 555 citations
    Computer Science · #Semantic Web and Ontologies
  3. Gemma 3 Technical Report
    2025/03/25 by Aishwarya Kamath, Gemma Team, Johan Ferret +418 · 3 voices · 469 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #cs.AI #cs.CL
  4. Vid2Seq: Large-Scale Pretraining of a Visual Language Model for Dense Video Captioning
    2023/02/27 by Antoine Yang, Yang, Antoine, Arsha Nagrani +13 · 40 citations
    Computer Science · #Multimodal Machine Learning Applications #Human Pose and Action Recognition #Video Analysis and Summarization
  5. Just Ask: Learning to Answer Questions from Millions of Narrated Videos
    2020/12/01 by Yang, Antoine, Miech, Antoine, Sivic, Josef +2 · 15 citations
    #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  6. TubeDETR: Spatio-Temporal Video Grounding with Transformers
    2022/03/30 by Yang, Antoine, Miech, Antoine, Sivic, Josef +2 · 14 citations
    #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  7. Zero-Shot Video Question Answering via Frozen Bidirectional Language Models
    2022/06/16 by Antoine Yang, Yang, Antoine, Antoine Miech +7 · 14 citations
    Computer Science · #Multimodal Machine Learning Applications #Domain Adaptation and Few-Shot Learning #Advanced Image and Video Retrieval Techniques
  8. VidChapters-7M: Video Chapters at Scale
    2023/09/25 by Yang, Antoine, Nagrani, Arsha, Laptev, Ivan +2 · 6 citations
    #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  9. Chapter-Llama: Efficient Chaptering in Hour-Long Videos with LLMs
    2025/03/31 by Lucas Ventura, Ventura, Lucas, Antoine Yang +5 · 2 voices · 3 citations
    #cs.CV
  10. Learning to Answer Visual Questions from Web Videos
    2022/05/10 by Antoine Yang, Antoine Miech, Yang, Antoine +7 · 1 citation
    Computer Science · #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications #Topic Modeling