2017/01/01 by Yonatan Belinkov, Nadir Durrani, Fahim Dalvi +2 · 283 citations
Computer Science · #Artificial intelligence #Artificial neural network #Character (mathematics) #Computer science #Encoder #Encoding (memory) #Granularity #Identity (music) #Language model #Linguistics #Machine translation #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Natural language processing #Process (computing) #Programming language #Topic Modeling #Word (group theory) #cs.CL
paper · pdf · doi:10.18653/v1/p17-1080
published as ACL 55 (2017), volume 1, 861-872 · Updated decoder experiments
openalex publication_date 2017/01/01 · arxiv created 2018/10/22 · arxiv updated 2018/10/23 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05
Neural machine translation (MT) models obtain state-of-the-art performance while maintaining a simple, end-to-end architecture. However, little is known about what these models learn about source and target languages during the training process.