Neural Machine Translation by Jointly Learning to Align and Translate
2014/09/01 by Dzmitry Bahdanau, Kyunghyun Cho, Bahdanau, Dzmitry +3 · 9 voices · 451 citations
Computer Science · #Natural Language Processing Techniques #Topic Modeling #Multimodal Machine Learning Applications
paper · pdf · doi:10.48550/arxiv.1409.0473
Abstract
Neural machine translation is a recently proposed approach to machine translation. Unlike the traditional statistical machine translation, the neural machine translation aims at building a single neural network that can be jointly tuned to maximize the translation performance. The models proposed recently for neural machine translation often belong to a family of encoder-decoders and consists of an encoder that encodes a source sentence into a fixed-length vector from which a decoder generates a translation. In this paper, we conjecture that the use of a fixed-length vector is a bottleneck in improving the performance of this basic encoder-decoder architecture, and propose to extend this by allowing a model to automatically (soft-)search for parts of a source sentence that are relevant to predicting a target word, without having to form these parts as a hard segment explicitly. With this new approach, we achieve a translation performance comparable to the existing state-of-the-art phrase-based system on the task of English-to-French translation. Furthermore, qualitative analysis reveals that the (soft-)alignments found by the model agree well with our intuition.
Citations
Cited by
- Biomedical Machine Translation for Low-Resource Arabic-Script Languages via Cross-Lingual Transfer and LoRA Adapter Merging
- Pretraining Recurrent Networks without Recurrence
- Breaking the Data Barrier in Learning Symbolic Computation: A Case Study on Variable Ordering Suggestion for Cylindrical Algebraic Decomposition
- Graph-Theoretic Neural Network Fragmentation with Covariant Direct Molecular Force Learning: Enabling Coupled-Cluster Accuracy AIMD for Fluxional Systems
- A Deep Reinforcement Learning Algorithm for the Vehicle Routing Problem with Stochastic Demands and Outsourcing
- Neuro-Symbolic Meta-Policies for Temporal Knowledge-Graph Memory under Partial Observability
- An Explicit World Model Based on Data-First Ontology: DaoQL Multimodal Storage Validation and Counterfactual Reasoning Evaluation
- Jet quenching identification via supervised learning in simulated heavy-ion collisions
- MATNet: multi-level fusion transformer-based model for day-ahead PV generation forecasting
- Attention to Mamba: A Recipe for Cross-Architecture Distillation
- Mamba-3: Improved Sequence Modeling using State Space Principles
- Remapping and navigation of an embedding space via error minimization: a fundamental organizational principle of cognition in natural and artificial systems
- PHOTON: Hierarchical Autoregressive Modeling for Lightspeed and Memory-Efficient Language Generation
- Fast weight programming and linear transformers: from machine learning to neurobiology
- QiMeng: Fully Automated Hardware and Software Design for Processor Chip
- Transformers are Graph Neural Networks
- Blending Complementary Memory Systems in Hybrid Quadratic-Linear Transformers
- ATLAS: Learning to Optimally Memorize the Context at Test Time
- Log-Linear Attention
- NNN: Next-Generation Neural Networks for Marketing Measurement
- Multi-Token Attention
- Chapter-Llama: Efficient Chaptering in Hour-Long Videos with LLMs
- CalibQuant: 1-Bit KV Cache Quantization for Multimodal LLMs
- Learning When Not to Attend Globally
- AudioGAN: A Compact and Efficient Framework for Real-Time High-Fidelity Text-to-Audio Generation
- Attention Residuals
- Towards High-Level Semantic Intelligence
- Selecting Language Models for Social Science: Start Small, Start Open, and Validate
- Understanding Human-like Solutions in Combinatorial Optimization via Learning and Search
- Anchor Attention, Small Cache: Code Generation with Large Language Models
- PLATO: Pointer Learner for Agent and Task Openness
- Optimizing Transformer Neural Network for Real-Time Outlier Detection on FPGAs
- Statistical vs. Deep Learning Models for Estimating Substance Overdose Excess Mortality in the US
- SPOT!: Map-Guided LLM Agent for Unsupervised Multi-CCTV Dynamic Object Tracking
- Explainable time-series forecasting with sampling-free SHAP for Transformers
- BiCoR-Seg: Bidirectional Co-Refinement Framework for High-Resolution Remote Sensing Image Segmentation
- RIS-Enabled Smart Wireless Environments: Fundamentals and Distributed Optimization
- Asynchronous Pipeline Parallelism for Real-Time Multilingual Lip Synchronization in Video Communication Systems
- Adversarial Robustness of Vision in Open Foundation Models
- InfinityEBSD : Metrics-Guided Infinite-Size EBSD Map Generation With Diffusion Models
- KOSS: Kalman-Optimal Selective State Spaces for Long-Term Sequence Modeling
- Evaluating OpenAI GPT Models for Translation of Endangered Uralic Languages: A Comparison of Reasoning and Non-Reasoning Architectures
- SegGraph: Leveraging Graphs of SAM Segments for Few-Shot 3D Part Segmentation
- From Minutes to Days: Scaling Intracranial Speech Decoding with Supervised Pretraining
- An Empirical Study on Chinese Character Decomposition in Multiword Expression-Aware Neural Machine Translation
- Incentives or Ontology? A Structural Rebuttal to OpenAI's Hallucination Thesis
- Advancing Bangla Machine Translation Through Informal Datasets
- Towards Deep Learning Surrogate for the Forward Problem in Electrocardiology: A Scalable Alternative to Physics-Based Models
- WAY: Estimation of Vessel Destination in Worldwide AIS Trajectory
- Behavior and Representation in Large Language Models for Combinatorial Optimization: From Feature Extraction to Algorithm Selection
- HydroDiffusion: Diffusion-Based Probabilistic Streamflow Forecasting with a State Space Backbone
- Kinetic Mining in Context: Few-Shot Action Synthesis via Text-to-Motion Distillation
- Unlocking the Address Book: Dissecting the Sparse Semantic Structure of LLM Key-Value Caches via Sparse Autoencoders
- Rates and architectures for learning geometrically non-trivial operators
- PVeRA: Probabilistic Vector-Based Random Matrix Adaptation
- Towards a Relationship-Aware Transformer for Tabular Data
- QL-LSTM: A Parameter-Efficient LSTM for Stable Long-Sequence Modeling
- Learning the Cosmic Web: Graph-based Classification of Simulated Galaxies by their Dark Matter Environments
- HTR-ConvText: Leveraging Convolution and Textual Information for Handwritten Text Recognition
- Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models
- ProtoEFNet: Dynamic Prototype Learning for Inherently Interpretable Ejection Fraction Estimation in Echocardiography
- Robust Probabilistic Load Forecasting for a Single Household: A Comparative Study from SARIMA to Transformers on the REFIT Dataset
- Interpretable Alzheimer's Diagnosis via Multimodal Fusion of Regional Brain Experts
- BanglaSentNet: An Explainable Hybrid Deep Learning Framework for Multi-Aspect Sentiment Analysis with Cross-Domain Transfer Learning
- "When Data is Scarce, Prompt Smarter"... Approaches to Grammatical Error Correction in Low-Resource Settings
- Hierarchical Spatio-Temporal Attention Network with Adaptive Risk-Aware Decision for Forward Collision Warning in Complex Scenarios
- SimDiff: Simpler Yet Better Diffusion Model for Time Series Point Forecasting
- Large-Scale In-Game Outcome Forecasting for Match, Team and Players in Football using an Axial Transformer Neural Network
- GContextFormer: A global context-aware hybrid multi-head attention approach with scaled additive aggregation for multimodal trajectory prediction
- Foundations of Artificial Intelligence Frameworks: Notion and Limits of AGI
- Attention-Guided Feature Fusion (AGFF) Model for Integrating Statistical and Semantic Features in News Text Classification
- Computation for Epidemic Prediction with Graph Neural Network by Model Combination
- What Really Counts? Examining Step and Token Level Attribution in Multilingual CoT Reasoning
- LAYA: Layer-wise Attention Aggregation for Interpretable Depth-Aware Neural Networks
- EchoLSTM: A Self-Reflective Recurrent Network for Stabilizing Long-Range Memory
- MTMed3D: A Multi-Task Transformer-Based Model for 3D Medical Imaging
- Evaluation of Attention Mechanisms in U-Net Architectures for Semantic Segmentation of Brazilian Rock Art Petroglyphs
- Additive Large Language Models for Semi-Structured Text
- Scaling Open-Weight Large Language Models for Hydropower Regulatory Information Extraction: A Systematic Analysis
- TEDxTN: A Three-way Speech Translation Corpus for Code-Switched Tunisian Arabic - English
- Know Your Limits: Entropy Estimation Modeling for Compression and Generalization
- MVSMamba: Multi-View Stereo with State Space Model
- H-Model: Dynamic Neural Architectures for Adaptive Processing
- A Circular Argument : Does RoPE need to be Equivariant for Vision?
- Reduced Density Matrices Through Machine Learning
- A Picture is Worth a Thousand (Correct) Captions: A Vision-Guided Judge-Corrector System for Multimodal Machine Translation
- Sensitivity of Small Language Models to Fine-tuning Data Contamination
- Vocabulary In-Context Learning in Transformers: Benefits of Positional Encoding
- Attention and Compression is all you need for Controllably Efficient Language Models
- It Takes Two: A Dual Stage Approach for Terminology-Aware Translation
- NAPS: Attention-Based Fusion of Heterogeneous Physiological Signals
- Neural Network Interoperability Across Platforms
- Vibe Learning: Education in the age of AI
- A Dual-Use Framework for Clinical Gait Analysis: Attention-Based Sensor Optimization and Automated Dataset Auditing
- TransAlign: Machine Translation Encoders are Strong Word Aligners, Too
- Hybrid Quantum-Classical Recurrent Neural Networks
- Agentic Economic Modeling
- Transformer Atomic Cluster Expansion: TRACE
- Improving the sample-efficiency of neural architecture search with reinforcement learning
- A neural attention model for speech command recognition
- Utilizing AI questionnaire translations in cross-cultural and intercultural research: Insights and recommendations
- Acoustic source localization by deep-learning attention-based modulation of microphone array data
- ProGen2: Exploring the boundaries of protein language models
- Deep Graph Memory Networks for Forgetting-Robust Knowledge Tracing
- Personalized Route Recommendation With Neural Network Enhanced Search Algorithm
- Lightweight Transformer for EEG Classification via Balanced Signed Graph Algorithm Unrolling
- Sequential Routing Framework: Fully Capsule Network-based Speech Recognition
- Benchmarking DNA large language models on quadruplexes
- Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling
- A Survey of Model Compression and Acceleration for Deep Neural Networks
- OpenNMT: Neural Machine Translation Toolkit
- BadNets: Identifying Vulnerabilities in the Machine Learning Model Supply Chain
- ROULETTE: A neural attention multi-output model for explainable Network Intrusion Detection
- Recurrent Neural Networks as Weighted Language Recognizers
- SeqGAN: Sequence Generative Adversarial Nets with Policy Gradient
- Sequential Recommendation with Graph Neural Networks
- Talking-Heads Attention
- Neural Melody Composition from Lyrics
- Image Captioning with Semantic Attention
- simNet: Stepwise Image-Topic Merging Network for Generating Detailed and Comprehensive Image Captions
- Distance-based Self-Attention Network for Natural Language Inference
- HRM-Text: Efficient Pretraining Beyond Scaling
- GRAM: Graph-based Attention Model for Healthcare Representation Learning
- SkipFlow: Incorporating Neural Coherence Features for End-to-End Automatic Text Scoring
- Quantifying Concentration Phenomena of Mean-Field Transformers in the Low-Temperature Regime
- A Syntactic Neural Model for General-Purpose Code Generation
- Augmenting Librispeech with French Translations: A Multimodal Corpus for Direct Speech Translation Evaluation
- Stochastic Dynamics for Video Infilling
- SGM: Sequence Generation Model for Multi-label Classification
- A Joint Model for Question Answering and Question Generation
- Sequential Variational Autoencoders for Collaborative Filtering
- Improving Neural Machine Translation with Pre-trained Representation
- Word Embeddings via Tensor Factorization
- MemEIC: A Step Toward Continual and Compositional Knowledge Editing
- Dual-Domain Constraints: Designing Covert and Efficient Adversarial Examples for Secure Communication
- An Empirical Study on End-to-End Singing Voice Synthesis with Encoder-Decoder Architectures
- Large Language Models, Agency, and Why Speech Acts are Beyond Them (For Now) – A Kantian-Cum-Pragmatist Case
- Education Paradigm Shift To Maintain Human Competitive Advantage Over AI
- Normalization in Attention Dynamics
- Flight Delay Prediction via Cross-Modality Adaptation of Large Language Models and Aircraft Trajectory Representation
- Towards Neural Decompilation
- Tibetan Language and AI: A Comprehensive Survey of Resources, Methods and Challenges
- SemiAdapt and SemiLoRA: Efficient Domain Adaptation for Transformer-based Low-Resource Language Translation with a Case Study on Irish
- Hyperbolic Attention Networks
- Lingua Custodi's participation at the WMT 2025 Terminology shared task
- Transformer-Based Low-Resource Language Translation: A Study on Standard Bengali to Sylheti
- Renaissance of RNNs in Streaming Clinical Time Series: Compact Recurrence Remains Competitive with Transformers
- Resolution-Aware Retrieval Augmented Zero-Shot Forecasting
- Exact-K Recommendation via Maximal Clique Optimization
- Extending Audio Context for Long-Form Understanding in Large Audio-Language Models
- Unsupervised Cyberbullying Detection via Time-Informed Gaussian Mixture Model
- Predicting the Unpredictable: Reproducible BiLSTM Forecasting of Incident Counts in the Global Terrorism Database (GTD)
- Watermarking for Factuality: Guiding Vision-Language Models Toward Truth via Tri-layer Contrastive Decoding
- MatchAttention: Matching the Relative Positions for High-Resolution Cross-View Matching
- A Survey on Explainable Artificial Intelligence (XAI): Toward Medical XAI
- Breadcrumbs Reasoning: Memory-Efficient Reasoning with Compression Beacons
- Simplicial Embeddings Improve Sample Efficiency in Actor-Critic Agents
- CMIS-Net: A Cascaded Multi-Scale Individual Standardization Network for Backchannel Agreement Estimation
- Multi-Agent Design Assistant for the Simulation of Inertial Fusion Energy
- Pyramid Stereo Matching Network
- Analysis of Bag-of-n-grams Representation's Properties Based on Textual Reconstruction
- Attribution Graphs and Causal Probing for Mechanistic Discovery and Bias Repair in Multimodal Generative Learning
- LLM Reasoning for Machine Translation: Synthetic Data Generation over Thinking Tokens
- Generating an Overview Report over Many Documents
- Fast Decoding in Sequence Models using Discrete Latent Variables
- GenCNER: A Generative Framework for Continual Named Entity Recognition
- Modeling Programs Hierarchically with Stack-Augmented LSTM
- Structured Spectral Graph Representation Learning for Multi-label Abnormality Analysis from 3D CT Scans
- Grounding Referring Expressions in Images by Variational Context
- Multimodal Attention for Neural Machine Translation
- From Explainability to Action: A Generative Operational Framework for Integrating XAI in Clinical Mental Health Screening
- Multigrid Neural Memory
- Approximating meta-heuristics with homotopic recurrent neural networks
- LLM Based Long Code Translation using Identifier Replacement
- TinySpeech: Attention Condensers for Deep Speech Recognition Neural Networks on Edge Devices
- Assessing the Helpfulness of Review Content for Explaining Recommendations
- Hamming OCR: A Locality Sensitive Hashing Neural Network for Scene Text Recognition
- Artificial Hippocampus Networks for Efficient Long-Context Modeling
- A Cluster Ranking Model for Full Anaphora Resolution
- RheOFormer: A generative transformer model for simulation of complex fluids and flows
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- HoRA: Cross-Head Low-Rank Adaptation with Joint Hypernetworks
- Exact Causal Attention with 10% Fewer Operations
- RACE Attention: A Strictly Linear-Time Attention Layer for Training on Outrageously Large Contexts
- Domain-Adapted Granger Causality for Real-Time Cross-Slice Attack Attribution in 6G Networks
- Understanding Neural Networks through Representation Erasure
- Pose-conditioned Spatio-Temporal Attention for Human Action Recognition
- Avoiding Latent Variable Collapse With Generative Skip Models
- The Transformer Cookbook
- FAME: Adaptive Functional Attention with Expert Routing for Function-on-Function Regression
- Time-series forecasting with deep learning: a survey
- DescribeEarth: Describe Anything for Remote Sensing Images
- A multiscale analysis of mean-field transformers in the moderate interaction regime
- Bundle Network: a Machine Learning-Based Bundle Method
- Metamorphic Testing for Audio Content Moderation Software
- AI Methods for Antimicrobial Peptides: Progress and Challenges
- A multi-label classification method using a hierarchical and transparent representation for paper-reviewer recommendation
- Large-Scale Constraint Generation -- Can LLMs Parse Hundreds of Constraints?
- Memory-Efficient Fine-Tuning via Low-Rank Activation Compression
- Generative power of a protein language model trained on multiple sequence alignments
- A model of errors in transformers
- Bridging Fairness and Explainability: Can Input-Based Explanations Promote Fairness in Hate Speech Detection?
- Extracting Parallel Sentences with Bidirectional Recurrent Neural Networks to Improve Machine Translation
- Extrapolating Phase-Field Simulations in Space and Time with Purely Convolutional Architectures
- Adaptive Event-Triggered Policy Gradient for Multi-Agent Reinforcement Learning
- Deep Learning for Clouds and Cloud Shadow Segmentation in Methane Satellite and Airborne Imaging Spectroscopy
- Mamba Modulation: On the Length Generalization of Mamba
- A State-of-the-art Survey of Artificial Neural Networks for Whole-slide Image Analysis:from Popular Convolutional Neural Networks to Potential Visual Transformers
- Character-Level Neural Translation for Multilingual Media Monitoring in the SUMMA Project
- A short trajectory is all you need: A transformer-based model for long-time dissipative quantum dynamics
- Causal Discovery with Inverted Self-attention for Multivariate Time Series
- Why language clouds our ascription of understanding, intention and consciousness
- Disruption in the Chinese E-Commerce During COVID-19
- An Exploration of Word Embedding Initialization in Deep-Learning Tasks
- MSVD-Turkish: A Comprehensive Multimodal Dataset for Integrated Vision and Language Research in Turkish
- LSTMVis: A Tool for Visual Analysis of Hidden State Dynamics in\n Recurrent Neural Networks
- Limitations of Normalization in Attention Mechanism
- Cascaded Text Generation with Markov Transformers
- Deep Learning in Multiple Multistep Time Series Prediction
- Drug-Drug Interaction Extraction from Biomedical Text Using Long Short Term Memory Network
- Climate-Adaptive and Cascade-Constrained Machine Learning Prediction for Sea Surface Height under Greenhouse Warming
- CCQA: Generating Question from Solution Can Improve Inference-Time Reasoning in SLMs
- Multi Task Deep Morphological Analyzer: Context Aware Joint Morphological Tagging and Lemma Prediction
- MSGAT-GRU: A Multi-Scale Graph Attention and Recurrent Model for Spatiotemporal Road Accident Prediction
- Modeling Human Motion with Quaternion-Based Neural Networks
- Cross-Attention is Half Explanation in Speech-to-Text Models
- Specification-Aware Machine Translation and Evaluation for Purpose Alignment
- Time Series Forecasting Using a Hybrid Deep Learning Method: A Bi-LSTM Embedding Denoising Auto Encoder Transformer
- SeaPearl: A Constraint Programming Solver guided by Reinforcement\n Learning
- Navigational Instruction Generation as Inverse Reinforcement Learning\n with Neural Machine Translation
- TANDEM: Temporal Attention-guided Neural Differential Equations for Missingness in Time Series Classification
- EdgeRec: Recommender System on Edge in Mobile Taobao
- Generative AI Meets Wireless Sensing: Towards Wireless Foundation Model
- Evaluating the Impact of Verbal Multiword Expressions on Machine Translation
- DyWPE: Signal-Aware Dynamic Wavelet Positional Encoding for Time Series Transformers
- Compositional Vector Space Models for Knowledge Base Completion
- SSL-SSAW: Self-Supervised Learning with Sigmoid Self-Attention Weighting for Question-Based Sign Language Translation
- ZTFed-MAS2S: A Zero-Trust Federated Learning Framework with Verifiable Privacy and Trust-Aware Aggregation for Wind Power Data Imputation
- Relational Collaborative Filtering:Modeling Multiple Item Relations for Recommendation
- Learning Effective Representations for Person-Job Fit by Feature Fusion
- Explainable Unsupervised Multi-Anomaly Detection and Temporal Localization in Nuclear Times Series Data with a Dual Attention-Based Autoencoder
- Integrating Attention-Enhanced LSTM and Particle Swarm Optimization for Dynamic Pricing and Replenishment Strategies in Fresh Food Supermarkets
- Dynamic Relational Priming Improves Transformer in Multivariate Time Series
- Artificial Neural Networks for Neuroscientists: A Primer
- Unrolling Graph-based Douglas-Rachford Algorithm for Image Interpolation with Informed Initialization
- Compositional Generalization by Learning Analytical Expressions
- Literaturwissenschaft und Informatik
- Deriving the Scaled-Dot-Function via Maximum Likelihood Estimation and Maximum Entropy Approach
- Quantum Graph Attention Networks: Trainable Quantum Encoders for Inductive Graph Learning
- Towards a Neural Network Approach to Abstractive Multi-Document Summarization
- Optimal message passing for molecular prediction is simple, attentive and spatial
- Hierarchical Text Generation and Planning for Strategic Dialogue
- Predicting void nucleation in microstructure with convolutional neural networks
- Unsupervised Word Segmentation from Speech with Attention
- LLM Architecture, Scaling Laws, and Economics: A Quick Summary
- Streaming Sequence-to-Sequence Learning with Delayed Streams Modeling
- Improved Receiver Chain Performance via Error Location Inference
- Customizing the Inductive Biases of Softmax Attention using Structured Matrices
- Exploiting Multi-domain Visual Information for Fake News Detection
- The Natural Language Decathlon: Multitask Learning as Question Answering
- Can Neural Networks Understand Logical Entailment?
- Geometric Dynamics of Consumer Credit Cycles: A Multivector-based Linear-Attention Framework for Explanatory Economic Analysis
- Towards Abstraction from Extraction: Multiple Timescale Gated Recurrent Unit for Summarization
- Frame-Semantic Parsing with Softmax-Margin Segmental RNNs and a Syntactic Scaffold
- Learning to Extract Coherent Summary via Deep Reinforcement Learning
- Video-Based MPAA Rating Prediction: An Attention-Driven Hybrid Architecture Using Contrastive Learning
- Neural Turing Machines
- Learning to Teach with Dynamic Loss Functions
- CLARA: Clinical Report Auto-completion
- Causal Multi-fidelity Surrogate Forward and Inverse Models for ICF Implosions
- Recomposer: Event-roll-guided generative audio editing
- Hunyuan-MT Technical Report
- Attend to the beginning: A study on using bidirectional attention for\n extractive summarization
- Domain-Constrained Advertising Keyword Generation
- Handwriting Imagery EEG Classification based on Convolutional Neural Networks
- Informer: Beyond Efficient Transformer for Long Sequence Time-Series Forecasting
- Stability-Aware Joint Communication and Control for Nonlinear Control-Non-Affine Wireless Networked Control Systems
- Whisper based Cross-Lingual Phoneme Recognition between Vietnamese and English
- A human-inspired recognition system for premodern Japanese historical documents
- Learning Longitudinal Stress Dynamics from Irregular Self-Reports via Time Embeddings
- ArabEmoNet: A Lightweight Hybrid 2D CNN-BiLSTM Model with Attention for Robust Arabic Speech Emotion Recognition
- Exploiting Unlabeled Data for Neural Grammatical Error Detection
- REVELIO -- Universal Multimodal Task Load Estimation for Cross-Domain Generalization
- Missing Data Imputation using Neural Cellular Automata
- Aspect Based Sentiment Analysis with Gated Convolutional Networks
- Beyond Individual Input for Deep Anomaly Detection on Tabular Data
- LegalChainReasoner: A Legal Chain-guided Framework for Criminal Judicial Opinion Generation
- Why Stop at Words? Unveiling the Bigger Picture through Line-Level OCR
- Adaptive Computation Time for Recurrent Neural Networks
- Axiomatic Attribution for Deep Networks
- Bridging Minds and Machines: Toward an Integration of AI and Cognitive Science
- Robust Scene Text Recognition with Automatic Rectification
- Dual Attention Networks for Multimodal Reasoning and Matching
- Bangla-Bayanno: A 52K-Pair Bengali Visual Question Answering Dataset with LLM-Assisted Translation Refinement
- Hybrid Decoding: Rapid Pass and Selective Detailed Correction for Sequence Models
- On Surjectivity of Neural Networks: Can you elicit any behavior from your model?
- Quantum-Classical Hybrid Molecular Autoencoder for Advancing Classical Decoding
- Attention-Based Explainability for Structure-Property Relationships
- Mini-Batch Robustness Verification of Deep Neural Networks
- Dial2Desc: End-to-end Dialogue Description Generation
- Constraint Matters: Multi-Modal Representation for Reducing Mixed-Integer Linear programming
- A New NMT Model for Translating Clinical Texts from English to Spanish
- Integral Transformer: Denoising Attention, Not Too Much Not Too Little
- Psychlab: A Psychology Laboratory for Deep Reinforcement Learning Agents
- PCR-CA: Parallel Codebook Representations with Contrastive Alignment for Multiple-Category App Recommendation
- Emergence of Compositional Language with Deep Generational Transmission
- Generative Model-Based Feature Attention Module for Video Action Analysis
- A fully-programmable integrated photonic processor for both domain-specific and general-purpose computing
- A Neural Conversational Model
- A Three Step Training Approach with Data Augmentation for Morphological Inflection
- AttnGAN: Fine-Grained Text to Image Generation with Attentional Generative Adversarial Networks
- Recognizing Handwritten Mathematical Expressions as LaTex Sequences Using a Multiscale Robust Neural Network
- OPTIC-ER: A Reinforcement Learning Framework for Real-Time Emergency Response and Equitable Resource Allocation in Underserved African Communities
- Incremental processing of noisy user utterances in the spoken language\n understanding task
- Unsupervised Predictive Memory in a Goal-Directed Agent
- Itinerary-aware Personalized Deep Matching at Fliggy
- Handwritten Text Recognition of Historical Manuscripts Using Transformer-Based Models
- Reference Points in LLM Sentiment Analysis: The Role of Structured Context
- Master Thesis: Neural Sign Language Translation by Learning Tokenization
- Overcoming Low-Resource Barriers in Tulu: Neural Models and Corpus Creation for OffensiveLanguage Identification
- Are AI Machines Making Humans Obsolete?
- Privacy-Aware Detection of Fake Identity Documents: Methodology, Benchmark, and Improved Algorithms (FakeIDet2)
- A Transformer-Based Approach for DDoS Attack Detection in IoT Networks
- A Self-Attention Network for Hierarchical Data Structures with an\n Application to Claims Management
- Imperial College London Submission to VATEX Video Captioning Task
- DepressLLM: Interpretable domain-adapted language model for depression detection from real-world narratives
- VulTriNet: A software vulnerability detection method based on tri-channel network
- Hierarchical Recurrent Neural Encoder for Video Representation with Application to Captioning
- Image Captioning at Will: A Versatile Scheme for Effectively Injecting Sentiments into Image Descriptions
- Multi-task Adversarial Attacks against Black-box Model with Few-shot Queries
- Bag-of-Vector Embeddings of Dependency Graphs for Semantic Induction
- SiMiC: Context-aware silicon microstructure characterization using attention-based convolutional neural networks for field-emission tip analysis
- Audio Does Matter: Importance-Aware Multi-Granularity Fusion for Video Moment Retrieval
- Excavate the potential of Single-Scale Features: A Decomposition Network for Water-Related Optical Image Enhancement
- CORE-ReID V2: Advancing the Domain Adaptation for Object Re-Identification with Optimized Training and Ensemble Fusion
- User Perception of Attention Visualizations: Effects on Interpretability Across Evidence-Based Medical Documents
- Tricks and Plug-ins for Gradient Boosting with Transformers
- Residual Switching Network for Portfolio Optimization
- A Machine Learning Approach to Routing
- Folksonomication: Predicting Tags for Movies from Plot Synopses Using Emotion Flow Encoded Neural Network
- Acquiring Knowledge from Pre-trained Model to Neural Machine Translation
- A Semantic Relevance Based Neural Network for Text Summarization and Text Simplification
- Speech2Vec: A Sequence-to-Sequence Framework for Learning Word Embeddings from Speech
- Object-driven Text-to-Image Synthesis via Adversarial Training
- SequenceLayers: Sequence Processing and Streaming Neural Networks Made Easy
- Are Transformers Effective for Time Series Forecasting?
- On the Expressiveness of Softmax Attention: A Recurrent Neural Network Perspective
- Learning When to Concentrate or Divert Attention: Self-Adaptive Attention Temperature for Neural Machine Translation
- Type-driven Neural Programming by Example
- Efficient Spatial-Temporal Modeling for Real-Time Video Analysis: A Unified Framework for Action Recognition and Object Tracking
- AttnMove: History Enhanced Trajectory Recovery via Attentional Network
- Can large language models assist choice modelling? Insights into prompting strategies and current models capabilities
- Entropy-Enhanced Multimodal Attention Model for Scene-Aware Dialogue Generation
- Memorization in Fine-Tuned Large Language Models
- Order Matters: Sequence to sequence for sets
- A copula-based visualization technique for a neural network
- Why Pay More When You Can Pay Less: A Joint Learning Framework for Active Feature Acquisition and Classification
- Advancing Dialectal Arabic to Modern Standard Arabic Machine Translation
- Comparative Visual Analytics for Assessing Medical Records with Sequence Embedding
- EcoTransformer: Attention without Multiplication
- Stealing Links from Graph Neural Networks
- NERO: A Neural Rule Grounding Framework for Label-Efficient Relation Extraction
- Transfer Learning for Clinical Time Series Analysis using Recurrent\n Neural Networks
- Chunk-Based Bi-Scale Decoder for Neural Machine Translation
- Endowing Deep 3D Models with Rotation Invariance Based on Principal Component Analysis
- Grounded Conversation Generation as Guided Traverses in Commonsense Knowledge Graphs
- Patch Pruning Strategy Based on Robust Statistical Measures of Attention Weight Diversity in Vision Transformers
- GIIFT: Graph-guided Inductive Image-free Multimodal Machine Translation
- Restoring Rhythm: Punctuation Restoration Using Transformer Models for Bangla, a Low-Resource Language
- Micromobility Flow Prediction: A Bike Sharing Station-level Study via Multi-level Spatial-Temporal Attention Neural Network
- DCFFSNet: Deep Connectivity Feature Fusion Separation Network for Medical Image Segmentation
- Adapting the Neural Encoder-Decoder Framework from Single to Multi-Document Summarization
- Understanding Hidden Memories of Recurrent Neural Networks
- Deep Recurrent Neural Network for Protein Function Prediction from Sequence
- Higher Order Recurrent Neural Networks
- The Origin of Self-Attention: Pairwise Affinity Matrices in Feature Selection and the Emergence of Self-Attention
- Machine Translation between Vietnamese and English: an Empirical Study
- Recurrent Attention Unit
- Grading video interviews with fairness considerations
- The Costs and Benefits of Goal-Directed Attention in Deep Convolutional Neural Networks
- Neural Machine Translation with Monolingual Translation Memory
- Guiding Neural Machine Translation with Retrieved Translation Pieces
- On Inductive Biases for Machine Learning in Data Constrained Settings
- A Data-Efficient Deep Learning Based Smartphone Application For Detection Of Pulmonary Diseases Using Chest X-rays
- Doubly Attentive Transformer Machine Translation
- PARK: Personalized academic retrieval with knowledge-graphs
- The Accidental Pump and Dump: When Agentic AI Meets Autonomous Trading
- Generative AI in Qualitative Research and Related Transparency Problems: A Novel Heuristic for Disclosing Uses of AI
- Modeling Attention Flow on Graphs
- Context-Aware Sequence-to-Sequence Models for Conversational Systems
- Transformation Networks for Target-Oriented Sentiment Classification
- Accelerating Asynchronous Stochastic Gradient Descent for Neural Machine\n Translation
- Achieving Robust Channel Estimation Neural Networks by Designed Training Data
- From Time-series Generation, Model Selection to Transfer Learning: A Comparative Review of Pixel-wise Approaches for Large-scale Crop Mapping
- Emerging Properties in Self-Supervised Vision Transformers
- On the Definition of Japanese Word
- Incremental Adaptation of NMT for Professional Post-editors: A User Study
- Seq2seq Translation Model for Sequential Recommendation
- OSU Multimodal Machine Translation System Report
- Translating Natural Language to SQL using Pointer-Generator Networks and How Decoding Order Matters
- Recurrent Neural Networks for Time Series Forecasting: Current status and future directions
- Feedforward Sequential Memory Networks: A New Structure to Learn Long-term Dependency
- Paraphrase Generation as Unsupervised Machine Translation
- Domain-Adaptive Small Language Models for Structured Tax Code Prediction
- Time-series forecasting for nonlinear high-dimensional system using hybrid method combining autoencoder and multi-parallelized quantum long short-term memory and gated recurrent unit
- Understood in Translation, Transformers for Domain Understanding
- Multimodal Dialogue State Tracking By QA Approach with Data Augmentation
- Multi-Head Attention: Collaborate Instead of Concatenate
- Visual Question Answering with Memory-Augmented Networks
- Semantic-Unit-Based Dilated Convolution for Multi-Label Text Classification
- CycleGAN-Driven Transfer Learning for Electronics Response Emulation in High-Purity Germanium Detectors
- Queue up for takeoff: a transferable deep learning framework for flight delay prediction
- Towards Extracting Software Requirements from App Reviews using Seq2seq Framework
- Emergent Natural Language with Communication Games for Improving Image Captioning Capabilities without Additional Data
- Universal Approximation Theorem for a Single-Layer Transformer
- Generative Language-Grounded Policy in Vision-and-Language Navigation with Bayes' Rule
- Probabilistic Transformers
- SAS: Simulated Attention Score
- A multi-task encoder-dual-decoder framework for mixed frequency data prediction
- KPFlow: An Operator Perspective on Dynamic Collapse Under Gradient Descent Training of Recurrent Networks
- PhoniTale: Phonologically Grounded Mnemonic Generation for Typologically Distant Language Pairs
- A Survey of Pun Generation: Datasets, Evaluations and Methodologies
- PRISM: Pointcloud Reintegrated Inference via Segmentation and Cross-attention for Manipulation
- Multi-node Bert-pretraining: Cost-efficient Approach
- ViTaL: A Multimodality Dataset and Benchmark for Multi-pathological Ovarian Tumor Recognition
- Relational inductive biases on attention mechanisms
- Graph Collaborative Attention Network for Link Prediction in Knowledge Graphs
- RECA-PD: A Robust Explainable Cross-Attention Method for Speech-based Parkinson's Disease Classification
- Automatically Generating Commit Messages from Diffs using Neural Machine Translation
- Memory Mosaics at scale
- Linear Attention with Global Context: A Multipole Attention Mechanism for Vision and Physics
- Can Artificial Intelligence solve the blockchain oracle problem? Unpacking the Challenges and Possibilities
- Multi-hop Inference for Question-driven Summarization
- Listen, Attend, and Walk: Neural Mapping of Navigational Instructions to Action Sequences
- A Sketch-Based Neural Model for Generating Commit Messages from Diffs
- A Deep Neural Network for Unsupervised Anomaly Detection and Diagnosis in Multivariate Time Series Data
- A Review on Sound Source Localization in Robotics: Focusing on Deep Learning Methods
- Examining Reject Relations in Stimulus Equivalence Simulations
- Thinking About Thinking: SAGE-nano's Inverse Reasoning for Self-Aware Language Models
- Minimizing the Bag-of-Ngrams Difference for Non-Autoregressive Neural Machine Translation
- Relevance-Promoting Language Model for Short-Text Conversation
- Opportunities and challenges in developing deep learning models using electronic health records data: a systematic review
- Concept Pinpoint Eraser for Text-to-image Diffusion Models via Residual Attention Gate
- Improved and Robust Controversy Detection in General Web Pages Using\n Semantic Approaches under Large Scale Conditions
- Object Based Attention Through Internal Gating
- SU-RUG at the CoNLL-SIGMORPHON 2017 shared task: Morphological\n Inflection with Attentional Sequence-to-Sequence Models
- Offensive Language Detection on Social Media Using XLNet
- GoIRL: Graph-Oriented Inverse Reinforcement Learning for Multimodal Trajectory Prediction
- Communicating Smartly in Molecular Communication Environments: Neural Networks in the Internet of Bio-Nano Things
- Variational Digital Twins
- Program Language Translation Using a Grammar-Driven Tree-to-Tree Model
Discussions
- Me rationalising what I am currently doing: There is no shame in designing a recipish, clunky, deep architecture kept together with duct tape. This is how we move into new territories. The Transformer [bsky, 26 points, 1 comments]
- interesting connection, but the transformer paper didn’t invent attention? arxiv.org/abs/1409.0473 [bsky, 23 points, 1 comments]
- For those of you who don't know (or don't remember), the original paper introducing "attention" was Bahdanau, Cho, & Bengio's "Neural machine translation by jointly learning to align and translate" (2 [bsky, 8 points, 1 comments]
- Neural Machine Translation by Jointly Learning to Align and Translate [hn, 2 points, 0 comments]
- וזה המאמר שתאר לראשונה-Attention עבור רשתות ניורונים: arxiv.org/abs/1409.0473 אבל בתכלס עדיף לקרוא דברים מאוחרים יותר. [bsky, 2 points, 0 comments]
- CDS Prof. @kyunghyuncho.bsky.social's 2014 "attention" paper was recently the Runner-Up for the ICLR 2025 Test of Time Award. The paper introduced dynamic attention in machine translation, laying the [bsky, 2 points, 0 comments]
- We've come full circle [lemmy, 1 points, 0 comments]
- Neural Machine Translation by Jointly Learning to Align and Translate (2015) [hn, 1 points, 0 comments]
- what a nice paper this is arxiv.org/pdf/1409.0473 [bsky, 1 points, 0 comments]
Related