Neural Machine Translation by Jointly Learning to Align and Translate
2014/09/01 by Dzmitry Bahdanau, Kyunghyun Cho, Bahdanau, Dzmitry +3 · 9 voices · 14,622 citations
Computer Science · Mathematics · #Artificial intelligence #Artificial neural network #Bottleneck #Computer science #Encoder #Example-based machine translation #Machine translation #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Natural language processing #Phrase #Sentence #Speech recognition #Topic Modeling #Transfer-based machine translation #Translation (biology) #Word (group theory) #cs.CL #cs.LG #cs.NE #stat.ML
paper · pdf · doi:10.48550/arxiv.1409.0473
published in arXiv (Cornell University) (Cornell University) · Accepted at ICLR 2015 as oral presentation
openalex publication_date 2014/09/01 · arxiv created 2016/05/19 · arxiv updated 2016/05/23 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05
Abstract
Neural machine translation is a recently proposed approach to machine translation. Unlike the traditional statistical machine translation, the neural machine translation aims at building a single neural network that can be jointly tuned to maximize the translation performance. The models proposed recently for neural machine translation often belong to a family of encoder-decoders and consists of an encoder that encodes a source sentence into a fixed-length vector from which a decoder generates a translation. In this paper, we conjecture that the use of a fixed-length vector is a bottleneck in improving the performance of this basic encoder-decoder architecture, and propose to extend this by allowing a model to automatically (soft-)search for parts of a source sentence that are relevant to predicting a target word, without having to form these parts as a hard segment explicitly. With this new approach, we achieve a translation performance comparable to the existing state-of-the-art phrase-based system on the task of English-to-French translation. Furthermore, qualitative analysis reveals that the (soft-)alignments found by the model agree well with our intuition.
Citations
Cited by
- Biomedical Machine Translation for Low-Resource Arabic-Script Languages via Cross-Lingual Transfer and LoRA Adapter Merging
- Pretraining Recurrent Networks without Recurrence
- Breaking the Data Barrier in Learning Symbolic Computation: A Case Study on Variable Ordering Suggestion for Cylindrical Algebraic Decomposition
- Graph-Theoretic Neural Network Fragmentation with Covariant Direct Molecular Force Learning: Enabling Coupled-Cluster Accuracy AIMD for Fluxional Systems
- A Deep Reinforcement Learning Algorithm for the Vehicle Routing Problem with Stochastic Demands and Outsourcing
- Neuro-Symbolic Meta-Policies for Temporal Knowledge-Graph Memory under Partial Observability
- An Explicit World Model Based on Data-First Ontology: DaoQL Multimodal Storage Validation and Counterfactual Reasoning Evaluation
- Jet quenching identification via supervised learning in simulated heavy-ion collisions
- MATNet: multi-level fusion transformer-based model for day-ahead PV generation forecasting
- Attention to Mamba: A Recipe for Cross-Architecture Distillation
- Mamba-3: Improved Sequence Modeling using State Space Principles
- Remapping and navigation of an embedding space via error minimization: a fundamental organizational principle of cognition in natural and artificial systems
- PHOTON: Hierarchical Autoregressive Modeling for Lightspeed and Memory-Efficient Language Generation
- Fast weight programming and linear transformers: from machine learning to neurobiology
- QiMeng: Fully Automated Hardware and Software Design for Processor Chip
- Transformers are Graph Neural Networks
- Blending Complementary Memory Systems in Hybrid Quadratic-Linear Transformers
- ATLAS: Learning to Optimally Memorize the Context at Test Time
- Log-Linear Attention
- NNN: Next-Generation Neural Networks for Marketing Measurement
- Multi-Token Attention
- Chapter-Llama: Efficient Chaptering in Hour-Long Videos with LLMs
- CalibQuant: 1-Bit KV Cache Quantization for Multimodal LLMs
- Learning When Not to Attend Globally
- AudioGAN: A Compact and Efficient Framework for Real-Time High-Fidelity Text-to-Audio Generation
- Attention Residuals
- Towards High-Level Semantic Intelligence
- Selecting Language Models for Social Science: Start Small, Start Open, and Validate
- Understanding Human-like Solutions in Combinatorial Optimization via Learning and Search
- Anchor Attention, Small Cache: Code Generation with Large Language Models
- PLATO: Pointer Learner for Agent and Task Openness
- Optimizing Transformer Neural Network for Real-Time Outlier Detection on FPGAs
- Statistical vs. Deep Learning Models for Estimating Substance Overdose Excess Mortality in the US
- SPOT!: Map-Guided LLM Agent for Unsupervised Multi-CCTV Dynamic Object Tracking
- Explainable time-series forecasting with sampling-free SHAP for Transformers
- BiCoR-Seg: Bidirectional Co-Refinement Framework for High-Resolution Remote Sensing Image Segmentation
- RIS-Enabled Smart Wireless Environments: Fundamentals and Distributed Optimization
- Asynchronous Pipeline Parallelism for Real-Time Multilingual Lip Synchronization in Video Communication Systems
- Adversarial Robustness of Vision in Open Foundation Models
- InfinityEBSD : Metrics-Guided Infinite-Size EBSD Map Generation With Diffusion Models
- KOSS: Kalman-Optimal Selective State Spaces for Long-Term Sequence Modeling
- Evaluating OpenAI GPT Models for Translation of Endangered Uralic Languages: A Comparison of Reasoning and Non-Reasoning Architectures
- SegGraph: Leveraging Graphs of SAM Segments for Few-Shot 3D Part Segmentation
- From Minutes to Days: Scaling Intracranial Speech Decoding with Supervised Pretraining
- An Empirical Study on Chinese Character Decomposition in Multiword Expression-Aware Neural Machine Translation
- Incentives or Ontology? A Structural Rebuttal to OpenAI's Hallucination Thesis
- Advancing Bangla Machine Translation Through Informal Datasets
- Towards Deep Learning Surrogate for the Forward Problem in Electrocardiology: A Scalable Alternative to Physics-Based Models
- WAY: Estimation of Vessel Destination in Worldwide AIS Trajectory
- Behavior and Representation in Large Language Models for Combinatorial Optimization: From Feature Extraction to Algorithm Selection
- HydroDiffusion: Diffusion-Based Probabilistic Streamflow Forecasting with a State Space Backbone
- Kinetic Mining in Context: Few-Shot Action Synthesis via Text-to-Motion Distillation
- Unlocking the Address Book: Dissecting the Sparse Semantic Structure of LLM Key-Value Caches via Sparse Autoencoders
- Rates and architectures for learning geometrically non-trivial operators
- PVeRA: Probabilistic Vector-Based Random Matrix Adaptation
- Towards a Relationship-Aware Transformer for Tabular Data
- QL-LSTM: A Parameter-Efficient LSTM for Stable Long-Sequence Modeling
- Learning the Cosmic Web: Graph-based Classification of Simulated Galaxies by their Dark Matter Environments
- HTR-ConvText: Leveraging Convolution and Textual Information for Handwritten Text Recognition
- Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models
- ProtoEFNet: Dynamic Prototype Learning for Inherently Interpretable Ejection Fraction Estimation in Echocardiography
- Robust Probabilistic Load Forecasting for a Single Household: A Comparative Study from SARIMA to Transformers on the REFIT Dataset
- Interpretable Alzheimer's Diagnosis via Multimodal Fusion of Regional Brain Experts
- BanglaSentNet: An Explainable Hybrid Deep Learning Framework for Multi-Aspect Sentiment Analysis with Cross-Domain Transfer Learning
- "When Data is Scarce, Prompt Smarter"... Approaches to Grammatical Error Correction in Low-Resource Settings
- Hierarchical Spatio-Temporal Attention Network with Adaptive Risk-Aware Decision for Forward Collision Warning in Complex Scenarios
- SimDiff: Simpler Yet Better Diffusion Model for Time Series Point Forecasting
- Large-Scale In-Game Outcome Forecasting for Match, Team and Players in Football using an Axial Transformer Neural Network
- GContextFormer: A global context-aware hybrid multi-head attention approach with scaled additive aggregation for multimodal trajectory prediction
- Foundations of Artificial Intelligence Frameworks: Notion and Limits of AGI
- Attention-Guided Feature Fusion (AGFF) Model for Integrating Statistical and Semantic Features in News Text Classification
- Computation for Epidemic Prediction with Graph Neural Network by Model Combination
- What Really Counts? Examining Step and Token Level Attribution in Multilingual CoT Reasoning
- LAYA: Layer-wise Attention Aggregation for Interpretable Depth-Aware Neural Networks
- EchoLSTM: A Self-Reflective Recurrent Network for Stabilizing Long-Range Memory
- MTMed3D: A Multi-Task Transformer-Based Model for 3D Medical Imaging
- Evaluation of Attention Mechanisms in U-Net Architectures for Semantic Segmentation of Brazilian Rock Art Petroglyphs
- Additive Large Language Models for Semi-Structured Text
- Scaling Open-Weight Large Language Models for Hydropower Regulatory Information Extraction: A Systematic Analysis
- TEDxTN: A Three-way Speech Translation Corpus for Code-Switched Tunisian Arabic - English
- Know Your Limits: Entropy Estimation Modeling for Compression and Generalization
- MVSMamba: Multi-View Stereo with State Space Model
- H-Model: Dynamic Neural Architectures for Adaptive Processing
- A Circular Argument : Does RoPE need to be Equivariant for Vision?
- Reduced Density Matrices Through Machine Learning
- A Picture is Worth a Thousand (Correct) Captions: A Vision-Guided Judge-Corrector System for Multimodal Machine Translation
- Sensitivity of Small Language Models to Fine-tuning Data Contamination
- Vocabulary In-Context Learning in Transformers: Benefits of Positional Encoding
- Controllably Efficient Language Models
- It Takes Two: A Dual Stage Approach for Terminology-Aware Translation
- NAPS: Attention-Based Fusion of Heterogeneous Physiological Signals
- Neural Network Interoperability Across Platforms
- Vibe Learning: Education in the age of AI
- A Dual-Use Framework for Clinical Gait Analysis: Attention-Based Sensor Optimization and Automated Dataset Auditing
- TransAlign: Machine Translation Encoders are Strong Word Aligners, Too
- Hybrid Quantum-Classical Recurrent Neural Networks
- Agentic Economic Modeling
- Transformer Atomic Cluster Expansion: TRACE
- Improving the sample-efficiency of neural architecture search with reinforcement learning
- A neural attention model for speech command recognition
- Utilizing AI questionnaire translations in cross-cultural and intercultural research: Insights and recommendations
- Acoustic source localization by deep-learning attention-based modulation of microphone array data
- ProGen2: Exploring the boundaries of protein language models
- Deep Graph Memory Networks for Forgetting-Robust Knowledge Tracing
- Personalized Route Recommendation With Neural Network Enhanced Search Algorithm
- Lightweight Transformer for EEG Classification via Balanced Signed Graph Algorithm Unrolling
- Sequential Routing Framework: Fully Capsule Network-based Speech Recognition
- Benchmarking DNA large language models on quadruplexes
- Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling
- A Survey of Model Compression and Acceleration for Deep Neural Networks
- OpenNMT: Neural Machine Translation Toolkit
- BadNets: Identifying Vulnerabilities in the Machine Learning Model Supply Chain
- ROULETTE: A neural attention multi-output model for explainable Network Intrusion Detection
- Recurrent Neural Networks as Weighted Language Recognizers
- SeqGAN: Sequence Generative Adversarial Nets with Policy Gradient
- Sequential Recommendation with Graph Neural Networks
- Talking-Heads Attention
- Neural Melody Composition from Lyrics
- Image Captioning with Semantic Attention
- simNet: Stepwise Image-Topic Merging Network for Generating Detailed and Comprehensive Image Captions
- Distance-based Self-Attention Network for Natural Language Inference
- HRM-Text: Efficient Pretraining Beyond Scaling
- GRAM: Graph-based Attention Model for Healthcare Representation Learning
- SkipFlow: Incorporating Neural Coherence Features for End-to-End Automatic Text Scoring
- Quantifying Concentration Phenomena of Mean-Field Transformers in the Low-Temperature Regime
- A Syntactic Neural Model for General-Purpose Code Generation
- Augmenting Librispeech with French Translations: A Multimodal Corpus for Direct Speech Translation Evaluation
- Stochastic Dynamics for Video Infilling
- SGM: Sequence Generation Model for Multi-label Classification
- A Joint Model for Question Answering and Question Generation
- Sequential Variational Autoencoders for Collaborative Filtering
- Improving Neural Machine Translation with Pre-trained Representation
- Word Embeddings via Tensor Factorization
- MemEIC: A Step Toward Continual and Compositional Knowledge Editing
- Dual-Domain Constraints: Designing Covert and Efficient Adversarial Examples for Secure Communication
- An Empirical Study on End-to-End Singing Voice Synthesis with Encoder-Decoder Architectures
- Large Language Models, Agency, and Why Speech Acts are Beyond Them (For Now) – A Kantian-Cum-Pragmatist Case
- Education Paradigm Shift To Maintain Human Competitive Advantage Over AI
- Normalization in Attention Dynamics
- Flight Delay Prediction via Cross-Modality Adaptation of Large Language Models and Aircraft Trajectory Representation
- Towards Neural Decompilation
- Tibetan Language and AI: A Comprehensive Survey of Resources, Methods and Challenges
- SemiAdapt and SemiLoRA: Efficient Domain Adaptation for Transformer-based Low-Resource Language Translation with a Case Study on Irish
- Hyperbolic Attention Networks
- Lingua Custodi's participation at the WMT 2025 Terminology shared task
- Transformer-Based Low-Resource Language Translation: A Study on Standard Bengali to Sylheti
- Renaissance of RNNs in Streaming Clinical Time Series: Compact Recurrence Remains Competitive with Transformers
- Resolution-Aware Retrieval Augmented Zero-Shot Forecasting
- Exact-K Recommendation via Maximal Clique Optimization
- Extending Audio Context for Long-Form Understanding in Large Audio-Language Models
- Unsupervised Cyberbullying Detection via Time-Informed Gaussian Mixture Model
- Predicting the Unpredictable: Reproducible BiLSTM Forecasting of Incident Counts in the Global Terrorism Database (GTD)
- Watermarking for Factuality: Guiding Vision-Language Models Toward Truth via Tri-layer Contrastive Decoding
- MatchAttention: Matching the Relative Positions for High-Resolution Cross-View Matching
- A Survey on Explainable Artificial Intelligence (XAI): Towards Medical XAI
- Breadcrumbs Reasoning: Memory-Efficient Reasoning with Compression Beacons
- Simplicial Embeddings Improve Sample Efficiency in Actor-Critic Agents
- CMIS-Net: A Cascaded Multi-Scale Individual Standardization Network for Backchannel Agreement Estimation
- Multi-Agent Design Assistant for the Simulation of Inertial Fusion Energy
- Pyramid Stereo Matching Network
- Analysis of Bag-of-n-grams Representation's Properties Based on Textual Reconstruction
- Attribution Graphs and Causal Probing for Mechanistic Discovery and Bias Repair in Multimodal Generative Learning
- LLM Reasoning for Machine Translation: Synthetic Data Generation over Thinking Tokens
- Generating an Overview Report over Many Documents
- Fast Decoding in Sequence Models using Discrete Latent Variables
- GenCNER: A Generative Framework for Continual Named Entity Recognition
- Modeling Programs Hierarchically with Stack-Augmented LSTM
- Structured Spectral Graph Representation Learning for Multi-label Abnormality Analysis from 3D CT Scans
- Grounding Referring Expressions in Images by Variational Context
- Multimodal Attention for Neural Machine Translation
- From Explainability to Action: A Generative Operational Framework for Integrating XAI in Clinical Mental Health Screening
- Multigrid Neural Memory
- Approximating meta-heuristics with homotopic recurrent neural networks
- LLM Based Long Code Translation using Identifier Replacement
- TinySpeech: Attention Condensers for Deep Speech Recognition Neural Networks on Edge Devices
- Assessing the Helpfulness of Review Content for Explaining Recommendations
- Hamming OCR: A Locality Sensitive Hashing Neural Network for Scene Text Recognition
- Artificial Hippocampus Networks for Efficient Long-Context Modeling
- A Cluster Ranking Model for Full Anaphora Resolution
- RheOFormer: A generative transformer model for simulation of complex fluids and flows
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Hypernetwork-Driven Low-Rank Adaptation Across Attention Heads
- Exact Causal Attention with 10% Fewer Operations
- RACE Attention: A Strictly Linear-Time Attention Layer for Training on Outrageously Large Contexts
- Domain-Adapted Granger Causality for Real-Time Cross-Slice Attack Attribution in 6G Networks
- Understanding Neural Networks through Representation Erasure
- Pose-conditioned Spatio-Temporal Attention for Human Action Recognition
- Avoiding Latent Variable Collapse With Generative Skip Models
- The Transformer Cookbook
- FAME: Adaptive Functional Attention with Expert Routing for Function-on-Function Regression
- Time Series Forecasting With Deep Learning: A Survey
- DescribeEarth: Describe Anything for Remote Sensing Images
- A multiscale analysis of mean-field transformers in the moderate interaction regime
- Bundle Network: a Machine Learning-Based Bundle Method
- Metamorphic Testing for Audio Content Moderation Software
- AI Methods for Antimicrobial Peptides: Progress and Challenges
- A multi-label classification method using a hierarchical and transparent representation for paper-reviewer recommendation
- Large-Scale Constraint Generation -- Can LLMs Parse Hundreds of Constraints?
- Memory-Efficient Fine-Tuning via Low-Rank Activation Compression
- Generative power of a protein language model trained on multiple sequence alignments
- A model of errors in transformers
- Bridging Fairness and Explainability: Can Input-Based Explanations Promote Fairness in Hate Speech Detection?
- Extracting Parallel Sentences with Bidirectional Recurrent Neural Networks to Improve Machine Translation
- Extrapolating Phase-Field Simulations in Space and Time with Purely Convolutional Architectures
- Adaptive Event-Triggered Policy Gradient for Multi-Agent Reinforcement Learning
- Deep Learning for Clouds and Cloud Shadow Segmentation in Methane Satellite and Airborne Imaging Spectroscopy
- Mamba Modulation: On the Length Generalization of Mamba
- A State-of-the-art Survey of Artificial Neural Networks for Whole-slide Image Analysis:from Popular Convolutional Neural Networks to Potential Visual Transformers
- Character-Level Neural Translation for Multilingual Media Monitoring in the SUMMA Project
- A short trajectory is all you need: A transformer-based model for long-time dissipative quantum dynamics
- Causal Discovery with Inverted Self-attention for Multivariate Time Series
- Why language clouds our ascription of understanding, intention and consciousness
- Disruption in the Chinese E-Commerce During COVID-19
- An Exploration of Word Embedding Initialization in Deep-Learning Tasks
- MSVD-Turkish: A Comprehensive Multimodal Dataset for Integrated Vision and Language Research in Turkish
- LSTMVis: A Tool for Visual Analysis of Hidden State Dynamics in Recurrent Neural Networks
- Limitations of Normalization in Attention Mechanism
- Cascaded Text Generation with Markov Transformers
- Deep Learning in Multiple Multistep Time Series Prediction
- Drug-Drug Interaction Extraction from Biomedical Text Using Long Short Term Memory Network
- Climate-Adaptive and Cascade-Constrained Machine Learning Prediction for Sea Surface Height under Greenhouse Warming
- CCQA: Generating Question from Solution Can Improve Inference-Time Reasoning in SLMs
- Multi Task Deep Morphological Analyzer: Context Aware Joint Morphological Tagging and Lemma Prediction
- MSGAT-GRU: A Multi-Scale Graph Attention and Recurrent Model for Spatiotemporal Road Accident Prediction
- Modeling Human Motion with Quaternion-based Neural Networks
- Cross-Attention is Half Explanation in Speech-to-Text Models
- Specification-Aware Machine Translation and Evaluation for Purpose Alignment
- Time Series Forecasting Using a Hybrid Deep Learning Method: A Bi-LSTM Embedding Denoising Auto Encoder Transformer
- SeaPearl: A Constraint Programming Solver guided by Reinforcement Learning
- Navigational Instruction Generation as Inverse Reinforcement Learning with Neural Machine Translation
- TANDEM: Temporal Attention-guided Neural Differential Equations for Missingness in Time Series Classification
- EdgeRec: Recommender System on Edge in Mobile Taobao
- Generative AI Meets Wireless Sensing: Towards Wireless Foundation Model
- Evaluating the Impact of Verbal Multiword Expressions on Machine Translation
- DyWPE: Signal-Aware Dynamic Wavelet Positional Encoding for Time Series Transformers
- Compositional Vector Space Models for Knowledge Base Completion
- SSL-SSAW: Self-Supervised Learning with Sigmoid Self-Attention Weighting for Question-Based Sign Language Translation
- ZTFed-MAS2S: A Zero-Trust Federated Learning Framework with Verifiable Privacy and Trust-Aware Aggregation for Wind Power Data Imputation
- Relational Collaborative Filtering:Modeling Multiple Item Relations for Recommendation
- Learning Effective Representations for Person-Job Fit by Feature Fusion
- Explainable Unsupervised Multi-Anomaly Detection and Temporal Localization in Nuclear Times Series Data with a Dual Attention-Based Autoencoder
- Integrating Attention-Enhanced LSTM and Particle Swarm Optimization for Dynamic Pricing and Replenishment Strategies in Fresh Food Supermarkets
- Dynamic Relational Priming Improves Transformer in Multivariate Time Series
- Artificial neural networks for neuroscientists: A primer
- Unrolling Graph-based Douglas-Rachford Algorithm for Image Interpolation with Informed Initialization
- Compositional Generalization by Learning Analytical Expressions
- Literaturwissenschaft und Informatik
- Deriving the Scaled-Dot-Function via Maximum Likelihood Estimation and Maximum Entropy Approach
- Quantum Graph Attention Networks: Trainable Quantum Encoders for Inductive Graph Learning
- Towards a Neural Network Approach to Abstractive Multi-Document Summarization
- Optimal message passing for molecular prediction is simple, attentive and spatial
- Hierarchical Text Generation and Planning for Strategic Dialogue
- Predicting void nucleation in microstructure with convolutional neural networks
- Unsupervised Word Segmentation from Speech with Attention
- LLM Architecture, Scaling Laws, and Economics: A Quick Summary
- Streaming Sequence-to-Sequence Learning with Delayed Streams Modeling
- Improved Receiver Chain Performance via Error Location Inference
- Customizing the Inductive Biases of Softmax Attention using Structured Matrices
- Exploiting Multi-domain Visual Information for Fake News Detection
- The Natural Language Decathlon: Multitask Learning as Question Answering
- Can Neural Networks Understand Logical Entailment?
- Geometric Dynamics of Consumer Credit Cycles: A Multivector-based Linear-Attention Framework for Explanatory Economic Analysis
- Towards Abstraction from Extraction: Multiple Timescale Gated Recurrent Unit for Summarization
- Frame-Semantic Parsing with Softmax-Margin Segmental RNNs and a Syntactic Scaffold
- Learning to Extract Coherent Summary via Deep Reinforcement Learning
- Video-Based MPAA Rating Prediction: An Attention-Driven Hybrid Architecture Using Contrastive Learning
- Neural Turing Machines
- Learning to Teach with Dynamic Loss Functions
- CLARA: Clinical Report Auto-completion
- Causal Multi-fidelity Surrogate Forward and Inverse Models for ICF Implosions
- Recomposer: Event-roll-guided generative audio editing
- Hunyuan-MT Technical Report
- Attend to the beginning: A study on using bidirectional attention for extractive summarization
- Domain-Constrained Advertising Keyword Generation
- Handwriting Imagery EEG Classification based on Convolutional Neural Networks
- Informer: Beyond Efficient Transformer for Long Sequence Time-Series Forecasting
- Stability-Aware Joint Communication and Control for Nonlinear Control-Non-Affine Wireless Networked Control Systems
- Whisper based Cross-Lingual Phoneme Recognition between Vietnamese and English
- A human-inspired recognition system for premodern Japanese historical documents
- Learning Longitudinal Stress Dynamics from Irregular Self-Reports via Time Embeddings
- ArabEmoNet: A Lightweight Hybrid 2D CNN-BiLSTM Model with Attention for Robust Arabic Speech Emotion Recognition
- Exploiting Unlabeled Data for Neural Grammatical Error Detection
- REVELIO -- Universal Multimodal Task Load Estimation for Cross-Domain Generalization
- Missing Data Imputation using Neural Cellular Automata
- Aspect Based Sentiment Analysis with Gated Convolutional Networks
- Beyond Individual Input for Deep Anomaly Detection on Tabular Data
- LegalChainReasoner: A Legal Chain-guided Framework for Criminal Judicial Opinion Generation
- Why Stop at Words? Unveiling the Bigger Picture through Line-Level OCR
- Survival Analysis as Imprecise Classification with Trainable Kernels
- Adaptive Computation Time for Recurrent Neural Networks
- Axiomatic Attribution for Deep Networks
- Bridging Minds and Machines: Toward an Integration of AI and Cognitive Science
- Robust Scene Text Recognition with Automatic Rectification
- Dual Attention Networks for Multimodal Reasoning and Matching
- Bangla-Bayanno: A 52K-Pair Bengali Visual Question Answering Dataset with LLM-Assisted Translation Refinement
- Hybrid Decoding: Rapid Pass and Selective Detailed Correction for Sequence Models
- On Surjectivity of Neural Networks: Can you elicit any behavior from your model?
- Quantum-Classical Hybrid Molecular Autoencoder for Advancing Classical Decoding
- Attention-Based Explainability for Structure-Property Relationships
- Mini-Batch Robustness Verification of Deep Neural Networks
- Dial2Desc: End-to-end Dialogue Description Generation
- Improving Generalization Performance by Switching from Adam to SGD
- Constraint Matters: Multi-Modal Representation for Reducing Mixed-Integer Linear programming
- A New NMT Model for Translating Clinical Texts from English to Spanish
- Integral Transformer: Denoising Attention, Not Too Much Not Too Little
- Psychlab: A Psychology Laboratory for Deep Reinforcement Learning Agents
- PCR-CA: Parallel Codebook Representations with Contrastive Alignment for Multiple-Category App Recommendation
- Emergence of Compositional Language with Deep Generational Transmission
- Generative Model-Based Feature Attention Module for Video Action Analysis
- A fully-programmable integrated photonic processor for both domain-specific and general-purpose computing
- A Neural Conversational Model
- A Three Step Training Approach with Data Augmentation for Morphological Inflection
- AttnGAN: Fine-Grained Text to Image Generation with Attentional Generative Adversarial Networks
- Recognizing Handwritten Mathematical Expressions as LaTex Sequences Using a Multiscale Robust Neural Network
- OPTIC-ER: A Reinforcement Learning Framework for Real-Time Emergency Response and Equitable Resource Allocation in Underserved African Communities
- Incremental processing of noisy user utterances in the spoken language understanding task
- Unsupervised Predictive Memory in a Goal-Directed Agent
- Itinerary-aware Personalized Deep Matching at Fliggy
- Handwritten Text Recognition of Historical Manuscripts Using Transformer-Based Models
- Reference Points in LLM Sentiment Analysis: The Role of Structured Context
- Master Thesis: Neural Sign Language Translation by Learning Tokenization
- Overcoming Low-Resource Barriers in Tulu: Neural Models and Corpus Creation for OffensiveLanguage Identification
- Are AI Machines Making Humans Obsolete?
- Privacy-Aware Detection of Fake Identity Documents: Methodology, Benchmark, and Improved Algorithms (FakeIDet2)
- A Transformer-Based Approach for DDoS Attack Detection in IoT Networks
- A Self-Attention Network for Hierarchical Data Structures with an Application to Claims Management
- Imperial College London Submission to VATEX Video Captioning Task
- DepressLLM: Interpretable domain-adapted language model for depression detection from real-world narratives
- VulTriNet: A software vulnerability detection method based on tri-channel network
- Hierarchical Recurrent Neural Encoder for Video Representation with Application to Captioning
- Image Captioning at Will: A Versatile Scheme for Effectively Injecting Sentiments into Image Descriptions
- Multi-task Adversarial Attacks against Black-box Model with Few-shot Queries
- Bag-of-Vector Embeddings of Dependency Graphs for Semantic Induction
- SiMiC: Context-aware silicon microstructure characterization using attention-based convolutional neural networks for field-emission tip analysis
- Audio Does Matter: Importance-Aware Multi-Granularity Fusion for Video Moment Retrieval
- Excavate the potential of Single-Scale Features: A Decomposition Network for Water-Related Optical Image Enhancement
- CORE-ReID V2: Advancing the Domain Adaptation for Object Re-Identification with Optimized Training and Ensemble Fusion
- User Perception of Attention Visualizations: Effects on Interpretability Across Evidence-Based Medical Documents
- Tricks and Plug-ins for Gradient Boosting with Transformers
- Residual Switching Network for Portfolio Optimization
- A Machine Learning Approach to Routing
- Folksonomication: Predicting Tags for Movies from Plot Synopses Using Emotion Flow Encoded Neural Network
- Acquiring Knowledge from Pre-trained Model to Neural Machine Translation
- A Semantic Relevance Based Neural Network for Text Summarization and Text Simplification
- Speech2Vec: A Sequence-to-Sequence Framework for Learning Word Embeddings from Speech
- Object-driven Text-to-Image Synthesis via Adversarial Training
- SequenceLayers: Sequence Processing and Streaming Neural Networks Made Easy
- Are Transformers Effective for Time Series Forecasting?
- On the Expressiveness of Softmax Attention: A Recurrent Neural Network Perspective
- Learning When to Concentrate or Divert Attention: Self-Adaptive Attention Temperature for Neural Machine Translation
- Type-driven Neural Programming by Example
- Efficient Spatial-Temporal Modeling for Real-Time Video Analysis: A Unified Framework for Action Recognition and Object Tracking
- AttnMove: History Enhanced Trajectory Recovery via Attentional Network
- Can large language models assist choice modelling? Insights into prompting strategies and current models capabilities
- Entropy-Enhanced Multimodal Attention Model for Scene-Aware Dialogue Generation
- Memorization in Fine-Tuned Large Language Models
- Order Matters: Sequence to sequence for sets
- A copula-based visualization technique for a neural network
- Why Pay More When You Can Pay Less: A Joint Learning Framework for Active Feature Acquisition and Classification
- Advancing Dialectal Arabic to Modern Standard Arabic Machine Translation
- Comparative Visual Analytics for Assessing Medical Records with Sequence Embedding
- EcoTransformer: Attention without Multiplication
- Stealing Links from Graph Neural Networks
- NERO: A Neural Rule Grounding Framework for Label-Efficient Relation Extraction
- Transfer Learning for Clinical Time Series Analysis using Recurrent Neural Networks
- Chunk-Based Bi-Scale Decoder for Neural Machine Translation
- Endowing Deep 3D Models with Rotation Invariance Based on Principal Component Analysis
- Grounded Conversation Generation as Guided Traverses in Commonsense Knowledge Graphs
- Patch Pruning Strategy Based on Robust Statistical Measures of Attention Weight Diversity in Vision Transformers
- GIIFT: Graph-guided Inductive Image-free Multimodal Machine Translation
- Restoring Rhythm: Punctuation Restoration Using Transformer Models for Bangla, a Low-Resource Language
- Micromobility Flow Prediction: A Bike Sharing Station-level Study via Multi-level Spatial-Temporal Attention Neural Network
- DCFFSNet: Deep Connectivity Feature Fusion Separation Network for Medical Image Segmentation
- Adapting the Neural Encoder-Decoder Framework from Single to Multi-Document Summarization
- Understanding Hidden Memories of Recurrent Neural Networks
- Deep Recurrent Neural Network for Protein Function Prediction from Sequence
- Higher Order Recurrent Neural Networks
- The Origin of Self-Attention: Pairwise Affinity Matrices in Feature Selection and the Emergence of Self-Attention
- Machine Translation between Vietnamese and English: an Empirical Study
- Recurrent Attention Unit
- Grading video interviews with fairness considerations
- The Costs and Benefits of Goal-Directed Attention in Deep Convolutional Neural Networks
- Neural Machine Translation with Monolingual Translation Memory
- Guiding Neural Machine Translation with Retrieved Translation Pieces
- QiMeng-MuPa: Mutual-Supervised Learning for Sequential-to-Parallel Code Translation
- On Inductive Biases for Machine Learning in Data Constrained Settings
- A Data-Efficient Deep Learning Based Smartphone Application For Detection Of Pulmonary Diseases Using Chest X-rays
- Doubly Attentive Transformer Machine Translation
- PARK: Personalized academic retrieval with knowledge-graphs
- The Accidental Pump and Dump: When Agentic AI Meets Autonomous Trading
- Generative AI in Qualitative Research and Related Transparency Problems: A Novel Heuristic for Disclosing Uses of AI
- Modeling Attention Flow on Graphs
- Context-Aware Sequence-to-Sequence Models for Conversational Systems
- Transformation Networks for Target-Oriented Sentiment Classification
- Accelerating Asynchronous Stochastic Gradient Descent for Neural Machine Translation
- Achieving Robust Channel Estimation Neural Networks by Designed Training Data
- From Time-series Generation, Model Selection to Transfer Learning: A Comparative Review of Pixel-wise Approaches for Large-scale Crop Mapping
- Analyzing Emotions in Bangla Social Media Comments Using Machine Learning and LIME
- Emerging Properties in Self-Supervised Vision Transformers
- On the Definition of Japanese Word
- Incremental Adaptation of NMT for Professional Post-editors: A User Study
- Deep Learning: Our Miraculous Year 1990-1991
- Seq2seq Translation Model for Sequential Recommendation
- OSU Multimodal Machine Translation System Report
- Translating Natural Language to SQL using Pointer-Generator Networks and How Decoding Order Matters
- Recurrent Neural Networks for Time Series Forecasting: Current Status and Future Directions
- Feedforward Sequential Memory Networks: A New Structure to Learn Long-term Dependency
- Paraphrase Generation as Unsupervised Machine Translation
- Domain-Adaptive Small Language Models for Structured Tax Code Prediction
- Time-series forecasting for nonlinear high-dimensional system using hybrid method combining autoencoder and multi-parallelized quantum long short-term memory and gated recurrent unit
- TSRec: Enhancing Repeat-Aware Recommendation from a Temporal-Sequential Perspective
- Understood in Translation, Transformers for Domain Understanding
- Multimodal Dialogue State Tracking By QA Approach with Data Augmentation
- Multi-Head Attention: Collaborate Instead of Concatenate
- Visual Question Answering with Memory-Augmented Networks
- Semantic-Unit-Based Dilated Convolution for Multi-Label Text Classification
- CycleGAN-Driven Transfer Learning for Electronics Response Emulation in High-Purity Germanium Detectors
- Queue up for takeoff: a transferable deep learning framework for flight delay prediction
- TACTIC: Translation Agents with Cognitive-Theoretic Interactive Collaboration
- Towards Extracting Software Requirements from App Reviews using Seq2seq Framework
- Emergent Natural Language with Communication Games for Improving Image Captioning Capabilities without Additional Data
- Universal Approximation Theorem for a Single-Layer Transformer
- Generative Language-Grounded Policy in Vision-and-Language Navigation with Bayes' Rule
- Quantifying Mix Network Privacy Erosion with Generative Models
- Probabilistic Transformers
- SAS: Simulated Attention Score
- Cross-Lingual Transfer for Machine Translation in Turkic Languages
- A multi-task encoder-dual-decoder framework for mixed frequency data prediction
- KPFlow: An Operator Perspective on Dynamic Collapse Under Gradient Descent Training of Recurrent Networks
- PhoniTale: Phonologically Grounded Mnemonic Generation for Typologically Distant Language Pairs
- A Survey of Pun Generation: Datasets, Evaluations and Methodologies
- PRISM: Pointcloud Reintegrated Inference via Segmentation and Cross-attention for Manipulation
- Multi-node Bert-pretraining: Cost-efficient Approach
- ViTaL: A Multimodality Dataset and Benchmark for Multi-pathological Ovarian Tumor Recognition
- Relational inductive biases on attention mechanisms
- Graph Collaborative Attention Network for Link Prediction in Knowledge Graphs
- RECA-PD: A Robust Explainable Cross-Attention Method for Speech-based Parkinson's Disease Classification
- Automatically Generating Commit Messages from Diffs using Neural Machine Translation
- Memory Mosaics at scale
- Linear Attention with Global Context: A Multipole Attention Mechanism for Vision and Physics
- Can Artificial Intelligence solve the blockchain oracle problem? Unpacking the Challenges and Possibilities
- Multi-hop Inference for Question-driven Summarization
- Listen, Attend, and Walk: Neural Mapping of Navigational Instructions to Action Sequences
- A Sketch-Based Neural Model for Generating Commit Messages from Diffs
- A Deep Neural Network for Unsupervised Anomaly Detection and Diagnosis in Multivariate Time Series Data
- Memory-augmented Dense Predictive Coding for Video Representation Learning
- A Review on Sound Source Localization in Robotics: Focusing on Deep Learning Methods
- Examining Reject Relations in Stimulus Equivalence Simulations
- Learning quantum phase transition in parametrized quantum circuits with an attention mechanism
- Thinking About Thinking: SAGE-nano's Inverse Reasoning for Self-Aware Language Models
- Multilingual Hate Speech Detection in Social Media Using Translation-Based Approaches with Large Language Models
- Minimizing the Bag-of-Ngrams Difference for Non-Autoregressive Neural Machine Translation
- Relevance-Promoting Language Model for Short-Text Conversation
- Hierarchical Lexical Graph for Enhanced Multi-Hop Retrieval
- Opportunities and challenges in developing deep learning models using electronic health records data: a systematic review
- Concept Pinpoint Eraser for Text-to-image Diffusion Models via Residual Attention Gate
- Improved and Robust Controversy Detection in General Web Pages Using Semantic Approaches under Large Scale Conditions
- Object Based Attention Through Internal Gating
- SU-RUG at the CoNLL-SIGMORPHON 2017 shared task: Morphological Inflection with Attentional Sequence-to-Sequence Models
- Offensive Language Detection on Social Media Using XLNet
- GoIRL: Graph-Oriented Inverse Reinforcement Learning for Multimodal Trajectory Prediction
- Communicating Smartly in Molecular Communication Environments: Neural Networks in the Internet of Bio-Nano Things
- Variational Digital Twins
- Program Language Translation Using a Grammar-Driven Tree-to-Tree Model
- Beyond the Sentence: A Survey on Context-Aware Machine Translation with Large Language Models
- daVinciNet: Joint Prediction of Motion and Surgical State in Robot-Assisted Surgery
- DeepSense: A Unified Deep Learning Framework for Time-Series Mobile Sensing Data Processing
- Enhanced Hybrid Transducer and Attention Encoder Decoder with Text Data
- FPETS : Fully Parallel End-to-End Text-to-Speech System
- An Interpretable Transformer-Based Foundation Model for Cross-Procedural Skill Assessment Using Raw fNIRS Signals
- Sequential Interpretability: Methods, Applications, and Future Direction for Understanding Deep Learning Models in the Context of Sequential Data
- A Survey of State Representation Learning for Deep Reinforcement Learning
- VeriLocc: End-to-End Cross-Architecture Register Allocation via LLM
- Sequence-to-Sequence Models with Attention Mechanistically Map to the Architecture of Human Memory Search
- Reinforced Mnemonic Reader for Machine Reading Comprehension
- Fine-Tuning by Curriculum Learning for Non-Autoregressive Neural Machine Translation
- Aligning where to see and what to tell: image caption with region-based attention and scene factorization
- Zero-Shot Cognitive Impairment Detection from Speech Using AudioLLM
- Better Language Model Inversion by Compactly Representing Next-Token Distributions
- RobustART: Benchmarking Robustness on Architecture Design and Training Techniques
- A Hybrid DeBERTa and Gated Broad Learning System for Cyberbullying Detection in English Text
- Non-Projective Dependency Parsing via Latent Heads Representation (LHR)
- Passing the Turing Test in Political Discourse: Fine-Tuning LLMs to Mimic Polarized Social Media Comments
- Understanding LLM-Centric Challenges for Deep Learning Frameworks: An Empirical Analysis
- Investigation of learning abilities on linguistic features in sequence-to-sequence text-to-speech synthesis
- AlphaEvolve: A coding agent for scientific and algorithmic discovery
- Improved English to Russian Translation by Neural Suffix Prediction
- Detecting Hard-Coded Credentials in Software Repositories via LLMs
- Churn Intent Detection in Multilingual Chatbot Conversations and Social Media
- End-to-End Non-Autoregressive Neural Machine Translation with Connectionist Temporal Classification
- The Helsinki Neural Machine Translation System
- Wat zei je? Detecting Out-of-Distribution Translations with Variational Transformers
- Automated proof synthesis for propositional logic with deep neural networks
- Bridging the Digital Divide: Small Language Models as a Pathway for Physics and Photonics Education in Underdeveloped Regions
- Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs
- FAD-Net: Frequency-Domain Attention-Guided Diffusion Network for Coronary Artery Segmentation using Invasive Coronary Angiography
- Retrieval-Augmented Code Review Comment Generation
- Principled Approaches for Extending Neural Architectures to Function Spaces for Operator Learning
- Interior-Point Vanishing Problem in Semidefinite Relaxations for Neural Network Verification
- Solving a New 3D Bin Packing Problem with Deep Reinforcement Learning Method
- Enhancing Medical Dialogue Generation through Knowledge Refinement and Dynamic Prompt Adjustment
- Controllable Coupled Image Generation via Diffusion Models
- How To Evaluate Your Dialogue System: Probe Tasks as an Alternative for Token-level Evaluation Metrics
- From Symbolic to Neural and Back: Exploring Knowledge Graph-Large Language Model Synergies
- Cross-Media Keyphrase Prediction: A Unified Framework with Multi-Modality Multi-Head Attention and Image Wordings
- Video Paragraph Captioning Using Hierarchical Recurrent Neural Networks
- Transformative or Conservative? Conservation laws for ResNets and Transformers
- Zoneout: Regularizing RNNs by Randomly Preserving Hidden Activations
- Automatic Text Summarization of COVID-19 Medical Research Articles using BERT and GPT-2
- Power Law Guided Dynamic Sifting for Efficient Attention
- A Novel Transformer-Based Method for Full Lower-Limb Joint Angles and Moments Prediction in Gait Using sEMG and IMU data
- MesaNet: Sequence Modeling by Locally Optimal Test-Time Training
- Domain Control for Neural Machine Translation
- ConECT Dataset: Overcoming Data Scarcity in Context-Aware E-Commerce MT
- Regression as Classification: Influence of Task Formulation on Neural Network Features
- Voice Imitating Text-to-Speech Neural Networks
- A Deep Learning Approach with an Attention Mechanism for Automatic Sleep Stage Classification
- Discriminative Neural Clustering for Speaker Diarisation
- Jointly Trained Sequential Labeling and Classification by Sparse Attention Neural Networks
- On Extractive and Abstractive Neural Document Summarization with Transformer Language Models
- LCSTS: A Large Scale Chinese Short Text Summarization Dataset
- Adversarially Robust Neural Architectures
- On the Dimensionality of Word Embedding
- Backpropagation through time and the brain
- A Skeleton-Based Model for Promoting Coherence Among Sentences in Narrative Story Generation
- Sequential Context Encoding for Duplicate Removal
- To Compress, or Not to Compress: Characterizing Deep Learning Model Compression for Embedded Inference
- Sequence-to-Set Semantic Tagging: End-to-End Multi-label Prediction using Neural Attention for Complex Query Reformulation and Automated Text Categorization
- Adversarial Ranking for Language Generation
- Reconstruction Network for Video Captioning
- ReDecode Framework for Iterative Improvement in Paraphrase Generation
- Attention-Aware Compositional Network for Person Re-identification
- Sounding that Object: Interactive Object-Aware Image to Audio Generation
- Shift-Reduce Constituent Parsing with Neural Lookahead Features
- CodeSum: Translate Program Language to Natural Language
- SynLang and Symbiotic Epistemology: A Manifesto for Conscious Human-AI Collaboration
- End-to-end Silent Speech Recognition with Acoustic Sensing
- Learning to Optimize in Swarms
- Natural, Artificial, and Human Intelligences
- MVAN: Multi-View Attention Networks for Fake News Detection on Social Media
- Conditional Variational Autoencoder for Neural Machine Translation
- Visual Concept Reasoning Networks
- SynPO: Synergizing Descriptiveness and Preference Optimization for Video Detailed Captioning
- Visual Agreement Regularized Training for Multi-Modal Machine Translation
- A Deep Learning Approach to Automate High-Resolution Blood Vessel Reconstruction on Computerized Tomography Images With or Without the Use of Contrast Agent
- Channel-Imposed Fusion: A Simple yet Effective Method for Medical Time Series Classification
- Tensor Programs I: Wide Feedforward or Recurrent Neural Networks of Any Architecture are Gaussian Processes
- Subjective Bias in Abstractive Summarization
- Improving Language and Modality Transfer in Translation by Character-level Modeling
- Two failure modes of deep transformers and how to avoid them: a unified theory of signal propagation at initialisation
- Generation of Synthetic Electronic Medical Record Text
- Machine Translation Approaches and Survey for Indian Languages
- Evaluating Syntactic Properties of Seq2seq Output with a Broad Coverage HPSG: A Case Study on Machine Translation
- Explainable Depression Detection using Masked Hard Instance Mining
- Relational Neural Expectation Maximization: Unsupervised Discovery of Objects and their Interactions
- A crossover code for high-dimensional composition
- Learning to Encode Evolutionary Knowledge for Automatic Commenting Long Novels
- Using holistic event information in the trigger
- Real-time low-resource phoneme recognition on edge devices
- Untangling tradeoffs between recurrence and self-attention in neural networks
- Towards one-shot learning for rare-word translation with external experts
- Structured-based Curriculum Learning for End-to-end English-Japanese Speech Translation
- Two-Stage Feature Generation with Transformer and Reinforcement Learning
- Towards Robust Assessment of Pathological Voices via Combined Low-Level Descriptors and Foundation Model Representations
- Identifying Super Spreaders in Multilayer Networks
- A Physics-Augmented GraphGPS Framework for the Reconstruction of 3D Riemann Problems from Sparse Data
- S-OHEM: Stratified Online Hard Example Mining for Object Detection
- Decomposing Complex Questions Makes Multi-Hop QA Easier and More Interpretable
- VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation
- Joint Training for Neural Machine Translation Models with Monolingual Data
- Towards Proof Synthesis Guided by Neural Machine Translation for Intuitionistic Propositional Logic
- The pointer network for reward maximisation in multi-target space mission sequence selection
- Design Challenges in Named Entity Transliteration
- Dual Ask-Answer Network for Machine Reading Comprehension
- CSVideoNet: A Real-time End-to-end Learning Framework for High-frame-rate Video Compressive Sensing
- Detecting Interrogative Utterances with Recurrent Neural Networks
- Multimodal Memory Modelling for Video Captioning
- Classification of assembly tasks combining multiple primitive actions using Transformers and xLSTMs
- LMask: Learn to Solve Constrained Routing Problems with Lazy Masking
- Selection Mechanisms for Sequence Modeling using Linear State Space Models
- Low-Resource NMT: A Case Study on the Written and Spoken Languages in Hong Kong
- DeepTransport: Learning Spatial-Temporal Dependency for Traffic Condition Forecasting
- A Multi-Head Attention Soft Random Forest for Interpretable Patient No-Show Prediction
- Attention with Trained Embeddings Provably Selects Important Tokens
- Self-attention based end-to-end Hindi-English Neural Machine Translation
- Physical models realizing the transformer architecture of large language models
- A Framework for Non-Linear Attention via Modern Hopfield Networks
- SUS backprop: linear backpropagation algorithm for long inputs in transformers
- Cross-Modal Alignment with Mixture Experts Neural Network for Intral-City Retail Recommendation
- Secrets Everywhere: Auditing Memorization in Mobility Prediction Models
- Tversky Neural Networks: Psychologically Plausible Deep Learning with Differentiable Tversky Similarity
- Grounded Recurrent Neural Networks
- Abstract Syntax Networks for Code Generation and Semantic Parsing
- Regularizing Output Distribution of Abstractive Chinese Social Media Text Summarization for Improved Semantic Consistency
- Challenges and Thrills of Legal Arguments
- Continuous Space Reordering Models for Phrase-based MT
- TransBench: Benchmarking Machine Translation for Industrial-Scale Applications
- Probabilistic Attention for Interactive Segmentation
- Augmenting human innovation teams with artificial intelligence: Exploring transformer‐based language models
- Going deep into schizophrenia with artificial intelligence
- Attention-based clustering
- Lightweight and Interpretable Transformer via Mixed Graph Algorithm Unrolling for Traffic Forecast
- Meta-Learning a Dynamical Language Model
- Combining the Best of Both Worlds: A Method for Hybrid NMT and LLM Translation
- Improving Next-Application Prediction with Deep Personalized-Attention Neural Network
- SchoenbAt: Rethinking Attention with Polynomial basis
- Learning to Dissipate Energy in Oscillatory State-Space Models
- Is Neural Machine Translation Ready for Deployment? A Case Study on 30 Translation Directions
- A Survey on Large Language Models in Multimodal Recommender Systems
- Schedule repair for flexible job shops under machine breakdowns by deep reinforcement learning
- Towards a safe and efficient clinical implementation of machine learning in radiation oncology by exploring model interpretability, explainability and data-model dependency
- Recursive Recurrent Nets with Attention Modeling for OCR in the Wild
- Unsupervised Neural Text Simplification
- Multi-Component Graph Convolutional Collaborative Filtering
- Putting It All into Context: Simplifying Agents with LCLMs
- EmotionX-DLC: Self-Attentive BiLSTM for Detecting Sequential Emotions in Dialogue
- Overflow Prevention Enhances Long-Context Recurrent LLMs
- What value do explicit high level concepts have in vision to language problems?
- Learning Penalty for Optimal Partitioning via Automatic Feature Extraction
- Security through the Eyes of AI: How Visualization is Shaping Malware Detection
- Joint Intent Detection And Slot Filling Based on Continual Learning Model
- Attention-based Mixture Density Recurrent Networks for History-based Recommendation
- Attention-Enhanced Reservoir Computing as a Multiple Dynamical System Approximator
- Graph Constrained Reinforcement Learning for Natural Language Action Spaces
- FNED
- Generative Models for Long Time Series: Approximately Equivariant Recurrent Network Structures for an Adjusted Training Scheme
- Geometric Analysis of Token Selection in Multi-Head Attention
- Skeleton Key: Image Captioning by Skeleton-Attribute Decomposition
- A Multi-task Selected Learning Approach for Solving 3D Flexible Bin Packing Problem
- A Comprehensive Analysis of Adversarial Attacks against Spam Filters
- Multimodal Machine Learning: A Survey and Taxonomy
- k-Nearest Neighbor Augmented Neural Networks for Text Classification
- QiMeng-Xpiler: Transcompiling Tensor Programs for Deep Learning Systems with a Neural-Symbolic Approach
- Concept Tagging for Natural Language Understanding: Two Decadelong Algorithm Development
- SAC: Accelerating and Structuring Self-Attention via Sparse Adaptive Connection
- Syntactic Scaffolds for Semantic Structures
- GeloVec: Higher Dimensional Geometric Smoothing for Coherent Visual Feature Extraction in Image Segmentation
- Applications of deep learning in stock market prediction: recent progress
- Harnessing Structured Knowledge: A Concept Map-Based Approach for High-Quality Multiple Choice Question Generation with Effective Distractors
- Detecting the Root Cause Code Lines in Bug-Fixing Commits by Heterogeneous Graph Learning
- Dynamic Forecasting and Temporal Feature Evolution of Stock Repurchases in Listed Companies Using Attention-Based Deep Temporal Networks
- SA-GAT-SR: Self-Adaptable Graph Attention Networks with Symbolic Regression for high-fidelity material property prediction
- The LAMBADA dataset: Word prediction requiring a broad discourse context
- Relative Positional Encoding for Transformers with Linear Complexity
- Personalizing Search Results Using Hierarchical RNN with Query-aware Attention
- Dual-Primal Graph Convolutional Networks
- Show and tell: A neural image caption generator
- Adversarial Examples in the Physical World
- Automatic Acrostic Couplet Generation with Three-Stage Neural Network Pipelines
- The Neuroscience of Transformers
- In-context learning emerges in chemical reaction networks without attention
- Machine Translation: A Literature Review
- Exploiting Linguistic Resources for Neural Machine Translation Using Multi-task Learning
- QCRI Machine Translation Systems for IWSLT 16
- A Multi-Object Rectified Attention Network for Scene Text Recognition
- Learning Spatial-Semantic Context with Fully Convolutional Recurrent Network for Online Handwritten Chinese Text Recognition
- A contextual hierarchical attention network for detecting mental health disorders using social media
- On Subquadratic Architectures: From Applications to Principles
- A Reinforced Generation of Adversarial Examples for Neural Machine Translation
- Modeling of Rakugo Speech and Its Limitations: Toward Speech Synthesis That Entertains Audiences
- Mutual Information Scaling and Expressive Power of Sequence Models
- Image to Video Domain Adaptation Using Web Supervision
- Evaluating Sequence-to-Sequence Learning Models for If-Then Program Synthesis
- Signed Dual Attention: Capturing Signed Dependencies in Time Series Forecasting
- When Deep Learning Meets Information Retrieval-based Bug Localization: A Survey
- Learning better discourse representation for implicit discourse relation recognition via attention networks
- Strategies for Training Large Vocabulary Neural Language Models
- Fast Transformer Inference on ARM-Based HMPSoCs
- Translators as Invisible Teachers of AI: Copyright, Translation Memory, and the Political Economy of Linguistic Data
- Polish - English Speech Statistical Machine Translation Systems for the\n IWSLT 2014
- Polysemy of Synthetic Neurons Towards a New Type of Explanatory Categorical Vector Spaces
- Hierarchical Multi-Label Generation with Probabilistic Level-Constraint
- A comparative study of deep learning and ensemble learning to extend the horizon of traffic forecasting
- Jekyll-and-Hyde Tipping Point in an AI's Behavior
- Softpick: No Attention Sink, No Massive Activations with Rectified Softmax
- Characterizing the Expressivity of Local Attention in Transformers
- Leveraging Depth Maps and Attention Mechanisms for Enhanced Image Inpainting
- Enhancing event reconstruction for γ-ray particle detector arrays using transformers
- The State-Prediction Separation Hypothesis
- Position-aware deep multi-task learning for drug–drug interaction extraction
- Improving Generalization of Transfer Learning Across Domains Using Spatio-Temporal Features in Autonomous Driving
- Nested Learning: The Illusion of Deep Learning Architectures
- Enhanced Partially Relevant Video Retrieval through Inter- and Intra-Sample Analysis with Coherence Prediction
- Hierarchical Reinforcement Learning in Multi-Goal Spatial Navigation with Autonomous Mobile Robots
- Spatial Speech Translation: Translating Across Space With Binaural Hearables
- Myocardial Infarction Severity Stages Classification From ECG Signals Using Attentional Recurrent Neural Network
- Escaping the BLEU Trap: A Signal-Grounded Framework with Decoupled Semantic Guidance for EEG-to-Text Decoding
- Deep Learning
- A Deep Prediction Network for Understanding Advertiser Intent and Satisfaction
- Low-Resource Neural Machine Translation Using Recurrent Neural Networks and Transfer Learning: A Case Study on English-to-Igbo
- You Need Better Attention Priors
- Simplify-then-Translate: Automatic Preprocessing for Black-Box Machine Translation
- Cross-Level Cross-Scale Cross-Attention Network for Point Cloud Representation
- Instance-aware Image and Sentence Matching with Selective Multimodal LSTM
- Semi-Supervised Confidence Network aided Gated Attention based Recurrent Neural Network for Clickbait Detection
- The Sparse Frontier: Sparse Attention Trade-offs in Transformer LLMs
- Amobee at SemEval-2018 Task 1: GRU Neural Network with a CNN Attention Mechanism for Sentiment Classification
- Deep Choice Model Using Pointer Networks for Airline Itinerary Prediction
- A General Survey on Attention Mechanisms in Deep Learning
- Representation Learning for Natural Language Processing
- Inf-VAE
- Deep Multi-View Learning for Tire Recommendation
- Attention-based Clinical Note Summarization
- End-to-end Continuous Speech Recognition using Attention-based Recurrent NN: First Results
- Wavelet-Enhanced Sequence-to-Sequence Modeling with Attention Mechanism for Short-Term Wind Power Forecasting
- A Spatial-Temporal Graph Neural Network Framework for Automated Software Bug Triaging
- Semantic Graphs for Generating Deep Questions
- Capturing AI's Attention: Physics of Repetition, Hallucination, Bias and Beyond
- A Novel Hybrid Approach Using an Attention-Based Transformer + GRU Model for Predicting Cryptocurrency Prices
- Multivariate LSTM-FCNs for time series classification
- A Hybrid Convolutional Variational Autoencoder for Text Generation
- Set-to-Sequence Methods in Machine Learning: A Review
- Neural Rating Regression with Abstractive Tips Generation for Recommendation
- Attention-Based Multimodal Fusion for Video Description
- SCA-CNN: Spatial and Channel-wise Attention in Convolutional Networks for Image Captioning
- NUIG-Shubhanker@Dravidian-CodeMix-FIRE2020: Sentiment Analysis of Code-Mixed Dravidian text using XLNet
- GADS: A Super Lightweight Model for Head Pose Estimation
- Teaching Machines to Code: Neural Markup Generation with Visual Attention
- A survey of state-of-the-art approaches for emotion recognition in text
- Multi-Level Attention Pooling for Graph Neural Networks: Unifying Graph Representations with Multiple Localities
- What do Neural Machine Translation Models Learn about Morphology?
- Quantitative Clustering in Mean-Field Transformer Models
- Self-Explaining Structures Improve NLP Models
- Learning to Attribute with Attention
- Multimodal time-aware attention networks for depression detection
- Automatic Commit Message Generation: A Critical Review and Directions for Future Work
- Generating captions without looking beyond objects
- Prior Knowledge Driven Label Embedding for Slot Filling in Natural Language Understanding
- Insights on Neural Representations for End-to-End Speech Recognition
- From Large to Super-Tiny: End-to-End Optimization for Cost-Efficient LLMs
- Question Answering over Knowledge Base using Language Model Embeddings
- Artificial intelligence and deep learning algorithms for epigenetic sequence analysis: A review for epigeneticists and AI experts
- Identifying Harm Events in Clinical Care through Medical Narratives
- Sequence-to-sequence models for workload interference prediction on batch processing datacenters
- Local minima in training of neural networks
- Experiment Segmentation in Scientific Discourse as Clause-level Structured Prediction using Recurrent Neural Networks
- Tag-less back-translation
- End-to-End ASR-Free Keyword Search From Speech
- Spatial-temporal Conv-sequence Learning with Accident Encoding for Traffic Flow Prediction
- Environmental sound classification using convolution neural networks with different integrated loss functions
- Structure-Tags Improve Text Classification for Scholarly Document Quality Prediction
- Denoising Multi-Source Weak Supervision for Neural Text Classification
- Sparks of Science: Hypothesis Generation Using Structured Paper Data
- Why and How LLMs Hallucinate: Connecting the Dots with Subsequence Associations
- Explainable Scene Understanding with Qualitative Representations and Graph Neural Networks
- Multilingual Contextualization of Large Language Models for Document-Level Machine Translation
- Clarifying Ambiguities: on the Role of Ambiguity Types in Prompting Methods for Clarification Generation
- SemDiff: Generating Natural Unrestricted Adversarial Examples via Semantic Attributes Optimization in Diffusion Models
- Predicting Wave Dynamics using Deep Learning with Multistep Integration Inspired Attention and Physics-Based Loss Decomposition
- Action-Attending Graphic Neural Network
- Pay Attention to What and Where? Interpretable Feature Extractor in Vision-based Deep Reinforcement Learning
- Siamese Network with Dual Attention for EEG-Driven Social Learning: Bridging the Human-Robot Gap in Long-Tail Autonomous Driving
- Ordinary Least Squares as an Attention Mechanism
- Modeling Time Series Similarity with Siamese Recurrent Networks
- Large Language Models and Attention-Based AI for Hardware Design and Security: Progress, Challenges, and Opportunities
- SRVP: Strong Recollection Video Prediction Model Using Attention-Based Spatiotemporal Correlation Fusion
- PROPEL: Supervised and Reinforcement Learning for Large-Scale Supply Chain Planning
- Predicting Temporal Sets with Deep Neural Networks
- Show and Tell: Lessons Learned from the 2015 MSCOCO Image Captioning Challenge
- MHSA-Net: Multihead Self-Attention Network for Occluded Person Re-Identification
- Hogwild! Inference: Parallel LLM Generation via Concurrent Attention
- Separator Injection Attack: Uncovering Dialogue Biases in Large Language Models Caused by Role Separators
Discussions
- Me rationalising what I am currently doing: There is no shame in designing a recipish, clunky, deep architecture kept together with duct tape. This is how we move into new territories. The Transformer [bsky, 26 points, 1 comments]
- interesting connection, but the transformer paper didn’t invent attention? arxiv.org/abs/1409.0473 [bsky, 23 points, 1 comments]
- For those of you who don't know (or don't remember), the original paper introducing "attention" was Bahdanau, Cho, & Bengio's "Neural machine translation by jointly learning to align and translate" (2 [bsky, 8 points, 1 comments]
- Neural Machine Translation by Jointly Learning to Align and Translate [hn, 2 points, 0 comments]
- וזה המאמר שתאר לראשונה-Attention עבור רשתות ניורונים: arxiv.org/abs/1409.0473 אבל בתכלס עדיף לקרוא דברים מאוחרים יותר. [bsky, 2 points, 0 comments]
- CDS Prof. @kyunghyuncho.bsky.social's 2014 "attention" paper was recently the Runner-Up for the ICLR 2025 Test of Time Award. The paper introduced dynamic attention in machine translation, laying the [bsky, 2 points, 0 comments]
- We've come full circle [lemmy, 1 points, 0 comments]
- Neural Machine Translation by Jointly Learning to Align and Translate (2015) [hn, 1 points, 0 comments]
- what a nice paper this is arxiv.org/pdf/1409.0473 [bsky, 1 points, 0 comments]
Related