Attention Is All You Need
The Transformer, an attention-only architecture, beat prior best translation results while training far faster than RNN or CNN models.
2017/06/12 by Ashish Vaswani, Vaswani, Ashish, Noam Shazeer +13 · 85 voices · 6522 citations
#cs.CL #cs.LG
paper · pdf · doi:10.48550/arxiv.1706.03762
Abstract
The dominant sequence transduction models are based on complex recurrent or convolutional neural networks in an encoder-decoder configuration. The best performing models also connect the encoder and decoder through an attention mechanism. We propose a new simple network architecture, the Transformer, based solely on attention mechanisms, dispensing with recurrence and convolutions entirely. Experiments on two machine translation tasks show these models to be superior in quality while being more parallelizable and requiring significantly less time to train. Our model achieves 28.4 BLEU on the WMT 2014 English-to-German translation task, improving over the existing best results, including ensembles by over 2 BLEU. On the WMT 2014 English-to-French translation task, our model establishes a new single-model state-of-the-art BLEU score of 41.8 after training for 3.5 days on eight GPUs, a small fraction of the training costs of the best models from the literature. We show that the Transformer generalizes well to other tasks by applying it successfully to English constituency parsing both with large and limited training data.
Summary
The paper introduces the Transformer, a sequence-to-sequence architecture built entirely out of attention mechanisms, with no recurrence or convolution at all. On two machine translation benchmarks it beat the previous best published models (including ensembles) while training in a fraction of the time, and a version of it also did well on English constituency parsing.
machine-generated · claude-sonnet-5
Outline
- Introduction — Recurrent models process sequences step by step, which blocks parallelization; the paper proposes replacing recurrence with attention to fix this.
- Background — Prior work (ByteNet, ConvS2S) used convolutions to parallelize sequence processing but still struggled to relate distant positions; the Transformer is presented as the first model to rely entirely on self-attention.
- Model Architecture — Stacked encoder and decoder, each built from self-attention and feed-forward sub-layers with residual connections and layer normalization.
- Attention — Defines scaled dot-product attention and multi-head attention, and describes the three places attention is used in the model.
- Position-wise Feed-Forward Networks — A two-layer ReLU network applied identically to each sequence position.
- Embeddings and Softmax — Input/output tokens and the pre-softmax layer share a learned embedding matrix, scaled by sqrt(dmodel).
- Positional Encoding — Since there's no recurrence or convolution, fixed sine/cosine signals are added to embeddings to encode token position.
- Why Self-Attention — Compares self-attention, recurrent, and convolutional layers on computational cost, parallelizability, and path length between distant positions.
- Training — Details on data (WMT 2014 En-De and En-Fr), hardware (8 P100 GPUs), the Adam optimizer with a warm-up learning-rate schedule, and dropout/label-smoothing regularization.
- Results: Machine Translation — 28.4 BLEU on En-De and 41.8 BLEU on En-Fr, both new state-of-the-art, at a fraction of the training cost of prior best models.
- Results: Model Variations — Ablations showing effects of head count, key/value dimension, model size, dropout, and positional encoding choice.
- Results: English Constituency Parsing — A 4-layer Transformer applied to parsing without task-specific tuning outperforms most prior parsers.
- Conclusion — States plans to extend the architecture to non-text modalities, explore restricted/local attention for long sequences, and make generation less sequential.
machine-generated · claude-sonnet-5
Claims
- The Transformer (big) reaches 28.4 BLEU on the WMT 2014 English-to-German task, more than 2 BLEU above the previous best result including ensembles. [experiment]
- The Transformer (big) reaches 41.8 BLEU on the WMT 2014 English-to-French task, a new single-model state of the art, after 3.5 days of training on 8 GPUs. [experiment]
- A self-attention layer connects any two sequence positions with a constant number of sequential operations, versus O(n) for a recurrent layer, and is computationally cheaper than a recurrent layer when sequence length is smaller than representation dimension. [argument]
- Scaling dot products by 1/sqrt(dk) counteracts the softmax being pushed into low-gradient regions for large key dimensions. [argument]
- Multi-head attention outperforms single-head attention at matched computational cost, but quality degrades again with too many heads. [experiment]
- Learned positional embeddings produce translation quality nearly identical to the fixed sinusoidal positional encoding. [experiment]
- A 4-layer Transformer trained only on the 40K-sentence WSJ Penn Treebank outperforms the BerkeleyParser and most other previously reported constituency parsers. [experiment]
machine-generated · claude-sonnet-5
Key figure
Figure 1 — A diagram of the full Transformer architecture: the encoder stack on the left and the decoder stack on the right, each made of repeated blocks of multi-head attention and feed-forward layers.
machine-generated · claude-sonnet-5
Glossary
- Attention mechanism
- A way for a model to weigh and combine information from different parts of its input based on relevance, rather than processing it strictly in order.
- Self-attention
- Attention applied within a single sequence, letting each position look at every other position in that same sequence to build its representation.
- Multi-head attention
- Running several attention computations in parallel on different learned projections of the input, then combining the results, so the model can attend to different kinds of relationships at once.
- Encoder-decoder
- A model structure where one component (encoder) turns an input sequence into an internal representation, and another (decoder) generates the output sequence from that representation.
- BLEU score
- An automated metric that scores machine translation output by comparing overlapping word sequences against human reference translations; higher is better.
- Positional encoding
- A signal added to each token's embedding so the model can tell tokens apart by their position in the sequence, since attention alone has no built-in sense of order.
- Residual connection / layer normalization
- A shortcut that adds a layer's input to its output (helping gradients flow through deep networks), followed by rescaling the result to stabilize training.
- Byte-pair / word-piece encoding
- A way of splitting text into subword units so a fixed, moderate-size vocabulary can still represent rare and unseen words.
- Beam search
- A decoding strategy that keeps several of the most likely partial output sequences at each step instead of only the single best one.
- Label smoothing
- A training trick that discourages the model from becoming overconfident in its predictions, which lowers perplexity but improves BLEU and accuracy.
machine-generated · claude-sonnet-5
Audience
Machine learning practitioners and researchers working on sequence modeling, NLP, or machine translation who want to understand the origin of the Transformer architecture underlying most modern language models.
prerequisites: Basic neural network training concepts, Familiarity with recurrent and convolutional sequence-to-sequence models, Linear algebra (matrix multiplication, vector spaces), Some prior exposure to attention mechanisms in encoder-decoder models
machine-generated · claude-sonnet-5
Open questions
- Can self-attention be restricted to a local neighborhood of size r without hurting quality, to reduce its quadratic cost on very long sequences?
The paper notes this as future work needed to handle inputs and outputs like images, audio, and video efficiently. - Does the Transformer architecture transfer well to input and output modalities other than text?
The authors state they plan to extend the Transformer to non-text modalities but present no evidence here that it works beyond text. - Can output generation be made less sequential than the current autoregressive, one-token-at-a-time decoding?
The authors list this as an explicit research goal, implying the current decoding process is still a bottleneck.
machine-generated · claude-sonnet-5
Supplementary links
machine-generated · claude-sonnet-5
Cited by
- Learning to Rank for Selected Configuration Interaction
- Adaptive Learned State Estimation based on KalmanNet
- Numerical Fragility in Transformers: A Layer-wise Theory for Risk Estimation and Selective Stabilization
- CorVS+: Correspondence-Driven Association of Video Trajectories and Sensors for Identity-Aware Person Localization in Warehouses
- Spatially-Enhanced Temporal Fusion Transformer: Interpretable Multi-Output Prediction for Parametric Dynamical Systems with Time-Varying Inputs
- LiMuon: Light and Fast Muon Optimizer for Large Models
- Automatic Stability and Recovery for Neural Network Training
- \kappa-LoRA: Condition Numbers Reveal Which LoRA Matrices Worth Updating
- PCS-UQ: Uncertainty Quantification via the Predictability-Computability-Stability Framework
- Unsupervised Multimodal Intent Discovery via MLLM-Guided Concept Generation and Semantic Propagation
- DCS: A Unified Conditional Sensitivity Framework for Cross-Modal Copyright Infringement Detection
- CODA: Cascaded Online Discontinuity-Aware Alignment for Real-Time Image-Based Score Following
- Low-Altitude Channel Multipath Prediction via Panoramic Perception and Vision-Language Model
- When Machines Lie Differently: Detecting AI vs Human Fake News
- Universal Quantum Transformer
- A Structural Interpretation of GELU and Threshold-Transmission Activations via the First-Order Loss Function
- Transforming Keystroke Noise to Text: Self-Supervised Acoustic Eavesdropping Attacks on Keyboards
- InterOCF: Spatio-Temporal 2D-3D Interaction for Camera-Only 4D Occupancy Forecasting
- An Empirical Study of OpenPangu Quantization on Ascend NPUs
- Do Transformers Actually Help Intrusion Detection? A Temporal Sequence Evaluation on CIC-IDS2017
- RIS-Kernel: A Model-Agnostic Architecture for Long-Context LLM Inference via Sparse Attention
- Toward High-Fidelity 3D Point-Cloud Learning for Brain Folding Morphology Prediction Using Trans-Unet
- A Foundation Model for Cross-Band CSI Reconstruction
- Dynamic Commonsense Coordination for Empathetic Response Generation
- IFCLoRA: Topology-Aware Rank Allocation for Parameter-Efficient Fine-Tuning
- A Dual Path Framework with Hotspot Guided Fusion for Three Dimensional CT to PET Synthesis in Head and Neck Cancer
- An Introduction to Bayesian and Frequentist Simulation-Based Inference with Machine Learning
- IQ-JEPA: A Joint-Embedding Predictive Architecture with a Hermitian Vision Transformer for Sound Speed and Attenuation Estimation from Ultrasound IQ Data
- Interior interpretability with attention rollout: contraction and propagation profiles in Transformers
- Exact Neural-Network Representations of the Motzkin States
- Short-Term-to-Long-Term Memory Transfer for Knowledge Graphs under Partial Observability
- Pretraining Recurrent Networks without Recurrence
- Breaking the Data Barrier in Learning Symbolic Computation: A Case Study on Variable Ordering Suggestion for Cylindrical Algebraic Decomposition
- Self-Guided Process Reward Optimization with Redefined Step-wise Advantage for Process Reinforcement Learning
- Analyzing the Ethical Logic of Eight Large Language Models
- Climate-resilient electric vehicle charging infrastructure for sustainable cities: An interpretable causal-ensemble framework for preventive maintenance and low-carbon mobility
- Graph-Theoretic Neural Network Fragmentation with Covariant Direct Molecular Force Learning: Enabling Coupled-Cluster Accuracy AIMD for Fluxional Systems
- PolyGraphPy: A unified Python framework for atomistic simulation and machine learning-driven polymer design
- Mechanistic study of mixed lithium halides solid state electrolytes
- Ordered Action Tokens for Visuomotor Policy Learning
- Safe In-Context Reinforcement Learning
- IR275K: A Benchmark for Infrared Multi-Frame Super-Resolution Toward Efficient Remote Sensing
- Learning Structural Convergence: A Neuro-Symbolic Benchmark for Temporal Reasoning
- Vibe Coding: An Experiment with Test-Driven Development
- StateFormer: A Multivariate Transformer for Learning History-Dependent Battery State Dynamics and Long-Horizon Health Forecasting
- PRIMS: Physics-guided Representation for Fluid Identification in Multimodal Sensing
- RED-PIM: Reducing Data Movement for Transformers using Processing-in-Memory
- MV-GEL: Language-Driven Multi-View Geometric Entity Localization on Meshes
- Hierarchical Self-Supervised Representation Learning Framework for Multivariate Time Series Grounded in ECG Analysis
- From Vector Autoregressions to AI-based Time Series Forecasting: A Review
- What Do Temporal Graph Learning Models Learn?
- A Temporal Machine Learning-Based Time-to-Event Model for Predicting ALS Progression and Healthcare Utilization
- A Weak Penalty Neural ODE for Learning Chaotic Dynamics from Noisy Time Series
- Stabilizing Native Low-Rank LLM Pretraining
- A machine learning benchmarking framework for lipid nanoparticle transfection efficiency prediction
- LaViDa: A Large Diffusion Language Model for Multimodal Understanding
- DCVC-MB: Neural B-Frame Video Compression using State Space Models
- DyneTrion: A Spatio-temporally Coherent Generative Emulator for Protein Dynamics Across Timescales
- Neural Architectures for Amortized Bayesian Inference: Statistical Foundations and Empirical Assessments
- Trainable Log-linear Sparse Attention for Efficient Diffusion Transformers
- Multi-Task Learning for Heterogeneous Prediction from Video Game State with Transfer Learning
- Bridging Behavior and Implementation: Automated Java Glue Code Generation for Behavior-Driven Development
- PC-Edit: Prompt-Contrastive Region Discovery and Region-Guided Editing
- CHERRY: Compressed Hierarchical Experts with Recurrent Representational Yield
- From Static Bibliometrics to Dynamic Knowledge Graphs: An LLM-Powered Framework for Modernizing Science, Technology, and Innovation (STI) Analytics
- Efficient Nonlinear Multiscale Prediction for Unseen Polycrystalline Textures via Self-Supervised Microstructure Pretraining
- SHIFT: Self-reconstruction Harnesses Implicit Fine-grained Thinking for Retrieval
- LaSEr-Edit: Localized Span-level Error Editing with Energy-based Localization
- Look Less, Think Faster: Joint Token-Compute Adaptation for Multimodal LLMs
- MoAKE: Toward Unified All-in-One Action Quality Assessment via Mixture of Action Knowledge Experts
- Defer to Plan: Adaptive Multi-Agent Fusion for End-to-End V2X Driving
- Reading and Steering Representations of Materials-Science Mechanisms in an Open-Weight Language Model
- Two-Step Occupation Coding
- HijackKV: New Threat in Position-Independent KV Cache Reuse
- Test Case Prioritization for DNNs via Neural Collapse Instability
- Cumsum-Composable Phase Transport for Low-Cost Streaming Keyword Spotting
- HeadCast: Casting Attention Heads for Efficient Autoregressive Video Generation
- GPU-to-Grid: Voltage Regulation via GPU Utilization Control
- TempoNet: Slack-Quantized Transformer-Guided Reinforcement Scheduler for Adaptive Deadline-Centric Real-Time Dispatchs
- OC-Distill: Ontology-aware Contrastive Learning with Cross-Modal Distillation for ICU Risk Prediction
- Simultaneous Speech-to-Speech Translation Without Aligned Data
- Harmonia: Algorithm-Hardware Co-Design for Memory- and Compute-Efficient BFP-based LLM Inference
- Pixel-Space Diffusion Transformers
- DQAOA-GPT: AI-Accelerated Distributed Quantum Optimization for Combinatorial Problems
- Microlensing Detection and Inference via Learned Bayes Factors
- Multi-modal transformer for signal classification in nanopore blockade experiments
- ELSAA: Efficient Low-Rank and Sparse Attention Approximation for Training Transformers
- Enhancing Transformer-based Routing by Encoding Distance via Relative Positional Encoding
- IMMoE: Incomplete Multi-View Anomaly Detection via Mixture of View Experts Fusion
- NaviAIS: A Scenario-Level Vessel Trajectory Prediction Dataset withVectorized Lane Priors and the NaviLane Forecasting Framework
- Integration Matters: Rollout-Based Training for Constrained Diffusion Models
- FlexNGIA 2.0: Redesigning the Internet with Agentic AI -- Protocols, Services, and Traffic Engineering Designed, Deployed, and Managed by AI
- WorldPack: Dynamic Frame Compression for Long-context Video World Modeling
- Exposure is Optional: Learning Unlike Coordination in Language Models
- Environment-Aware Channel Inference via Cross-Modal Flow: From Multimodal Sensing to Wireless Channels
- LSRM: High-Fidelity Object-Centric Reconstruction via Scaled Context Windows
- Understanding the Impact of Linguistic Realization Choices on LLM Stance with Causal Tracing
- SR-TTT Does Not Learn Retrieval: A Correction and Mechanistic Post-Mortem of Surprisal-Aware Residual Test-Time Training
- ClothTransformer: Unified Latent-Space Transformers for Scalable Cloth Simulation
- The Power of Attention: Bridging Cognitive Load, Multimedia Learning, and AI
- Post-Training in Time Series Foundation Models: A Unifying Framework
- GaugeQuant: Online Learning of Quantization-Optimal Bases from LLM Symmetries
- SPECTRA: State-Space Exogenous Context and Temporal-Frequency Resolution Architecture for Probabilistic Energy Forecasting
- Masked Topology Modeling for Self-Supervised Learning on Parametric CAD
- FELT: Generating Tactile Signals from Vision for Visuo-Tactile Manipulation
- Attribution Markets: A Fisher-Market Formulation for Fractional Credit Assignment Between Planned Tasks and Performed Actions
- Evaluating Uncertainty and Quality of Visual Language Action-enabled Robots
- Towards Full Candidate Interaction: A Comprehensive Comparison Network for Better Route Recommendation
- MoA-Structured Decode Attention DNF Derivation, KV-Cache Accumulation, GQA/MQA, and OpenACC Kernel
- Equivariant Conditional Diffusion Model for Head and Neck CT Image Synthesis from CBCT
- Cognitive Dual-Process Planning for Autonomous Driving with Structured Scene Knowledge and Verifiable Reasoning-Action Consistency
- RPPNet: Perceptually-Grouped Rhythm-Pitch Primitives for Long-Term Structure Melody Generation via Boundary-Aware Modeling
- PhenSPINE: A Standardized Benchmark for Spine Pathology Diagnosis
- How Fast Can Reward Models Score? A Systems Study of C++ and PyTorch Inference Runtimes for RLHF
- Leaky Language Models: Stealing Architecture and Inference Optimizations via Per-Token Timing
- Self-Attention Transformer-Based Detector for Faster-than-Nyquist Signaling
- ChineseBERT: Chinese Pretraining Enhanced by Glyph and Pinyin Information
- Glyce: Glyph-vectors for Chinese Character Representations
- Pain in 3D: Generating Controllable Synthetic Faces for Automated Pain Assessment
- A Self-Supervised Framework for Space Object Behaviour Characterisation
- StructGen: Disambiguating Multi-Reference Image Generation via Structured Context Modeling
- HiCI: Hierarchical Construction-Integration for Long-Context Attention
- HyCoRec: Hypergraph-Enhanced Multi-Preference Learning for Alleviating Matthew Effect in Conversational Recommendation
- Unsupervised Keypoints for Real-Time Fall Detection: Comparative Analysis Under Real-world Conditions with Predictive Bandwidth Reduction
- Insecure Coding Preferences in Long-Term Memory: Security Risks for LLM-based Code Generation
- LFM: Leveraging Foundation Models for Source-Free Universal Domain Adaptation
- Beyond Objective Expressivity: Geometry Preservation in Multimodal Contrastive Learning
- LaCache: Exact Caching and Precision-Adaptive Inference for Diffusion Large Language Models
- Stochastic Thermodynamics for Autoregressive Generative Models: A Non-Markovian Perspective
- CompilerKV: Risk-Adaptive KV Compression via Offline Experience Compilation
- SKETCH: Semantic Key-Point Conditioning for Long-Horizon Vessel Trajectory Prediction
- Efficient Multi-round LLM Inference over Disaggregated Serving
- ImplicitRDP: An End-to-End Visual-Force Diffusion Policy with Structural Slow-Fast Learning
- Trustworthy Privacy-Preserving Multimodal Federated Learning for Personalised Breast Cancer Prediction
- Pathologist Attention-Aligned Report Generation for Prostate Histopathology
- MATANet: A Multi-context Attention and Taxonomy-Aware Network for Fine-Grained Underwater Recognition of Marine Species
- A Deep Reinforcement Learning Algorithm for the Vehicle Routing Problem with Stochastic Demands and Outsourcing
- Chebyshev Manifold Adaptation
- Occlusion-Aware Panoptic Segmentation with Joint Position Embedding and Occlusion-Level Attention
- Context-Masked Truncated Reasoning Audits for Answer-Key Dependence in LLM Tutors
- Geometric Capacity of Transformers: A Tropical Geometry Perspective
- JD-BP: A Joint-Decision Generative Framework for Auto-Bidding and Pricing
- Teaching Robots to Interpret Social Interactions through Lexically-guided Dynamic Graph Learning
- Zero-Shot DINOv3-Based Image Matching via Many-to-Many Association
- Information in Many-body Eigenstates: A Question of Learnability
- Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery
- CSAR: Containerized System Architecture for Robotics
- Robust Betatron-Tune Measurement from Schottky Spectra: Complementary Classical and Deep-Learning Paradigms
- Team RAS in 11th ABAW Competition: Multimodal Ambivalence Recognition Approach
- Beyond Noisy Signals: Dual-Level Denoising for Multi-modal Sequential Recommendation
- Moving Alphabet: A Controlled Study of Training Data for Text-to-Video Generation
- Algebraic Signatures for Structural Learning in Probability Tensors
- Mitigating Matthew Effect: Multi-Hypergraph Boosted Multi-Interest Self-Supervised Learning for Conversational Recommendation
- Norm or Direction? Decoding Vision Mambas for High-Resolution Vision
- Strategy-Following Multi-Agent Deep Reinforcement Learning Considering Control Strategies Provided to Other Agents
- Find Before You Fine-Tune: A Diagnostic Study of Small LLMs for Cybersecurity QA
- DeforM: Reasoning-Guided Physics-Aware Video Generation via Spatial-Temporal Masking
- Dual Attention Residuals
- LLM-Driven Cost-Effective Requirements Change Impact Analysis
- Think Sparse, Predict Dense: Continuous Thought Machines for Image Super-Resolution
- Breaking Feedback-Blindness: Utility-Augmented Transformer for Sequential Decision Making
- Visual Semantic Decoding of Electrocorticography from Video Stimuli using End-to-End Deep Learning
- AutoJourn: Multi-Perspective Summarisation, Bias Detection and Bias Neutralisation for LLM-Generated News in Automated Journalism
- Advancing Multimodal Fusion on Heterogeneous Medical Data with Hybrid Geometry Attention
- CGCE: Classifier-Guided Concept Erasure in Generative Models
- GUIDED Network-Agnostic Feature Initialization for Spatial Transferability in GNN-based Models
- Towards anomaly detection searches for new physics signatures including Higgs bosons with weakly supervised machine learning
- IBoxCLA: Towards Robust Box-supervised Segmentation of Polyp via Improved Box-dice and Contrastive Latent-anchors
- Can Interpretation Predict Behavior on Unseen Data?
- Aligned Stable Inpainting: Mitigating Unwanted Object Insertion and Preserving Color Consistency
- Retrieval-Augmented Interpretable Learning: Towards Task-Specific Zero-Shot Models in Healthcare
- An Adjoint-Sensitivity Framework for Lost-in-the-Middle Phenomena in Causal Residual Transformers
- LieBN: Batch Normalization over Lie Groups
- EmbeddedKittens: An Evaluation of Code Embeddings for Scratch
- Residual-Guided Multi-Resolution Refinement of Foundation Models: A Case Study in Drought Forecasting
- Can Transformers Really Do It All? On the Compatibility of Inductive Biases Across Tasks
- CoCurve: Cross-Module Co-Pruning Curvature for Training-Free Structured LLM Pruning
- Mambalaya: Einsum-Based Fusion Optimizations on State-Space Models
- Rethinking Heterogeneous LLM Merging: A Weighted Model Averaging Perspective
- FF-ProCams: Feed-Forward Gaussian Splatting for Projector-Camera System
- Lightweight Wrappers for Adapting Time Series Foundation Models to Regional Drought Forecasting
- Brain-Aligned Multi-Stream Video Transformers with Sparse Self-Selection
- DMSNet: Cross-Band Learning for Multi-Target Sensing in Multi-Band ISAC
- BMFA: Boundary-Minority Free-Energy Adaptive Screening
- Early Yield Prediction for Sugar Beet Fields using Satellite Data -- Learnings from Specialized Vision Transformers
- The Aura in the Machine: Genealogy and the Status of the Work of Art in the Generative Era
- GeneSpeak-FP: Target and Compound Retrieval from Observed Cell-Level Perturbation Signatures
- Neuro-Symbolic Meta-Policies for Temporal Knowledge-Graph Memory under Partial Observability
- Participatory provenance as representational auditing for AI-mediated public consultation
- Patch Policy: Efficient Embodied Control via Dense Visual Representations
- Don't Blame the Large Language Model: How Agent Harness Evolution Shapes Coding Agent Quality
- Mobius Learning: Cyclic Depth Folding in Transformers
- DirPA: Addressing Prior Shift in Imbalanced Few-shot Crop-type Classification
- EVOLVE: Efficient Learned Volume Compression with Variable-Rate Encoding on a Cross-Domain Database
- Entity-Relation Extraction as Multi-Turn Question Answering
- The Art of Not Forgetting
- Topological Signatures of Context-Level Reliability in TabPFN
- HAS: Highlight-guided Attention Steering for Multimodal LLM Video Summarization
- AInimation: Animating from Prompt to AI-Generated Responses
- Capability from Access Structure, Not Scale: Lower Bounds and Pre-Registered Tests for Hybrid Sequence Models
- An Analysis of Residual-Stream Geometry Across Transformer Depth
- Vectorizing the Trie: Efficient Constrained Decoding for LLM-based Generative Retrieval on Accelerators
- Higher-Order Cell Tracking Transformer
- FedDP-PALD: A Privacy-Preserving Federated Latent Diffusion Framework with Prototype Aggregation for Medical Data Synthesis
- Assisting or resisting patriarchy? a critical discourse analysis of chatgpt’s responses on feminism
- Position: Stop Reactively Patching Your Model Every Time and Start Proactive Test-Driven AI Development
- De+e-ffusion: Capturing the Beam-Beam Physics of e+e- Collisions with Diffusion Models
- Censoring-Aware In-Context Learning for Generalized Supplier Lead Time Estimation in Supply Chain Planning
- Opportunistic Lower-Terahertz Rainfall Estimation with DSD-Constrained Channel Characterization
- Decode-Time Grammars: Constrained LLM Generation over a Refinement Order of Grammar Fragments
- A Controlled Study of Attention-Only Transformers
- Perturbation is All You Need for Extrapolating Language Models
- Beyond Parallel Tracking: Interactive Multi-Feature Fusion Drives Semantic Reconstruction from Non-invasive Brain Recordings
- When Fewer Layers Break More Chains: Layer Pruning Harms Test-Time Scaling in LLMs
- BanClickThumb: A Multimodal Dataset and Transformer Fusion Benchmarks for Clickbait Detection in Bengali YouTube Videos
- Artificially intelligent agents in the social and behavioral sciences: A history and outlook
- Kernelized Linear Attention: Breaking the Capacity Wall with Symmetric Cones
- The Geometry of Semantic Space: A Continuous Geometric Framework for the Transformer Architecture
- Learning MMSE Filters for OFDM Channel Estimation: Attention Transformer Gains at Linear Inference
- Explaining and Tuning Transformer-based LLMs in Arithmetic Tasks with Human Strategies
- Reward-Driven LLM Agent Workflows: Synthesizing POMDP Routing and Self-Correction for Autonomous Decision-Making
- Dynamic Interaction-Aware and Causality-Disentangled Framework for Multimodal Sentiment Analysis
- Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins
- Overlapping Schwarz Attention: Hierarchical Attention via Domain Decomposition
- Every Component is a Lookup: Token Attribution and Composition from a Single Decomposition
- Feasibility-Aware Security-Constrained Unit Commitment via Hybrid Soft Actor-Critic with Quantum-Sampled Features
- A Diagnostic Framework for AI Agent Behavior
- Cross-Coordinate Correspondence Pruning for Image-to-Point Cloud Registration
- WHALE: A Scalable Unified Model for Recommendation with Wukong-HSTU Architecture
- Cognitive-YOLO: LLM-Driven Architecture Synthesis from First Principles of Data for Object Detection
- Node4All: Learning Node Representation Beyond Datasets
- UPLiFT: Efficient Pixel-Dense Feature Upsampling with Local Attenders
- SubdivAR: Autoregressive Next-Scale Prediction for Neural Mesh Subdivision
- Gradient-Free Privacy Leakage in Federated Language Models through Selective Weight Tampering
- KEPLA: A Knowledge-Enhanced Deep Learning Framework for Accurate Protein-Ligand Binding Affinity Prediction
- Hybrid Mamba-Attention Neural Architecture for Channel Estimation
- DynImmune-BERT: Dynamic Immune Repertoire Modeling with Neural ODE Driven Continuous Transformers
- ChemFusion: A Multimodal Cross-Attention Network for Reaction Yield Prediction
- Measuring and Evaluating the Performance of Generative AI Models for Scam Detection
- Transition-Aware Backend Dispatch for Edge LLM Inference
- Attentions Under the Microscope: A Comparative Study of Resource Utilization for Variants of Self-Attention
- ReactiveBFM: Reactive Closed-Loop Motion Planning Towards Universal Humanoid Whole-Body Control
- RoVE: Rotary Value Embeddings Attention for Relative Position-dependent Value Pathways
- LLM-Guided Task-Semantic Field Factorization for Industrial Process Forecasting
- QuReC: All-in-One Image Restoration with Query-Specific Guidance and Local-Global Response Calibration
- Data-driven Machine Learning Cannot Reach Symbolic-level Logical Reasoning -- The Limit of the Scaling Law
- Do Value Vectors in Deep Layers Need Context from the Residual Stream?
- Phantom Transitions in Language Model Fine-Tuning: A Density-Matrix Analysis
- HE-LRM: Encrypted Deep Learning Recommendation Models using Fully Homomorphic Encryption
- OmniMotion-X: Versatile Multimodal Whole-Body Motion Generation
- PI-H2T: Enhancing Long-Tailed Visual Recognition with Permutation-Invariant and Head-to-Tail Feature Fusion
- Machine Learning for Electrode Materials: Property Prediction via Composition
- NABEATs: Noise-Aware Audio Representation Learning
- NIRVANA: Structured Pruning Reimagined for Large Language Model Compression
- Supervised Reward Inference
- DRIFT: Difficulty-aware Rectified Flows for Through-plane MRI Super-Resolution
- InertialAR: Autoregressive 3D Molecule Generation with Inertial Frames
- Visual Access Boundaries in Vision-Language Model Reasoning
- Unsupervised Multimodal Clustering for Semantics Discovery in Multimodal Utterances
- Points as Tori: Fast Pointwise Signed Distance for Point Clouds
- Time-Aware Prior Fitted Networks for Zero-Shot Forecasting with Exogenous Variables
- NanoZK: Privacy-Preserving Verifiable Inference for Large Language Models via Layerwise Zero-Knowledge Proofs
- Dropout and Random Gradient Masking Are Asymptotically Equivalent in Large ResNets
- Hybrid Machine Learning for Articulation Angle Estimation of Truck-Semitrailer Combinations
- Curvature-Adaptive Consistency Flow Matching: Autonomous Trajectory Optimization via Reinforcement Learning
- Look Clearly Before Answering: Mitigating Hallucinations in LVLMs via Saliency-Driven Perceptual Realignment
- Hierarchical Wireless Foundation Model for Multi-Task Optimization
- DORS: Dynamic Attention Routing for Diffusion-based Object Removal in Dense Scenes
- Dependency-Guided Code Generation: Structured Matrix Decomposition and Consistency-Guided Refinement
- A Benchmark for Electrical Load Forecasting Across Grid Levels: Time-Series Transformers Outperform Established Methods
- WREN: Low Light Image Enhancement Using Retinex theory-based Double U-Net-like Structures
- HyBDM: Multi-Scale Hybrid Experts for Time Series Forecasting with Bidirectional Dependency Modeling
- Distributed solar generation forecasting using attention-based deep neural networks for cloud movement prediction
- TVGL-CFM:Generating and Forecasting Time-Varying Trajectories of Dynamic Networks with Conditional Flow Matching
- Spectral Heat Flow for Conservative Token Condensation in Vision-Language Models
- Efficient Audio-Visual Event Recognition via Knowledge Distillation and Dynamic INT8 Quantization of a Hybrid Cross-Attention Network
- A JoLT for the KV Cache: Near-Lossless KV Cache Compression via Joint Tucker and JL-Residual Allocation for LLMs
- Code-Poisoning Property Inference Attacks
- Posts of Peril: Detecting Information About Hazards in Text
- DELUGE: Towards Continental-Scale Daily Pluvial Flood Damage Prediction via Interpretable Conditioning on Foundation Model Embeddings
- A Blueprint for Equilibrium-Based Differentiable Continuous-Variable Thermodynamic Computing
- RTL-Sequencer: Towards Scalable RTL Timing Prediction with the Sequence-based Paradigm
- Spatio-Temporal Prediction of Unsteady Airfoil Aerodynamics Using Augmented Graph Neural Ordinary Differential Equations with Exogenous Controls
- An expressivity analysis of hierarchical modelling in deep transformers via bounded-depth grammars
- GeoRouteNet: A Geometry-Aware Non-Autoregressive Neural Solver for the Euclidean Traveling Salesman Problem
- Knowledge-Assisted Multi-Graph Dependency Learning for Multivariate Time Series Anomaly Detection in Multi-Stage Industrial Processes
- Cost-efficient generative AI summarization for scalable automated essay scoring in educational assessment
- HCIG: A Hierarchical Cross-Modal Incongruity Graph Network for Multimodal Sarcasm and Cyberbullying Detection
- Kolmogorov--Arnold Networks for Small Language Models
- T2MLR: Transformer with Temporal Middle-Layer Recurrence
- CaloTrilogy: Toward a Breakthrough in One-Step, End-to-End, Physics-Guided Shower Generation for Modern Calorimeters
- Measurement of the branching ratio of the K+→π+νν decay
- Learning Standard Model structure from LHC data with Riemannian flow matching
- Debiasing Text-to-Image Evaluation via Implicit Cultural Alignment Reward Modeling
- In-context learning of closed form solution to simple linear regression task using transformer with linear self-attention
- Decoupled Alignment for Robust Plug-and-Play Adaptation
- Self-Attention as Transport: Limits of Symmetric Spectral Diagnostics
- Stop the Sampler! Classifier-Based Adaptive Stopping for Sampling Kernels
- Geometry-Enhanced Portion Estimation for Multimodal LLMs
- Auto-Scaling Heterogeneous Neural Processing Units for Energy and Cost-Efficient LLM Serving
- Hierarchical Domain Generalization
- Mechanistic Interpretability of Cognitive Complexity in LLMs via Linear Probing using Bloom's Taxonomy
- EgoExoMoCap: Distributed Ego-Exo Human Motion Capture
- Agentic Calibration of Grey-Box Simulation Models: An LLM-Driven Alternative
- Rate-Utility Frontiers for Language Encodings: Comparing Tokens, Bytes, and Pixels Under Controlled Linguistic Content
- Song Aesthetics Evaluation with Multi-Stem Attention and Hierarchical Uncertainty Modeling
- Encoding EEG Signals to Examine Human-Like Next-Word Prediction Behaviour in Language Models
- (MPO)2: Multivariate Polynomial Optimization based on Matrix Product Operators
- Day-Ahead Forecasting of Largest Single Infeed/Outfeed on the Irish Power Grid: A Generative Artificial Intelligence Approach
- An MLIR-Based Compilation Method for Large Language Models
- Exo2EgoPose: Leveraging Exocentric Demonstrations for Vision-Language guided Egocentric 3D Hand Pose Forecasting
- Knowledge-Guided Cross-Modal Fusion for Adult-to-Pediatric ECG Transfer via Label-Conditioned Contrastive Alignment
- Language Identification via Compositional Data Analysis: A Linear-Time Classifier Based on Log-Ratio Geometry
- Concept-Guided Spatial Regularization for World Models in Atari Pong
- Language as a Wave Phenomenon: Semantic Phase Locking and Interference in Neural Networks
- Sample Is Feature: Beyond Item-Level, Toward Sample-Level Tokens for Unified Large Recommender Models
- VIABench: A Comprehensive Video Benchmark Collected from Blind Individuals for Visual Impairment Assistance
- Variational Inference for Bird's Eye View Segmentation in Autonomous Driving
- LLM4EHR: Aligning Clinical Time Series with Medical Event Sequences via Large Language Models
- 3D Geometric Tooth Alignment Planning via Deep Reinforcement Learning
- The CRAFT principles for the responsible use of large language models in policymaking
- MxGPS: Multiplex Graph Transformers for a Power Grid Foundation Model
- Implicit Virtual Leader: Decentralized Vision-Only Relative Pose Estimation for Multi-Robot Formations
- EduGuard: A Safe RAG-Based LLM Tutor for Programming Education
- Robust Filter Attention: Self-Attention as Precision-Weighted State Estimation
- Online Neural Space Time Memory for Dynamic Novel View Synthesis
- AI-Conducted Interviews in Empirical Software Engineering: An Experience Report
- Probabilistic Physics-Informed Neural Networks for Estimating Heterogeneous Elastic Properties from Low-Resolution and Noisy Displacement Data
- Sharp Stability Threshold and Certification for Designing Stable Residual Architectures
- Latent Trajectory Discrimination for AI-Generated Text Detection
- Harnessing Machine Learning for Hybrid Constitutive Modelling of Viscoelastic Fluid Flows in Computational Rheology
- Expanding the Lexicon of Ge'ez Based African Languages: A Comparative Study of Amharic and Tigrinya
- MVFusion-GS: Motion-Variance Guided Temporal Attention for High-Quality Dynamic Gaussian Splatting
- The RG-Flow Transformer: Encoding Scale-Free Dynamics in Scarce EEG
- AHEAD: Anticipatory Hand-Driven Teleoperation via Human Intent Prediction
- NIFA: Nonlinear IMC enhanced FPGA for efficient ML inference
- Multimodal Semantic-Aware Contrastive Learning For False Negative Mitigation in 3D Medical Imaging
- VAN-AD: Visual Masked Autoencoder with Normalizing Flow For Time Series Anomaly Detection
- Bifocal Attention: Harmonizing Geometric and Spectral Positional Embeddings for Algorithmic Generalization
- Converting T1-weighted MRI from 3T to 7T quality using deep learning
- PhasorFlow: A Python Library for Unit Circle Based Computing
- Introspective Attention Modulation for Safe Text-to-Image Generation
- GS-RealBlur: A Flexible Data Acquisition Framework for Real-World Image Deblurring
- AI Trading: Evaluating Large Language Models for Technical Market Analysis
- xHC: Expanded Hyper-Connections
- Robust Explanations for User Trust in Enterprise NLP Systems
- VarRate: Training-Free Variable-Rate KV Cache Compression for Long-Context LLMs
- Physics-Informed Diffusion for Biomechanically Plausible 3D Sign Language Generation
- FlashDecoder: Real-Time Latent-to-Pixel Streaming Decoder with Transformers
- Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning
- Reflex: Real-Time VLA Control through Streaming Inference
- Does generative AI supersede supervised XMLC? A Benchmark Study on Automated Subject Indexing with German Scientific Literature
- SEMA: a Scalable and Efficient Mamba like Attention via Token Localization and Averaging
- Automated Trading System for Straddle-Option Based on Deep Q-Learning
- NeuronSoup: Evolving Asynchronous, Shared-Neuron Temporal Graphs without Backpropagation
- Attention Degradation, Function Token Anchoring, and the Limits of Attention-Based Intervention in Large Language Models
- Do Active SAE Feature Planes Carry More Holonomy? A Preregistered Reversal in Gemma
- Structured Latent Space Modeling over Multi-Scale Temporal Patches for Multivariate Time Series Forecasting
- TEDDY: A Pediatric Foundation Model for Risk Forewarning from ICD-Coded Diagnostic Histories
- Reproducing Recurrent Transformers: The CoTFormer
- Gemma 4 Technical Report
- Large Language Model-Enhanced Multi-hop Parallel Image Semantic Communication
- Advancing bioinformatics with language models: components, applications, and perspectives
- Self-Attention Dynamics with Rotary Position Embeddings: Twisted States and Explicit Consensus Rates on the Sphere
- Native Multi-Dimensional Subquadratic Operators via Input Dependent Long Convolutions
- STN-TGAT: Top-K Portfolio Construction via Prior-Guided Graph Attention with Learnable Soft-Threshold Sparsification
- Edge-Efficient Transformer for End-to-End RF Spectrum Monitoring
- Generating consensus and dissent on massive discussion platforms with a semantic-vector model
- Structure of the Circular-Dyadic Convolution Error
- Beyond Output-Space Calibration: Spectral Evidence Bundling for Selective Reliability Estimation in Time-Series Classification
- Physics-driven local-whole elastic deformation modeling for point cloud representation learning
- High-accuracy Low-Bit KV-Cache Quantization via Local Distribution Restoration
- Token-Level Cross-Modal Transformer with Contrastive Multi-Task Learning for Breast Cancer Subtype Classification and Survival Prediction
- Efficient and Training-Free Single-Image Diffusion Models
- The Failures of Marginal Influence-Based Attribution Methods for Global Time Series Explanations
- Workload-Aware Caching for Multi-Agent Systems
- Explaining Attention with Program Synthesis
- The Market in the Model: Latent Diffusion as Neural Economy
- Scaling quantum machine learning without tricks: full-resolution and diverse image generation
- AlphaOracle: Oracle bone script decipherment via human-workflow-inspired deep learning
- Trajectory-Aware Clinical Risk Prediction via Severity-Grounded Knowledge Graphs and Retrieval-Augmented Generation
- RelGT-AC: A Relational Graph Transformer for Autocomplete Tasks in Relational Databases
- Do Transformers Need Three Projections? Systematic Study of QKV Variants
- LISA: Linear-Indexed Sparse Attention for Efficient Long-Context Reasoning
- FBLayout: Optimizing Memory Layout for Efficient LLM Finetuning on Mobile GPUs
- Rotary GPU: Exploring Local Execution Paths for Large Mixture-of-Experts Models Under Limited GPU Memory
- MATNet: multi-level fusion transformer-based model for day-ahead PV generation forecasting
- Do Language Models Need Sleep? Offline Recurrence for Improved Online Inference
- Forgetting in Language Models: Capacity, Optimization, and Self-Generated Replay
- What Makes Linguistic Representations Good Models of High-Level Visual Perception in the Human Brain?
- A Systematic Survey on Image Description Techniques for STEM Domains
- Represented Is Not Computed: A Causal Test of Candidate Algorithmic Intermediates in a Transformer
- Collocational bootstrapping: A hypothesis about the learning of subject-verb agreement in humans and neural networks
- Multi-Stream LLMs: Unblocking Language Models with Parallel Streams of Thoughts, Inputs and Outputs
- A Survey on <scp>GNN</scp> ‐Based Link Prediction: Techniques, Applications, and Challenges
- Toward manifest relationality in transformers via symmetry reduction
- Semantic Field Theory: Historical Origin, Higher-Order Interaction, and Stabilized Semantic Inference
- Domyn-Small: A European 10B Reasoning Language Model
- Beyond Accuracy and Cost: Latency-Aware LLM Query Routing for Dynamic Workloads
- Knowledge Injection Exists in MoE? Exploring Expert-Aware Contrast Decoding in MoE for Mitigating LLMs'Hallucinations
- Non-Markovianity and memory enhancement in quantum reservoir computing
- FineServe: A Fine-Grained Dataset and Characterization of Global LLM Serving Workloads
- Attention to Mamba: A Recipe for Cross-Architecture Distillation
- CPGRec+: A Balance-oriented Framework for Personalized Video Game Recommendations
- "AI Psychosis" in Context: How Conversation History Shapes LLM Responses to Delusional Beliefs
- PR-CAD: Progressive Refinement for Unified Controllable and Faithful Text-to-CAD Generation with Large Language Models
- Convergent Evolution: How Different Language Models Learn Similar Number Representations
- PoM: A Linear-Time Replacement for Attention with the Polynomial Mixer
- Don't Waste Bits! Adaptive KV-Cache Quantization for Lightweight On-Device LLMs
- Temporal structure of the language hierarchy within small cortical patches
- Effective Distillation to Hybrid xLSTM Architectures
- Mamba-3: Improved Sequence Modeling using State Space Principles
- From transcription to semantic corpus analysis: unsupervised learning of sentence representations for ancient languages
- M2RNN: Non-Linear RNNs with Matrix-Valued States for Scalable Language Modeling
- Tiny neural networks for multi-object tracking in a modular Kalman framework
- ARC-AGI-3: A New Challenge for Frontier Agentic Intelligence
- Sparser, Faster, Lighter Transformer Language Models
- GraphDiffMed: Knowledge-Informed Differential Attention with Pharmacological Graph Priors for Medication Recommendation
- Why the Brain Consolidates: Predictive Forgetting for Optimal Generalisation
- Inhibitory normalization of error signals improves learning in neural circuits
- Transformers are Bayesian Networks
- Semantic Novelty Trajectories in 80,000 Books: A Cross-Corpus Embedding Analysis
- Position: Modular Memory is the Key to Continual Learning Agents
- Generative AI & Fictionality: How Novels Power Large Language Models
- Transformers for dynamical systems learn transfer operators in-context
- Scalable Explainability-as-a-Service (XaaS) for Edge AI Systems
- Memory Caching: RNNs with Growing Memory
- Genomic perplexity and the evolution of context-dependent function
- FMPose3D: monocular 3D pose estimation via flow matching
- Spelling Bee Embeddings for Language Modeling
- Strategies for Span Labeling with Large Language Models
- Do Models Hear Like Us? Probing the Representational Alignment of Audio LLMs and Naturalistic EEG
- A Unified View of Attention and Residual Sinks: Outlier-Driven Rescaling is Essential for Transformer Training
- BreastDCEDL: A standardized deep learning-ready breast DCE-MRI dataset of 2,070 patients
- Remapping and navigation of an embedding space via error minimization: a fundamental organizational principle of cognition in natural and artificial systems
- 2026 International Conference on 3D Vision (3DV)
- Generative Autonomous Grid Control: Integrating Decision Transformers with a Two-Stage Safety Stack
- Chameleon: A Multiplier-Free Temporal Convolutional Network Accelerator for End-to-End Few-Shot and Continual Learning from Sequential Data
- Benchmarking for practice: Few-shot time-series crop-type classification on the EuroCropsML dataset
- Multi-Agent Inverted Transformer for Flight Trajectory Prediction
- mHC: Manifold-Constrained Hyper-Connections
- LLM-Based Agentic Systems for Software Engineering: Challenges and Opportunities
- Shared sensitivity to data distribution during learning in humans and transformer networks
- Compressing What Matters: Neuron Importance Meets Data-Aware Low Rank Approximation for Language Model Compression
- Universal Reasoning Model
- RePo: Language Models with Context Re-Positioning
- On the Approximation of Phylogenetic Distance Functions by Artificial Neural Networks
- The Mean-Field Dynamics of Transformers
- Iterative Compositional Data Generation for Robot Control
- PHOTON: Hierarchical Autoregressive Modeling for Lightspeed and Memory-Efficient Language Generation
- Epistemological Fault Lines Between Human and Artificial Intelligence
- Evaluating Pretrained Protein Language Model Embeddings as Proxies for Functional Similarity
- Reduced order modeling with shallow recurrent decoder networks
- Professional Software Developers Don't Vibe, They Control: AI Agent Use for Coding in 2025
- Extending the Context of Pretrained LLMs by Dropping Their Positional Embeddings
- The 4/δ Bound: Designing Predictable LLM-Verifier Systems for Formal Method Guarantee
- From Attention to Frequency: Integration of Vision Transformer and FFT-ReLU for Enhanced Image Deblurring
- Artificial intelligence for risk assessment and outcome prediction in malignant haematology
- LLMs can hide text in other text of the same length
- Ten Simple Rules for AI-Assisted Coding in Science
- Do LLMs Truly “Understand” when a Precedent Is Overruled?
- Butter-Bench: Evaluating LLM Controlled Robots for Practical Intelligence
- Kinaema: a recurrent sequence model for memory and pose in motion
- Effectiveness of LLMs in Temporal User Profiling for Recommendation
- Kimi Linear: An Expressive, Efficient Attention Architecture
- Cosine capital: Large language models and the embedding of all things
- Every Language Model Has a Forgery-Resistant Signature
- To model human linguistic prediction, make LLMs less superhuman
- Expected Attention: KV Cache Compression by Estimating Attention from Future Queries Distribution
- Transformers are Inherently Succinct
- StutterZero and StutterFormer: End-to-End Speech Conversion for Stuttering Transcription and Correction
- The Impossibility of Inverse Permutation Learning in Transformer Models
- ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference
- Artificial Phantasia: Emergent Mental Imagery in Large Language Models
- Extract-0: A Specialized Language Model for Document Information Extraction
- LLaDA-MoE: A Sparse MoE Diffusion Language Model
- Platonic Transformers: A Solid Choice For Equivariance
- FlyTrap: Physical Distance-Pulling Attack Towards Camera-based Autonomous Target Tracking Systems
- SimpleFold: Folding Proteins is Simpler than You Think
- scPortrait integrates single-cell images into multimodal modeling
- Latent learning: episodic memory complements parametric learning by enabling flexible reuse of experiences
- Decoupling the "What" and "Where" With Polar Coordinate Positional Embeddings
- Towards a Physics Foundation Model
- Invisible Ears at Your Fingertips: Acoustic Eavesdropping via Mouse Sensors
- AISSISTANT: Human-AI Collaborative Review and Perspective Research Workflows in Data Science
- Toward comprehensive cellular characterization of H&E slides
- Flashzoi: an enhanced Borzoi for accelerated genomic analysis
- Biophysics-based protein language models for protein engineering
- uGMM-NN: Univariate Gaussian Mixture Model Neural Network
- If generative AI is the answer, what is the question?
- SHFormer: Dynamic spectral filtering convolutional neural network and high-pass kernel generation transformer for adaptive MRI reconstruction
- Arnold: a generalist muscle transformer policy
- Jet-Nemotron: Efficient Language Model with Post Neural Architecture Search
- From basic affordances to symbolic thought: A computational phylogenesis of biological intelligence.
- Universal Learning of Nonlinear Dynamics
- A Survey on Diffusion Language Models
- Do Language Models Agree with Human Perceptions of Suspense in Stories?
- Fast weight programming and linear transformers: from machine learning to neurobiology
- Whither symbols in the era of advanced neural networks?
- Cameras as Relative Positional Encoding
- Markov Chain Estimation with In-Context Learning
- The wall confronting large language models
- Technological folie à deux: Feedback Loops Between AI Chatbots and Mental Illness
- AlphaGo Moment for Model Architecture Discovery
- Learning without training: The implicit dynamics of in-context learning
- A Sitewise Model of Natural Selection on Individual Antibodies via a Transformer–Encoder
- Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
- What Neuroscience Can Teach AI About Learning in Continuously Changing Environments
- Dynamic Chunking for End-to-End Hierarchical Sequence Modeling
- Position: We Need An Algorithmic Understanding of Generative AI
- Hallucination Stations: On Some Basic Limitations of Transformer-Based Language Models
- Small Batch Size Training for Language Models: When Vanilla SGD Works, and Why Gradient Accumulation Is Wasteful
- Preemptive Solving of Future Problems: Multitask Preplay in Humans and Machines
- Beyond Sensor Data: Foundation Models of Behavioral Data from Wearables Improve Health Predictions
- Adopting a human developmental visual diet yields robust, shape-based AI vision
- Fast and Simplex: 2-Simplicial Attention in Triton
- Hierarchical Reasoning Model
- Universal pre-training by iterated random computation
- Who Does What in Deep Learning? Multidimensional Game-Theoretic Attribution of Function of Neural Units
- From Memories to Maps: Mechanisms of In-Context Reinforcement Learning in Transformers
- Anwendungen, Herausforderungen und ein vertrauenswürdiger Umgang mit künstlicher Intelligenz im Bereich Public Health
- Log-Normal Multiplicative Dynamics for Stable Low-Precision Training of Large Networks
- Theoretical Analysis of Positional Encodings in Transformer Models: Impact on Expressiveness and Generalization
- Misinformation by Omission: The Need for More Environmental Transparency in AI
- POCO: Scalable Neural Forecasting through Population Conditioning
- Mercury: Ultra-Fast Language Models Based on Diffusion
- Tensor Manipulation Unit (TMU): Reconfigurable, Near-Memory Tensor Manipulation for High-Throughput AI SoC
- Leveraging Natural Language Processing to Unravel the Mystery of Life: A Review of NLP Approaches in Genomics, Transcriptomics, and Proteomics
- Self-supervised learning of molecular representations from millions of tandem mass spectra using DreaMS
- FrontierCO: Real-World and Large-Scale Evaluation of Machine Learning Solvers for Combinatorial Optimization
- Transformers are Graph Neural Networks
- AbsenceBench: Language Models Can't Tell What's Missing
- Blending Complementary Memory Systems in Hybrid Quadratic-Linear Transformers
- ATLAS: Learning to Optimally Memorize the Context at Test Time
- Stochastic Chameleons: Irrelevant Context Hallucinations Reveal Class-Based (Mis)Generalization in LLMs
- Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decoding
- QuARI: Query Adaptive Retrieval Improvement
- Parkour in the Wild: Learning a General and Extensible Agile Locomotion Policy Using Multi-expert Distillation and RL Fine-tuning
- Log-Linear Attention
- Relational reasoning and inductive bias in transformers and large language models
- Leaner Transformers: More Heads, Less Depth
- DataRater: Meta-Learned Dataset Curation
- Gradient-Based Program Repair: Fixing Bugs in Continuous Program Spaces
- Breaking Quadratic Barriers: A Non-Attention LLM for Ultra-Long Context Horizons
- CaRL: Learning Scalable Planning Policies with Simple Rewards
- Paper2Code: Automating Code Generation from Scientific Papers in Machine Learning
- Fourier-mixed window attention for efficient and robust long sequence time-series forecasting
- Reasoning by Superposition: A Theoretical Perspective on Chain of Continuous Thought
- Can LLMs Credibly Transform the Creation of Panel Data from Diverse Historical Tables?
- Questioning Representational Optimism in Deep Learning: The Fractured Entangled Representation Hypothesis
- Beyond Attention: Toward Machines with Intrinsic Higher Mental States
- Coral Protocol: Open Infrastructure Connecting The Internet of Agents
- Perception Encoder: The best visual embeddings are not at the output of the network
- BitNet b1.58 2B4T Technical Report
- Visual Language Models show widespread visual deficits on neuropsychological tests
- NNN: Next-Generation Neural Networks for Marketing Measurement
- Mixture of Experts Made Intrinsically Interpretable
- Dynamic Markov Blanket Detection for Macroscopic Physics Discovery
- Optimizing Biophysical Large-Scale Brain Circuit Models With Deep Neural Networks
- RANa: Retrieval-Augmented Navigation
- LLM Social Simulations Are a Promising Research Method
- Multi-Token Attention
- Layers at Similar Depths Generate Similar Activations Across LLM Architectures
- NeuRaLaTeX: A machine learning library written in pure LaTeX
- TransMamba: A Sequence-Level Hybrid Transformer-Mamba Language Model
- What Is Artificial General Intelligence?
- VGGT: Visual Geometry Grounded Transformer
- Transformers without Normalization
- Probing LLMs for Multilingual Discourse Generalization Through a Unified Label Set
- ZSMerge: Zero-Shot KV Cache Compression for Memory-Efficient Long-Context LLMs
- MatplotAlt: A Python Library for Adding Alt Text to Matplotlib Figures in Computational Notebooks
- Palatable Conceptions of Disembodied Being: Terra Incognita in the Space of Possible Minds
- SPECTRE: An FFT-Based Efficient Drop-In Replacement to Self-Attention for Long Contexts
- Erwin: A Tree-based Hierarchical Transformer for Large-scale Physical Systems
- Large Language Diffusion Models
- Foundation neural-networks quantum states as a unified Ansatz for multiple hamiltonians
- Inverse problems with experiment-guided AlphaFold
- InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU
- DUNE: Distilling a Universal Encoder from Heterogeneous 2D and 3D Teachers
- TLOB: A Novel Transformer Model with Dual Attention for Price Trend Prediction with Limit Order Book Data
- TransMLA: Multi-Head Latent Attention Is All You Need
- From Thought to Action: How a Hierarchy of Neural Dynamics Supports Language Production
- Memory Instance Gated Transformer Reinforcement Learning for Portfolio Management
- Adaptive fusion of multi-modal remote sensing data for optimal sub-field crop yield prediction
- Extending the RANGE of Graph Neural Networks: Relaying Attention Nodes for Global Encoding
- Smaller But Better: Unifying Layout Generation with Smaller Large Language Models
- Matryoshka Quantization
- Infinite Retrieval: Attention Enhanced LLMs in Long-Context Processing
- Idiosyncrasies in Large Language Models
- LLMs as a synthesis between symbolic and distributed approaches to language
- Block Diffusion: Interpolating Between Autoregressive and Diffusion Language Models
- Protein Structure Tokenization: Benchmarking and New Recipe
- Training Large Neural Networks With Low-Dimensional Error Feedback
- Shades of Zero: Distinguishing Impossibility from Inconceivability
- Emergent Symbolic Mechanisms Support Abstract Reasoning in Large Language Models
- Teasing Apart Architecture and Initial Weights as Sources of Inductive Bias in Neural Networks
- Language Models Grow Less Humanlike beyond Phase Transition
- Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention
- CalibQuant: 1-Bit KV Cache Quantization for Multimodal LLMs
- High-Fidelity Simultaneous Speech-To-Speech Translation
- The Cake that is Intelligence and Who Gets to Bake it: An AI Analogy and its Implications for Participation
- Training Dynamics of In-Context Learning in Linear Attention
- VLMaterial: Procedural Material Generation with Large Vision-Language Models
- Self-Improving Transformers Overcome Easy-to-Hard and Length Generalization Challenges
- MambaGlue: Fast and Robust Local Feature Matching With Mamba
- Fast3R: Towards 3D Reconstruction of 1000+ Images in One Forward Pass
- Topological constraints on self-organization in locally interacting systems
- Evolution and The Knightian Blindspot of Machine Learning
- Decoding-based Regression
- The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training
- MedicoSAM: Robust Improvement of SAM for Medical Imaging
- Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation
- Context-Aware Token Pruning and Discriminative Selective Attention for Transformer Tracking
- Vision-Enhanced Large Language Models for High-Resolution Image Synthesis and Multimodal Data Interpretation
- Accelerated Multi-Modal Motion Planning Using Context-Conditioned Diffusion Models
- Context-Selective State Space Models: Feedback is All You Need
- On the accuracy of implicit neural representations for cardiovascular anatomies and hemodynamic fields
- Exploring Depth Generalization in Large Language Models for Solving Recursive Logic Tasks
- Unison: A Fully Automatic, Task-Universal, and Low-Cost Framework for Unified Understanding and Generation
- On Conditional Stochastic Interpolation for Generative Nonlinear Sufficient Dimension Reduction
- Utilizing Multi-Agent Reinforcement Learning with Encoder-Decoder Architecture Agents to Identify Optimal Resection Location in Glioblastoma Multiforme Patients
- Time Series Foundation Models for Process Model Forecasting
- Empirical Investigation of the Impact of Phase Information on Fault Diagnosis of Rotating Machinery
- Omni-Attribute: Open-vocabulary Attribute Encoder for Visual Concept Personalization
- The Appeal and Reality of Recycling LoRAs with Adaptive Merging
- When Machines Get It Wrong: Large Language Models Perpetuate Autism Myths More Than Humans Do
- Exploring the integration of large language models in industrial test maintenance processes
- Improving IR-based bug localization with semantics-driven query reduction
- Coupling Experts and Routers in Mixture-of-Experts via an Auxiliary Loss
- Do You Have Freestyle? Expressive Humanoid Locomotion via Audio Control
- Image Denoising Using Global and Local Circulant Representation
- Measuring complex constructs in large-scale text with computational social mixed methods
- Reinforcement Learning via Self-Distillation
- Foundation models for electrocardiogram interpretation: clinical implications
- SC-Net: Robust Correspondence Learning via Spatial and Cross-Channel Context
- Osmotic Learning: A Self-Supervised Paradigm for Decentralized Contextual Data Representation
- Towards end-to-end automation of AI research
- A new adaptive two-layer model for opinion spread in hypergraphs: parameter sensitivity and estimation
- ECG-RAMBA: Zero-Shot ECG Generalization by Morphology-Rhythm Disentanglement and Long-Range Modeling
- Deep learning for pedestrians: backpropagation in Transformers
- PCR-ORB: Enhanced ORB-SLAM3 with Point Cloud Refinement Using Deep Learning-Based Dynamic Object Filtering
- Plug-and-Play Fidelity Optimization for Diffusion Transformer Acceleration via Cumulative Error Minimization
- Robust LLM-based Column Type Annotation via Prompt Augmentation with LoRA Tuning
- Text-Routed Sparse Mixture-of-Experts Model with Explanation and Temporal Alignment for Multi-Modal Sentiment Analysis
- PGOT: A Physics-Geometry Operator Transformer for Complex PDEs
- From Model Choice to Model Belief: Establishing a New Measure for LLM-Based Research
- Taking time seriously: Predicting conflict fatalities using temporal fusion transformers
- Computational structuralism: Toward a formal theory of meaning in the age of digital intelligence
- When science meets geopolitics: global AI research network transformation (2000–2025)
- MAPLE: Efficient and Diverse Multi-Alpha Generation for Portfolio Construction
- Reservoir Computing inspired Matrix Multiplication-free Language Model
- Graph Neural Networks with Transformer Fusion of Brain Connectivity Dynamics and Tabular Data for Forecasting Future Tobacco Use
- Breaking the Memory Wall: Exact Analytical Differentiation via Tiled Operator-Space Evolution
- An Architecture-Led Hybrid Report on Body Language Detection Project
- With Great Context Comes Great Prediction Power: Classifying Objects via Geo-Semantic Scene Graphs
- JADAI: Jointly Amortizing Adaptive Design and Bayesian Inference
- Long short-term memory network with adapted attention mechanism for credit risk modeling
- ConDiff: Conditional graph diffusion model for recommendation
- Fusion or Confusion? Multimodal Complexity Is Not All You Need
- Fast Organic Crystal Structure Prediction with Unit Cell Flow Matching
- Revisiting [CLS] and Patch Token Interaction in Vision Transformers
- Argus: Token Aware Distributed LLM Inference Optimization
- Towards Understanding Steering Strength
- A Brain-like Synergistic Core in LLMs Drives Behaviour and Learning
- Stream-DiffVSR: Low-Latency Streamable Video Super-Resolution via Auto-Regressive Diffusion
- A Neural Network-Based Real-time Casing Collar Recognition System for Downhole Instruments
- M-ErasureBench: A Comprehensive Multimodal Evaluation Benchmark for Concept Erasure in Diffusion Models
- OptiNIC: A Resilient and Tail-Optimal RDMA NIC for Distributed ML Workloads
- EgoReAct: Egocentric Video-Driven 3D Human Reaction Generation
- Improved cystic hygroma detection from prenatal imaging using ultrasound-specific self-supervised representation learning
- GHaLIB: A Multilingual Framework for Hope Speech Detection in Low-Resource Languages
- What Matters in Deep Learning for Time Series Forecasting?
- EEGForceFusion: Joint Tokenised-Continuous Representation Learning for Subject-Independent Grasp Force Decoding
- Tree Meets Transformer: A Hybrid Architecture for Scalable Power Allocation in Cell-Free Networks
- Raven: Mining Defensive Patterns in Ethereum via Semantic Transaction Revert Invariants Categories
- Enhancing Noise Resilience in Face Clustering via Sparse Differential Transformer
- LLM Agents as VC investors: Predicting Startup Success via RolePlay-Based Collective Simulation
- Energy-Guided Flow Matching Enables Few-Step Conformer Generation and Ground-State Identification
- KV-Tracker: Real-Time Pose Tracking with Transformers
- Lessons from Neuroscience for AI: How integrating Actions, Compositional Structure and Episodic Memory could enable Safe, Interpretable and Human-Like AI
- Learning When Not to Attend Globally
- TimePerceiver: An Encoder-Decoder Framework for Generalized Time-Series Forecasting
- SPECTRE: Spectral Pre-training Embeddings with Cylindrical Temporal Rotary Position Encoding for Fine-Grained sEMG-Based Movement Decoding
- Pose-Guided Residual Refinement for Interpretable Text-to-Motion Generation and Editing
- AFA-LoRA: Enabling Non-Linear Adaptations in LoRA with Activation Function Annealing
- LLM-Based vs. Lexicon-Based Sentiment Signals for Tail-Risk Detection in Meme Stocks
- Completed Hyperparameter Transfer across Modules, Width, Depth, Batch and Duration
- LLMBoost: Make Large Language Models Stronger with Boosting
- Agentic Software Issue Resolution with Large Language Models: A Survey
- Meta-information Guided Cross-domain Synergistic Diffusion Model for Low-dose PET Reconstruction
- Hierarchical Geometry of Cognitive States in Transformer Embedding Spaces
- On the Existence and Behavior of Secondary Attention Sinks
- Transformer Reconstructed with Dynamic Value Attention
- Emotion-Inspired Learning Signals (EILS): A Homeostatic Framework for Adaptive Autonomous Agents
- Latent Sculpting for Zero-Shot Generalization: A Manifold Learning Approach to Out-of-Distribution Anomaly Detection
- BitFlipScope: Scalable Fault Localization and Recovery for Bit-Flip Corruptions in LLMs
- AudioGAN: A Compact and Efficient Framework for Real-Time High-Fidelity Text-to-Audio Generation
- Neural ocean forecasting from sparse satellite-derived observations: a case-study for SSH dynamics and altimetry data
- Rethinking Leveraging Pre-Trained Multi-Layer Representations for Speaker Verification
- GPU-Virt-Bench: A Comprehensive Benchmarking Framework for Software-Based GPU Virtualization Systems
- Blood Pressure Estimation from PPG: A Comparative Study of Direct and ECG-Mediated Deep Learning Pipelines
- PointCHR: Point Cloud Analysis via Curvature-Aware Hyperbolic Rectification
- Joint Flow Matching for Generator-Consistent Classification
- DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention
- Attention Residuals
- Towards High-Level Semantic Intelligence
- BeyondFusion: Self-Aligned Latent Diffusion for Calibration-Free Infrared Super-Resolution and Infrared-Visible Fusion
- RayRoPE: Projective Ray Positional Encoding for Multi-view Attention
- LLM-Generated or Human-Written? Comparing Review and Non-Review Papers on ArXiv
- Impute On-Demand: Adaptive Correlated Time Series Imputation for Changing Environments
- NeurGO: Learning to Generate Elite Candidates for Meta-Black-Box Expensive Optimization
- PlanCraft: Sketch, Refine, and Furnish for Architect-Inspired Progressive 3D Residential Scene Generation
- ATLAS: Automated Approximation of Transformers for Efficient Homomorphic Inference in One Hour
- A Multi-stage Constrained Optimization Framework for Data-driven Problems
- Compressing Observation History into Agent Memory: Distilling Transformers into Recurrent Transformers
- Selecting Language Models for Social Science: Start Small, Start Open, and Validate
- Trainable Nonexpansive Denoisers for Contractive Image Reconstruction
- Variational-Ising-Attention (VIA):TailoredAttentionMattersfor Science
- SketchMamba: A Lightweight State-Space Model for Joint Progressive Sketch Classification and Stroke Auto-Completion
- MS-GPT: Rethinking MS/MS De Novo Structure Elucidation as Spectrum-Induced Posterior Querying of a Molecule-Language Model
- Restoration Flow Matching-Based Channel Refinement and Equalization Correction for MIMO Semantic Communications
- Multi-Modal Object Re-Identification with Prompt-S6 and Semantic-Aware Knowledge Guidance
- BERT-based Models vs. Large Language Models for Low-Resource Named Entity Recognition: A Comparative Study on Marathi
- PATCH-FFT: Unmasking Dormant Hardware Trojans with Patch-Based Frequency-Domain Transformers
- Omni-Prune: Query-Aware Unified Token Pruning for Efficient Omnimodal Large Language Models
- Express Language Modeling
- A satellite foundation model for improved wealth monitoring
- Full event interpretation with machine-learning-based particle-flow reconstruction in the CMS detector
- Identifying Dolphin Whistle Producers With Deep Learning: Moving Beyond Signature Whistles
- End-to-End 3-D Spatiotemporal Perception with Multimodal Fusion and V2X Collaboration
- Accelerating Language Model Workflows with Prompt Choreography
- Enhancing Code Understanding for Impact Analysis by Combining Transformers and Program Dependence Graphs
- Controllable Diversity in Normalization-Based Implicit Ensembles via Softmax-Temperature Modulation
- Latent-LoRA: Compact Latent-Space Adapters with Gradient-Free Routing for Continual Learning
- Kalypso: Relational LLM Serving
- Focus Is All You Need: Adaptive Goal-aware Attention Orchestration for Multi-Agent Graph Systems
- Machine Learning Decoding of Circuit-Level Noise for Bivariate Bicycle Codes
- Asleep at the Keyboard? Assessing the Security of GitHub Copilot's Code Contributions
- Deep reinforcement learning for tracking a moving target in jellyfish-like swimming
- Tracking by Detection and Query: An Efficient End-to-End Framework for Multi-Object Tracking
- RhythmFormer: Extracting Patterned rPPG Signals based on Periodic Sparse Attention
- Generative AI for Requirements Engineering: A Systematic Literature Review
- Generative Artificial Intelligence for Software Engineering -- A Research Agenda
- ConnectomeDiffuser: Generative AI Enables Brain Network Construction from Diffusion Tensor Imaging
- Anchor Attention, Small Cache: Code Generation with Large Language Models
- A Closer Look into Transformer-Based Code Intelligence Through Code Transformation: Challenges and Opportunities
- Generative AI and the future of scientometrics: current topics and future questions
- Decoding Phone Pairs from MEG Signals Across Speech Modalities
- A transformer-based multi-stream approach for isolated iranian sign language recognition
- SpecFormer: Mitigating Embedding and Attention Collapse via Spectral-Aware Transformer for Recommendation
- Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification
- Modeling Stylistic Co-evolution in Symbolic Music Heritage Collections
- The Semantic Least-Energy Principle: A Hypothesis for Intelligence
- Toward Alias-Free Channel Extrapolation in Upper Mid-Band Systems: A Spatial-Frequency-Temporal Tensor Learning Approach
- UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models
- Physics-informed token transformer methodology for nonlinear balance laws. I. Schwarzschild--Burgers fluid flows
- Forecasting the Emergence and Evolution of Crash Hotspots: A Unified Deep Learning Framework for Proactive Traffic Safety
- LLM-based Source Code Compression via Thresholded Symbol Ranking
- Strategy-Aware Parameter-Efficient Adaptation for LLM-based Auto-Bidding
- Efficient Ultrasound Image Segmentation with Token-Conditioned Neural Cellular Automata
- Unifying Generative Recall and Multi-Objective Ranking in a Single Decoder-Only Sequence
- DSCH-Loss: A Dynamic Semantic Channel Objective for Deep Semantic Hashing
- Evaluating the Impact of Explainable AI on Trust in AI-Assisted Code Review
- A Coulomb Particle Model for Learning Kernel Attention in Transformers
- What CLIP Knows but Cannot Say: Recovering Negation from Frozen Intermediate Features
- PI-GINOT: Data-free geometry-informed neural operator learning for finite-strain hyperelasticity on parametric DogBone specimens
- Libra: Taming Attention Workload Skew in Long-Context LLM Training with Bounded Sequence Pool
- Training with (Swap) Regret Loss in a Single-Layer Self-Attention Model: A Case Study on the Probability Simplex
- Mini-batch Noise Lowers Sharpness via Dominant-Subspace Fluctuations
- Backend-Aware Graph Learning for Denoising Outcome Distributions in Quantum Program Testing
- AI Empowered Communication and Radar Modulation Recognition: A Survey
- Structure over Depth: A Single-Block Spatio-Temporal Transformer for Multi-Entity Reasoning
- Foundation Models and Fine-Tuning: Toward a New Generation of Models for Time Series Forecasting
- GPT-like transformer model for silicon tracking detector simulation
- Quantum Transformer BSDE Solver via Multi-Layer Fully-Connected Variational Quantum Circuits
- TopoGR: Revealing and Preserving Latent Structure of Semantic ID in Generative Recommendation
- WHTMix: Efficient Stereo Depth Estimation via Walsh-Hadamard Token Mixing
- OmniMate: Open-Ended Real-Time Streaming Audio-Visual Generation for Interactive Avatars
- A scaling law of contextual persistence in human language
- Beyond Single-Episode Optimization: Sliding-Window Aware Generative Auto-Bidding for Long-Term Advertising Effectiveness
- Tokenizing Numerical and Embedding Features for LLM RecSys
- CADENCE: A Cardiac Atom Dictionary for Interpretable Neural Concept Extraction from ECG Foundation Models
- Automated Numerical Stability Analysis of Deep Learning Operators
- PreDiff-LM: Pretrained Discrete Masked Diffusion Language Modeling with Hybrid Attention
- Memory for Large Language Models
- Gaussian Volumetric Representation for Efficient Shear-Warp Visualization
- Grevo: A Unified Generative Recommendation Framework with Evolutionary Item Indexing
- Many-body Tipping Dynamics of ChatGPT-like AIs
- Raven: High-Recall Sequence Modeling with Sparse Memory Routing
- Beyond Predictive Accuracy: A Reliability-Aware Audit of Molecular Representations for Human Olfaction
- Phase Structure in Rotary Attention: A Spectral Framework for Semantic Continuity and Execution-Boundary Governance
- ReLATE: Reliability-Guided Evidence Fusion for Robust UAV--Satellite cross-view Geo-Localization
- Emergent Latent-State Computation under Stochastic Volatility
- AMPBench-MT: A Homology-Controlled Benchmark for Antimicrobial Peptide Potency, Spectrum, and Safety Prediction
- Beyond Prefill-Decode Disaggregation: Dissecting LLM Inference for Heterogeneous Platforms via Dynamic Operator Scheduling
- Intrinsic and Triangulation-Agnostic Attention: A Simple and Powerful Approach for Learning on Meshes
- Inferring Missing Trajectory Data with Temporal Convolutional Networks
- Eliminating Propagation Delay: Attention-Based Spatial-Temporal Fusion Graph Convolution Network for Traffic Flow Prediction
- PLATO: Pointer Learner for Agent and Task Openness
- Out-of-Length Scene Text Recognition: A Two-Axis Diagnosis and a Training-Free Fix
- A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions
- LithoFormer: A Robust Framework for Stratigraphic Inference via Transformers
- Multimodal Surface EMG Hand Gesture Recognition Using Query-Based Transformers for Prosthetic Control
- What Softmax Throws Away: Mass-Aware Attention for Evidence Accumulation
- Hidden Boundary Motion in Transformer Optimization: Function-Space Orthogonalization of Affine Weight and Bias Updates
- LC-SEPLM: long-range contact-supervised adaptation for sequence-only protein representation learning
- Evaluating and Mitigating the Misguidance Effect of Buggy Code in LLM-Generated Unit Tests
- Optimizing Transformer Neural Network for Real-Time Outlier Detection on FPGAs
- Learning to Access Computation: Accessibility Plasticity as a Principle of Adaptive Intelligence
- CrossSpine: Multi-scale Cross-sequence Attention with Anatomical Priors for Automated Pfirrmann Grading
- Open Your Model’s Eyes: Video and Context-Aware Multimodal Backchannel Prediction
- CausalGate: Causal Importance Distillation for Transformer Module Pruning
- StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design
- Visible-Light Imaging Diagnosis of Neutral Particle Emission Tomography in the Tokamak Divertor: An Efficient Transformer-based Surrogate Model
- RoleMix: Unifying Sequential and Non-Sequential Features via Semantic Tokenization for Post-Click Conversion Rate Prediction
- Imprompt: A Language Framework for Prompt Programming
- DOSA: A Tree-Guided, Self-Regressive Framework for Long Document Structure Analysis
- Bumblebee: Interleaved Mixed-Layer Building Blocks for Large-Scale Recommendation Systems
- Reading Without a Reader: Large Language Models Collapse Reading and Writing into a Single Entangled Code
- Do modern speech LLMs and re-scoring techniques improve bilingual ASR performance for Basque and Spanish in domain-specific contexts?
- T2LDM++: A Self-Conditioned Representation Guided Diffusion Model for Realistic Text-to-LiDAR Scene Generation
- Neural Machine Translation for Low-Resource Tangkhul--English
- Diffusion models for SU(2) lattice gauge theory in two dimensions
- Retrieval, not hallucinations, will be the limiting factor for LLM-based clinical AI tools
- GLIDE: Guided Layerwise Hybrid Attention for Efficient LLM Inference
- PTStore (Prefix Tensor Store): Distributed Prefix Caching and Replication for High Throughput Inference Serving
- GrocLM: Grocery Category Recommendation in E-Commerce with Large Language Models
- Evaluating LLMs as Interpretable Controllers for Dynamical Systems
- Differencing the Diffusion Trajectory toward Uncertain Components for Time Series Forecasting
- Atmospheric Diffusion-Guided Spatio-Temporal Transformer for Nuclear Radiation Forecasting
- Lexical discovery in unknown environments orchestrated by Large Language Models
- TriSP: Tri-Signal Structured Pruning for Large Language Models
- FRIGID: Scaling Diffusion-Based Molecular Generation from Mass Spectra at Training and Inference Time
- Key-Value Means: Transformers with Expandable Block-Recurrent Compressed Memory
- Learning to Optimize at Scale: A Benders Decomposition-TransfORmers Framework for Stochastic Combinatorial Optimization
- Kan Extension Transformers: A Categorical Unification of Attention, Diffusion, and Predict-Detach Self-Conditioning
- Bridging MARL to SARL: An Order-Independent Multi-Agent Transformer via Latent Consensus
- Accelerating Frequency Domain Diffusion Models with Error-Feedback Event-Driven Caching
- Training-Driven Representational Geometry Modularization Predicts Brain Alignment in Language Models
- Forecasting Ionospheric Irregularities on GNSS Lines of Sight Using Dynamic Graphs with Ephemeris Conditioning
- MOCA: A Transformer-based Modular Causal Inference Framework with One-way Cross-attention and Cutting Feedback
- Latent Space Probing for Adult Content Detection in Video Generative Models
- Scalable Neural Decoders for Practical Fault-Tolerant Quantum Computation
- Deep Generative Spatiotemporal Engression for Probabilistic Forecasting of Epidemics
- VEMamba: Efficient Isotropic Reconstruction of Volume Electron Microscopy with Axial-Lateral Consistent Mamba
- ORGAN: Object-Centric Representation Learning Using Cycle Consistent Generative Adversarial Networks
- Exact and Asymptotically Complete Robust Verifications of Neural Networks via Ising Solvers
- A General Deep Learning Framework for Wireless Resource Allocation under Discrete Constraints
- The Convergence Frontier: Integrating Machine Learning and High Performance Quantum Computing for Next-Generation Drug Discovery
- ZipMap: Linear-Time Stateful 3D Reconstruction via Test-Time Training
- StegaFFD: Privacy-preserving Face Forgery Detection via Fine-grained Steganographic Domain Lifting
- Reverso: Efficient Time Series Foundation Models for Zero-shot Forecasting
- Plain Transformers are Surprisingly Powerful Link Predictors
- FUSCO: High-Performance Distributed Data Shuffling via Transformation-Communication Fusion
- Benchmarking deep learning models for Raman spectroscopy across open-source datasets
- Reconstructing Item Characteristic Curves using Fine-Tuned Large Language Models
- CellMamba: Adaptive Mamba for Accurate and Efficient Cell Detection
- Self-attention vector output similarities reveal how machines pay attention
- Optimizing Resource Allocation for Geographically-Distributed Inference by Large Language Models
- Reloc-VGGT: Visual Re-localization with Geometry Grounded Transformer
- DPAR: Dynamic Patchification for Efficient Autoregressive Visual Generation
- TimeBill: Time-Budgeted Inference for Large Language Models
- UniStateDLO: Unified Generative State Estimation and Tracking of Deformable Linear Objects Under Occlusion for Constrained Manipulation
- Knot Forcing: Taming Autoregressive Video Diffusion Models for Real-time Infinite Interactive Portrait Animation
- Analyzing the Mechanism of Attention Collapse in VGGT from a Dynamics Perspective
- RIPCN: A Road Impedance Principal Component Network for Probabilistic Traffic Flow Forecasting
- Contrastive Graph Modeling for Cross-Domain Few-Shot Medical Image Segmentation
- MoE-TransMov: A Transformer-based Model for Next POI Prediction in Familiar & Unfamiliar Movements
- AVP-Fusion: Adaptive Multi-Modal Fusion and Contrastive Learning for Two-Stage Antiviral Peptide Identification
- Vision Transformers are Circulant Attention Learners
- GoldenFuzz: Generative Golden Reference Hardware Fuzzing
- Efficient MoE Inference with Fine-Grained Scheduling of Disaggregated Expert Parallelism
- RefineBridge: Generative Bridge Models Improve Financial Forecasting by Foundation Models
- Hyperion: Low-Latency Ultra-HD Video Analytics via Collaborative Vision Transformer Inference
- Dynamic Attention (DynAttn): Interpretable High-Dimensional Spatio-Temporal Forecasting (with Application to Conflict Fatalities)
- LLM-Driven Feature-Level Adversarial Attacks on Android Malware Detectors
- GraviBERT: Transformer-based inference for gravitational-wave time series
- Fast SAM2 with Text-Driven Token Pruning
- TICON: A Slide-Level Tile Contextualizer for Histopathology Representation Learning
- Parallel Token Prediction for Language Models
- Surgical Scene Segmentation using a Spike-Driven Video Transformer with Real-Time Potential
- SMART SLM: Structured Memory and Reasoning Transformer, A Small Language Model for Accurate Document Assistance
- GriDiT: Factorized Grid-Based Diffusion for Efficient Long Image Sequence Generation
- TGC-Net: A Structure-Aware and Semantically-Aligned Framework for Text-Guided Medical Image Segmentation
- SparScene: Efficient Traffic Scene Representation via Sparse Graph Learning for Large-Scale Trajectory Generation
- STLDM: Spatio-Temporal Latent Diffusion Model for Precipitation Nowcasting
- A Mechanistic Analysis of Transformers for Dynamical Systems
- Beyond Context: Large Language Models Failure to Grasp Users Intent
- FreeInpaint: Tuning-free Prompt Alignment and Visual Rationality Enhancement in Image Inpainting
- Uncovering Hierarchical Structure in LLM Embeddings with δ-Hyperbolicity, Ultrametricity, and Neighbor Joining
- Hierarchical Modeling Approach to Fast and Accurate Table Recognition
- Understanding Scaling Laws in Deep Neural Networks via Feature Learning Dynamics
- Language-Guided Grasp Detection with Coarse-to-Fine Learning for Robotic Manipulation
- Next-Scale Prediction: A Self-Supervised Approach for Real-World Image Denoising
- Critical Points of Degenerate Metrics on Algebraic Varieties: A Tale of Overparametrization
- FluencyVE: Marrying Temporal-Aware Mamba with Bypass Attention for Video Editing
- Linear Attention for Joint Power Optimization and User-Centric Clustering in Cell-Free Networks
- Distilling the Essence: Efficient Reasoning Distillation via Sequence Truncation
- CoSeNet: A Novel Approach for Optimal Segmentation of Correlation Matrices
- Generalised Linear Models in Deep Bayesian RL with Learnable Basis Functions
- Reflection Pretraining Enables Token-Level Self-Correction in Biological Sequence Models
- Foundation Model-based Evaluation of Neuropsychiatric Disorders: A Lifespan-Inclusive, Multi-Modal, and Multi-Lingual Study
- Embodied AI-Enhanced IoMT Edge Computing: UAV Trajectory Optimization and Task Offloading with Mobility Prediction
- Architectural Trade-offs in Small Language Models Under Compute Constraints
- USE: A Unified Model for Universal Sound Separation and Extraction
- A Community-Enhanced Graph Representation Model for Link Prediction
- Assessing the Software Security Comprehension of Large Language Models
- Learning the Macroeconomic Language
- Symbolic regression for defect interactions in 2D materials
- Stabilizing Multimodal Autoencoders: A Theoretical and Empirical Analysis of Fusion Strategies
- SA-DiffuSeq: Addressing Computational and Scalability Challenges in Long-Document Generation with Sparse Attention
- Advancing Multimodal Teacher Sentiment Analysis:The Large-Scale T-MED Dataset & The Effective AAM-TSA Model
- Explainable time-series forecasting with sampling-free SHAP for Transformers
- Recurrent Off-Policy Deep Reinforcement Learning Doesn't Have to be Slow
- milliMamba: Specular-Aware Human Pose Estimation via Dual mmWave Radar with Multi-Frame Mamba Fusion
- GeoTransolver: Learning Physics on Irregular Domains Using Multi-scale Geometry Aware Physics Attention Transformer
- Field-Space Attention for Structure-Preserving Earth System Transformers
- Toward Explaining Large Language Models in Software Engineering Tasks
- SpidR: Learning Fast and Stable Linguistic Units for Spoken Language Models Without Supervision
- Structured Visualization Design Knowledge for Grounding Generative Reasoning and Situated Feedback
- UbiQVision: Quantifying Uncertainty in XAI for Image Recognition
- Debate-Enhanced Pseudo Labeling and Frequency-Aware Progressive Debiasing for Weakly-Supervised Camouflaged Object Detection with Scribble Annotations
- HGAN-SDEs: Learning Neural Stochastic Differential Equations with Hermite-Guided Adversarial Training
- Evolutionary Neural Architecture Search with Dual Contrastive Learning
- UMAMI: Unifying Masked Autoregressive Models and Deterministic Rendering for View Synthesis
- A Novel Graph-Sequence Learning Model for Inductive Text Classification
- Jensen-Shannon Divergence Message-Passing for Rich-Text Graph Representation Learning
- Item Region-based Style Classification Network (IRSN): A Fashion Style Classifier Based on Domain Knowledge of Fashion Experts
- QE-Catalytic: A Graph-Language Multimodal Base Model for Relaxed-Energy Prediction in Catalytic Adsorption
- Self-motion as a structural prior for coherent and robust formation of cognitive maps
- DecoKAN: Interpretable Decomposition for Forecasting Cryptocurrency Market Dynamics
- SegEarth-R2: Towards Comprehensive Language-guided Segmentation for Remote Sensing Images
- LoFT-LLM: Low-Frequency Time-Series Forecasting with Large Language Models
- A Dual-Branch Local-Global Framework for Cross-Resolution Land Cover Mapping
- Block-Recurrent Dynamics in Vision Transformers
- Edge-Served Congestion Control for Wireless Multipath Transmission with a Transformer Agent
- BiCoR-Seg: Bidirectional Co-Refinement Framework for High-Resolution Remote Sensing Image Segmentation
- LiteFusion: Taming 3D Object Detectors from Vision-Based to Multi-Modal with Minimal Adaptation
- Memory as Resonance: A Biomimetic Architecture for Infinite Context Memory on Ergodic Phonetic Manifolds
- MoE-DiffuSeq: Enhancing Long-Document Diffusion Models with Sparse Attention and Mixture of Experts
- LADLE-MM: Limited Annotation based Detector with Learned Ensembles for Multimodal Misinformation
- Attention Is Not What You Need
- Generative Krylov Subspace Representations for Scalable Quantum Eigensolvers
- Non-Contrast CT Esophageal Varices Grading through Clinical Prior-Enhanced Multi-Organ Analysis
- Generating the Past, Present and Future from a Motion-Blurred Image
- Over++: Generative Video Compositing for Layer Interaction Effects
- No Data? No Problem: Robust Vision-Tabular Learning with Missing Values
- The Epistemological Consequences of Large Language Models: Rethinking collective intelligence and institutional knowledge
- ActAvatar: Temporally-Aware Precise Action Control for Talking Avatars
- StoryMem: Multi-shot Long Video Storytelling with Memory
- Multi-Modal Soccer Scene Analysis with Masked Pre-Training
- Deep Learning for Unrelated-Machines Scheduling: Handling Variable Dimensions
- Dynamic Stream Network for Combinatorial Explosion Problem in Deformable Medical Image Registration
- Kunnafonidilaw ka Cadeau: an ASR dataset of present-day Bambara
- Efficient Spike-driven Transformer for High-performance Drone-View Geo-Localization
- ReasonCD: A Multimodal Reasoning Large Model for Implicit Change-of-Interest Semantic Mining
- STAR: Semantic-Traffic Alignment and Retrieval for Zero-Shot HTTPS Website Fingerprinting
- Generative Human-Object Interaction Detection via Differentiable Cognitive Steering of Multi-modal LLMs
- PEDESTRIAN: An Egocentric Vision Dataset for Obstacle Detection on Pavements
- Practical Quantum-Classical Feature Fusion for complex data Classification
- CascadeInfer: Length-Aware Scheduling of LLM Serving with Low Latency and Load Balancing
- Beyond Sliding Windows: Learning to Manage Memory in Non-Markovian Environments
- RP-CATE: Recurrent Perceptron-based Channel Attention Transformer Encoder for Industrial Hybrid Modeling
- SAP: Syntactic Attention Pruning for Transformer-based Language Models
- HyperLoad: A Cross-Modality Enhanced Large Language Model-Based Framework for Green Data Center Cooling Load Prediction
- Can abstract concepts from LLM improve SLM performance?
- VLNVerse: A Benchmark for Vision-Language Navigation with Versatile, Embodied, Realistic Simulation and Evaluation
- R-GenIMA: Integrating Neuroimaging and Genetics with Interpretable Multimodal AI for Alzheimer's Disease Progression
- DeepGESI: A Non-Intrusive Objective Evaluation Model for Predicting Speech Intelligibility in Hearing-Impaired Listeners
- Self-Attention with State-Object Weighted Combination for Compositional Zero Shot Learning
- OmniEgoCap: Camera-Agnostic Sequence-Level Egocentric Motion Reconstruction
- CASA: Cross-Attention over Self-Attention for Efficient Vision-Language Fusion
- DeltaMIL: Gated Memory Integration for Efficient and Discriminative Whole Slide Image Analysis
- CienaLLM: Generative Climate-Impact Extraction from News Articles with Autoregressive LLMs
- HARMON-E: Hierarchical Agentic Reasoning for Multimodal Oncology Notes to Extract Structured Data
- MEEA: Mere Exposure Effect-Driven Confrontational Optimization for LLM Jailbreaking
- IPCV: Information-Preserving Compression for MLLM Visual Encoders
- Alternative positional encoding functions for neural transformers
- Enhancing Fully Formatted End-to-End Speech Recognition with Knowledge Distillation via Multi-Codebook Vector Quantization
- Brain-Gen: Towards Interpreting Neural Signals for Stimulus Reconstruction Using Transformers and Latent Diffusion Models
- HARBOR: Holistic Adaptive Risk assessment model for BehaviORal healthcare
- Misbehavior Forecasting for Focused Autonomous Driving Systems Testing
- Revealing Perception and Generation Dynamics in LVLMs: Mitigating Hallucinations via Validated Dominance Correction
- RIS-Enabled Smart Wireless Environments: Fundamentals and Distributed Optimization
- From Scratch to Fine-Tuned: A Comparative Study of Transformer Training Strategies for Legal Machine Translation
- Overcoming Spectral Bias via Cross-Attention
- From Shortcut to Induction Head: How Data Diversity Shapes Algorithm Selection in Transformers
- A Comparative Study of Light-weight Language Models for PII Masking and their Deployment for Real Conversational Texts
- ShibuyaSocial: Multi-scale Model of Pedestrian Flows in Scramble Crossing
- KeenKT: Knowledge Mastery-State Disambiguation for Knowledge Tracing
- MeanFlow-TSE: One-Step Generative Target Speaker Extraction with Mean Flow
- Data-Driven Calibration of Large Liquid Detectors with Unsupervised Learning
- Phoneme-based speech recognition driven by large language models and sampling marginalization
- Enhancing 3D Semantic Scene Completion with a Refinement Module
- Large Language Models as Discounted Bayesian Filters
- Plasticine: A Traceable Diffusion Model for Medical Image Translation
- Secret mixtures of experts inside your LLM
- On the Universality of Transformer Architectures; How Much Attention Is Enough?
- MeniMV: A Multi-view Benchmark for Meniscus Injury Severity Grading
- Learning-Based Estimation of Spatially Resolved Scatter Radiation Fields in Interventional Radiology
- HyDRA: Hierarchical and Dynamic Rank Adaptation for Mobile Vision Language Model
- InstructNet: A Novel Approach for Multi-Label Instruction Classification through Advanced Deep Learning
- UniMPR: A Unified Framework for Multimodal Place Recognition with Heterogeneous Sensor Configurations
- When Does Learning Renormalize? Sufficient Conditions for Power Law Spectral Dynamics
- External Hippocampus: Topological Cognitive Maps for Guiding Large Language Model Reasoning
- MACE-Dance: Motion-Appearance Cascaded Experts for Music-Driven Dance Video Generation
- Estimating Solvation Free Energies with Boltzmann Generators
- A two-stream network with global-local feature fusion for bone age assessment
- Name That Part: 3D Part Segmentation and Naming
- Unifying Deep Predicate Invention with Pre-trained Foundation Models
- Adversarial Robustness of Vision in Open Foundation Models
- InfinityEBSD : Metrics-Guided Infinite-Size EBSD Map Generation With Diffusion Models
- MambaMIL+: Modeling Long-Term Contextual Patterns for Gigapixel Whole Slide Image
- Learning Spatio-Temporal Feature Representations for Video-Based Gaze Estimation
- StereoMV2D: A Sparse Temporal Stereo-Enhanced Framework for Robust Multi-View 3D Object Detection
- Parameter-Efficient Fine-Tuning for HAR: Integrating LoRA and QLoRA into Transformer Models
- GreedySnake: Accelerating SSD-Offloaded LLM Training with Efficient Scheduling and Optimizer Step Overlapping
- Electric Vehicle Charging Load Forecasting: An Experimental Comparison of Machine Learning Methods
- Accelerating Multi-modal LLM Gaming Performance via Input Prediction and Mishit Correction
- Spatially-informed transformers: Injecting geostatistical covariance biases into self-attention for spatio-temporal forecasting
- KV Admission: Learning What to Write for Efficient Long-Context Inference
- A Systematic Reproducibility Study of BSARec for Sequential Recommendation
- Improving Cardiac Risk Prediction Using Data Generation Techniques
- Towards Deeper Emotional Reflection: Crafting Affective Image Filters with Generative Priors
- AdvJudge-Zero: Binary Decision Flips in LLM-as-a-Judge via Adversarial Control Tokens
- Multi-level distortion-aware deformable network for omnidirectional image super-resolution
- Bridging Natural Language and Formal Specification--Automated Translation of Software Requirements to LTL via Hierarchical Semantics Decomposition Using LLMs
- SynergyWarpNet: Attention-Guided Cooperative Warping for Neural Portrait Animation
- Deep Learning-Based Surrogate Creep Modelling in Inconel 625: A High-Temperature Alloy Study
- Explanation Beyond Intuition: A Testable Criterion for Inherent Explainability
- Vision-Language Model Guided Image Restoration
- WDFFU-Mamba: A Wavelet-guided Dual-attention Feature Fusion Mamba for Breast Tumor Segmentation in Ultrasound Images
- BEOL Ferroelectric Compute-in-Memory Ising Machine for Simulated Bifurcation
- From Fake Focus to Real Precision: Confusion-Driven Adversarial Attention Learning in Transformers
- Physics of Language Models: Part 4.1, Architecture Design and the Magic of Canon Layers
- Attention Distance: A Novel Metric for Directed Fuzzing with Large Language Models
- Confidence-Credibility Aware Weighted Ensembles of Small LLMs Outperform Large LLMs in Emotion Detection
- DNAMotifTokenizer: Towards Biologically Informed Tokenization of Genomic Sequences
- A systematic assessment of Large Language Models for constructing two-level fractional factorial designs
- DGH: Dynamic Gaussian Hair
- Adversarial VR: An Open-Source Testbed for Evaluating Adversarial Robustness of VR Cybersickness Detection and Mitigation
- In-Context Algebra
- Training Together, Diagnosing Better: Federated Learning for Collagen VI-Related Dystrophies
- TOGGLE: Temporal Logic-Guided Large Language Model Compression for Edge
- PrivateXR: Defending Privacy Attacks in Extended Reality Through Explainable AI-Guided Differential Privacy
- LLMCache: Layer-Wise Caching Strategies for Accelerated Reuse in Transformer Inference
- What Do Prosody and Text Convey? Characterizing How Meaningful Information is Distributed Across Multiple Channels
- The Colombian legislative process, 2014-2025: networks, topics, and polarization
- KineST: A Kinematics-guided Spatiotemporal State Space Model for Human Motion Tracking from Sparse Signals
- FlowDet: Unifying Object Detection and Generative Transport Flows
- NRGPT: An Energy-based Alternative for GPT
- VERM: Leveraging Foundation Models to Create a Virtual Eye for Efficient 3D Robotic Manipulation
- KOSS: Kalman-Optimal Selective State Spaces for Long-Term Sequence Modeling
- Towards Reproducibility in Predictive Process Mining: SPICE -- A Deep Learning Library
- PoseMoE: Mixture-of-Experts Network for Monocular 3D Human Pose Estimation
- Smile on the Face, Sadness in the Eyes: Bridging the Emotion Gap with a Multimodal Dataset of Eye and Facial Behaviors
- Efficient CPU-GPU Collaborative Inference for MoE-based LLMs on Memory-Limited Systems
- Synthelite: Chemist-aligned and feasibility-aware synthesis planning with LLMs
- Kascade: A Practical Sparse Attention Method for Long-Context LLM Inference
- EverybodyDance: Bipartite Graph-Based Identity Correspondence for Multi-Character Animation
- Collaborative Edge-to-Server Inference for Vision-Language Models
- Bunch-by-Bunch Prediction of Beam Transverse Position, Phase, and Length in a Storage Ring Using Neural Networks
- Can Transformers overcome the lack of data in the simulation of history-dependent flows?
- Real-Time Human-Robot Interaction Intent Detection Using RGB-based Pose and Emotion Cues with Cross-Camera Model Generalization
- QuadSentinel: Sequent Safety for Machine-Checkable Control in Multi-agent Systems
- Sigma-MoE-Tiny Technical Report
- The Evolution of Reranking Models in Information Retrieval: From Heuristic Methods to Large Language Models
- ARMFlow: AutoRegressive MeanFlow for Online 3D Human Reaction Generation
- LAPX: Lightweight Hourglass Network with Global Context
- In-Context Multi-Operator Learning with DeepOSets
- Convolutional Lie Operator for Sentence Classification
- Predictive Modeling of Maritime Radar Data Using Transformer Architecture
- Open Ad-hoc Categorization with Contextualized Feature Learning
- Evaluation of Generative Models for Emotional 3D Animation Generation in VR
- On Recommending Category: A Cascading Approach
- Higher-Order LaSDI: Reduced Order Modeling with Multiple Time Derivatives
- Lyapunov-based Adaptive Transformer (LyAT) for Control of Stochastic Nonlinear Systems
- Dynamic Rank Reinforcement Learning for Adaptive Low-Rank Multi-Head Self Attention in Large Language Models
- In-Context Semi-Supervised Learning
- DSO: Direct Steering Optimization for Bias Mitigation
- DiffusionVL: Translating Any Autoregressive Models into Diffusion Vision Language Models
- Dynamic Rebatching for Efficient Early-Exit Inference with DREX
- End-to-End Training for Autoregressive Video Diffusion via Self-Resampling
- Multi-Modal Semantic Communication
- Explaining the Reasoning of Large Language Models Using Attribution Graphs
- SoFlow: Solution Flow Models for One-Step Generative Modeling
- Characterizing Mamba's Selective Memory using Auto-Encoders
- OccSTeP: Benchmarking 4D Occupancy Spatio-Temporal Persistence
- Autoregressive Language Models are Secretly Energy-Based Models: Insights into the Lookahead Capabilities of Next-Token Prediction
- How Smoothing is N-simplicial Attention?
- FlexAvatar: Learning Complete 3D Head Avatars with Partial Supervision
- Bolmo: Byteifying the Next Generation of Language Models
- Reducing Pilots in Channel Estimation with Predictive Foundation Models
- Learning inflection classes using Adaptive Resonance Theory
- DeX-Portrait: Disentangled and Expressive Portrait Animation via Explicit and Latent Motion Representations
- RUMPL: Ray-Based Transformers for Universal Multi-View 2D to 3D Human Pose Lifting
- AnySleep: a channel-agnostic deep learning system for high-resolution sleep staging in multi-center cohorts
- ST-DETrack: Identity-Preserving Branch Tracking in Entangled Plant Canopies via Dual Spatiotemporal Evidence
- Towards Seamless Interaction: Causal Turn-Level Modeling of Interactive 3D Conversational Head Dynamics
- BBNet: accurate neural network emulator for primordial light element abundances
- O-EENC-SD: Efficient Online End-to-End Neural Clustering for Speaker Diarization
- Yes-MT's Submission to the Low-Resource Indic Language Translation Shared Task in WMT 2024
- SLCFormer: Spectral-Local Context Transformer with Physics-Grounded Flare Synthesis for Nighttime Flare Removal
- Explainable Action Form Assessment by Exploiting Multimodal Chain-of-Thoughts Reasoning
- Attention in Motion: Secure Platooning via Transformer-based Misbehavior Detection
- FADTI: Fourier and Attention Driven Diffusion for Multivariate Time Series Imputation
- How Many Heads Make an SSM? A Unified Framework for Attention and State Space Models
- CF-Net: A Cross-Feature Reconstruction Network for High-Accuracy 1-Bit Target Classification
- Beyond Fast and Slow: Cognitive-Inspired Elastic Reasoning for Large Language Models
- SigMA: Path Signatures and Multi-head Attention for Learning Parameters in fBm-driven SDEs
- EMFusion: Uncertainty-Aware Conditional Diffusion Model for Multivariate Narrow-band Exposure Forecasting
- LADY: Linear Attention for Autonomous Driving Efficiency without Transformers
- Adaptive Multimodal Person Recognition: A Robust Framework for Handling Missing Modalities
- SeBERTis: A Framework for Producing Classifiers of Security-Related Issue Reports
- Dense Associative Memories with Analog Circuits
- Efficient Nudged Elastic Band Method using Neural Network Bayesian Algorithm Execution
- Mixture of Attention Schemes (MoAS): Learning to Route Between MHA, GQA, and MQA
- Cross-modal ultra-scale learning with tri-modalities of renal biopsy images for glomerular multi-disease auxiliary diagnosis
- Physics-driven human-like working memory outperforms digital networks in dynamic vision
- ISS Policy : Scalable Diffusion Policy with Implicit Scene Supervision
- An Empirical Study on Chinese Character Decomposition in Multiword Expression-Aware Neural Machine Translation
- ScamSweeper: Detecting Illegal Accounts in Web3 Scams via Transactions Analysis
- Multiscale Aggregated Hierarchical Attention (MAHA): A Game Theoretic and Optimization Driven Approach to Efficient Contextual Modeling in Large Language Models
- Task Matrices: Linear Maps for Cross-Model Finetuning Transfer
- Isolated Sign Language Recognition with Segmentation and Pose Estimation
- T5Gemma 2: Seeing, Reading, and Understanding Longer
- ART: Articulated Reconstruction Transformer
- Focus: A Streaming Concentration Architecture for Efficient Vision-Language Models
- Segmental Attention Decoding With Long Form Acoustic Encodings
- TiME: Tiny Monolingual Encoders for Efficient NLP Pipelines
- AMD-HookNet++: Evolution of AMD-HookNet with Hybrid CNN-Transformer Feature Enhancement for Glacier Calving Front Segmentation
- PruneX: A Hierarchical Communication-Efficient System for Distributed CNN Training with Structured Pruning
- Incentives or Ontology? A Structural Rebuttal to OpenAI's Hallucination Thesis
- ParaFormer: A Generalized PageRank Graph Transformer for Graph Representation Learning
- FUSION: Forecast-Embedded Agent Scheduling with Service Incentive Optimization over Distributed Air-Ground Edge Networks
- Hybrid Iterative Solvers with Geometry-Aware Neural Preconditioners for Parametric PDEs
- LLM-driven Knowledge Enhancement for Multimodal Cancer Survival Prediction
- Towards Nepali-language LLMs: Efficient GPT training with a Nepali BPE tokenizer
- Residual GRU+MHSA: A Lightweight Hybrid Recurrent Attention Model for Cardiovascular Disease Detection
- Dual-objective Language Models: Training Efficiency Without Overfitting
- HiFi-Portrait: Zero-shot Identity-preserved Portrait Generation with High-fidelity Multi-face Fusion
- C-ing Clearly: Enhanced Binary Code Explanations using C code
- IaC Generation with LLMs: An Error Taxonomy and A Study on Configuration Knowledge Injection
- Relaying Signal When Monitoring Traffic: Double Use of Aerial Vehicles Towards Intelligent Low-Altitude Networking
- Attention-Based Foundation Model for Quantum States
- PSMamba: Progressive Self-supervised Vision Mamba for Plant Disease Recognition
- TUN: Detecting Significant Points in Persistence Diagrams with Deep Learning
- Gödel's Poetry
- Elastic3D: Controllable Stereo Video Conversion with Guided Latent Decoding
- 4D-RaDiff: Latent Diffusion for 4D Radar Point Cloud Generation
- Georeferencing complex relative locality descriptions with large language models
- End-to-End Learning-based Video Streaming Enhancement Pipeline: A Generative AI Approach
- FastDDHPose: Towards Unified, Efficient, and Disentangled 3D Human Pose Estimation
- Particulate: Feed-Forward 3D Object Articulation
- Softmax as Linear Attention in the Large-Prompt Regime: a Measure-based Perspective
- Prompt Governance? On Governing Technologies Governed by Natural Language
- Citation importance-aware document representation learning for large-scale science mapping
- ReflCtrl: Controlling LLM Reflection via Representation Engineering
- SASQ: Static Activation Scaling for Quantization-Aware Training in Large Language Models
- PathFinder: Advancing Path Loss Prediction for Single-to-Multi-Transmitter Scenario
- TorchTraceAP: A New Benchmark Dataset for Detecting Performance Anti-Patterns in Computer Vision Models
- Neurosymbolic Inference On Foundation Models For Remote Sensing Text-to-image Retrieval With Complex Queries
- A Unified Sparse Attention via Multi-Granularity Compression
- SonicMoE: Accelerating MoE with IO and Tile-aware Optimizations
- What Affects the Effective Depth of Large Language Models?
- Context Representation via Action-Free Transformer encoder-decoder for Meta Reinforcement Learning
- AsarRec: Adaptive Sequential Augmentation for Robust Self-supervised Sequential Recommendation
- Dynamic stacking ensemble learning with investor knowledge representations for stock market index prediction based on multi-source financial data
- From Feature Interaction to Feature Generation: A Generative Paradigm of CTR Prediction Models
- Unleashing the Power of Image-Tabular Self-Supervised Learning via Breaking Cross-Tabular Barriers
- Accelerating MHC-II Epitope Discovery via Multi-Scale Prediction in Antigen Presentation
- A Single Architecture for Representing Invariance Under Any Space Group
- KLO-Net: A Dynamic K-NN Attention U-Net with CSP Encoder for Efficient Prostate Gland Segmentation from MRI
- Tractable Model for Tunable Non-Markovian Dynamics
- VajraV1 -- The most accurate Real Time Object Detector of the YOLO family
- 3D Human-Human Interaction Anomaly Detection
- Verifying Rumors via Stance-Aware Structural Modeling
- Toward Agentic Environments: GenAI and the Convergence of AI, Sustainability, and Human-Centric Spaces
- MoLingo: Motion-Language Alignment for Text-to-Human Motion Generation
- Machine learning discovers new champion codes
- BlossomRec: Block-level Fused Sparse Attention Mechanism for Sequential Recommendations
- Beyond surface form: A pipeline for semantic analysis in Alzheimer's Disease detection from spontaneous speech
- A Scientific Reasoning Model for Organic Synthesis Procedure Generation
- Scaling Up AI-Generated Image Detection with Generator-Aware Prototypes
- StutterFuse: Mitigating Modality Collapse in Stuttering Detection with Jaccard-Weighted Metric Learning and Gated Fusion
- ReFusion: A Diffusion Large Language Model with Parallel Autoregressive Decoding
- Improving Recursive Transformers with Mixture of LoRAs
- Pancakes: Consistent Multi-Protocol Image Segmentation Across Biomedical Domains
- How Low Can You Go? The Data-Light SE Challenge
- A Deep Learning Model of Mental Rotation Informed by Interactive VR Experiments
- From Zipf's Law to Neural Scaling through Heaps' Law and Hilberg's Hypothesis
- Non-Resolution Reasoning (NRR): A Computational Framework for Contextual Identity and Ambiguity Preservation
- FIN-bench-v2: A Unified and Robust Benchmark Suite for Evaluating Finnish Large Language Models
- Security and Detectability Analysis of Unicode Text Watermarking Methods Against Large Language Models
- Dual-Qubit Hierarchical Fuzzy Neural Network for Image Classification: Enabling Relational Learning via Quantum Entanglement
- DisCo-Speech: Controllable Zero-Shot Speech Generation with A Disentangled Speech Codec
- Ego-EXTRA: video-language Egocentric Dataset for EXpert-TRAinee assistance
- Improving the Plausibility of Pressure Distributions Synthesized from Depth Image through Generative Modeling
- WAY: Estimation of Vessel Destination in Worldwide AIS Trajectory
- Enhancing Node-Level Graph Domain Adaptation by Alleviating Local Dependency
- Towards Unified Co-Speech Gesture Generation via Hierarchical Implicit Periodicity Learning
- Towards Practical Large-scale Dynamical Heterogeneous Graph Embedding: Cold-start Resilient Recommendation
- Autoregressive Neural Network Extrapolation of Quantum Spin Dynamics Across Time and Space
- Large Language Newsvendor: Decision Biases and Cognitive Mechanisms
- Time-aware UNet and super-resolution deep residual networks for spatial downscaling
- Scaling Bidirectional Spans and Span Violations in Attention Mechanism
- The algorithmic muse and the public domain: Why copyrights legal philosophy precludes protection for generative AI outputs
- Does Tone Change the Answer? Evaluating Prompt Politeness Effects on Modern LLMs: GPT, Gemini, and LLaMA
- From Small to Large: Generalization Bounds for Transformers on Variable-Size Inputs
- VLCache: Computing 2% Vision Tokens and Reusing 98% for Vision-Language Inference
- BLADE: A Behavior-Level Data Augmentation Framework with Dual Fusion Modeling for Multi-Behavior Sequential Recommendation
- Let the Model Learn to Feel: Mode-Guided Tonality Injection for Symbolic Music Emotion Recognition
- UAGLNet: Uncertainty-Aggregated Global-Local Fusion Network with Cooperative CNN-Transformer for Building Extraction
- Towards High-Fidelity Gaussian Splatting with Queried-Convolution Neural Networks
- Behavior and Representation in Large Language Models for Combinatorial Optimization: From Feature Extraction to Algorithm Selection
- Practical Hybrid Quantum Language Models with Observable Readout on Real Hardware
- DrivePI: Spatial-aware 4D MLLM for Unified Autonomous Driving Understanding, Perception, Prediction and Planning
- A Rule-Aware Prompt Framework for Structured Numeric Reasoning in Cyber-Physical Systems
- State over Tokens: Characterizing the Role of Reasoning Tokens
- GenieDrive: Towards Physics-Aware Driving World Model with 4D Occupancy Guided Video Generation
- FuXi-γ: Efficient Sequential Recommendation with Exponential-Power Temporal Encoder and Diagonal-Sparse Positional Mechanism
- Progressive Conditioned Scale-Shift Recalibration of Self-Attention for Online Test-time Adaptation
- Integrating Fourier Neural Operator with Diffusion Model for Autoregressive Predictions of Three-dimensional Turbulence
- Exact Flow Linear Attention: Exact Solution from Continuous-Time Dynamics
- TF-MCL: Time-frequency Fusion and Multi-domain Cross-Loss for Self-supervised Depression Detection
- TwinFormer: A Dual-Level Transformer for Long-Sequence Time-Series Forecasting
- GrowTAS: Progressive Expansion from Small to Large Subnets for Efficient ViT Architecture Search
- Anatomy Guided Coronary Artery Segmentation from CCTA Using Spatial Frequency Joint Modeling
- AttenGW: A Lightweight Attention-Based Multi-Detector Gravitational-Wave Detection Pipeline
- SafeGen: Embedding Ethical Safeguards in Text-to-Image Generation
- Breaking the Curse of Dimensionality: On the Stability of Modern Vector Retrieval
- Boosting Monocular Metric Depth Estimation via Bokeh Rendering
- Interference Effects in Resonant Standard Model di-Higgs Production and Decay into 4b Final States: the Role of Machine Learning Analysis
- Large and Small Model Collaboration for Air Interface
- BaRISTA: Brain Scale Informed Spatiotemporal Representation of Human Intracranial Neural Activity
- Neural CDEs as Correctors for Learned Time Series Models
- Advancing Cache-Based Few-Shot Classification via Patch-Driven Relational Gated Graph Attention
- Moment and Highlight Detection via MLLM Frame Segmentation
- VEGAS: Mitigating Hallucinations in Large Vision-Language Models via Vision-Encoder Attention Guided Adaptive Steering
- SigTime: Learning and Visually Explaining Time Series Signatures
- Machine learning methods for subpixel trajectory reconstruction in discretized position detectors
- Hold Onto That Thought: Assessing KV Cache Compression On Reasoning
- Fully Inductive Node Representation Learning via Graph View Transformation
- Multi-temporal Calving Front Segmentation
- 3DTeethSAM: Taming SAM2 for 3D Teeth Segmentation
- ACCOR: Attention-Enhanced Complex-Valued Contrastive Learning for Occluded Object Classification Using mmWave Radar IQ Signals
- Graph Embedding with Mel-spectrograms for Underwater Acoustic Target Recognition
- All-in-One ASR: Unifying Encoder-Decoder Models of CTC, Attention, and Transducer in Dual-Mode ASR
- xGR: Efficient Generative Recommendation Serving at Scale
- Contrastive Time Series Forecasting with Anomalies
- DREAM-B3P: Dual-Stream Transformer Network Enhanced by Feedback Diffusion Model for Blood-Brain Barrier Penetrating Peptide Prediction
- TSkel-Mamba: Temporal Dynamic Modeling via State Space Model for Human Skeleton-based Action Recognition
- Temporal-Anchor3DLane: Enhanced 3D Lane Detection with Multi-Task Losses and LSTM Fusion
- On the Dangers of Bootstrapping Generation for Continual Learning and Beyond
- Three methods, one problem: Classical and AI approaches to no-three-in-line
- AgentBalance: Backbone-then-Topology Design for Cost-Effective Multi-Agent Systems under Budget Constraints
- Sliced ReLU attention: Quasi-linear contextual expressivity via sorting
- PhraseVAE and PhraseLDM: Latent Diffusion for Full-Song Multitrack Symbolic Music Generation
- QGEC : Quantum Golay Code Error Correction
- AdaSD: Adaptive Speculative Decoding for Efficient Language Model Inference
- Do We Need Reformer for Vision? An Experimental Comparison with Vision Transformers
- Multi-Intent Spoken Language Understanding: Methods, Trends, and Challenges
- A Simple Generalisation of the Implicit Dynamics of In-Context Learning
- FAIR: Focused Attention Is All You Need for Generative Recommendation
- Task-Aware Multi-Expert Architecture For Lifelong Deep Learning
- ReactorFold: Generative discovery of nuclear reactor cores via emergent physical reasoning
- FutureX: Enhance End-to-End Autonomous Driving via Latent Chain-of-Thought World Model
- Seeing to Act, Prompting to Specify: A Bayesian Factorization of Vision Language Action Policy
- Transformer Embeddings for Fast Microlensing Inference
- Automating Historical Insight Extraction from Large-Scale Newspaper Archives via Neural Topic Modeling
- Surveillance Video-Based Traffic Accident Detection Using Transformer Architecture
- CAT: Can Trust be Predicted with Context-Awareness in Dynamic Heterogeneous Networks?
- Network and Compiler Optimizations for Efficient Linear Algebra Kernels in Private Transformer Inference
- Neuronal Attention Circuit (NAC) for Representation Learning
- Group Diffusion: Enhancing Image Generation by Unlocking Cross-Sample Collaboration
- Bidirectional Normalizing Flow: From Data to Noise and Back
- Towards Efficient and Effective Multi-Camera Encoding for End-to-End Driving
- MeViS: A Multi-Modal Dataset for Referring Motion Expression Video Segmentation
- GaussianHeadTalk: Wobble-Free 3D Talking Heads with Audio Driven Gaussian Splatting
- Stronger Normalization-Free Transformers
- Asynchronous Reasoning: Training-Free Interactive Thinking LLMs
- SparseSwaps: Tractable LLM Pruning Mask Refinement at Scale
- UrbanAI 2025 Challenge: Linear vs Transformer Models for Long-Horizon Exogenous Temperature Forecasting
- SplatPainter: Interactive Authoring of 3D Gaussians from 2D Edits via Test-Time Training
- Large Language Models for Superconductor Discovery
- An Elementary Proof of the Near Optimality of LogSumExp Smoothing
- Natural Language Interface for Firewall Configuration
- Template-Free Retrosynthesis with Graph-Prior Augmented Transformers
- LDP: Parameter-Efficient Fine-Tuning of Multimodal LLM for Medical Report Generation
- PMB-NN: Physiology-Centred Hybrid AI for Personalized Hemodynamic Monitoring from Photoplethysmography
- Error-Propagation-Free Learned Video Compression With Dual-Domain Progressive Temporal Alignment
- TransLocNet: Cross-Modal Attention for Aerial-Ground Vehicle Localization with Contrastive Learning
- The Best of the Two Worlds: Harmonizing Semantic and Hash IDs for Sequential Recommendation
- GPG: Generalized Policy Gradient Theorem for Transformer-based Policies
- Hybrid Transformer-Mamba for Weakly Supervised Volumetric Medical Segmentation
- Zero-shot Adaptation of Stable Diffusion via Plug-in Hierarchical Degradation Representation for Real-World Super-Resolution
- Point2Pose: A Generative Framework for 3D Human Pose Estimation with Multi-View Point Cloud Dataset
- Tracking large chemical reaction networks and rare events by neural networks
- Independent Density Estimation
- GDKVM: Echocardiography Video Segmentation via Spatiotemporal Key-Value Memory with Gated Delta Rule
- Robustness of Probabilistic Models to Low-Quality Data: A Multi-Perspective Analysis
- SWAA: Sliding Window Attention Adaptation for Efficient and Quality Preserving Long Context Processing
- Beyond Endpoints: Path-Centric Reasoning for Vectorized Off-Road Network Extraction
- FALCON: Few-step Accurate Likelihoods for Continuous Flows
- HybridVFL: Disentangled Feature Learning for Edge-Enabled Vertical Federated Multimodal Classification
- Scaling Behavior of Discrete Diffusion Language Models
- Mull-Tokens: Modality-Agnostic Latent Thinking
- Token Sample Complexity of Attention
- Generative Modeling of Entangled Polymers with a Distance-Based Variational Autoencoder
- DynaIP: Dynamic Image Prompt Adapter for Scalable Zero-shot Personalized Text-to-Image Generation
- DynaMate: An Autonomous Agent for Protein-Ligand Molecular Dynamics Simulations
- Unambiguous Representations in Neural Networks: An Information-Theoretic Approach to Intentionality
- Closing the Train-Test Gap in World Models for Gradient-Based Planning
- Simultaneous Tactile-Visual Perception for Learning Multimodal Robot Manipulation
- Composing Concepts from Images and Videos via Concept-prompt Binding
- Defining Cost Function of Steganography with Large Language Models
- Circuits, Features, and Heuristics in Molecular Transformers
- Mixture of Lookup Key-Value Experts
- Hands-on Evaluation of Visual Transformers for Object Recognition and Detection
- CUBE: A Cardinality Estimator Based on Neural CDF
- CS3D: An Efficient Facial Expression Recognition via Event Vision
- Supporting Dynamic Agentic Workloads: How Data and Agents Interact
- REASAN: Learning Reactive Safe Navigation for Legged Robots
- Transformers for Tabular Data: A Training Perspective of Self-Attention via Optimal Transport
- NeuroSketch: An Effective Framework for Neural Decoding via Systematic Architectural Optimization
- Temporal-Spatial Tubelet Embedding for Cloud-Robust MSI Reconstruction using MSI-SAR Fusion: A Multi-Head Self-Attention Video Vision Transformer Approach
- ODMA: On-Demand Memory Allocation Framework for LLM Serving on LPDDR-Class Accelerators
- Rates and architectures for learning geometrically non-trivial operators
- Self Distillation Fine-Tuning of Protein Language Models Improves Versatility in Protein Design
- A Distributed Framework for Privacy-Enhanced Vision Transformers on the Edge
- GoodSpeed: Optimizing Fair Goodput with Adaptive Speculative Decoding in Distributed Edge Inference
- FoundIR-v2: Optimizing Pre-Training Data Mixtures for Image Restoration Foundation Model
- Impact of Positional Encoding: Clean and Adversarial Rademacher Complexity for Transformers under In-Context Regression
- Supervised learning pays attention
- Efficient Continual Learning in Neural Machine Translation: A Low-Rank Adaptation Approach
- Stanford Sleep Bench: Evaluating Polysomnography Pre-training Methods for Sleep Foundation Models
- A Hybrid Model for Stock Market Forecasting: Integrating News Sentiment and Time Series Data with Graph Neural Networks
- Understanding Mental States in Active and Autonomous Driving with EEG
- Learning Patient-Specific Disease Dynamics with Latent Flow Matching for Longitudinal Imaging Generation
- Understanding temperature tuning in energy-based models
- DraCo: Draft as CoT for Text-to-Image Preview and Rare Concept Generation
- Language-Conditioned Safe Trajectory Generation for Spacecraft Rendezvous
- Modular Deep-Learning-Based Early Warning System for Deadly Heatwave Prediction
- KD-OCT: Efficient Knowledge Distillation for Clinical-Grade Retinal OCT Classification
- Efficiently Reconstructing Dynamic Scenes One D4RT at a Time
- SAQ: Stabilizer-Aware Quantum Error Correction Decoder
- InfiniteVL: Synergizing Linear and Sparse Attention for Highly-Efficient, Unlimited-Input Vision-Language Models
- Do Depth-Grown Models Overcome the Curse of Depth? An In-Depth Analysis
- Quantum Decision Transformers (QDT): Synergistic Entanglement and Interference for Offline Reinforcement Learning
- Data-Driven Dynamic Parameter Learning of manipulator robots
- A Multi-Robot Platform for Robotic Triage Combining Onboard Sensing and Foundation Models
- Skewness-Guided Pruning of Multimodal Swin Transformers for Federated Skin Lesion Classification on Edge Devices
- LaMoSys3.5D: Enabling 3.5D-IC-Based Large Language Model Inference Serving Systems via Hardware/Software Co-Design
- Engagement in Code Review: Emotional, Behavioral, and Cognitive Dimensions in Peer vs. LLM Interactions
- What really matters for person re-identification? A Mixture-of-Experts Framework for Semantic Attribute Importance
- Dual-Branch Center-Surrounding Contrast: Rethinking Contrastive Learning for 3D Point Clouds
- Protein Secondary Structure Prediction Using Transformers
- DMAGT: Unveiling miRNA-Drug Associations by Integrating SMILES and RNA Sequence Structures through Graph Transformer Models
- SensHRPS: Sensing Comfortable Human-Robot Proxemics and Personal Space With Eye-Tracking
- HATSolver: Learning Groebner Bases with Hierarchical Attention Transformers
- NeurIDA: Dynamic Modeling for Effective In-Database Analytics
- Soft Inductive Bias Approach via Explicit Reasoning Perspectives in Inappropriate Utterance Detection Using Large Language Models
- Solving Oversmoothing in GNNs via Nonlocal Message Passing: Algebraic Smoothing and Depth Scalability
- Transformers for Multimodal Brain State Decoding: Integrating Functional Magnetic Resonance Imaging Data and Medical Metadata
- Attention is All You Need to Defend Against Indirect Prompt Injection Attacks in LLMs
- The Unseen Bias: How Norm Discrepancy in Pre-Norm MLLMs Leads to Visual Information Loss
- Age-Inclusive 3D Human Mesh Recovery for Action-Preserving Data Anonymization
- PointDico: Contrastive 3D Representation Learning Guided by Diffusion Models
- Probabilistic Multi-Agent Aircraft Landing Time Prediction
- Is GPT-OSS All You Need? Benchmarking Large Language Models for Financial Intelligence and the Surprising Efficiency Paradox
- EgoX: Egocentric Video Generation from a Single Exocentric Video
- Distilling Future Temporal Knowledge with Masked Feature Reconstruction for 3D Object Detection
- Chopper: A Multi-Level GPU Characterization Tool & Derived Insights Into LLM Training Inefficiency
- SOP2: Transfer Learning with Scene-Oriented Prompt Pool on 3D Object Detection
- Restoring Network Evolution from Static Structure
- LayerPipe2: Multistage Pipelining and Weight Recompute via Improved Exponential Moving Average for Training Neural Networks
- Trajectory Densification and Depth from Perspective-based Blur
- Selfi: Self Improving Reconstruction Engine via 3D Geometric Feature Alignment
- Is Generation Required for Data-Efficient Perception?
- Understanding the Failure Modes of Transformers through the Lens of Graph Neural Networks
- ValuePilot: A Two-Phase Framework for Value-Driven Decision-Making
- Meta Lattice: Model Space Redesign for Cost-Effective Industry-Scale Ads Recommendations
- GeoDiffMM: Geometry-Guided Conditional Diffusion for Motion Magnification
- Scalable Offline Model-Based RL with Action Chunks
- Short-Context Dominance: How Much Local Context Natural Language Actually Needs?
- LUNA: Linear Universal Neural Attention with Generalization Guarantees
- Large Language Models for Education and Research: An Empirical and User Survey-based Analysis
- Bridging Code Graphs and Large Language Models for Better Code Understanding
- Benchmarking Offline Multi-Objective Reinforcement Learning in Critical Care
- Beyond Unified Models: A Service-Oriented Approach to Low Latency, Context Aware Phonemization for Real Time TTS
- Forecasting Dark Matter Subhalo Constraints from Stellar Streams using Implicit Likelihood Inference
- OneStory: Coherent Multi-Shot Video Generation with Adaptive Memory
- GatedFWA: Linear Flash Windowed Attention with Gated Associative Memory
- Enhancing Medical Cross-Modal Hashing Retrieval using Dropout-Voting Mixture-of-Experts Fusion
- The Native Spiking Microarchitecture: From Iontronic Primitives to Bit-Exact FP8 Arithmetic
- In-Context and Few-Shots Learning for Forecasting Time Series Data based on Large Language Models
- PVeRA: Probabilistic Vector-Based Random Matrix Adaptation
- A Large-Scale Multimodal Dataset and Benchmarks for Human Activity Scene Understanding and Reasoning
- MoCoRP: Modeling Consistent Relations between Persona and Response for Persona-based Dialogue
- Incorporating Structure and Chord Constraints in Symbolic Transformer-based Melodic Harmonization
- Weighted Contrastive Learning for Anomaly-Aware Time-Series Forecasting
- Flash Multi-Head Feed-Forward Network
- Beyond Real: Imaginary Extension of Rotary Position Embeddings for Long-Context LLMs
- AutoICE: Automatically Synthesizing Verifiable C Code via LLM-driven Evolution
- Amulet: Fast TEE-Shielded Inference for On-Device Model Protection
- InterAgent: Physics-based Multi-agent Command Execution via Diffusion on Interaction Graphs
- GlimmerNet: A Lightweight Grouped Dilated Depthwise Convolutions for UAV-Based Emergency Monitoring
- MultiAPI Spoof: A Multi-API Dataset and Local-Attention Network for Speech Anti-spoofing Detection
- Multi-Scale Protein Structure Modelling with Geometric Graph U-Nets
- PrivORL: Differentially Private Synthetic Dataset for Offline Reinforcement Learning
- M-STAR: Multi-Scale Spatiotemporal Autoregression for Human Mobility Modeling
- Radiance-Field Reinforced Pretraining: Scaling Localization Models with Unlabeled Wireless Signals
- Effective Attention-Guided Multi-Scale Medical Network for Skin Lesion Segmentation
- Unified Camera Positional Encoding for Controlled Video Generation
- Towards Benchmarking Design Pattern Detection Under Obfuscation: Reproducing and Evaluating Attention-Based Detection Method
- Attention as Binding: A Vector-Symbolic Perspective on Transformer Reasoning
- PICKT: Practical Interlinked Concept Knowledge Tracing for Personalized Learning using Knowledge Map Concept Relations
- Materium: An Autoregressive Approach for Material Generation
- TRACE: A Generalizable Drift Detector for Streaming Data-Driven Optimization
- Dual-Stream Cross-Modal Representation Learning via Residual Semantic Decorrelation
- SJD++: Improved Speculative Jacobi Decoding for Training-free Acceleration of Discrete Auto-regressive Text-to-Image Generation
- VulnLLM-R: Specialized Reasoning LLM with Agent Scaffold for Vulnerability Detection
- Enhancing Knowledge Transfer in Hyperspectral Image Classification via Cross-scene Knowledge Integration
- Multi-view Pyramid Transformer: Look Coarser to See Broader
- ThinkTrap: Denial-of-Service Attacks against Black-box LLM Services via Infinite Thinking
- FOAM: Blocked State Folding for Memory-Efficient LLM Training
- Leveraging KV Similarity for Online Structured Pruning in LLMs
- Dual Refinement Cycle Learning: Unsupervised Text Classification of Mamba and Community Detection on Text Attributed Graph
- CERNet: Class-Embedding Predictive-Coding RNN for Unified Robot Motion, Recognition, and Confidence Estimation
- Power of Boundary and Reflection: Semantic Transparent Object Segmentation using Pyramid Vision Transformer with Transparent Cues
- A Hetero-Associative Sequential Memory Model Utilizing Neuromorphic Signals: Validated on a Mobile Manipulator
- Task adaptation of Vision-Language-Action model: 1st Place Solution for the 2025 BEHAVIOR Challenge
- Progress Ratio Embeddings: An Impatience Signal for Robust Length Control in Neural Text Generation
- Large Language Models and Forensic Linguistics: Navigating Opportunities and Threats in the Age of Generative AI
- Towards Robust Pseudo-Label Learning in Semantic Segmentation: An Encoding Perspective
- Quantifying Memory Use in Reinforcement Learning with Temporal Range
- STAR-GO: Improving Protein Function Prediction by Learning to Hierarchically Integrate Ontology-Informed Semantic Embeddings
- SparseCoop: Cooperative Perception with Kinematic-Grounded Queries
- Optimal and Diffusion Transports in Machine Learning
- LLM as a Neural Architect: Controlled Generation of Image Captioning Models Under Strict API Contracts
- TreeQ: Pushing the Quantization Boundary of Diffusion Transformer via Tree-Structured Mixed-Precision Search
- A Patient-Doctor-NLP-System to contest inequality for less privileged
- Enhancing Interpretability of AR-SSVEP-Based Motor Intention Recognition via CNN-BiLSTM and SHAP Analysis on EEG Data
- Distribution-Aware Exploration for Adaptive HNSW Search
- FlatFormer: A Flat Transformer Knowledge Tracing Model Based on Cognitive Bias Injection
- Singing Timbre Popularity Assessment Based on Multimodal Large Foundation Model
- Always Keep Your Promises: A Model-Agnostic Attribution Algorithm for Neural Networks
- On Memory: A comparison of memory mechanisms in world models
- Scaling Zero-Shot Reference-to-Video Generation
- Graph-Regularized Sparse Autoencoders for LLM Safety Steering
- TextMamba: Scene Text Detector with Mamba
- Personalized Image Descriptions from Attention Sequences
- Hybrid Quantum-Classical Ensemble Learning for S&P 500 Directional Prediction
- SUGAR: A Sweeter Spot for Generative Unlearning of Many Identities
- BitStopper: An Efficient Transformer Attention Accelerator via Stage-fusion and Early Termination
- KANFormer for Predicting Fill Probabilities via Survival Analysis in Limit Order Books
- Efficient Text Classification with Conformal In-Context Learning
- Teaching Language Models Mechanistic Explainability Through Arrow-Pushing
- When Gender is Hard to See: Multi-Attribute Support for Long-Range Recognition
- Rethinking Training Dynamics in Scale-wise Autoregressive Generation
- AgenticCyber: A GenAI-Powered Multi-Agent System for Multimodal Threat Detection and Adaptive Response in Cybersecurity
- The Red Queen's Trap: Limits of Deep Evolution in High-Frequency Trading
- CLUENet: Cluster Attention Makes Neural Networks Have Eyes
- Control-Oriented System Identification: Classical, Learning, and Physics-Informed Approaches
- Chemistry Integrated Language Model using Hierarchical Molecular Representation for Polymer Informatics
- Unleashing the Intrinsic Visual Representation Capability of Multimodal Large Language Models
- From Remote Sensing to Multiple Time Horizons Forecasts: Transformers Model for CyanoHAB Intensity in Lake Champlain
- QL-LSTM: A Parameter-Efficient LSTM for Stable Long-Sequence Modeling
- Beyond Token-level Supervision: Unlocking the Potential of Decoding-based Regression via Reinforcement Learning
- Hierarchical geometric deep learning enables scalable analysis of molecular dynamics
- Physics-Informed Neural Koopman Machine for Interpretable Longitudinal Personalized Alzheimer's Disease Forecasting
- Correspondence-Oriented Imitation Learning: Flexible Visuomotor Control with 3D Conditioning
- KQ-SVD: Compressing the KV Cache with Provable Guarantees on Attention Fidelity
- SCAIL: Towards Studio-Grade Character Animation via In-Context Learning of 3D-Consistent Pose Representations
- TinyMyo: a Tiny Foundation Model for Flexible EMG Signal Processing at the Edge
- Context-Aware Model Predictive Control for Microgrid Energy Management via LLMs
- Evolutionary System 2 Reasoning: An Empirical Proof
- HQ-DM: Single Hadamard Transformation-Based Quantization-Aware Training for Low-Bit Diffusion Models
- Beyond Data Filtering: Knowledge Localization for Capability Removal in LLMs
- Fast SceneScript: Accurate and Efficient Structured Language Model via Multi-Token Prediction
- Ontology Learning with LLMs: A Benchmark Study on Axiom Identification
- Learning High-Fidelity Cloth Animation via Skinning-Free Image Transfer
- Automated Identification of Incidentalomas Requiring Follow-Up: A Multi-Anatomy Evaluation of LLM-Based and Supervised Approaches
- Know-Show: Benchmarking Video-Language Models on Spatio-Temporal Grounded Reasoning
- Rethinking Infrared Small Target Detection: A Foundation-Driven Efficient Paradigm
- Knowing Your Uncertainty -- On the application of LLM in social sciences
- IdealTSF: Can Non-Ideal Data Contribute to Enhancing the Performance of Time Series Forecasting Models?
- Smart Timing for Mining: A Deep Learning Framework for Bitcoin Hardware ROI Prediction
- ChipMind: Retrieval-Augmented Reasoning for Long-Context Circuit Design Specifications
- AI & Human Co-Improvement for Safer Co-Superintelligence
- OMTRA: A Multi-Task Generative Model for Structure-Based Drug Design
- BulletTime: Decoupled Control of Time and Camera Pose for Video Generation
- Fine-Tuning BERT for Domain-Specific Question Answering: Toward Educational NLP Resources at University Scale
- CNN on `Top': In Search of Scalable & Lightweight Image-based Jet Taggers
- Dual-Path Region-Guided Attention Network for Ground Reaction Force and Moment Regression
- RAMEN: Resolution-Adjustable Multimodal Encoder for Earth Observation
- DNA: Dual-branch Network with Adaptation for Open-Set Online Handwriting Generation
- Self-supervised prior learning improves structured illumination microscopy resolution
- GeoPE:A Unified Geometric Positional Embedding for Structured Tensors
- STELLA: Guiding Large Language Models for Time Series Forecasting with Semantic Abstractions
- Ask Safely: Privacy-Aware LLM Query Generation for Knowledge Graphs
- DAMASHA: Detecting AI in Mixed Adversarial Texts via Segmentation with Human-interpretable Attribution
- Setting up for failure: automatic discovery of the neural mechanisms of cognitive errors
- Matrix-Free Photoacoustic Image Reconstruction via Sensor-Token Self-Attention
- Generative Recursive Reasoning
- ArchSym: Detecting 3D-Grounded Architectural Symmetries in the Wild
- Group Selection as a Safeguard Against AI Substitution
- Parabolic Position Encoding: Vision-Centric, Principled, Extrapolatable, General
- QuaMo: Quaternion Motions for Vision-based 3D Human Kinematics Capture
- Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models
- Order Matters: 3D Shape Generation from Sequential VR Sketches
- Reward Forcing: Efficient Streaming Video Generation with Rewarded Distribution Matching Distillation
- Disentangling Progress in Medical Image Registration: Beyond Trend-Driven Architectures towards Domain-Specific Strategies
- Denoise to Track: Harnessing Video Diffusion Priors for Robust Correspondence
- Neural Decoding of Overt Speech from ECoG Using Vision Transformers and Contrastive Representation Learning
- QoSDiff: An Implicit Topological Embedding Learning Framework Leveraging Denoising Diffusion and Adversarial Attention for Robust QoS Prediction
- Shift-Window Meets Dual Attention: A Multi-Model Architecture for Specular Highlight Removal
- Controllable Long-term Motion Generation with Extended Joint Targets
- Self-Paced and Self-Corrective Masked Prediction for Movie Trailer Generation
- Solving LLM Repetition Problem in Production: A Comprehensive Study of Multiple Solutions
- Learnt Microwave Image Reconstruction with A Conformal Antenna Array
- Airport Passenger Flow Forecasting via Deformable Temporal-Spectral Transformer Approach
- FMA-Net++: Motion- and Exposure-Aware Real-World Joint Video Super-Resolution and Deblurring
- Vision and Causal Learning Based Channel Estimation for THz Communications
- RapidUn: Influence-Driven Parameter Reweighting for Efficient Large Language Model Unlearning
- Long-Horizon Model-Based Offline Reinforcement Learning Without Explicit Conservatism
- Detection and imaging of chemicals and hidden explosives using terahertz time-domain spectroscopy and deep learning
- Polynomiogram: An Integrated Framework for Root Visualization and Generative Art
- Training and Evaluation of Guideline-Based Medical Reasoning in LLMs
- Decentralized Social Media and Artificial Intelligence in Digital Public Health Monitoring
- Addressing Logical Fallacies In Scientific Reasoning From Large Language Models: Towards a Dual-Inference Training Framework
- MoReGen: Multi-Agent Motion-Reasoning Engine for Code-based Text-to-Video Synthesis
- Beyond Flicker: Detecting Kinematic Inconsistencies for Generalizable Deepfake Video Detection
- C3G: Learning Compact 3D Representations with 2K Gaussians
- On the Temporality for Sketch Representation Learning
- "All You Need" is Not All You Need for a Paper Title: On the Origins of a Scientific Meme
- Performance and efficiency of a transformer-based quark/gluon jet tagger in the ATLAS experiment
- GAOT: Generating Articulated Objects Through Text-Guided Diffusion Models
- A Theoretical Framework for Auxiliary-Loss-Free Load Balancing of Sparse Mixture-of-Experts in Large-Scale AI Models
- Reconstructing KV Caches with Cross-layer Fusion For Enhanced Transformers
- Heatmap Pooling Network for Action Recognition from RGB Videos
- EfficientECG: Cross-Attention with Feature Fusion for Efficient Electrocardiogram Classification
- LSRS: Latent Scale Rejection Sampling for Visual Autoregressive Modeling
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- CaFTRA: Frequency-Domain Correlation-Aware Feedback-Free MIMO Transmission and Resource Allocation for 6G and Beyond
- Tutorial on Large Language Model-Enhanced Reinforcement Learning for Wireless Networks
- Context-Aware Hierarchical Learning: A Two-Step Paradigm towards Safer LLMs
- VAT: Vision Action Transformer by Unlocking Full Representation of ViT
- FFTrainer: Fast Failover in Large-Language Model Training with Almost-Free State Management
- Cognitive Mirrors: Exploring the Diverse Functional Roles of Attention Heads in LLM Reasoning
- The promising potential of vision language models for the generation of textual weather forecasts
- SELF: A Robust Singular Value and Eigenvalue Approach for LLM Fingerprinting
- KVNAND: Efficient On-Device Large Language Model Inference Using DRAM-Free In-Flash Computing
- Observation-driven correction of numerical weather prediction for marine winds
- Memory-Guided Point Cloud Completion for Dental Reconstruction
- DF-Mamba: Deformable State Space Modeling for 3D Hand Pose Estimation in Interactions
- AdaPower: Specializing World Foundation Models for Predictive Manipulation
- Data-Free Pruning of Self-Attention Layers in LLMs
- Cross-Space Synergy: A Unified Framework for Multimodal Emotion Recognition in Conversation
- Exploiting Domain Properties in Language-Driven Domain Generalization for Semantic Segmentation
- SweetDeep: A Wearable AI Solution for Real-Time Non-Invasive Diabetes Screening
- Nexus: Higher-Order Attention Mechanisms in Transformers
- ShelfGaussian: Shelf-Supervised Open-Vocabulary Gaussian-based 3D Scene Understanding
- Idea-Gated Transformers: Enforcing Semantic Coherence via Differentiable Vocabulary Pruning
- Network of Theseus (like the ship)
- SeeU: Seeing the Unseen World via 4D Dynamics-aware Generation
- PretrainZero: Reinforcement Active Pretraining
- PSA: Pyramid Sparse Attention for Efficient Video Understanding and Generation
- DisentangleFormer: Spatial-Channel Decoupling for Multi-Channel Vision
- When, How Long and How Much? Interpretable Neural Networks for Time Series Regression by Learning to Mask and Aggregate
- When Do Symbolic Solvers Enhance Reasoning in Large Language Models?
- Decision Tree Embedding by Leaf-Means
- Nonlinear diffusion limit of non-local interactions on a sphere
- Video4Spatial: Towards Visuospatial Intelligence with Context-Guided Video Generation
- AutoBrep: Autoregressive B-Rep Generation with Unified Topology and Geometry
- Invasive Context Engineering to Control Large Language Models
- Flexible Gravitational-Wave Parameter Estimation with Transformers
- MSPT: Efficient Large-Scale Physical Modeling via Parallelized Multi-Scale Attention
- CoT4AD: A Vision-Language-Action Model with Explicit Chain-of-Thought Reasoning for Autonomous Driving
- Benchmarking machine learning models for multi-class state recognition in double quantum dot data
- TrackNetV5: Residual-Driven Spatio-Temporal Refinement and Motion Direction Decoupling for Fast Object Tracking
- PEFT-Factory: Unified Parameter-Efficient Fine-Tuning of Autoregressive Large Language Models
- PaCo-RL: Advancing Reinforcement Learning for Consistent Image Generation with Pairwise Reward Modeling
- Reasoning-Aware Multimodal Fusion for Hateful Video Detection
- Zero-Shot Instruction Following in RL via Structured LTL Representations
- G-PIFNN: A Generalizable Physics-informed Fourier Neural Network Framework for Electrical Circuits
- Leveraging Large-Scale Pretrained Spatial-Spectral Priors for General Zero-Shot Pansharpening
- Feedback Loops and Code Perturbations in LLM-based Software Engineering: A Case Study on a C-to-Rust Translation System
- Deep Learning-Based Joint Uplink-Downlink CSI Acquisition for Next-Generation Upper Mid-Band Systems
- DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
- AID: Agent Intent from Diffusion for Multi-Agent Informative Path Planning
- ESACT: An End-to-End Sparse Accelerator for Compute-Intensive Transformers via Local Similarity
- Attention-guided reference point shifting for Gaussian-mixture-based partial point set registration
- Temporal Dynamics Enhancer for Directly Trained Spiking Object Detectors
- Training Data Attribution for Image Generation using Ontology-Aligned Knowledge Graphs
- From Detection to Association: Learning Discriminative Object Embeddings for Multi-Object Tracking
- Safeguarded Stochastic Polyak Step Sizes for Non-smooth Optimization: Robust Performance Without Small (Sub)Gradients
- Associative Memory using Attribute-Specific Neuron Groups-1: Learning between Multiple Cue Balls
- Leveraging generative adversarial networks with spatially adaptive denormalization for multivariate stochastic seismic data inversion
- The brain-AI convergence: Predictive and generative world models for general-purpose computation
- TabGRU: An Enhanced Design for Urban Rainfall Intensity Estimation Using Commercial Microwave Links
- Q-KVComm: Efficient Multi-Agent Communication Via Adaptive KV Cache Compression
- Beyond N-grams: A Hierarchical Reward Learning Framework for Clinically-Aware Medical Report Generation
- Empathy Level Prediction in Multi-Modal Scenario with Supervisory Documentation Assistance
- SAGE: Style-Adaptive Generalization for Privacy-Constrained Semantic Segmentation Across Domains
- VLM-Pruner: Buffering for Spatial Sparsity in an Efficient VLM Centrifugal Token Pruning Paradigm
- LeechHijack: Covert Computational Resource Exploitation in Intelligent Agent Systems
- PULSE-ICU: A Pretrained Unified Long-Sequence Encoder for Multi-task Prediction in Intensive Care Units
- Spatiotemporal Pyramid Flow Matching for Climate Emulation
- StructuredDNA: A Bio-Physical Framework for Energy-Aware Transformer Routing
- WhAM: Towards A Translative Model of Sperm Whale Vocalization
- Story2MIDI: Emotionally Aligned Music Generation from Text
- Context-Enriched Contrastive Loss: Enhancing Presentation of Inherent Sample Connections in Contrastive Learning Framework
- Feature Selection Empowered BERT for Detection of Hate Speech with Vocabulary Augmentation
- Generative Video Motion Editing with 3D Point Tracks
- Improved Mean Flows: On the Challenges of Fastforward Generative Models
- Ensemble Privacy Defense for Knowledge-Intensive LLMs against Membership Inference Attacks
- KV Pareto: Systems-Level Optimization of KV Cache and Model Compression for Long Context Inference
- Script: Graph-Structured and Query-Conditioned Semantic Token Pruning for Multimodal Large Language Models
- SVRG and Beyond via Posterior Correction
- Rectifying LLM Thought from Lens of Optimization
- Latent Debate: A Surrogate Framework for Interpreting LLM Thinking
- New Spiking Architecture for Multi-Modal Decision-Making in Autonomous Vehicles
- Testing Transformer Learnability on the Arithmetic Sequence of Rooted Trees
- Cross-Lingual Interleaving for Speech Language Models
- Topological Order in Deep State
- OpenREAD: Reinforced Open-Ended Reasoning for End-to-End Autonomous Driving with LLM-as-Critic
- Generative Action Tell-Tales: Assessing Human Motion in Synthesized Videos
- Robust Rigid and Non-Rigid Medical Image Registration Using Learnable Edge Kernels
- Probabilistic Neuro-Symbolic Reasoning for Sparse Historical Data: A Framework Integrating Bayesian Inference, Causal Models, and Game-Theoretic Allocation
- A unified framework for geometry-independent operator learning in cardiac electrophysiology simulations
- ViT3: Unlocking Test-Time Training in Vision
- Accelerated Machine Learning Force Field for Predicting Thermal Conductivity of Organic Liquids
- Parallel Delayed Memory Units for Enhanced Temporal Modeling in Biomedical and Bioacoustic Signal Analysis
- In-Context Learning for Deep Joint Source-Channel Coding Over MIMO Channels
- RoMe: Row Granularity Access Memory System for Large Language Models
- FlashVGGT: Efficient and Scalable Visual Geometry Transformers with Compressed Descriptor Attention
- Neural Networks for Predicting Permeability Tensors of 2D Porous Media: Comparison of Convolution- and Transformer-based Architectures
- Heuristic algorithms for the stochastic critical node detection problem
- CourtMotion: Learning Event-Driven Motion Representations from Skeletal Data for Basketball
- Differentiable Weightless Controllers: Learning Logic Circuits for Continuous Control
- MEGConformer: Conformer-Based MEG Decoder for Robust Speech and Phoneme Classification
- Syndrome-Flow Consistency Model Achieves One-step Denoising Error Correction Codes
- Handwritten Text Recognition for Low Resource Languages
- Securing Large Language Models (LLMs) from Prompt Injection Attacks
- Panda: Self-distillation of Reusable Sensor-level Representations for High Energy Physics
- Data-Driven Learnability Transition of Measurement-Induced Entanglement
- AI-Enabled grading with near-domain data for scaling feedback with human-level accuracy
- KVReviver: Reversible KV Cache Compression with Sketch-Based Token Reconstruction
- Efficient Training of Diffusion Mixture-of-Experts Models: A Practical Recipe
- SceneProp: Combining Neural Network and Markov Random Field for Scene-Graph Grounding
- Learning to Reconstruct Temperature Field from Sparse Observations with Implicit Physics Priors
- Teaching by Failure: Counter-Example-Driven Curricula for Transformer Self-Improvement
- MindFuse: Towards GenAI Explainability in Marketing Strategy Co-Creation
- KidSpeak: A General Multi-purpose LLM for Kids' Speech Recognition and Screening
- A Self-explainable Model of Long Time Series by Extracting Informative Structured Causal Patterns
- MDiff4STR: Mask Diffusion Model for Scene Text Recognition
- Masked Symbol Modeling for Demodulation of Oversampled Baseband Communication Signals in Impulsive Noise-Dominated Channels
- fMRI2GES: Co-speech Gesture Reconstruction from fMRI Signal with Dual Brain Decoding Alignment
- Fantastic Features and Where to Find Them: A Probing Method to combine Features from Multiple Foundation Models
- Structural Prognostic Event Modeling for Multimodal Cancer Survival Analysis
- Estimation of Kinematic Motion from Dashcam Footage
- Learning Eigenstructures of Unstructured Data Manifolds
- Testing the Machine Consciousness Hypothesis
- CycleManip: Enabling Cyclic Task Manipulation via Effective Historical Perception and Understanding
- Upper Approximation Bounds for Neural Oscillators
- FDRMFL:Multi-modal Federated Feature Extraction Model Based on Information Maximization and Contrastive Learning
- MM-ACT: Learn from Multimodal Parallel Generation to Act
- Table as a Modality for Large Language Models
- LAHNet: Local Attentive Hashing Network for Point Cloud Registration
- Dialect Identification Using Resource-Efficient Fine-Tuning Approaches
- TalkingPose: Efficient Face and Gesture Animation with Feedback-guided Diffusion Model
- Accelerating Streaming Video Large Language Models via Hierarchical Token Compression
- Feed-Forward 3D Gaussian Splatting Compression with Long-Context Modeling
- Robust Probabilistic Load Forecasting for a Single Household: A Comparative Study from SARIMA to Transformers on the REFIT Dataset
- Auxiliary-Hyperparameter-Free Sampling: Entropy Equilibrium for Text Generation
- Sigma: The Key for Vision-Language-Action Models toward Telepathic Alignment
- Dynamic-eDiTor: Training-Free Text-Driven 4D Scene Editing with Multimodal Diffusion Transformer
- MPR-GUI: Benchmarking and Enhancing Multilingual Perception and Reasoning in GUI Agents
- Beyond Topology: A Morphological Symmetry Graph Representation for Locomotion Policy Learning
- SpeContext: Enabling Efficient Long-context Reasoning with Speculative Context Sparsity in LLMs
- VFM-ISRefiner: Towards Better Adapting Vision Foundation Models for Interactive Segmentation of Remote Sensing Images
- DPWMixer: Dual-Path Wavelet Mixer for Long-Term Time Series Forecasting
- Interpretable Alzheimer's Diagnosis via Multimodal Fusion of Regional Brain Experts
- Hierarchical Molecular Language Models (HMLMs)
- Silhouette-based Gait Foundation Model
- LAP: Fast LAtent Diffusion Planner for Autonomous Driving
- Bias Testing and Mitigation in Black Box LLMs using Metamorphic Relations
- Relightable Holoported Characters: Capturing and Relighting Dynamic Human Performance from Sparse Views
- TrendGNN: Towards Understanding of Epidemics, Beliefs, and Behaviors
- MambaScope: Coarse-to-Fine Scoping for Efficient Vision Mamba
- Financial Text Classification Based On rLoRA Finetuning On Qwen3-8B model
- G-KV: Decoding-Time KV Cache Eviction with Global Attention
- CACARA: Cross-Modal Alignment Leveraging a Text-Centric Approach for Cost-Effective Multimodal and Multilingual Learning
- Digital Surfactant
- Simplex-Optimized Hybrid Ensemble for Large Language Model Text Detection Under Generative Distribution Drif
- Structured Context Learning for Generic Event Boundary Detection
- GreenPlanner: Practical Floorplan Layout Generation via an Energy-Aware and Function-Feasible Generative Framework
- SelfAI: Building a Self-Training AI System with LLM Agents
- Efficient and Programmable Exploration of Synthesizable Chemical Space
- EnzyCLIP: A Cross-Attention Dual Encoder Framework with Contrastive Learning for Predicting Enzyme Kinetic Constants
- The Information Theory of Similarity
- Vision Transformer for Classification of UAV and Helicopters Using Micro-Doppler Spectrograms in Surveillance Radar
- Comparative Analysis of 47 Context-Based Question Answer Models Across 8 Diverse Datasets
- TGSFormer: Scalable Temporal Gaussian Splatting for Embodied Semantic Scene Completion
- FiCoTS: Fine-to-Coarse LLM-Enhanced Hierarchical Cross-Modality Interaction for Time Series Forecasting
- BioArc: Discovering Optimal Neural Architectures for Biological Foundation Models
- InvarDiff: Cross-Scale Invariance Caching for Accelerated Diffusion Models
- HIMOSA: Efficient Remote Sensing Image Super-Resolution with Hierarchical Mixture of Sparse Attention
- Quantum Private Distributed Matrix Multiplication With Degree Tables
- FR-TTS: Test-Time Scaling for NTP-based Image Generation with Effective Filling-based Reward Signal
- Dressing composite fermions with artificial intelligence
- CodeFlowLM: Incremental Just-In-Time Defect Prediction with Pretrained Language Models and Exploratory Insights into Defect Localization
- ReactionMamba: Generating Short &Long Human Reaction Sequences
- Start Making Sense(s): A Developmental Probe of Attention Specialization Using Lexical Ambiguity
- Quantum Circuit Reasoning Models: A Variational Framework for Differentiable Logical Inference
- MOTIF-RF: Multi-template On-chip Transformer Synthesis Incorporating Frequency-domain Self-transfer Learning for RFIC Design Automation
- Video-R2: Reinforcing Consistent and Grounded Reasoning in Multimodal Language Models
- AnyTalker: Scaling Multi-Person Talking Video Generation with Interactivity Refinement
- Heterogeneous Multi-Agent Reinforcement Learning with Attention for Cooperative and Scalable Feature Transformation
- Rethinking AI Evaluation in Education: The TEACH-AI Framework and Benchmark for Generative AI Assistants
- Optimizing Multimodal Language Models through Attention-based Interpretability
- Distributed Dynamic Associative Memory via Online Convex Optimization
- Memory-Amortized Inference: A Topological Unification of Search, Closure, and Structure
- The Geometry of Certainty: Recursive Topological Condensation and the Limits of Inference
- Robust HRRP Recognition under Interrupted Sampling Repeater Jamming using a Prior Jamming Information-Guided Network
- Towards Understanding Transformers in Learning Random Walks
- Unlocking Multilingual Reasoning Capability of LLMs and LVLMs through Representation Engineering
- Instruction Tuning of Large Language Models for Tabular Data Generation-in One Day
- Vision Bridge Transformer at Scale
- Data-Efficient Motor Condition Monitoring with Time Series Foundation Models
- Fast Multi-view Consistent 3D Editing with Video Priors
- An Empirical Study on the Security Vulnerabilities of GPTs
- Estimating the Event-Related Potential from Few EEG Trials
- Cascaded Robust Rectification for Arbitrary Document Images
- LUMOS: Large User MOdels for User Behavior Prediction
- Standard Occupation Classifier -- A Natural Language Processing Approach
- Delta-XAI: A Unified Framework for Explaining Prediction Changes in Online Time Series Monitoring
- Breaking Scale Anchoring: Frequency Representation Learning for Accurate High-Resolution Inference from Low-Resolution Training
- Masked Diffusion for Generative Recommendation
- Analysis of Invasive Breast Cancer in Mammograms Using YOLO, Explainability, and Domain Adaptation
- Convolutional Feature Noise Reduction for 2D Cardiac MR Image Segmentation
- Pooling Attention: Evaluating Pretrained Transformer Embeddings for Deception Classification
- Training-Free Loosely Speculative Decoding: Accepting Semantically Correct Drafts Beyond Exact Match
- GSPN-2: Efficient Parallel Sequence Modeling
- A Customer Journey in the Land of Oz: Leveraging the Wizard of Oz Technique to Model Emotions in Customer Service Interactions
- CoordSpeaker: Exploiting Gesture Captioning for Coordinated Caption-Empowered Co-Speech Gesture Generation
- PerfMamba: Performance Analysis and Pruning of Selective State Space Models
- The Immutable Tensor Architecture: A Pure Dataflow Approach for Secure, Energy-Efficient AI Inference
- Intelligent Neural Networks: From Layered Architectures to Graph-Organized Intelligence
- Robust Image Self-Recovery against Tampering using Watermark Generation with Pixel Shuffling
- Language-conditioned world model improves policy generalization by reading environmental descriptions
- Artwork Interpretation with Vision Language Models: A Case Study on Emotions and Emotion Symbols
- InstanceV: Instance-Level Video Generation
- Every Token Counts: Generalizing 16M Ultra-Long Context in Large Language Models
- Closed-Loop Transformers: Autoregressive Modeling as Iterative Latent Equilibrium
- Alzheimer's Disease Prediction Using EffNetViTLoRA and BiLSTM with Multimodal Longitudinal MRI Data
- Modeling Romanized Hindi and Bengali: Dataset Creation and Multilingual LLM Integration
- A3T-GCN for FTSE100 Components Price Forecasting
- CoFiRec: Coarse-to-Fine Tokenization for Generative Recommendation
- Invisible Hands: Gray-Box Bit Flip Attack for Steering LLMs Without Knowledge of Gradients, Data, and Weights
- Generative Anchored Fields: Controlled Data Generation via Emergent Velocity Fields and Transport Algebra
- Modèles de Fondation et Ajustement : Vers une Nouvelle Génération de Modèles pour la Prévision des Séries Temporelles
- Generative models for crystalline materials
- OOCO: Latency-disaggregated Architecture for Online-Offline Co-locate LLM Serving
- Hard Spatial Gating for Precision-Driven Brain Metastasis Segmentation: Addressing the Over-Segmentation Paradox in Deep Attention Networks
- TransCoder: A Neural-Enhancement Framework for Channel Codes
- Fast3Dcache: Training-free 3D Geometry Synthesis Acceleration
- Orthogonal Disentanglement with Projected Feature Alignment for Multimodal Emotion Recognition in Conversation
- IPDiff: Diffusion-driven ORSI Salient Object Detection with Information Reconstruction and Multi-Prior Guidance
- Enhancing Trustworthiness with Mixed Precision: Benchmarks, Opportunities, and Challenges
- Adversarial Flow Models
- ARPGNet: Appearance- and Relation-aware Parallel Graph Attention Fusion Network for Facial Expression Recognition
- Angle-Optimized Partial Disentanglement for Multimodal Emotion Recognition in Conversation
- Seeing without Pixels: Perception from Camera Trajectories
- Efficient-Husformer: Efficient Multimodal Transformer Hyperparameter Optimization for Stress and Cognitive Loads
- FADiff: Fusion-Aware Differentiable Optimization for DNN Scheduling on Tensor Accelerators
- Unexplored flaws in multiple-choice VQA evaluations
- Swarms of Large Language Model Agents for Protein Sequence Design with Experimental Validation
- Improving Sparse IMU-based Motion Capture with Motion Label Smoothing
- Fourier-Enhanced Recurrent Neural Networks for Electrical Load Time Series Downscaling
- Shoe Style-Invariant and Ground-Aware Learning for Dense Foot Contact Estimation
- Harmonic-Percussive Disentangled Neural Audio Codec for Bandwidth Extension
- The Age-specific Alzheimer 's Disease Prediction with Characteristic Constraints in Nonuniform Time Span
- IntAttention: A Fully Integer Attention Pipeline for Efficient Edge Inference
- SAM Guided Semantic and Motion Changed Region Mining for Remote Sensing Change Captioning
- DiverseVAR: Balancing Diversity and Quality of Next-Scale Visual Autoregressive Models
- Biomimetic Metamaterial-based Interface for Decoding Heterogeneous Mechanodermal Activity
- CanKD: Cross-Attention-based Non-local operation for Feature-based Knowledge Distillation
- Frequency-Aware Token Reduction for Efficient Vision Transformer
- MobileI2V: Fast and High-Resolution Image-to-Video on Mobile Devices
- Hierarchical Ranking Neural Network for Long Document Readability Assessment
- One-Step Generative Policies with Q-Learning: A Reformulation of MeanFlow
- SUPN: Shallow Universal Polynomial Networks
- Subjective Depth and Timescale Transformers: Learning Where and When to Compute
- SetAD: Semi-Supervised Anomaly Learning in Contextual Sets
- Controlling changes to attention logits
- Evaluation of Large Language Models for Numeric Anomaly Detection in Power Systems
- Are Graph Transformers Necessary? Efficient Long-Range Message Passing with Fractal Nodes in MPNNs
- Infinite-Story: A Training-Free Consistent Text-to-Image Generation
- Exploring Automated Recognition of Instructional Activity and Discourse from Multimodal Classroom Data
- PathMamba: A Hybrid Mamba-Transformer for Topologically Coherent Road Segmentation in Satellite Imagery
- PEFT-Bench: A Parameter-Efficient Fine-Tuning Methods Benchmark
- Co-Training Vision Language Models for Remote Sensing Multi-task Learning
- Unlocking Zero-shot Potential of Semi-dense Image Matching via Gaussian Splatting
- AuthenLoRA: Entangling Stylization with Imperceptible Watermarks for Copyright-Secure LoRA Adapters
- Transformer Driven Visual Servoing for Fabric Texture Matching Using Dual-Arm Manipulator
- Referring Video Object Segmentation with Cross-Modality Proxy Queries
- Enhancing Burmese News Classification with Kolmogorov-Arnold Network Head Fine-tuning
- Context-Aware Pragmatic Metacognitive Prompting for Sarcasm Detection
- A Unified Understanding of Offline Data Selection and Online Self-refining Generation for Post-training LLMs
- RED-F: Reconstruction-Elimination based Dual-stream Contrastive Forecasting for Multivariate Time Series Anomaly Prediction
- FlowerDance: MeanFlow for Efficient and Refined 3D Dance Generation
- Gated KalmaNet: A Fading Memory Layer Through Test-Time Ridge Regression
- DOPD: A Dynamic PD-Disaggregation Architecture for Maximizing Goodput in LLM Inference Serving
- Semantic Superiority vs. Forensic Efficiency: A Comparative Analysis of Deep Learning and Psycholinguistics for Business Email Compromise Detection
- Chatty-KG: A Multi-Agent AI System for On-Demand Conversational Question Answering over Knowledge Graphs
- On the Origin of Algorithmic Progress in AI
- LAPA: Log-Domain Prediction-Driven Dynamic Sparsity Accelerator for Transformer Model
- G2VLM: Geometry Grounded Vision Language Model with Unified 3D Reconstruction and Spatial Reasoning
- RLM: A Vision-Language Model Approach for Radar Scene Understanding
- Towards an Effective Action-Region Tracking Framework for Fine-grained Video Action Recognition
- Primal: A Unified Deterministic Framework for Quasi-Orthogonal Hashing and Manifold Learning
- RefTr: Recurrent Refinement of Confluent Trajectories for 3D Vascular Tree Centerline Graphs
- Intriguing Properties of Dynamic Sampling Networks
- Physics Steering: Causal Control of Cross-Domain Concepts in a Physics Foundation Model
- Dynamical Properties of Tokens in Self-Attention and Effects of Positional Encoding
- Revisiting KRISP: A Lightweight Reproduction and Analysis of Knowledge-Enhanced Vision-Language Models
- Adversarial Multi-Task Learning for Liver Tumor Segmentation, Dynamic Enhancement Regression, and Classification
- Unleashing the Power of Vision-Language Models for Long-Tailed Multi-Label Visual Recognition
- Image2Gcode: Image-to-G-code Generation for Additive Manufacturing Using Diffusion-Transformer Model
- Evaluating the Performance of Deep Learning Models in Whole-body Dynamic 3D Posture Prediction During Load-reaching Activities
- Adaptive Hopfield Network: Rethinking Similarities in Associative Memory
- The Driver-Blindness Phenomenon: Why Deep Sequence Models Default to Autocorrelation in Blood Glucose Forecasting
- From Words to Wisdom: Discourse Annotation and Baseline Models for Student Dialogue Understanding
- Mistake Attribution: Fine-Grained Mistake Understanding in Egocentric Videos
- MoRE: Batch-Robust Multi-Omics Representations from Frozen Pre-trained Transformers
- Fluid Intelligence: A Forward Look on AI Foundation Models in Computational Fluid Dynamics
- BRIC: Bridging Kinematic Plans and Physical Control at Test Time
- Automated Histopathologic Assessment of Hirschsprung Disease Using a Multi-Stage Vision Transformer Framework
- Block Cascading: Training Free Acceleration of Block-Causal Video Models
- TaCo: Capturing Spatio-Temporal Semantic Consistency in Remote Sensing Change Detection
- CrossEarth-Gate: Fisher-Guided Adaptive Tuning Engine for Efficient Adaptation of Cross-Domain Remote Sensing Semantic Segmentation
- AI-Designed Photonics Gratings with Experimental Verification
- LLM-Driven Transient Stability Assessment: From Automated Simulation to Neural Architecture Design
- Interpretable Air Pollution Forecasting by Physics-Guided Spatiotemporal Decoupling
- HHFT: Hierarchical Heterogeneous Feature Transformer for Recommendation Systems
- Denoising gravitational wave with deep learning in the time-frequency domain
- FINE: Factorized multimodal sentiment analysis via mutual INformation Estimation
- UltraViCo: Breaking Extrapolation Limits in Video Diffusion Transformers
- V-Attack: Targeting Disentangled Value Features for Controllable Adversarial Attacks on LVLMs
- Softmax Transformers are Turing-Complete
- Cross-Contrastive Clustering for Multimodal Attributed Graphs with Dual Graph Filtering
- Foundry: Distilling 3D Foundation Models for the Edge
- ACIT: Attention-Guided Cross-Modal Interaction Transformer for Pedestrian Crossing Intention Prediction
- Pedestrian Crossing Intention Prediction Using Multimodal Fusion Network
- REWA: A General Theory of Witness-Based Similarity
- Operator Learning at Machine Precision
- Hierarchical Spatio-Temporal Attention Network with Adaptive Risk-Aware Decision for Forward Collision Warning in Complex Scenarios
- Low-Resolution Editing is All You Need for High-Resolution Editing
- AI/ML based Joint Source and Channel Coding for HARQ-ACK Payload
- LLM-EDT: Large Language Model Enhanced Cross-domain Sequential Recommendation with Dual-phase Training
- MHB: Multimodal Handshape-aware Boundary Detection for Continuous Sign Language Recognition
- Frailty-Aware Transformer for Recurrent Survival Modeling of Driver Retention in Ride-Hailing Platforms
- 3M-TI: High-Quality Mobile Thermal Imaging via Calibration-free Multi-Camera Cross-Modal Diffusion
- Rectified SpaAttn: Revisiting Attention Sparsity for Efficient Video Generation
- Language-Independent Sentiment Labelling with Distant Supervision: A Case Study for English, Sepedi and Setswana
- Zero-Knowledge Proof Based Verifiable Inference of Models
- Explicit Tonal Tension Conditioning via Dual-Level Beam Search for Symbolic Music Generation
- SSA: Sparse Sparse Attention by Aligning Full and Sparse Attention Outputs in Feature Space
- AdaCap: An Adaptive Contrastive Approach for Small-Data Neural Networks
- In-Context Compositional Learning via Sparse Coding Transformer
- Mispronunciation Detection and Diagnosis Without Model Training: A Retrieval-Based Approach
- Latent Collaboration in Multi-Agent Systems
- Enhancing low energy reconstruction and classification in KM3NeT/ORCA with transformers
- Learning to Solve Weighted Maximum Satisfiability with a Co-Training Architecture
- TiCT: A Synthetically Pre-Trained Foundation Model for Time Series Classification
- TREASURE: The Visa Payment Foundation Model for High-Volume Transaction Understanding
- Pretraining Transformer-Based Models on Diffusion-Generated Synthetic Graphs for Alzheimer's Disease Prediction
- On the Utility of Foundation Models for Fast MRI: Vision-Language-Guided Image Reconstruction
- Flow Map Distillation Without Data
- Learning Massively Multitask World Models for Continuous Control
- Neural surrogates for designing gravitational wave detectors
- Understanding the Staged Dynamics of Transformers in Learning Latent Structure
- MapFormer: Self-Supervised Learning of Cognitive Maps with Input-Dependent Positional Embeddings
- Solar-GECO: Perovskite Solar Cell Property Prediction with Geometric-Aware Co-Attention
- SimDiff: Simpler Yet Better Diffusion Model for Time Series Point Forecasting
- SENTINEL: A Fully End-to-End Language-Action Model for Humanoid Whole Body Control
- In Machina N400: Pinpointing Where a Causal Language Model Detects Semantic Violations
- Percept-WAM: Perception-Enhanced World-Awareness-Action Model for Robust End-to-End Autonomous Driving
- Graph-based 3D Human Pose Estimation using WiFi Signals
- Extracting Robust Register Automata from Neural Networks over Data Sequences
- Large Language Models as Search Engines: Societal Challenges
- Changes in Gaza: DINOv3-Powered Multi-Class Change Detection for Damage Assessment in Conflict Zones
- Dynamic Granularity Matters: Rethinking Vision Transformers Beyond Fixed Patch Splitting
- 3D Dynamic Radio Map Prediction Using Vision Transformers for Low-Altitude Wireless Networks
- Eevee: Towards Close-up High-resolution Video-based Virtual Try-on
- Compressor-VLA: Instruction-Guided Visual Token Compression for Efficient Robotic Manipulation
- On the Role of Hidden States of Modern Hopfield Network in Transformer
- Medal S: Spatio-Textual Prompt Model for Medical Segmentation
- Parallel Vision Token Scheduling for Fast and Accurate Multimodal LMMs Inference
- DualGazeNet: A Biologically Inspired Dual-Gaze Query Network for Salient Object Detection
- Generating Reading Comprehension Exercises with Large Language Models for Educational Applications
- Cognitive Alpha Mining via LLM-Driven Code-Based Evolution
- Large Language Models for the Summarization of Czech Documents: From History to the Present
- Personalized Federated Segmentation with Shared Feature Aggregation and Boundary-Focused Calibration
- DiP: Taming Diffusion Models in Pixel Space
- SAOT: An Enhanced Locality-Aware Spectral Transformer for Solving PDEs
- Think How Your Teammates Think: Active Inference Can Benefit Decentralized Execution
- SpeedAug: Policy Acceleration via Tempo-Enriched Policy and RL Fine-Tuning
- Robust Multimodal Sentiment Analysis of Image-Text Pairs by Distribution-Based Feature Recovery and Fusion
- Multimodal Large Language Models with Adaptive Preference Optimization for Sequential Recommendation
- Large-Scale In-Game Outcome Forecasting for Match, Team and Players in Football using an Axial Transformer Neural Network
- LATTICE: Democratize High-Fidelity 3D Generation at Scale
- When and What to Recommend: Joint Modeling of Timing and Content for Active Sequential Recommendation
- Modality-Collaborative Low-Rank Decomposers for Few-Shot Video Domain Adaptation
- Multimodal Real-Time Anomaly Detection and Industrial Applications
- Phase-Aligned RoPE for Mixed-Resolution Diffusion Transformer
- Towards Efficient VLMs: Information-Theoretic Driven Compression via Adaptive Structural Pruning
- GRIT-LP: Graph Transformer with Long-Range Skip Connection and Partitioned Spatial Graphs for Accurate Ice Layer Thickness Prediction
- Unsupervised Multi-View Visual Anomaly Detection via Progressive Homography-Guided Alignment
- Artificial Intelligence Driven Workflow for Accelerating Design of Novel Photosensitizers
- An Invariant Latent Space Perspective on Language Model Inversion
- GContextFormer: A global context-aware hybrid multi-head attention approach with scaled additive aggregation for multimodal trajectory prediction
- CoD: A Diffusion Foundation Model for Image Compression
- Leveraging Language Models for Interpretable Analysis of Narratives in a Large Corpus
- Dealing with the Hard Facts of Low-Resource African NLP
- TimePre: Bridging Accuracy, Efficiency, and Stability in Probabilistic Time-Series Forecasting
- Foundations of Artificial Intelligence Frameworks: Notion and Limits of AGI
- Adaptive Mesh-Quantization for Neural PDE Solvers
- Gaze Beyond the Frame: Forecasting Egocentric 3D Visual Span
- A Systematic Study of Compression Ordering for Large Language Models
- RNN as Linear Transformer: A Closer Investigation into Representational Potentials of Visual Mamba Models
- Progressive Localisation in Localist LLMs
- The Catastrophic Paradox of Human Cognitive Frameworks in Large Language Model Evaluation: A Comprehensive Empirical Analysis of the CHC-LLM Incompatibility
- RoadSceneVQA: Benchmarking Visual Question Answering in Roadside Perception Systems for Intelligent Transportation System
- Vision Token Masking Alone Cannot Prevent PHI Leakage in Medical Document OCR: A Systematic Evaluation
- Coherent Multi-Agent Trajectory Forecasting in Team Sports with CausalTraj
- Protein Set Transformer: a protein-based genome language model to power high-diversity viromics
- PeriodNet: Boosting the Potential of Attention Mechanism for Time Series Forecasting
- MimiCAT: Mimic with Correspondence-Aware Cascade-Transformer for Category-Free 3D Pose Transfer
- Typing Reinvented: Towards Hands-Free Input via sEMG
- Partial multivariate transformer as a tool for cryptocurrencies time series prediction
- MEDIC: a network for monitoring data quality in collider experiments
- Compact neural networks for astronomy with optimal transport bias correction
- Together, Then Apart: Revisiting Multimodal Survival Analysis via a Min-Max Perspective
- Noise-Adaptive Quantum Circuit Mapping for Multi-Chip NISQ Systems via Deep Reinforcement Learning
- Diffusion-based Surrogate Model for Time-varying Underwater Acoustic Channels
- X-ReID: Multi-granularity Information Interaction for Video-Based Visible-Infrared Person Re-Identification
- Multi-speaker Attention Alignment for Multimodal Social Interaction
- DUPLE: An Intelligent Cross-Deployment Recognition Framework for Fiber-Optic Perimeter Security under Scarce Target Labels
- MobileVLA-R1: Reinforcing Vision-Language-Action for Mobile Robots
- MINDiff: Mask-Integrated Negative Attention for Controlling Overfitting in Text-to-Image Personalization
- Equivalence of Context and Parameter Updates in Modern Transformer Blocks
- Towards a future space-based, highly scalable AI infrastructure system design
- HyM-UNet: Synergizing Local Texture and Global Context via Hybrid CNN-Mamba Architecture for Medical Image Segmentation
- Controllability Analysis of State Space-based Language Model
- Generative Model Predictive Control in Manufacturing Processes: A Review
- ENGRAM: Effective, Lightweight Memory Orchestration for Conversational Agents
- SD-PSFNet: Sequential and Dynamic Point Spread Function Network for Image Deraining
- Signal: Selective Interaction and Global-local Alignment for Multi-Modal Object Re-Identification
- A Lightweight Approach to Detection of AI-Generated Texts Using Stylometric Features
- Blu-WERP (Web Extraction and Refinement Pipeline): A Scalable Pipeline for Preprocessing Large Language Model Datasets
- The Rapid Growth of AI Foundation Model Usage in Science
- Radar2Shape: 3D Shape Reconstruction from High-Frequency Radar using Multiresolution Signed Distance Functions
- Enhancing Quranic Learning: A Multimodal Deep Learning Approach for Arabic Phoneme Recognition
- Unmasking Airborne Threats: Guided-Transformers for Portable Aerosol Mass Spectrometry
- Multi-Agent Pointer Transformer: Seq-to-Seq Reinforcement Learning for Multi-Vehicle Dynamic Pickup-Delivery Problems
- MCMoE: Completing Missing Modalities with Mixture of Experts for Incomplete Multimodal Action Quality Assessment
- Selective Rotary Position Embedding
- R2PS: Worst-Case Robust Real-Time Pursuit Strategies under Partial Observability
- DeepCoT: Deep Continual Transformers for Real-Time Inference on Data Streams
- ReBaPL: Repulsive Bayesian Prompt Learning
- Mesh RAG: Retrieval Augmentation for Autoregressive Mesh Generation
- UI-Styler: Ultrasound Image Style Transfer with Class-Aware Prompts for Cross-Device Diagnosis Using a Frozen Black-Box Inference Network
- LangMark: A Multilingual Dataset for Automatic Post-Editing
- Difficulty-Controlled Simplification of Piano Scores with Synthetic Data for Inclusive Music Education
- Reconstruction of Surface EMG Signal using IMU data for Upper Limb Actions
- VLA-4D: Embedding 4D Awareness into Vision-Language-Action Models for SpatioTemporally Coherent Robotic Manipulation
- Steering in the Shadows: Causal Amplification for Activation Space Attacks in Large Language Models
- Attention-Guided Feature Fusion (AGFF) Model for Integrating Statistical and Semantic Features in News Text Classification
- A Cross-Cultural Assessment of Human Ability to Detect LLM-Generated Fake News about South Africa
- Diversity Has Always Been There in Your Visual Autoregressive Models
- Anatomy of an Idiom: Tracing Non-Compositionality in Language Models
- OmniPT: Unleashing the Potential of Large Vision Language Models for Pedestrian Tracking and Understanding
- RacketVision: A Multiple Racket Sports Benchmark for Unified Ball and Racket Analysis
- Identifying Quantum Structure in AI Language: Evidence for Evolutionary Convergence of Human and Artificial Cognition
- Energy Scaling Laws for Diffusion Models: Quantifying Compute and Carbon Emissions in Image Generation
- Feature Partitioning and Semantic Equalization for Intrinsic Robustness in Semantic Communication under Packet Loss
- Real-Time Cooked Food Image Synthesis and Visual Cooking Progress Monitoring on Edge Devices
- Fine-grained MoE Load Balancing with Linear Programming
- Aggregating Direct and Indirect Neighbors through Graph Linear Transformations
- A Little More Like This: Text-to-Image Retrieval with Vision-Language Models Using Relevance Feedback
- DiffRefiner: Coarse to Fine Trajectory Planning via Diffusion Refinement with Semantic Interaction for End to End Autonomous Driving
- UAM: A Unified Attention-Mamba Backbone of Multimodal Framework for Tumor Cell Classification
- MorphSeek: Fine-grained Latent Representation-Level Policy Optimization for Deformable Image Registration
- MolSight: Optical Chemical Structure Recognition with SMILES Pretraining, Multi-Granularity Learning and Reinforcement Learning
- Loomis Painter: Reconstructing the Painting Process
- AngioDG: Interpretable Channel-informed Feature-modulated Single-source Domain Generalization for Coronary Vessel Segmentation in X-ray Angiography
- Innovation by Displacement
- Predicting one-year clinical instability and mortality in heart failure patients using sequence modeling
- ManifoldFormer: Geometric Deep Learning for Neural Dynamics on Riemannian Manifolds
- CLAWDIA: A dictionary learning framework for gravitational-wave data analysis
- VersaPants: A Loose-Fitting Textile Capacitive Sensing System for Lower-Body Motion Capture
- Evolution Strategies at the Hyperscale
- TFCDiff: Robust ECG Denoising via Time-Frequency Complementary Diffusion
- TimeViper: A Hybrid Mamba-Transformer Vision-Language Model for Efficient Long Video Understanding
- gfnx: Fast and Scalable Library for Generative Flow Networks in JAX
- Boosting Predictive Performance on Tabular Data through Data Augmentation with Latent-Space Flow-Based Diffusion
- Flow and Depth Assisted Video Prediction with Latent Transformer
- A Comparison Between Decision Transformers and Traditional Offline Reinforcement Learning Algorithms
- Beyond Visual Cues: Leveraging General Semantics as Support for Few-Shot Segmentation
- From generative AI to the brain: five takeaways
- CAMS: Towards Compositional Zero-Shot Learning via Gated Cross-Attention and Multi-Space Disentanglement
- Unsupervised Graph Neural Network Framework for Balanced Multipatterning in Advanced Electronic Design Automation Layouts
- VTinker: Guided Flow Upsampling and Texture Mapping for High-Resolution Video Frame Interpolation
- Global Cross-Time Attention Fusion for Enhanced Solar Flare Prediction from Multivariate Time Series
- InEKFormer: A Hybrid State Estimator for Humanoid Robots
- Memory-DD: A Low-Complexity Dendrite-Inspired Neuron for Temporal Prediction Tasks
- FT-NCFM: An Influence-Aware Data Distillation Framework for Efficient VLA Models
- UT-OSANet: A Multimodal Deep Learning model for Evaluating and Classifying Obstructive Sleep Apnea
- Simba: Towards High-Fidelity and Geometrically-Consistent Point Cloud Completion via Transformation Diffusion
- Enhancing Nuclear Reactor Core Simulation through Data-Based Surrogate Models
- On 10x Better Scalability: KV Stores Scale Up KV Cache
- CoSP: Reconfigurable Multi-State Metamaterial Inverse Design via Contrastive Pretrained Large Language Model
- SUNAC: Source-aware Unified Neural Audio Codec
- BlockCert: Certified Blockwise Extraction of Transformer Mechanisms
- LLMs-based Augmentation for Domain Adaptation in Long-tailed Food Datasets
- Synergizing Deconfounding and Temporal Generalization For Time-series Counterfactual Outcome Estimation
- CARE: Turning LLMs Into Causal Reasoning Expert
- A Spatial Semantics and Continuity Perception Attention for Remote Sensing Water Body Change Detection
- MamTiff-CAD: Multi-Scale Latent Diffusion with Mamba+ for Complex Parametric Sequence
- Beyond Generative AI: World Models for Clinical Prediction, Counterfactuals, and Planning
- Joint Semantic-Channel Coding and Modulation for Token Communications
- CoordAR: One-Reference 6D Pose Estimation of Novel Objects via Autoregressive Coordinate Map Generation
- CASL: Curvature-Augmented Self-supervised Learning for 3D Anomaly Detection
- TB or Not TB: Coverage-Driven Direct Preference Optimization for Verilog Stimulus Generation
- Pharmacophore-based design by learning on voxel grids
- What Does It Take to Be a Good AI Research Agent? Studying the Role of Ideation Diversity
- Real-Time Optimal Control via Transformer Networks and Bernstein Polynomials
- A Decade of Systems for Human Data Interaction
- B+ANN: A Fast Billion-Scale Disk-based Nearest-Neighbor Index
- Standardising the NLP Workflow: A Framework for Reproducible Linguistic Analysis
- A Review of Machine Learning for Cavitation Intensity Recognition in Complex Industrial Systems
- Computation for Epidemic Prediction with Graph Neural Network by Model Combination
- A time for monsters: Organizational knowing after LLMs
- Small Language Models for Phishing Website Detection: Cost, Performance, and Privacy Trade-Offs
- Building Robust and Scalable Multilingual ASR for Indian Languages
- RRT*former: Environment-Aware Sampling-Based Motion Planning using Transformer
- Communication-Pipelined Split Federated Learning for Foundation Model Fine-Tuning in UAV Networks
- EVA-Net: Interpretable Anomaly Detection for Brain Health via Learning Continuous Aging Prototypes from One-Class EEG Cohorts
- TTF: A Trapezoidal Temporal Fusion Framework for LTV Forecasting in Douyin
- Taming Generative Synthetic Data for X-ray Prohibited Item Detection
- STREAM-VAE: Dual-Path Routing for Slow and Fast Dynamics in Vehicle Telemetry Anomaly Detection
- Graph Query Networks for Object Detection with Automotive Radar
- Learning Where, What and How to Transfer: A Multi-Role Reinforcement Learning Approach for Evolutionary Multitasking
- GloTok: Global Perspective Tokenizer for Image Reconstruction and Generation
- Efficient Score Pre-computation for Diffusion Models via Cross-Matrix Krylov Projection
- Unbiased Semantic Decoding with Vision Foundation Models for Few-shot Segmentation
- A Comprehensive Study on Visual Token Redundancy for Discrete Diffusion-based Multimodal Large Language Models
- MambaTrack3D: A State Space Model Framework for LiDAR-Based Object Tracking under High Temporal Variation
- Rectifying Distribution Shift in Cascaded Precipitation Nowcasting
- WiCo-MG: Wireless Channel Foundation Model for Multipath Generation via Synesthesia of Machines
- Context Cascade Compression: Exploring the Upper Limits of Text Compression
- HEAD-QA v2: Expanding a Healthcare Benchmark for Reasoning
- MoDES: Accelerating Mixture-of-Experts Multimodal Large Language Models via Dynamic Expert Skipping
- Transparent Early ICU Mortality Prediction with Clinical Transformer and Per-Case Modality Attribution
- Transformer-Guided Deep Reinforcement Learning for Optimal Takeoff Trajectory Design of an eVTOL Drone
- GeoSceneGraph: Geometric Scene Graph Diffusion Model for Text-guided 3D Indoor Scene Synthesis
- B-Rep Distance Functions (BR-DF): How to Represent a B-Rep Model by Volumetric Distance Functions?
- Graph Neural Networks for Vehicular Social Networks: Trends, Challenges, and Opportunities
- FreeSwim: Revisiting Sliding-Window Attention Mechanisms for Training-Free Ultra-High-Resolution Video Generation
- Attention via Synaptic Plasticity is All You Need: A Biologically Inspired Spiking Neuromorphic Transformer
- M-CALLM: Multi-level Context Aware LLM Framework for Group Interaction Prediction
- Adapformer: Adaptive Channel Management for Multivariate Time Series Forecasting
- IMSE: Efficient U-Net-based Speech Enhancement using Inception Depthwise Convolution and Amplitude-Aware Linear Attention
- Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech
- Dynamic Nested Hierarchies: Pioneering Self-Evolution in Machine Learning Architectures for Lifelong Intelligence
- DeepBlip: Estimating Conditional Average Treatment Effects Over Time
- LiteCache: A Query Similarity-Driven, GPU-Centric KVCache Subsystem for Efficient LLM Inference
- Parameter Aware Mamba Model for Multi-task Dense Prediction
- Enhancing End-to-End Autonomous Driving with Risk Semantic Distillaion from VLM
- Towards Stable and Structured Time Series Generation with Perturbation-Aware Flow Matching
- CompEvent: Complex-valued Event-RGB Fusion for Low-light Video Enhancement and Deblurring
- From Topology to Behavioral Semantics: Enhancing BGP Security by Understanding BGP's Language with LLMs
- Hyperion: Hierarchical Scheduling for Parallel LLM Acceleration in Multi-tier Networks
- Self-Supervised Multisensory Pretraining for Contact-Rich Robot Reinforcement Learning
- Continuous Vision-Language-Action Co-Learning with Semantic-Physical Alignment for Behavioral Cloning
- IBGS: Image-Based Gaussian Splatting
- Unified Multimodal Vessel Trajectory Prediction with Explainable Navigation Intention
- Object-Centric World Models for Causality-Aware Reinforcement Learning
- ArbESC+: Arabic Enhanced Edit Selection System Combination for Grammatical Error Correction Resolving conflict and improving system combination in Arabic GEC
- StreamingTalker: Audio-driven 3D Facial Animation with Autoregressive Diffusion Model
- Segment Anything Across Shots: A Method and Benchmark
- SymLoc: Symbolic Localization of Hallucination across HaluEval and TruthfulQA
- iGaussian: Real-Time Camera Pose Estimation via Feed-Forward 3D Gaussian Splatting Inversion
- Unifying Convolution and Attention via Convolutional Nearest Neighbors
- Compute-in-Memory Implementation of State Space Models for Event Sequence Processing
- From Projection to Prediction: Beyond Logits for Scalable Language Models
- CD-DPE: Dual-Prompt Expert Network Based on Convolutional Dictionary Feature Decoupling for Multi-Contrast MRI Super-Resolution
- FlakyGuard: Automatically Fixing Flaky Tests at Industry Scale
- WebRec: Enhancing LLM-based Recommendations with Attention-guided RAG from Web
- Can You Learn to See Without Images? Procedural Warm-Up for Vision Transformers
- ActVAR: Activating Mixtures of Weights and Tokens for Efficient Visual Autoregressive Generation
- VLMs Guided Interpretable Decision Making for Autonomous Driving
- MergeDNA: Context-aware Genome Modeling with Dynamic Tokenization through Token Merging
- Adaptive Multi-Scale Integration Unlocks Robust Cell Annotation in Histopathology Images
- ForgeDAN: An Evolutionary Framework for Jailbreaking Aligned Large Language Models
- Robust Defense Strategies for Multimodal Contrastive Learning: Efficient Fine-tuning Against Backdoor Attacks
- GREAT: Generalizable Representation Enhancement via Auxiliary Transformations for Zero-Shot Environmental Prediction
- Uni-Hand: Universal Hand Motion Forecasting in Egocentric Views
- PAST: A Primary-Auxiliary Spatio-Temporal Network for Traffic Time Series Imputation
- TripleFDS: Triple Feature Disentanglement and Synthesis for Scene Text Editing
- Generalized Denoising Diffusion Codebook Models (gDDCM): Tokenizing images using a pre-trained diffusion model
- Tab-PET: Graph-Based Positional Encodings for Tabular Transformers
- Whistledown: Combining User-Level Privacy with Conversational Coherence in LLMs
- PASE: Leveraging the Phonological Prior of WavLM for Low-Hallucination Generative Speech Enhancement
- View-aware Cross-modal Distillation for Multi-view Action Recognition
- Uncovering and Mitigating Transient Blindness in Multimodal Model Editing
- End-to-End Multi-Person Pose Estimation with Pose-Aware Video Transformer
- RefineVAD: Semantic-Guided Feature Recalibration for Weakly Supervised Video Anomaly Detection
- Dissecting and Re-architecting 3D NAND Flash PIM Arrays for Efficient Single-Batch Token Generation in LLMs
- GenTract: Generative Global Tractography
- Translation Entropy: A Statistical Framework for Evaluating Translation Systems
- HDW-SR: High-Frequency Guided Diffusion Model based on Wavelet Decomposition for Image Super-Resolution
- Inverse Electromagnetic Scattering for Doubly-Connected Cylinders using Convolutional Neural Networks
- Automated Road Distress Detection Using Vision Transformersand Generative Adversarial Networks
- Soft Conflict-Resolution Decision Transformer for Offline Multi-Task Reinforcement Learning
- Region-Point Joint Representation for Effective Trajectory Similarity Learning
- F.A.C.U.L.: Language-Based Interaction with AI Companions in Gaming
- NuBench: An Open Benchmark for Deep Learning-Based Event Reconstruction in Neutrino Telescopes
- Transformer-Based Multi-Agent Reinforcement Learning for Networked Systems with Long-Range Interactions
- MACKO: Sparse Matrix-Vector Multiplication for Low Sparsity
- DiffPixelFormer: Differential Pixel-Aware Transformer for RGB-D Indoor Scene Segmentation
- MGCA-Net: Multi-Grained Category-Aware Network for Open-Vocabulary Temporal Action Localization
- The Future of Food: How Artificial Intelligence is Transforming Food Manufacturing
- Functional Mean Flow in Hilbert Space
- TacEleven: generative tactic discovery for football open play
- BIRD: Bronze Inscription Restoration and Dating
- Federated Cyber Defense: Privacy-Preserving Ransomware Detection Across Distributed Systems
- Gated Fusion Enhanced Multi-Scale Hierarchical Graph Convolutional Network for Stock Movement Prediction
- Learning what to say and how precisely: Efficient Communication via Differentiable Discrete Communication Learning
- SAGA: Source Attribution of Generative AI Videos
- On the Cross-lingual Transferability of Pre-trained wav2vec2-based Models
- Improving the Generalisation of Learned Reconstruction Frameworks
- LAYA: Layer-wise Attention Aggregation for Interpretable Depth-Aware Neural Networks
- FSDAM: Few-Shot Driving Attention Modeling via Vision-Language Coupling
- Counting Through Occlusion: Framework for Open World Amodal Counting
- EchoLSTM: A Self-Reflective Recurrent Network for Stabilizing Long-Range Memory
- Terrain-Enhanced Resolution-aware Refinement Attention for Off-Road Segmentation
- UniSOT: A Unified Framework for Multi-Modality Single Object Tracking
- COFAP: A Universal Framework for COFs Adsorption Prediction through Designed Multi-Modal Extraction and Cross-Modal Synergy
- Floor Plan-Guided Visual Navigation Incorporating Depth and Directional Cues
- Neural process model predictive control
- A knowledge-driven approach for automated fire safety compliance checking in operational buildings
- OmniSparse: Training-Aware Fine-Grained Sparse Attention for Long-Video MLLMs
- X-VMamba: Explainable Vision Mamba
- Attention-Enhanced Convolutional Autoencoder and Structured Delay Embeddings for Weather Prediction
- CSV-Decode: Certifiable Sub-Vocabulary Decoding for Efficient Large Language Model Inference
- Fine-Grained Representation for Lane Topology Reasoning
- LMM-IR: Large-Scale Netlist-Aware Multimodal Framework for Static IR-Drop Prediction
- Can Small GenAI Language Models Rival Large Language Models in Understanding Application Behavior?
- Collaborative Charging Optimization for Wireless Rechargeable Sensor Networks via Heterogeneous Mobile Chargers
- One Request, Multiple Experts: LLM Orchestrates Domain Specific Models via Adaptive Task Routing
- BdSL-SPOTER: A Transformer-Based Framework for Bengali Sign Language Recognition with Cultural Adaptation
- Global-Lens Transformers: Adaptive Token Mixing for Dynamic Link Prediction
- From Black Box to Bijection: Interpreting Machine Learning to Build a Zeta Map Algorithm
- VLA-R: Vision-Language Action Retrieval toward Open-World End-to-End Autonomous Driving
- MSLoRA: Multi-Scale Low-Rank Adaptation via Attention Reweighting
- On the Dimension-Free Approximation of Deep Neural Networks for Symmetric Korobov Functions
- Which Way from B to A: The role of embedding geometry in image interpolation for Stable Diffusion
- Adaptive Focus Memory for Language Models
- GenSIaC: Toward Security-Aware Infrastructure-as-Code Generation with Large Language Models
- MTMed3D: A Multi-Task Transformer-Based Model for 3D Medical Imaging
- Hi-DREAM: Brain Inspired Hierarchical Diffusion for fMRI Reconstruction via ROI Encoder and visuAl Mapping
- CLAReSNet: When Convolution Meets Latent Attention for Hyperspectral Image Classification
- Learning Time in Static Classifiers
- MMSense: Adapting Vision-based Foundation Model for Multi-task Multi-modal Wireless Sensing
- Evaluating Latent Generative Paradigms for High-Fidelity 3D Shape Completion from a Single Depth Image
- Mixture of States: Routing Token-Level Dynamics for Multimodal Generation
- Scaling Law Analysis in Federated Learning: How to Select the Optimal Model Size?
- Data-Efficient Self-Supervised Algorithms for Fine-Grained Birdsong Analysis
- Open Banking Foundational Model: Learning Language Representations from Few Financial Transactions
- A Digital SRAM-Based Compute-In-Memory Macro for Weight-Stationary Dynamic Matrix Multiplication in Transformer Attention Score Computation
- Efficient Mathematical Reasoning Models via Dynamic Pruning and Knowledge Distillation
- Teaching Prompts to Coordinate: Hierarchical Layer-Grouped Prompt Tuning for Continual Learning
- A Robust and Explainable Transformer-Based Framework for Phishing Email Detection
- Intelligent Collaborative Optimization for Rubber Tyre Film Production Based on Multi-path Differentiated Clipping Proximal Policy Optimization
- DeiTFake: Deepfake Detection Model using DeiT Multi-Stage Training
- DCMM-Transformer: Degree-Corrected Mixed-Membership Attention for Medical Imaging
- BudgetLeak: Membership Inference Attacks on RAG Systems via the Generation Budget Side Channel
- Improved Masked Image Generation with Knowledge-Augmented Token Representations
- STAMP: Spatial-Temporal Adapter with Multi-Head Pooling
- Evaluation of Attention Mechanisms in U-Net Architectures for Semantic Segmentation of Brazilian Rock Art Petroglyphs
- Learning the relative composition of EEG signals using pairwise relative shift pretraining
- KVSwap: Disk-aware KV Cache Offloading for Long-Context On-device Inference
- Arcee: Differentiable Recurrent State Chain for Generative Vision Modeling with Mamba SSMs
- MP-GFormer: A 3D-Geometry-Aware Dynamic Graph Transformer Approach for Machining Process Planning
- SOTFormer: A Minimal Transformer for Unified Object Tracking and Trajectory Prediction
- On the Notion that Language Models Reason
- PAS : Prelim Attention Score for Detecting Object Hallucinations in Large Vision--Language Models
- CURENet: Combining Unified Representations for Efficient Chronic Disease Prediction
- TMDC: A Two-Stage Modality Denoising and Complementation Framework for Multimodal Sentiment Analysis with Missing and Noisy Modalities
- Large Language Models and 3D Vision for Intelligent Robotic Perception and Autonomy
- Reverberation: Learning the Latencies Before Forecasting Trajectories
- Enhancing Graph Representations with Neighborhood-Contextualized Message-Passing
- ChemFixer: Correcting Invalid Molecules to Unlock Previously Unseen Chemical Space
- VitalBench: A Rigorous Multi-Center Benchmark for Long-Term Vital Sign Prediction in Intraoperative Care
- LEMUR: Large scale End-to-end MUltimodal Recommendation
- Text-guided Weakly Supervised Framework for Dynamic Facial Expression Recognition
- Spatial Reasoning in Multimodal Large Language Models: A Survey of Tasks, Benchmarks and Methods
- Transformers vs. Recurrent Models for Estimating Forest Gross Primary Production
- Seeing the Forest and the Trees: Query-Aware Tokenizer for Long-Video Multimodal Language Models
- Optimizing Mixture of Block Attention
- Viper-F1: Fast and Fine-Grained Multimodal Understanding with Cross-Modal State-Space Modulation
- GFT: Graph Feature Tuning for Efficient Point Cloud Analysis
- Faster Algorithms for Structured Matrix Multiplication via Flip Graph Search
- Quantifying vacuum-like jets in heavy-ion collisions: a Machine Learning study
- Know Your Limits: Entropy Estimation Modeling for Compression and Generalization
- Belief Net: A Filter-Based Framework for Learning Hidden Markov Models from Observations
- Moirai 2.0: When Less Is More for Time Series Forecasting
- Graph Neural Networks, Deep Reinforcement Learning and Probabilistic Topic Modeling for Strategic Multiagent Settings
- STAGE: A Symbolic Tensor grAph GEnerator for distributed AI system co-design
- H3Former: Hypergraph-based Semantic-Aware Aggregation via Hyperbolic Hierarchical Contrastive Loss for Fine-Grained Visual Classification
- VISTA: A Vision and Intent-Aware Social Attention Framework for Multi-Agent Trajectory Prediction
- Beyond MSE: Ordinal Cross-Entropy for Probabilistic Time Series Forecasting
- Data-driven multi-species heat flux closures for two-stream-unstable plasmas with nonlinear sparse regression
- GPR: Towards a Generative Pre-trained One-Model Paradigm for Large-Scale Advertising Recommendation
- Do Not Merge My Model! Safeguarding Open-Source LLMs Against Unauthorized Model Merging
- FabasedVC: Enhancing Voice Conversion with Text Modality Fusion and Phoneme-Level SSL Features
- LoG3D: Ultra-High-Resolution 3D Shape Modeling via Local-to-Global Partitioning
- 2.5D Transformer: An Efficient 3D Seismic Interpolation Method without Full 3D Training
- AdaptViG: Adaptive Vision GNN with Exponential Decay Gating
- Simulator and Experience Enhanced Diffusion Model for Comprehensive ECG Generation
- Regional Attention-Enhanced Swin Transformer for Clinically Relevant Medical Image Captioning
- EnchTable: Unified Safety Alignment Transfer in Fine-tuned Large Language Models
- IPCD: Intrinsic Point-Cloud Decomposition
- ConSurv: Multimodal Continual Learning for Survival Analysis
- Steering Pretrained Drafters during Speculative Decoding
- SMoFi: Step-wise Momentum Fusion for Split Federated Learning on Heterogeneous Data
- A Methodology for Developing Foundational Transformer Models in Collider Physics Analysis
- A3: Attention-Aware Accurate KV Cache Fusion for Fast Large Language Model Serving
- Retrieval as a Decision: Training-Free Adaptive Gating for Efficient RAG
- Readability Measures and Automatic Text Simplification: In the Search of a Construct
- An ultrafast plenoptic-camera system for high-resolution 3D particle tracking in unsegmented scintillators
- Context-Aware Multimodal Representation Learning for Spatio-Temporally Explicit Environmental Modelling
- Transformer Semantic Genetic Programming for d-dimensional Symbolic Regression Problems
- Making Every Head Count: Sparse Attention Without the Speed-Performance Trade-off
- EPSegFZ: Efficient Point Cloud Semantic Segmentation for Few- and Zero-Shot Scenarios with Language Guidance
- End-to-end Contrastive Language-Speech Pretraining Model For Long-form Spoken Question Answering
- Deep Learning for Metabolic Rate Estimation from Biosignals: A Comparative Study of Architectures and Signal Selection
- Video Echoed in Music: Semantic, Temporal, and Rhythmic Alignment for Video-to-Music Generation
- The Learning Dynamics of Subword Segmentation for Morphologically Diverse Languages
- Iterated Population Based Training with Task-Agnostic Restarts
- Instrumental variables system identification with Lp consistency
- PressTrack-HMR: Pressure-Based Top-Down Multi-Person Global Human Mesh Recovery
- RGMP: Recurrent Geometric-prior Multimodal Policy for Generalizable Humanoid Robot Manipulation
- π-Attention: Periodic Sparse Transformers for Efficient Long-Context Modeling
- Automatic Minds: Cognitive Parallels Between Hypnotic States and Large Language Model Processing
- SpiralThinker: Latent Reasoning through an Iterative Process with Text-Latent Interleaving
- Spatio-Temporal Data Enhanced Vision-Language Model for Traffic Scene Understanding
- Selective Sinkhorn Routing for Improved Sparse Mixture of Experts
- Blurred Encoding for Trajectory Representation Learning
- Free-Boundary Quasiconformal Maps via a Least-squares Operator in Diffeomorphism Optimization
- TransactionGPT
- CMI-MTL: Cross-Mamba interaction based multi-task learning for medical visual question answering
- Weaver: Kronecker Product Approximations of Spatiotemporal Attention for Traffic Network Forecasting
- Hierarchical Memorization in Large Language Models: Evidence from Citation Generation
- Decomposition of Small Transformer Models
- A centroid based framework for text classification in itsm environments
- NeuroLingua: A Language-Inspired Hierarchical Framework for Multimodal Sleep Stage Classification Using EEG and EOG
- Synergistic Feature Fusion for Latent Lyrical Classification: A Gated Deep Learning Architecture
- SSMRadNet : A Sample-wise State-Space Framework for Efficient and Ultra-Light Radar Segmentation and Object Detection
- Hey Pentti, We Did (More of) It!: A Vector-Symbolic Lisp With Residue Arithmetic
- Optimal control of the future via prospective learning with control
- How Modality Shapes Perception and Reasoning: A Study of Error Propagation in ARC-AGI
- MACIE: Multi-Agent Causal Intelligence Explainer for Collective Behavior Understanding
- SENCA-st: Integrating Spatial Transcriptomics and Histopathology with Cross Attention Shared Encoder for Region Identification in Cancer Pathology
- Machine Learning-Guided Memory Optimization for DLRM Inference on Tiered Memory
- Vision Transformer Based User Equipment Positioning
- Why Should the Server Do It All?: A Scalable, Versatile, and Model-Agnostic Framework for Server-Light DNN Inference over Massively Distributed Clients via Training-Free Intermediate Feature Compression
- RDTE-UNet: A Boundary and Detail Aware UNet for Precise Medical Image Segmentation
- MVSMamba: Multi-View Stereo with State Space Model
- Leveraging unlabelled data for generalizable neural population decoding
- Moebius: 0.2B Lightweight Image Inpainting Framework with 10B-Level Performance
- Understanding semantic impairments in schizophrenia from a predictive coding perspective
- TurkEmbed: Turkish Embedding Model on NLI & STS Tasks
- One Model for All: Universal Pre-training for EEG based Emotion Recognition across Heterogeneous Datasets and Paradigms
- NeuCLIP: Efficient Large-Scale CLIP Training with Neural Normalizer Optimization
- EMAformer: Enhancing Transformer through Embedding Armor for Time Series Forecasting
- Hybrid Quantum-Classical Selective State Space Artificial Intelligence
- H-Model: Dynamic Neural Architectures for Adaptive Processing
- Test-time Diverse Reasoning by Riemannian Activation Steering
- HiLoMix: Robust High- and Low-Frequency Graph Learning Framework for Mixing Address Association
- Navigating the Wild: Pareto-Optimal Visual Decision-Making in Image Space
- Data-Driven Discovery of Feature Groups in Clinical Time Series
- A Unified Geometric Field Theory Framework for Transformers: From Manifold Embeddings to Kernel Modulation
- Fill the gaps: continuous in time interpolation of fluid dynamical simulations
- Evaluating Gemini LLM in Food Image-Based Recipe and Nutrition Description with EfficientNet-B4 Visual Backbone
- Encoder Fine-tuning with Stochastic Sampling Outperforms Open-weight GPT in Astronomy Knowledge Extraction
- Multi-Granularity Mutual Refinement Network for Zero-Shot Learning
- Filtering Jump Markov Systems with Partially Known Dynamics: A Model-Based Deep Learning Approach
- Gate-level boolean evolutionary geometric attention neural networks
- On the Interplay between Positional Encodings, Morphological Complexity, and Word Order Flexibility
- UniFormer: Unified and Efficient Transformer for Reasoning Across General and Custom Computing
- OTSNet: A Neurocognitive-Inspired Observation-Thinking-Spelling Pipeline for Scene Text Recognition
- A Small Leak Sinks All: Exploring the Transferable Vulnerability of Source Code Models
- Quantizing Whisper-small: How design choices affect ASR performance
- Information Capacity: Evaluating the Efficiency of Large Language Models via Text Compression
- Generalized-Scale Object Counting with Gradual Query Aggregation
- From Sequential to Recursive: Enhancing Decision-Focused Learning with Bidirectional Feedback
- Generalizable Insights for Graph Transformers in Theory and Practice
- High-Quality Proposal Encoding and Cascade Denoising for Imaginary Supervised Object Detection
- State of the Art in Text Classification for South Slavic Languages: Fine-Tuning or Prompting?
- Versatile and Risk-Sensitive Cardiac Diagnosis via Graph-Based ECG Signal Representation
- Unified Work Embeddings: Contrastive Learning of a Bidirectional Multi-task Ranker
- Multi-objective Hyperparameter Optimization in the Age of Deep Learning
- The Impact of Longitudinal Mammogram Alignment on Breast Cancer Risk Assessment
- Computational Blueprints: Generating Isomorphic Mathematics Problems with Large Language Models
- CellARC: Measuring Intelligence with Cellular Automata
- Data Descriptions from Large Language Models with Influence Estimation
- SpikCommander: A High-performance Spiking Transformer with Multi-view Learning for Efficient Speech Command Recognition
- Parallel Sampling via Autospeculation
- MonoCLUE : Object-Aware Clustering Enhances Monocular 3D Object Detection
- A General Method for Proving Networks Universal Approximation Property
- CloudMamba: Grouped Selective State Spaces for Point Cloud Analysis
- LAD-BNet: Lag-Aware Dual-Branch Networks for Real-Time Energy Forecasting on Edge Devices
- Streaming Tensor Program: A streaming abstraction for dynamic parallelism
- DWFF-Net : A Multi-Scale Farmland System Habitat Identification Method with Adaptive Dynamic Weight
- National Institute on Aging PREPARE Challenge: Early Detection of Cognitive Impairment Using Speech -- The SpeechCARE Solution
- A Circular Argument : Does RoPE need to be Equivariant for Vision?
- Filtered-ViT: A Robust Defense Against Multiple Adversarial Patch Attacks
- Auto-US: An Ultrasound Video Diagnosis Agent Using Video Classification Framework and LLMs
- oboro: Text-to-Image Synthesis on Limited Data using Flow-based Diffusion Transformer with MMH Attention
- Generative Artificial Intelligence in Qualitative Research Methods: Between Hype and Risks?
- HQ-SVC: Towards High-Quality Zero-Shot Singing Voice Conversion in Low-Resource Scenarios
- Rethinking Generative Image Pretraining: How Far Are We From Scaling Up Next-Pixel Prediction?
- Quantum-centric machine learning for molecular dynamics
- Misaligned by Design: Incentive Failures in Machine Learning
- Adaptive Graph Learning with Transformer for Multi-Reservoir Inflow Prediction
- TurkEmbed4Retrieval: Turkish Embedding Model for Retrieval Task
- Beyond Fact Retrieval: Episodic Memory for RAG with Generative Semantic Workspaces
- Forecasting implied volatility surface with generative diffusion models
- On the Creativity of AI Agents
- S-MUSt3R: Sliding Multi-view 3D Reconstruction
- Optimal Attention Temperature Enhances In-Context Learning under Distribution Shift
- Adaptation of Foundation Models for Medical Image Analysis: Strategies, Challenges, and Future Directions
- Conditional Diffusion as Latent Constraints for Controllable Symbolic Music Generation
- Adversarial Spatio-Temporal Attention Networks for Epileptic Seizure Forecasting
- Extending QAOA-GPT to Higher-Order Quantum Optimization Problems
- TNT: Improving Chunkwise Training for Test-Time Memorization
- The Value of Personalized Recommendations: Evidence from Netflix
- StreamKV: Streaming Video Question-Answering with Segment-based KV Cache Retrieval and Compression
- Revisiting the Neural Tangent Kernel: the role of large width and depth
- Omni-AVSR: Towards Unified Multimodal Speech Recognition with Large Language Models
- 4DSTR: Advancing Generative 4D Gaussians with Spatial-Temporal Rectification for High-Quality and Consistent 4D Generation
- Neural Directional Filtering Using a Compact Microphone Array
- Guiding Generative Models to Uncover Diverse and Novel Crystals via Reinforcement Learning
- Hi-WaveTST: A Hybrid High-Frequency Wavelet-Transformer for Time-Series Classification
- ProcGen3D: Learning Neural Procedural Graph Representations for Image-to-3D Reconstruction
- Two Heads are Better than One: Distilling Large Language Model Features Into Small Models with Feature Decomposition and Mixture
- Green AI: A systematic review and meta-analysis of its definitions, lifecycle models, hardware and measurement attempts
- LLM Driven Processes to Foster Explainable AI
- LeCoT: revisiting network architecture for two-view correspondence pruning
- KAT-GNN: A Knowledge-Augmented Temporal Graph Neural Network for Risk Prediction in Electronic Health Records
- Evaluating Large Language Models for Anxiety, Depression, and Stress Detection: Insights into Prompting Strategies and Synthetic Data
- Wavelet Enhanced Adaptive Frequency Filter for Sequential Recommendation
- Diffolio: A Diffusion Model for Multivariate Probabilistic Financial Time-Series Forecasting and Portfolio Construction
- A Picture is Worth a Thousand (Correct) Captions: A Vision-Guided Judge-Corrector System for Multimodal Machine Translation
- Learning from the Right Patches: A Two-Stage Wavelet-Driven Masked Autoencoder for Histopathology Representation Learning
- Learning to Focus: Prioritizing Informative Histories with Structured Attention Mechanisms in Partially Observable Reinforcement Learning
- COGNOS: Universal Enhancement for Time Series Anomaly Detection via Constrained Gaussian-Noise Optimization and Smoothing
- Beyond Observations: Reconstruction Error-Guided Irregularly Sampled Time Series Representation Learning
- Distillation Dynamics: Towards Understanding Feature-Based Distillation in Vision Transformers
- Learning to Focus: Focal Attention for Selective and Scalable Transformers
- Controllable Flow Matching for Online Reinforcement Learning
- Learning to Fast Unrank in Collaborative Filtering Recommendation
- Recursive Dynamics in Fast-Weights Homeostatic Reentry Networks: Toward Reflective Intelligence
- Resource Efficient Sleep Staging via Multi-Level Masking and Prompt Learning
- Learning Biomolecular Motion: The Physics-Informed Machine Learning Paradigm
- ASTER: Attention-based Spiking Transformer Engine for Event-driven Reasoning
- K-Stain: Keypoint-Driven Correspondence for H&E-to-IHC Virtual Staining
- Can LLM Annotations Replace User Clicks for Learning to Rank?
- Dual-branch Spatial-Temporal Self-supervised Representation for Enhanced Road Network Learning
- Spatial-Frequency Enhanced Mamba for Multi-Modal Image Fusion
- Attention and Compression is all you need for Controllably Efficient Language Models
- Rethinking Parameter Sharing as Graph Coloring for Structured Compression
- Superhuman AI for Stratego Using Self-Play Reinforcement Learning and Test-Time Search
- Sensitivity of Small Language Models to Fine-tuning Data Contamination
- DETECT: Data-Driven Evaluation of Treatments Enabled by Classification Transformers
- On the diameter of subgradient sequences in o-minimal structures
- Breaking the Stealth-Potency Trade-off in Clean-Image Backdoors with Generative Trigger Optimization
- DeepBooTS: Dual-Stream Residual Boosting for Drift-Resilient Time-Series Forecasting
- MobileLLM-Pro Technical Report
- You Had One Job: Per-Task Quantization Using LLMs' Hidden Representations
- Hybrid CNN-ViT Framework for Motion-Blurred Scene Text Restoration
- Route Experts by Sequence, not by Token
- Diagnose Like A REAL Pathologist: An Uncertainty-Focused Approach for Trustworthy Multi-Resolution Multiple Instance Learning
- Learning the Inverse Ryu--Takayanagi Formula with Transformers
- Vocabulary In-Context Learning in Transformers: Benefits of Positional Encoding
- Reaction Prediction via Interaction Modeling of Symmetric Difference Shingle Sets
- GazeVLM: A Vision-Language Model for Multi-Task Gaze Understanding
- EcoSpa: Efficient Transformer Training with Coupled Sparsity
- Improving Multimodal Sentiment Analysis via Modality Optimization and Dynamic Primary Modality Selection
- Seq2Seq Models Reconstruct Visual Jigsaw Puzzles without Seeing Them
- Learning Topology-Driven Multi-Subspace Fusion for Grassmannian Deep Network
- SAMora: Enhancing SAM through Hierarchical Self-Supervised Pre-Training for Medical Images
- Transolver is a Linear Transformer: Revisiting Physics-Attention through the Lens of Linear Attention
- COTN: A Chaotic Oscillatory Transformer Network for Complex Volatile Systems under Extreme Conditions
- LaneDiffusion: Improving Centerline Graph Learning via Prior Injected BEV Feature Generation
- Spatially-Aware Mixture of Experts with Log-Logistic Survival Modeling for Whole-Slide Images
- AdaDrive: Self-Adaptive Slow-Fast System for Language-Grounded Autonomous Driving
- LUT-LLM: Efficient Large Language Model Inference with Memory-based Computations on FPGAs
- Breaking the Modality Barrier: Generative Modeling for Accurate Molecule Retrieval from Mass Spectra
- Synheart Emotion: Privacy-Preserving On-Device Emotion Recognition from Biosignals
- Resilience Inference for Supply Chains with Hypergraph Neural Network
- Optimizing Long-context LLM Serving via Fine-grained Sequence Parallelism
- Exploiting Inter-Session Information with Frequency-enhanced Dual-Path Networks for Sequential Recommendation
- Interaction-Centric Knowledge Infusion and Transfer for Open-Vocabulary Scene Graph Generation
- Enhancing Multimodal Misinformation Detection by Replaying the Whole Story from Image Modality Perspective
- Scaling Laws and In-Context Learning: A Unified Theoretical Framework
- VLDrive: Vision-Augmented Lightweight MLLMs for Efficient Language-grounded Autonomous Driving
- Guardian-regularized Safe Offline Reinforcement Learning for Smart Weaning of Mechanical Circulatory Devices
- Forecasting Thermospheric Density with Transformers for Multi-Satellite Orbit Management
- Neodragon: Mobile Video Generation using Diffusion Transformer
- How Particle-System Random Batch Methods Enhance Graph Transformer: Memory Efficiency and Parallel Computing Strategy
- ITPP: Learning Disentangled Event Dynamics in Marked Temporal Point Processes
- Lethe: Layer- and Time-Adaptive KV Cache Pruning for Reasoning-Intensive LLM Serving
- MALeR: Improving Compositional Fidelity in Layout-Guided Generation
- DiA-gnostic VLVAE: Disentangled Alignment-Constrained Vision Language Variational AutoEncoder for Robust Radiology Reporting with Missing Modalities
- Adaptive Agent Selection and Interaction Network for Image-to-point cloud Registration
- Next-Latent Prediction Transformers Learn Compact World Models
- Retrieval-Augmented Generation in Medicine: A Scoping Review of Technical Implementations, Clinical Applications, and Ethical Considerations
- A Remarkably Efficient Paradigm to Multimodal Large Language Models for Sequential Recommendation
- Predicting the Future by Retrieving the Past
- EGG-SR: Embedding Symbolic Equivalence into Symbolic Regression via Equality Graph
- Securing UAV Communications by Fusing Cross-Layer Fingerprints
- MARAuder's Map: Motion-Aware Real-time Activity Recognition with Layout-Based Trajectories
- Sign language recognition from skeletal data using graph and recurrent neural networks
- Hilbert-Guided Sparse Local Attention
- Causal Tracing of Object Representations in Large Vision Language Models: Mechanistic Interpretability and Hallucination Mitigation
- Toward Better Generalization in Few-Shot Learning through the Meta-Component Combination
- Temporal convolutional and fusional transformer model with Bi-LSTM encoder-decoder for multi-time-window remaining useful life prediction
- AWEMixer: Adaptive Wavelet-Enhanced Mixer Network for Long-Term Time Series Forecasting
- DGTN: Graph-Enhanced Transformer with Diffusive Attention Gating Mechanism for Enzyme DDG Prediction
- A Metamorphic Testing Perspective on Knowledge Distillation for Language Models of Code: Does the Student Deeply Mimic the Teacher?
- Predicting Grain Growth in Polycrystalline Materials Using Deep Learning Time Series Models
- How Many Tokens Do 3D Point Cloud Transformer Architectures Really Need?
- Accurate online action and gesture recognition system using detectors and Deep SPD Siamese Networks
- Global Feature Enhancing and Fusion Framework for Strain Gauge Time Series Classification
- Reasoning on Time-Series for Financial Technical Analysis
- ForecastGAN: A Decomposition-Based Adversarial Framework for Multi-Horizon Time Series Forecasting
- Multi-period Learning for Financial Time Series Forecasting
- Rethinking Metrics and Diffusion Architecture for 3D Point Cloud Generation
- Code Review Automation using Retrieval Augmented Generation
- QueStER: Query Specification for Generative keyword-based Retrieval
- Embedding-Space Data Augmentation to Prevent Membership Inference Attacks in Clinical Time Series Forecasting
- Integrating Score-Based Diffusion Models with Machine Learning-Enhanced Localization for Advanced Data Assimilation in Geological Carbon Storage
- An End-to-End Deep Reinforcement Learning Approach for Solving the Traveling Salesman Problem with Drones
- ADPretrain: Advancing Industrial Anomaly Detection via Anomaly Representation Pretraining
- FreeControl: Efficient, Training-Free Structural Control via One-Step Attention Extraction
- Evaluating Spatio-Temporal Forecasting Trade-offs Between Graph Neural Networks and Foundation Models
- From Linear Probing to Joint-Weighted Token Hierarchy: A Foundation Model Bridging Global and Cellular Representations in Biomarker Detection
- Efficient representation of 3D spatial data for defense-related applications
- Iterative Layer-wise Distillation for Efficient Compression of Large Language Models
- Order-Level Attention Similarity Across Language Models: A Latent Commonality
- Medical Referring Image Segmentation via Next-Token Mask Prediction
- Implementation of transformer-based LLMs with large-scale optoelectronic neurons on a CMOS image sensor platform
- UHDRes: Ultra-High-Definition Image Restoration via Dual-Domain Decoupled Spectral Modulation
- Deep Progressive Training: scaling up depth capacity of zero/one-layer models
- Pattern-Aware Diffusion Synthesis of fMRI/dMRI with Tissue and Microstructural Refinement
- The Future of Fully Homomorphic Encryption System: from a Storage I/O Perspective
- Search Is Not Retrieval: Decoupling Semantic Matching from Contextual Assembly in RAG
- BudgetMem: Learning Selective Memory Policies for Cost-Efficient Long-Context Processing in Language Models
- Multi-Agent Collaborative Framework For Math Problem Generation
- Self-Supervised Implicit Attention Priors for Point Cloud Reconstruction
- 3D Gaussian Point Encoders
- Conditional Neural ODE for Longitudinal Parkinson's Disease Progression Forecasting
- When Data Falls Short: Grokking Below the Critical Threshold
- Nowcast3D: Reliable precipitation nowcasting via gray-box learning
- Eyes on Target: Gaze-Aware Object Detection in Egocentric Video
- wa-hls4ml: A Benchmark and Surrogate Models for hls4ml Resource and Latency Estimation
- Integrating Temporal and Structural Context in Graph Transformers for Relational Deep Learning
- Probabilistic Textual Time Series Depression Detection
- Temporal Action Selection for Action Chunking
- MacroNav: Multi-Task Context Representation Learning Enables Efficient Navigation in Unknown Environments
- Can Context Bridge the Reality Gap? Sim-to-Real Transfer of Context-Aware Policies
- On the Brittleness of CLIP Text Encoders
- The Strong Lottery Ticket Hypothesis for Multi-Head Attention Mechanisms
- AStF: Motion Style Transfer via Adaptive Statistics Fusor
- Forget BIT, It is All about TOKEN: Towards Semantic Information Theory for LLMs
- Decomposable Neuro Symbolic Regression
- When Swin Transformer Meets KANs: An Improved Transformer Architecture for Medical Image Segmentation
- Transfer Learning for Transformer-Based Modeling of Nonlinear Pulse Evolution in Er-Doped Fiber Amplifiers
- An LLM-based Framework for Human-Swarm Teaming Cognition in Disaster Search and Rescue
- Towards Scalable Meta-Learning of near-optimal Interpretable Models via Synthetic Model Generations
- TwIST: Rigging the Lottery in Transformers with Independent Subnetwork Training
- PETRA: Pretrained Evolutionary Transformer for SARS-CoV-2 Mutation Prediction
- MoSa: Motion Generation with Scalable Autoregressive Modeling
- ForeRobo: Unlocking Infinite Simulation Data for 3D Goal-driven Robotic Manipulation
- AI as We Describe It: How Large Language Models and Their Applications in Health are Represented Across Channels of Public Discourse
- GraphCliff: Short-Long Range Gating for Subtle Differences but Critical Changes
- Sketch-Augmented Features Improve Learning Long-Range Dependencies in Graph Neural Networks
- A Transferable Machine Learning Approach to Predict Quantum Circuit Parameters for Electronic Structure Problems
- Open Source State-Of-the-Art Solution for Romanian Speech Recognition
- Part-Aware Bottom-Up Group Reasoning for Fine-Grained Social Interaction Detection
- Signal Intensity-weighted coordinate channels improve learning stability and generalisation in 1D and 2D CNNs in localisation tasks on biomedical signals
- A systematic review of relation extraction task since the emergence of Transformers
- Imitation Learning in the Deep Learning Era: A Novel Taxonomy and Recent Advances
- Bearing Syntactic Fruit with Stack-Augmented Neural Networks
- From Prompts to Power: Measuring the Energy Footprint of LLM Inference
- NAPS: Attention-Based Fusion of Heterogeneous Physiological Signals
- THD-BAR: Topology Hierarchical Derived Brain Autoregressive Modeling for EEG Generic Representations
- SyMuPe: Affective and Controllable Symbolic Music Performance
- Segmentation Beyond Defaults: Asymmetrical Byte Pair Encoding for Optimal Machine Translation Performance
- GEMMA-SQL: A Novel Text-to-SQL Model Based on Large Language Models
- Enhancing composition-based materials property prediction by cross-modal knowledge transfer
- Decoupling Augmentation Bias in Prompt Learning for Vision-Language Models
- Unified Long Video Inpainting and Outpainting via Overlapping High-Order Co-Denoising
- Comparing the Performance of LLMs in RAG-based Question-Answering: A Case Study in Computer Science Literature
- Enhancing Medical Image Segmentation via Heat Conduction Equation
- MvBody: Multi-View-Based Hybrid Transformer Using Optical 3D Body Scan for Explainable Cesarean Section Prediction
- Cross-Modal Alignment via Variational Copula Modelling
- Efficient Linear Attention for Multivariate Time Series Modeling via Entropy Equality
- Periodic Skill Discovery
- EGMOF: Efficient Generation of Metal-Organic Frameworks Using a Hybrid Diffusion-Transformer Architecture
- SnapStream: Efficient Long Sequence Decoding on Dataflow Accelerators
- A Computer Vision Based Proxy for Political Polarization in Religious Countries: A Turkiye Case Study
- COMPASS: Context-Modulated PID Attention Steering System for Hallucination Mitigation
- Large Language Models as Information Sources: Distinctive Characteristics and Types of Low-Quality Information
- Beyond Citations: Measuring Idea-level Knowledge Diffusion from Research to Journalism and Policy-making
- Analyzing the Power of Chain of Thought through Memorization Capabilities
- Benchmark Datasets for Lead-Lag Forecasting on Social Platforms
- AILA--First Experiments with Localist Language Models
- Conditional Diffusion Model-Enabled Scenario-Specific Neural Receivers for Superimposed Pilot Schemes
- The Curved Spacetime of Transformer Architectures
- Fine-Tuning Vision-Language Models for Multimodal Polymer Property Prediction
- Discrete Bayesian Sample Inference for Graph Generation
- Scalable Single-Cell Gene Expression Generation with Latent Diffusion Models
- Automatic Machine Translation Detection Using a Surrogate Multilingual Translation Model
- Generative Hints
- Test-Time Steering for Lossless Text Compression via Weighted Product of Experts
- MVAFormer: RGB-based Multi-View Spatio-Temporal Action Recognition with Transformer
- Modeling Hawkish-Dovish Latent Beliefs in Multi-Agent Debate-Based LLMs for Monetary Policy Decision Classification
- Apriel-H1: Towards Efficient Enterprise Reasoning Models
- Federated Attention: A Distributed Paradigm for Collaborative LLM Inference over Edge Networks
- Condition Numbers and Eigenvalue Spectra of Shallow Networks on Spheres
- Resource-efficient Automatic Refinement of Segmentations via Weak Supervision from Light Feedback
- SEAL - A Symmetry EncourAging Loss for High Energy Physics
- Dynamic Priors in Bayesian Optimization for Hyperparameter Optimization
- An End-to-End Learning Approach for Solving Capacitated Location-Routing Problems
- Keeping it Local, Tiny and Real: Automated Report Generation on Edge Computing Devices for Mechatronic-Based Cognitive Systems
- Object Detection as an Optional Basis: A Graph Matching Network for Cross-View UAV Localization
- Improving DF-Conformer Using Hydra For High-Fidelity Generative Speech Enhancement on Discrete Codec Token
- Chronic Kidney Disease Prognosis Prediction Using Transformer
- Differentiable Hierarchical Visual Tokenization
- High-Resolution Magnetic Particle Imaging System Matrix Recovery Using a Vision Transformer with Residual Feature Network
- KGBridge: Knowledge-Guided Prompt Learning for Non-overlapping Cross-Domain Recommendation
- DoFlow: Causal Generative Flows for Interventional and Counterfactual Time-Series Prediction
- Using Span Queries to Optimize for Cache and Attention Locality
- Dexterous Robotic Piano Playing at Scale
- BRAINS: A Retrieval-Augmented System for Alzheimer's Detection and Monitoring
- Sample-Efficient Language Modeling with Linear Attention and Lightweight Enhancements
- A Roadmap for Predictive Human Immunology
- Optimizing Attention on GPUs by Exploiting GPU Architectural NUMA Effects
- DL4Proteins Jupyter Notebooks Teach how to use Artificial Intelligence for Biomolecular Structure Prediction and Design
- Vibe Learning: Education in the age of AI
- Flashlight: PyTorch Compiler Extensions to Accelerate Attention Variants
- Shared Parameter Subspaces and Cross-Task Linearity in Emergently Misaligned Behavior
- Machine and Deep Learning for Indoor UWB Jammer Localization
- No-rank Tensor Decomposition Using Metric Learning
- Fractional Diffusion Bridge Models
- Random Initialization of Gated Sparse Adapters
- LM-Fix: Lightweight Bit-Flip Detection and Rapid Recovery Framework for Language Models
- Lightweight Learning from Actuation-Space Demonstrations via Flow Matching for Whole-Body Soft Robotic Grasping
- Wonder3D++: Cross-domain Diffusion for High-fidelity 3D Generation from a Single Image
- HGFreNet: Hop-hybrid GraphFomer for 3D Human Pose Estimation with Trajectory Consistency in Frequency Domain
- SemBench: A Benchmark for Semantic Query Processing Engines
- The Ghost in the Keys: A Disklavier Demo for Human-AI Musical Co-Creativity
- Rescuing the Unpoisoned: Efficient Defense against Knowledge Corruption Attacks on RAG Systems
- Anatomically Constrained Transformers for Echocardiogram Analysis
- Hyper Hawkes Processes: Interpretable Models of Marked Temporal Point Processes
- Improving Romanian LLM Pretraining Data using Diversity and Quality Filtering
- GeoToken: Hierarchical Geolocalization of Images via Next Token Prediction
- Knowledge Elicitation with Large Language Models for Interpretable Cancer Stage Identification from Pathology Reports
- On the Emergence of Induction Heads for In-Context Learning
- Seed-Induced Uniqueness in Transformer Models: Subspace Alignment Governs Subliminal Transfer
- SARIMAX-Based Power Outage Prediction During Extreme Weather Events
- None To Optima in Few Shots: Bayesian Optimization with MDP Priors
- Transformer-Based Decoding in Concatenated Coding Schemes Under Synchronization Errors
- GauDP: Reinventing Multi-Agent Collaboration through Gaussian-Image Synergy in Diffusion Policies
- MID: A Self-supervised Multimodal Iterative Denoising Framework
- Hydra: Dual Exponentiated Memory for Multivariate Time Series Analysis
- VesSAM: Efficient Multi-Prompting for Segmenting Complex Vessel
- Keys in the Weights: Transformer Authentication Using Model-Bound Latent Representations
- Motion-Robust Multimodal Fusion of PPG and Accelerometer Signals for Three-Class Heart Rhythm Classification
- Dynamic Population Distribution Aware Human Trajectory Generation with Diffusion Model
- Transformers as Intrinsic Optimizers: Forward Inference through the Energy Principle
- FlexiCache: Leveraging Temporal Stability of Attention Heads for Efficient KV Cache Management
- Neural Green's Functions
- Parameter Interpolation Adversarial Training for Robust Image Classification
- Reconstruction of Black Hole Ringdown Signals with Data Gaps using a Deep-Learning Framework
- Linear Differential Vision Transformer: Learning Visual Contrasts via Pairwise Differentials
- OMEGA: Optimized Multimodal Position Encoding Index Derivation with Global Adaptive Scaling for Vision-Language Models
- GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding
- Effective Series Decomposition and Components Learning for Time Series Generation
- Equilibrium Policy Generalization: A Reinforcement Learning Framework for Cross-Graph Zero-Shot Generalization in Pursuit-Evasion Games
- Occlusion-Aware Diffusion Model for Pedestrian Intention Prediction
- Validating Deep Models for Alzheimer's 18F-FDG PET Diagnosis Across Populations: A Study with Latin American Data
- Structurally Refined Graph Transformer for Multimodal Recommendation
- TRACES: Temporal Recall with Contextual Embeddings for Real-Time Video Anomaly Detection
- FTT-GRU: A Hybrid Fast Temporal Transformer with GRU for Remaining Useful Life Prediction
- Improving Robustness to Out-of-Distribution States in Imitation Learning via Deep Koopman-Boosted Diffusion Policy
- Temporal Fusion Transformer for Multi-Horizon Probabilistic Forecasting of Weekly Retail Sales
- MIFO: Learning and Synthesizing Multi-Instance from One Image
- Air Pollution Forecasting in Bucharest
- ID-Crafter: VLM-Grounded Online RL for Compositional Multi-Subject Video Generation
- Leveraging the Cross-Domain & Cross-Linguistic Corpus for Low Resource NMT: A Case Study On Bhili-Hindi-English Parallel Corpus
- ToxicTextCLIP: Text-Based Poisoning and Backdoor Attacks on CLIP Pre-training
- Superpositional Gradient Descent: Harnessing Quantum Principles for Model Training
- TRISKELION-1: Unified Descriptive-Predictive-Generative AI
- Region-Aware Reconstruction Strategy for Pre-training fMRI Foundation Model
- MambaNetLK: Enhancing Colonoscopy Point Cloud Registration with Mamba
- BeetleFlow: An Integrative Deep Learning Pipeline for Beetle Image Processing
- Machine learning-based cloud resource allocation algorithms: a comprehensive comparative review
- Generative Modeling Enables Molecular Structure Retrieval from Coulomb Explosion Imaging
- Challenges in Credit Assignment for Multi-Agent Reinforcement Learning in Open Agent Systems
- NegoCollab: A Common Representation Negotiation Approach for Heterogeneous Collaborative Perception
- Validity Is What You Need
- TetraJet-v2: Accurate NVFP4 Training for Large Language Models with Oscillation Suppression and Outlier Control
- BiSparse-AAS: Bilinear Sparse Attention and Adaptive Spans Framework for Scalable and Efficient Text Summarization
- FedAdamW: A Communication-Efficient Optimizer with Convergence and Generalization Guarantees for Federated Large Models
- From the Rock Floor to the Cloud: A Systematic Survey of State-of-the-Art NLP in Battery Life Cycle
- A Sensing Whole Brain Zebrafish Foundation Model for Neuron Dynamics and Behavior
- ThoughtProbe: Classifier-Guided LLM Thought Space Exploration via Probing Representations
- MedM2T: A MultiModal Framework for Time-Aware Modeling with Electronic Health Record and Electrocardiogram Data
- Parameterized Prompt for Incremental Object Detection
- Higher-order Linear Attention
- Languages are Modalities: Cross-Lingual Alignment via Encoder Injection
- Consciousness-ECG Transformer for Conscious State Estimation System with Real-Time Monitoring
- Sparse Model Inversion: Efficient Inversion of Vision Transformers for Data-Free Applications
- Probability-Biased Attention over Directed Bipartite Graphs for Long-Tail ICD Coding
- M3Detection: Multi-Frame Multi-Level Feature Fusion for Multi-Modal 3D Object Detection with Camera and 4D Imaging Radar
- AFM-Net: Advanced Fusing Hierarchical CNN Visual Priors with Global Sequence Modeling for Remote Sensing Image Scene Classification
- Exploring Landscapes for Better Minima along Valleys
- Improving Cross-view Object Geo-localization: A Dual Attention Approach with Cross-view Interaction and Multi-Scale Spatial Features
- Functional embeddings enable Aggregation of multi-area SEEG recordings over subjects and sessions
- Language Modeling With Factorization Memory
- Integrating ConvNeXt and Vision Transformers for Enhancing Facial Age Estimation
- Spiking Neural Networks: The Future of Brain-Inspired Computing
- Addressing Longstanding Challenges in Cognitive Science with Language Models
- ParaScopes: What do Language Models Activations Encode About Future Text?
- Mitigating Semantic Collapse in Partially Relevant Video Retrieval
- Balancing Knowledge Updates: Toward Unified Modular Editing in LLMs
- Spatial-SSRL: Enhancing Spatial Understanding via Self-Supervised Reinforcement Learning
- POSESTITCH-SLT: Linguistically Inspired Pose-Stitching for End-to-End Sign Language Translation
- GeneFlow: Translation of Single-cell Gene Expression to Histopathological Images via Rectified Flow
- Calibration Across Layers: Understanding Calibration Evolution in LLMs
- H-FA: A Hybrid Floating-Point and Logarithmic Approach to Hardware Accelerated FlashAttention
- Scalable Processing-Near-Memory for 1M-Token LLM Inference: CXL-Enabled KV-Cache Management Beyond GPU Limits
- Mixture-of-Transformers Learn Faster: A Theoretical Study on Classification Problems
- MoME: Mixture of Visual Language Medical Experts for Medical Imaging Segmentation
- Semantic Frame Aggregation-based Transformer for Live Video Comment Generation
- Understanding and Enhancing Mamba-Transformer Hybrids for Memory Recall and Language Modeling
- Integrating Ontologies with Large Language Models for Enhanced Control Systems in Chemical Engineering
- Scaling Image Geo-Localization to Continent Level
- LLMs Process Lists With General Filter Heads
- Determination of the initial condition for the Balitsky-Kovchegov equation with transformers
- Cross-Platform Evaluation of Reasoning Capabilities in Foundation Models
- MSAD: A Deep Dive into Model Selection for Time series Anomaly Detection
- Encoder-Decoder or Decoder-Only? Revisiting Encoder-Decoder Large Language Model
- Spiking Patches: Asynchronous, Sparse, and Efficient Tokens for Event Cameras
- SA2Net: Scale-Adaptive Structure-Affinity Transformation for Spine Segmentation from Ultrasound Volume Projection Imaging
- The Structure of Relation Decoding Linear Operators in Large Language Models
- Hebrew Diacritics Restoration using Visual Representation
- CyberNER: A Harmonized STIX Corpus for Cybersecurity Named Entity Recognition
- Context Engineering 2.0: The Context of Context Engineering
- Scaffolding Creativity: How Divergent and Convergent LLM Personas Shape Human Machine Creative Problem-Solving
- SecureReviewer: Enhancing Large Language Models for Secure Code Review through Secure-aware Fine-tuning
- Barlow Twins for Sequential Recommendation
- SPG-CDENet: Spatial Prior-Guided Cross Dual Encoder Network for Multi-Organ Segmentation
- Adaptive Context Length Optimization with Low-Frequency Truncation for Multi-Agent Reinforcement Learning
- UniTok-Audio: A Unified Audio Generation Framework via Generative Modeling on Discrete Codec Tokens
- Towards Explainable and Reliable AI in Finance
- Generative Artificial Intelligence for Air Shower Simulation
- Modeling strategies for speech enhancement in the latent space of a neural audio codec
- Towards Realistic Earth-Observation Constellation Scheduling: Benchmark and Methodology
- Pragmatic Theories Enhance Understanding of Implied Meanings in LLMs
- Questionnaire meets LLM: A Benchmark and Empirical Study of Structural Skills for Understanding Questions and Responses
- Hybrid Quantum-Classical Recurrent Neural Networks
- Similarity-Distance-Magnitude Language Models
- MossNet: Mixture of State-Space Experts is a Multi-Head Attention
- MoTDiff: High-resolution Motion Trajectory estimation from a single blurred image using Diffusion models
- MV-MLM: Bridging Multi-View Mammography and Language for Breast Cancer Diagnosis and Risk Prediction
- StructLayoutFormer:Conditional Structured Layout Generation via Structure Serialization and Disentanglement
- Denoising Refinement Diffusion Models for Simultaneous Generation of Multi-scale Mobile Network Traffic
- Security Risk of Misalignment between Text and Image in Multi-modal Model
- Nirvana: A Specialized Generalist Model With Task-Aware Memory Mechanism
- Group-Equivariant Diffusion Models for Lattice Field Theory
- Towards Scaling Laws for Symbolic Regression
- Electric Vehicle Charging Load Modeling: A Survey, Trends, Challenges and Opportunities
- Predicate Renaming via Large Language Models
- Loquetier: A Virtualized Multi-LoRA Framework for Unified LLM Fine-tuning and Serving
- Cloning Deterministic Worlds: The Critical Role of Latent Geometry in Long-Horizon World Models
- AttnCache: Accelerating Self-Attention Inference for LLM Prefill via Attention Cache
- SoK: Honeypots & LLMs, More Than the Sum of Their Parts?
- Gaperon: A Peppered English-French Generative Language Model Suite
- DiagramEval: Evaluating LLM-Generated Diagrams via Graphs
- How Data Mixing Shapes In-Context Learning: Asymptotic Equivalence for Transformers with MLPs
- ScaleDiff: Higher-Resolution Image Synthesis via Efficient and Model-Agnostic Diffusion
- Interpreting LLMs as Credit Risk Classifiers: Do Their Feature Explanations Align with Classical ML?
- PureKV: Plug-and-Play KV Cache Optimization with Spatial-Temporal Sparse Attention for Vision-Language Large Models
- Leveraging an Atmospheric Foundational Model for Subregional Sea Surface Temperature Forecasting
- Transformers Provably Learn Directed Acyclic Graphs via Kernel-Guided Mutual Information
- Off-policy Reinforcement Learning with Model-based Exploration Augmentation
- TempoPFN: Synthetic Pre-training of Linear RNNs for Zero-shot Time Series Forecasting
- Hallucinations in Bibliographic Recommendation: Citation Frequency as a Proxy for Training Data Redundancy
- MMEdge: Accelerating On-device Multimodal Inference via Pipelined Sensing and Encoding
- Beyond Epistemia: Epistemic Schizologia and Large Language Models as Techno-Semiotic Machines
- Systematic Analysis of Large Language Models and Transformer-Based Machine Translation for English-Tamil and Tamil-English Across Diverse Datasets
- F(AI)2R: Who Did What, and Who Checked? Verifiable AI Provenance as an Executable Skill
- Scaling universal Fermi network toward ground states: A diffusion-Monte-Carlo assessment
- WALoMA: A Multitask Wireless Foundation Model via Adaptive Low-Rank Masked Autoencoders
- SignDeepSC: A Semantic Signature-based Approach for Robust Semantic Communication
- Transformer Atomic Cluster Expansion: TRACE
- Faster, Higher, Stronger? The Impact of GenAI on Knowledge Work Productivity - Evidence from the Field
- Audio-Native Speech Recognition with a Frozen Discrete-Diffusion Language Model
- Auger Spectroscopy via Generative Quantum Eigensolver: A Quantum Approach to Molecular Excitations
- ClockRoPE: Random Fourier Rotations for Temporal Routine Modeling
- The Art of Not Forgetting A Local Learning Architecture for Continual Learning
- LLMET: Enabling Cross-Layer Evaluation of Emerging M3D Memories for Energy-Efficient LLM Serving
- PSG: Pair-Space Generation for Efficient Generative Reranking
- Phoneme- vs. Character-Level Targets and Selective State-Space Models for Intracortical Brain-to-Text
- CASIAL: Geometric Distortion Robust Image Watermarking
- An Attention-Based Framework for Alzheimers Disease Classification Using Resting-State fMRI
- AlphaSchema: Exploring the Space of Trading Semantics for LLM-Based Alpha Mining
- Unfolded Recursive Expectation-Maximization Neural Network For Speaker Tracking
- Steering Instruction Hierarchies at Inference Time
- Cognitive Convergence: Deep Similarities Between Large Language Models and Human Cognition
- Dynamic Parameterization Is Not Dynamic Inference
- Lag-aware cross-hand alignment for dual-hand action segmentation
- Data Fusion and Contrastive Alignment for Unconstrained IR Molecular Structure Elucidation
- DVPSFormer: Efficient Online Depth-aware Video Panoptic Segmentation for Autonomous Driving
- Weight and Height Estimation from a Single Human Image Captured in the Wild
- Shape-Based Inductive Bias for Glioma Grading from Tumor Contours
- Guess Where You Go: Generative Next Point-of-Interest Recommendation in Amap
- GuidedRAG: Semantic Steering of Retrieval-Augmented Generation
- Towards Flexible, Natural, Efficient Interaction for Conversational Talking Face Generation
- RedKnot: Efficient Long-Context LLM Serving with Head-Aware KV Reuse and SegPagedAttention
- An Empirical Audit of Input Encoders for Multi-Channel Signal Transformers
- Keyless Attention: Value-Space Routing and Value-Only Caching for Efficient Transformers
- Mixture-of-Depths Attention
- Distance-aware Soft Prompt Learning for Multimodal Valence-Arousal Estimation
- Lightweight Transformer for EEG Classification via Balanced Signed Graph Algorithm Unrolling
- TS-Reasoner: Aligning Time Series Foundation Models with LLM Reasoning
- ExplainRec: Towards Explainable Multi-Modal Zero-Shot Recommendation with Preference Attribution and Large Language Models
- Product-Quantised Image Representation for High-Quality Image Synthesis
- Paris: A Decentralized Trained Open-Weight Diffusion Model
- A Unified Deep Reinforcement Learning Approach for Close Enough Traveling Salesman Problem
- Tools Are Not Islands: Set-Level Tool Retrieval for LLM Agents via Query-Conditioned Hyperedge Prediction
- MRCoder: An Efficient Context Selecting Approach for Repository-Level Code Generation
- From Uncertainty to Determinism: Coarse-to-Fine Visual Floorplan Localization without Ray Matching
- Reactive Transformer (RxT) -- Stateful Real-Time Processing for Event-Driven Reactive Language Models
- Don't Just Chase "Highlighted Tokens" in MLLMs: Revisiting Visual Holistic Context Retention
- TridentServe: A Stage-level Serving System for Diffusion Pipelines
- ZeroShotOpt: Towards Zero-Shot Pretrained Models for Efficient Black-Box Optimization
- The Curious Case of In-Training Compression of State Space Models
- Mitigating Modal Imbalance in Multimodal Reasoning
- AURA: Adaptive Unified Reasoning and Automation with LLM-Guided MARL for NextG Cellular Networks
- Truth-Aware Decoding: A Program-Logic Approach to Factual Language Generation
- Multi-scale Autoregressive Models are Laplacian, Discrete, and Latent Diffusion Models in Disguise
- What Does It Take to Detect an AI Agent? Minimal Feature Sets for Behavioral Detection under Browser Automation
- Ideas Have Genomes: Benchmarking Scientific Lineage Reasoning and Lineage-Grounded Idea Generation
- Bi Directional Feedback Fusion for Activity Aware Forecasting of Indoor CO2 and PM2.5
- LLM and Human Modes of Representation
- Generalization from Low- to Moderate-Resolution Spectra with Neural Networks for Stellar Parameter Estimation: A Case Study with DESI
- NOVA3R: Non-pixel-aligned Visual Transformer for Amodal 3D Reconstruction
- TweetyBERT: Automated parsing of birdsong through self-supervised machine learning
- MedFusionT5: Cross-Modal Attention Boosts Semantic Quality and Reduces Hallucinations in Dental AI
- Towards Verifiable Transformers: Solver-Checkable Circuit Explanations
- AlphaCrafter: Harnessing Multi-Agent Workflows for Cross-Sectional Quantitative Trading
- The Scaling Properties of Implicit Deductive Reasoning in Transformers
- Convolutional neural networks in Vis–NIR chemometrics: From contradiction to conditional design
- Enhancing Multi-Agent Communication through Attention Steering with Context Relevance
- Dissecting the Black Box: Circuit-Level Analysis of LLM Vulnerability Detection
- Symbol-Equivariant Recurrent Reasoning Models
- MANDO-LLM: Heterogeneous Graph Transformers with Large Language Models for Smart Contract Vulnerability Detection
- ThriftAttention: Selective Mixed Precision for Long-Context FP4 Attention
- HRM-Text: Efficient Pretraining Beyond Scaling
- SSV: Sparse Speculative Verification for Efficient LLM Inference
- Understanding Wacky Weights: A Dissection of SPLADE's Learned Term Importance
- Quantifying Concentration Phenomena of Mean-Field Transformers in the Low-Temperature Regime
- VGGT-Det: Mining VGGT Internal Priors for Sensor-Geometry-Free Multi-View Indoor 3D Object Detection
- OpenMAP-BrainAge: generalizable and interpretable brain age predictor from MRI
- A Theory of Generalization in Deep Learning
- CommFuse: Hiding Tail Latency via Communication Decomposition and Fusion for Distributed LLM Training
- How Does Machine Learning Manage Complexity?
- CatPath‐GPT: A Mixture of Experts System for Computational Catalyst Design
- Exploring the Intersection of AI, Language, and Law: A Bibliometric Analysis
- Low-latency Event-based Object Detection with Spatially-Sparse Linear Attention
- Calibri: Enhancing Diffusion Transformers via Parameter-Efficient Calibration
- Copy-Augmented Representation for Structure Invariant Template-Free Retrosynthesis
- ViT-Transformer: Self-attention mechanism based constitutive modeling for nonlinear heterogeneous materials
- Arcee Trinity Large Technical Report
- Protein Language Model Embeddings Improve Generalization of Implicit Transfer Operators
- Speech-XL: Towards Long-Form Speech Understanding in Large Speech Language Models
- Understanding and Optimizing Attention-Based Sparse Matching for Diverse Local Features
- Towards Real-World Industrial-Scale Verification: LLM-Driven Theorem Proving on seL4
- NeurIPT: Foundation Model for Neural Interfaces
- On the Use of Large Language Models for Qualitative Synthesis
- Architecture, Simulation and Software Stack to Support Post-CMOS Accelerators: The ARCHYTAS Project
- A robust and stable hybrid neural network/finite element method for 2D flows that generalizes to different geometries
- An Empirical Study of Foundation Models for Variability-Induced Compilation Errors in Configurable C Code
- Better Call Grep: Evaluating and Improving Grep-Like Lexical Retrieval for Repository-Level Code Completion
- EDVD-LLaMA: Explainable Deepfake Video Detection via Multimodal Large Language Model Reasoning
- D2 Actor Critic: Diffusion Actor Meets Distributional Critic
- State of the Art of LLM-Enabled Interaction with Visualization
- ProbFM: Probabilistic Time Series Foundation Model with Uncertainty Decomposition
- Ministral 3
- DeepTaxa: a hybrid CNN-BERT framework for 16S rRNA taxonomic classification
- Generalizable and scalable protein stability prediction with rewired protein generative models
- Stay Tuned: Improving Sentiment Analysis and Stance Detection Using Large Language Models
- Artificial intelligence in bioinformatics: a survey
- Sparse Transformer Architectures via Regularized Wasserstein Proximal Operator with L1 Prior
- GaTector+: A Unified Head-free Framework for Gaze Object and Gaze Following Prediction
- Revisiting scalable sequential recommendation with Multi-Embedding Approach and Mixture-of-Experts
- IBNorm: Information-Bottleneck Inspired Normalization for Representation Learning
- TV-Rec: Time-Variant Convolutional Filter for Sequential Recommendation
- Sharpness-Guided Group Relative Policy Optimization via Probability Shaping
- MGTS-Net: Exploring Graph-Enhanced Multimodal Fusion for Augmented Time Series Forecasting
- GReF: A Unified Generative Framework for Efficient Reranking via Ordered Multi-token Prediction
- mLR: Scalable Laminography Reconstruction based on Memoization
- Optimizing Knowledge Utilization for Multi-Intent Comment Generation with Large Language Models
- A Study on Inference Latency for Vision Transformers on Mobile Devices
- Target-Guided Bayesian Flow Networks for Quantitatively Constrained CAD Generation
- Large Language Model for Verilog Code Generation: Literature Review and the Road Ahead
- EnzyControl: Adding Functional and Substrate-Specific Control for Enzyme Backbone Generation
- Bridging the Divide: End-to-End Sequence-Graph Learning
- GAPMAP: Mapping Scientific Knowledge Gaps in Biomedical Literature Using Large Language Models
- Breast Cancer VLMs: Clinically Practical Vision-Language Train-Inference Models
- PitchFlower: A flow-based neural audio codec with pitch controllability
- MaGNet: A Mamba Dual-Hypergraph Network for Stock Prediction via Temporal-Causal and Global Relational Learning
- Resource Allocation in Hybrid Radio-Optical IoT Networks using GNN with Multi-task Learning
- The use of LLMs to annotate data in management research: Foundational guidelines and warnings
- Future of AI Models: A Computational perspective on Model collapse
- Emergence of Minimal Circuits for Indirect Object Identification in Attention-Only Transformers
- What Really Matters in Matrix-Whitening Optimizers?
- Enhancing Hierarchical Reinforcement Learning through Change Point Detection in Time Series
- FaRAccel: FPGA-Accelerated Defense Architecture for Efficient Bit-Flip Attack Resilience in Transformer Models
- Exponential Dynamic Energy Network for High Capacity Sequence Memory
- SCOUT: A Lightweight Framework for Scenario Coverage Assessment in Autonomous Driving
- Decoupled MeanFlow: Turning Flow Models into Flow Maps for Accelerated Sampling
- Non-Singularity of the Gradient Descent map for Neural Networks with Piecewise Analytic Activations
- Modality-Aware SAM: Sharpness-Aware-Minimization Driven Gradient Modulation for Harmonized Multimodal Learning
- Understanding Multi-View Transformers
- Proper Body Landmark Subset Enables More Accurate and 5X Faster Recognition of Isolated Signs in LIBRAS
- FruitProm: Probabilistic Maturity Estimation and Detection of Fruits and Vegetables
- SARC: Sentiment-Augmented Deep Role Clustering for Fake News Detection
- Does Object Binding Naturally Emerge in Large Pretrained Vision Transformers?
- Dissecting Role Cognition in Medical LLMs via Neuronal Ablation
- Multi-Agent Scenario Generation in Roundabouts with a Transformer-enhanced Conditional Variational Autoencoder
- Particle-level transformers for 95 GeV Higgs boson searches at future e+e- Higgs factories
- Group Relative Attention Guidance for Image Editing
- Long-Context Modeling with Dynamic Hierarchical Sparse Attention for On-Device LLMs
- Leveraging Scale Separation and Stochastic Closure for Data-Driven Prediction of Chaotic Dynamics
- What Limits Agentic Systems Efficiency?
- Time-Embedded Algorithm Unrolling for Computational MRI
- Generative AI for Healthcare: Fundamentals, Challenges, and Perspectives
- MIMIC-Sepsis: A Curated Benchmark for Modeling and Learning from Sepsis Trajectories in the ICU
- Mitigating Hallucination in Large Language Models (LLMs): An Application-Oriented Survey on RAG, Reasoning, and Agentic Systems
- On the Provable Importance of Gradients for Language-Assisted Image Clustering
- Cerberus: Real-Time Video Anomaly Detection via Cascaded Vision-Language Models
- HiMAE: Hierarchical Masked Autoencoders Discover Resolution-Specific Structure in Wearable Time Series
- MiniOneRec: An Open-Source Framework for Scaling Generative Recommendation
- CodeWiki: Evaluating AI's Ability to Generate Holistic Documentation for Large-Scale Codebases
- Uncovering Gaps Between RFC Updates and TCP/IP Implementations: LLM-Facilitated Differential Checks on Intermediate Representations
- Metadata-Driven Retrieval-Augmented Generation for Financial Question Answering
- Synergizing chemical and AI communities for advancing laboratories of the future
- Text Simplification with Sentence Embeddings
- A Unified Geometric Space Bridging AI Models and the Human Brain
- Few-Shot Remote Sensing Image Scene Classification with CLIP and Prompt Learning
- Transformers can do Bayesian Clustering
- MAGNET: A Multi-Graph Attentional Network for Code Clone Detection
- PRIVET: Privacy Metric Based on Extreme Value Theory
- ETC: training-free diffusion models acceleration with Error-aware Trend Consistency
- Causal Convolutional Neural Networks as Finite Impulse Response Filters
- Closing Gaps: An Imputation Analysis of ICU Vital Signs
- SCOPE: Saliency-Coverage Oriented Token Pruning for Efficient Multimodel LLMs
- Blindfolded Experts Generalize Better: Insights from Robotic Manipulation and Videogames
- Exploring the Influence of Relevant Knowledge for Natural Language Generation Interpretability
- DualCap: Enhancing Lightweight Image Captioning via Dual Retrieval with Similar Scenes Visual Prompts
- EddyFormer: Accelerated Neural Simulations of Three-Dimensional Turbulence at Scale
- UniPlanner: A Unified Motion Planning Framework for Autonomous Vehicle Decision-Making Systems via Multi-Dataset Integration
- Fixed Point Neural Acceleration and Inverse Surrogate Model for Battery Parameter Identification
- PFEA: An LLM-based High-Level Natural Language Planning and Feedback Embodied Agent for Human-Centered AI
- ZTRS: Zero-Imitation End-to-end Autonomous Driving with Trajectory Scoring
- Language-Conditioned Representations and Mixture-of-Experts Policy for Robust Multi-Task Robotic Manipulation
- Mitigating Negative Transfer via Reducing Environmental Disagreement
- MeCeFO: Enhancing LLM Training Robustness via Fault-Tolerant Optimization
- NeuroPathNet: Dynamic Path Trajectory Learning for Brain Functional Connectivity Analysis
- An efficient probabilistic hardware architecture for diffusion-like models
- Neural USD: An object-centric framework for iterative editing and control
- Key and Value Weights Are Probably All You Need: On the Necessity of the Query, Key, Value weight Triplet in Decoder-Only Transformers
- Learning-based Spectral Regression for Cocoa Bean Physicochemical Property Prediction
- Hybrid Modeling, Sim-to-Real Reinforcement Learning, and Large Language Model Driven Control for Digital Twins
- MoPHES:Leveraging on-device LLMs as Agent for Mobile Psychological Health Evaluation and Support
- Opportunities for artificial intelligence and synthetic biology in designing living drug delivery systems
- LightGlueStick: a Fast and Robust Glue for Joint Point-Line Matching
- Does GenAI Rewrite How We Write? An Empirical Study on Two-Million Preprints
- Edge-Based Speech Transcription and Synthesis for Kinyarwanda and Swahili Languages
- Fine-tuning of Large Language Models for Constituency Parsing Using a Sequence to Sequence Approach
- FourierCompress: Layer-Aware Spectral Activation Compression for Efficient and Accurate Collaborative LLM Inference
- Structured Temporal Causality for Interpretable Multivariate Time Series Anomaly Detection
- Can LLMs Narrate Tabular Data? An Evaluation Framework for Natural Language Representations of Text-to-SQL System Outputs
- A Neural Model for Contextual Biasing Score Learning and Filtering
- Image Categorization and Search via a GAT Autoencoder and Representative Models
- Sentiment and Volatility in Financial Markets: A Review of BERT and GARCH Applications during Geopolitical Crises
- Sequence Modeling with Spectral Mean Flows
- A Survey on Efficient Vision-Language-Action Models
- A U-Net and Transformer Pipeline for Multilingual Image Translation
- Parallel BiLSTM-Transformer networks for forecasting chaotic dynamics
- Protein Folding with Neural Ordinary Differential Equations
- Education Paradigm Shift To Maintain Human Competitive Advantage Over AI
- Exploring Vulnerability in AI Industry
- EMTSF:Extraordinary Mixture of SOTA Models for Time Series Forecasting
- One-Timestep is Enough: Achieving High-performance ANN-to-SNN Conversion via Scale-and-Fire Neurons
- Revisiting Multimodal Positional Encoding in Vision-Language Models
- Large language model-based task planning for service robots: A review
- SwiftEmbed: Ultra-Fast Text Embeddings via Static Token Lookup for Real-Time Applications
- Predicting symbolic ODEs from multiple trajectories
- PAHQ: Accelerating Automated Circuit Discovery through Mixed-Precision Inference Optimization
- Deep Active Inference with Diffusion Policy and Multiple Timescale World Model for Real-World Exploration and Navigation
- Provable test-time adaptivity and distributional robustness of in-context learning
- Autoregressive Styled Text Image Generation, but Make it Reliable
- VR-Drive: Viewpoint-Robust End-to-End Driving with Feed-Forward 3D Gaussian Splatting
- MATCH: Task-Driven Code Evaluation through Contrastive Learning
- Implicit Modeling for Transferability Estimation of Vision Foundation Models
- DQ3D: Depth-guided Query for Transformer-Based 3D Object Detection in Traffic Scenarios
- DeepSalt: Bridging Laboratory and Satellite Spectra through Domain Adaptation and Knowledge Distillation for Large-Scale Soil Salinity Estimation
- Neural Emulator Superiority: When Machine Learning for PDEs Surpasses its Training Data
- Awakening Facial Emotional Expressions in Human-Robot
- Knocking-Heads Attention
- SwiftTS: A Swift Selection Framework for Time Series Pre-trained Models via Multi-task Meta-Learning
- Nested AutoRegressive Models
- Softmax is 1/2-Lipschitz: A tight bound across all ℓp norms
- Can Language Models Compose Skills In-Context?
- Capsule Network-Based Multimodal Fusion for Mortgage Risk Assessment from Unstructured Data Sources
- QoSGMAA: A Robust Multi-Order Graph Attention and Adversarial Framework for Sparse QoS Prediction
- Hankel Singular Value Regularization for Highly Compressible State Space Models
- Intelligent Multimodal Multi-Sensor Fusion-Based UAV Identification, Localization, and Countermeasures for Safeguarding Low-Altitude Economy
- Positional Preservation Embedding for Multimodal Large Language Models
- ESCA: Enabling Seamless Codec Avatar Execution through Algorithm and Hardware Co-Optimization for Virtual Reality
- Diffuse to Detect: A Generalizable Framework for Anomaly Detection with Diffusion Models Applications to UAVs and Beyond
- Modeling Political Discourse with Sentence-BERT and BERTopic
- Exploring Structures of Inferential Mechanisms through Simplistic Digital Circuits
- Neural Architecture Search for global multi-step Forecasting of Energy Production Time Series
- Sub-microsecond Transformers for Jet Tagging on FPGAs
- A Review of End-to-End Precipitation Prediction Using Remote Sensing Data: from Divination to Machine Learning
- Encoder-Decoder Diffusion Language Models for Efficient Training and Inference
- Semantic Surgery: Zero-Shot Concept Erasure in Diffusion Models
- Exploration of Summarization by Generative Language Models for Automated Scoring of Long Essays
- MAGIC-Talk: Motion-aware Audio-Driven Talking Face Generation with Customizable Identity Control
- MedXplain-VQA: Multi-Component Explainable Medical Visual Question Answering
- FOXES: A Framework For Operational X-ray Emission Synthesis
- Agentic Meta-Orchestrator for Multi-task Copilots
- SeeDNorm: Self-Rescaled Dynamic Normalization
- Beyond Semantics: How Temporal Biases Shape Retrieval in Transformer and State-Space Models
- Scalable Neural Decoders for Practical Real-Time Quantum Error Correction
- Step2Motion: Locomotion Reconstruction from Pressure Sensing Insoles
- WaveMAE: Wavelet decomposition Masked Auto-Encoder for Remote Sensing
- SALSA: Single-pass Autoregressive LLM Structured Classification
- Random Search Neural Networks for Efficient and Expressive Graph Learning
- Transformers from Compressed Representations
- Alias-Free ViT: Fractional Shift Invariance via Linear Attention
- If You Want to Be Robust, Be Wary of Initialization
- DynaCausal: Dynamic Causality-Aware Root Cause Analysis for Distributed Microservices
- JiuTian Chuanliu: A Large Spatiotemporal Model for General-purpose Dynamic Urban Sensing
- Mitigating Attention Sinks and Massive Activations in Audio-Visual Speech Recognition with LLMs
- Publication Trend Analysis and Synthesis via Large Language Model: A Case Study of Engineering in PNAS
- An End-to-End Generative Diffusion Model for Heavy-Ion Collisions
- TLSQKT: A Question-Aware Dual-Channel Transformer for Literacy Tracing from Learning Sequences
- PromptReverb: Multimodal Room Impulse Response Generation Through Latent Rectified Flow Matching
- Extragradient Method for (L0, L1)-Lipschitz Root-finding Problems
- TERRA: A Transformer-Enabled Recursive R-learner for Longitudinal Heterogeneous Treatment Effect Estimation
- Johnson-Lindenstrauss Lemma Beyond Euclidean Geometry
- NetBurst: Event-Centric Forecasting of Bursty, Intermittent Time Series
- Top-Down Semantic Refinement for Image Captioning
- Memory-based Language Models: An Efficient, Explainable, and Eco-friendly Approach to Large Language Modeling
- Cross-Enhanced Multimodal Fusion of Eye-Tracking and Facial Features for Alzheimer's Disease Diagnosis
- PatenTEB: A Comprehensive Benchmark and Model Family for Patent Text Embedding
- Accident Anticipation via Temporal Occurrence Prediction
- SteerX: Disentangled Steering for LLM Personalization
- HPC-Driven Modeling with ML-Based Surrogates for Magnon-Photon Dynamics in Hybrid Quantum Systems
- Enpowering Your Pansharpening Models with Generalizability: Unified Distribution is All You Need
- TrajGATFormer: A Graph-Based Transformer Approach for Worker and Obstacle Trajectory Prediction in Off-site Construction Environments
- Multi-dataset Joint Pre-training of Emotional EEG Enables Generalizable Affective Computing
- M-CIF: Multi-Scale Alignment For CIF-Based Non-Autoregressive ASR
- Graph Identification of Proteins in Tomograms (GRIP-Tomo) 2.0: Topologically Aware Classification for Proteins
- Edit Less, Achieve More: Dynamic Sparse Neuron Masking for Lifelong Knowledge Editing in LLMs
- When UAV Swarm Meets IRS: Collaborative Secure Communications in Low-altitude Wireless Networks
- Gradual Forgetting: Logarithmic Compression for Extending Transformer Context Windows
- STAR-RIS-assisted Collaborative Beamforming for Low-altitude Wireless Networks
- Streaming Generation for Music Accompaniment
- Deep Gaussian Processes for Functional Maps
- Energy-Efficient Domain-Specific Artificial Intelligence Models and Agents: Pathways and Paradigms
- Massive Memorization with Hundreds of Trillions of Parameters for Sequential Transducer Generative Recommenders
- Caption-Driven Explainability: Probing CNNs for Bias via CLIP
- Normalization in Attention Dynamics
- From Black-box to Causal-box: Towards Building More Interpretable Models
- Performance Trade-offs of Optimizing Small Language Models for E-Commerce
- Parallel Sampling from Masked Diffusion Models via Conditional Independence Testing
- Transformer Based Linear Attention with Optimized GPU Kernel Implementation
- AutoSciDACT: Automated Scientific Discovery through Contrastive Embedding and Hypothesis Testing
- Spatially Aware Linear Transformer (SAL-T) for Particle Jet Tagging
- Visual Diffusion Models are Geometric Solvers
- BachVid: Training-Free Video Generation with Consistent Background and Character
- StylePitcher: Generating Style-Following and Expressive Pitch Curves for Versatile Singing Tasks
- LLM-Generated Negative News Headlines Dataset: Creation and Benchmarking Against Real Journalism
- DAP-MAE: Domain-Adaptive Point Cloud Masked Autoencoder for Effective Cross-Domain Learning
- Few-Shot Knowledge Distillation of LLMs With Counterfactual Explanations
- S3OD: Towards Generalizable Salient Object Detection with Synthetic Data
- Unveiling the Spatial-temporal Effective Receptive Fields of Spiking Neural Networks
- Integrating Genomics into Multimodal EHR Foundation Models
- MATrack: Efficient Multiscale Adaptive Tracker for Real-Time Nighttime UAV Operations
- REVE: A Foundation Model for EEG -- Adapting to Any Setup with Large-Scale Pretraining on 25,000 Subjects
- A Unified Model for Multi-Task Drone Routing in Post-Disaster Road Assessment
- Surrogate-based quantification of policy uncertainty in generative flow networks
- ITC-RWKV: Interactive Tissue-Cell Modeling with Recurrent Key-Value Aggregation for Histopathological Subtyping
- Unified token representations for sequential decision models
- Redefining Retrieval Evaluation in the Era of LLMs
- Flight Delay Prediction via Cross-Modality Adaptation of Large Language Models and Aircraft Trajectory Representation
- Compressing Quaternion Convolutional Neural Networks for Audio Classification
- Cost-Sensitive Freeze-thaw Bayesian Optimization for Efficient Hyperparameter Tuning
- Dynamic Semantic-Aware Correlation Modeling for UAV Tracking
- α-LoRA: Effective Fine-Tuning via Base Model Rescaling
- CausalRec: A CausalBoost Attention Model for Sequential Recommendation
- Chronos-2: From Univariate to Universal Forecasting
- Cavity Duplexer Tuning with 1d Resnet-like Neural Networks
- LLM-Powered Detection of Price Manipulation in DeFi
- Sparser Block-Sparse Attention via Token Permutation
- Relieving the Over-Aggregating Effect in Graph Transformers
- Randomized-MLP Regularization Improves Domain Adaptation and Interpretability in DINOv2
- Model Merging with Functional Dual Anchors
- Embedding Explainable AI in NHS Clinical Safety: The Explainability-Enabled Clinical Safety Framework (ECSF)
- Mitra: Mixed Synthetic Priors for Enhancing Tabular Foundation Models
- Quantum Neural Network Architectures for Multivariate Time-Series Forecasting
- LLMComp: A Language Modeling Paradigm for Error-Bounded Scientific Data Compression (Technical Report)
- Large Language Models Meet Text-Attributed Graphs: A Survey of Integration Frameworks and Applications
- Robust Distortion-Free Watermark for Autoregressive Audio Generation Models
- Bridging Language Gaps with Adaptive RAG: Improving Indonesian Language Question Answering
- FSRF: Factorization-guided Semantic Recovery for Incomplete Multimodal Sentiment Analysis
- xMem: A CPU-Based Approach for Accurate Estimation of GPU Memory in Deep Learning Training Workloads
- From Questions to Queries: An AI-powered Multi-Agent Framework for Spatial Text-to-SQL
- Attention Sinks in Diffusion Language Models
- Generative AI in Depth: A Survey of Recent Advances, Model Variants, and Real-World Applications
- Learning Grouped Lattice Vector Quantizers for Low-Bit LLM Compression
- Memory Constrained Dynamic Subnetwork Update for Transfer Learning
- Irish-BLiMP: A Linguistic Benchmark for Evaluating Human and Language Model Performance in a Low-Resource Setting
- Generative Point Tracking with Flow Matching
- Focal Modulation and Bidirectional Feature Fusion Network for Medical Image Segmentation
- Stateful KV Cache Management for LLMs: Balancing Space, Time, Accuracy, and Positional Fidelity
- Towards General Modality Translation with Contrastive and Predictive Latent Diffusion Bridge
- Video Prediction of Dynamic Physical Simulations With Pixel-Space Spatiotemporal Transformers
- ARGenSeg: Image Segmentation with Autoregressive Image Generation Model
- Alleviating Forgetfulness of Linear Attention by Hybrid Sparse Attention and Contextualized Learnable Token Eviction
- C-NAV: Towards Self-Evolving Continual Object Navigation in Open World
- A Scalable, Causal, and Energy Efficient Framework for Neural Decoding with Spiking Neural Networks
- Towards the Formalization of a Trustworthy AI for Mining Interpretable Models explOiting Sophisticated Algorithms
- BUSTED at AraGenEval Shared Task: A Comparative Study of Transformer-Based Models for Arabic AI-Generated Text Detection
- Speaking Clearly: A Simplified Whisper-Based Codec for Low-Bitrate Speech Coding
- NeuralTouch: Neural Descriptors for Precise Sim-to-Real Tactile Robot Control
- Positional Encoding Field
- NeuPerm: Disrupting Malware Hidden in Neural Network Parameters by Leveraging Permutation Symmetry
- GRATING: Low-Latency and Memory-Efficient Semantic Selection on Device
- A Transformer Inspired AI-based MIMO receiver
- Calibrating Multimodal Consensus for Emotion Recognition
- Teaching Language Models to Reason with Tools
- The Mirror Loop: Recursive Non-Convergence in Generative Reasoning Systems
- DB-FGA-Net: Dual Backbone Frequency Gated Attention Network for Multi-Class Brain Tumor Classification with Grad-CAM Interpretability
- FlexiReID: Adaptive Mixture of Expert for Multi-Modal Person Re-Identification
- EditInfinity: Image Editing with Binary-Quantized Generative Models
- Vox-Evaluator: Enhancing Stability and Fidelity for Zero-shot TTS with A Multi-Level Evaluator
- Diffusion Bridge Networks Simulate Clinical-grade PET from MRI for Dementia Diagnostics
- Think Parallax: Solving Multi-Hop Problems via Multi-View Knowledge-Graph-Based Retrieval-Augmented Generation
- Deep Learning Based Joint Space-Time-Frequency Domain Channel Prediction for Cell-Free Massive MIMO Systems
- Inverse Image-Based Rendering for Light Field Generation from Single Images
- SpeechAgent: An End-to-End Mobile Infrastructure for Speech Impairment Assistance
- AsyncHZP: Hierarchical ZeRO Parallelism with Asynchronous Scheduling for Scalable LLM Training
- Attentive Convolution: Unifying the Expressivity of Self-Attention with Convolutional Efficiency
- Hierarchical Dual-Head Model for Suicide Risk Assessment via MentalRoBERTa
- DyPE: Dynamic Position Extrapolation for Ultra High Resolution Diffusion
- Classical Feature Embeddings Help in BERT-Based Human Mobility Prediction
- InvDec: Inverted Decoder for Multivariate Time Series Forecasting with Separated Temporal and Variate Modeling
- ShapeX: Shapelet-Driven Post Hoc Explanations for Time Series Classification Models
- Meta-Learning for Cross-Task Generalization in Protein Mutation Property Prediction
- A Concrete Roadmap towards Safety Cases based on Chain-of-Thought Monitoring
- Re-evaluating Minimum Bayes Risk Decoding for Automatic Speech Recognition
- HybridEP: Scaling Expert Parallelism to Cross-Datacenter Scenario via Hybrid Expert/Data Transmission
- Exploring Conditions for Diffusion models in Robotic Control
- Rotate Both Ways: Time-and-Order RoPE for Generative Recommendation
- Enhancing Reasoning Skills in Small Persian Medical Language Models Can Outperform Large-Scale Data Training
- Speculative Sampling for Parametric Temporal Point Processes
- Forging GEMs: Advancing Greek NLP through Quality-Based Corpus Curation
- On Interaction Effects in Greybox Fuzzing
- Guiding diffusion models to reconstruct flow fields from sparse data
- Seed3D 1.0: From Images to High-Fidelity Simulation-Ready 3D Assets
- Data-Centric Lessons To Improve Speech-Language Pretraining
- Theoretical Refinement of CLIP by Utilizing Linear Structure of Optimal Similarity
- Deep Sequence-to-Sequence Models for GNSS Spoofing Detection
- A Survey on Cache Methods in Diffusion Models: Toward Efficient Multi-Modal Generation
- From Optimization to Prediction: Transformer-Based Path-Flow Estimation to the Traffic Assignment Problem
- SEMPO: Lightweight Foundation Models for Time Series Forecasting
- KARIPAP: Quantum-Inspired Tensor Network Compression of Large Language Models Using Infinite Projected Entangled Pair States and Tensor Renormalization Group
- Study of Training Dynamics for Memory-Constrained Fine-Tuning
- Unraveling Emotions with Pre-Trained Models
- RatioWaveNet: A Learnable RDWT Front-End for Robust and Interpretable EEG Motor-Imagery Classification
- Decomposed Attention Fusion in MLLMs for Training-Free Video Reasoning Segmentation
- Multi-modal Co-learning for Earth Observation: Enhancing single-modality models via modality collaboration
- DAIL: Beyond Task Ambiguity for Language-Conditioned Reinforcement Learning
- Demonstrating Real Advantage of Machine-Learning-Enhanced Monte Carlo for Combinatorial Optimization
- Automated HIV Screening on Dutch Electronic Health Records with Large Language Models
- Exploring "Many in Few" and "Few in Many" Properties in Long-Tailed, Highly-Imbalanced IC Defect Classification
- Yang-Mills Meets Data
- NeSyPr: Neurosymbolic Proceduralization For Efficient Embodied Reasoning
- ToMMeR -- Efficient Entity Mention Detection from Large Language Models
- AMAuT: A Flexible and Efficient Multiview Audio Transformer Framework Trained from Scratch
- Sign Language Translation with Sentence Embedding Supervision
- DARE: A Deformable Adaptive Regularization Estimator for Learning-Based Medical Image Registration
- An Experimental Study of Real-Life LLM-Proposed Performance Improvements
- Loopholing Discrete Diffusion: Deterministic Bypass of the Sampling Wall
- Difficulty-Controllable Multiple-Choice Question Generation Using Large Language Models and Direct Preference Optimization
- Selecting and Combining Large Language Models for Scalable Code Clone Detection
- Brain-Inspired Perspective on Configurations: Unsupervised Similarity and Early Cognition
- No Intelligence Without Statistics: The Invisible Backbone of Artificial Intelligence
- Aligning Multilingual News for Stock Return Prediction
- PruneHal: Reducing Hallucinations in Multi-modal Large Language Models through Adaptive KV Cache Pruning
- News-Aware Direct Reinforcement Trading for Financial Markets
- HAMLOCK: HArdware-Model LOgically Combined attacK
- Spatio-temporal Sign Language Representation and Translation
- Defending Against Prompt Injection with DataFilter
- Transformers are almost optimal metalearners for linear classification
- Overlap-weighted orthogonal meta-learner for treatment effect estimation over time
- Synthesizability Prediction of Crystalline Structures with a Hierarchical Transformer and Uncertainty Quantification
- PRGCN: A Graph Memory Network for Cross-Sequence Pattern Reuse in 3D Human Pose Estimation
- Coupled Transformer Autoencoder for Disentangling Multi-Region Neural Latent Dynamics
- Every Attention Matters: An Efficient Hybrid Architecture for Long-Context Reasoning
- GRASPLAT: Enabling dexterous grasping through novel view synthesis
- UniHPR: Unified Human Pose Representation via Singular Value Contrastive Learning
- A Graph Signal Processing Framework for Hallucination Detection in Large Language Models
- Advancing Brain Tumor Segmentation via Attention-based 3D U-Net Architecture and Digital Image Processing
- Cyberattack Detection in Critical Infrastructure and Supply Chains
- QKCV Attention: Enhancing Time Series Forecasting with Static Categorical Embeddings for Both Lightweight and Pre-trained Foundation Models
- When LRP Diverges from Leave-One-Out in Transformers
- Degeneracy-Aware Pulsar Parameter Estimation from Light Curves via Deep Learning and Test-Time Optimization
- Protein generation with embedding learning for motif diversification
- Topoformer: brain-like topographic organization in Transformer language models through spatial querying and reweighting
- Diffusion Buffer for Online Generative Speech Enhancement
- SemiAdapt and SemiLoRA: Efficient Domain Adaptation for Transformer-based Low-Resource Language Translation with a Case Study on Irish
- SITS-DECO: A Generative Decoder Is All You Need For Multitask Satellite Image Time Series Modelling
- Reasoning Language Model Inference Serving Unveiled: An Empirical Study
- On-Device Inference versus Wireless Streaming: Energy-Efficient Multi-Modal Deep Learning for Wearable Cardiovascular Patches
- Image augmentation with invertible networks in interactive satellite image change detection
- MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval
- Learning Time-Varying Turn-Taking Behavior in Group Conversations
- Informed Learning for Estimating Drought Stress at Fine-Scale Resolution Enables Accurate Yield Prediction
- Optimality and NP-Hardness of Transformers in Learning Markovian Dynamical Functions
- VAPU: System for Autonomous Legacy Code Modernization
- StarBench: A Turn-Based RPG Benchmark for Agentic Multimodal Decision-Making and Information Seeking
- Simple and Efficient Heterogeneous Temporal Graph Neural Network
- Embodied Navigation with Auxiliary Task of Action Description Prediction
- Adamas: Hadamard Sparse Attention for Efficient Long-Context Inference
- Benchmarking On-Device Machine Learning on Apple Silicon with MLX
- Biomechanically consistent real-time action recognition for human-robot interaction
- Training Diverse Graph Experts for Ensembles: A Systematic Empirical Study
- GPTFace: Generative Pre-training of Facial-Linguistic Transformer by Span Masking and Weakly Correlated Text-image Data
- Enhancing Few-Shot Classification of Benchmark and Disaster Imagery with ATTBHFA-Net
- Higher Embedding Dimension Creates a Stronger World Model for a Simple Sorting Task
- ParaStyleTTS: Toward Efficient and Robust Paralinguistic Style Control for Expressive Text-to-Speech Generation
- BrailleLLM: Braille Instruction Tuning with Large Language Models for Braille Domain Tasks
- Text or Pixels? It Takes Half: On the Token Efficiency of Visual Text Inputs in Multimodal LLMs
- DiffGRM: Diffusion-based Generative Recommendation Model
- Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs
- LIME: Link-based user-item Interaction Modeling with decoupled xor attention for Efficient test time scaling
- Unifying and Enhancing Graph Transformers via a Hierarchical Mask Framework
- Ensembling Pruned Attention Heads For Uncertainty-Aware Efficient Transformers
- MoGA: Mixture-of-Groups Attention for End-to-End Long Video Generation
- Condition-Invariant fMRI Decoding of Speech Intelligibility with Deep State Space Model
- Δt-Mamba3D: A Time-Aware Spatio-Temporal State-Space Model for Breast Cancer Risk Prediction
- Multi-Resolution Analysis of the Convective Structure of Tropical Cyclones for Short-Term Intensity Guidance
- Provable Generalization Bounds for Deep Neural Networks with Momentum-Adaptive Gradient Dropout
- Rethinking PCA Through Duality
- Any-Depth Alignment: Unlocking Innate Safety Alignment of LLMs to Any-Depth
- MEG-GPT: A transformer-based foundation model for magnetoencephalography data
- Cross-Domain Long-Term Forecasting: Radiation Dose from Sparse Neutron Sensor via Spatio-Temporal Operator Network
- Towards Robust Zero-Shot Reinforcement Learning
- Cortical-SSM: A Deep State Space Model for EEG and ECoG Motor Imagery Decoding
- Benchmarking Probabilistic Time Series Forecasting Models on Neural Activity
- Transformer Redesign for Late Fusion of Audio-Text Features on Ultra-Low-Power Edge Hardware
- ViBED-Net: Video Based Engagement Detection Network Using Face-Aware and Scene-Aware Spatiotemporal Cues
- Universal Spectral Tokenization via Self-Supervised Panchromatic Representation Learning
- OmniCast: A Masked Latent Diffusion Model for Weather Forecasting Across Time Scales
- Seeing but Not Believing: Probing the Disconnect Between Visual Attention and Answer Correctness in VLMs
- An Enhanced Dual Transformer Contrastive Network for Multimodal Sentiment Analysis
- Contextual Attention Modulation: Towards Efficient Multi-Task Adaptation in Large Language Models
- AI-Boosted Video Annotation: Assessing the Process Enhancement
- Towards 3D Objectness Learning in an Open World
- AtlasKV: Augmenting LLMs with Billion-Scale Knowledge Graphs in 20GB VRAM
- Lingua Custodi's participation at the WMT 2025 Terminology shared task
- Split-Fuse-Transport: Annotation-Free Saliency via Dual Clustering and Optimal Transport Alignment
- ReXMoE: Reusing Experts with Minimal Overhead in Mixture-of-Experts
- AdapTrack: Constrained Decoding without Distorting LLM's Output Intent
- Semantic-E2VID: a Semantic-Enriched Paradigm for Event-to-Video Reconstruction
- Explainability of Large Language Models: Opportunities and Challenges toward Generating Trustworthy Explanations
- A Prototypical Network with an Attention-based Encoder for Drivers Identification Application
- Round Outcome Prediction in VALORANT Using Tactical Features from Video Analysis
- SOLE: Hardware-Software Co-design of Softmax and LayerNorm for Efficient Transformer Inference
- ParaVul: A Parallel Large Language Model and Retrieval-Augmented Framework for Smart Contract Vulnerability Detection
- Temporally Detailed Hypergraph Neural ODEs for Disease Progression Modeling
- An Evaluation of LLMs Inference on Popular Single-board Computers
- Soft-Masked Diffusion Language Models
- One Dinomaly2 Detect Them All: A Unified Framework for Full-Spectrum Unsupervised Anomaly Detection
- HGAdapter: Hypergraph-based Adapters in Language Models for Code Summarization and Clone Detection
- Fighter: Unveiling the Graph Convolutional Nature of Transformers in Time Series Modeling
- Explainable Heterogeneous Anomaly Detection in Financial Networks via Adaptive Expert Routing
- Transformer-Based Low-Resource Language Translation: A Study on Standard Bengali to Sylheti
- Matricial Free Energy as a Gaussianizing Regularizer: Enhancing Autoencoders for Gaussian Code Generation
- KineDiff3D: Kinematic-Aware Diffusion for Category-Level Articulated Object Shape Reconstruction and Generation
- Saber: An Efficient Sampling with Adaptive Acceleration and Backtracking Enhanced Remasking for Diffusion Language Model
- Conveying Meaning through Gestures: An Investigation into Semantic Co-Speech Gesture Generation
- Localist LLMs with Recruitment Learning
- Person Re-Identification via Generalized Class Prototypes
- DINO-CVA: A Multimodal Goal-Conditioned Vision-to-Action Model for Autonomous Catheter Navigation
- An empirical study of the effect of video encoders on Temporal Video Grounding
- Adaptive Deterministic Flow Matching for Target Speaker Extraction
- L-MoE: End-to-End Training of a Lightweight Mixture of Low-Rank Adaptation Experts
- EventFormer: A Node-graph Hierarchical Attention Transformer for Action-centric Video Event Prediction
- Connecting Domains and Contrasting Samples: A Ladder for Domain Generalization
- DeepChem Equivariant: SE(3)-Equivariant Support in an Open-Source Molecular Machine Learning Library
- Uncovering Brain-Like Hierarchical Patterns in Vision-Language Models through fMRI-Based Neural Encoding
- Neuronal Group Communication for Efficient Neural representation
- An Efficient Semantic Segmentation Decoder for In-Car or Distributed Applications
- Efficient High-Accuracy PDEs Solver with the Linear Attention Neural Operator
- Domain Generalizable Continual Learning
- ProtoMol: Enhancing Molecular Property Prediction via Prototype-Guided Multimodal Learning
- Zero-Shot Performance Prediction for Probabilistic Scaling Laws
- Renaissance of RNNs in Streaming Clinical Time Series: Compact Recurrence Remains Competitive with Transformers
- Domain-Contextualized Concept Graphs: A Computable Framework for Knowledge Representation
- 3D-GSRD: 3D Molecular Graph Auto-Encoder with Selective Re-mask Decoding
- Closing the Curvature Gap: Full Transformer Hessians and Their Implications for Scaling Laws
- The Sherpa.ai Blind Vertical Federated Learning Paradigm to Minimize the Number of Communications
- All You Need is One: Capsule Prompt Tuning with a Single Vector
- Improving Model Representation and Reducing KV Cache via Skip Connections with First Value Heads
- WaMaIR: Image Restoration via Multiscale Wavelet Convolutions and Mamba-based Channel Modeling with Texture Enhancement
- U-Codec: Ultra Low Frame-rate Neural Speech Codec for Fast High-fidelity Speech Generation
- MOSAIC: Masked Objective with Selective Adaptation for In-domain Contrastive Learning
- Mixed-Precision Quantization for Language Models: Techniques and Prospects
- Universal and Transferable Attacks on Pathology Foundation Models
- Advances in Pre-trained Language Models for Domain-Specific Text Classification: A Systematic Review
- TraceCoder: Towards Traceable ICD Coding via Multi-Source Knowledge Integration
- Deep learning for flash drought forecasting and interpretation
- Multi-dimensional Data Analysis and Applications Basing on LLM Agents and Knowledge Graph Interactions
- Attention-guided few-shot learning for metal surface defect classification
- HEADER: Hierarchical Robot Exploration via Attention-Based Deep Reinforcement Learning with Expert-Guided Reward
- AB-UPT for Automotive and Aerospace Applications
- Dynamic Recalibration in LiDAR SLAM: Integrating AI and Geometric Methods with Real-Time Feedback Using INAF Fusion
- Deep Compositional Phase Diffusion for Long Motion Sequence Generation
- Extending Audio Context for Long-Form Understanding in Large Audio-Language Models
- Early-stopping for Transformer model training
- Rethinking Convergence in Deep Learning: The Predictive-Corrective Paradigm for Anatomy-Informed Brain MRI Segmentation
- SADCHER: Scheduling using Attention-based Dynamic Coalitions of Heterogeneous Robots in Real-Time
- TeamFormer: Shallow Parallel Transformers with Progressive Approximation
- VO-DP: Semantic-Geometric Adaptive Diffusion Policy for Vision-Only Robotic Manipulation
- Enhancing reliability in AI inference services: An empirical study on real production incidents
- The Right to Be Remembered: Preserving Maximally Truthful Digital Memory in the Age of AI
- Exploring the Synergy of Quantitative Factors and Newsflow Representations from Large Language Models for Stock Return Prediction
- SNOO: Step-K Nesterov Outer Optimizer - The Surprising Effectiveness of Nesterov Momentum Applied to Pseudo-Gradients
- Auditing Algorithmic Bias in Transformer-Based Trading
- Polarization based direction of arrival estimation using a radio interferometric array
- Measuring the Effect of Disfluency in Multilingual Knowledge Probing Benchmarks
- PoTS: Proof-of-Training-Steps for Backdoor Detection in Large Language Models
- Decorrelation Speeds Up Vision Transformers
- Terra: Explorable Native 3D World Model with Point Latents
- Ponimator: Unfolding Interactive Pose for Versatile Human-human Interaction Animation
- Attention Is All You Need for KV Cache in Diffusion LLMs
- AI-Powered Early Diagnosis of Mental Health Disorders from Real-World Clinical Conversations
- DeLeaker: Dynamic Inference-Time Reweighting For Semantic Leakage Mitigation in Text-to-Image Models
- A Hard-Label Black-Box Evasion Attack against ML-based Malicious Traffic Detection Systems
- Multi-modal video data-pipelines for machine learning with minimal human supervision
- LightQANet: Quantized and Adaptive Feature Learning for Low-Light Image Enhancement
- To Infinity and Beyond: Tool-Use Unlocks Length Generalization in State Space Models
- Capturing Context-Aware Route Choice Semantics for Trajectory Representation Learning
- Tackling Time-Series Forecasting Generalization via Mitigating Concept Drift
- Rethinking Hebbian Principle: Low-Dimensional Structural Projection for Unsupervised Learning
- A Novel GPT-Based Framework for Anomaly Detection in System Logs
- Inpainting the Red Planet: Diffusion Models for the Reconstruction of Martian Environments in Virtual Reality
- Cross-Layer Feature Self-Attention Module for Multi-Scale Object Detection
- Tawa: Automatic Warp Specialization for Modern GPUs with Asynchronous References
- Causality Enhancement for Cross-Domain Recommendation
- Adapting Self-Supervised Representations as a Latent Space for Efficient Generation
- GOPLA: Generalizable Object Placement Learning via Synthetic Augmentation of Human Arrangement
- ScalePool: Hybrid XLink-CXL Fabric for Composable Resource Disaggregation in Unified Scale-up Domains
- State-Space Models for Tabular Prior-Data Fitted Networks
- Spatially Aware Self-Supervised Models for Multi-Channel Neural Speaker Diarization
- Hybrid Autoencoder-Based Framework for Early Fault Detection in Wind Turbines
- EcoScaleNet: A Lightweight Multi Kernel Network for Long Sequence 12 lead ECG Classification
- QuASH: Using Natural-Language Heuristics to Query Visual-Language Robotic Maps
- A Deep State-Space Model Compression Method using Upper Bound on Output Error
- Exploring Image Representation with Decoupled Classical Visual Descriptors
- Towards Generalist Intelligence in Dentistry: Vision Foundation Models for Oral and Maxillofacial Radiology
- Vision Mamba for Permeability Prediction of Porous Media
- Enhancing Time Series Forecasting through Selective Representation Spaces: A Patch Perspective
- Stealthy Dual-Trigger Backdoors: Attacking Prompt Tuning in LM-Empowered Graph Foundation Models
- Pruning Overparameterized Multi-Task Networks for Degraded Web Image Restoration
- Even Faster Kernel Matrix Linear Algebra via Density Estimation
- Hierarchical Semantic Retrieval with Cobweb
- DCMIL: A Progressive Representation Learning of Whole Slide Images for Cancer Prognosis Analysis
- ASecond-Order SpikingSSM for Wearables
- DRBD-Mamba for Robust and Efficient Brain Tumor Segmentation with Analytical Insights
- Leveraging Cycle-Consistent Anchor Points for Self-Supervised RGB-D Registration
- Vision-Centric Activation and Coordination for Multimodal Large Language Models
- Beyond a Single Perspective: Towards a Realistic Evaluation of Website Fingerprinting Attacks
- MorphoBench: A Benchmark with Difficulty Adaptive to Model Reasoning
- MatchAttention: Matching the Relative Positions for High-Resolution Cross-View Matching
- Synergistic Integration and Discrepancy Resolution of Contextualized Knowledge for Personalized Recommendation
- A Physics Prior-Guided Dual-Stream Attention Network for Motion Prediction of Elastic Bragg Breakwaters
- PIA: Deepfake Detection Using Phoneme-Temporal and Identity-Dynamic Analysis
- LOTA: Bit-Planes Guided AI-Generated Image Detection
- Joint Modeling of Big Five and HEXACO for Multimodal Apparent Personality-trait Recognition
- Predicting sequence-specific amplification efficiency in multi-template PCR with deep learning
- TGT: Text-Grounded Trajectories for Locally Controlled Video Generation
- pi-Flow: Policy-Based Few-Step Generation via Imitation Distillation
- Beat Tracking as Object Detection
- Inferred global dense residue transition graphs from primary structure sequences enable protein interaction prediction via directed graph convolutional neural networks
- Learning Multi-Index Models with Hyper-Kernel Ridge Regression
- Unraveling Syntax: Language Modeling and the Substructure of Grammars
- Capture, Canonicalize, Splat: Zero-Shot 3D Gaussian Avatars from Unstructured Phone Images
- Robotic Classification of Divers' Swimming States using Visual Pose Keypoints as IMUs
- Breadcrumbs Reasoning: Memory-Efficient Reasoning with Compression Beacons
- PriorGuide: Test-Time Prior Adaptation for Simulation-Based Inference
- Adam or Gauss-Newton? A Comparative Study In Terms of Basis Alignment and SGD Noise
- APRIL: Auxiliary Physically-Redundant Information in Loss - A physics-informed framework for parameter estimation with a gravitational-wave case study
- Message Passing on the Edge: Towards Scalable and Expressive GNNs
- NOSA: Native and Offloadable Sparse Attention
- EEGChaT: A Transformer-Based Modular Channel Selector for SEEG Analysis
- Element2Vec: Build Chemical Element Representation from Text for Property Prediction
- LLM one-shot style transfer for Authorship Attribution and Verification
- Neural Sum-of-Squares: Certifying the Nonnegativity of Polynomials with Transformers
- F-BFQ: Flexible Block Floating-Point Quantization Accelerator for LLMs
- A New Perspective on Transformers in Online Reinforcement Learning for Continuous Control
- AOAD-MAT: Transformer-based multi-agent deep reinforcement learning model considering agents' order of action decisions
- STAR: Boosting Time Series Foundation Models for Anomaly Detection through State-aware Adapter
- BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Disaggregated LLM Serving in AI Infrastructure
- Approximate Bilevel Graph Structure Learning for Histopathology Image Classification
- CMIS-Net: A Cascaded Multi-Scale Individual Standardization Network for Backchannel Agreement Estimation
- Multi-Agent Design Assistant for the Simulation of Inertial Fusion Energy
- GatePro: Parameter-Free Expert Selection Optimization for Mixture-of-Experts Models
- Transformer-based Scalable Beamforming Optimization via Deep Residual Learning
- Reciprocal Space Attention for Learning Long-Range Interactions
- Higher Satisfaction, Lower Cost: A Technical Report on How LLMs Revolutionize Meituan's Intelligent Interaction Systems
- DMTrack: Deformable State-Space Modeling for UAV Multi-Object Tracking with Kalman Fusion and Uncertainty-Aware Association
- Universal Image Restoration Pre-training via Masked Degradation Classification
- The Mechanistic Emergence of Symbol Grounding in Language Models
- OS-HGAdapter: Open Semantic Hypergraph Adapter for Large Language Models Assisted Entropy-Enhanced Image-Text Alignment
- Reasoning in Space via Grounding in the World
- AVAR-Net: A Lightweight Audio-Visual Anomaly Recognition Framework with a Benchmark Dataset
- Mirror Speculative Decoding: Breaking the Serial Barrier in LLM Inference
- MedREK: Retrieval-Based Editing for Medical LLMs with Key-Aware Prompts
- Novel Class Discovery for Point Cloud Segmentation via Joint Learning of Causal Representation and Reasoning
- Behavioral Embeddings of Programs: A Quasi-Dynamic Approach for Optimization Prediction
- Beyond Static LLM Policies: Imitation-Enhanced Reinforcement Learning for Recommendation
- Reinforcement Learning Meets Masked Generative Models: Mask-GRPO for Text-to-Image Generation
- PIShield: Detecting Prompt Injection Attacks via Intrinsic LLM Features
- On the Reasoning Abilities of Masked Diffusion Language Models
- GRIDAI: Generating and Repairing Intrusion Detection Rules via Collaboration among Multiple LLM-based Agents
- Don't Be Greedy, Just Relax! Pruning LLMs via Frank-Wolfe
- VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models
- Litespark Technical Report: High-Throughput, Energy-Efficient LLM Training Framework
- A Critical Review of the Need for Knowledge-Centric Evaluation of Quranic Recitation
- CSI-4CAST: A Hybrid Deep Learning Model for CSI Prediction with Comprehensive Robustness and Generalization Testing
- Behavioral Biometrics for Automatic Detection of User Familiarity in VR
- Computationally Efficient Neural Receivers via Axial Self-Attention
- Narrow Finetuning Leaves Clearly Readable Traces in Activation Differences
- CAMNet: Leveraging Cooperative Awareness Messages for Vehicle Trajectory Prediction
- What If : Understanding Motion Through Sparse Interactions
- Convolutional Attention in Betting Exchange Markets
- Which Word Orders Facilitate Length Generalization in LMs? An Investigation with GCG-Based Artificial Languages
- Heterogeneous Graph Representation of Stiffened Panels with Non-Uniform Boundary Conditions and Loads
- Assessing the Potential for Catastrophic Failure in Dynamic Post-Training Quantization
- CoRA: Covariate-Aware Adaptation of Time Series Foundation Models
- On the Use of Hierarchical Vision Foundation Models for Low-Cost Human Mesh Recovery and Pose Estimation
- On Foundation Models for Temporal Point Processes to Accelerate Scientific Discovery
- Learning-To-Measure: In-context Active Feature Acquisition
- LayerSync: Self-aligning Intermediate Layers
- Learning Human Motion with Temporally Conditional Mamba
- K-frames: Scene-Driven Any-k Keyframe Selection for long video understanding
- Learning to Recognize Correctly Completed Procedure Steps in Egocentric Assembly Videos through Spatio-Temporal Modeling
- Self-attention enabled quantum path analysis of high-harmonic generation in solids
- Efficient Adaptive Transformer: An Empirical Study and Reproducible Framework
- A Hierarchical Quantized Tokenization Framework for Task-Adaptive Graph Representation Learning
- Simple Projection Variants Improve ColBERT Performance
- Tensor Logic: The Language of AI
- HiLoRA: Adaptive Hierarchical LoRA Routing for Training-Free Domain Generalization
- A Gradient Guided Diffusion Framework for Chance Constrained Programming
- SDGraph: Multi-Level Sketch Representation Learning by Sparse-Dense Graph Architecture
- iCodeReviewer: Improving Secure Code Review with Mixture of Prompts
- CompoDistill: Attention Distillation for Compositional Reasoning in Multimodal LLMs
- State Space Prompting via Gathering and Spreading Spatio-Temporal Information for Video Understanding
- Self-Verifying Reflection Helps Transformers with CoT Reasoning
- Probabilistic Super-Resolution for Urban Micrometeorology via a Schrödinger Bridge
- Credal Transformer: A Principled Approach for Quantifying and Mitigating Hallucinations in Large Language Models
- Chimera: State Space Models Beyond Sequences
- DRL: Discriminative Representation Learning with Parallel Adapters for Class Incremental Learning
- SpikePool: Event-driven Spiking Transformer with Pooling Attention
- Compressibility Measures Complexity: Minimum Description Length Meets Singular Learning Theory
- CrossRay3D: Geometry and Distribution Guidance for Efficient Multimodal 3D Detection
- Mamba Can Learn Low-Dimensional Targets In-Context via Test-Time Feature Learning
- SAIL-Embedding Technical Report: Omni-modal Embedding Foundation Model
- Multi-Action Self-Improvement for Neural Combinatorial Optimization
- SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models
- TFGA-Net: Temporal-Frequency Graph Attention Network for Brain-Controlled Speaker Extraction
- There is No VAE: End-to-End Pixel-Space Generative Modeling via Self-Supervised Pre-training
- Ethic-BERT: An Enhanced Deep Learning Model for Ethical and Non-Ethical Content Classification
- SMEC: Rethinking Matryoshka Representation Learning for Retrieval Embedding Compression
- DSAS: A Universal Plug-and-Play Framework for Attention Optimization in Multi-Document Question Answering
- Attribution Graphs and Causal Probing for Mechanistic Discovery and Bias Repair in Multimodal Generative Learning
- PAINT: Parallel-in-time Neural Twins for Dynamical System Reconstruction
- MosaicDiff: Training-free Structural Pruning for Diffusion Model Acceleration Reflecting Pretraining Dynamics
- Catch Your Breath: Adaptive Computation for Self-Paced Sequence Production
- Inpainting the Neural Picture: Inferring Unrecorded Brain Area Dynamics from Multi-Animal Datasets
- LLM Reasoning for Machine Translation: Synthetic Data Generation over Thinking Tokens
- Robust Adversarial Reinforcement Learning in Stochastic Games via Sequence Modeling
- Bound on entanglement in neural quantum states
- Dimension-Free Minimax Rates for Learning Pairwise Interactions in Attention-Style Models
- DiT360: High-Fidelity Panoramic Image Generation via Hybrid Training
- Continual Personalization for Diffusion Models
- Beyond 'Templates': Category-Agnostic Object Pose, Size, and Shape Estimation from a Single View
- NV3D: Leveraging Spatial Shape Through Normal Vector-based 3D Object Detection
- Attention Factors for Statistical Arbitrage
- Deconstructing Attention: Investigating Design Principles for Effective Language Modeling
- MS-Mix: Sentiment-Guided Adaptive Augmentation for Multimodal Sentiment Analysis
- Massive Activations are the Key to Local Detail Synthesis in Diffusion Transformers
- An Encoder-Integrated PhoBERT with Graph Attention for Vietnamese Token-Level Classification
- Hallucination Detection via Internal States and Structured Reasoning Consistency in Large Language Models
- LikePhys: Evaluating Intuitive Physics Understanding in Video Diffusion Models via Likelihood Preference
- Numerical Methods for Kernel Slicing
- Unifying Deductive and Abductive Reasoning in Knowledge Graphs with Masked Diffusion Model
- Joint Discriminative-Generative Modeling via Dual Adversarial Training
- Event-Aware Prompt Learning for Dynamic Graphs
- DiffStyleTS: Diffusion Model for Style Transfer in Time Series
- Next Interest Flow: A Generative Pre-training Paradigm for Recommender Systems by Modeling All-domain Movelines
- FOSSIL: Harnessing Feedback on Suboptimal Samples for Data-Efficient Generalisation with Imitation Learning for Embodied Vision-and-Language Tasks
- Channel-Aware Deep Learning for Superimposed Pilot Power Allocation and Receiver Design
- CNSocialDepress: A Chinese Social Media Dataset for Depression Risk Detection and Structured Analysis
- Investigating Identity Signals in Conversational Facial Dynamics via Disentangled Expression Features
- Fairness Metric Design Exploration in Multi-Domain Moral Sentiment Classification using Transformer-Based Models
- An Explorative Study on Distributed Computing Techniques in Training and Inference of Large Language Models
- Evaluating Line-level Localization Ability of Learning-based Code Vulnerability Detection Models
- Efficient In-Memory Acceleration of Sparse Block Diagonal LLMs
- Reliable Cross-modal Alignment via Prototype Iterative Construction
- CoPRS: Learning Positional Prior from Chain-of-Thought for Reasoning Segmentation
- A Comprehensive Forecasting-Based Framework for Time Series Anomaly Detection: Benchmarking on the Numenta Anomaly Benchmark (NAB)
- Robust Photoplethysmography Signal Denoising via Mamba Networks
- Compositional Zero-Shot Learning: A Survey
- VCB Bench: An Evaluation Benchmark for Audio-Grounded Large Language Model Conversational Agents
- Text-Enhanced Panoptic Symbol Spotting in CAD Drawings
- ROFI: A Deep Learning-Based Ophthalmic Sign-Preserving and Reversible Patient Face Anonymizer
- Decoupled Multimodal Fusion for User Interest Modeling in Click-Through Rate Prediction
- PruneGCRN: Minimizing and explaining spatio-temporal problems through node pruning
- Rethinking deep learning: linear regression remains a key benchmark in predicting terrestrial water storage
- High-Resolution Spatiotemporal Modeling with Global-Local State Space Models for Video-Based Human Pose Estimation
- Automating Structural Engineering Workflows with Large Language Model Agents
- ShishuLM: Lightweight Language Model with Hybrid Decoder-MLP Architecture and Paired Weight Sharing
- In-Context Learning Is Provably Bayesian Inference: A Generalization Theory for Meta-Learning
- Paving the Way Towards Kinematic Assessment Using Monocular Video: A Preclinical Benchmark of State-of-the-Art Deep-Learning-Based 3D Human Pose Estimators Against Inertial Sensors in Daily Living Activities
- SceneTextStylizer: A Training-Free Scene Text Style Transfer Framework with Diffusion Model
- xLSTM Scaling Laws: Competitive Performance with Linear Time-Complexity
- QLENS: Towards A Quantum Perspective of Language Transformers
- Enhancing Large Language Model Reasoning via Selective Critical Token Fine-Tuning
- Does LLM Focus on the Right Words? Mitigating Context Bias in LLM-based Recommenders
- HoMer: Addressing Heterogeneities by Modeling Sequential and Set-wise Contexts for CTR Prediction
- Audio-Guided Visual Perception for Audio-Visual Navigation
- Evaluating the Explainability of Vision Transformers in Medical Imaging
- Flow Matching-Based Autonomous Driving Planning with Advanced Interactive Behavior Modeling
- Multimodal Disease Progression Modeling via Spatiotemporal Disentanglement and Multiscale Alignment
- PAGE: Prompt Augmentation for text Generation Enhancement
- PESTalk: Speech-Driven 3D Facial Animation with Personalized Emotional Styles
- Cognitive Load Traces as Symbolic and Visual Accounts of Deep Model Cognition
- Direct Multi-Token Decoding
- Learning the syntax of plant assemblages
- Visual Odometry with Transformers
- Early Detection and Reduction of Memorisation for Domain Adaptation and Instruction Tuning
- BioOSS: A Bio-Inspired Oscillatory State System with Spatio-Temporal Dynamics
- Structured Spectral Graph Representation Learning for Multi-label Abnormality Analysis from 3D CT Scans
- Structure Over Signal: A Globalized Approach to Multi-relational GNNs for Stock Prediction
- How to Find Fantastic AI Papers: Self-Rankings as a Powerful Predictor of Scientific Impact Beyond Peer Review
- Attention-Enhanced LSTM Modeling for Improved Temperature and Rainfall Forecasting in Bangladesh
- Action-Dynamics Modeling and Cross-Temporal Interaction for Online Action Understanding
- From Craft to Constitution: A Governance-First Paradigm for Principled Agent Engineering
- Interactive Atmospheric Composition Emulation for Next-Generation Earth System Models
- Trustworthy Retrosynthesis: Eliminating Hallucinations with a Diverse Ensemble of Reaction Scorers
- Catalyst GFlowNet for electrocatalyst design: A hydrogen evolution reaction case study
- GraphTARIF: Linear Graph Transformer with Augmented Rank and Improved Focus
- dN/dx Reconstruction with Deep Learning for High-Granularity TPCs
- ADiP: Adaptive Precision Systolic Array for Matrix Multiplication Acceleration
- DCP: Addressing Input Dynamism In Long-Context Training via Dynamic Context Parallelism
- Self-Attention to Operator Learning-based 3D-IC Thermal Simulation
- A Simple and Better Baseline for Visual Grounding
- Injecting Frame-Event Complementary Fusion into Diffusion for Optical Flow in Challenging Scenes
- Deep semi-supervised approach based on consistency regularization and similarity learning for weeds classification
- BitMar: Low-Bit Multimodal Fusion with Episodic Memory for Edge Devices
- Multi-scale Frequency-Aware Adversarial Network for Parkinson's Disease Assessment Using Wearable Sensors
- Self-Supervised Representation Learning with ID-Content Modality Alignment for Sequential Recommendation
- MCE: Towards a General Framework for Handling Missing Modalities under Imbalanced Missing Rates
- Layout-Independent License Plate Recognition via Integrated Vision and Language Models
- Assessing Large Language Models for Structured Medical Order Extraction
- Personalized Motion Guidance Framework for Athlete-Centric Coaching
- The Hidden DNA of LLM-Generated JavaScript: Structural Patterns Enable High-Accuracy Authorship Attribution
- Head-wise Adaptive Rotary Positional Encoding for Fine-Grained Image Generation
- SASER: Stego attacks on open-source LLMs
- CAPSim: A Fast CPU Performance Simulator Using Attention-based Predictor
- MSF-Mamba: Motion-aware State Fusion Mamba for Efficient Micro-Gesture Recognition
- Harnessing Consistency for Robust Test-Time LLM Ensemble
- Taming a Retrieval Framework to Read Images in Humanlike Manner for Augmenting Generation of MLLMs
- Softmax ≥ Linear: Transformers may learn to classify in-context by kernel gradient descent
- Knowing Unknowns in an Age of Information Overload
- SoundReactor: Frame-level Online Video-to-Audio Generation
- GrifFinNet: A Graph-Relation Integrated Transformer for Financial Predictions
- PRISM: Enhancing Protein Inverse Folding through Fine-Grained Retrieval on Structure-Sequence Multimodal Representations
- Latent Retrieval Augmented Generation of Cross-Domain Protein Binders
- ParsVoice: A Large-Scale Multi-Speaker Persian Speech Corpus for Text-to-Speech Synthesis
- When Tracking Fails: Analyzing Failure Modes of SAM2 for Point-Based Tracking in Surgical Videos
- Preconditioned Norms: A Unified Framework for Steepest Descent, Quasi-Newton and Adaptive Methods
- UltraLLaDA: Scaling the Context Length to 128K for Diffusion Large Language Models
- A3RNN: Bi-directional Fusion of Bottom-up and Top-down Process for Developmental Visual Attention in Robots
- Bridging Semantics & Structure for Software Vulnerability Detection using Hybrid Network Models
- Embodiment in multimodal large language models
- SP-MoE: Speculative Decoding and Prefetching for Accelerating MoE-based Model Inference
- MetaBreak: Jailbreaking Online LLM Services via Special Token Manipulation
- Enhancing the Cross-Size Generalization for Solving Vehicle Routing Problems via Continual Learning
- Are Video Models Emerging as Zero-Shot Learners and Reasoners in Medical Imaging?
- Learning to Guarantee Type Correctness in Code Generation through Type-Guided Program Synthesis
- CauchyNet: Compact and Data-Efficient Learning using Holomorphic Activation Functions
- A Survey of Inductive Reasoning for Large Language Models
- LLMs are All You Need? Improving Fuzz Testing for MOJO with Large Language Models
- Peransformer: Improving Low-informed Expressive Performance Rendering with Score-aware Discriminator
- A Unified Frequency Domain Decomposition Framework for Interpretable and Robust Time Series Forecasting
- YOLOv11-Litchi: Efficient Litchi Fruit Detection based on UAV-Captured Agricultural Imagery in Complex Orchard Environments
- CacheClip: Accelerating RAG with Effective KV Cache Reuse
- Breaking the Likelihood Trap: Consistent Generative Recommendation with Graph-structured Model
- Translution: Unifying Self-attention and Convolution for Adaptive and Relative Modeling
- PANTHER: Generative Pretraining Beyond Language for Sequential User Behavior Modeling
- Automated Glaucoma Report Generation via Dual-Attention Semantic Parallel-LSTM and Multimodal Clinical Data Integration
- Automatic Speech Recognition in the Modern Era: Architectures, Training, and Evaluation
- Mathematical Modeling and Convergence Analysis of Deep Neural Networks with Dense Layer Connectivities in Deep Learning
- Variational Secret Common Randomness Extraction
- Conformal Sparsification for Bandwidth-Efficient Edge-Cloud Speculative Decoding
- ACT: Automatically Generating Compiler Backends from Tensor Accelerator ISA Descriptions
- ClustViT: Clustering-based Token Merging for Semantic Segmentation
- Inverse Language Modeling towards Robust and Grounded LLMs
- PVDetector: Detecting Prompt Injection Attacks on Purpose-Specific LLM Agents through Policy-Violation Concept Analysis
- GPU-Tile-Sim: A Tile-Centric GPU Simulation Framework for LLM Hardware-Software Co-Design
- Flow-Matching Guided Deep Unfolding for Hyperspectral Image Reconstruction
- Exploring Large Language Models for Financial Applications: Techniques, Performance, and Challenges with FinMA
- Multimodal Foundation Models for Early Disease Detection
- Tapered Language Models
- A Modular Theory of Subjective Consciousness for Natural and Artificial Minds
- Microscaling Floating Point Formats for Large Language Models
- Compositional meta-learning through probabilistic task inference
- Cheap Reward Hacking Detection
- Forget Attention: Importance-Aware Attention Is All You Need
- Towards Speeding up Program Repair with Non-Autoregressive Model
- Sparse Query Attention (SQA): A Computationally Efficient Attention Mechanism with Query Heads Reduction
- Rethinking the shape convention of an MLP
- Attention Once Is All You Need: Efficient Streaming Inference with Stateful Transformers
- Towards Photonic Band Diagram Generation with Transformer-Latent Diffusion Models
- Emotional Text-To-Speech Based on Mutual-Information-Guided Emotion-Timbre Disentanglement
- OntoLogX: Ontology-Guided Knowledge Graph Extraction from Cybersecurity Logs with Large Language Models
- Accelerating Attention with Basis Decomposition
- PyramidStyler: Transformer-Based Neural Style Transfer with Pyramidal Positional Encoding and Reinforcement Learning
- On the Quantization Robustness of Diffusion Language Models in Coding Benchmarks
- TinyTorch: Building Machine Learning Systems from First Principles
- The Geometry of Forgetting
- SOL-ExecBench: Speed-of-Light Benchmarking for Real-World GPU Kernels Against Hardware Limits
- NCCL EP: Towards a Unified Expert Parallel Communication API for NCCL
- Benchmarking Language Modeling for Lossless Compression of Full-Fidelity Audio
- PolyBlocks: A Compiler Infrastructure for AI Chips and Programming Frameworks
- DualPath: Breaking the Storage Bandwidth Bottleneck in Agentic LLM Inference
- Adversarial News and Lost Profits: Manipulating Headlines in LLM-Driven Algorithmic Trading
- Dynamic Stress Detection: A Study of Temporal Progression Modelling of Stress in Speech
- Bidirectional Time-Frequency Pyramid Network for Enhanced Robust EEG Classification
- Advancing Intoxication Detection: A Smartwatch-Based Approach
- Holistic Order Prediction in Natural Scenes
- TAWRMAC: A Novel Dynamic Graph Representation Learning Method
- DELTA: Dynamic Layer-Aware Token Attention for Efficient Long-Context Reasoning
- Token Is All You Price
- Domain Knowledge Infused Conditional Generative Models for Accelerating Drug Discovery
- Latent-Feature-Informed Neural ODE Modeling for Lightweight Stability Evaluation of Black-box Grid-Tied Inverters
- A Generic Machine Learning Framework for Radio Frequency Fingerprinting
- PromptGuard at BLP-2025 Task 1: A Few-Shot Classification Framework Using Majority Voting and Keyword Similarity for Bengali Hate Speech Detection
- Hierarchical Multi-Modal Threat Intelligence Fusion Without Aligned Data: A Practical Framework for Real-World Security Operations
- HeSRN: Representation Learning On Heterogeneous Graphs via Slot-Aware Retentive Network
- FSP-DETR: Few-Shot Prototypical Parasitic Ova Detection
- Architecture Induces Structural Invariant Manifolds of Neural Network Training Dynamics
- Titans Revisited: A Lightweight Reimplementation and Critical Analysis of a Test-Time Memory Model
- Improving AGI Evaluation: A Data Science Perspective
- On the Representations of Entities in Auto-regressive Large Language Models
- ARROW: An Adaptive Rollout and Routing Method for Global Weather Forecasting
- Design Principles for Sequence Models via Coefficient Dynamics
- Randomized HyperSteiner: A Stochastic Delaunay Triangulation Heuristic for the Hyperbolic Steiner Minimal Tree
- Entropy Meets Importance: A Unified Head Importance-Entropy Score for Stable and Efficient Transformer Pruning
- Tag-Enriched Multi-Attention with Large Language Models for Cross-Domain Sequential Recommendation
- AdaPM: a Partial Momentum Algorithm for LLM Training
- 3D Reconstruction from Transient Measurements with Time-Resolved Transformer
- Online Video Depth Anything: Temporally-Consistent Depth Prediction with Low Memory Consumption
- Efficient Resource-Constrained Training of Vision Transformers via Subspace Optimization
- Cross-Representation Benchmarking in Time-Series Electronic Health Records for Clinical Outcome Prediction
- LitE-SQL: A Lightweight and Efficient Text-to-SQL Framework with Vector-based Schema Linking and Execution-Guided Self-Correction
- From Explainability to Action: A Generative Operational Framework for Integrating XAI in Clinical Mental Health Screening
- Value-State Gated Attention for Mitigating Extreme-Token Phenomena in Transformers
- Group-Adaptive Adversarial Learning for Robust Fake News Detection Against Malicious Comments
- PHyCLIP: ℓ1-Product of Hyperbolic Factors Unifies Hierarchy and Compositionality in Vision-Language Representation Learning
- Psyzkaller: Learning from Historical and On-the-Fly Execution Data for Smarter Seed Generation in OS kernel Fuzzing
- Verifying Chain-of-Thought Reasoning via Its Computational Graph
- CARLE: A Hybrid Deep-Shallow Learning Framework for Robust and Explainable RUL Estimation of Rolling Element Bearings
- TIT: A Tree-Structured Instruction Tuning Approach for LLM-Based Code Translation
- Analytical Survey of Learning with Low-Resource Data: From Analysis to Investigation
- Efficient Autoregressive Inference for Transformer Probabilistic Models
- Fundamentals of Building Autonomous LLM Agents
- MambaH-Fit: Rethinking Hyper-surface Fitting-based Point Cloud Normal Estimation via State Space Modelling
- Goal-oriented Backdoor Attack against Vision-Language-Action Models via Physical Objects
- Modeling Time-Lapse Trajectories to Characterize Cranberry Growth
- Boosting Multi-modal Keyphrase Prediction with Dynamic Chain-of-Thought in Vision-Language Models
- Task-Level Insights from Eigenvalues across Sequence Models
- Progressive Uncertainty-Guided Evidential U-KAN for Trustworthy Medical Image Segmentation
- Cross-attention Secretly Performs Orthogonal Alignment in Recommendation Models
- Geo-Aware Models for Stream Temperature Prediction across Different Spatial Regions and Scales
- NL2GenSym: Natural Language to Generative Symbolic Rules for SOAR Cognitive Architecture via Large Language Models
- LLM Based Long Code Translation using Identifier Replacement
- Spatially-Augmented Sequence-to-Sequence Neural Diarization for Meetings
- Discrete Facial Encoding: : A Framework for Data-driven Facial Display Discovery
- Chain-of-Influence: Tracing Interdependencies Across Time and Features in Clinical Predictive Modelings
- PairSem: LLM-Guided Pairwise Semantic Matching for Scientific Document Retrieval
- Understanding and Predicting Temporal Visual Attention Influenced by Dynamic Highlights in Monitoring Task
- ReSplat: Learning Recurrent Gaussian Splats
- X2Video: Adapting Diffusion Models for Multimodal Controllable Neural Video Rendering
- Have We Scene It All? Scene Graph-Aware Deep Point Cloud Compression
- AI-Driven Radiology Report Generation for Traumatic Brain Injuries
- SummDiff: Generative Modeling of Video Summarization with Diffusion
- gLSTM: Mitigating Over-Squashing by Increasing Storage Capacity
- LTCA: Long-range Temporal Context Attention for Referring Video Object Segmentation
- Neuron-Level Analysis of Cultural Understanding in Large Language Models
- pyGinkgo: A Sparse Linear Algebra Operator Framework for Python
- Memory Retrieval and Consolidation in Large Language Models through Function Tokens
- RASALoRE: Region Aware Spatial Attention with Location-based Random Embeddings for Weakly Supervised Anomaly Detection in Brain MRI Scans
- AILoRA: Function-Aware Asymmetric Initialization for Low-Rank Adaptation of Large Language Models
- Do We Really Need Permutations? Impact of Width Expansion on Linear Mode Connectivity
- SHE: Stepwise Hybrid Examination Reinforcement Learning Framework for E-commerce Search Relevance
- From Defender to Devil? Unintended Risk Interactions Induced by LLM Defenses
- SatFusion: A Unified Framework for Enhancing Satellite IoT Images via Multi-Temporal and Multi-Source Data Fusion
- FlowLensing: Simulating Gravitational Lensing with Flow Matching
- Multilingual Generative Retrieval via Cross-lingual Semantic Compression
- IKNet: Interpretable Stock Price Prediction via Keyword-Guided Integration of News and Technical Indicators
- IntentionVLA: Generalizable and Efficient Embodied Intention Reasoning for Human-Robot Interaction
- Deep Neural Networks Inspired by Differential Equations
- From Noisy to Native: LLM-driven Graph Restoration for Test-Time Graph Domain Adaptation
- Parallel Test-Time Scaling for Latent Reasoning Models
- GeoGen: A Two-stage Coarse-to-Fine Framework for Fine-grained Synthetic Location-based Social Network Trajectory Generation
- Value Flows
- AutoQual: An LLM Agent for Automated Discovery of Interpretable Features for Review Quality Assessment
- Selection, Reflection and Self-Refinement: Revisit Reasoning Tasks via a Causal Lens
- Understanding the Geospatial Reasoning Capabilities of LLMs: A Trajectory Recovery Perspective
- Queries Are Not Alone: Clustering Text Embeddings for Video Search
- RayFusion: Ray Fusion Enhanced Collaborative Visual Perception
- New Machine Learning Approaches for Intrusion Detection in ADS-B
- MeSH: Memory-as-State-Highways for Recursive Transformers
- Vision-Enabled LLMs in Historical Lexicography: Digitising and Enriching Estonian-German Dictionaries from the 17th and 18th Centuries
- SPAD: Specialized Prefill and Decode Hardware for Disaggregated LLM Inference
- DGTEN: A Robust Deep Gaussian based Graph Neural Network for Dynamic Trust Evaluation with Uncertainty-Quantification Support
- An Adaptive Multi Agent Bitcoin Trading System
- Interacting cortico-basal ganglia-thalamocortical loops shape behavioral control through cognitive maps and shortcuts
- Learning What to Remember: Adaptive Probabilistic Memory Retention for Memory-Efficient Language Models
- Single layer tiny Co4 outpaces GPT-2 and GPT-BERT
- ZeroCard: Cardinality Estimation with Zero Dependence on Target Databases -- No Data, No Query, No Retraining
- Language Models Do Not Embed Numbers Continuously
- From Tokens to Layers: Redefining Stall-Free Scheduling for MoE Serving with Layered Prefill
- Long-Tailed Recognition via Information-Preservable Two-Stage Learning
- DACIP-RC: Domain Adaptive Continual Instruction Pre-Training via Reading Comprehension on Business Conversations
- Bridging the Physics-Data Gap with FNO-Guided Conditional Flow Matching: Designing Inductive Bias through Hierarchical Physical Constraints
- In-Context Clustering with Large Language Models
- Bidirectional Representations Augmented Autoregressive Biological Sequence Generation
- Post-Norm can Resharpen Attention
- CAT: Curvature-Adaptive Transformers for Geometry-Aware Learning
- Locality-Sensitive Hashing-Based Efficient Point Transformer for Charged Particle Reconstruction
- Knowledge-Aware Mamba for Joint Change Detection and Classification from MODIS Times Series
- Symbolic-Diffusion: Deep Learning Based Symbolic Regression with D3PM Discrete Token Diffusion
- Multi-Task Pre-Finetuning of Lightweight Transformer Encoders for Text Classification and NER
- Investigating Thematic Patterns and User Preferences in LLM Interactions using BERTopic
- Phase Diagram of Dropout for Two-Layer Neural Networks in the Mean-Field Regime
- MLLM4TS: Leveraging Vision and Multimodal Language Models for General Time-Series Analysis
- Populism Meets AI: Advancing Populism Research with LLMs
- Exploring Teachers' Perceptions of ChatGPT Through Prompt Engineering
- Lemma Dilemma: On Lemma Generation Without Domain- or Language-Specific Training Data
- Attention to Order: Transformers Discover Phase Transitions via Learnability
- Temporal Prompting Matters: Rethinking Referring Video Object Segmentation
- Artificial Hippocampus Networks for Efficient Long-Context Modeling
- Accelerating Inference for Multilayer Neural Networks with Quantum Computers
- Resolution scaling governs DINOv3 transfer performance in chest radiograph classification
- Biasless Language Models Learn Unnaturally: How LLMs Fail to Distinguish the Possible from the Impossible
- Exposing LLM User Privacy via Traffic Fingerprint Analysis: A Study of Privacy Risks in LLM Agent Interactions
- Encode, Think, Decode: Scaling test-time reasoning with recursive latent thoughts
- A Multi-Agent Framework for Stateful Inference-Time Search
- Generative World Modelling for Humanoids: 1X World Model Challenge Technical Report
- HTMformer: Hybrid Time and Multivariate Transformer for Time Series Forecasting
- Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications
- Search-R3: Unifying Reasoning and Embedding in Large Language Models
- Tool-Augmented Policy Optimization: Synergizing Reasoning and Adaptive Tool Use with Reinforcement Learning
- Diffusing Trajectory Optimization Problems for Recovery During Multi-Finger Manipulation
- Native Hybrid Attention for Efficient Sequence Modeling
- Textual interpretation of transient image classifications from large language models
- Label-frugal satellite image change detection with generative virtual exemplar learning
- A Denoising Diffusion-Based Evolutionary Algorithm Framework: Application to the Maximum Independent Set Problem
- Online Generic Event Boundary Detection
- Hexagon-MLIR: An AI Compilation Stack For Qualcomm's Neural Processing Units (NPUs)
- SoftMatcha 2: A Fast and Soft Pattern Matcher for Trillion-Scale Corpora
- Syncopate: Efficient Multi-GPU AI Kernels via Automatic Chunk-Centric Compute-Communication Overlap
- Investigating Industry--Academia Collaboration in Artificial Intelligence: PDF-Based Bibliometric Analysis from Leading Conferences
- TetriServe: Efficiently Serving Mixed DiT Workloads
- AI Foundation Model for Time Series with Innovations Representation
- Consistent Assistant Domains Transformer for Source-free Domain Adaptation
- Data augmentation in a triple transformer loop retrosynthesis model
- BOTANIC-0: a series of foundation models for plant genomic data
- Iterative design of a NAND hybrid riboswitch by deep batch Bayesian optimization
- Clinical-grade AI model for molecular subtyping of endometrial cancer: a multi-center cohort study in China
- Vectorized FlashAttention with Low-cost Exponential Computation in RISC-V Vector Processors
- Recurrence-Complete Frame-based Action Models
- DiFFPO: Training Diffusion LLMs to Reason Fast and Furious via Reinforcement Learning
- Pool Me Wisely: On the Effect of Pooling in Transformer-Based Models
- StyleKeeper: Prevent Content Leakage using Negative Visual Query Guidance
- Mid-Training of Large Language Models: A Survey
- Extreme Amodal Face Detection
- Study on LLMs for Promptagator-Style Dense Retriever Training
- Verifying Memoryless Sequential Decision-making of Large Language Models
- Reproducing and Extending Causal Insights Into Term Frequency Computation in Neural Rankers
- Is the Hard-Label Cryptanalytic Model Extraction Really Polynomial?
- Gaussian Equivalence for Self-Attention: Asymptotic Spectral Analysis of Attention Matrix
- TimeFormer: Transformer with Attention Modulation Empowered by Temporal Characteristics for Time Series Forecasting
- Heptapod: Language Modeling on Visual Signals
- The Effect of Attention Head Count on Transformer Approximation
- Rethinking Nonlinearity: Trainable Gaussian Mixture Modules for Modern Neural Architectures
- Distilling Lightweight Language Models for C/C++ Vulnerabilities
- A Comparative Analysis of Contextual Representation Flow in State-Space and Transformer Architectures
- RheOFormer: A generative transformer model for simulation of complex fluids and flows
- Reusing Overtrained Language Models Saturates Scaling
- Benchmarking Fake Voice Detection in the Fake Voice Generation Arms Race
- Associative Memory Model with Neural Networks: Memorizing multiple images with one neuron
- BACHI: Boundary-Aware Symbolic Chord Recognition Through Masked Iterative Decoding on Pop and Classical Music
- scPPDM: A Diffusion Model for Single-Cell Drug-Response Prediction
- Cross-Modal Attention Guided Unlearning in Vision-Language Models
- AsyncSpade: Efficient Test-Time Scaling with Asynchronous Sparse Decoding
- From Condensation to Rank Collapse: A Two-Stage Analysis of Transformer Training Dynamics
- GyroSwin: 5D Surrogates for Gyrokinetic Plasma Turbulence Simulations
- AudioMarathon: A Comprehensive Benchmark for Long-Context Audio Understanding and Efficiency in Audio LLMs
- metabeta -- A fast neural model for Bayesian mixed-effects regression
- DPA-Net: A Dual-Path Attention Neural Network for Inferring Glycemic Control Metrics from Self-Monitored Blood Glucose Data
- TGM: a Modular and Efficient Library for Machine Learning on Temporal Graphs
- OWL: Overcoming Window Length-Dependence in Speculative Decoding for Long-Context Inputs
- Evaluation of LLMs for Process Model Analysis and Optimization
- Cluster Paths: Navigating Interpretability in Neural Networks
- CLAQS: Compact Learnable All-Quantum Token Mixer with Shared-ansatz for Text Classification
- Grouped Differential Attention
- HEMERA: A Human-Explainable Transformer Model for Estimating Lung Cancer Risk using GWAS Data
- Text2Interact: High-Fidelity and Diverse Text-to-Two-Person Interaction Generation
- GUIDE: Guided Initialization and Distillation of Embeddings
- Mitigating Surgical Data Imbalance with Dual-Prediction Video Diffusion Model
- Relational Transformer: Toward Zero-Shot Foundation Models for Relational Data
- Flexible Swarm Learning May Outpace Foundation Models in Essential Tasks
- SpotDiff: Spotting and Disentangling Interference in Feature Space for Subject-Preserving Image Generation
- SDAR: A Synergistic Diffusion-AutoRegression Paradigm for Scalable Sequence Generation
- TabPFN-Wide: Continued Pre-Training for Extreme Feature Counts
- Discrete Diffusion Models with MLLMs for Unified Medical Multimodal Generation
- Distributional Semantics Tracing: A Framework for Explaining Hallucinations in Large Language Models
- Edit-Based Flow Matching for Temporal Point Processes
- Optimal Batched Scheduling of Stochastic Processing Networks Using Atomic Action Decomposition
- Shaken or Stirred? An Analysis of MetaFormer's Token Mixing for Medical Imaging
- The Anatomy of a Triton Attention Kernel
- Untangling Component Imbalance in Hybrid Linear Attention Conversion Methods
- Multimodal Trajectory Representation Learning for Travel Time Estimation
- Towards Robust and Realible Multimodal Misinformation Recognition with Incomplete Modality
- BlockGPT: Spatio-Temporal Modelling of Rainfall via Frame-Level Autoregression
- Toward a Safer Web: Multilingual Multi-Agent LLMs for Mitigating Adversarial Misinformation Attacks
- Oracle-Guided Masked Contrastive Reinforcement Learning for Visuomotor Policies
- vAttention: Verified Sparse Attention
- When and How to Cut Classical Concerts? A Multimodal Automated Video Editing Approach
- Efficient Conditional Generation on Scale-based Visual Autoregressive Models
- Traj-Transformer: Diffusion Models with Transformer for GPS Trajectory Generation
- When Does Global Attention Help? A Unified Empirical Study on Atomistic Graph Learning
- Mnemosyne: An Unsupervised, Human-Inspired Long-Term Memory Architecture for Edge-Based LLMs
- Activation-Informed Pareto-Guided Low-Rank Compression for Efficient LLM/VLM
- H1B-KV: Hybrid One-Bit Caches for Memory-Efficient Large Language Model Inference
- Human Action Recognition from Point Clouds over Time
- Ocular-Induced Abnormal Head Posture: Diagnosis and Missing Data Imputation
- Human3R: Everyone Everywhere All at Once
- InforME: Improving Informativeness of Abstractive Text Summarization With Informative Attention Guided by Named Entity Salience
- Refusal Falls off a Cliff: How Safety Alignment Fails in Reasoning?
- MASA: Rethinking the Representational Bottleneck in LoRA with Multi-A Shared Adaptation
- SD-MVSum: Script-Driven Multimodal Video Summarization Method and Datasets
- Life at the extremes: maximally divergent microbes with similar genomic signatures linked to extreme environments
- Latent Speech-Text Transformer
- AgeBooth: Controllable Facial Aging and Rejuvenation via Diffusion Models
- D2E: Scaling Vision-Action Pretraining on Desktop Data for Transfer to Embodied AI
- The New Quant: A Survey of Large Language Models in Financial Prediction and Trading
- Deep Generative Model for Human Mobility Behavior
- Permutation-Invariant Representation Learning for Robust and Privacy-Preserving Feature Selection
- RamPINN: Recovering Raman Spectra From Coherent Anti-Stokes Spectra Using Embedded Physics
- Patterns behind Chaos: Forecasting Data Movement for Efficient Large-Scale MoE LLM Inference
- Midway Network: Learning Representations for Recognition and Motion from Latent Dynamics
- Anchor: Reducing Temporal and Spatial Output Performance Variability on Quantum Computers
- A Hierarchical Geometry-guided Transformer for Histological Subtyping of Primary Liver Cancer
- QDeepGR4J: Quantile-based ensemble of deep learning and GR4J hybrid rainfall-runoff models for extreme flow prediction with uncertainty quantification
- UnitTenX: Generating Tests for Legacy Packages with AI Agents Powered by Formal Verification
- ResCP: Reservoir Conformal Prediction for Time Series Forecasting
- Decoding Partial Differential Equations: Cross-Modal Adaptation of Decoder-only Models to PDEs
- AUREXA-SE: Audio-Visual Unified Representation Exchange Architecture with Cross-Attention and Squeezeformer for Speech Enhancement
- Adjusting the Output of Decision Transformer with Action Gradient
- Attention-Enhanced Prototypical Learning for Few-Shot Infrastructure Defect Segmentation
- Rivaling Transformers: Multi-Scale Structured State-Space Mixtures for Agentic 6G O-RAN
- A Data-Driven Prism: Multi-View Source Separation with Diffusion Model Priors
- Finish First, Perfect Later: Test-Time Token-Level Cross-Validation for Diffusion Large Language Models
- Resource-Efficient Fine-Tuning of LLaMA-3.2-3B for Medical Chain-of-Thought Reasoning
- Aligning Perception, Reasoning, Modeling and Interaction: A Survey on Physical AI
- SSDD: Single-Step Diffusion Decoder for Efficient Image Tokenization
- Bidirectional Mammogram View Translation with Column-Aware and Implicit 3D Conditional Diffusion
- On Structured State-Space Duality
- Visual Representations inside the Language Model
- A Comparative Study of Vision Transformers and CNNs for Few-Shot Rigid Transformation and Fundamental Matrix Estimation
- Bridging Reasoning to Learning: Unmasking Illusions using Complexity Out of Distribution Generalization
- Noise or Signal? Deconstructing Contradictions and An Adaptive Remedy for Reversible Normalization in Time Series Forecasting
- IMLP: An Energy-Efficient Continual Learning Method for Tabular Data Streams
- SFANet: Spatial-Frequency Attention Network for Deepfake Detection
- Language Model Based Text-to-Audio Generation: Anti-Causally Aligned Collaborative Residual Transformers
- Poolformer: Recurrent Networks with Pooling for Long-Sequence Modeling
- Diffusion Transformers for Imputation: Statistical Efficiency and Uncertainty Quantification
- Improving Large-Scale Recommender Systems with Auxiliary Learning
- The Role of Feature Interactions in Graph-based Tabular Deep Learning
- 3Dify: a Framework for Procedural 3D-CG Generation Assisted by LLMs Using MCP and RAG
- Asynchronous Denoising Diffusion Models for Aligning Text-to-Image Generation
- Multi-Agent Collaborative Intelligence: Dual-Dial Control for Reliable LLM Reasoning
- TBStar-Edit: From Image Editing Pattern Shifting to Consistency Enhancement
- HRTFformer: A Spatially-Aware Transformer for Individual HRTF Upsampling in Immersive Audio Rendering
- Compressed Convolutional Attention: Efficient Attention in a Compressed Latent Space
- Partial Information Decomposition via Normalizing Flows in Latent Gaussian Distributions
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Watch and Learn: Learning to Use Computers from Online Videos
- Personalizing Retrieval using Joint Embeddings or "the Return of Fluffy"
- Residualized Similarity for Faithfully Explainable Authorship Verification
- MHA-RAG: Improving Efficiency, Accuracy, and Consistency by Encoding Exemplars as Soft Prompts
- SafeGuider: Robust and Practical Content Safety Control for Text-to-Image Models
- On the Limitations and Capabilities of Position Embeddings for Length Generalization
- Learning on the Job: Test-Time Curricula for Targeted Reinforcement Learning
- Progressive Gaussian Transformer with Anisotropy-aware Sampling for Open Vocabulary Occupancy Prediction
- Exploring the Power of Diffusion Large Language Models for Software Engineering: An Empirical Investigation
- Bridging Text and Video Generation: A Survey
- Benchmarking M-LTSF: Frequency and Noise-Based Evaluation of Multivariate Long Time Series Forecasting Models
- How Different from the Past? Spatio-Temporal Time Series Forecasting with Self-Supervised Deviation Learning
- Glocal Information Bottleneck for Time Series Imputation
- Hybrid Architectures for Language Models: Systematic Analysis and Design Insights
- Time Is Effort: Estimating Human Post-Editing Time for Grammar Error Correction Tool Evaluation
- Representation Potentials of Foundation Models for Multimodal Alignment: A Survey
- Adaptive Weighted Loss for Sequential Recommendations on Sparse Domains
- DoRAN: Stabilizing Weight-Decomposed Low-Rank Adaptation via Noise Injection and Auxiliary Networks
- FoilDiff: A Hybrid Transformer Backbone for Diffusion-based Modelling of 2D Airfoil Flow Fields
- HoRA: Cross-Head Low-Rank Adaptation with Joint Hypernetworks
- Scaling Sequence-to-Sequence Generative Neural Rendering
- Probing Whisper for Dysarthric Speech in Detection and Assessment
- Enhancing Talent Search Ranking with Role-Aware Expert Mixtures and LLM-based Fine-Grained Job Descriptions
- Detecting Semantic Clones of Unseen Functionality
- A Mathematical Explanation of Transformers for Large Language Models and GPTs
- Developing a Sequential Deep Learning Pipeline to Model Alaskan Permafrost Thaw Under Climate Change
- RLRF: Competitive Search Agent Design via Reinforcement Learning from Ranker Feedback
- Modeling Time Series Dynamics with Fourier Ordinary Differential Equations
- Evaluation of Clinical Trials Reporting Quality using Large Language Models
- Beyond Next-Token Prediction: A Performance Characterization of Diffusion versus Autoregressive Language Models
- Activation Steering with a Feedback Controller
- Wave-PDE Nets: Trainable Wave-Equation Layers as an Alternative to Attention
- Toward Uncertainty-Aware and Generalizable Neural Decoding for Quantum LDPC Codes
- Unmasking Backdoors: An Explainable Defense via Gradient-Attention Anomaly Scoring for Pre-trained Language Models
- Large Language Models Hallucination: A Comprehensive Survey
- Exact Causal Attention with 10% Fewer Operations
- MoME: Mixture of Matryoshka Experts for Audio-Visual Speech Recognition
- RACE Attention: A Strictly Linear-Time Attention Layer for Training on Outrageously Large Contexts
- An Early Exploration of Deep-Learning-Driven Prefetching for Far Memory
- From News to Returns: A Granger-Causal Hypergraph Transformer on the Sphere
- OpenTSLM: Time-Series Language Models for Reasoning over Multivariate Medical Text- and Time-Series Data
- Don't Pass@k: A Bayesian Framework for Large Language Model Evaluation
- Sliding Window Attention for Learned Video Compression
- Multimodal Learning with Augmentation Techniques for Natural Disaster Assessment
- Allocation of Parameters in Transformers
- A Benchmark Study of Deep Learning Methods for Multi-Label Pediatric Electrocardiogram-Based Cardiovascular Disease Classification
- Trajectory prediction for heterogeneous agents: A performance analysis on small and imbalanced datasets
- Adapting Diarization-Conditioned Whisper for End-to-End Multi-Talker Speech Recognition
- Person-Centric Annotations of LAION-400M: Auditing Bias and Its Transfer to Models
- Evolutionary Computation as Natural Generative AI
- Understanding the Role of Training Data in Test-Time Scaling
- On the Empirical Power of Goodness-of-Fit Tests in Watermark Detection
- Rainbow Padding: Mitigating Early Termination in Instruction-Tuned Diffusion LLMs
- Towards Sampling Data Structures for Tensor Products in Turnstile Streams
- XG-Attention-WGAN PIC: Utilizing XGboost-Attention-WGAN for Photonics Integrated Circuit Design
- Towards Unsupervised Speech Recognition at the Syllable-Level
- Efficient Test-Time Scaling for Small Vision-Language Models
- CrossLag: Predicting Major Dengue Outbreaks with a Domain Knowledge Informed Transformer
- A Qualitative Comparative Evaluation of Cognitive and Generative Theories
- Neural Correlates of Language Models Are Specific to Human Language
- Signature-Informed Transformer for Asset Allocation
- Cross-Modal Reconstruction Pretraining for Ramp Flow Prediction at Highway Interchanges
- What Drives Compositional Generalization in Visual Generative Models?
- Modeling Quantum Geometry for Fractional Chern Insulators with unsupervised learning
- Self-Reflective Generation at Test Time
- Improving Adversarial Robustness of Zero-Shot CLIP with Confidence-Aware Weighting
- Attack via Overfitting: 10-shot Benign Fine-tuning to Jailbreak LLMs
- Development of Deep Neural Network First-Level Hardware Track Trigger for the Belle II Experiment
- Bayesian Test-time Adaptation for Object Recognition and Detection with Vision-language Models
- Hyperparameter Loss Surfaces Are Simple Near their Optima
- Kolmogorov-Arnold Networks in Thermoelectric Materials Design
- HALO: Memory-Centric Heterogeneous Accelerator with 2.5D Integration for Low-Batch LLM Inference
- AutoMaAS: Self-Evolving Multi-Agent Architecture Search for Large Language Models
- Deep Generative Continual Learning using Functional LoRA: FunLoRA
- A Trajectory Generator for High-Density Traffic and Diverse Agent-Interaction Scenarios
- Longitudinal Flow Matching for Trajectory Modeling
- Visual Language Model as a Judge for Object Detection in Industrial Diagrams
- Coevolutionary Continuous Discrete Diffusion: Make Your Diffusion Language Model a Latent Reasoner
- Homophily-induced Emergence of Biased Structures in LLM-based Multi-Agent AI Systems
- Automatic Building Code Review: A Case Study
- Viability-Preserving Passive Torque Control
- Brain-Language Model Alignment: Insights into the Platonic Hypothesis and Intermediate-Layer Advantage
- Image Generation Based on Image Style Extraction
- CosmoUiT: A Vision Transformer-UNet Hybrid for Fast and Accurate Emulation of 21-cm Maps from the Epoch of Reionization
- Theory of Scaling Laws for In-Context Regression: Depth, Width, Context and Time
- Syntax-Guided Diffusion Language Models with User-Integrated Personalization
- TextCAM: Explaining Class Activation Map with Text
- Spiralformer: Low Latency Encoder for Streaming Speech Recognition with Circular Layer Skipping and Early Exiting
- Deep Learning-Based Approach for Improving Relational Aggregated Search
- A Neuro-Fuzzy System for Interpretable Long-Term Stock Market Forecasting
- GLAI: GreenLightningAI for Accelerated Training through Knowledge Decoupling
- Gather-Scatter Mamba: Accelerating Propagation with Efficient State Space Model
- Feature Identification for Hierarchical Contrastive Learning
- GeoGraph: Geometric and Graph-based Ensemble Descriptors for Intrinsically Disordered Proteins
- UniverSR: Unified and Versatile Audio Super-Resolution via Vocoder-Free Flow Matching
- LEAP: Local ECT-Based Learnable Positional Encodings for Graphs
- CroSTAta: Cross-State Transition Attention Transformer for Robotic Manipulation
- DEAP DIVE: Dataset Investigation with Vision transformers for EEG evaluation
- HAMLET: Switch your Vision-Language-Action Model into a History-Aware Policy
- Inclusive Easy-to-Read Generation for Individuals with Cognitive Impairments
- Erased, But Not Forgotten: Erased Rectified Flow Transformers Still Remain Unsafe Under Concept Attack
- Composer: A Search Framework for Hybrid Neural Architecture Design
- Linguistic Characteristics of AI-Generated Text: A Survey
- Understanding Sensitivity of Differential Attention through the Lens of Adversarial Robustness
- Affordance-Guided Diffusion Prior for 3D Hand Reconstruction
- Normal-Abnormal Guided Generalist Anomaly Detection
- Structuring Reasoning for Complex Rules Beyond Flat Representations
- TimeEmb: A Lightweight Static-Dynamic Disentanglement Framework for Time Series Forecasting
- LongCodeZip: Compress Long Context for Code Language Models
- TokMem: Tokenized Procedural Memory for Large Language Models
- Domain-Specialized Interactive Segmentation Framework for Meningioma Radiotherapy Planning
- Physics-Informed Neural Controlled Differential Equations for Scalable Long Horizon Multi-Agent Motion Forecasting
- SAGE-Music: Low-Latency Symbolic Music Generation via Attribute-Specialized Key-Value Head Sharing
- The Transformer Cookbook
- Continual Learning with Query-Only Attention
- FAME: Adaptive Functional Attention with Expert Routing for Function-on-Function Regression
- Generalized Parallel Scaling with Interdependent Generations
- Fast, Secure, and High-Capacity Image Watermarking with Autoencoded Text Vectors
- Fast frequency reconstruction using Deep Learning for event recognition in ring laser data
- Efficient Multi-modal Large Language Models via Progressive Consistency Distillation
- Multi-Category Materials Information Extraction (Composition, Processing, Microstructure, Properties)
- A Scene is Worth a Thousand Features: Feed-Forward Camera Localization from a Collection of Image Features
- Local Linear Attention: An Optimal Interpolation of Linear and Softmax Attention For Test-Time Regression
- SAGE-LD: Towards Scalable and Generalizable End-to-End Language Diarization via Simulated Data Augmentation
- AttentionDep: Domain-Aware Attention for Explainable Depression Severity Assessment
- Tenyidie Syllabification corpus creation and deep learning applications
- Flow Autoencoders are Effective Protein Tokenizers
- Cutting the Skip: Training Residual-Free Transformers
- A phase-aware AI car-following model for electric vehicles with adaptive cruise control: Development and validation using real-world data
- Post-Training Quantization for Audio Diffusion Transformers
- Looking Beyond the Known: Towards a Data Discovery Guided Open-World Object Detection
- Free Draft-and-Verification: Toward Lossless Parallel Decoding for Diffusion Large Language Models
- Data driven approaches in nanophotonics: A review of AI-enabled metadevices
- Delayed Attention Training Improves Length Generalization in Transformer--RNN Hybrids
- The Pitfalls of KV Cache Compression
- Thoughtbubbles: an Unsupervised Method for Parallel Thinking in Latent Space
- Quantifying Semantic Shift in Financial NLP: Robust Metrics for Market Prediction Stability
- Train Large, Deploy Compact: Structured Compression for Compact Low-Rank Adaptation
- Object-Centric Case-Based Reasoning via Argumentation
- Atlas-free Brain Network Transformer
- Large Language Models Inference Engines based on Spiking Neural Networks
- Learning Generalizable Shape Completion with SIM(3) Equivariance
- Attention as a Compass: Efficient Exploration for Process-Supervised RL in Reasoning Models
- DA2: Depth Anything in Any Direction
- Video Object Segmentation-Aware Audio Generation
- DiffCamera: Arbitrary Refocusing on Images
- Probing the Critical Point (CritPt) of AI Reasoning: a Frontier Physics Research Benchmark
- TASP: Topology-aware Sequence Parallelism
- Rearchitecting Datacenter Lifecycle for AI: A TCO-Driven Framework
- Contrastive Diffusion Guidance for Spatial Inverse Problems
- Combining Knowledge Graphs and NLP to Analyze Instant Messaging Data in Criminal Investigations
- TVS Sidekick: Challenges and Practical Insights from Deploying Large Language Models in the Enterprise
- IR-UWB Radar-Based Contactless Silent Speech Recognition with Attention-Enhanced Temporal Convolutional Networks
- TrackCore-F: Deploying Transformer-Based Subatomic Particle Tracking on FPGAs
- Latent Thinking Optimization: Your Latent Reasoning Language Model Secretly Encodes Reward Signals in Its Latent Thoughts
- SlimPack: Fine-Grained Asymmetric Packing for Balanced and Efficient Variable-Length LLM Training
- Event Tokenization and Next-Token Prediction for Anomaly Detection at the Large Hadron Collider
- The silence of the weights: an investigation of structural pruning strategies for attention-based audio signal architectures
- VietBinoculars: A Zero-Shot Approach for Detecting Vietnamese LLM-Generated Text
- EntroPE: Entropy-Guided Dynamic Patch Encoder for Time Series Forecasting
- Accelerating Transformers in Online RL
- AiDE-Q: Synthetic Labeled Datasets Can Enhance Learning Models for Quantum Property Estimation
- Evaluating the Use of Large Language Models as Synthetic Social Agents in Social Science Research
- RE2: Improving Chinese Grammatical Error Correction via Retrieving Appropriate Examples with Explanation
- Indirect Attention: Turning Context Misalignment into a Feature
- Using GPT to build a Project Management assistant for Jira environments
- S3E: Self-Supervised State Estimation for Radar-Inertial System
- Neural Network State-Space Estimators
- Towards Intuitive Human-Robot Interaction through Embodied Gesture-Driven Control with Woven Tactile Skins
- Accelerating LLM Inference with Precomputed Query Storage
- Bringing Emerging Architectures to Sequence Labeling in NLP
- Understanding the Mixture-of-Experts with Nadaraya-Watson Kernel
- PAT: Pattern-Perceptive Transformer for Error Detection in Relational Databases
- Revoking Amnesia: RL-based Trajectory Optimization to Resurrect Erased Concepts in Diffusion Models
- HiStyle: Hierarchical Style Embedding Predictor for Text-Prompt-Guided Controllable Speech Synthesis
- Kairos: Towards Adaptive and Generalizable Time Series Foundation Models
- Adapting SAM with Dynamic Similarity Graphs for Few-Shot Parameter-Efficient Small Dense Object Detection: A Case Study of Chickpea Pods in Field Conditions
- Better with Less: Small Proprietary Models Surpass Large Language Models in Financial Transaction Understanding
- Sharpness of Minima in Deep Matrix Factorization: Exact Expressions
- Autonomy-Aware Clustering: When Local Decisions Supersede Global Prescriptions
- Galton's Law of Mediocrity: Why Large Language Models Regress to the Mean and Fail at Creativity in Advertising
- Chinese vs. World Bank Development Projects: Insights from Earth Observation and Computer Vision on Wealth Gains in Africa, 2002-2013
- Deep set based operator learning with uncertainty quantification
- LieHMR: Autoregressive Human Mesh Recovery with SO(3) Diffusion
- Transformer-Based Neural Networks Backflow for Strongly Correlated Electronic Structure
- Guiding Mixture-of-Experts with Temporal Multimodal Interactions
- TraceDet: Hallucination Detection from the Decoding Trace of Diffusion Large Language Models
- DescribeEarth: Describe Anything for Remote Sensing Images
- Iterative Residual Cross-Attention Mechanism: An Integrated Approach for Audio-Visual Navigation Tasks
- Unsupervised Detection of Spatiotemporal Anomalies in PMU Data Using Transformer-Based BiGAN
- Transformers through the lens of support-preserving maps between measures
- Effective Model Pruning
- Leveraging Scene Context with Dual Networks for Sequential User Behavior Modeling
- Submodular Context Partitioning and Compression for In-Context Learning
- Landmark-Guided Knowledge for Vision-and-Language Navigation
- Hierarchical Reasoning Models: Perspectives and Misconceptions
- Interpret, prune and distill Donut : towards lightweight VLMs for VQA on document
- Thinking Sparks!: Emergent Attention Heads in Reasoning Models During Post Training
- TTT3R: 3D Reconstruction as Test-Time Training
- SoREX: Towards Self-Explainable Social Recommendation with Relevant Ego-Path Extraction
- SafePassage: High-Fidelity Information Extraction with Black Box LLMs
- MoReFlow: Motion Retargeting Learning through Unsupervised Flow Matching
- Causal Autoencoder-like Generation of Feedback Fuzzy Cognitive Maps with an LLM Agent
- AttentionViG: Cross-Attention-Based Dynamic Neighbor Aggregation in Vision GNNs
- Scalable Boltzmann Generators for equilibrium sampling of large-scale materials
- TDHook: A Lightweight Framework for Interpretability
- Joint Embeddings Go Temporal
- Boolean Satisfiability via Imitation Learning
- A Cartography of Open Collaboration in Open Source AI: Mapping Practices, Motivations, and Governance in 14 Open Large Language Model Projects
- Spontaneous High-Order Generalization in Neural Theory-of-Mind Networks
- Guided Diffusion for the Discovery of New Superconductors
- VideoAnchor: Reinforcing Subspace-Structured Visual Cues for Coherent Visual-Spatial Reasoning
- Scaling with Collapse: Efficient and Predictable Training of LLM Families
- GRACE-MoE: Grouping and Replication with Locality-Aware Routing for Efficient Distributed MoE Inference
- A multiscale analysis of mean-field transformers in the moderate interaction regime
- LVT: Large-Scale Scene Reconstruction via Local View Transformers
- Path Diffuser: Diffusion Model for Data-Driven Traffic Simulator
- Learning Relationships Between Separate Audio Tracks for Creative Applications
- On-the-Fly Data Augmentation for Brain Tumor Segmentation
- VIVALDy: A Hybrid Generative Reduced-Order Model for Turbulent Flows, Applied to Vortex-Induced Vibrations
- When Autonomous Vehicle Meets V2X Cooperative Perception: How Far Are We?
- Inductive Bias and Spectral Properties of Single-Head Attention in High Dimensions
- Attention Surgery: An Efficient Recipe to Linearize Your Video Diffusion Transformer
- Accurate Cobb Angle Estimation via SVD-Based Curve Detection and Vertebral Wedging Quantification
- Identifying Information-Transfer Nodes in a Recurrent Neural Network Reveals Dynamic Representations
- Room Impulse Response Prediction with Neural Networks: From Energy Decay Curves to Perceptual Validation
- SemShareKV: Efficient KVCache Sharing for Semantically Similar Prompts via Token-Level LSH Matching
- DSAT-HD: Dual-Stream Adaptive Transformer with Hybrid Decomposition for Multivariate Time Series Forecasting
- ClustRecNet: A Novel End-to-End Deep Learning Framework for Clustering Algorithm Recommendation
- Vision Function Layer in Multimodal LLMs
- Fidel-TS: A High-Fidelity Benchmark for Multimodal Time Series Forecasting
- Deep Learning-Based Prediction of Energy Decay Curves from Room Geometry and Material Properties
- From Ambiguity to Verdict: A Semiotic-Grounded Multi-Perspective Agent for LLM Logical Reasoning
- In-Context Learning of Temporal Point Processes with Foundation Inference Models
- Spatial-Functional awareness Transformer-based graph archetype contrastive learning for Decoding Visual Neural Representations from EEG
- ProxyAttn: Guided Sparse Attention via Representative Heads
- Bundle Network: a Machine Learning-Based Bundle Method
- Unsupervised Machine Learning for Anomaly Detection in LHC Collider Searches
- Can you SPLICE it together? A Human Curated Benchmark for Probing Visual Reasoning in VLMs
- Word-Level Emotional Expression Control in Zero-Shot Text-to-Speech Synthesis
- HiKE: Hierarchical Evaluation Framework for Korean-English Code-Switching Speech Recognition
- U-DiT Policy: U-shaped Diffusion Transformers for Robotic Manipulation
- EOE: Evolutionary Optimization of Experts for Training Language Models
- BiHDTrans: binary hyperdimensional transformer for efficient multivariate time series classification
- Multi-Item-Query Attention for Stable Sequential Recommendation
- CLQ: Cross-Layer Guided Orthogonal-based Quantization for Diffusion Transformers
- UniFlow-Audio: Unified Flow Matching for Audio Generation from Omni-Modalities
- Vid-LLM: A Compact Video-based 3D Multimodal LLM with Reconstruction-Reasoning Synergy
- HarmMetric Eval: Benchmarking Metrics and Judges for LLM Harmfulness Assessment
- RServe: Overlapping Encoding and Prefill for Efficient LMM Inference
- Agentic Services Computing
- DINOReg: Strong Point Cloud Registration with Vision Foundation Model
- United We Stand: Towards End-to-End Log-based Fault Diagnosis via Interactive Multi-Task Learning
- An Enhanced Pyramid Feature Network Based on Long-Range Dependencies for Multi-Organ Medical Image Segmentation
- Walk the Talk: Is Your Log-based Software Reliability Maintenance System Really Reliable?
- Comparing Open-Source and Commercial LLMs for Domain-Specific Analysis and Reporting: Software Engineering Challenges and Design Trade-offs
- Hyperspherical Latents Improve Continuous-Token Autoregressive Generation
- Training Dynamics of Parametric and In-Context Knowledge Utilization in Language Models
- PEARL: Performance-Enhanced Aggregated Representation Learning
- SVGThinker: Instruction-Aligned and Reasoning-Driven Text-to-SVG Generation
- VeriLLM: A Lightweight Framework for Publicly Verifiable Decentralized Inference
- Graph Optimization Foundation Model: Tokenizing Graph via A Language-Model Paradigm
- MAESTRO : Adaptive Sparse Attention and Robust Learning for Multimodal Dynamic Time Series
- Tumor Synthesis conditioned on Radiomics
- BladderFormer: A Streaming Transformer for Real-Time Urological State Monitoring
- ASTROCO: Self-Supervised Conformer-Style Transformers for Light-Curve Embeddings
- InfLLM-V2: Dense-Sparse Switchable Attention for Seamless Short-to-Long Adaptation
- DyMoDreamer: World Modeling with Dynamic Modulation
- Emergent World Representations in OpenVLA
- BALF: Budgeted Activation-Aware Low-Rank Factorization for Fine-Tuning-Free Model Compression
- PanoWorld-X: Generating Explorable Panoramic Worlds via Sphere-Aware Video Diffusion
- AGNOMIN -- Architecture Agnostic Multi-Label Function Name Prediction
- Ultra-Fast Language Generation via Discrete Diffusion Divergence Instruct
- Harnessing Deep Learning in Searching Wild Relatives of Domestic Animals
- International Intellectual Property Law in the Age of AI
- Cross-species gene redesign leveraging ortholog information and generative modeling
- Deep learning-based morphological classification of ceramics: A case study of 3D point cloud analysis for Sue ware, Japan
- Muon: Training and Trade-offs with Latent Attention and MoE
- Multi-modal Spatio-Temporal Transformer for High-resolution Land Subsidence Prediction
- CURA: Size Isnt All You Need -- A Compact Universal Architecture for On-Device Intelligence
- SCOPE: Semantic Conditioning for Sim2Real Category-Level Object Pose Estimation in Robotics
- Training Agents Inside of Scalable World Models
- From Score Distributions to Balance: Plug-and-Play Mixture-of-Experts Routing
- The Rise of AfricaNLP: A Survey of Contributions, Contributors, Community Impact, and Bibliometric Analysis
- Pretraining with hierarchical memories: separating long-tail and common knowledge
- Code-switching Speech Recognition Under the Lens: Model- and Data-Centric Perspectives
- Optimizing Privacy-Preserving Primitives to Support LLM-Scale Applications
- Learning to Parallel: Accelerating Diffusion Large Language Models via Learnable Parallel Decoding
- Brain Harmony: A Multimodal Foundation Model Unifying Morphology and Function into 1D Tokens
- An Agent-Based Framework for Automated Higher-Voice Harmony Generation
- One-Prompt Strikes Back: Sparse Mixture of Experts for Prompt-based Continual Learning
- Echo Flow Networks
- GeoFunFlow: Geometric Function Flow Matching for Inverse Operator Learning over Complex Geometries
- ADAPT: Lightweight, Long-Range Machine Learning Force Fields Without Graphs
- ResFormer: All-Time Reservoir Memory for Long Sequence Classification
- Discovering Transformer Circuits via a Hybrid Attribution and Pruning Framework
- SLA: Beyond Sparsity in Diffusion Transformers via Fine-Tunable Sparse-Linear Attention
- TREAT-Net: Tabular-Referenced Echocardiography Analysis for Acute Coronary Syndrome Treatment Prediction
- DNABERT-2: Fine-Tuning a Genomic Language Model for Colorectal Gene Enhancer Classification
- AutoPrune: Each Complexity Deserves a Pruning Policy
- Graph Mixing Additive Networks
- Token Painter: Training-Free Text-Guided Image Inpainting via Mask Autoregressive Models
- From Neural Networks to Logical Theories: The Correspondence between Fibring Modal Logics and Fibring Neural Networks
- Q-FSRU: Quantum-Augmented Frequency-Spectral For Medical Visual Question Answering
- Differentiable Sparsity via D-Gating: Simple and Versatile Structured Penalization
- Disentangling Score Content and Performance Style for Joint Piano Rendering and Transcription
- Sim-DETR: Unlock DETR for Temporal Sentence Grounding
- AgentGuard: Runtime Verification of AI Agents
- Adversarial Diffusion for Robust Reinforcement Learning
- LORT: Locally Refined Convolution and Taylor Transformer for Monaural Speech Enhancement
- Influence-Guided Concolic Testing of Transformer Robustness
- Sequence Pathfinder for Multi-Agent Pickup and Delivery in the Warehouse
- NeuSO: Neural Optimizer for Subgraph Queries
- Texture Vector-Quantization and Reconstruction Aware Prediction for Generative Super-Resolution
- A Modality-Tailored Graph Modeling Framework for Urban Region Representation via Contrastive Learning
- Transparent Visual Reasoning via Object-Centric Agent Collaboration
- Time-Shifted Token Scheduling for Symbolic Music Generation
- LocoFormer: Generalist Locomotion via Long-context Adaptation
- HieraTok: Multi-Scale Visual Tokenizer Improves Image Reconstruction and Generation
- FastViDAR: Real-Time Omnidirectional Depth Estimation via Alternative Hierarchical Attention
- FraudTransformer: Time-Aware GPT for Transaction Fraud Detection
- StrucADT: Generating Structure-controlled 3D Point Clouds with Adjacency Diffusion Transformer
- Graph Neural Networks with Diversity-aware Neighbor Selection and Dynamic Multi-scale Fusion for Multivariate Time Series Forecasting
- Virtual Nodes based Heterogeneous Graph Convolutional Neural Network for Efficient Long-Range Information Aggregation
- MotionVerse: A Unified Multimodal Framework for Motion Comprehension, Generation and Editing
- Emergent Slow Thinking in LLMs as Inverse Tree Freezing
- DiffInk: Glyph- and Style-Aware Latent Diffusion Transformer for Text to Online Handwriting Generation
- Generalizable Speech Deepfake Detection via Information Bottleneck Enhanced Adversarial Alignment
- BioVessel-Net and RetinaMix: Unsupervised Retinal Vessel Segmentation from OCTA Images
- Efficient Audio-Visual Speech Separation with Discrete Lip Semantics and Multi-Scale Global-Local Attention
- Improving the Efficiency of LLM Agent Systems through Trajectory Reduction
- Channel, Trend and Periodic-Wise Representation Learning for Multivariate Long-term Time Series Forecasting
- Automatic Speech Recognition for Greek Medical Dictation
- End-to-End Deep Learning for Predicting Metric Space-Valued Outputs
- A Computational Perspective on NeuroAI and Synthetic Biological Intelligence
- HunyuanImage 3.0 Technical Report
- From What to Why: A Multi-Agent System for Evidence-based Chemical Reaction Condition Reasoning
- MemMamba: Rethinking Memory Patterns in State Space Model
- QuadEnhancer: Leveraging Quadratic Transformations to Enhance Deep Neural Networks
- Artificial Intelligence-Powered Assessment Framework for Skill-Oriented Engineering Lab Education
- An Senegalese Legal Texts Structuration Using LLM-augmented Knowledge Graph
- ViTSP: A Vision Language Models Guided Framework for Large-Scale Traveling Salesman Problems
- LOTFormer: Doubly-Stochastic Linear Attention via Low-Rank Optimal Transport
- FracDetNet: Advanced Fracture Detection via Dual-Focus Attention and Multi-scale Calibration in Medical X-ray Imaging
- Mind the Links: Cross-Layer Attention for Link Prediction in Multiplex Networks
- AI-Based Stroke Rehabilitation Domiciliary Assessment System with STGCN Attention
- Generative Modeling of Shape-Dependent Self-Contact Human Poses
- A Predictive and Synergistic Two-Layer Scheduling Framework for LLM Serving
- MimiTalk: Revolutionizing Qualitative Research with Dual-Agent AI
- UniPose: Unified Cross-modality Pose Prior Propagation towards RGB-D data for Weakly Supervised 3D Human Pose Estimation
- Language, Culture, and Ideology: Personalizing Offensiveness Detection in Political Tweets with Reasoning LLMs
- SDQ-LLM: Sigma-Delta Quantization for 1-bit LLMs of any size
- Seeing the Unseen in Low-light Spike Streams
- TimeExpert: Boosting Long Time Series Forecasting with Temporal Mix of Experts
- Adaptive Token-Weighted Differential Privacy for LLMs: Not All Tokens Require Equal Protection
- Fact Grounded Attention: Eliminating Hallucination in Large Language Models Through Attention Level Knowledge Integration
- One-Shot Multi-Label Causal Discovery in High-Dimensional Event Sequences
- PDE-Transformer: A Continuous Dynamical Systems Approach to Sequence Modeling
- Understanding and Enhancing the Planning Capability of Language Models via Multi-Token Prediction
- Limit Analysis for Symbolic Multi-step Reasoning Tasks with Information Propagation Rules Based on Transformers
- WARBERT: A Hierarchical BERT-based Model for Web API Recommendation
- Stochastic Interpolants via Conditional Dependent Coupling
- RHYTHM: Reasoning with Hierarchical Temporal Tokenization for Human Mobility
- AttAnchor: Guiding Cross-Modal Token Alignment in VLMs with Attention Anchors
- How to Make Large Language Models Generate 100% Valid Molecules?
- Follow-Your-Preference: Towards Preference-Aligned Image Inpainting
- Cross-Domain Multi-Person Human Activity Recognition via Near-Field Wi-Fi Sensing
- CLAD-Net: Continual Activity Recognition in Multi-Sensor Wearable Systems
- FMC-DETR: Frequency-Decoupled Multi-Domain Coordination for Aerial-View Object Detection
- ABConformer: Physics-inspired Sliding Attention for Antibody-Antigen Interface Prediction
- URS: A Unified Neural Routing Solver for Cross-Problem Zero-Shot Generalization
- Quant-dLLM: Post-Training Extreme Low-Bit Quantization for Diffusion Large Language Models
- From Noise to Laws: Regularized Time-Series Forecasting via Denoised Dynamic Graphs
- Memory-Efficient Fine-Tuning via Low-Rank Activation Compression
- Vid-Freeze: Protecting Images from Malicious Image-to-Video Generation via Temporal Freezing
- Revisiting Multivariate Time Series Forecasting with Missing Values
- AudioFuse: Unified Spectral-Temporal Learning via a Hybrid ViT-1D CNN Architecture for Robust Phonocardiogram Classification
- Impute-MACFM: Imputation based on Mask-Aware Flow Matching
- Steering Prepositional Phrases in Language Models: A Case of with-headed Adjectival and Adverbial Complements in Gemma-2
- Functional Critic Modeling for Provably Convergent Off-Policy Actor-Critic
- Reinforcement Learning with Discrete Diffusion Policies for Combinatorial Action Spaces
- Convolutional Set Transformer
- Learning to Detect Relevant Contexts and Knowledge for Response Selection in Retrieval-based Dialogue Systems
- Introducing Multimodal Paradigm for Learning Sleep Staging PSG via General-Purpose Model
- What Do They Fix? LLM-Aided Categorization of Security Patches for Critical Memory Bugs
- When Can AI Models Explain Learning? Validity Criteria for AI as Cognitive Models in Education
- IONext: Unlocking the Next Era of Inertial Odometry
- Transformers Can Learn Connectivity in Some Graphs but Not Others
- Orochi: Versatile Biomedical Image Processor
- The Philosophy of Language Models
- CCNeXt: An Effective Self-Supervised Stereo Depth Estimation Approach
- LLM-Augmented and Fair Machine Learning Framework for University Admission Prediction
- JointDiff: Bridging Continuous and Discrete in Multi-Agent Trajectory Generation
- Data-Driven Temperature Modelling of Machine Tools by Neural Networks: A Benchmark
- Learning the Neighborhood: Contrast-Free Multimodal Self-Supervised Molecular Graph Pretraining
- IIET: Efficient Numerical Transformer via Implicit Iterative Euler Method
- Physics-informed GNN for medium-high voltage AC power flow with edge-aware attention and line search correction operator
- U-MAN: U-Net with Multi-scale Adaptive KAN Network for Medical Image Segmentation
- Machines in the Margins: A Systematic Review of Automated Content Generation for Wikipedia
- Partial Parameter Updates for Efficient Distributed Training
- Explaining multimodal LLMs via intra-modal token interactions
- A model of errors in transformers
- GPT-4 for Occlusion Order Recovery
- Zero-Effort Image-to-Music Generation: An Interpretable RAG-based VLM Approach
- HierLight-YOLO: A Hierarchical and Lightweight Object Detection Network for UAV Photography
- Stochastic activations
- RAPID3: Tri-Level Reinforced Acceleration Policies for Diffusion Transformer
- Distributed Associative Memory via Online Convex Optimization
- Adaptive Policy Backbone via Shared Network
- Johnson-Lindenstrauss Lemma Guided Network for Efficient 3D Medical Segmentation
- SoDaDE: Solvent Data-Driven Embeddings with Small Transformer Models
- Aurora: Towards Universal Generative Multimodal Time Series Forecasting
- Structured Sparse Transition Matrices to Enable State Tracking in State-Space Models
- Red Teaming Quantum-Resistant Cryptographic Standards: A Penetration Testing Framework Integrating AI and Quantum Security
- DHAGrasp: Synthesizing Affordance-Aware Dual-Hand Grasps with Text Instructions
- Mixture of Detectors: A Compact View of Machine-Generated Text Detection
- Multi-Agent Path Finding via Offline RL and LLM Collaboration
- FoodSEM: Large Language Model Specialized in Food Named-Entity Linking
- Multilingual Vision-Language Models, A Survey
- Mind the Missing: Variable-Aware Representation Learning for Irregular EHR Time Series using Large Language Models
- Does Generative Retrieval Overcome the Limitations of Dense Retrieval?
- SK2Decompile: LLM-based Two-Phase Binary Decompilation from Skeleton to Skin
- Self-driving cars: Are we there yet?
- Factor-Based Conditional Diffusion Model for Portfolio Optimization
- COSPADI: Compressing LLMs via Calibration-Guided Sparse Dictionary Learning
- The Rogue Scalpel: Activation Steering Compromises LLM Safety
- High-Quality Sound Separation Across Diverse Categories via Visually-Guided Generative Modeling
- Teaching Transformers to Solve Combinatorial Problems through Efficient Trial & Error
- Task-Adaptive Parameter-Efficient Fine-Tuning for Weather Foundation Models
- FlowDrive: moderated flow matching with data balancing for trajectory planning
- Statistical Advantage of Softmax Attention: Insights from Single-Location Regression
- Extracting Actionable Insights from Building Energy Data using Vision LLMs on Wavelet and 3D Recurrence Representations
- DynaNav: Dynamic Feature and Layer Selection for Efficient Visual Navigation
- SAGE: Scene Graph-Aware Guidance and Execution for Long-Horizon Manipulation Tasks
- SingRef6D: Monocular Novel Object Pose Estimation with a Single RGB Reference
- Taming Flow-based I2V Models for Creative Video Editing
- A High-Capacity and Secure Disambiguation Algorithm for Neural Linguistic Steganography
- Elastic MoE: Unlocking the Inference-Time Scalability of Mixture-of-Experts
- StableDub: Taming Diffusion Prior for Generalized and Efficient Visual Dubbing
- FastEnhancer: Speed-Optimized Streaming Neural Speech Enhancement
- Dynamic Novel View Synthesis in High Dynamic Range
- Zeppelin: Balancing Variable-length Workloads in Data Parallel Large Model Training
- SoK: Potentials and Challenges of Large Language Models for Reverse Engineering
- FastGRPO: Accelerating Policy Optimization via Concurrency-aware Speculative Decoding and Online Draft Learning
- Prompt-guided Disentangled Representation for Action Recognition
- SynerGen: Contextualized Generative Recommender for Unified Search and Recommendation
- Backdoor Attribution: Elucidating and Controlling Backdoor in Language Models
- Rethinking RoPE Scaling in Quantized LLM: Theory, Outlier, and Channel-Band Analysis with Weight Rescaling
- A Unifying Framework for Parallelizing Sequential Models with Linear Dynamical Systems
- Bridging Kolmogorov Complexity and Deep Learning: Asymptotically Optimal Description Length Objectives for Transformers
- An LLM-Powered Agent for Real-Time Analysis of the Vietnamese IT Job Market
- Triple-BERT: Do We Really Need MARL for Order Dispatch on Ride-Sharing Platforms?
- AI Brown and AI Koditex: LLM-Generated Corpora Comparable to Traditional Corpora of English and Czech Texts
- PartSAM: A Scalable Promptable Part Segmentation Model Trained on Native 3D Data
- Large Material Gaussian Model for Relightable 3D Generation
- Rethinking Large Language Model Distillation: A Constrained Markov Decision Process Perspective
- Compute-Optimal Quantization-Aware Training
- RLP: Reinforcement as a Pretraining Objective
- The GDN-CC Dataset: Automatic Corpus Clarification for AI-enhanced Democratic Citizen Consultations
- Towards Transparent AI: A Survey on Explainable Language Models
- AI for Sustainable Future Foods
- DroneFL: Federated Learning for Multi-UAV Visual Target Tracking
- Contrastive Mutual Information Learning: Toward Robust Representations without Positive-Pair Augmentations
- New Algorithmic Directions in Optimal Transport and Applications for Product Spaces
- SlimDiff: Training-Free, Activation-Guided Hands-free Slimming of Diffusion Models
- GraphPFN: A Prior-Data Fitted Graph Foundation Model
- Dual-Head Reasoning Distillation: Improving Classifier Accuracy with Train-Time-Only Reasoning
- ARTI-6: Towards Six-dimensional Articulatory Speech Encoding
- Forecasting Seismic Waveforms: A Deep Learning Approach for Einstein Telescope
- Nova: Real-Time Agentic Vision-Language Model Serving with Adaptive Cross-Stage Parallelization
- No Prior, No Leakage: Revisiting Reconstruction Attacks in Trained Neural Networks
- Bounds of Chain-of-Thought Robustness: Reasoning Steps, Embed Norms, and Beyond
- SuperOffload: Unleashing the Power of Large-Scale LLM Training on Superchips
- A Causality-Aware Spatiotemporal Model for Multi-Region and Multi-Pollutant Air Quality Forecasting
- Semantic Edge-Cloud Communication for Real-Time Urban Traffic Surveillance with ViT and LLMs over Mobile Networks
- RetoVLA: Reusing Register Tokens for Spatial Reasoning in Vision-Language-Action Models
- Explaining Fine Tuned LLMs via Counterfactuals A Knowledge Graph Driven Framework
- SlideMamba: Entropy-Based Adaptive Fusion of GNN and Mamba for Enhanced Representation Learning in Digital Pathology
- Go With The Flow: Churn-Tolerant Decentralized Training of Large Language Models
- Differential-Integral Neural Operator for Long-Term Turbulence Forecasting
- Towards Foundation Models for Zero-Shot Time Series Anomaly Detection: Leveraging Synthetic Data and Relative Context Discrepancy
- EvoMail: Self-Evolving Cognitive Agents for Adaptive Spam and Phishing Email Defense
- Fine-tuning of Large Language Models for Domain-Specific Cybersecurity Knowledge
- Reinforcement Learning Fine-Tuning Enhances Activation Intensity and Diversity in the Internal Circuitry of LLMs
- LayerNorm Induces Recency Bias in Transformer Decoders
- A Formal Comparison Between Chain-of-Thought and Latent Thought
- PhenoMoler: Phenotype-Guided Molecular Optimization via Chemistry Large Language Model
- Decoding the Surgical Scene: A Scoping Review of Scene Graphs in Surgery
- Autoregressive End-to-End Planning with Time-Invariant Spatial Alignment and Multi-Objective Policy Refinement
- Conditionally Whitened Generative Models for Probabilistic Time Series Forecasting
- AIBA: Attention-based Instrument Band Alignment for Text-to-Audio Diffusion
- ArchGPT: Understanding the World's Architectures with Large Multimodal Models
- FHRFormer: A Self-supervised Transformer Approach for Fetal Heart Rate Inpainting and Forecasting
- Explaining Grokking and Information Bottleneck through Neural Collapse Emergence
- Measuring LLM Sensitivity in Transformer-based Tabular Data Synthesis
- IoT-MCP: Bridging LLMs and IoT Systems Through Model Context Protocol
- Decoupled-Value Attention for Prior-Data Fitted Networks: GP Inference for Physical Equations
- TyphoonMLA: A Mixed Naive-Absorb MLA Kernel For Shared Prefix
- Why Attention Fails: The Degeneration of Transformers into MLPs in Time Series Forecasting
- Fast-SEnSeI: Lightweight Sensor-Independent Cloud Masking for On-board Multispectral Sensors
- Mixture of Thoughts: Learning to Aggregate What Experts Think, Not Just What They Say
- Concepts in Motion: Temporal Concept Bottleneck Model for Interpretable Video Classification
- ArtUV: Artist-style UV Unwrapping
- GEP: A GCG-Based method for extracting personally identifiable information from chatbots built on small language models
- ARMesh: Autoregressive Mesh Generation via Next-Level-of-Detail Prediction
- MolCluster: Integrating Graph Neural Network with Community Detection for Coarse-Grained Mapping
- TasselNetV4: A vision foundation model for cross-scene, cross-scale, and cross-species plant counting
- SimDiff: Simulator-constrained Diffusion Model for Physically Plausible Motion Generation
- TSKAN: Interpretable Machine Learning for QoE modeling over Time Series Data
- MDBench: Benchmarking Data-Driven Methods for Model Discovery
- InsightGUIDE: An Opinionated AI Assistant for Guided Critical Reading of Scientific Literature
- CoSupFormer : A Contrastive Supervised learning approach for EEG signal Classification
- Document Summarization with Conformal Importance Guarantees
- Spatio-Temporal Directed Graph Learning for Account Takeover Fraud Detection
- Adaptive Event-Triggered Policy Gradient for Multi-Agent Reinforcement Learning
- Graph Variate Neural Networks
- A HyperGraphMamba-Based Multichannel Adaptive Model for ncRNA Classification
- Low-Resource English-Tigrinya MT: Leveraging Multilingual Models, Custom Tokenizers, and Clean Evaluation Benchmarks
- Can LLMs Forecast Internet Traffic from Social Media?
- Unleashing the Potential of the Semantic Latent Space in Diffusion Models for Image Dehazing
- Can Constructions "SCAN" Compositionality ?
- PS3: A Multimodal Transformer Integrating Pathology Reports with Histology Images and Biological Pathways for Cancer Survival Prediction
- Embodied AI: From LLMs to World Models
- Learning Robust Penetration Testing Policies under Partial Observability: A systematic evaluation
- The Knowledge-Behaviour Disconnect in LLM-based Chatbots
- Improving Generalizability and Undetectability for Targeted Adversarial Attacks on Multimodal Pre-trained Models
- From Samples to Scenarios: A New Paradigm for Probabilistic Forecasting
- OmniScene: Attention-Augmented Multimodal 4D Scene Understanding for Autonomous Driving
- An effective control of large systems of active particles: An application to evacuation problem
- Generalist Robot Manipulation beyond Action Labeled Data
- Multimodal-enhanced Federated Recommendation: A Group-wise Fusion Approach
- AMLA: MUL by ADD in FlashAttention Rescaling
- AJAHR: Amputated Joint Aware 3D Human Mesh Recovery
- Towards Self-Supervised Foundation Models for Critical Care Time Series
- CollaPipe: Adaptive Segment-Optimized Pipeline Parallelism for Collaborative LLM Training in Heterogeneous Edge Networks
- Enhancing Linear Attention with Residual Learning
- BurstEngine: an Efficient Distributed Framework for Training Transformers on Extremely Long Sequences of over 1M Tokens
- Polarity Detection of Sustainable Detection Goals in News Text
- Sobolev acceleration for neural networks
- SMILES-Inspired Transfer Learning for Quantum Operators in Generative Quantum Eigensolver
- Linear Transformers Implicitly Discover Unified Numerical Algorithms
- A Unified Noise-Curvature View of Loss of Trainability
- Diffusion-Based Impedance Learning for Contact-Rich Manipulation Tasks
- Anatomically Constrained Transformers for Cardiac Amyloidosis Classification
- Enhancing Transformer-Based Vision Models: Addressing Feature Map Anomalies Through Novel Optimization Strategies
- RoboSSM: Scalable In-context Imitation Learning via State-Space Models
- IBiT: Utilizing Inductive Biases to Create a More Data Efficient Attention Mechanism
- Analyzing Generalization in Pre-Trained Symbolic Regression
- Blueprint-Bench: Comparing spatial intelligence of LLMs, agents and image models
- Interpretable Kolmogorov-Arnold networks for enzyme commission number prediction
- MSCoD: An Enhanced Bayesian Updating Framework with Multi-Scale Information Bottleneck and Cooperative Attention for Structure-Based Drug Design
- Myosotis: structured computation for attention like layer
- Hierarchical Resolution Transformers: A Wavelet-Inspired Architecture for Multi-Scale Language Understanding
- Are We Scaling the Right Thing? A System Perspective on Test-Time Scaling
- TIMED: Adversarial and Autoregressive Refinement of Diffusion-Based Time Series Generation
- Mamba Modulation: On the Length Generalization of Mamba
- Evaluating Language Translation Models by Playing Telephone
- ExPe: Exact Positional Encodings for Generative Transformer Models with Extrapolating Capabilities
- Confidence Calibration in Large Language Model-Based Entity Matching
- Cognitive Load Limits in Large Language Models: Benchmarking Multi-Hop Reasoning
- Building a User Foundation Model for the Open Web
- HyperClaim: Fine-Grained Cross-Modal Hypergraph Reasoning for Video Misinformation Detection
- SAFViT: Spatial Attention Fusion Gating for Vision Transformer-Based Nucleus Segmentation and Classification
- scPlantAnnotate: an accurate and robust transformer-based model for plant cell type annotation
- Looped Transformers with Source-Centered State Evolution
- RefineSVG: Visual Feedback-Driven Reinforcement Learning for Image-to-SVG Generation
- AWARE-FX: An Auditable Knowledge-Guided AI System for Measuring Corporate Foreign-Exchange Hedging Disclosure
- HealthCAT: An Interpretable Encoder-only Transformer Framework for Health Indicator Prediction and Temporal Interpretation of Wearable Sensor Data
- ProgFormer: Hierarchical Voxel Diffusion Transformer for Longitudinal Brain MRI Prediction
- A Montage-Agnostic Encoder for Calibration-Light Cross-User Gesture Recognition from Surface Electromyography
- Recognition and Label-Free Adaptation Across Recording Sessions in Surface-EMG Gesture Decoding
- Back from the Future: Key-Value Cache Management by Counter-Causal Surprise
- MedXplore: Towards Reliable and Unbiased Generalized Category Discovery in Medical Imaging
- SiamJEPA: On the Role of Siamese Student Encoders in JEPA
- Σ-Mem: An Online Reliability Memory for LLM-based Multi-Agent Systems
- What Makes Graph Unified? Principles and Generative Sliding-Window Transformer for Graph Foundation Models
- Specification-Guided Synthesis of Deadlock-Free Communication Protocol Refinements with Large Language Models
- Chem World: A Large-Scale Benchmark and Physics-Informed Framework for Trustworthy Chemical Property Prediction
- FootprintNet: State-Transition-Guided Dynamic Footprint Learning for Multi-temporal Remote Sensing Change Detection
- Semantic-Aligned Structural Abstraction for Multimodal Sentiment Analysis
- STEREODISCO: Discovering Stereotypicality in LLMs
- It's All Just Vectorization: einx, a Universal Notation for Tensor Operations
- SemPIC: Learning Semantic Position-Independent KV Caches
- Temporal Poisoning: Clean-Label Backdoors via Event Redistribution in SNNs
- Causal Discovery with Inverted Self-attention for Multivariate Time Series
- SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer
- Improving Mental Health Screening and Early Risk Detection in Spanish
- Oracle-Budgeted Molecular Optimization with Short-Term Graph Memory
- A report-grounded vision-language foundation model for colonoscopy from 280000 routine reports
- AuricularWorld: Hierarchical Action-Guided World Modeling for Fine-Grained Auricular Structure Segmentation from CT Scans
- Collaborative feature aggregation for face super-resolution and robust re-identification
- Towards joint scaling laws with optimal batch size schedules
- LM-GRASP: Instance-Specific Language Models for Combinatorial Construction via Online Imitation Learning
- OptGraph: Large Language Models Enhanced Evolutionary Optimization Via Graph Retrieval-Augmented Generation
- Exact Action Values Are Not Enough: Rollout-Verified Reinforcement Fine-Tuning of a Reasoning Model for Multi-Zone VAV Control
- EgoGVAE: Ego-body Mesh Reconstruction via Guided Variational Autoencoder
- ROCS: Request-Oriented Compute Sharing for Efficient Large-Scale Recommendation
- SpecCal: Ambiguity-Aware Candidate Calibration for Infrared Spectrum-Based Molecular Structure Reconstruction
- A Sparse Glimpse of the Whole: Train-Free Self-Speculative Decoding
- Study of ttH and tH production in the H→ττ channel in pp collisions at √(s)=13 TeV and 13.6 TeV with the ATLAS detector
- Generalization Bounds on Optimal Control for Transformer Training and Wasserstein Distributional Robustness
- Are Three Matrices All You Need To Beat the Market? Observable Matrix Dynamics for Portfolio Optimization
- Beyond KV Reconstruction: Functional Reconstruction for MLA Draft Models in Speculative Decoding
- Modeling Decisions in Blockchain Analytics: A Leakage-Aware Evaluation of Tree-Based vs. Sequential Models
- Context-Informed Ship Trajectory Prediction via Conditional Attention
- Multi-Head Attention Residuals
- Escaping the Self-Confirmation Trap: An Execute-Distill-Verify Paradigm for Agentic Experience Learning
- Benchmarking Open-Ended Multi-Agent Coordination in Language Agents
- Edge Prediction for Roof Wireframe Reconstruction with Transformers
- Successive convex optimization for transformer encoder model predictive control
- Thinking in Scales: Accelerating Gigapixel Pathology Image Analysis via Adaptive Continuous Reasoning
- NorBERTo: A ModernBERT Model Trained for Portuguese with 331 Billion Tokens Corpus
- mHC-lite: You Don't Need 20 Sinkhorn-Knopp Iterations
- Persuading large language models to comply with objectionable requests
- EMO: Pretraining Mixture of Experts for Emergent Modularity
- When AI Meets Science: Research Diversity, Interdisciplinarity, Visibility, and Retractions across Disciplines in a Global Surge
- Improving the Performance and Learning Stability of Parallelizable RNNs Designed for Ultra-Low Power Applications
- Training Continuously‐Coupled Reconfigurable Photonic Chips with Quantum Machine Learning
- Retrieval from Within: An Intrinsic Capability of Attention-Based Models
- Deep learning-based 3D morphological segmentation and quantitative growth analysis of field-grown cabbage across the full cycle
- GCVA: A Multiview Fusion Mechanism for Heterogeneous Data Representations
- The Topological Trouble With Transformers
- Driving on Registers
- A transformer-based multi-task deep learning model for urban livability evaluation by fusing remote sensing and textual geospatial data
- Homodyne Photonic Tensor Processor exceeds 1,000-TOPS
- RELISH: LLM REgression with a Latent Iterative State Head
- How ready are we to use artificial intelligence in our fight against antimicrobial resistance? An ESGAID and EAAS perspective
- Detecting Positive Selection by Modeling Structure Within Images of Genetic Variation
- Energy-Driven Adaptive Visual Token Pruning for Efficient Vision-Language Models
- A Theory of Appropriateness That Accounts for Norms of Rationality
- Prediction of motions and mooring tensions for the OC3 spar in short-crested seas using a LSTM NN model, with application to fatigue damage assessment
- GradMAP: Faster Layer Pruning with Gradient Metric and Projection Compensation
- Harnessing Synthetic Data from Generative AI for Statistical Inference
- OSCAgent: Accelerating the Discovery of Organic Solar Cells with LLM Agents
- GCT: A Granger-Causal Transformer for Multivariate Traffic Analysis in Smart Villages
- Amortising Inference and Meta-Learning Priors in Neural Networks
- Use What You Know: Causal Foundation Models with Partial Graphs
- epiGPTope: A Machine Learning-Based Epitope Generator and Classifier
- SHINE: A Scalable In-Context Hypernetwork for Mapping Context to LoRA in a Single Pass
- Beyond computational equivalence: the behavioral inference principle for machine consciousness
- When and how to disclose AI use in academic publishing: AMEE Guide No.192
- Previously on... Automating Code Review
- Topology Aware Neural Interpolation of Scalar Fields
- Automating Conflict-Aware ACL Configurations with Natural Language Intents
- SAIL-Recon: Large SfM by Augmenting Scene Regression with Localization
- A Unified Transformer Architecture for Low-Latency and Scalable Wireless Signal Processing
- Optimizing protein tokenization: reduced amino acid alphabets for efficient and accurate protein language models
- Understanding the hippocampus as an apex of the cortical hierarchy and self-supervised predictive learning engine
- Not all linguistic variation is equally predictable
- ALIMA – Ein RAG-basiertes System zur LLM-gestützten Sacherschließung: Prototypentwicklung und erste Erfahrungen aus der Praxis
- Recent Advances and Future Perspectives of AI-Based Mineral Exploration: A Review of Machine Learning, Deep Learning, and Geologically Informed Approaches
- Learning a Sampling-Free Variational DNN Plugin from Tiny Training Sets to Refine OOD Segmentation With Uncertainty Estimation
- A study of word embedding models for measuring topic coherence
- Global Forecasting of Tropical Cyclone Intensity Using Neural Weather Models
- GRIPHIN: grids of pharmacophore interaction fields for affinity prediction
- Exploring temporal dynamics in digital trace data: mining user-sequences for communication research
- ISALux: Illumination and Segmentation Aware Transformer Employing Mixture of Experts for Low Light Image Enhancement
- Ada-TransGNN: An Air Quality Prediction Model Based On Adaptive Graph Convolutional Networks
- Algebraic Approach to Ridge-Regularized Mean Squared Error Minimization in Minimal ReLU Neural Network
- scGMB: A scRNA‐seq Cell Classification Method Combining GCN and Mamba
- Reading the unreadable: creating a dataset of 19th century English newspapers using image-to-text language models
- Novel hybrid machine learning framework for high-fidelity prediction of fly ash-based geopolymer concrete strength
- LexSemBridge: Fine-Grained Dense Representation Enhancement through Token-Aware Embedding Augmentation
- VISA: Group-wise Visual Token Selection and Aggregation via Graph Summarization for Efficient MLLMs Inference
- Helixer: ab initio prediction of primary eukaryotic gene models combining deep learning and a hidden Markov model
- Exploiting individual differences to bootstrap communication
- Limitations of Normalization in Attention Mechanism
- Zero-shot Context Biasing with Trie-based Decoding using Synthetic Multi-Pronunciation
- Optimizing Informer with Whale Optimization Algorithm for Enhanced Ship Trajectory Prediction
- AI-CNet3D: An Anatomically-Informed Cross-Attention Network with Multi-Task Consistency Fine-tuning for 3D Glaucoma Classification
- DiffusionGS: Generative Search with Query Conditioned Diffusion in Kuaishou
- Transformer Modeling for Both Scalability and Performance in Multivariate Time Series
- Evaluation-Aware Reinforcement Learning
- Two-Timescale Learning for Pilot-Free ISAC Systems
- DroneKey: Drone 3D Pose Estimation in Image Sequences using Gated Key-representation and Pose-adaptive Learning
- Lyra: Generative 3D Scene Reconstruction via Video Diffusion Model Self-Distillation
- Defending against Stegomalware in Deep Neural Networks with Permutation Symmetry
- Adversarially-Refined VQ-GAN with Dense Motion Tokenization for Spatio-Temporal Heatmaps
- CompLLM: Compression for Long Context Q&A
- MsFIN: Multi-scale Feature Interaction Network for Traffic Accident Anticipation
- Systematic Comparative Analysis of Large Pretrained Language Models on Contextualized Medication Event Extraction
- GSTM-HMU: Generative Spatio-Temporal Modeling for Human Mobility Understanding
- Improving Credit Card Fraud Detection through Transformer-Enhanced GAN Oversampling
- Layerwise Importance Analysis of Feed-Forward Networks in Transformer-based Language Models
- LLM-based Agents Suffer from Hallucinations: A Survey of Taxonomy, Methods, and Directions
- Audio-Driven Universal Gaussian Head Avatars
- Advancing Metallic Surface Defect Detection via Anomaly-Guided Pretraining on a Large Industrial Dataset
- Enhancing the Effectiveness and Durability of Backdoor Attacks in Federated Learning through Maximizing Task Distinction
- Dynamic Multi-Target Fusion for Efficient Audio-Visual Navigation
- Text Slider: Efficient and Plug-and-Play Continuous Concept Control for Image/Video Synthesis via LoRA Adapters
- Reconstruction of Optical Coherence Tomography Images from Wavelength-space Using Deep-learning
- Climate-Adaptive and Cascade-Constrained Machine Learning Prediction for Sea Surface Height under Greenhouse Warming
- Knowledge Transfer from Interaction Learning
- LEAF-Mamba: Local Emphatic and Adaptive Fusion State Space Model for RGB-D Salient Object Detection
- CATformer: Contrastive Adversarial Transformer for Image Super-Resolution
- MolMark: Safeguarding Molecular Structures through Learnable Atom-Level Watermarking
- Database Normalization via Dual-LLM Self-Refinement
- Scalable Evaluation for Audio Identification via Synthetic Latent Fingerprint Generation
- MLF-4DRCNet: Multi-Level Fusion with 4D Radar and Camera for 3D Object Detection in Autonomous Driving
- SSCM: A Spatial-Semantic Consistent Model for Multi-Contrast MRI Super-Resolution
- LCMF: Lightweight Cross-Modality Mambaformer for Embodied Robotics VQA
- Explainable Graph Neural Networks: Understanding Brain Connectivity and Biomarkers in Dementia
- Automatic coherence-driven inference on arguments
- Dynamical Modeling of Behaviorally Relevant Spatiotemporal Patterns in Neural Imaging Data
- Machine learning approach to single-shot multiparameter estimation for the non-linear Schrödinger equation
- Single-Branch Network Architectures to Close the Modality Gap in Multimodal Recommendation
- False Friends Are Not Foes: Investigating Vocabulary Overlap in Multilingual Language Models
- Enabling Plant Phenotyping in Weedy Environments using Multi-Modal Imagery via Synthetic and Generated Training Data
- Experience Scaling: Post-Deployment Evolution For Large Language Models
- OmniBridge: Unified Multimodal Understanding, Generation, and Retrieval via Latent Space Alignment
- Trace Is In Sentences: Unbiased Lightweight ChatGPT-Generated Text Detector
- Live-E2T: Real-time Threat Monitoring in Video via Deduplicated Event Reasoning and Chain-of-Thought
- M4SER: Multimodal, Multirepresentation, Multitask, and Multistrategy Learning for Speech Emotion Recognition
- Lightweight Vision Transformer with Window and Spatial Attention for Food Image Classification
- Characterizing the Behavior of Training Mamba-based State Space Models on GPUs
- An overview of neural architectures for self-supervised audio representation learning from masked spectrograms
- Towards Practical Multi-label Causal Discovery in High-Dimensional Event Sequences via One-Shot Graph Aggregation
- Flow marching for a generative PDE foundation model
- LiDAR Point Cloud Image-based Generation Using Denoising Diffusion Probabilistic Models
- Reading Images Like Texts: Sequential Image Understanding in Vision-Language Models
- WolBanking77: Wolof Banking Speech Intent Classification Dataset
- Bounded PCTL Model Checking of Large Language Model Outputs
- MOMEMTO: Patch-based Memory Gate Model in Time Series Foundation Model
- Zero-Shot Visual Deepfake Detection: Can AI Predict and Prevent Fake Content Before It's Created?
- GluMind: Multimodal Parallel Attention and Knowledge Retention for Robust Cross-Population Blood Glucose Forecasting
- Explicit Path CGR: Maintaining Sequence Fidelity in Geometric Representations
- Policy Gradient with Self-Attention for Model-Free Distributed Nonlinear Multi-Agent Games
- Generalization and the Rise of System-level Creativity in Science
- A Graph-Neural-Network-Entropy model of vital node identification on network attack and propagation
- Learning to Condition: A Neural Heuristic for Scalable MPE Inference
- Tracing the Techno-Supremacy Doctrine: A Critical Discourse Analysis of the AI Executive Elite
- TensLoRA: Tensor Alternatives for Low-Rank Adaptation
- SSNet: Flexible and robust channel extrapolation for fluid antenna systems enabled by an self-supervised learning framework
- Degradation-Aware All-in-One Image Restoration via Latent Prior Encoding
- Turk-LettuceDetect: A Hallucination Detection Models for Turkish RAG Applications
- SD-VLM: Spatial Measuring and Understanding with Depth-Encoded Vision-Language Models
- SeqBattNet: A Discrete-State Physics-Informed Neural Network with Aging Adaptation for Battery Modeling
- Visual Instruction Pretraining for Domain-Specific Foundation Models
- Exploring Machine Learning Models for Physical Dose Calculation in Carbon Ion Therapy Using Heterogeneous Imaging Data -- A Proof of Concept Study
- Pre-Trained CNN Architecture for Transformer-Based Image Caption Generation Model
- Cronus: Efficient LLM inference on Heterogeneous GPU Clusters via Partially Disaggregated Prefill
- LLaVul: A Multimodal LLM for Interpretable Vulnerability Reasoning about Source Code
- DiffQ: Unified Parameter Initialization for Variational Quantum Algorithms via Diffusion Models
- Clotho: Measuring Task-Specific Pre-Generation Test Adequacy for LLM Inputs
- Rational Multi-Modal Transformers for TCR-pMHC Prediction
- Attention-based Mixture of Experts for Robust Speech Deepfake Detection
- BlurBall: Joint Ball and Motion Blur Estimation for Table Tennis Ball Tracking
- Interpreting Attention Heads for Image-to-Text Information Flow in Large Vision-Language Models
- Speech Vecalign: an Embedding-based Method for Aligning Parallel Speech Documents
- ControlEchoSynth: Boosting Ejection Fraction Estimation Models via Controlled Video Diffusion
- DINOv3-Diffusion Policy: Self-Supervised Large Visual Model for Visuomotor Diffusion Policy Learning
- Learning to vary: Teaching LMs to reproduce human linguistic variability in next-word prediction
- Neural Network-Driven Direct CBCT-Based Dose Calculation for Head-and-Neck Proton Treatment Planning
- CSDformer: A Conversion Method for Fully Spike-Driven Transformer
- Audio Super-Resolution with Latent Bridge Models
- ReDepress: A Cognitive Framework for Detecting Depression Relapse from Social Media
- An Artificial Intelligence Value at Risk Approach: Metrics and Models
- Graph Enhanced Trajectory Anomaly Detection
- CPT-4DMR: Continuous sPatial-Temporal Representation for 4D-MRI Reconstruction
- Towards Provable Emergence of In-Context Reinforcement Learning
- Achilles' Heel of Mamba: Essential difficulties of the Mamba architecture demonstrated by synthetic data
- Medical priority fusion: achieving dual optimization of sensitivity and interpretability in nipt anomaly detection
- Automatic Intermodal Loading Unit Identification using Computer Vision: A Scoping Review
- Cross-Attention is Half Explanation in Speech-to-Text Models
- Elucidating the Design Space of FP4 training
- SilentStriker:Toward Stealthy Bit-Flip Attacks on Large Language Models
- StableGuard: Towards Unified Copyright Protection and Tamper Localization in Latent Diffusion Models
- GEM-T: Generative Tabular Data via Fitting Moments
- RoboSeek: You Need to Interact with Your Objects
- GraDeT-HTR: A Resource-Efficient Bengali Handwritten Text Recognition System utilizing Grapheme-based Tokenizer and Decoder-only Transformer
- Incorporating the Refractory Period into Spiking Neural Networks through Spike-Triggered Threshold Dynamics
- Understanding Post-Training Structural Changes in Large Language Models
- DA-Mamba: Dialogue-aware selective state-space model for multimodal engagement estimation
- OnePiece: Bringing Context Engineering and Reasoning to Industrial Cascade Ranking System
- DriveDPO: Policy Learning via Safety DPO For End-to-End Autonomous Driving
- Clothing agnostic Pre-inpainting Virtual Try-ON
- FRET-guided selection of RNA 3D structures
- Can Agents Judge Systematic Reviews Like Humans? Evaluating SLRs with LLM-based Multi-Agent System
- nDNA -- the Semantic Helix of Artificial Cognition
- Point-RTD: Replaced Token Denoising for Pretraining Transformer Models on Point Clouds
- MAST: Multi-Agent Spatial Transformer for Learning to Collaborate
- Dendritic Resonate-and-Fire Neuron for Effective and Efficient Long Sequence Modeling
- Exploring Synthesizable Chemical Space with Iterative Pathway Refinements
- MTS-DMAE: Dual-Masked Autoencoder for Unsupervised Multivariate Time Series Representation Learning
- LEMs: A Primer On Large Execution Models
- MaskVCT: Masked Voice Codec Transformer for Zero-Shot Voice Conversion With Increased Controllability via Multiple Guidances
- Data-Driven Reconstruction of Significant Wave Heights from Sparse Observations
- TSGym: Design Choices for Deep Multivariate Time-Series Forecasting
- CoPAD : Multi-source Trajectory Fusion and Cooperative Trajectory Prediction with Anchor-oriented Decoder in V2X Scenarios
- Learning Attribute-Aware Hash Codes for Fine-Grained Image Retrieval via Query Optimization
- Attentive AV-FusionNet: Audio-Visual Quality Prediction with Hybrid Attention
- ExpertWeave: Efficiently Serving Expert-Specialized Fine-Tuned Adapters at Scale
- Quantum Adaptive Self-Attention for Financial Rebalancing: An Empirical Study on Automated Market Makers in Decentralized Finance
- Equip Pre-ranking with Target Attention by Residual Quantization
- Cross-Attention with Confidence Weighting for Multi-Channel Audio Alignment
- Advancing Audio-Visual Navigation Through Multi-Agent Collaboration in 3D Environments
- End2Race: Efficient End-to-End Imitation Learning for Real-Time F1Tenth Racing
- seqBench: A Tunable Benchmark to Quantify Sequential Reasoning Limits of LLMs
- Towards Interpretable and Efficient Attention: Compressing All by Contracting a Few
- From Prediction to Understanding: Will AI Foundation Models Transform Brain Science?
- Time Series Forecasting Using a Hybrid Deep Learning Method: A Bi-LSTM Embedding Denoising Auto Encoder Transformer
- Clinical Multi-modal Fusion with Heterogeneous Graph and Disease Correlation Learning for Multi-Disease Prediction
- Scalable Multi Agent Diffusion Policies for Coverage Control
- Neural Network Based Framework for Passive Intermodulation Cancellation in MIMO Systems
- Prospective Multi-Graph Cohesion for Multivariate Time Series Anomaly Detection
- Deep Learning Inductive Biases for fMRI Time Series Classification during Resting-state and Movie-watching
- Conditional Policy Generator for Dynamic Constraint Satisfaction and Optimization
- ScenGAN: Attention-Intensive Generative Model for Uncertainty-Aware Renewable Scenario Forecasting
- RALLM-POI: Retrieval-Augmented LLM for Zero-shot Next POI Recommendation with Geographical Reranking
- Analyzing Memory Effects in Large Language Models through the lens of Cognitive Psychology
- DecipherGuard: Understanding and Deciphering Jailbreak Prompts for a Safer Deployment of Intelligent Software Systems
- Learning Hyperspectral Images with Curated Text Prompts for Efficient Multimodal Alignment
- DiffEye: Diffusion-Based Continuous Eye-Tracking Data Generation Conditioned on Natural Images
- Angular Dispersion Accelerates k-Nearest Neighbors Machine Translation
- MedCutMix: A Data-Centric Approach to Improve Radiology Vision-Language Pre-training with Disease Awareness
- Self-Supervised Learning of Graph Representations for Network Intrusion Detection
- Learn to Rank Risky Investors: A Case Study of Predicting Retail Traders' Behaviour and Profitability
- A Novel Differential Feature Learning for Effective Hallucination Detection and Classification
- KV-Efficient VLA: A Method to Speed up Vision Language Models with RNN-Gated Chunked KV Cache
- DA-Font: Few-Shot Font Generation via Dual-Attention Hybrid Integration
- Thermal Imaging-based Real-time Fall Detection using Motion Flow and Attention-enhanced Convolutional Recurrent Architecture
- Assessing Classical Machine Learning and Transformer-based Approaches for Detecting AI-Generated Research Text
- GWM: Towards Scalable Gaussian World Models for Robotic Manipulation
- ADVEDM:Fine-grained Adversarial Attack against VLM-based Embodied Agents
- Surgical-MambaLLM: Mamba2-enhanced Multimodal Large Language Model for VQLA in Robotic Surgery
- ViTCAE: ViT-based Class-conditioned Autoencoder
- Efficient Rectified Flow for Image Fusion
- Enhancing Scientific Visual Question Answering via Vision-Caption aware Supervised Fine-Tuning
- DISCO: Disentangled Communication Steering for Large Language Models
- AHA -- Predicting What Matters Next: Online Highlight Detection Without Looking Ahead
- OS-DiffVSR: Towards One-step Latent Diffusion Model for High-detailed Real-world Video Super-Resolution
- Causality-Induced Positional Encoding for Transformer-Based Representation Learning of Non-Sequential Features
- Question Answering with LLMs and Learning from Answer Sets
- DoubleGen: Debiased Generative Modeling of Counterfactuals
- CGTGait: Collaborative Graph and Transformer for Gait Emotion Recognition
- EG-MLA: Embedding-Gated Multi-head Latent Attention for Scalable and Efficient LLMs
- HypeMARL: Multi-Agent Reinforcement Learning For High-Dimensional, Parametric, and Distributed Systems
- Dynamic Objects Relocalization in Changing Environments with Flow Matching
- Improving Deep Tabular Learning
- Recovering Parametric Scenes from Very Few Time-of-Flight Pixels
- RadarGaussianDet3D: An Efficient and Effective Gaussian-based 3D Detector with 4D Automotive Radars
- Randomized Smoothing Meets Vision-Language Models
- FMD-TransUNet: Abdominal Multi-Organ Segmentation Based on Frequency Domain Multi-Axis Representation Learning and Dual Attention Mechanisms
- Sequential Token Merging: Revisiting Hidden States
- Sampling String Vacua Using Generative Models
- Interpreting the Role of Visemes in Audio-Visual Speech Recognition
- DAFTED: Decoupled Asymmetric Fusion of Tabular and Echocardiographic Data for Cardiac Hypertension Diagnosis
- BEFT: Bias-Efficient Fine-Tuning of Language Models
- A multi-temporal multi-spectral attention-augmented deep convolution neural network with contrastive learning for crop yield prediction
- ArchesClimate: Probabilistic Decadal Ensemble Generation With Flow Matching
- Efficient Pre-Training of LLMs via Topology-Aware Communication Alignment on More Than 9600 GPUs
- PAN: Pillars-Attention-Based Network for 3D Object Detection
- Global Regulation and Excitation via Attention Tuning for Stereo Matching
- Monte Carlo Tree Diffusion with Multiple Experts for Protein Design
- CBPNet: A Continual Backpropagation Prompt Network for Alleviating Plasticity Loss on Edge Devices
- UniGist: Towards General and Hardware-aligned Sequence-level Long Context Compression
- A Memory Efficient Adjoint Method to Enable Billion Parameter Optimization on a Single GPU in Dynamic Problems
- Understanding Embedding Scaling in Collaborative Filtering
- CCrepairBench: A High-Fidelity Benchmark and Reinforcement Learning Framework for C++ Compilation Repair
- Interpretable Modeling of Articulatory Temporal Dynamics from real-time MRI for Phoneme Recognition
- Blind Source Separation of Radar Signals in Time Domain Using Deep Learning
- Bridging Graph and State-Space Modeling for Intensive Care Unit Length of Stay Prediction
- Adversarially Robust Assembly Language Model for Packed Executables Detection
- Mental Accounts for Actions: EWA-Inspired Attention in Decision Transformers
- Using machine learning to automate data annotation in corpus linguistics
- Escaping saddle points without Lipschitz smoothness: the power of nonlinear preconditioning
- Diffusion-Based Cross-Modal Feature Extraction for Multi-Label Classification
- Deep Learning Empowered Super-Resolution: A Comprehensive Survey and Future Prospects
- TractoTransformer: Diffusion MRI Streamline Tractography using CNN and Transformer Networks
- Localmax dynamics for attention in transformers and its asymptotic behavior
- Language-Instructed Reasoning for Group Activity Detection via Multimodal Large Language Model
- Chunk Based Speech Pre-training with High Resolution Finite Scalar Quantization
- How Large Language Models are Designed to Hallucinate
- AdaSports-Traj: Role- and Domain-Aware Adaptation for Multi-Agent Trajectory Modeling in Sports
- Enhancing Generative Auto-bidding with Offline Reward Evaluation and Policy Search
- Attention Schema-based Attention Control (ASAC): A Cognitive-Inspired Approach for Attention Management in Transformers
- LightCode: Compiling LLM Inference for Photonic-Electronic Systems
- SolarCrossFormer: Improving day-ahead Solar Irradiance Forecasting by Integrating Satellite Imagery and Ground Sensors
- RangeSAM: On the Potential of Visual Foundation Models for Range-View represented LiDAR segmentation
- Efficient Extractive Text Summarization for Online News Articles Using Machine Learning
- Approximate Fiber Products of Schemes and Their Étale Homotopical Invariants
- Prostate Capsule Segmentation from Micro-Ultrasound Images using Adaptive Focal Loss
- Neural Architecture Search Algorithms for Quantum Autoencoders
- Region-Aware Deformable Convolutions
- Accurate typhoon intensity forecasts using a non-iterative spatiotemporal transformer model
- SPH-Net: A Co-Attention Hybrid Model for Accurate Stock Price Prediction
- MaskAttn-SDXL: Controllable Region-Level Text-To-Image Generation
- CoDoL: Conditional Domain Prompt Learning for Out-of-Distribution Generalization
- Kuramoto Orientation Diffusion Models
- Lightweight and Accurate Multi-View Stereo with Confidence-Aware Diffusion Model
- Out-of-Sight Trajectories: Tracking, Fusion, and Prediction
- TITAN: A Trajectory-Informed Technique for Adaptive Parameter Freezing in Large-Scale VQE
- Language Modeling with Learned Meta-Tokens
- SMARTER: A Data-efficient Framework to Improve Toxicity Detection with Explanation via Self-augmenting Large Language Models
- Leveraging Geometric Visual Illusions as Perceptual Inductive Biases for Vision Models
- Self-Improving Embodied Foundation Models
- OmniMRI: A Unified Vision--Language Foundation Model for Generalist MRI Interpretation
- OmniSegmentor: A Flexible Multi-Modal Learning Framework for Semantic Segmentation
- From Patterns to Predictions: A Shapelet-Based Framework for Directional Forecasting in Noisy Financial Markets
- TANDEM: Temporal Attention-guided Neural Differential Equations for Missingness in Time Series Classification
- GateTS: Versatile and Efficient Forecasting via Attention-Inspired routed Mixture-of-Experts
- Social-MAE: A Transformer-Based Multimodal Autoencoder for Face and Voice
- Multimodal Representation Learning Conditioned on Semantic Relations
- Attention Beyond Neighborhoods: Reviving Transformer for Graph Clustering
- Beyond Random Masking: A Dual-Stream Approach for Rotation-Invariant Point Cloud Masked Autoencoders
- FAWN: A MultiEncoder Fusion-Attention Wave Network for Integrated Sensing and Communication Indoor Scene Inference
- A Novel Task-Driven Diffusion-Based Policy with Affordance Learning for Generalizable Manipulation of Articulated Objects
- Searching for Lorentz invariance violation with artificial neural networks
- Fracture interactive geodesic active contours for bone segmentation
- Trade-offs in Cross-Domain Generalization of Foundation Model Fine-Tuned for Biometric Applications
- Pseudo-Label Enhanced Cascaded Framework: 2nd Technical Report for LSVOS 2025 VOS Track
- Partial Column Generation with Graph Neural Networks for Team Formation and Routing
- Llama-Mimi: Speech Language Models with Interleaved Semantic and Acoustic Tokens
- Exploring the Global-to-Local Attention Scheme in Graph Transformers: An Empirical Study
- MeanFlowSE: one-step generative speech enhancement via conditional mean flow
- UMind: A Unified Multitask Network for Zero-Shot M/EEG Visual Decoding
- Subject Matter Expertise vs Professional Management in Collective Sequential Decision Making
- MoE-Inference-Bench: Performance Evaluation of Mixture of Expert Large Language and Vision Models
- Generative AI Meets Wireless Sensing: Towards Wireless Foundation Model
- SpeechMLC: Speech Multi-label Classification
- Attention Lattice Adapter: Visual Explanation Generation for Visual Foundation Model
- A Noninvasive and Dispersive Framework for Estimating Nonuniform Conductivity of Brain Tumor in Patient-Specific Head Models
- DIPP: Discriminative Impact Point Predictor for Catching Diverse In-Flight Objects
- UMA-Split: unimodal aggregation for both English and Mandarin non-autoregressive speech recognition
- SALT4Decompile: Inferring Source-level Abstract Logic Tree for LLM-Based Binary Decompilation
- Evaluating the Impact of Verbal Multiword Expressions on Machine Translation
- DeCoP: Enhancing Self-Supervised Time Series Representation with Dependency Controlled Pre-training
- Efficient 3D Perception on Embedded Systems via Interpolation-Free Tri-Plane Lifting and Volume Fusion
- DyWPE: Signal-Aware Dynamic Wavelet Positional Encoding for Time Series Transformers
- Automating Modelica Module Generation Using Large Language Models: A Case Study on Building Control Description Language
- LSTC-MDA: A Unified Framework for Long-Short Term Temporal Convolution and Mixed Data Augmentation in Skeleton-Based Action Recognition
- Leveraging IndoBERT and DistilBERT for Indonesian Emotion Classification in E-Commerce Reviews
- HybridMamba: A Dual-domain Mamba for 3D Medical Image Segmentation
- Structure-Preserving Margin Distribution Learning for High-Order Tensor Data with Low-Rank Decomposition
- Frequency-Aware Ensemble Learning for BraTS 2025 Pediatric Brain Tumor Segmentation
- VisMoDAl: Visual Analytics for Evaluating and Improving Corruption Robustness of Vision-Language Models
- Controlling Language Difficulty in Dialogues with Linguistic Features
- Introducing OmniGEC: A Silver Multilingual Dataset for Grammatical Error Correction
- Super-Linear: A Lightweight Pretrained Mixture of Linear Experts for Time Series Forecasting
- Data from Discrete Flow-Based Generative Models for Measurement Optimization in Quantum Computing
- Asymptotic Study of In-context Learning with Random Transformers through Equivalent Models
- Bidirectional Feature-aligned Motion Transformation for Efficient Dynamic Point Cloud Compression
- Deep learning and abstractive summarisation for radiological reports: an empirical study for adapting the PEGASUS models' family with scarce data
- Beyond Marginals: Learning Joint Spatio-Temporal Patterns for Multivariate Anomaly Detection
- Hierarchical Self-Attention: Generalizing Neural Attention Mechanics to Multi-Scale Problems
- DF-LLaVA: Unlocking MLLMs for Synthetic Image Detection via Knowledge Injection and Conflict-Driven Self-Reflection
- Not What the Doctor Ordered: Surveying LLM-based De-identification and Quantifying Clinical Information Loss
- Synthetic bootstrapped pretraining
- VLHSA: Vision-Language Hierarchical Semantic Alignment for Jigsaw Puzzle Solving with Eroded Gaps
- A Framework for Generating Artificial Datasets to Validate Absolute and Relative Position Concepts
- A Variational Framework for Residual-Based Adaptivity in Neural PDE Solvers and Operator Learning
- An Attention-Based Stochastic Simulator for Multisite Extremes to Evaluate Nonstationary, Cascading Flood Risk
- Resolving the Body-Order Paradox of Machine Learning Interatomic Potentials
- MoSA: Motion-Coherent Human Video Generation via Structure-Appearance Decoupling
- Large Language Models as Universal Predictors? An Empirical Study on Small Tabular Datasets
- VSE-MOT: Multi-Object Tracking in Low-Quality Video Scenes Guided by Visual Semantic Enhancement
- SSL-SSAW: Self-Supervised Learning with Sigmoid Self-Attention Weighting for Question-Based Sign Language Translation
- RepCaM++: Exploring Transparent Visual Prompt With Inference-Time Re-Parameterization for Neural Video Delivery
- MOCHA: Multi-modal Objects-aware Cross-arcHitecture Alignment
- Inverse Design of Amorphous Materials with Targeted Properties
- White Aggregation and Restoration for Few-shot 3D Point Cloud Semantic Segmentation
- Detecting Struggling Student Programmers using Proficiency Taxonomies
- Noise Supervised Contrastive Learning and Feature-Perturbed for Anomalous Sound Detection
- Pre-Manipulation Alignment Prediction with Parallel Deep State-Space and Transformer Models
- Masked Feature Modeling Enhances Adaptive Segmentation
- Improving Generalized Visual Grounding with Instance-aware Joint Learning
- State Space Models over Directed Graphs
- FLAME: A Serving System Optimized for Large-Scale Generative Recommendation with Efficiency
- Re-purposing SAM into Efficient Visual Projectors for MLLM-Based Referring Image Segmentation
- Deep Lookup Network
- PROFUSEme: PROstate Cancer Biochemical Recurrence Prediction via FUSEd Multi-modal Embeddings
- VocSegMRI: Multimodal Learning for Precise Vocal Tract Segmentation in Real-time MRI
- CLMTracing: Black-box User-level Watermarking for Code Language Model Tracing
- Understanding and Tackling Over-Dilution in Graph Neural Networks
- Learning the natural history of human disease with generative transformers
- You Are What You Train: Effects of Data Composition on Training Context-aware Machine Translation Models
- FishBEV: Distortion-Resilient Bird's Eye View Segmentation with Surround-View Fisheye Cameras
- StyleProtect: Safeguarding Artistic Identity in Fine-tuned Diffusion Models
- Do Large Language Models Understand Word Senses?
- Local-Canonicalization Equivariant Graph Neural Networks for Sample-Efficient and Generalizable Swarm Robot Control
- Agile in the Face of Delay: Asynchronous End-to-End Learning for Real-World Aerial Navigation
- ShortListing Model: A Streamlined SimplexDiffusion for Discrete Variable Generation
- Sequential Data Augmentation for Generative Recommendation
- Risk Assessment and Security Analysis of Large Language Models
- PosBridge: Multi-View Positional Embedding Transplant for Identity-Aware Image Editing
- LLM on a Budget: Active Knowledge Distillation for Efficient Classification of Large Text Corpora
- TreeIRL: Safe Urban Driving with Tree Search and Inverse Reinforcement Learning
- Learning Nonlinear Responses in PET Bottle Buckling with a Hybrid DeepONet-Transolver Framework
- ERF-BA-TFD+: A Multimodal Model for Audio-Visual Deepfake Detection
- Accelerating Protein Molecular Dynamics Simulation with DeepJump
- A Lightweight Architecture for Multi-instrument Transcription with Practical Optimizations
- A Design Co-Pilot for Task-Tailored Manipulators
- MSR-Codec: A Low-Bitrate Multi-Stream Residual Codec for High-Fidelity Speech Generation with Information Disentanglement
- HAM: Hierarchical Adapter Merging for Scalable Continual Learning
- MAGIC-Enhanced Keyword Prompting for Zero-Shot Audio Captioning with CLIP Models
- Channel Estimation for Rydberg Atomic Quantum Receivers
- LadderSym: A Multimodal Interleaved Transformer for Music Practice Error Detection
- The Few-shot Dilemma: Over-prompting Large Language Models
- From Next Token Prediction to (STRIPS) World Models -- Preliminary Results
- Green Recommender Systems: Understanding and Minimizing the Carbon Footprint of AI-Powered Personalization
- HPIM: Heterogeneous Processing-In-Memory-based Accelerator for Large Language Models Inference
- SHREC 2025: Protein surface shape retrieval including electrostatic potential
- MTNet: Learning modality-aware representation with transformer for RGBT tracking
- Reversible Deep Equilibrium Models
- LTA-thinker: Latent Thought-Augmented Training Framework for Large Language Models on Complex Reasoning
- GRATE: a Graph transformer-based deep Reinforcement learning Approach for Time-efficient autonomous robot Exploration
- Similarity-Distance-Magnitude Activations
- HQCNN: A Hybrid Quantum-Classical Neural Network for Medical Image Classification
- Fast reconstruction of degenerate populations of conductance-based neuron models from spike times
- Performance is not All You Need: Sustainability Considerations for Algorithms
- Provable Generalization in Overparameterized Neural Nets
- A Novel Recurrent Neural Network Framework for Prediction and Treatment of Oncogenic Mutation Progression
- SPGen: Spherical Projection as Consistent and Flexible Representation for Single Image 3D Shape Generation
- Valuation of Exotic Options and Counterparty Games Based on Conditional Diffusion
- Harnessing the Power of AI in Qualitative Research: Role Assignment, Engagement, and User Perceptions of AI-Generated Follow-Up Questions in Semi-Structured Interviews
- Soft Graph Transformer for MIMO Detection
- Large Language Model Scaling Laws for Neural Quantum States in Quantum Chemistry
- MFAF: An EVA02-Based Multi-scale Frequency Attention Fusion Method for Cross-View Geo-Localization
- Beyond Artificial Misalignment: Detecting and Grounding Semantic-Coordinated Multimodal Manipulations
- Positional Encoding via Token-Aware Phase Attention
- Yet Another Watermark for Large Language Models
- LEAF: Knowledge Distillation of Text Embedding Models with Teacher-Aligned Representations
- Multimodal Hate Detection Using Dual-Stream Graph Neural Networks
- BATR-FST: Bi-Level Adaptive Token Refinement for Few-Shot Transformers
- A Data-Aware Fourier Neural Operator for Modeling Spatiotemporal Electromagnetic Fields
- TFANet: Three-Stage Image-Text Feature Alignment Network for Robust Referring Image Segmentation
- Module-Aware Parameter-Efficient Machine Unlearning on Transformers
- Multi-Metric Preference Alignment for Generative Speech Restoration
- GPG-HT: Generalized Policy Gradient with History-Aware Decision Transformer for Probabilistic Path Planning
- The Lifecycle Principle: Stabilizing Dynamic Neural Networks with State Memory
- Retrieve-and-Verify: A Table Context Selection Framework for Accurate Column Annotations
- Multi-scale Scanning Network for Machine Anomalous Sound Detection
- Scaling Graph Transformers: A Comparative Study of Sparse and Dense Attention
- FusionMAE: large-scale pretrained model to optimize and simplify diagnostic and control of fusion plasma
- Enhancing Video Large Language Models with Structured Multi-Video Collaborative Reasoning
- Match Chat: Real Time Generative AI and Generative Computing for Tennis
- Large Language Models Imitate Logical Reasoning, but at what Cost?
- T-SiamTPN: Temporal Siamese Transformer Pyramid Networks for Robust and Efficient UAV Tracking
- End4: End-to-end Denoising Diffusion for Diffusion-Based Inpainting Detection
- A Multimodal Foundation Model to Enhance Generalizability and Data Efficiency for Pan-cancer Prognosis Prediction
- CECT-Mamba: a Hierarchical Contrast-enhanced-aware Model for Pancreatic Tumor Subtyping from Multi-phase CECT
- Instant prediction of relaxation in moiré superlattices using neural networks
- Ensembling Large Language Models for Code Vulnerability Detection: An Empirical Evaluation
- Generative AI for Multimedia Communication: Recent Advances, An Information-Theoretic Framework, and Future Opportunities
- What Developers Ask to ChatGPT in GitHub Pull Requests? an Exploratory Study
- DyGLNet: Hybrid Global-Local Feature Fusion with Dynamic Upsampling for Medical Image Segmentation
- Geolocation-Aware Robust Spoken Language Identification
- NEFT: A Unified Transformer Framework for Efficient Near-Field CSI Feedback in XL-MIMO Systems
- Road Obstacle Video Segmentation
- Multimodal Graph Network Modeling for Human-Object Interaction Detection with PDE Graph Diffusion
- A comparison of pipelines for the translation of a low resource language based on transformers
- Text Adaptation to Plain Language and Easy Read via Automatic Post-Editing Cycles
- Towards Foundational Models for Single-Chip Radar
- Neural-Quantum-States Impurity Solver for Quantum Embedding Problems
- Causal-Symbolic Meta-Learning (CSML): Inducing Causal World Models for Few-Shot Generalization
- Explainable Unsupervised Multi-Anomaly Detection and Temporal Localization in Nuclear Times Series Data with a Dual Attention-Based Autoencoder
- Entanglement and optimization within autoregressive neural quantum states
- DS@GT AnimalCLEF: Triplet Learning over ViT Manifolds with Nearest Neighbor Classification for Animal Re-identification
- Integrating Attention-Enhanced LSTM and Particle Swarm Optimization for Dynamic Pricing and Replenishment Strategies in Fresh Food Supermarkets
- Dynamic Relational Priming Improves Transformer in Multivariate Time Series
- U-Mamba2: Scaling State Space Models for Dental Anatomy Segmentation in CBCT
- Embodied Navigation Foundation Model
- GTA: Supervised-Guided Reinforcement Learning for Text Classification with Large Language Models
- End-to-End 4D Heart Mesh Recovery Across Full-Stack and Sparse Cardiac MRI
- Asterisk Operator
- Low-rank Orthogonalization for Large-scale Matrix Optimization with Applications to Foundation Model Training
- Examining the Relationship between Scientific Publishing Activity and Hype-Driven Financial Bubbles: A Comparison of the Dot-Com and AI Eras
- A Straightforward Pipeline for Targeted Entailment and Contradiction Detection
- LEGO: Spatial Accelerator Generation and Optimization for Tensor Applications
- Query-Focused Extractive Summarization for Sentiment Explanation
- Neuromorphic Intelligence
- EEND-SAA: Enrollment-Less Main Speaker Voice Activity Detection Using Self-Attention Attractors
- Generative AI in Game Development: A Qualitative Research Synthesis
- Meta-Learning Neural Process for Implied Volatility Surfaces with SABR-induced Priors
- Unrolling Graph-based Douglas-Rachford Algorithm for Image Interpolation with Informed Initialization
- EgoMem: Lifelong Memory Agent for Full-duplex Omnimodal Models
- SugarcaneShuffleNet: A Very Fast, Lightweight Convolutional Neural Network for Diagnosis of 15 Sugarcane Leaf Diseases
- NeuroStrike: Neuron-Level Attacks on Aligned LLMs
- Enhancing Knowledge Tracing through Leakage-Free and Recency-Aware Embeddings
- Detecting Multilevel Manipulation from Limit Order Book via Cascaded Contrastive Representation Learning
- MSMA: Multi-Scale Feature Fusion For Multi-Attribute 3D Face Reconstruction From Unconstrained Images
- Making Judicial Reasoning Visible: Structured Annotation of Holding, Evidentiary Considerations, and Subsumption in Criminal Judgments
- Attention-Enhanced Learning for Sensing-Assisted Long-Term Beam Tracking in mmWave Communications
- Beyond Regularity: Modeling Chaotic Mobility Patterns for Next Location Prediction
- Scaling to Multimodal and Multichannel Heart Sound Classification with Synthetic and Augmented Biosignals
- Dynamic Adaptive Parsing of Temporal and Cross-Variable Patterns for Network State Classification
- POT: Inducing Overthinking in LLMs via Black-Box Iterative Optimization
- Bhaasha, Bhasa, Zaban: A Survey for Low-Resourced Languages in South Asia -- Current Stage and Challenges
- Preservation of Language Understanding Capabilities in Speech-aware Large Language Models
- AdaSTI: Conditional Diffusion Models with Adaptive Dependency Modeling for Spatio-Temporal Imputation
- When marine radar target detection meets pretrained large language models
- Bridging the Gap Between Sparsity and Redundancy: A Dual-Decoding Framework with Global Context for Map Inference
- Diffusion-Based Generation and Imputation of Driving Scenarios from Limited Vehicle CAN Data
- SparseDoctor: Towards Efficient Chat Doctor with Mixture of Experts Enhanced Large Language Models
- TheUse of Conditional Variational Autoencoders in Generating Stellar Spectra
- GRAID: Synthetic Data Generation with Geometric Constraints and Multi-Agentic Reflection for Harmful Content Detection
- Wavelet-SARIMA-Transformer: A Hybrid Model for Rainfall Forecasting
- Joint-octamamba:an octa joint segmentation network based on feature enhanced mamba
- Multimodal Regression for Enzyme Turnover Rates Prediction
- Knowledge Distillation for Sensing-Assisted Long-Term Beam Tracking in mmWave Communications
- Quantum Graph Attention Networks: Trainable Quantum Encoders for Inductive Graph Learning
- Transformer Enhanced Relation Classification: A Comparative Analysis of Contextuality, Data Efficiency and Sequence Complexity
- Toward Next-generation Medical Vision Backbones: Modeling Finer-grained Long-range Visual Dependency
- Holographic Transformers for Complex-Valued Signal Processing: Integrating Phase Interference into Self-Attention
- AlignKT: Explicitly Modeling Knowledge State for Knowledge Tracing with Ideal State Alignment
- Auto-Slides: An Interactive Multi-Agent System for Creating and Customizing Research Presentations
- BERT4beam: Large AI Model Enabled Generalized Beamforming Optimization
- Human Activity Recognition Based on Electrocardiogram Data Only
- California Wildfire Inventory (CAWFI): An Extensive Dataset for Predictive Techniques based on Artificial Intelligence
- DMLDroid: Deep Multimodal Fusion Framework for Android Malware Detection with Resilience to Code Obfuscation and Adversarial Perturbations
- FEWT: Improving Humanoid Robot Perception with Frequency-Enhanced Wavelet-based Transformers
- SelectMix: Enhancing Label Noise Robustness through Targeted Sample Mixing
- Knowledge-Guided Adaptive Mixture of Experts for Precipitation Prediction
- Quantum and Classical Machine Learning in Decentralized Finance: Comparative Evidence from Multi-Asset Backtesting of Automated Market Makers
- GCN-TULHOR: Trajectory-User Linking Leveraging GCNs and Higher-Order Spatial Representations
- Leveraging Geometric Priors for Unaligned Scene Change Detection
- Learned Structure in Cartridges: Keys as Shareable Routers in Self-Studied Representations
- Predictability Enables Parallelization of Nonlinear State Space Models
- GFS: A Preemption-aware Scheduling Framework for GPU Clusters with Predictive Spot Instance Management
- RephraseTTS: Dynamic Length Text based Speech Insertion with Speaker Style Transfer
- CCoMAML: Efficient Cattle Identification Using Cooperative Model-Agnostic Meta-Learning
- Weakly Supervised Vulnerability Localization via Multiple Instance Learning
- UDFS: Lightweight Representation-Driven Open World Robust Encrypted Traffic Classification
- AI-Generated Content in Cross-Domain Applications: Research Trends, Challenges and Propositions
- MIS-LSTM: Multichannel Image-Sequence LSTM for Sleep Quality and Stress Prediction
- Filling the Gaps: A Multitask Hybrid Multiscale Generative Framework for Missing Modality in Remote Sensing Semantic Segmentation
- Forecasting Self-Similar User Traffic Demand Using Transformers in LEO Satellite Networks
- Point-Plane Projections for Accurate LiDAR Semantic Segmentation in Small Data Scenarios
- Why Bonds Fail Differently? Explainable Multimodal Learning for Multi-Class Default Prediction
- OpenHA: A Series of Open-Source Hierarchical Agentic Models in Minecraft
- A Traditional Approach to Symbolic Piano Continuation
- Long Context Automated Essay Scoring with Language Models
- Transformer Networks for Continuous Gravitational-wave Searches
- Reasoning Under Uncertainty: Exploring Probabilistic Reasoning Capabilities of LLMs
- A Survey on Retrieval And Structuring Augmented Generation with Large Language Models
- Kalman Bayesian Transformer
- An Efficient Dual-Line Decoder Network with Multi-Scale Convolutional Attention for Multi-organ Segmentation
- Context Copying Modulation: The Role of Entropy Neurons in Managing Parametric and Contextual Knowledge Conflicts
- Safety and Security Analysis of Large Language Models: Benchmarking Risk Profile and Harm Potential
- GC-VLN: Instruction as Graph Constraints for Training-free Vision-and-Language Navigation
- WhisTLE: Deeply Supervised, Text-Only Domain Adaptation for Pretrained Speech Recognition Transformers
- GraphPPD: Posterior Predictive Modelling for Graph-Level Inference
- Compressed Video Quality Enhancement: Classifying and Benchmarking over Standards
- Developer-LLM Conversations: An Empirical Study of Interactions and Generated Code Quality
- Flow Straight and Fast in Hilbert Space: Functional Rectified Flow
- I-Segmenter: Integer-Only Vision Transformer for Efficient Semantic Segmentation
- ARMA Block: A CNN-Based Autoregressive and Moving Average Module for Long-Term Time Series Forecasting
- ExDoS: Expert-Guided Dual-Focus Cross-Modal Distillation for Smart Contract Vulnerability Detection
- Data-independent Beamforming for End-to-end Multichannel Multi-speaker ASR
- Benchmark of stylistic variation in LLM-generated texts
- A Symmetry-Integrated Approach to Surface Code Decoding
- BenchECG and xECG: a benchmark and baseline for ECG foundation models
- Generating Energy-Efficient Code via Large-Language Models -- Where are we now?
- BEVTraj: Map-Free End-to-End Trajectory Prediction in Bird's-Eye View with Deformable Attention and Sparse Goal Proposals
- Predictive Spike Timing Enables Distributed Shortest Path Computation in Spiking Neural Networks
- Opening the Black Box: Interpretable LLMs via Semantic Resonance Architecture
- Decoding Alignment: A Critical Survey of LLM Development Initiatives through Value-setting and Data-centric Lens
- Few-Part-Shot Font Generation
- Development of Automated Software Design Document Review Methods Using Large Language Models
- Adaptive Token Merging for Efficient Transformer Semantic Communication at the Edge
- Byte by Byte: Unmasking Browser Fingerprinting at the Function Level Using V8 Bytecode Transformers
- LLM Bazaar: A Service Design for Supporting Collaborative Learning with an LLM-Powered Multi-Party Collaboration Infrastructure
- The Hidden Width of Deep ResNets: Tight Error Bounds and Phase Diagram
- Event Camera Guided Visual Media Restoration & 3D Reconstruction: A Survey
- Quantum-Inspired Stacked Integrated Concept Graph Model (QISICGM) for Diabetes Risk Prediction
- OnlineHOI: Towards Online Human-Object Interaction Generation and Perception
- Can we use automated approaches to measure the quality of online political discussion? How to (not) measure interactivity, diversity, rationality, and incivility in online comments to the news
- MCBP: A Memory-Compute Efficient LLM Inference Accelerator Leveraging Bit-Slice-enabled Sparsity and Repetitiveness
- DyKen-Hyena: Dynamic Kernel Generation via Cross-Modal Attention for Multimodal Intent Recognition
- Explainable Fraud Detection with GNNExplainer and Shapley Values
- An Autoencoder and Vision Transformer-based Interpretability Analysis of the Differences in Automated Staging of Second and Third Molars
- DOSA: Differentiable Model-Based One-Loop Search for DNN Accelerators
- Loc2: Interpretable Cross-View Localization via Depth-Lifted Local Feature Matching
- Purge-Gate: Backpropagation-Free Test-Time Adaptation for Point Clouds Classification via Token Purging
- AI Wellbeing
- All for One: LLMs Solve Mental Math at the Last Token With Information Transferred From Other Tokens
- LAVa: Layer-wise KV Cache Eviction with Dynamic Budget Allocation
- Conditioning on PDE Parameters to Generalise Deep Learning Emulation of Stochastic and Chaotic Dynamics
- ObjectReact: Learning Object-Relative Control for Visual Navigation
- Boosting Embodied AI Agents through Perception-Generation Disaggregation and Asynchronous Pipeline Execution
- DualTrack: Sensorless 3D Ultrasound needs Local and Global Context
- Region-Specific Audio Tagging for Spatial Sound
- Combating the Memory Walls: Optimization Pathways for Long-Context Agentic LLM Inference
- Region-Wise Correspondence Prediction between Manga Line Art Images
- Prompt Pirates Need a Map: Stealing Seeds helps Stealing Prompts
- Attention Layers Add Into Low-Dimensional Residual Subspaces
- AquaCast: Urban Water Dynamics Forecasting with Precipitation-Informed Multi-Input Transformer
- Semantic Concentration for Self-Supervised Dense Representations Learning
- D-CAT: Decoupled Cross-Attention Transfer between Sensor Modalities for Unimodal Inference
- Efficient Transformer-Based Piano Transcription With Sparse Attention Mechanisms
- Cross-Domain Evaluation of Transformer-Based Vulnerability Detection on Open & Industry Data
- What You Code Is What We Prove: Translating BLE App Logic into Formal Models with LLMs for Vulnerability Detection
- RENet: Fault-Tolerant Motion Control for Quadruped Robots via Redundant Estimator Networks under Visual Collapse
- Constructing a Question-Answering Simulator through the Distillation of LLMs
- When and How to Express Empathy in Human-Robot Interaction Scenarios
- Vejde: A Framework for Inductive Deep Reinforcement Learning Based on Factor Graph Color Refinement
- DeCodec: Rethinking Audio Codecs as Universal Disentangled Representation Learners
- MGTraj: Multi-Granularity Goal-Guided Human Trajectory Prediction with Recursive Refinement Network
- DeepAries: Adaptive Rebalancing Interval Selection for Enhanced Portfolio Selection
- Dark-ISP: Enhancing RAW Image Processing for Low-Light Object Detection
- MoLEx: Mixture of LoRA Experts in Speech Self-Supervised Models for Audio Deepfake Detection
- Adaptive Pareto-Optimal Token Merging for Edge Transformer Models in Semantic Communication
- Target-oriented Multimodal Sentiment Classification with Counterfactual-enhanced Debiasing
- OCELOT 2023: Cell Detection from Cell-Tissue Interaction Challenge
- Peering Partner Recommendation for ISPs using Machine Learning
- ViRanker: A BGE-M3 & Blockwise Parallel Transformer Cross-Encoder for Vietnamese Reranking
- IRDFusion: Iterative Relation-Map Difference guided Feature Fusion for Multispectral Object Detection
- Enhancing 3D Medical Image Understanding with Pretraining Aided by 2D Multimodal Large Language Models
- LLM Architecture, Scaling Laws, and Economics: A Quick Summary
- WAVE-DETR Multi-Modal Visible and Acoustic Real-Life Drone Detector
- MimicDroid: In-Context Learning for Humanoid Robot Manipulation from Human Play Videos
- LLMs as Agentic Cooperative Players in Multiplayer UNO
- CoDiCodec: Unifying Continuous and Discrete Compressed Representations of Audio
- IMDMR: An Intelligent Multi-Dimensional Memory Retrieval System for Enhanced Conversational AI
- Generative quantum advantage for classical and quantum problems
- Fast attention mechanisms: a tale of parallelism
- When FinTech Meets Privacy: Securing Financial LLMs with Differential Private Fine-Tuning
- iMatcher: Improve matching in point cloud registration via local-to-global geometric consistency learning
- CoSwin: Convolution Enhanced Hierarchical Shifted Window Attention For Small-Scale Vision
- Diffusion-Based Action Recognition Generalizes to Untrained Domains
- MoVoC: Morphology-Aware Subword Construction for Geez Script Languages
- Do All Autoregressive Transformers Remember Facts the Same Way? A Cross-Architecture Analysis of Recall Mechanisms
- Streaming Sequence-to-Sequence Learning with Delayed Streams Modeling
- Universal Graph Learning for Power System Reconfigurations: Transfer Across Topology Variations
- UOPSL: Unpaired OCT Predilection Sites Learning for Fundus Image Diagnosis Augmentation
- Modular PE-Structured Learning for Cross-Task Wireless Communications
- QFrCoLA: a Quebec-French Corpus of Linguistic Acceptability Judgments
- scI2CL: Effectively Integrating Single-cell Multi-omics by Intra- and Inter-omics Contrastive Learning
- Beyond Emotion Recognition: A Multi-Turn Multimodal Emotion Understanding and Reasoning Benchmark
- CNN-ViT Hybrid for Pneumonia Detection: Theory and Empiric on Limited Data without Pretraining
- Improved Receiver Chain Performance via Error Location Inference
- NinA: Normalizing Flows in Action. Training VLA Models with Normalizing Flows
- Vision-Language Semantic Aggregation Leveraging Foundation Model for Generalizable Medical Image Segmentation
- Chirality in Action: Time-Aware Video Representation Learning by Latent Straightening
- Too Helpful, Too Harmless, Too Honest or Just Right?
- Adapting Vision-Language Models for Neutrino Event Classification in High-Energy Physics
- Sparse Transformer for Ultra-sparse Sampled Video Compressive Sensing
- Semantic Causality-Aware Vision-Based 3D Occupancy Prediction
- Predator-Prey Model: Driven Hunt for Accelerated Grokking
- Bitrate-Controlled Diffusion for Disentangling Motion and Content in Video
- AVEC: Bootstrapping Privacy for Local LLMs
- Automatic Detection of Inauthentic Templated Responses in English Language Assessments
- Few-shot Personalization via In-Context Learning for Speech Emotion Recognition based on Speech-Language Model
- Accelerating Mixture-of-Expert Inference with Adaptive Expert Split Mechanism
- Practice on Long Behavior Sequence Modeling in Tencent Advertising
- Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism
- S2Transformer: Scalable Structured Transformers for Global Station Weather Forecasting
- Advancing Few-Shot Pediatric Arrhythmia Classification with a Novel Contrastive Loss and Multimodal Learning
- Dual-Thresholding Heatmaps to Cluster Proposals for Weakly Supervised Object Detection
- ALIGNS: Unlocking nomological networks in psychological measurement through a large language model
- ARTPS: Depth-Enhanced Hybrid Anomaly Detection and Learnable Curiosity Score for Autonomous Rover Target Prioritization
- Physics-Inspired Spatial Temporal Graph Neural Networks for Predicting Industrial Chain Resilience
- ReProCon: Scalable and Resource-Efficient Few-Shot Biomedical Named Entity Recognition
- Towards Scalable and Structured Spatiotemporal Forecasting
- Lightweight Deep Unfolding Networks with Enhanced Robustness for Infrared Small Target Detection
- Generative quantum eigensolver with constrained circuit-cutting overhead
- Recurrence Meets Transformers for Universal Multimodal Retrieval
- Ubiquitous Intelligence Via Wireless Network-Driven LLMs Evolution
- Rollout-LaSDI: Enhancing the long-term accuracy of Latent Space Dynamics
- Selective Induction Heads: How Transformers Select Causal Structures In Context
- Neuromorphic Simulation of Drosophila Melanogaster Brain Connectome on Loihi 2
- Customizing the Inductive Biases of Softmax Attention using Structured Matrices
- One Model for All Tasks: Leveraging Efficient World Models in Multi-Task Planning
- Multimodal Contrastive Pretraining of CBCT and IOS for Enhanced Tooth Segmentation
- ScoreHOI: Physically Plausible Reconstruction of Human-Object Interaction via Score-Guided Diffusion
- Object-level Correlation for Few-Shot Segmentation
- Wind farm layout optimization using a novel machine learning approach
- Query Expansion in the Age of Pre-trained and Large Language Models: A Comprehensive Survey
- A Generalisable Generative Model for Multi-Detector Calorimeter Simulation
- BREATH: A Bio-Radar Embodied Agent for Tonal and Human-Aware Diffusion Music Generation
- E2E Learning Massive MIMO for Multimodal Semantic Non-Orthogonal Transmission and Fusion
- BALI: Enhancing Biomedical Language Representations through Knowledge Graph and Language Model Alignment
- Improving Machine Learning-Based Robot Self-Collision Checking with Input Positional Encoding
- FLeW: Facet-Level and Adaptive Weighted Representation Learning of Scientific Documents
- Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence
- TaDiCodec: Text-aware Diffusion Speech Tokenizer for Speech Language Modeling
- SA-OOSC: A Multimodal LLM-Distilled Semantic Communication Framework for Enhanced Coding Efficiency with Scenario Understanding
- EyeMulator: Improving Code Language Models by Mimicking Human Visual Attention
- Data-Efficient Fine-Tuning of Vision-Language Models for Diagnosis of Alzheimer's Disease
- RINO: Renormalization Group Invariance with No Labels
- AgentX: Towards Orchestrating Robust Agentic Workflow Patterns with FaaS-hosted MCP Services
- Universal Few-Shot Spatial Control for Diffusion Models
- DuoServe-MoE: Dual-Phase Expert Prefetch and Caching for LLM Inference QoS Assurance
- Causal Attention with Lookahead Keys
- Dual Knowledge-Enhanced Two-Stage Reasoner for Multimodal Dialog Systems
- Testing chatbots on the creation of encoders for audio conditioned image generation
- Progressive Facial Granularity Aggregation with Bilateral Attribute-based Enhancement for Face-to-Speech Synthesis
- NOWJ@COLIEE 2025: A Multi-stage Framework Integrating Embedding Models and Large Language Models for Legal Retrieval and Entailment
- Tracing and Mitigating Hallucinations in Multimodal LLMs via Dynamic Attention Localization
- PanoLAM: Large Avatar Model for Gaussian Full-Head Synthesis from One-shot Unposed Image
- In-Context Learning Enhanced Credibility Transformer
- Parse Graph-Based Visual-Language Interaction for Human Pose Estimation
- Mitigating Attention Localization in Small Scale: Self-Attention Refinement via One-step Belief Propagation
- Privacy Preserving Semantic Communications Using Vision Language Models: A Segmentation and Generation Approach
- From Detection to Mitigation: Addressing Gender Bias in Chinese Texts via Efficient Tuning and Voting-Based Rebalancing
- Attention Maps in 3D Shape Classification for Dental Stage Estimation with Class Node Graph Attention Networks
- Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm
- ALICE: An Interpretable Neural Architecture for Generalization in Substitution Ciphers
- Geometric Dynamics of Consumer Credit Cycles: A Multivector-based Linear-Attention Framework for Explanatory Economic Analysis
- The ML-SUPERB 2.0 Challenge: Towards Inclusive ASR Benchmarking for All Language Varieties
- SoK: Security and Privacy of AI Agents for Blockchain
- Short-Term Gaze Prediction: Analysis of Individual Differences, Typical and Extreme-Case Errors
- H2OT: Hierarchical Hourglass Tokenizer for Efficient Video Pose Transformers
- BIR-Adapter: A Low-Complexity Diffusion Model Adapter for Blind Image Restoration
- Beyond Memorization: Extending Reasoning Depth with Recurrence, Memory and Test-Time Compute Scaling
- Barlow-Swin: Toward a novel siamese-based segmentation architecture using Swin-Transformers
- floq: Training Critics via Flow-Matching for Scaling Compute in Value-Based RL
- Green Learning for STAR-RIS mmWave Systems with Implicit CSI
- Zero-shot 3D-Aware Trajectory-Guided image-to-video generation via Test-Time Training
- Probabilistic Modeling of Latent Agentic Substructures in Deep Neural Networks
- Hybrid Swin Attention Networks for Simultaneously Low-Dose PET and CT Denoising
- Integrated Detection and Tracking Based on Radar Range-Doppler Feature
- Lane Change Intention Prediction of two distinct Populations using a Transformer
- QualityFM: a Multimodal Physiological Signal Foundation Model with Self-Distillation for Signal Quality Challenges in Critically Ill Patients
- FSG-Net: Frequency-Spatial Synergistic Gated Network for High-Resolution Remote Sensing Change Detection
- Cross3DReg: Towards a Large-scale Real-world Cross-source Point Cloud Registration Benchmark
- Ban&Pick: Ehancing Performance and Efficiency of MoE-LLMs via Smarter Routing
- RecMind: LLM-Enhanced Graph Neural Networks for Personalized Consumer Recommendations
- AI-driven Remote Facial Skin Hydration and TEWL Assessment from Selfie Images: A Systematic Solution
- RL Fine-Tuning Heals OOD Forgetting in SFT
- Block-Sparse Global Attention for Efficient Multi-View Geometry Transformers
- TrajAware: Graph Cross-Attention and Trajectory-Aware for Generalisable VANETs under Partial Observations
- Learning spatially structured open quantum dynamics with regional-attention transformers
- Continuous Audio Language Models
- IGAff: Benchmarking Adversarial Iterative and Genetic Affine Algorithms on Deep Neural Networks
- Video-Based MPAA Rating Prediction: An Attention-Driven Hybrid Architecture Using Contrastive Learning
- Leveraging Generic Foundation Models for Multimodal Surgical Data Analysis
- Dimensionally Reduced Open-World Clustering: DROWCULA
- Lookup multivariate Kolmogorov-Arnold Networks
- Dato: A Task-Based Programming Model for Dataflow Accelerators
- From Noise to Narrative: Tracing the Origins of Hallucinations in Transformers
- Prototype-Aware Multimodal Alignment for Open-Vocabulary Visual Grounding
- A Spatiotemporal Adaptive Local Search Method for Tracking Congestion Propagation in Dynamic Networks
- Matching Shapes Under Different Topologies: A Topology-Adaptive Deformation Guided Approach
- Neurocognitive Modeling for Text Generation: Deep Learning Architecture for EEG Data
- FASL-Seg: Anatomy and Tool Segmentation of Surgical Scenes
- TinyDef-DETR: A Transformer-Based Framework for Defect Detection in Transmission Lines from UAV Imagery
- Micro-Expression Recognition via Fine-Grained Dynamic Perception
- Motion Aware ViT-based Framework for Monocular 6-DoF Spacecraft Pose Estimation
- Enhancing the Robustness of Contextual ASR to Varying Biasing Information Volumes Through Purified Semantic Correlation Joint Modeling
- ALPHA: LLM-Enabled Active Learning for Human-Free Network Anomaly Detection
- AttriPrompt: Dynamic Prompt Composition Learning for CLIP
- Modeling shopper interest broadness with entropy-driven dialogue policy in the context of arbitrarily large product catalogs
- Challenges in Deep Learning-Based Small Organ Segmentation: A Benchmarking Perspective for Medical Research with Limited Datasets
- From Eigenmodes to Proofs: Integrating Graph Spectral Operators with Symbolic Interpretable Reasoning
- RL Is Neither a Panacea Nor a Mirage: Understanding Supervised vs. Reinforcement Learning Fine-Tuning for LLMs
- Hyperbolic Large Language Models
- Unleashing Hierarchical Reasoning: An LLM-Driven Framework for Training-Free Referring Video Object Segmentation
- Llama-GENBA-10B: A Trilingual Large Language Model for German, English and Bavarian
- Language-guided Recursive Spatiotemporal Graph Modeling for Video Summarization
- TreeGPT: Pure TreeFFN Encoder-Decoder Architecture for Structured Reasoning Without Attention Mechanisms
- QCSE: A Pretrained Quantum Context-Sensitive Word Embedding for Natural Language Processing
- Adaptive Temporal Fusion Transformers for Cryptocurrency Price Prediction
- GraMFedDHAR: Graph Based Multimodal Differentially Private Federated HAR
- SpecPrune-VLA: Accelerating Vision-Language-Action Models via Action-Aware Self-Speculative Pruning
- The Token Tax: Systematic Bias in Multilingual Tokenization
- CURE: Controlled Unlearning for Robust Embeddings -- Mitigating Conceptual Shortcuts in Pre-Trained Language Models
- TripOptimizer: Generative 3D Shape Optimization and Drag Prediction using Triplane VAE Networks
- Graph-Based Spatio-temporal Attention and Multi-Scale Fusion for Clinically Interpretable, High-Fidelity Fetal ECG Extraction
- Recomposer: Event-roll-guided generative audio editing
- Towards Efficient Pixel Labeling for Industrial Anomaly Detection and Localization
- Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper
- Artificial intelligence for representing and characterizing quantum systems
- Scaling Law for Large-Scale Pre-Training Using Chaotic Time Series and Predictability in Financial Time Series
- PLaMo 2 Technical Report
- PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination
- VoltanaLLM: Feedback-Driven Frequency Control and State-Space Routing for Energy-Efficient LLM Serving
- SemSteDiff: Generative Diffusion Model-based Coverless Semantic Steganography Communication
- VARMA-Enhanced Transformer for Time Series Forecasting
- Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts
- Dynamic Group Detection using VLM-augmented Temporal Groupness Graph
- Imitation Learning Based on Disentangled Representation Learning of Behavioral Characteristics
- Reverse Browser: Vector-Image-to-Code Generator
- COGITAO: A Visual Reasoning Framework To Study Compositionality & Generalization
- KRAFT: A Knowledge Graph-Based Framework for Automated Map Conflation
- PRIM: Towards Practical In-Image Multilingual Machine Translation
- XMUspeech Systems for the ASVspoof 5 Challenge
- veScale: Consistent and Efficient Tensor Programming with Eager-Mode SPMD
- The Non-Determinism of Small LLMs: Evidence of Low Answer Consistency in Repetition Trials of Standard Multiple-Choice Benchmarks
- Towards Open World Detection: A Survey
- Phonological Representation Learning for Isolated Signs Improves Out-of-Vocabulary Generalization
- SpikingBrain: Spiking Brain-inspired Large Models
- Bilingual Word Level Language Identification for Omotic Languages
- Hunyuan-MT Technical Report
- VCMamba: Bridging Convolutions with Multi-Directional Mamba for Efficient Visual Representation
- A Node-Aware Dynamic Quantization Approach for Graph Collaborative Filtering
- ML-PWS: Estimating the Mutual Information Between Experimental Time Series Using Neural Networks
- MuST2-Learn: Multi-view Spatial-Temporal-Type Learning for Heterogeneous Municipal Service Time Estimation
- Durian: Dual Reference Image-Guided Portrait Animation with Attribute Transfer
- Manipulating Transformer-Based Models: Controllability, Steerability, and Robust Interventions
- Contextualized Token Discrimination for Speech Search Query Correction
- SSGaussian: Semantic-Aware and Structure-Preserving 3D Style Transfer
- Parking Availability Prediction via Fusing Multi-Source Data with A Self-Supervised Learning Enhanced Spatio-Temporal Inverted Transformer
- A RoBERTa-Based Functional Syntax Annotation Model for Chinese Texts
- Differential Morphological Profile Neural Networks for Semantic Segmentation
- Rethinking the long-range dependency in Mamba/SSM and transformer models
- Sailing Towards Zero-Shot State Estimation using Foundation Models Combined with a UKF
- COBRA: Multimodal Sensing Deep Learning Framework for Remote Chronic Obesity Management via Wrist-Worn Activity Monitoring
- Joint Modeling of Entities and Discourse Relations for Coherence Assessment
- Unobtrusive In-Situ Measurement of Behavior Change by Deep Metric Similarity Learning of Motion Patterns
- Wavelet Fourier Diffuser: Frequency-Aware Diffusion Model for Reinforcement Learning
- Detecting Regional Spurious Correlations in Vision Transformers via Token Discarding
- PracMHBench: Re-evaluating Model-Heterogeneous Federated Learning Based on Practical Edge Device Constraints
- NeuroBreak: Unveil Internal Jailbreak Mechanisms in Large Language Models
- SAC-MIL: Spatial-Aware Correlated Multiple Instance Learning for Histopathology Whole Slide Image Classification
- Expanding Foundational Language Capabilities in Open-Source LLMs through a Korean Case Study
- Multimodal Feature Fusion Network with Text Difference Enhancement for Remote Sensing Change Detection
- LMAE4Eth: Generalizable and Robust Ethereum Fraud Detection by Exploring Transaction Semantics and Masked Graph Embedding
- Attn-Adapter: Attention Is All You Need for Online Few-shot Learner of Vision-Language Model
- Human Motion Video Generation: A Survey
- Vehicle-to-Infrastructure Collaborative Spatial Perception via Multimodal Large Language Models
- NE-PADD: Leveraging Named Entity Knowledge for Robust Partial Audio Deepfake Detection via Attention Aggregation
- Learning Neural Decoding with Parallelism and Self-Coordination for Quantum Error Correction
- Causality-guided Prompt Learning for Vision-language Models via Visual Granulation
- Skywork UniPic 2.0: Building Kontext Model with Online RL for Unified Multimodal Model
- Spoken in Jest, Detected in Earnest: A Systematic Review of Sarcasm Recognition -- Multimodal Fusion, Challenges, and Future Prospects
- Attention is all you need to solve chiral superconductivity
- A Comprehensive Review of Multi-Agent Reinforcement Learning in Video Games
- Efficient Virtuoso: A Latent Diffusion Transformer Model for Goal-Conditioned Trajectory Planning
- The Optimiser Hidden in Plain Sight: Training with the Loss Landscape's Induced Metric
- Warming Up for Zeroth-Order Federated Pre-Training with Low Resource Clients
- Time-Scaling State-Space Models for Dense Video Captioning
- Generative Auto-Bidding in Large-Scale Competitive Auctions via Diffusion Completer-Aligner
- Temporal social network modeling of mobile connectivity data with graph neural networks
- NeuroKoop: Neural Koopman Fusion of Structural-Functional Connectomes for Identifying Prenatal Drug Exposure in Adolescents
- OneSearch: A Preliminary Exploration of the Unified End-to-End Generative Framework for E-commerce Search
- RTGMFF: Enhanced fMRI-based Brain Disorder Diagnosis via ROI-driven Text Generation and Multimodal Feature Fusion
- Tabular foundation model for GEOAI benchmark problems BM/AirportSoilProperties/2/2025
- Temporally-Aware Diffusion Model for Brain Progression Modelling with Bidirectional Temporal Regularisation
- Adaptive KV-Cache Compression without Manually Setting Budget
- PromptCOS: Towards Content-only System Prompt Copyright Auditing for LLMs
- TRELLIS-Enhanced Surface Features for Comprehensive Intracranial Aneurysm Analysis
- MedLiteNet: Lightweight Hybrid Medical Image Segmentation Model
- AR-KAN: Autoregressive-Weight-Enhanced Kolmogorov-Arnold Network for Time Series Forecasting
- Unsupervised Instance Segmentation with Superpixels
- CloudFormer: An Attention-based Performance Prediction for Public Clouds with Unknown Workload
- Advancing Minority Stress Detection with Transformers: Insights from the Social Media Datasets
- S2M2ECG: Spatio-temporal bi-directional State Space Model Enabled Multi-branch Mamba for ECG
- FoMEMO: Towards Foundation Models for Expensive Multi-objective Optimization
- Agentic AI Empowered Multi-UAV Trajectory Optimization in Low-Altitude Economy Networks
- LINKER: Learning Interactions Between Functional Groups and Residues With Chemical Knowledge-Enhanced Reasoning and Explainability
- Multi-level SSL Feature Gating for Audio Deepfake Detection
- Silent Until Sparse: Backdoor Attacks on Semi-Structured Sparsity
- Continuous Saudi Sign Language Recognition: A Vision Transformer Approach
- Unveiling the Response of Large Vision-Language Models to Visually Absent Tokens
- LatPhon: Lightweight Multilingual G2P for Romance Languages and English
- EverTracer: Hunting Stolen Large Language Models via Stealthy and Robust Probabilistic Fingerprint
- H-PRM: A Pluggable Hotword Pre-Retrieval Module for Various Speech Recognition Systems
- Heatmap Guided Query Transformers for Robust Astrocyte Detection across Immunostains and Resolutions
- CARPO: Leveraging Listwise Learning-to-Rank for Context-Aware Query Plan Optimization
- Lattice Annotated Temporal (LAT) Logic for Non-Markovian Reasoning
- Grocery to General Merchandise: A Cross-Pollination Recommender using LLMs and Real-Time Cart Context
- The Transparent Earth: A Multimodal Foundation Model for the Earth's Subsurface
- Attention Mechanism in Randomized Time Warping
- RNN Generalization to Omega-Regular Languages
- Anisotropic Fourier Features for Positional Encoding in Medical Imaging
- HydroGAT: Distributed Heterogeneous Graph Attention Transformer for Spatiotemporal Flood Prediction
- Generative AI for Crystal Structures: A Review
- Motion-Refined DINOSAUR for Unsupervised Multi-Object Discovery
- FLM-Audio: Natural Monologues Improves Native Full-Duplex Chatbots via Dual Training
- Probabilistic Pretraining for Neural Regression
- Efficient Pyramidal Analysis of Gigapixel Images on a Decentralized Modest Computer Cluster
- VASSO: Variance Suppression for Sharpness-Aware Minimization
- Cache Management for Mixture-of-Experts LLMs -- extended version
- Vision encoders should be image size agnostic and task driven
- Earthquake Source Depth Determination using Single Station Waveforms and Deep Learning
- Data-Driven Loss Functions for Inference-Time Optimization in Text-to-Image Generation
- Exploiting Information Redundancy in Attention Maps for Extreme Quantization of Vision Transformers
- JaParaPat: A Large-Scale Japanese-English Parallel Patent Application Corpus
- Assessing Consciousness-Related Behaviors in Large Language Models Using the Maze Test
- Should LLMs be WEIRD? Exploring WEIRDness and Human Rights in Large Language Models
- A Multimodal-Multitask Framework with Cross-modal Relation and Hierarchical Interactive Attention for Semantic Comprehension
- Human-Inspired Soft Anthropomorphic Hand System for Neuromorphic Object and Pose Recognition Using Multimodal Signals
- Speech transformer models for extracting information from baby cries
- An Epidemiological Knowledge Graph extracted from the World Health Organization's Disease Outbreak News
- DaCe AD: Unifying High-Performance Automatic Differentiation for Machine Learning and Scientific Computing
- DivMerge: A divergence-based model merging method for multi-tasking
- IRSAMap:Towards Large-Scale, High-Resolution Land Cover Map Vectorization
- LLMs that Understand Processes: Instruction-tuning for Semantics-Aware Process Mining
- Whisper based Cross-Lingual Phoneme Recognition between Vietnamese and English
- StructCoh: Structured Contrastive Learning for Context-Aware Text Semantic Matching
- MUSE-FM: Multi-task Environment-aware Foundation Model for Wireless Communications
- Exploring Diffusion Models for Generative Forecasting of Financial Charts
- Batch Query Processing and Optimization for Agentic Workflows
- Learnable Loss Geometries with Mirror Descent for Scalable and Convergent Meta-Learning
- Do LLM Modules Generalize? A Study on Motion Generation for Autonomous Driving
- Deep Reinforcement Learning for Drone Route Optimization in Post-Disaster Road Assessment
- Fake & Square: Training Self-Supervised Vision Transformers with Synthetic Data and Synthetic Hard Negatives
- OmniCache: A Trajectory-Oriented Global Perspective on Training-Free Cache Reuse for Diffusion Transformer Models
- Artificial intelligence in headache medicine: between automation and the doctor-patient relationship. A systematic review
- Unsupervised Training of Vision Transformers with Synthetic Negatives
- Generative Sequential Notification Optimization via Multi-Objective Decision Transformers
- Discrete Noise Inversion for Next-scale Autoregressive Text-based Image Editing
- Optimizing Paths for Adaptive Fly-Scan Microscopy: An Extended Version
- AudioRWKV: Efficient and Stable Bidirectional RWKV for Audio Pattern Recognition
- CellPainTR: Generalizable Representation Learning for Cross-Dataset Cell Painting Analysis
- FuXi-TC: A generative framework integrating deep learning and physics-based models for improved tropical cyclone forecasts
- On the Collapse Errors Induced by the Deterministic Sampler for Diffusion Models
- SSVD: Structured SVD for Parameter-Efficient Fine-Tuning and Benchmarking under Domain Shift in ASR
- GradES: Significantly Faster Training in Transformers with Gradient-Based Early Stopping
- Hardwired-Neurons Language Processing Units as General-Purpose Cognitive Substrates
- chDzDT: Word-level morphology-aware language model for Algerian social media text
- Learning to Ask: Decision Transformers for Adaptive Quantitative Group Testing
- Entropy-Driven Curriculum for Multi-Task Training in Human Mobility Prediction
- TransForSeg: A Multitask Stereo ViT for Joint Stereo Segmentation and 3D Force Estimation in Catheterization
- Learning Longitudinal Stress Dynamics from Irregular Self-Reports via Time Embeddings
- CommonKV: Compressing KV Cache with Cross-layer Parameter Sharing
- From Noise to Precision: A Diffusion-Driven Approach to Zero-Inflated Precipitation Prediction
- IR-Agent: Expert-Inspired LLM Agents for Structure Elucidation from Infrared Spectra
- ArabEmoNet: A Lightweight Hybrid 2D CNN-BiLSTM Model with Attention for Robust Arabic Speech Emotion Recognition
- CbLDM: A Diffusion Model for recovering nanostructure from atomic pair distribution function
- The AudioMOS Challenge 2025
- Multitask Battery Management with Flexible Pretraining
- Towards More Diverse and Challenging Pre-training for Point Cloud Learning: Self-Supervised Cross Reconstruction with Decoupled Views
- LobRA: Multi-tenant Fine-tuning over Heterogeneous Data
- StoxLSTM: A Stochastic Extended Long Short-Term Memory Network for Time Series Forecasting
- ADMP-GNN: Adaptive Depth Message Passing GNN
- MVTrajecter: Multi-View Pedestrian Tracking with Trajectory Motion Cost and Trajectory Appearance Cost
- FICGen: Frequency-Inspired Contextual Disentanglement for Layout-driven Degraded Image Generation
- An End-to-End Framework for Video Multi-Person Pose Estimation
- Lightening the Load: A Cluster-Based Framework for A Lower-Overhead, Provable Website Fingerprinting Defense
- Neural Scene Designer: Self-Styled Semantic Image Manipulation
- TransGAT: Transformer-Based Graph Neural Networks for Multi-Dimensional Automated Essay Scoring
- Optimal Parallel Scheduling under Concave Speedup Functions
- Re3: Learning to Balance Relevance & Recency for Temporal Information Retrieval
- Collaborative local–global context modeling for session-based recommendation
- Leveraging learned representations and multitask learning for lysine methylation site discovery
- Bidirectional Sparse Attention for Faster Video Diffusion Training
- Model Unmerging: Making Your Models Unmergeable for Secure Model Sharing
- Less Redundancy: Boosting Practicality of Vision Language Model in Walking Assistants
- A Unified Voxel Diffusion Module for Point Cloud 3D Object Detection
- Enabling 6G Through Multi-Domain Channel Extrapolation: Opportunities and Challenges of Generative Artificial Intelligence
- A Multimodal Deep Learning Framework for Early Diagnosis of Liver Cancer via Optimized BiLSTM-AM-VMD Architecture
- Multi-vessel Interaction-Aware Trajectory Prediction and Collision Risk Assessment
- Ultra Fast Warm Start Solution for Graph Recommendations
- AttnBoost: Retail Supply Chain Sales Insights via Gradient Boosting Perspective
- Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation
- A Multi-target Bayesian Transformer Framework for Predicting Cardiovascular Disease Biomarkers during Pandemics
- A Continuous-Time Consistency Model for 3D Point Cloud Generation
- Advances and Trends in the 3D Reconstruction of the Shape and Motion of Animals
- Generative Foundation Model for Structured and Unstructured Electronic Health Records
- Enhancing Large Language Model for Knowledge Graph Completion via Structure-Aware Alignment-Tuning
- Causal Sensitivity Identification using Generative Learning
- Enhancing Partially Relevant Video Retrieval with Robust Alignment Learning
- Mic Drop or Data Flop? Evaluating the Fitness for Purpose of AI Voice Interviewers for Data Collection within Quantitative & Qualitative Research Contexts
- Wrong Model, Right Uncertainty: Spatial Associations for Discrete Data with Misspecification
- Deep Learning-Based Rock Particulate Classification Using Attention-Enhanced ConvNeXt
- LongCat-Flash Technical Report
- Fluid Antenna Port Prediction based on Large Language Models
- REVELIO -- Universal Multimodal Task Load Estimation for Cross-Domain Generalization
- AHAMask: Reliable Task Specification for Large Audio Language Models without Instructions
- Efficient Large Language Models with Zero-Shot Adjustable Acceleration
- Learning Correlation-aware Aleatoric Uncertainty for 3D Hand Pose Estimation
- QueryBandits for Hallucination Mitigation: Exploiting Semantic Features for No-Regret Rewriting
- A case study of forensic psychiatry experts' reports analysis through large language models
- MEPT: Mixture of Expert Prompt Tuning as a Manifold Mapper
- LLM-Guided Semantic Relational Reasoning for Multimodal Intent Recognition
- Optimal Dynamic Regret by Transformers for Non-Stationary Reinforcement Learning
- Imputing Missing Long-Term Spatiotemporal Multivariate Atmospheric Data with CNN-Transformer Machine Learning
- Testing the assumptions about the geometry of sentence embedding spaces: the cosine measure need not apply
- SpecVLM: Enhancing Speculative Decoding of Video LLMs via Verifier-Guided Token Pruning
- Prospects of Imitating Trading Agents in the Stock Market
- Self-Exploring Language Models for Explainable Link Forecasting on Temporal Graphs via Reinforcement Learning
- SCOUT: Toward Sub-Quadratic Attention via Segment Compression for Optimized Utility in Transformers
- DeepSeasons: a Deep Learning scale-selecting approach to Seasonal Forecasts
- Supervised In-Context Fine-Tuning for Generative Sequence Labeling
- DarkVRAI: Capture-Condition Conditioning and Burst-Order Selective Scan for Low-light RAW Video Denoising
- SeLeRoSa: Sentence-Level Romanian Satire Detection Dataset
- Optical Music Recognition of Jazz Lead Sheets
- Exploring Over-stationarization in Deep Learning-based Bus/Tram Arrival Time Prediction: Analysis and Non-stationary Effect Recovery
- MarkSplatter: Generalizable Watermarking for 3D Gaussian Splatting Model via Splatter Image Structure
- Multi-Level CLS Token Fusion for Contrastive Learning in Endoscopy Image Classification
- Enhancing Fairness in Skin Lesion Classification for Medical Diagnosis Using Prune Learning
- Efficient Graph Understanding with LLMs via Structured Context Injection
- Resting-state fMRI Analysis using Quantum Time-series Transformer
- Designing LMS and Instructional Strategies for Integrating Generative-Conversational AI
- Why Pool When You Can Flow? Active Learning with GFlowNets
- CSFMamba: Cross State Fusion Mamba Operator for Multimodal Remote Sensing Image Classification
- Automatic Identification and Description of Jewelry Through Computer Vision and Neural Networks for Translators and Interpreters
- IndiaWeatherBench: A Dataset and Benchmark for Data-Driven Regional Weather Forecasting over India
- Missing Data Imputation using Neural Cellular Automata
- MV-SSM: Multi-View State Space Modeling for 3D Human Pose Estimation
- Speaker-Conditioned Phrase Break Prediction for Text-to-Speech with Phoneme-Level Pre-trained Language Model
- Neuro-Symbolic Predictive Process Monitoring
- UrbanInsight: A Distributed Edge Computing Framework with LLM-Powered Data Filtering for Smart City Digital Twins
- Predicting Multi-Type Talented Students in Secondary School Using Semi-Supervised Machine Learning
- Self-supervised neural operator for solving partial differential equations
- AImoclips: A Benchmark for Evaluating Emotion Conveyance in Text-to-Music Generation
- CaresAI at BioCreative IX Track 1 -- LLM for Biomedical QA
- Image-to-Brain Signal Generation for Visual Prosthesis with CLIP Guided Multimodal Diffusion Models
- DTRNet: Dynamic Token Routing Network to Reduce Quadratic Costs in Transformers
- NMR-Solver: Automated Structure Elucidation via Large-Scale Spectral Matching and Physics-Guided Fragment Optimization
- COMET: A Framework for Modeling Compound Operation Dataflows with Explicit Collectives
- Attention of a Kiss: Exploring Attention Maps in Video Diffusion for XAIxArts
- KVComp: A High-Performance, LLM-Aware, Lossy Compression Framework for KV Cache
- Cross-device Zero-shot Label Transfer via Alignment of Time Series Foundation Model Embeddings
- Modeling Long-term User Behaviors with Diffusion-driven Multi-interest Network for CTR Prediction
- Graph Convolutional Network With Pattern-Spatial Interactive and Regional Awareness for Traffic Forecasting
- LLM-Assisted Iterative Evolution with Swarm Intelligence Toward SuperBrain
- TRUST: Token-dRiven Ultrasound Style Transfer for Cross-Device Adaptation
- Universal Properties of Activation Sparsity in Modern Large Language Models
- A Deep Learning Framework for Joint Channel Acquisition and Communication Optimization in Movable Antenna Systems
- Memory Limitations of Prompt Tuning in Transformers
- Metis: Training LLMs with FP4 Quantization
- Deep Learning for Personalized Binaural Audio Reproduction
- A Study on the Framework for Evaluating the Ethics and Trustworthiness of Generative AI
- HERO-VQL: Hierarchical, Egocentric and Robust Visual Query Localization
- Visually Grounded Narratives: Reducing Cognitive Burden in Researcher-Participant Interaction
- Adaptive Point-Prompt Tuning: Fine-Tuning Heterogeneous Foundation Models for 3D Point Cloud Analysis
- TrimTokenator: Towards Adaptive Visual Token Pruning for Large Multimodal Models
- CoMET: A Contrastive-Masked Brain Foundation Model for Universal EEG Representation
- Gated Associative Memory: A Parallel O(N) Architecture for Efficient Sequence Modeling
- LLM-Driven Policy Diffusion: Enhancing Generalization in Offline Reinforcement Learning
- TimeCopilot
- Attribute Filtering in Approximate Nearest Neighbor Search: An In-depth Experimental Study
- Theory Foundation of Physics-Enhanced Residual Learning
- SHERPA: A Model-Driven Framework for Large Language Model Execution
- Quantum-Optimized Selective State Space Model for Efficient Time Series Prediction
- Standard vs. Modular Sampling: Best Practices for Reliable LLM Unlearning
- Simulation-based inference of yeast centromeres
- MMQ: Multimodal Mixture-of-Quantization Tokenization for Semantic ID Generation and User Behavioral Adaptation
- Category-level Text-to-Image Retrieval Improved: Bridging the Domain Gap with Diffusion Models and Vision Encoders
- VoCap: Video Object Captioning and Segmentation from Any Prompt
- Scaling Legal AI: Benchmarking Mamba and Transformers for Statutory Classification and Case Law Retrieval
- Learning Unified Representations from Heterogeneous Data for Robust Heart Rate Modeling
- Benchmarking GPT-5 in Radiation Oncology: Measurable Gains, but Persistent Need for Expert Oversight
- Why Stop at Words? Unveiling the Bigger Picture through Line-Level OCR
- Challenges and Applications of Large Language Models: A Comparison of GPT and DeepSeek family of models
- Efficient Diffusion-Based 3D Human Pose Estimation with Hierarchical Temporal Pruning
- Comparative Analysis of Large Language Models for the Machine-Assisted Resolution of User Intentions
- Spiking Decision Transformers: Local Plasticity, Phase-Coding, and Dendritic Routing for Low-Power Sequence Control
- ELV-Halluc: Benchmarking Semantic Aggregation Hallucinations in Long Video Understanding
- RepoMark: A Data-Usage Auditing Framework for Code Large Language Models
- DRASP: A Dual-Resolution Attentive Statistics Pooling Framework for Automatic MOS Prediction
- VeriLoRA: Fine-Tuning Large Language Models with Verifiable Security via Zero-Knowledge Proofs
- Quantum-Enhanced Natural Language Generation: A Multi-Model Framework with Hybrid Quantum-Classical Architectures
- Stage-Diff: Stage-wise Long-Term Time Series Generation Based on Diffusion Models
- A Financial Brain Scan of the LLM
- Scalable Equilibrium Propagation via Intermediate Error Signals for Deep Convolutional CRNNs
- Guess-and-Learn (G&L): Measuring the Cumulative Error Cost of Cold-Start Adaptation
- DMGIN: How Multimodal LLMs Enhance Large Recommendation Models for Lifelong User Post-click Behaviors
- AI Compute Architecture and Evolution Trends
- The Complexity Trap: Simple Observation Masking Is as Efficient as LLM Summarization for Agent Context Management
- Next Point-of-interest (POI) Recommendation Model Based on Multi-modal Spatio-temporal Context Feature Embedding
- From Canonical to Complex: Benchmarking LLM Capabilities in Undergraduate Thermodynamics
- Multi-Ontology Integration with Dual-Axis Propagation for Medical Concept Representation
- Maybe you don't need a U-Net: convolutional feature upsampling for materials micrograph segmentation
- Continuous Determination of Respiratory Rate in Hospitalized Patients using Machine Learning Applied to Electrocardiogram Telemetry
- Large Language Model Integration with Reinforcement Learning to Augment Decision-Making in Autonomous Cyber Operations
- Full-Frequency Temporal Patching and Structured Masking for Enhanced Audio Classification
- Decoding Memories: An Efficient Pipeline for Self-Consistency Hallucination Detection
- GCAV: A Global Concept Activation Vector Framework for Cross-Layer Consistency in Interpretability
- Deep Residual Echo State Networks: exploring residual orthogonal connections in untrained Recurrent Neural Networks
- InSQuAD: In-Context Learning for Efficient Retrieval via Submodular Mutual Information to Enforce Quality and Diversity
- Graph-Based Feature Augmentation for Predictive Tasks on Relational Datasets
- Scaling Neuro-symbolic Problem Solving: Solver-Free Learning of Constraints and Objectives
- E-ConvNeXt: A Lightweight and Efficient ConvNeXt Variant with Cross-Stage Partial Connections
- Meta-learning ecological priors from large language models explains human learning and decision making
- ATM-GAD: Adaptive Temporal Motif Graph Anomaly Detection for Financial Transaction Networks
- Multi-Lingual Implicit Discourse Relation Recognition with Multi-Label Hierarchical Learning
- Leveraging Discriminative Latent Representations for Conditioning GAN-Based Speech Enhancement
- SKGE-SWIN: End-To-End Autonomous Vehicle Waypoint Prediction and Navigation Using Skip Stage Swin Transformer
- Bridging Minds and Machines: Toward an Integration of AI and Cognitive Science
- Self-Composing Neural Operators with Depth and Accuracy Scaling via Adaptive Train-and-Unroll Approach
- Masked Autoencoders for Ultrasound Signals: Robust Representation Learning for Downstream Applications
- Trained Miniatures: Low cost, High Efficacy SLMs for Sales & Marketing
- Local Virtual Nodes for Alleviating Over-Squashing in Graph Neural Networks
- SemSR: Semantics aware robust Session-based Recommendations
- MERIT: Maximum-normalized Element-wise Ratio for Language Model Large-batch Training
- KCS: Diversify Multi-hop Question Generation with Knowledge Composition Sampling
- Dual-Stage Global and Local Feature Framework for Image Dehazing
- Molecular Machine Learning in Chemical Process Design
- Structure-aware Hypergraph Transformer for Diagnosis Prediction in Electronic Health Records
- Towards Inclusive Communication: A Unified Framework for Generating Spoken Language from Sign, Lip, and Audio
- Prediction of mortality and resource utilization in critical care: a deep learning approach using multimodal electronic health records with natural language processing techniques
- MSMVD: Exploiting Multi-scale Image Features via Multi-scale BEV Features for Multi-view Pedestrian Detection
- Uncovering the Spectral Bias in Diagonal State Space Models
- AI-SearchPlanner: Modular Agentic Search via Pareto-Optimal Multi-Objective Reinforcement Learning
- TrInk: Ink Generation with Transformer Network
- Understanding Incremental Learning with Closed-form Solution to Gradient Flow on Overparamerterized Matrix Factorization
- Mixture of Contexts for Long Video Generation
- MindGuard: Intrinsic Decision Inspection for Securing LLM Agents Against Metadata Poisoning
- STGAtt: A Spatial-Temporal Unified Graph Attention Network for Traffic Flow Forecasting
- Tutorial on the Probabilistic Unification of Estimation Theory, Machine Learning, and Generative AI
- Compositionality in Time Series: A Proof of Concept using Symbolic Dynamics and Compositional Data Augmentation
- End-to-End Analysis of Charge Stability Diagrams with Transformers
- OneRec-V2 Technical Report
- Sound event detection with audio-text models and heterogeneous temporal annotations
- SciTopic: Enhancing Topic Discovery in Scientific Literature through Advanced LLM
- Provable Benefits of In-Tool Learning for Large Language Models
- Tree-like Pairwise Interaction Networks
- Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning
- GUARD: Glocal Uncertainty-Aware Robust Decoding for Effective and Efficient Open-Ended Text Generation
- Beyond Transcription: Mechanistic Interpretability in ASR
- MobileCLIP2: Improving Multi-Modal Reinforced Training
- Neural Field Turing Machine: A Differentiable Spatial Computer
- Dynamics-Aligned Latent Imagination in Contextual World Models for Zero-Shot Generalization
- A Survey of Affective Recommender Systems: Modeling Attitudes, Emotions, and Moods for Personalization
- A Systematic Review on the Generative AI Applications in Human Medical Genomics
- Robustness Assessment and Enhancement of Text Watermarking for Google's SynthID
- FlowMalTrans: Unsupervised Binary Code Translation for Malware Detection Using Flow-Adapter Architecture
- What can we learn from signals and systems in a transformer? Insights for probabilistic modeling and inference architecture
- ProMSC-MIS: Prompt-based Multimodal Semantic Communication for Multi-Spectral Image Segmentation
- Visio-Verbal Teleimpedance Interface: Enabling Semi-Autonomous Control of Physical Interaction via Eye Tracking and Speech
- Automatic Pronunciation Error Detection and Correction of the Holy Quran's Learners Using Deep Learning
- ASCMamba: Multimodal Time-Frequency Mamba for Acoustic Scene Classification
- DCHO: A Decomposition-Composition Framework for Predicting Higher-Order Brain Connectivity to Enhance Diverse Downstream Applications
- SoK: Large Language Model Copyright Auditing via Fingerprinting
- FARM: Frame-Accelerated Augmentation and Residual Mixture-of-Experts for Physics-Based High-Dynamic Humanoid Control
- FlowletFormer: Network Behavioral Semantic Aware Pre-training Model for Traffic Classification
- The Next Layer: Augmenting Foundation Models with Structure-Preserving and Attention-Guided Learning for Local Patches to Global Context Awareness in Computational Pathology
- PersonaAnimator: Personalized Motion Transfer from Unconstrained Videos
- TrajFusionNet: Pedestrian Crossing Intention Prediction via Fusion of Sequential and Visual Trajectory Representations
- RATopo: Improving Lane Topology Reasoning via Redundancy Assignment
- The Art of Hide and Seek: Making Pickle-Based Model Supply Chain Poisoning Stealthy Again
- QuesGenie: Intelligent Multimodal Question Generation
- Tune My Adam, Please!
- Q-Align: Alleviating Attention Leakage in Zero-Shot Appearance Transfer via Query-Query Alignment
- Improving Generalization in Deepfake Detection with Face Foundation Models and Metric Learning
- An Effective Strategy for Modeling Score Ordinality and Non-uniform Intervals in Automated Speaking Assessment
- SPELUNKER: Item Similarity Search Using Large Language Models and Custom K-Nearest Neighbors
- Hybrid Decoding: Rapid Pass and Selective Detailed Correction for Sequence Models
- FinCast: A Foundation Model for Financial Time-Series Forecasting
- Understanding and Leveraging the Expert Specialization of Context Faithfulness in Mixture-of-Experts LLMs
- FlowDet: Overcoming Perspective and Scale Challenges in Real-Time End-to-End Traffic Detection
- Robustness is Important: Limitations of LLMs for Data Fitting
- Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference
- CSRD2025: A Large-Scale Synthetic Radio Dataset for Spectrum Sensing in Wireless Communications
- Emotion Transfer with Enhanced Prototype for Unseen Emotion Recognition in Conversation
- FLASepformer: Efficient Speech Separation with Gated Focused Linear Attention Transformer
- UNIFORM: Unifying Knowledge from Large-scale and Diverse Pre-trained Models
- Towards 6G Intelligence: The Role of Generative AI in Future Wireless Networks
- Fast 3D Diffusion for Scalable Granular Media Synthesis
- ERTACache: Error Rectification and Timesteps Adjustment for Efficient Diffusion
- ELIXIR: Efficient and LIghtweight model for eXplaIning Recommendations
- Task-Generalized Adaptive Cross-Domain Learning for Multimodal Image Fusion
- ScanMove: Motion Prediction and Transfer for Unregistered Body Meshes
- Enabling Transparent Cyber Threat Intelligence Combining Large Language Models and Domain Ontologies
- Stack Trace-Based Crash Deduplication with Transformer Adaptation
- On Surjectivity of Neural Networks: Can you elicit any behavior from your model?
- Two-dimensional electronic spectra from trajectory-based dynamics: pure-state Ehrenfest, spin-mapping, and mean classical path approaches
- Fine-Tuning Vision-Language Models for Neutrino Event Analysis in High-Energy Physics Experiments
- Amplifying Emotional Signals: Data-Efficient Deep Learning for Robust Speech Emotion Recognition
- Mitigating Clinician Information Overload: Generative AI for Integrated EHR and RPM Data Analysis
- VibeVoice Technical Report
- Grounding the Ungrounded: A Spectral-Graph Framework for Quantifying Hallucinations in Multimodal LLMs
- EffNetViTLoRA: An Efficient Hybrid Deep Learning Approach for Alzheimer's Disease Diagnosis
- Articulate3D: Zero-Shot Text-Driven 3D Object Posing
- Autoregressive Universal Video Segmentation Model
- SynthCoder: A Synthetical Strategy to Tune LLMs for Code Completion
- Attention-Based Explainability for Structure-Property Relationships
- JEDI-linear: Fast and Efficient Graph Neural Networks for Jet Tagging on FPGAs
- FastMesh: Efficient Artistic Mesh Generation via Component Decoupling
- Machine Learning Free Quotients of CICYs
- Federated Fine-Tuning of Sparsely-Activated Large Language Models on Resource-Constrained Devices
- ArgRAG: Explainable Retrieval Augmented Generation using Quantitative Bipolar Argumentation
- RoofSeg: An edge-aware transformer-based network for end-to-end roof plane segmentation
- MOSA: Mixtures of Simple Adapters Outperform Monolithic Approaches in LLM-based Multilingual ASR
- Enhancing compact convolutional transformers with super attention
- On the Generalisation of Koopman Representations for Chaotic System Control
- Novel Approaches to Artificial Intelligence Development Based on the Nearest Neighbor Method
- The GINN framework: a stochastic QED correspondence for stability and chaos in deep neural networks
- HierCVAE: Hierarchical Attention-Driven Conditional Variational Autoencoders for Multi-Scale Temporal Modeling
- Distance-informed Neural Processes
- Aligning Moments in Time using Video Queries
- DQEN: Dual Query Enhancement Network for DETR-based HOI Detection
- ClusterFusion: Expanding Operator Fusion Scope for LLM Inference via Cluster-Level Collective Primitive
- Boosting Micro-Expression Analysis via Prior-Guided Video-Level Regression
- Alljoined-1.6M: A Million-Trial EEG-Image Dataset for Evaluating Affordable Brain-Computer Interfaces
- EMind: A Foundation Model for Multi-task Electromagnetic Signals Understanding
- Insights into User Interface Innovations from a Design Thinking Workshop at deRSE25
- UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning
- M3HG: Multimodal, Multi-scale, and Multi-type Node Heterogeneous Graph for Emotion Cause Triplet Extraction in Conversations
- Drawing2CAD: Sequence-to-Sequence Learning for CAD Generation from Vector Drawings
- Beyond Tokens: Enhancing RTL Quality Estimation via Structural Graph Learning
- Class-wise Flooding Regularization for Imbalanced Image Classification
- LLM as an Execution Estimator: Recovering Missing Dependency for Practical Time-travelling Debugging
- Hierarchical Spatio-temporal Segmentation Network for Ejection Fraction Estimation in Echocardiography Videos
- Optimal Sparsity of Mixture-of-Experts Language Models for Reasoning Tasks
- The Sound of Risk: A Multimodal Physics-Informed Acoustic Model for Forecasting Market Volatility and Enhancing Market Interpretability
- ROSE: Remove Objects with Side Effects in Videos
- A New NMT Model for Translating Clinical Texts from English to Spanish
- What do language models model? Transformers, automata, and the format of thought
- Revisiting associative recall in modern recurrent models
- DIO: Refining Mutual Information and Causal Chain to Enhance Machine Abstract Reasoning Ability
- Improving Long-term Autoregressive Spatiotemporal Predictions: A Proof of Concept with Fluid Dynamics
- The Quasi-Creature and the Uncanny Valley of Agency: A Synthesis of Theory and Evidence on User Interaction with Inconsistent Generative AI
- EvoFormer: Learning Dynamic Graph-Level Representations with Structural and Temporal Bias Correction
- Integrating gender inclusivity into large language models via instruction tuning
- Vectorized Attention with Learnable Encoding for Quantum Transformer
- From Prediction to Simulation: AlphaFold 3 as a Differentiable Framework for Structural Biology
- LLM-Driven Intrinsic Motivation for Sparse Reward Reinforcement Learning
- Integral Transformer: Denoising Attention, Not Too Much Not Too Little
- DemoBias: An Empirical Study to Trace Demographic Biases in Vision Foundation Models
- MMTok: Multimodal Coverage Maximization for Efficient Inference of VLMs
- FSA: An Alternative Efficient Implementation of Native Sparse Attention Kernel
- Explain and Monitor Deep Learning Models for Computer Vision using Obz AI
- Frozen in Time: Parameter-Efficient Time Series Transformers via Reservoir-Induced Feature Expansion and Fixed Random Dynamics
- An Introduction to Silent Paralinguistics
- WISCA: A Lightweight Model Transition Method to Improve LLM Training via Weight Scaling
- Exploring Scaling Laws of CTR Model for Online Performance Improvement
- Training Transformers for Mesh-Based Simulations
- AQ-PCDSys: An Adaptive Quantized Planetary Crater Detection System for Autonomous Space Exploration
- Robust and Efficient Quantum Reservoir Computing with Discrete Time Crystal
- Dream 7B: Diffusion Large Language Models
- An Empirical Study of Knowledge Distillation for Code Understanding Tasks
- Curriculum Approximate Unlearning for Session-based Recommendation
- Survey of Vision-Language-Action Models for Embodied Manipulation
- Multi-Object Sketch Animation with Grouping and Motion Trajectory Priors
- CUPE: Contextless Universal Phoneme Encoder for Language-Agnostic Speech Processing
- Interface on demand: Towards AI native Control interfaces for 6G
- Towards Scalable and Interpretable Mobile App Risk Analysis via Large Language Models
- An Empirical Study on How Video-LLMs Answer Video Questions
- Efficient Identification of Critical Transitions via Flow Matching: A Scalable Generative Approach for Many-Body Systems
- Image-Conditioned 3D Gaussian Splat Quantization
- Measuring the environmental impact of delivering AI at Google Scale
- Frequency-adaptive tensor neural networks for high-dimensional multi-scale problems
- Discovering Hidden Algebraic Structures via Transformers with Rank-Aware Beam GRPO
- DesignCLIP: Multimodal Learning with CLIP for Design Patent Understanding
- Sensing, Social, and Motion Intelligence in Embodied Navigation: A Comprehensive Survey
- Lorentz-Equivariance without Limitations
- Graph Structure Learning with Temporal Graph Information Bottleneck for Inductive Representation Learning
- EventSSEG: Event-driven Self-Supervised Segmentation with Probabilistic Attention
- Leveraging Hardware-Aware Computation in Mixed-Precision Matrix Multiply: A Tile-Centric Approach
- Improving LLMs for Machine Translation Using Synthetic Preference Data
- Beyond Individuals: Collective Predictive Coding for Memory, Attention, and the Emergence of Language
- DualNILM: Energy Injection Identification Enabled Disaggregation with Deep Multi-Task Learning
- Reliable Smoke Detection via Optical Flow-Guided Feature Fusion and Transformer-Based Uncertainty Modeling
- Large Foundation Model for Ads Recommendation
- Controllable Latent Space Augmentation for Digital Pathology
- Improving OCR using internal document redundancy
- Machine learning revolution for exoplanet direct imaging detection: transformer architectures
- Artificial Intelligence-Based Multiscale Temporal Modeling for Anomaly Detection in Cloud Services
- DGenCTR: Towards a Universal Generative Paradigm for Click-Through Rate Prediction via Discrete Diffusion
- Global-Distribution Aware Scenario-Specific Variational Representation Learning Framework
- WeedSense: Multi-Task Learning for Weed Segmentation, Height Estimation, and Growth Stage Classification
- Distribution-Guided Auto-Encoder for User Multimodal Interest Cross Fusion
- Vivid-VR: Distilling Concepts from Text-to-Video Diffusion Transformer for Photorealistic Video Restoration
- A Laplace diffusion-based transformer model for heart rate forecasting within daily activity context
- Generalizable Engagement Estimation in Conversation via Domain Prompting and Parallel Attention
- MoCHA-former: Moiré-Conditioned Hybrid Adaptive Transformer for Video Demoiréing
- You Only Evaluate Once: A Tree-based Rerank Method at Meituan
- Taming Transformer for Emotion-Controllable Talking Face Generation
- HandCraft: Dynamic Sign Generation for Synthetic Data Augmentation
- Neural Network Quantization for Microcontrollers: A Comprehensive Survey of Methods, Platforms, and Applications
- Tinker: Diffusion's Gift to 3D--Multi-View Consistent Editing From Sparse Inputs without Per-Scene Optimization
- Making Pose Representations More Expressive and Disentangled via Residual Vector Quantization
- The Digital Sous Chef -- A Comparative Study on Fine-Tuning Language Models for Recipe Generation
- DeepEmoNet: Building Machine Learning Models for Automatic Emotion Recognition in Human Speeches
- The Prompting Brain: Neurocognitive Markers of Expertise in Guiding Large Language Models
- LLMs and Agentic AI in Insurance Decision-Making: Opportunities and Challenges For Africa
- Reconstruction Using the Invisible: Intuition from NIR and Metadata for Enhanced 3D Gaussian Splatting
- In-Context Iterative Policy Improvement for Dynamic Manipulation
- Seeing Further on the Shoulders of Giants: Knowledge Inheritance for Vision Foundation Models
- MoEcho: Exploiting Side-Channel Attacks to Compromise User Privacy in Mixture-of-Experts LLMs
- MultiFuzz: A Dense Retrieval-based Multi-Agent System for Network Protocol Fuzzing
- Pixels to Play: A Foundation Model for 3D Gameplay
- GALA: Guided Attention with Language Alignment for Open Vocabulary Gaussian Splatting
- Comparing energy consumption and accuracy in text classification inference
- StarStream: Live Video Analytics over Space Networking
- Reliability comparison of vessel trajectory prediction models via Probability of Detection
- Bites of Tomorrow: Personalized Recommendations for a Healthier and Greener Plate
- Ask Good Questions for Large Language Models
- Democratizing News Recommenders: Modeling Multiple Perspectives for News Candidate Generation with VQ-VAE
- DIME-Net: A Dual-Illumination Adaptive Enhancement Network Based on Retinex and Mixture-of-Experts
- LLMind 2.0: Distributed IoT Automation with Natural Language M2M Communication and Lightweight LLM Agents
- Multimodal Data Storage and Retrieval for Embodied AI: A Survey
- Distributed Distortion-Aware Robust Optimization for Movable Antenna-aided Cell-Free ISAC Systems
- Energy Management and Wake-up for IoT Networks Powered by Energy Harvesting
- A Fully Transformer Based Multimodal Framework for Explainable Cancer Image Segmentation Using Radiology Reports
- PENGUIN: Enhancing Transformer with Periodic-Nested Group Attention for Long-term Time Series Forecasting
- Refining Contrastive Learning and Homography Relations for Multi-Modal Recommendation
- Eliminating Rasterization: Direct Vector Floor Plan Generation with DiffPlanner
- Trans-XFed: An Explainable Federated Learning for Supply Chain Credit Assessment
- BiND: A Neural Discriminator-Decoder for Accurate Bimanual Trajectory Prediction in Brain-Computer Interfaces
- Minimizing the Weighted Number of Tardy Jobs: Data-Driven Heuristic for Single-Machine Scheduling
- MUFFIN: Mixture of User-Adaptive Frequency Filtering for Sequential Recommendation
- DeH4R: A Decoupled and Hybrid Method for Road Network Graph Extraction
- Interactive Query Answering on Knowledge Graphs with Soft Entity Constraints
- MACTAS: Self-Attention-Based Module for Inter-Agent Communication in Multi-Agent Reinforcement Learning
- In-Context Decision Making for Optimizing Complex AutoML Pipelines
- Text2Weight: Bridging Natural Language and Neural Network Weight Spaces
- Generative Model-Based Feature Attention Module for Video Action Analysis
- Equinox: Holistic Fair Scheduling in Serving Large Language Models
- GazeProphet: Software-Only Gaze Prediction for VR Foveated Rendering
- Fracture Detection and Localisation in Wrist and Hand Radiographs using Detection Transformer Variants
- Saudi-Dialect-ALLaM: LoRA Fine-Tuning for Dialectal Arabic Generation
- Evaluating Open-Source Vision Language Models for Facial Emotion Recognition against Traditional Deep Learning Models
- AI-Augmented Photon-Trapping Spectrometer-on-a-Chip on Silicon Platform with Extended Near-Infrared Sensitivity
- Cross-Cancer Knowledge Transfer in WSI-based Prognosis Prediction
- Autoregressive Typical Thermal States
- SVDformer: Direction-Aware Spectral Graph Embedding Learning via SVD and Transformer
- STPFormer: A State-of-the-Art Pattern-Aware Spatio-Temporal Transformer for Traffic Forecasting
- Extracting Structured Requirements from Unstructured Building Technical Specifications for Building Information Modeling
- CCFC: Core & Core-Full-Core Dual-Track Defense for LLM Jailbreak Protection
- A Fully Spectral Neuro-Symbolic Reasoning Architecture with Graph Signal Processing as the Computational Backbone
- DPad: Efficient Diffusion Language Models with Suffix Dropout
- Graph Concept Bottleneck Models
- GLASS: Global-Local Aggregation for Inference-time Sparsification of LLMs
- SPANER: Shared Prompt Aligner for Multimodal Semantic Representation
- LOOP: A Plug-and-Play Neuro-Symbolic Framework for Enhancing Planning in Autonomous Systems
- Overcoming Latency Bottlenecks in On-Device Speech Translation: A Cascaded Approach with Alignment-Based Streaming MT
- HiFo-Prompt: Prompting with Hindsight and Foresight for LLM-based Automatic Heuristic Design
- A Dual-Attention Graph Network for fMRI Data Classification
- Rapidly Adapting to New Voice Spoofing: Few-Shot Detection of Synthesized Speech Under Distribution Shifts
- Prune2Drive: A Plug-and-Play Framework for Accelerating Vision-Language Models in Autonomous Driving
- AI Agents for Photonic Integrated Circuit Design Automation
- Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy
- Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey
- SimGenHOI: Physically Realistic Whole-Body Humanoid-Object Interaction via Generative Modeling and Reinforcement Learning
- Point upsampling networks for single-photon sensing
- Arabic ASR on the SADA Large-Scale Arabic Speech Corpus with Transformer-Based Models
- MaskSem: Semantic-Guided Masking for Learning 3D Hybrid High-Order Motion Representation
- ONG: One-Shot NMF-based Gradient Masking for Efficient Model Sparsification
- Maximum Score Routing For Mixture-of-Experts
- Wavy Transformer
- MemorySim: An RTL-level, timing accurate simulator model for the Chisel ecosystem
- Prompt-Induced Linguistic Fingerprints for LLM-Generated Fake News Detection
- D2-Mamba: Dual-Scale Fusion and Dual-Path Scanning with SSMs for Shadow Removal
- Adaptive-basis sample-based neural diagonalization for quantum many-body systems
- A Neural-Network Framework for Tracking and Identification of Cosmic-Ray Nuclei in the RadMap Telescope
- Refine-and-Contrast: Adaptive Instance-Aware BEV Representations for Multi-UAV Collaborative Object Detection
- Multi-Granularity Distribution Modeling for Video Watch Time Prediction via Exponential-Gaussian Mixture Network
- Score-informed Neural Operator for Enhancing Ordering-based Causal Discovery
- Data-driven particle dynamics: Structure-preserving coarse-graining for emergent behavior in non-equilibrium systems
- PAPPL: Personalized AI-Powered Progressive Learning Platform
- FedSODA: Federated Fine-tuning of LLMs via Similarity Group Pruning and Orchestrated Distillation Alignment
- Adaptive Hybrid Caching for Efficient Text-to-Video Diffusion Model Acceleration
- The Application of Transformer-Based Models for Predicting Consequences of Cyber Attacks
- Multimodal Chain of Continuous Thought for Latent-Space Reasoning in Vision-Language Models
- Learning In-context \pmbn-grams with Transformers: Sub-\pmbn-grams Are Near-stationary Points
- EgoTwin: Dreaming Body and View in First Person
- Explainable AI-Based Feature Selection Approaches for Raman Spectroscopy
- MDPO: Overcoming the Training-Inference Divide of Masked Diffusion Language Models
- Hybrid Deep Reconstruction for Vignetting-Free Upconversion Imaging through Scattering in ENZ Materials
- CTFlow: Video-Inspired Latent Flow Matching for 3D CT Synthesis
- Accelerating LLM Inference via Dynamic KV Cache Placement in Heterogeneous Memory System
- A systematic evaluation of Dutch large language models’ surprisal estimates in sentence, paragraph and book reading
- Word Meanings in Transformer Language Models
- Dextr: Zero-Shot Neural Architecture Search with Singular Value Decomposition and Extrinsic Curvature
- FLARE: Fast Low-rank Attention Routing Engine
- A Hybrid Surrogate for Electric Vehicle Parameter Estimation and Power Consumption via Physics-Informed Neural Operators
- OPTIC-ER: A Reinforcement Learning Framework for Real-Time Emergency Response and Equitable Resource Allocation in Underserved African Communities
- An Initial Study of Bird's-Eye View Generation for Autonomous Vehicles using Cross-View Transformers
- LangVision-LoRA-NAS: Neural Architecture Search for Variable LoRA Rank in Vision Language Models
- IPGPhormer: Interpretable Pathology Graph-Transformer for Survival Analysis
- Is GPT-OSS Good? A Comprehensive Evaluation of OpenAI's Latest Open Source Models
- Uncovering Emergent Physics Representations Learned In-Context by Large Language Models
- LumiMAS: A Comprehensive Framework for Real-Time Monitoring and Enhanced Observability in Multi-Agent Systems
- CEM-Net: Cross-Emotion Memory Network for Emotional Talking Face Generation
- TaoSR1: The Thinking Model for E-commerce Relevance Search
- Polarization Reconfigurable Transmit-Receive Beam Alignment with Interpretable Transformer
- HuBERT-VIC: Improving Noise-Robust Automatic Speech Recognition of Speech Foundation Model via Variance-Invariance-Covariance Regularization
- Exploring Self-Supervised Audio Models for Generalized Anomalous Sound Detection
- Arabic Multimodal Machine Learning: Datasets, Applications, Approaches, and Challenges
- Towards Generalizable Human Activity Recognition: A Survey
- MBMamba: When Memory Buffer Meets Mamba for Structure-Aware Image Deblurring
- Bi-Axial Transformers: Addressing the Increasing Complexity of EHR Classification
- ZigzagAttention: Efficient Long-Context Inference with Exclusive Retrieval and Streaming Heads
- The Yokai Learning Environment: Tracking Beliefs Over Space and Time
- The Structural Sources of Verb Meaning Revisited: Large Language Models Display Syntactic Bootstrapping
- Scalable RF Simulation in Generative 4D Worlds
- Multi-Modal Drift Forecasting of Leeway Objects via Navier-Stokes-Guided CNN and Sequence-to-Sequence Attention-Based Models
- Exploring Efficiency Frontiers of Thinking Budget in Medical Reasoning: Scaling Laws between Computational Resources and Reasoning Quality
- VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine
- Enhancing 3D point accuracy of laser scanner through multi-stage convolutional neural network for applications in construction
- Generic Event Boundary Detection via Denoising Diffusion
- VimoRAG: Video-based Retrieval-augmented 3D Motion Generation for Motion Language Models
- Mitigating Jailbreaks with Intent-Aware LLMs
- AI Models for Depressive Disorder Detection and Diagnosis: A Review
- MOON: Generative MLLM-based Multimodal Representation Learning for E-commerce Product Understanding
- SafeCtrl: Region-Based Safety Control for Text-to-Image Diffusion via Detect-Then-Suppress
- Data Shift of Object Detection in Autonomous Driving
- SupraTok: Cross-Boundary Tokenization for Enhanced Language Model Performance
- Invitation Is All You Need! Promptware Attacks Against LLM-Powered Assistants in Production Are Practical and Dangerous
- Time-Scale Coupling Between States and Parameters in Recurrent Neural Networks
- Q-FSRU: Quantum-Augmented Frequency-Spectral Fusion for Medical Visual Question Answering
- BConformeR: A Conformer Based on Mutual Sampling for Unified Prediction of Continuous and Discontinuous Antibody Binding Sites
- Efficient Modular Learning through Naive LoRA Summation: Leveraging Orthogonality in High-Dimensional Models
- What Matters for Bioacoustic Encoding
- The Rise of Generative AI for Metal-Organic Framework Design and Synthesis
- OVSegDT: Segmenting Transformer for Open-Vocabulary Object Goal Navigation
- TinyTim: A Family of Language Models for Divergent Generation
- TrajSV: A Trajectory-based Model for Sports Video Representations and Applications
- MultiPark: Multimodal Parking Transformer with Next-Segment Prediction
- Language models align with brain regions that represent concepts across modalities
- Multi-State Tracker: Enhancing Efficient Object Tracking via Multi-State Specialization and Interaction
- Handwritten Text Recognition of Historical Manuscripts Using Transformer-Based Models
- Hierarchical Graph Feature Enhancement with Adaptive Frequency Modulation for Visual Recognition
- AI Agentic Programming: A Survey of Techniques, Challenges, and Opportunities
- Reference Points in LLM Sentiment Analysis: The Role of Structured Context
- EvoPSF: Online Evolution of Autonomous Driving Models via Planning-State Feedback
- Model Interpretability and Rationale Extraction by Input Mask Optimization
- NeMo: A Neuron-Level Modularizing-While-Training Approach for Decomposing DNN Models
- RegimeNAS: Regime-Aware Differentiable Architecture Search With Theoretical Guarantees for Financial Trading
- Delving into Dynamic Scene Cue-Consistency for Robust 3D Multi-Object Tracking
- CSGO: Generalized Optimization for Cold Start in Wireless Collaborative Edge LLM Systems
- Inference performance evaluation for LLMs on edge devices with a novel benchmarking framework and metric
- Generalized Decoupled Learning for Enhancing Open-Vocabulary Dense Perception
- Borrowing From the Future: Enhancing Early Risk Assessment through Contrastive Learning
- E-CaTCH: Event-Centric Cross-Modal Attention with Temporal Consistency and Class-Imbalance Handling for Misinformation Detection
- Representation Quantization for Collaborative Filtering Augmentation
- MobQA: A Benchmark Dataset for Semantic Understanding of Human Mobility Data through Question Answering
- Towards the Next-generation Bayesian Network Classifiers
- ADMIRE-BayesOpt: Accelerated Data MIxture RE-weighting for Language Models with Bayesian Optimization
- Actor-Critic for Continuous Action Chunks: A Reinforcement Learning Framework for Long-Horizon Robotic Manipulation with Sparse Reward
- CineTrans: Learning to Generate Videos with Cinematic Transitions via Masked Diffusion Models
- Copyright Protection for Large Language Models: A Survey of Methods, Challenges, and Trends
- Diffusion is a code repair operator and generator
- HierOctFusion: Multi-scale Octree-based 3D Shape Generation via Part-Whole-Hierarchy Message Passing
- Hybrid-Hierarchical Fashion Graph Attention Network for Compatibility-Oriented and Personalized Outfit Recommendation
- Predictive Multimodal Modeling of Diagnoses and Treatments in EHR
- Are AI Machines Making Humans Obsolete?
- Abundance-Aware Set Transformer for Microbiome Sample Embedding
- Puppeteer: Rig and Animate Your 3D Models
- Human-in-Context: Unified Cross-Domain 3D Human Motion Modeling via In-Context Learning
- MAESTRO: Masked AutoEncoders for Multimodal, Multitemporal, and Multispectral Earth Observation Data
- Generalizable Federated Learning using Client Adaptive Focal Modulation
- Advances in Speech Separation: Techniques, Challenges, and Future Trends
- A Multimodal Neural Network for Recognizing Subjective Self-Disclosure Towards Social Robots
- Memory-Augmented Transformers: A Systematic Review from Neuroscience Principles to Enhanced Model Architectures
- Dissecting Generalized Category Discovery: Multiplex Consensus under Self-Deconstruction
- Novel View Synthesis using DDIM Inversion
- Advancing Autonomous Incident Response: Leveraging LLMs and Cyber Threat Intelligence
- SurfaceLogicKV: Surface and Logic Attention Behaviors are All You Need for Robust KV Cache Compression
- Nonlinear filtering based on density approximation and deep BSDE prediction
- FuXi-β: Towards a Lightweight and Fast Large-Scale Generative Recommendation Model
- Reasoning in Computer Vision: Taxonomy, Models, Tasks, and Methodologies
- EgoMusic-driven Human Dance Motion Estimation with Skeleton Mamba
- Nonlocal Monte Carlo via Reinforcement Learning
- A Retrieval Augmented Spatio-Temporal Framework for Traffic Prediction
- Efficient Methods for Accurate Sparse Trajectory Recovery and Map Matching
- Exploring Cross-Utterance Speech Contexts for Conformer-Transducer Speech Recognition Systems
- NanoControl: A Lightweight Framework for Precise and Efficient Control in Diffusion Transformer
- Translation of Text Embedding via Delta Vector to Suppress Strongly Entangled Content in Text-to-Image Diffusion Models
- Super LiDAR Reflectance for Robotic Perception
- XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization
- Improving Generative Cross-lingual Aspect-Based Sentiment Analysis with Constrained Decoding
- Large Language Models for Summarizing Czech Historical Documents and Beyond
- Advancing Cross-lingual Aspect-Based Sentiment Analysis with LLMs and Constrained Decoding for Sequence-to-Sequence Models
- Quantum Wavefront Correction via Machine Learning for Satellite-to-Earth CV-QKD
- Beyond Semantic Understanding: Preserving Collaborative Frequency Components in LLM-based Recommendation
- SynBrain: Enhancing Visual-to-fMRI Synthesis via Probabilistic Representation Learning
- STRelay: A Universal Spatio-Temporal Relaying Framework for Location Prediction with Future Spatiotemporal Contexts
- Street Review: A Participatory AI-Based Framework for Assessing Streetscape Inclusivity
- Ask ChatGPT: Caveats and Mitigations for Individual Users of AI Chatbots
- MRFD: Multi-Region Fusion Decoding with Self-Consistency for Mitigating Hallucinations in LVLMs
- Failures to Surface Harmful Contents in Video Large Language Models
- A learning-driven automatic planning framework for proton PBS treatments of H&N cancers
- Deep Learning for Crack Detection: A Review of Learning Paradigms, Generalizability, and Datasets
- Meta-Metrics and Best Practices for System-Level Inference Performance Benchmarking
- Pruning and Malicious Injection: A Retraining-Free Backdoor Attack on Transformer Models
- CarAT: Carbon Atom Tracing across Industrial Chemical Value Chains via Chemistry Language Models
- Pre-trained Transformer-models using chronic invasive electrophysiology for symptom decoding without patient-individual training
- MANGO: Multimodal Attention-based Normalizing Flow Approach to Fusion Learning
- Nested-ReFT: Efficient Reinforcement Learning for Large Language Model Fine-Tuning via Off-Policy Rollouts
- Empowering Morphing Attack Detection using Interpretable Image-Text Foundation Model
- Specialised or Generic? Tokenization Choices for Radiology Language Models
- Stable Diffusion Models are Secretly Good at Visual In-Context Learning
- Amazon Nova AI Challenge -- Trusted AI: Advancing secure, AI-assisted software development
- Data-Driven Discovery of Interpretable Kalman Filter Variants through Large Language Models and Genetic Programming
- Story2Board: A Training-Free Approach for Expressive Storyboard Generation
- GBC: Generalized Behavior-Cloning Framework for Whole-Body Humanoid Imitation
- Quo Vadis Handwritten Text Generation for Handwritten Text Recognition?
- Teaching LLMs to Speak Spectroscopy
- Modern Neural Networks for Small Tabular Datasets: The New Default for Field-Scale Digital Soil Mapping?
- EEGDM: EEG Representation Learning via Generative Diffusion Model
- Reverse Convolution and Its Applications to Image Restoration
- Physical Autoregressive Model for Robotic Manipulation without Action Pretraining
- Automated Segmentation of Coronal Brain Tissue Slabs for 3D Neuropathology
- MUJICA: Reforming SISR Models for PBR Material Super-Resolution via Cross-Map Attention
- μ-Parametrization for Mixture of Experts
- Generative Modeling with Multi-Instance Reward Learning for E-commerce Creative Optimization
- CKFNet: Neural Network Aided Cubature Kalman filtering
- Improving ARDS Diagnosis Through Context-Aware Concept Bottleneck Models
- Temporal Anchoring in Deepening Embedding Spaces: Event-Indexed Projections, Drift, Convergence, and an Internal Computational Architecture
- On Negative-aware Preference Optimization for Recommendation
- TOTNet: Occlusion-Aware Temporal Tracking for Robust Ball Detection in Sports Videos
- Multi-Contrast Fusion Module: An attention mechanism integrating multi-contrast features for fetal torso plane classification
- MInDI-3D: Iterative Deep Learning in 3D for Sparse-view Cone Beam Computed Tomography
- OSUM-EChat: Enhancing End-to-End Empathetic Spoken Chatbot via Understanding-Driven Spoken Dialogue
- Edge General Intelligence Through World Models and Agentic AI: Fundamentals, Solutions, and Challenges
- Learning Spatial Decay for Vision Transformers
- Generation of Indian Sign Language Letters, Numbers, and Words
- UWBa at SemEval-2025 Task 7: Multilingual and Crosslingual Fact-Checked Claim Retrieval
- Cross-lingual Aspect-Based Sentiment Analysis: A Survey on Tasks, Approaches, and Challenges
- GeoMAE: Masking Representation Learning for Spatio-Temporal Graph Forecasting with Missing Values
- Animate-X++: Universal Character Image Animation with Dynamic Backgrounds
- Shadow in the Cache: Unveiling and Mitigating Privacy Risks of KV-cache in LLM Inference
- Security Analysis of ChatGPT: Threats and Privacy Risks
- Mitigating Distribution Shift in Stock Price Data via Return-Volatility Normalization for Accurate Prediction
- FM4NPP: A Scaling Foundation Model for Nuclear and Particle Physics
- RelayFormer: A Unified Local-Global Attention Framework for Scalable Image and Video Manipulation Localization
- LaajMeter: A Framework for LaaJ Evaluation
- DenoDet V2: Phase-Amplitude Cross Denoising for SAR Object Detection
- How Safe Will I Be Given What I Saw? Calibrated Prediction of Safety Chances for Image-Controlled Autonomy
- UltraLight Med-Vision Mamba for Classification of Neoplastic Progression in Tubular Adenomas
- Dynamic Survival Prediction using Longitudinal Images based on Transformer
- Fake-Mamba: Real-Time Speech Deepfake Detection Using Bidirectional Mamba as Self-Attention's Alternative
- HQ-OV3D: A High Box Quality Open-World 3D Detection Framework based on Diffision Model
- Can AI Keep a Secret? Contextual Integrity Verification: A Provable Security Architecture for LLMs
- Relative Pose Regression with Pose Auto-Encoders: Enhancing Accuracy and Data Efficiency for Retail Applications
- SinLlama -- A Large Language Model for Sinhala
- Utilizing Multilingual Encoders to Improve Large Language Models for Low-Resource Languages
- Cross-BCI, A Cross-BCI-Paradigm Classifica-tion Model Towards Universal BCI Applications
- Automated Charge Transition Detection in Quantum Dot Charge Stability Diagrams
- UniConvNet: Expanding Effective Receptive Field while Maintaining Asymptotically Gaussian Distribution for ConvNets of Any Scale
- NEFMind: Parameter-Efficient Fine-Tuning of Open-Source LLMs for Telecom APIs Automation
- Label Smoothing is a Pragmatic Information Bottleneck
- Lay2Story: Extending Diffusion Transformers for Layout-Togglable Story Generation
- Reveal-Bangla: A Dataset for Cross-Lingual Multi-Step Reasoning Evaluation
- Masked Clustering Prediction for Unsupervised Point Cloud Pre-training
- Oblivionis: A Lightweight Learning and Unlearning Framework for Federated Large Language Models
- Geometry-Aware Global Feature Aggregation for Real-Time Indirect Illumination
- Revisiting Efficient Semantic Segmentation: Learning Offsets for Better Spatial and Class Feature Alignment
- ReQuestNet: A Foundational Learning model for Channel Estimation
- Evaluating Podcast Recommendations with Profile-Aware LLM-as-a-Judge
- Visual Prompting for Robotic Manipulation with Annotation-Guided Pick-and-Place Using ACT
- Real-time forecasting of chaotic dynamics from sparse data and autoencoders
- Generative Modeling for Robust Deep Reinforcement Learning on the Traveling Salesman Problem
- Expert-Guided Diffusion Planner for Auto-Bidding
- PADReg: Physics-Aware Deformable Registration Guided by Contact Force for Ultrasound Sequences
- Closing the Performance Gap in Generative Recommenders with Collaborative Tokenization and Efficient Modeling
- Envisioning Generative Artificial Intelligence in Cartography and Mapmaking
- Prompt-Based Approach for Czech Sentiment Analysis
- Agentic Graph Neural Networks for Wireless Communications and Networking Towards Edge General Intelligence: A Survey
- Generative AI for Critical Infrastructure in Smart Grids: A Unified Framework for Synthetic Data Generation and Anomaly Detection
- Load Forecasting on A Highly Sparse Electrical Load Dataset Using Gaussian Interpolation
- Fine-grained Video Dubbing Duration Alignment with Segment Supervised Preference Optimization
- M3-Net: A Cost-Effective Graph-Free MLP-Based Model for Traffic Prediction
- Training Kindai OCR with parallel textline images and self-attention feature distance-based loss
- SegDAC: Visual Generalization in Reinforcement Learning via Dynamic Object Tokens
- DeepCodon: A deep learning codon-optimization model to enhance protein expression
- Scaling Learned Image Compression Models up to 1 Billion
- Profiling Large Language Model Inference on Apple Silicon: A Quantization Perspective
- RealisMotion: Decomposed Human Motion Control and Video Generation in the World Space
- SPARC: Soft Probabilistic Adaptive multi-interest Retrieval Model via Codebooks for recommender system
- Sound Signal Synthesis with Auxiliary Classifier GAN, COVID-19 cough as an example
- DiffVolume: Diffusion Models for Volume Generation in Limit Order Books
- MechaFormer: Sequence Learning for Kinematic Mechanism Design Automation
- P/D-Device: Disaggregated Large Language Model between Cloud and Devices
- DevNous: An LLM-Based Multi-Agent System for Grounding IT Project Management in Unstructured Conversation
- Revealing the Role of Audio Channels in ASR Performance Degradation
- Stationarity Exploration for Multivariate Time Series Forecasting
- Weakly Supervised Fine-grained Span-Level Framework for Chinese Radiology Report Quality Assurance
- Integrating attention into explanation frameworks for language and vision transformers
- Deep Neural Network Calibration by Reducing Classifier Shift with Stochastic Masking
- Pep2Prob Benchmark: Predicting Fragment Ion Probability for MS2-based Proteomics
- DeCAL Tokenwise Compression
- Benchmarking Federated Learning for Throughput Prediction in 5G Live Streaming Applications
- Towards Scalable Training for Handwritten Mathematical Expression Recognition
- Mamba-FCS: Joint Spatio- Frequency Feature Fusion, Change-Guided Attention, and SeK Loss for Enhanced Semantic Change Detection in Remote Sensing
- SAEMark: Multi-bit LLM Watermarking with Inference-Time Scaling
- Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions
- Follow-Your-Shape: Shape-Aware Image Editing via Trajectory-Guided Region Control
- ME-TST+: Micro-expression Analysis via Temporal State Transition with ROI Relationship Awareness
- From Source to Target: Leveraging Transfer Learning for Predictive Process Monitoring in Organizations
- TRIDE: A Text-assisted Radar-Image weather-aware fusion network for Depth Estimation
- DiffractGPT: Atomic Structure Determination from X-ray Diffraction Patterns using Generative Pre-trained Transformer
- Mem4D: Decoupling Static and Dynamic Memory for Dynamic Scene Reconstruction
- CleanCTG: A Deep Learning Model for Multi-Artefact Detection and Reconstruction in Cardiotocography
- Deep Space Weather Model: Long-Range Solar Flare Prediction from Multi-Wavelength Images
- DETACH: Cross-domain Learning for Long-Horizon Tasks via Mixture of Disentangled Experts
- Learning Robust Satellite Attitude Dynamics with Physics-Informed Normalising Flow
- Make Your MoVe: Make Your 3D Contents by Adapting Multi-View Diffusion Models to External Editing
- Undress to Redress: A Training-Free Framework for Virtual Try-On
- Towards Multimodal Sentiment Analysis via Contrastive Cross-modal Retrieval Augmentation and Hierachical Prompts
- Towards Comprehensible Recommendation with Large Language Model Fine-tuning
- Towards Theoretical Understanding of Transformer Test-Time Computing: Investigation on In-Context Linear Regression
- Large Language Models for Subjective Language Understanding: A Survey
- Large Language Models for Czech Aspect-Based Sentiment Analysis
- AD-AVSR: Asymmetric Dual-stream Enhancement for Robust Audio-Visual Speech Recognition
- Filling MIDI Velocity using U-Net Image Colorizer
- Temporal User Profiling with LLMs: Balancing Short-Term and Long-Term Preferences for Recommendations
- Using LLMs to Capture Users' Temporal Context for Recommendation
- Prototype-Guided Curriculum Learning for Zero-Shot Learning
- TAP: Parameter-efficient Task-Aware Prompting for Adverse Weather Removal
- Pareto Multi-Objective Alignment for Language Models
- Scaled-Dot-Product Attention as One-Sided Entropic Optimal Transport
- Disentangling Multiplex Spatial-Temporal Transition Graph Representation Learning for Socially Enhanced POI Recommendation
- Czech Dataset for Complex Aspect-Based Sentiment Analysis Tasks
- UrzaGPT: LoRA-Tuned Large Language Models for Card Selection in Collectible Card Games
- GraphCoT-VLA: A 3D Spatial-Aware Reasoning Vision-Language-Action Model for Robotic Manipulation with Ambiguous Instructions
- Training-Free ANN-to-SNN Conversion for High-Performance Spiking Transformer
- Dream4D: Lifting Camera-Controlled I2V towards Spatiotemporally Consistent 4D Generation
- Lightning Prediction under Uncertainty: DeepLight with Hazy Loss
- Scalable Controllable Accented TTS
- Leveraging GNN to Enhance MEF Method in Predicting ENSO
- A Spin Glass Characterization of Neural Networks
- Think Before You Talk: Enhancing Meaningful Dialogue Generation in Full-Duplex Speech Language Models with Planning-Inspired Text Guidance
- Keyword Mamba: Spoken Keyword Spotting with State Space Models
- Préface
- Planner-Refiner: Dynamic Space-Time Refinement for Vision-Language Alignment in Videos
- Multi-task Adversarial Attacks against Black-box Model with Few-shot Queries
- A Survey on Non-Intrusive ASR Refinement: From Output-Level Correction to Full-Model Distillation
- FLUID: Flow-Latent Unified Integration via Token Distillation for Expert Specialization in Multimodal Learning
- SUIT: Spatial-Spectral Union-Intersection Interaction Network for Hyperspectral Object Tracking
- ASM-UNet: Adaptive Scan Mamba Integrating Group Commonalities and Individual Variations for Fine-Grained Segmentation
- Selection and Exploitation of High-Quality Knowledge from Large Language Models for Recommendation
- CoopDiff: Anticipating 3D Human-object Interactions via Contact-consistent Decoupled Diffusion
- SketchConcept: Sketching-based Concept Recomposition for Product Design using Generative AI
- Enhancing Rumor Detection Methods with Propagation Structure Infused Language Model
- Strategies of Code-switching in Human-Machine Dialogs
- RORPCap: Retrieval-based Objects and Relations Prompt for Image Captioning
- DCoAR: Deep Concept Injection into Unified Autoregressive Models for Personalized Text-to-Image Generation
- FlexCTC: GPU-powered CTC Beam Decoding With Advanced Contextual Abilities
- Tasa: Thermal-aware 3D-Stacked Architecture Design with Bandwidth Sharing for LLM Inference
- Representation Understanding via Activation Maximization
- BrainATCL: Adaptive Temporal Brain Connectivity Learning for Functional Link Prediction and Age Estimation
- ForeSight: Multi-View Streaming Joint Object Detection and Trajectory Forecasting
- Improving Real-Time Concept Drift Detection using a Hybrid Transformer-Autoencoder Framework
- AR-GRPO: Training Autoregressive Image Generation Models via Reinforcement Learning
- eMotions: A Large-Scale Dataset and Audio-Visual Fusion Network for Emotion Analysis in Short-form Videos
- Score Before You Speak: Improving Persona Consistency in Dialogue Generation using Response Quality Scores
- ESNERA: Empirical and semantic named entity alignment for named entity dataset merging
- Towards Effective Prompt Stealing Attack against Text-to-Image Diffusion Models
- Multi-level Advantage Credit Assignment for Cooperative Multi-Agent Reinforcement Learning
- Remote Sensing Image Intelligent Interpretation with the Language-Centered Perspective: Principles, Methods and Challenges
- Data-Efficient Neural Training with Dynamic Connectomes
- Bridging Classical and Quantum Computing for Next-Generation Language Models
- Can Multitask Learning Enhance Model Explainability?
- Structure-Preserving Digital Twins via Conditional Neural Whitney Forms
- Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody
- CISO: Species Distribution Modeling Conditioned on Incomplete Species Observations
- gpt-oss-120b & gpt-oss-20b Model Card
- Fractal Language Modelling by Universal Sequence Maps (USM)
- Generalizing Scaling Laws for Dense and Sparse Large Language Models
- Position: Ideas Should be the Center of Machine Learning Research
- Approaching the integration of large language models in the parliamentary workspace
- HV-OCTAMamba: A high-order vision Mamba network for robust retinal vasculature segmentation in OCTA images
- AlphaFold reveals but sometimes distorts an organizational principle of protein folding
- AI with Symbolic Empathy: Shannon-Neumann Insight Guided Logic
- Learning the language of protein-protein interactions
- SiMiC: Context-aware silicon microstructure characterization using attention-based convolutional neural networks for field-emission tip analysis
- Intuition emerges in Maximum Caliber models at criticality
- Lifelong Learner: Discovering Versatile Neural Solvers for Vehicle Routing Problems
- SlimInfer: Accelerating Long-Context LLM Inference via Dynamic Token Pruning
- Cognitive Workspace: Active Memory Management for LLMs -- An Empirical Study of Functional Infinite Context
- Leveraging transfer learning for accurate estimation of ionic migration barriers in solids
- MotionSwap
- Acoustic Non-Stationarity Objective Assessment with Hard Label Criteria for Supervised Learning Models
- Tree-Based Deep Learning for Ranking Symbolic Integration Algorithms
- Text as Any-Modality for Zero-Shot Classification by Consistent Prompt Tuning
- SpeakerLM: End-to-End Versatile Speaker Diarization and Recognition with Multimodal Large Language Models
- From Explainable to Explanatory Artificial Intelligence: Toward a New Paradigm for Human-Centered Explanations through Generative AI
- Aligning Effective Tokens with Video Anomaly in Large Language Models
- AntiCheatPT: A Transformer-Based Approach to Cheat Detection in Competitive Computer Games
- ADT4Coupons: An Innovative Framework for Sequential Coupon Distribution in E-commerce
- DSConv: Dynamic Splitting Convolution for Pansharpening
- LLM Serving Optimization with Variable Prefill and Decode Lengths
- AdaptInfer: Adaptive Token Pruning for Vision-Language Model Inference with Dynamical Text Guidance
- Architecture-Aware Generalization Bounds for Temporal Networks: Theory and Fair Comparison Methodology
- Improved Sub-Visible Particle Classification in Flow Imaging Microscopy via Generative AI-Based Image Synthesis
- Hypergraph Neural Network with State Space Models for Node Classification
- Bifrost-1: Bridging Multimodal LLMs and Diffusion Models with Patch-level CLIP Latents
- Latent Policy Barrier: Learning Robust Visuomotor Policies by Staying In-Distribution
- Hybrid(Transformer+CNN)-based Polyp Segmentation
- ASAudio: A Survey of Advanced Spatial Audio Research
- Lightweight Quad Bayer HybridEVS Demosaicing via State Space Augmented Cross-Attention
- Matrix-Driven Identification and Reconstruction of LLM Weight Homology
- Recurrent Deep Differentiable Logic Gate Networks
- Accelerating Quantum Monte Carlo Calculations with Set-Equivariant Architectures and Transfer Learning
- AnomalyMoE: Towards a Language-free Generalist Model for Unified Visual Anomaly Detection
- GLIDR: Graph-Like Inductive Logic Programming with Differentiable Reasoning
- Role of Large Language Models and Retrieval-Augmented Generation for Accelerating Crystalline Material Discovery: A Systematic Review
- Testing the Limits of Machine Translation from One Book
- Panel-Scale Reconfigurable Photonic Interconnects for Scalable AI Computation
- Mask & Match: Learning to Recognize Handwritten Math with Self-Supervised Attention
- ME3-BEV: Mamba-Enhanced Deep Reinforcement Learning for End-to-End Autonomous Driving with BEV-Perception
- AGI for the Earth, the path, possibilities and how to evaluate intelligence of models that work with Earth Observation Data?
- Omni Geometry Representation Learning vs Large Language Models for Geospatial Entity Resolution
- DiffCap: Diffusion-based Real-time Human Motion Capture using Sparse IMUs and a Monocular Camera
- Distribution-Specific Learning for Joint Salient and Camouflaged Object Detection
- In-Context Reinforcement Learning via Communicative World Models
- Multi-view Gaze Target Estimation
- Towards Transparent Ethical AI: A Roadmap for Trustworthy Robotic Systems
- WeTok: Powerful Discrete Tokenization for High-Fidelity Visual Reconstruction
- Iterative Learning of Computable Phenotypes for Treatment Resistant Hypertension using Large Language Models
- Adapting Vision-Language Models Without Labels: A Comprehensive Survey
- EnergyPatchTST: Multi-scale Time Series Transformers with Uncertainty Estimation for Energy Forecasting
- Learning Geometric-Aware Quadrature Rules for Functional Minimization
- Large Language Models Transform Organic Synthesis From Reaction Prediction to Automation
- Smoothing Slot Attention Iterations and Recurrences
- A Scalable Pipeline for Enabling Non-Verbal Speech Generation and Understanding
- Deformable Attention Graph Representation Learning for Histopathology Whole Slide Image Analysis
- Cross-View Localization via Redundant Sliced Observations and A-Contrario Validation
- Can Language Models Critique Themselves? Investigating Self-Feedback for Retrieval Augmented Generation at BioASQ 2025
- Optimal Corpus Aware Training for Neural Machine Translation
- Reduction Techniques for Survival Analysis
- CoCAViT: Compact Vision Transformer with Robust Global Coordination
- Estimating Musical Surprisal from Audio in Autoregressive Diffusion Model Noise Spaces
- SONAR-LLM: Autoregressive Transformer that Thinks in Sentence Embeddings and Speaks in Tokens
- Fair Play in the Newsroom: Actor-Based Filtering Gender Discrimination in Text Corpora
- MoBE: Mixture-of-Basis-Experts for Compressing MoE-based LLMs
- ArbiViewGen: Controllable Arbitrary Viewpoint Camera Data Generation for Autonomous Driving via Stable Diffusion Models
- HiFi-Mamba: Dual-Stream W-Laplacian Enhanced Mamba for High-Fidelity MRI Reconstruction
- ReasoningTrack: Chain-of-Thought Reasoning for Long-term Vision-Language Tracking
- Advanced Hybrid Transformer LSTM Technique with Attention and TS Mixer for Drilling Rate of Penetration Prediction
- X-MoGen: Unified Motion Generation across Humans and Animals
- Digital Twin Channel-Aided CSI Prediction: An Environment-Based Subspace Extraction Approach for Achieving Low Overhead and High Robustness
- Deep Learning-based Animal Behavior Analysis: Insights from Mouse Chronic Pain Models
- Attention Basin: Why Contextual Position Matters in Large Language Models
- RAP: Real-time Audio-driven Portrait Animation with Video Diffusion Transformer
- BEE-RAG: Balanced Entropy Engineering for Retrieval-Augmented Generation
- An End-to-End Multi-objective Ensemble Ranking Framework for Video Recommendation
- PoseGen: In-Context LoRA Finetuning for Pose-Controllable Long Human Video Generation
- Align, Don't Divide: Revisiting the LoRA Architecture in Multi-Task Learning
- MetaDiT: Enabling Fine-grained Constraints in High-degree-of Freedom Metasurface Design
- TANGO: Graph Neural Dynamics via Learned Energy and Tangential Flows
- A Context-aware Attention and Graph Neural Network-based Multimodal Framework for Misogyny Detection
- A Study of the Framework and Real-World Applications of Language Embedding for 3D Scene Understanding
- Q-DPTS: Quantum Differentially Private Time Series Forecasting via Variational Quantum Circuits
- Evaluation of Finetuned LLMs in AMR Parsing
- Multimodal Causal-Driven Representation Learning for Generalizable Medical Image Segmentation
- Modeling Rapid Contextual Learning in the Visual Cortex with Fast-Weight Deep Autoencoder Networks
- Collaborative Learning-Enhanced Lightweight Models for Predicting Arterial Blood Pressure Waveform in a Large-scale Perioperative Dataset
- Sentiment-Aware Stock Price Prediction with Transformer and LLM-Generated Formulaic Alpha
- A Rose by Any Other Name Would Smell as Sweet: Categorical Homotopy Theory for Large Language Models
- Tesserae: Scalable Placement Policies for Deep Learning Workloads
- Training chord recognition models on artificially generated audio
- Surformer v1: Transformer-Based Surface Classification Using Tactile and Vision Features
- Salt-Rock Creep Deformation Forecasting Using Deep Neural Networks and Analytical Models for Subsurface Energy Storage Applications
- Why Attention Fails: A Taxonomy of Faults in Attention-Based Neural Networks
- Test-Time Adaptation for Video Highlight Detection Using Meta-Auxiliary Learning and Cross-Modality Hallucinations
- Advancing Hate Speech Detection with Transformers: Insights from the MetaHate
- Gaussian mixture layers for neural networks
- Keyword Spotting with Hyper-Matched Filters for Small Footprint Devices
- Fine-Tuning Small Language Models (SLMs) for Autonomous Web-based Geographical Information Systems (AWebGIS)
- Voost: A Unified and Scalable Diffusion Transformer for Bidirectional Virtual Try-On and Try-Off
- MambaITD: An Efficient Cross-Modal Mamba Network for Insider Threat Detection
- DMFI: Dual-Modality Fine-Tuning and Inference Framework for LLM-Based Insider Threat Detection
- Parity-Aware Byte-Pair Encoding: Improving Cross-lingual Fairness in Tokenization
- Uncertainty-aware Predict-Then-Optimize Framework for Equitable Post-Disaster Power Restoration
- Live Music Models
- HiD-VAE: Interpretable Generative Recommendation via Hierarchical and Disentangled Semantic IDs
- A Reproducible, Scalable Pipeline for Synthesizing Autoregressive Model Literature
- GraphProp: Training the Graph Foundation Models using Graph Properties
- Share Your Attention: Transformer Weight Sharing via Matrix-based Dictionary Learning
- DDTracking: A Deep Generative Framework for Diffusion MRI Tractography with Streamline Local-Global Spatiotemporal Modeling
- CLASP: Cross-modal Salient Anchor-based Semantic Propagation for Weakly-supervised Dense Audio-Visual Event Localization
- Two-Way Garment Transfer: Unified Diffusion Framework for Dressing and Undressing Synthesis
- Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding
- Balancing Stylization and Truth via Disentangled Representation Steering
- PRISM: Lightweight Multivariate Time-Series Classification through Symmetric Multi-Resolution Convolutional Layers
- Benchmarking Quantum and Classical Sequential Models for Urban Telecommunication Forecasting
- Zero-Residual Concept Erasure via Progressive Alignment in Text-to-Image Model
- GFocal: A Global-Focal Neural Operator for Solving PDEs on Arbitrary Geometries
- Small transformer architectures for task switching
- Decoding the Multimodal Maze: A Systematic Review on the Adoption of Explainability in Multimodal Attention-based Models
- Efficient Inter-Task Attention for Multitask Transformer Models
- Why are LLMs' abilities emergent?
- TSPO: Temporal Sampling Policy Optimization for Long-form Video Language Understanding
- LUST: A Multi-Modal Framework with Hierarchical LLM-based Scoring for Learned Thematic Significance Tracking in Multimedia Content
- Chain of Questions: Guiding Multimodal Curiosity in Language Models
- Deliberative Reasoning Network: An Uncertainty-Driven Paradigm for Belief-Tracked Inference with Pretrained Language Models
- Forgetting: A New Mechanism Towards Better Large Language Model Fine-tuning
- TDSNNs: Competitive Topographic Deep Spiking Neural Networks for Visual Cortex Modeling
- Revisiting Continual Semantic Segmentation with Pre-trained Vision Models
- WiFo-CF: Wireless Foundation Model for CSI Feedback
- T3Time: Tri-Modal Time Series Forecasting via Adaptive Multi-Head Alignment and Residual Fusion
- PA-RNet: Perturbation-Aware Reasoning Network for Multimodal Time Series Forecasting
- DP-GPT4MTS: Dual-Prompt Large Language Model for Textual-Numerical Time Series Forecasting
- Circuit-Aware SAT Solving: Guiding CDCL via Conditional Probabilities
- Gather and Trace: Rethinking Video TextVQA from an Instance-oriented Perspective
- RPCANet++: Deep Interpretable Robust PCA for Sparse Object Segmentation
- BadTime: An Effective Backdoor Attack on Multivariate Long-Term Time Series Forecasting
- Deeper Inside Deep ViT
- Bridging Search and Recommendation through Latent Cross Reasoning
- STARE: Predicting Decision Making Based on Spatio-Temporal Eye Movements
- Benefit from Rich: Tackling Search Interaction Sparsity in Search Enhanced Recommendation
- Excavate the potential of Single-Scale Features: A Decomposition Network for Water-Related Optical Image Enhancement
- Efficient Scaling for LLM-based ASR
- TNet: Terrace Convolutional Decoder Network for Remote Sensing Image Semantic Segmentation
- Towards Globally Predictable k-Space Interpolation: A White-box Transformer Approach
- Motion is the Choreographer: Learning Latent Pose Dynamics for Seamless Sign Language Generation
- Quantum Temporal Fusion Transformer
- FeDaL: Federated Dataset Learning for Time Series Foundation Models
- ZARA: Zero-shot Motion Time-Series Analysis via Knowledge and Retrieval Driven LLM Agents
- CORE-ReID V2: Advancing the Domain Adaptation for Object Re-Identification with Optimized Training and Ensemble Fusion
- Prototype-Driven Structure Synergy Network for Remote Sensing Images Segmentation
- CodonMoE: DNA Language Models for mRNA Analyses
- Investigating the Impact of Large-Scale Pre-training on Nutritional Content Estimation from 2D Images
- MiDashengLM: Efficient Audio Understanding with General Audio Captions
- Automated Deep Learning–based Segmentation of the Dentate Nucleus Using Quantitative Susceptibility Mapping MRI
- Beyond Hard Sharing: Efficient Multi-Task Speech-to-Text Modeling with Supervised Mixture of Experts
- Data and AI governance: Promoting equity, ethics, and fairness in large language models
- GP and LLMs for Program Synthesis: No Clear Winners
- Scaling Up Audio-Synchronized Visual Animation: An Efficient Training Paradigm
- Model Compression vs. Adversarial Robustness: An Empirical Study on Language Models for Code
- Point-Based Shape Representation Generation with a Correspondence-Preserving Diffusion Model
- Statistical Confidence Rescoring for Robust 3D Scene Graph Generation from Multi-View Images
- Comparing Normalization Methods for Portfolio Optimization with Reinforcement Learning
- Uncertainty-aware Accurate Elevation Modeling for Off-road Navigation via Neural Processes
- LiDARCrafter: Dynamic 4D World Modeling from LiDAR Sequences
- Veila: Panoramic LiDAR Generation from a Monocular RGB Image
- AttnTrace: Attention-based Context Traceback for Long-Context LLMs
- FairLangProc: A Python package for fairness in NLP
- Efficient Morphology-Aware Policy Transfer to New Embodiments
- AttZoom: Attention Zoom for Better Visual Features
- Hidden Dynamics of Massive Activations in Transformer Training
- Minimal Convolutional RNNs Accelerate Spatiotemporal Learning
- Block: Balancing Load in LLM Serving with Context, Knowledge and Predictive Scheduling
- Demystifying Sequential Recommendations: Counterfactual Explanations via Genetic Algorithms
- CADD: Context aware disease deviations via restoration of brain images using normative conditional diffusion models
- VITA: Variational Pretraining of Transformers for Climate-Robust Crop Yield Forecasting
- DeepFaith: A Domain-Free and Model-Agnostic Unified Framework for Highly Faithful Explanations
- Semantic-aware Graph-guided Behavior Sequences Generation with Large Language Models for Smart Homes
- GRASPing Anatomy to Improve Pathology Segmentation
- AVPDN: Learning Motion-Robust and Scale-Adaptive Representations for Video-Based Polyp Detection
- Cropping outperforms dropout as an augmentation strategy for training self-supervised text embeddings
- Data Overdose? Time for a Quadruple Shot: Knowledge Graph Construction using Enhanced Triple Extraction
- User Perception of Attention Visualizations: Effects on Interpretability Across Evidence-Based Medical Documents
- Variety Is the Spice of Life: Detecting Misinformation with Dynamic Environmental Representations
- Revisiting Heat Flux Analysis of Tungsten Monoblock Divertor on EAST using Physics-Informed Neural Network
- WaMo: Wavelet-Enhanced Multi-Frequency Trajectory Analysis for Fine-Grained Text-Motion Retrieval
- When Deep Learning Fails: Limitations of Recurrent Models on Stroke-Based Handwriting for Alzheimer's Disease Detection
- BaroPoser: Real-time Human Motion Tracking from IMUs and Barometers in Everyday Devices
- Understanding the Embedding Models on Hyper-relational Knowledge Graph
- Artificial Intelligence and Generative Models for Materials Discovery -- A Review
- Beyond Isolated Words: Diffusion Brush for Handwritten Text-Line Generation
- MVTOP: Multi-View Transformer-based Object Pose-Estimation
- Revisiting Deep Information Propagation: Fractal Frontier and Finite-size Effects
- BadBlocks: Low-Cost and Stealthy Backdoor Attacks Tailored for Text-to-Image Diffusion Models
- Monocular Depth Estimation with Global-Aware Discretization and Local Context Modeling
- Dual-disentangle Framework for Diversified Sequential Recommendation
- Rethinking Selectivity in State Space Models: A Minimal Predictive Sufficiency Approach
- SSFMamba: Symmetry-driven Spatial-Frequency Feature Fusion for 3D Medical Image Segmentation
- From Text to Trajectories: GPT-2 as an ODE Solver via In-Context
- ClinicalFMamba: Advancing Clinical Assessment using Mamba-based Multimodal Neuroimaging Fusion
- Adversarial Attention Perturbations for Large Object Detection Transformers
- Mamba-X: An End-to-End Vision Mamba Accelerator for Edge Computing Devices
- Inductive transfer learning from regression to classification in ECG analysis
- Understanding Transformers through the Lens of Pavlovian Conditioning
- Topos Theory for Generative AI and LLMs
- Toward Practical Equilibrium Propagation: Brain-inspired Recurrent Neural Network with Feedback Regulation and Residual Connections
- Inland-LOAM: Voxel-Based Structural Semantic LiDAR Odometry and Mapping for Inland Waterway Navigation
- Historical Prediction Attention Mechanism based Trajectory Forecasting for Proactive Work Zone Safety in a Digital Twin Environment
- AMD-Mamba: A Phenotype-Aware Multi-Modal Framework for Robust AMD Prognosis
- LLM-based IR-system for Bank Supervisors
- Tricks and Plug-ins for Gradient Boosting with Transformers
- Following Route Instructions using Large Vision-Language Models: A Comparison between Low-level and Panoramic Action Spaces
- Semantic Structure in Large Language Model Embeddings
- Evaluation and Analysis of Deep Neural Transformers and Convolutional Neural Networks on Modern Remote Sensing Datasets
- TransAM: Transformer-Based Agent Modeling for Multi-Agent Systems via Local Trajectory Encoding
- PyCAT4: A Hierarchical Vision Transformer-based Framework for 3D Human Pose Estimation
- Adaptive Knowledge Distillation for Device-Directed Speech Detection
- Breaking the Top-K Barrier: Advancing Top-K Ranking Metrics Optimization in Recommender Systems
- D2PPO: Diffusion Policy Policy Optimization with Dispersive Loss
- DeepKoopFormer: A Koopman Enhanced Transformer Based Architecture for Time Series Forecasting
- JustDense: Just using Dense instead of Sequence Mixer for Time Series analysis
- FinWorld: An All-in-One Open-Source Platform for End-to-End Financial AI Research and Deployment
- TreeRanker: Fast and Model-agnostic Ranking System for Code Suggestions in IDEs
- Glioblastoma Overall Survival Prediction With Vision Transformers
- Superior resilience to poisoning and amenability to unlearning in quantum machine learning
- Modality Bias in LVLMs: Analyzing and Mitigating Object Hallucination via Attention Lens
- Graph Embedding in the Graph Fractional Fourier Transform Domain
- Neural Network-Based Algorithmic Trading Systems: Multi-Timeframe Analysis and High-Frequency Execution in Cryptocurrency Markets
- VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo
- LaMPE: Length-aware Multi-grained Positional Encoding for Adaptive Long-context Scaling Without Training
- A French Version of the OLDI Seed Corpus
- SGAD: Semantic and Geometric-aware Descriptor for Local Feature Matching
- GaussianCross: Cross-modal Self-supervised 3D Representation Learning via Gaussian Splatting
- After the Party: Navigating the Mapping From Color to Ambient Lighting
- Trainable Dynamic Mask Sparse Attention
- Modular Transformer Architecture for Precision Agriculture Imaging
- Toward Efficient Spiking Transformers: Synapse Pruning Meets Synergistic Learning-Based Compensation
- Communicating Plans, Not Percepts: Scalable Multi-Agent Coordination with Embodied World Models
- MArgE: Meshing Argumentative Evidence from Multiple Large Language Models for Justifiable Claim Verification
- Multi-Class Human/Object Detection on Robot Manipulators using Proprioceptive Sensing
- Epi2-Net: Advancing Epidemic Dynamics Forecasting with Physics-Inspired Neural Networks
- Why Generate When You Can Transform? Unleashing Generative Attention for Dynamic Recommendation
- SpectraLLM: Uncovering the Ability of LLMs for Molecular Structure Elucidation from Multi-Spectral Data
- Contextual Graph Transformer: A Small Language Model for Enhanced Engineering Document Information Extraction
- Isolating Culture Neurons in Multilingual Large Language Models
- Localizing Audio-Visual Deepfakes via Hierarchical Boundary Modeling
- User Trajectory Prediction Unifying Global and Local Temporal Information
- Versatile yet Efficient Network Traffic Analysis: Offloading Network Foundation Model to SmartNIC
- Identifying actionable driver mutations in lung cancer using an efficient Asymmetric Transformer Decoder
- HGTS-Former: Hierarchical HyperGraph Transformer for Multivariate Time Series Analysis
- InfoSyncNet: Information Synchronization Temporal Convolutional Network for Visual Speech Recognition
- SUAD: Solid-Channel Ultrasound Injection Attack and Defense to Voice Assistants
- PRIME: Plasticity-Robust Incremental Model for Encrypted Traffic Classification in Dynamic Network Environments
- StarPose: 3D Human Pose Estimation via Spatial-Temporal Autoregressive Diffusion
- Parameter-Efficient Routed Fine-Tuning: Mixture-of-Experts Demands Mixture of Adaptation Modules
- Fine-grained Multiple Supervisory Network for Multi-modal Manipulation Detecting and Grounding
- Neuromorphic Computing with Multi-Frequency Oscillations: A Bio-Inspired Approach to Artificial Intelligence
- From Generation to Consumption: Personalized List Value Estimation for Re-ranking
- Large-Scale Model Enabled Semantic Communication Based on Robust Knowledge Distillation
- Text2Lip: Progressive Lip-Synced Talking Face Generation from Text via Viseme-Guided Rendering
- Quantum Machine Learning-based Test Oracle for Autonomous Mobile Robots
- GlaBoost: A multimodal Structured Framework for Glaucoma Risk Stratification
- Context Guided Transformer Entropy Modeling for Video Compression
- CTBench: Cryptocurrency Time Series Generation Benchmark
- Diffusion-based 3D Hand Motion Recovery with Intuitive Physics
- CSLRConformer: A Data-Centric Conformer Approach for Continuous Arabic Sign Language Recognition on the Isharah Datase
- "Energon": Unveiling Transformers from GPU Power and Thermal Side-Channels
- DMSC: Dynamic Multi-Scale Coordination Framework for Time Series Forecasting
- Rein++: Efficient Generalization and Adaptation for Semantic Segmentation with Vision Foundation Models
- Learning Unified System Representations for Microservice Tail Latency Prediction
- EAC-MoE: Expert-Selection Aware Compressor for Mixture-of-Experts Large Language Models
- Drift-aware Collaborative Assistance Mixture of Experts for Heterogeneous Multistream Learning
- Pulse Shape Discrimination Algorithms: Survey and Benchmark
- KANMixer: Can KAN Serve as a New Modeling Core for Long-term Time Series Forecasting?
- Empowering Tabular Data Preparation with Language Models: Why and How?
- MGCR-Net:Multimodal Graph-Conditioned Vision-Language Reconstruction Network for Remote Sensing Change Detection
- Social Media Information Operations
- CGCCE-Net:Change-Guided Cross Correlation Enhancement Network for Remote Sensing Building Change Detection
- Measuring and Predicting Where and When Pathologists Focus their Visual Attention while Grading Whole Slide Images of Cancer
- Neural Predictive Control to Coordinate Discrete- and Continuous-Time Models for Time-Series Analysis with Control-Theoretical Improvements
- The Art of Breaking Words: Rethinking Multilingual Tokenizer Design
- FluidFormer: Transformer with Continuous Convolution for Particle-based Fluid Simulation
- A Plug-and-Play Multi-Criteria Guidance for Diverse In-Betweening Human Motion Generation
- The potential and limitations of large language models for automatic classification of teachers' motivational messages in educational research
- VAGPO: Vision-augmented Asymmetric Group Preference Optimization for Graph Routing Problems
- LetheViT: Selective Machine Unlearning for Vision Transformers via Attention-Guided Contrastive Learning
- Enhancing Spectrogram Realism in Singing Voice Synthesis via Explicit Bandwidth Extension Prior to Vocoder
- Am I Blue or Is My Hobby Counting Teardrops? Expression Leakage in Large Language Models as a Symptom of Irrelevancy Disruption
- DiffusionFF: A Diffusion-based Framework for Joint Face Forgery Detection and Fine-Grained Artifact Localization
- Hedging with memory: shallow and deep learning with signatures
- MHARFedLLM: Multimodal Human Activity Recognition Using Federated Large Language Model
- Revenue Optimization in Wireless Video Caching Networks: A Privacy-Preserving Two-Stage Solution
- Practical, Generalizable and Robust Backdoor Attacks on Text-to-Image Diffusion Models
- M3LLM: Model Context Protocol-aided Mixture of Vision Experts For Multimodal LLMs in Networks
- Pi-SAGE: Permutation-invariant surface-aware graph encoder for binding affinity prediction
- The Vanishing Gradient Problem for Stiff Neural Differential Equations
- Frequency-Constrained Learning for Long-Term Forecasting
- Instruction-based Time Series Editing
- Training Dynamics of the Cooldown Stage in Warmup-Stable-Decay Learning Rate Scheduler
- HT-Transformer: Event Sequences Classification by Accumulating Prefix Information with History Tokens
- Can3Tok: Canonical 3D Tokenization and Latent Modeling of Scene-Level 3D Gaussians
- Fast and scalable retrosynthetic planning with a transformer neural network and speculative beam search
- Signals, Concepts, and Laws: Toward Universal, Explainable Time-Series Forecasting
- Spatial-Frequency Aware for Object Detection in RAW Image
- Effects of Feature Correlations on Associative Memory Capacity
- Lightweight Backbone Networks Only Require Adaptive Lightweight Self-Attention Mechanisms
- Predicting Video Slot Attention Queries from Random Slot-Feature Pairs
- SWAN: Synergistic Wavelet-Attention Network for Infrared Small Target Detection
- How Far Are LLMs from Symbolic Planners? An NLP-Based Perspective
- SGCap: Decoding Semantic Group for Zero-shot Video Captioning
- Soft Separation and Distillation: Toward Global Uniformity in Federated Unsupervised Learning
- Mitigating Information Loss under High Pruning Rates for Efficient Large Vision Language Models
- Multi-Cache Enhanced Prototype Learning for Test-Time Generalization of Vision-Language Models
- Perspective from a Broader Context: Can Room Style Knowledge Help Visual Floorplan Localization?
- Unified Generation-Refinement Planning: Bridging Guided Flow Matching and Sampling-Based MPC for Social Navigation
- RSPO: Risk-Seeking Policy Optimization for Pass@k and Max@k Metrics in Large Language Models
- Transformers in Pseudo-Random Number Generation: A Dual Perspective on Theory and Practice
- UniEgoMotion: A Unified Model for Egocentric Motion Reconstruction, Forecasting, and Generation
- Recovering Individual-Level Activity Sequences from Location-Based Service Data Using a Novel Transformer-Based Model
- TestWeaver: Execution-aware, Feedback-driven Regression Testing Generation with Large Language Models
- SaviorRec: Semantic-Behavior Alignment for Cold-Start Recommendation
- RelMap: Reliable Spatiotemporal Sensor Data Visualization via Imputative Spatial Interpolation
- COLLAGE: Adaptive Fusion-based Retrieval for Augmented Policy Learning
- Multimodal Attention-Aware Fusion for Diagnosing Distal Myopathy: Evaluating Model Interpretability and Clinician Trust
- Capturing More: Learning Multi-Domain Representations for Robust Online Handwriting Verification
- TEACH: Text Encoding as Curriculum Hints for Scene Text Recognition
- Graph-based Interaction Augmentation Network for Robust Multimodal Sentiment Analysis
- Multispin Physics of AI Tipping Points and Hallucinations
- Mobile U-ViT: Revisiting large kernel and U-shaped ViT for efficient medical image segmentation
- Connectivity Management in Satellite-Aided Vehicular Networks with Multi-Head Attention-Based State Estimation
- Teaching at Scale: Leveraging AI to Evaluate and Elevate Engineering Education
- Structured Spectral Graph Learning for Anomaly Classification in 3D Chest CT Scans
- Fusion of Modulation Spectrogram and SSL with Multi-head Attention for Fake Speech Detection
- Optimal Scheduling Algorithms for LLM Inference: Theory and Practice
- Beyond Fixed: Training-Free Variable-Length Denoising for Diffusion Large Language Models
- Interpreting Performance Profiles with Deep Learning
- Forecasting NCAA Basketball Outcomes with Deep Learning: A Comparative Study of LSTM and Transformer Models
- AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation
- IAMAP: Unlocking Deep Learning in QGIS for non-coders and limited computing resources
- LeakSealer: A Semisupervised Defense for LLMs Against Prompt Injection and Leakage Attacks
- Wukong Framework for Not Safe For Work Detection in Text-to-Image systems
- Session-Based Recommendation with Validated and Enriched LLM Intents
- HiPrune: Training-Free Visual Token Pruning via Hierarchical Attention in Vision-Language Models
- Text-Attributed Graph Anomaly Detection via Multi-Scale Cross- and Uni-Modal Contrastive Learning
- E2ATST: A Temporal-Spatial Optimized Energy-Efficient Architecture for Training Spiking Transformer
- HyPCV-Former: Hyperbolic Spatio-Temporal Transformer for 3D Point Cloud Video Anomaly Detection
- Transforming Credit Risk Analysis: A Time-Series-Driven ResE-BiLSTM Framework for Post-Loan Default Detection
- Cued-Agent: A Collaborative Multi-Agent System for Automatic Cued Speech Recognition
- EdgeInfinite-Instruct: Bridging SFT-Based Optimization and NPU-Level Efficiency for Edge Devices
- Representation Shift: Unifying Token Compression with FlashAttention
- Boosting Vision Semantic Density with Anatomy Normality Modeling for Medical Vision-language Pre-training
- Multi-grained spatial-temporal feature complementarity for accurate online cellular traffic prediction
- AniMer+: Unified Pose and Shape Estimation Across Mammalia and Aves via Family-Aware Transformer
- Integrating clinical reasoning into large language model-based diagnosis through etiology-aware attention steering
- ECGTwin: Personalized ECG Generation Using Controllable Diffusion Model
- Large AI Model-Enabled Secure Communications in Low-Altitude Wireless Networks: Concepts, Perspectives and Case Study
- Guided Depth Map Super-Resolution via Multi-Scale Fusion U-shaped Mamba Network
- Light-Weight Diffusion Multiplier and Uncertainty Quantification for Fourier Neural Operators
- Multi-Layer Attention is the Amplifier of Demonstration Effectiveness
- Soil moisture retrieval and spatiotemporal variation analysis based on deep learning
- UIS-Mamba: Exploring Mamba for Underwater Instance Segmentation via Dynamic Tree Scan and Hidden State Weaken
- Combining Observations and Models: A Review of the <scp>CARDAMOM</scp> Framework for Data‐Constrained Terrestrial Ecosystem Modeling
- CP-FREEZER: Latency Attacks against Vehicular Cooperative Perception
- DocTron-Formula: Generalized Formula Recognition in Complex and Structured Scenarios
- ITUNLP at SemEval-2025 Task 8: Question-Answering over Tabular Data: A Zero-Shot Approach using LLM-Driven Code Generation
- On Learning Closed-Loop Probabilistic Multi-Agent Simulator
- Cooperative Perception: A Resource-Efficient Framework for Multi-Drone 3D Scene Reconstruction Using Federated Diffusion and NeRF
- Minimum Data, Maximum Impact: 20 annotated samples for explainable lung nodule classification
- Invariant Graph Transformer for Out-of-Distribution Generalization
- DD-DeepONet: Domain decomposition and DeepONet for solving partial differential equations in three application scenarios
- GeoExplorer: Active Geo-localization with Curiosity-Driven Exploration
- GanitBench: A bi-lingual benchmark for evaluating mathematical reasoning in Vision Language Models
- Gaussian Variation Field Diffusion for High-fidelity Video-to-4D Synthesis
- Segmenting proto-halos with vision transformers
- XSpecMesh: Quality-Preserving Auto-Regressive Mesh Generation Acceleration via Multi-Head Speculative Decoding
- Search for t tt tW Production at √(s) = 13 TeV Using a Modified Graph Neural Network at the LHC
- CFDagent: A Language-Guided, Zero-Shot Multi-Agent System for Complex Flow Simulation
- One-Step Flow Policy Mirror Descent
- Efficient Masked Attention Transformer for Few-Shot Classification and Segmentation
- Deep Learning-based Prediction of Clinical Trial Enrollment with Uncertainty Estimates
- GraphRAG-R1: Graph Retrieval-Augmented Generation with Process-Constrained Reinforcement Learning
- Beyond Gloss: A Hand-Centric Framework for Gloss-Free Sign Language Translation
- 3D-MOOD: Lifting 2D to 3D for Monocular Open-Set Object Detection
- MoLAN: A Unified Modality-Aware Noise Dynamic Editing Framework for Multimodal Sentiment Analysis
- Mitigating Resolution-Drift in Federated Learning: Case of Keypoint Detection
- Enhanced Arabic Text Retrieval with Attentive Relevance Scoring
- FMIP: Joint Continuous-Integer Flow For Mixed-Integer Linear Programming
- Multi-Prompt Progressive Alignment for Multi-Source Unsupervised Domain Adaptation
- Simulation-based inference for Precision Neutrino Physics through Neural Monte Carlo tuning
- SequenceLayers: Sequence Processing and Streaming Neural Networks Made Easy
- Unveiling Super Experts in Mixture-of-Experts Large Language Models
- iLRM: An Iterative Large 3D Reconstruction Model
- Efficient Real-Time Aircraft ETA Prediction via Feature Tokenization Transformer
- Not Just What, But When: Integrating Irregular Intervals to LLM for Sequential Recommendation
- A Bayesian Hybrid Parameter-Efficient Fine-Tuning Method for Large Language Models
- BAR Conjecture: the Feasibility of Inference Budget-Constrained LLM Services with Authenticity and Reasoning
- How Far Are AI Scientists from Changing the World?
- RL as Regressor: A Reinforcement Learning Approach for Function Approximation
- Personalized Education with Ranking Alignment Recommendation
- From Image Captioning to Visual Storytelling
- Improved Algorithms for Kernel Matrix-Vector Multiplication Under Sparsity Assumptions
- Your Spending Needs Attention: Modeling Financial Habits with Transformers
- On LLM-Assisted Generation of Smart Contracts from Business Processes
- Feature Importance across Domains for Improving Non-Intrusive Speech Intelligibility Prediction in Hearing Aids
- AGA: An adaptive group alignment framework for structured medical cross-modal representation learning
- Benchmarking Partial Observability in Reinforcement Learning with a Suite of Memory-Improvable Domains
- BS-1-to-N: Diffusion-Based Environment-Aware Cross-BS Channel Knowledge Map Generation for Cell-Free Networks
- On the Expressiveness of Softmax Attention: A Recurrent Neural Network Perspective
- Distributed AI Agents for Cognitive Underwater Robot Autonomy
- Reinitializing weights vs units for maintaining plasticity in neural networks
- Foundation Models for Clean Energy Forecasting: A Comprehensive Review
- Details Matter for Indoor Open-vocabulary 3D Instance Segmentation
- Modeling Human Gaze Behavior with Diffusion Models for Unified Scanpath Prediction
- Synchronization of mean-field models on the circle
- Wall Shear Stress Estimation in Abdominal Aortic Aneurysms: Towards Generalisable Neural Surrogate Models
- Real-time News Story Identification
- Quantifying surprise in clinical care: Detecting highly informative events in electronic health records with foundation models
- Enhancing Multi-Agent Collaboration with Attention-Based Actor-Critic Policies
- Social-Pose: Enhancing Trajectory Prediction with Human Body Pose
- GestureHYDRA: Semantic Co-speech Gesture Synthesis via Hybrid Modality Diffusion Transformer and Cascaded-Synchronized Retrieval-Augmented Generation
- trAIce3D: A Prompt-Driven Transformer Based U-Net for Semantic Segmentation of Microglial Cells from Large-Scale 3D Microscopy Images
- Robust Adverse Weather Removal via Spectral-based Spatial Grouping
- Visual Language Models as Zero-Shot Deepfake Detectors
- Shallow Features Matter: Hierarchical Memory with Heterogeneous Interaction for Unsupervised Video Object Segmentation
- RCR-AF: Enhancing Model Generalization via Rademacher Complexity Reduction Activation Function
- Artificial intelligence engineering [wikipedia]
- Ashish Vaswani [wikipedia]
- Attention (machine learning) [wikipedia]
- Attention Is All You Need [wikipedia]
- GPT-3 [wikipedia]
- Generative AI [wikipedia]
- Geospatial foundation model [wikipedia]
- History of numerical weather prediction [wikipedia]
- Lukasz Kaiser [wikipedia]
- Medical image computing [wikipedia]
- Neural network (machine learning) [wikipedia]
- Noam Shazeer [wikipedia]
- Synthetic media [wikipedia]
- T5 (language model) [wikipedia]
- Timeline of machine learning [wikipedia]
- Transformer (deep learning) [wikipedia]
Discussions
- The paper that made ChatGPT possible [hn, 173 points, 55 comments]
- FOSAI Nexus Summer 2023 Edition [lemmy, 24 points, 0 comments]
- While you're studying up, paul, let me just say, if all y'all haven't read "Attention Is All You Need" (arxiv.org/abs/1706.03762) yet, you are sorely missing out on a truly riveting piece of modern hi [bsky, 19 points, 0 comments]
- The paper that blew up LLMs was Attention is All You Need. The key point is that you don't need a hidden variable stream - the prompt and previously generated text is ample information for a neural ne [bsky, 18 points, 2 comments]
- A single-head attention layer processes a sequence by assigning relevance scores between elements. It computes these scores to decide how much focus each part of the seq should receive and combines th [bsky, 15 points, 0 comments]
- Attention Is All You Need [lemmy, 14 points, 2 comments]
- Deep Neural Networks – Attention is all you need (No recurrence or convolutions) [hn, 12 points, 0 comments]
- Attention Is All You Need (Neural Networks) [hn, 8 points, 3 comments]
- Thema der Lerngruppe heute: Transformer-Modelle. Frühere Sequenzmodelle verarbeiten Daten Schritt für Schritt entlang der Sequenz. Transformer nutzen stattdessen Self-Attention: Dabei kann jedes Token [bsky, 7 points, 2 comments]
- 2017: arxiv.org/abs/1706.03762 2026: www.anthropic.com/news/confide... [bsky, 6 points, 0 comments]
- Attention Is All You Need [hn, 6 points, 2 comments]
- Tudo que conhecemos hoje como modelos generativos, GenAI, chatgpt, papa de jaco, tem origem da arquitetura de Transformers, que por sua vez tem origem das redes neurais simples. Então sim, tudo nasceu [bsky, 5 points, 0 comments]
- But we know that is what it is doing, because we know how it works. And the d&d attributes are named for concepts, like Wisdom, where there is a lot of text to draw from. arxiv.org/abs/1706.03762 [bsky, 5 points, 1 comments]
- arxiv.org/abs/1706.03762 [bsky, 4 points, 1 comments]
- Attention Is All You Need [hn, 4 points, 0 comments]
- But in all seriousness the main breakthrough that led to the architecture behind LLMs was a bit of a happy accident as Bob Ross would have said. Get a bunch of smart and well funded people flailing aw [bsky, 4 points, 1 comments]
- Attention is all you need (2017) [hn, 3 points, 0 comments]
- Attention is all you need paper is offline due to invalid certificate [hn, 3 points, 0 comments]
- if you're diving into generative ai, make sure to understand the transformer architecture - it is the backbone of all modern LLMs these resources are amazing and free: - Attention Is All You Need: ar [bsky, 3 points, 1 comments]
- Attention Is All You Need [hn, 3 points, 0 comments]
- the paper that proposed the transformer architecture everyone uses for everything now developed it for machine translation tasks arxiv.org/abs/1706.03762 [bsky, 3 points, 0 comments]
- À ce sujet, il faut différencier «qui a développé l’ia » ( c’est eux : arxiv.org/abs/1706.03762 ) de «qui a entraîné une IA sur tous le contenu des artistes pour en vendre la production » [bsky, 3 points, 1 comments]
- Our paper club recently revisited some of the earlier language modeling papers. Here's a one-liner for each. --- Attention: Query, Key, and Value are all you need* *Also position embeddings, mult [bsky, 3 points, 1 comments]
- Attention Is All You Need [hn, 3 points, 0 comments]
- feels like somebody should invent the conspiracy theory that the paper introducing transformer architectures was called "Attention Is All You Need" as a reference to the capitalistic impoverishment of [bsky, 3 points, 1 comments]
- just read up on the transformer architecture in general; I suppose the classic paper-that-started-it-all is arxiv.org/abs/1706.03762 [bsky, 3 points, 2 comments]
- Attention Is All You Need [hn, 3 points, 0 comments]
- Attention Is All You Need [pdf] [hn, 2 points, 0 comments]
- Attention is All You Need - Shoshana Zuboff👍 👇 arxiv.org/abs/1706.03762 [bsky, 2 points, 1 comments]
- There have been Language Models before #LLMs. Then came the Transformer architecture by Vaswani et al. in their paper “Attention is All You Need” that revolutionized #NLP in 2017. The Transformer m [bsky, 2 points, 1 comments]
- Attention Is All You Need was published in June 2017 — the paper introduced the Transformer architecture, which laid the foundation for modern large language models arxiv.org/abs/1706.03762 #paper #ll [bsky, 2 points, 2 comments]
- Is it possible that no one shared this link? doi.org/10.48550/arX... [bsky, 2 points, 0 comments]
- Attention is All You Need by Google Brain [hn, 2 points, 1 comments]
- Attention Is All You Need [hn, 2 points, 0 comments]
- Attention Is All You Need [hn, 2 points, 0 comments]
- Eine Art "Gedächtnis" haben Rekurrente Neuronale Netze, weil sie Rückkopplungen zwischen den Schichten zulassen. Die Transformer-Architektur dagegen, auf denen die meisten Large Language Models (und a [bsky, 2 points, 0 comments]
- כן, זה השם של המאמר שהמציא את הטכנולוגיה הזו. arxiv.org/abs/1706.03762 [bsky, 2 points, 2 comments]
- The # arXiv as the public record and version control system of a scientific manuscript: 7 versions spanning 6 years. "[Submitted on 12 Jun 2017 (v1), last revised 2 Aug 2023 (this version, v7)]" "Atte [mastodon, 1 points, 0 comments]
- In that article above you'll find Yale's bio dept usecase for LLMs - cell-level research. For generative AI in general there's usecases you probably use often Google Translate (which is where the basi [bsky, 1 points, 1 comments]
- arxiv.org/abs/1706.03762 [bsky, 1 points, 0 comments]
- arxiv.org/abs/1706.03762 [bsky, 1 points, 0 comments]
- arxiv.org/abs/1706.03762 [bsky, 1 points, 0 comments]
- We're coming up next year on the 10 year anniversary of the Attention is All you Need paper release which kicked off this LLM madness. And it's still very much a technology which is in dire need of mo [bsky, 1 points, 1 comments]
- hm, i wonder what google translate also is based on arxiv.org/pdf/1706.03762 [bsky, 1 points, 1 comments]
- Attention is all Elon needs. And we know Attention is All You Need [1] Therefore... [1] https://arxiv.org/abs/1706.03762 [bsky, 1 points, 0 comments]
- I used Audio Overview with this as the only source and it clearly combined it with lots of other (more recent) info. Like how the architecture is used for protein folding and all kinds of things not m [bsky, 1 points, 0 comments]
- Hi, computer scientist here- transformers were *INVENTED* for machine translation. It is the *INTENDED USE CASE*. Chatbots are a side effect. Here, read the dang paper: arxiv.org/abs/1706.03762 [bsky, 1 points, 0 comments]
- If anybody is at all interested in understanding how we got where we are with GenAI, it's because somebody looked at an incredibly complex problem and came up with an answer that was clear, simple, an [bsky, 1 points, 1 comments]
- youtu.be/z2tuflnRly4?... [bsky, 1 points, 0 comments]
- "Attention Is All You Need" 2017, cet article revêt une importance considérable, après sa publication les Transformers vont naître et tout va s'accélérer... arxiv.org/abs/1706.03762 #cornerstone #IA [bsky, 1 points, 0 comments]
- 15 AI Research Papers Every AI Engineer Should Read 1. Attention Is All You Need (Transformers) arxiv.org/abs/1706.03762 2. LoRA: Low-Rank Adaptation arxiv.org/abs/2106.09685 3. PEFT (Parameter-Effici [bsky, 1 points, 1 comments]
- Hmmm. Not really. The transformer architecture was originally intended for use in translation - an information reformatting task, if a complicated one. Search is not the only thing Google does. The or [bsky, 1 points, 2 comments]
- Time for more architecture innovation. "Attention is all you need" was published in 2017 after all. *2017* arxiv.org/abs/1706.03762 [bsky, 1 points, 1 comments]
- If you're genuinely curious, this might give you some insight into how they work and why. This was the original whitepaper that defines the tensor model of how they think, more or less. Everything els [bsky, 1 points, 1 comments]
- "Attention Is All You Need" arxiv.org/pdf/1706.03762 [bsky, 1 points, 0 comments]
- arxiv.org/abs/1706.03762 [bsky, 1 points, 1 comments]
- Attention Is All You Need Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, Illia Polosukhin Source: arxiv.org/abs/1706.03762 [bsky, 1 points, 0 comments]
- In Vaswani, Ashish, Noam Shazeer, Niki Parmar, et al. " Attention Is All You Need ." arXiv:1706.03762. Preprint, arXiv , August 2, 2023. the word "transduction" is used passim . Wiktionary defines it [stackexchange, 1 points]
- For a deep dive, the modern chatbots all developed from Google's Attention Is All You Need paper. arxiv.org/pdf/1706.03762 [bsky, 1 points, 1 comments]
- So step one the intro paper here is fine actually, and everything later is scale and hacks rather than fundamental changes to architecture. The critical point is that the model; the _coefficients_, ar [bsky, 1 points, 1 comments]
- Attention Is All You Need (2017) [hn, 1 points, 0 comments]
- > complex recurrent https://arxiv.org/abs/1706.03762 [bsky, 0 points, 0 comments]
- huh. meng (paraphrased) thinks an "unlock" similar to transformers will happen. me sharing...meng is referring to arxiv.org/abs/1706.03762 [bsky, 0 points, 1 comments]
- Try looking at the function definitions in arxiv.org/pdf/1706.03762 (the paper that started Transformers) sections 3.2.1, 3.2.2, 3.3 Then you might be able to understand how each "layer" is composed [bsky, 0 points, 1 comments]
- Oops arxiv.org/abs/1706.03762 [bsky, 0 points, 1 comments]
- Se me acumulan los papers para leer pero gracias! Le echaré un ojo 👁️ el último que me guardé es este arxiv.org/pdf/1706.03762 [bsky, 0 points, 0 comments]
- Reading the so famous whitepaper "Attention is all you need" (arxiv.org/pdf/1706.03762) by Vaswani et al. #AI #Transformers #Attention #SelfAttention #algorithms [bsky, 0 points, 0 comments]
- It’s the research paper that started everything arxiv.org/abs/1706.03762 [bsky, 0 points, 0 comments]
- Yeah things have, uh. Changed a bit since then. arxiv.org/abs/1706.03762 [bsky, 0 points, 2 comments]
- You know computers do that now. arxiv.org/abs/1706.03762 [bsky, 0 points, 1 comments]
- [1706.03762] Attention Is All You Need | arXiv | Cornell University arxiv.org/abs/1706.03762 [bsky, 0 points, 1 comments]
- 💎 「Attention Is All You Need」の共著者なのも本当。 https://arxiv.org/abs/1706.03762 なので肩書き盛りのニュースではない。ちゃんと重い人事。 [bsky, 0 points, 1 comments]
- The core of any LLM and GPT in particular was described in 2017 by Google engineers. And this is one of the best reading of the year for me. arxiv.org/pdf/1706.037... [bsky, 0 points, 0 comments]
- Anyway if you’re actually interested, this is a starting point for quite a lot of research just to get up to par. arxiv.org/abs/1706.03762 If you’d like more of an overview, this is quite good: youtu. [bsky, 0 points, 0 comments]
- 5/ ¿Son extraordinamiente open-source? No, y sí. Toda estos modelos se basan en la arquitectura de transformadores, un desarrollo publicado en abierto por Google en 2017. Bastantes avances de Open-A [bsky, 0 points, 1 comments]
- The "Attention Is All You Need" article arxiv.org/pdf/1706.03762 [bsky, 0 points, 0 comments]
- The study that changed everything in ai/llm : google brain 2017 arxiv.org/abs/1706.037... [bsky, 0 points, 0 comments]
- Correct. LLMs are doing clever math. I’d direct you to a paper called ‘Attention is All You Need’ which explains detail This does not mean they can’t interpret contextual relevance. Once again I’ll d [bsky, 0 points, 0 comments]
- Sorry? What has quoting the name of the paper which kicked this all off (arxiv.org/abs/1706.03762) have to do with homophobia? You can apologise for that. I'm not having anyone libel my by wrongly cal [bsky, 0 points, 1 comments]
- (self) Attention is all you need #medinfo2025 arxiv.org/pdf/1706.03762 [bsky, 0 points, 0 comments]
- Nah. It's all because of one paper. arxiv.org/abs/1706.03762 The entire explosion of modern generative AI is because of that paper. [bsky, 0 points, 0 comments]
- Okay, Trump wants us to call it Cisformer from now on? arxiv.org/pdf/1706.03762 [bsky, 0 points, 1 comments]
- Ich meine, dass der Hype um die neue Modellfamilie o1 nicht übertrieben ist und diese Verbindung des Transformers mit dem Reinforcement Learning (siehe AlphaGo oder Robotik) einen ähnlich bedeutenden [bsky, 0 points, 1 comments]
- やっぱりLLMのこと理解するにはAttention Is All You Need読んどくべきなんか arxiv.org/pdf/1706.03762 [bsky, 0 points, 1 comments]
- 앞서 만든 분류기도 그렇지만 진짜 이 논문이 다 했다… 중간 로직을 만들지 않고, 집중할 데이터만 잘 제공하면 로직을 만들어줘! [bsky, 0 points, 0 comments]
Related