BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
2018/10/11 by Jacob Devlin, Devlin, Jacob, Ming-Wei Chang +5 · 7 voices · 3044 citations
#cs.CL
paper · pdf · doi:10.48550/arxiv.1810.04805
Abstract
We introduce a new language representation model called BERT, which stands for Bidirectional Encoder Representations from Transformers. Unlike recent language representation models, BERT is designed to pre-train deep bidirectional representations from unlabeled text by jointly conditioning on both left and right context in all layers. As a result, the pre-trained BERT model can be fine-tuned with just one additional output layer to create state-of-the-art models for a wide range of tasks, such as question answering and language inference, without substantial task-specific architecture modifications. BERT is conceptually simple and empirically powerful. It obtains new state-of-the-art results on eleven natural language processing tasks, including pushing the GLUE score to 80.5% (7.7% point absolute improvement), MultiNLI accuracy to 86.7% (4.6% absolute improvement), SQuAD v1.1 question answering Test F1 to 93.2 (1.5 point absolute improvement) and SQuAD v2.0 Test F1 to 83.1 (5.1 point absolute improvement).
Cited by
- Moral Semantics Survive Machine Translation: Cross-Lingual Evidence from Moral Foundations Corpora
- What Matters When Building Universal Multilingual Named Entity Recognition Models?
- LMEB: Long-horizon Memory Embedding Benchmark
- \kappa-LoRA: Condition Numbers Reveal Which LoRA Matrices Worth Updating
- FSE: Continual Learning for Named Entity Recognition by Fast-Slow Experts
- Unsupervised Multimodal Intent Discovery via MLLM-Guided Concept Generation and Semantic Propagation
- MedKGent: A Large Language Model Agent Framework for Constructing Temporally Evolving Medical Knowledge Graph
- IDSTune: A Multi-Agent Collaborative Framework for Integrated Database System Tuning
- Transforming Keystroke Noise to Text: Self-Supervised Acoustic Eavesdropping Attacks on Keyboards
- Reasoning Denoiser: Denoising Reasoning Traces for Hallucination Detection in Large Reasoning Models
- Do Transformers Actually Help Intrusion Detection? A Temporal Sequence Evaluation on CIC-IDS2017
- Improving Large Vision-Language Models' Understanding for Flow Field Data
- IQ-JEPA: A Joint-Embedding Predictive Architecture with a Hermitian Vision Transformer for Sound Speed and Attenuation Estimation from Ultrasound IQ Data
- Exact Neural-Network Representations of the Motzkin States
- Math Education Digital Shadows for Investigating Learning with GenAI: Mathematics Performance, Anxiety, and Confidence in LLMs
- Breaking the Data Barrier in Learning Symbolic Computation: A Case Study on Variable Ordering Suggestion for Cylindrical Algebraic Decomposition
- A Factorial Study of Synthetic Data Generation for Low-Resource Machine Translation using Grammar Books
- Vibe Coding: An Experiment with Test-Driven Development
- Encoding Invisible Causation for Bridge Diagnostic Agents: Triple-Guided Retrieval-Augmented Fine-Tuning with QLoRA
- A Unified Moral-Value Dataset for Instruction Tuning
- Hierarchical Self-Supervised Representation Learning Framework for Multivariate Time Series Grounded in ECG Analysis
- A machine learning benchmarking framework for lipid nanoparticle transfection efficiency prediction
- LaViDa: A Large Diffusion Language Model for Multimodal Understanding
- Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models
- MonteRET: AI Agent Enhancing Multimodal LLMs with Multi-granularity Knowledge Retrieval for Chest CT Report Generation
- Dysco: Dynamic Subspace Boosting to Mitigate LoRA Interference in Federated Learning
- Lavida-O: Elastic Large Masked Diffusion Models for Unified Multimodal Understanding and Generation
- Portfolio Optimization under Dynamic Rebalancing via Topological Data Analysis and News Sentiments
- From Static Bibliometrics to Dynamic Knowledge Graphs: An LLM-Powered Framework for Modernizing Science, Technology, and Innovation (STI) Analytics
- Efficient Nonlinear Multiscale Prediction for Unseen Polycrystalline Textures via Self-Supervised Microstructure Pretraining
- SHIFT: Self-reconstruction Harnesses Implicit Fine-grained Thinking for Retrieval
- LoRA-Tuned Large Language Models for Dementia Detection via Multi-View Speech-Derived Features
- Sentence Splitter: Uncovering Latent Factual Structure for Self-Supervised Learning
- Two-Step Occupation Coding
- OC-Distill: Ontology-aware Contrastive Learning with Cross-Modal Distillation for ICU Risk Prediction
- Echoes Across Vietnam's Highlands, Delta, and Coast: A Multilingual Corpus for Cham, Khmer, and Tay-Nung
- OLEDLM: A Unified Language Model for OLED Molecular Design
- Multi-modal transformer for signal classification in nanopore blockade experiments
- The Maskability Index: Predicting Task-Objective Alignment in Pretrained Language Models
- Enhancing Transformer-based Routing by Encoding Distance via Relative Positional Encoding
- OntoBook: Ontology-Grounded Synthetic Textbooks for Medical Encoder Pretraining
- FlexNGIA 2.0: Redesigning the Internet with Agentic AI -- Protocols, Services, and Traffic Engineering Designed, Deployed, and Managed by AI
- surprisal is Not a Theory
- Post-Training in Time Series Foundation Models: A Unifying Framework
- BitNet Text Embeddings
- FSDBN: Foreground-Aware EEG-Visual Alignment via Dynamic Brain Networks
- How Fast Can Reward Models Score? A Systems Study of C++ and PyTorch Inference Runtimes for RLHF
- ChineseBERT: Chinese Pretraining Enhanced by Glyph and Pinyin Information
- Glyce: Glyph-vectors for Chinese Character Representations
- A Self-Supervised Framework for Space Object Behaviour Characterisation
- Not All Errors Are Created Equal: ASCoT Addresses Late-Stage Fragility in Efficient LLM Reasoning
- From Bit-Position Sensitivity to Unequal Error Protection for DNN Inference Memory
- HyCoRec: Hypergraph-Enhanced Multi-Preference Learning for Alleviating Matthew Effect in Conversational Recommendation
- D3VL: Understanding Driving Scenes from 3D Time Series Data and Video with Language Models
- Teaching Robots to Interpret Social Interactions through Lexically-guided Dynamic Graph Learning
- Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery
- Team RAS in 11th ABAW Competition: Multimodal Ambivalence Recognition Approach
- Beyond Noisy Signals: Dual-Level Denoising for Multi-modal Sequential Recommendation
- Mitigating Matthew Effect: Multi-Hypergraph Boosted Multi-Interest Self-Supervised Learning for Conversational Recommendation
- AILQA: Evaluating AI-Driven Legal Question Answering Systems for the Indian Legal System
- Masked Visual Actions for Unified World Modeling
- Parameter-Efficient Continual Fine-Tuning: A Survey
- When Benchmarks Mislead: Shortcut Learning, Length Confounds, and the Limits of Cross-Dataset Generalization in Multilingual Fake News and Sarcasm Detection
- SALT: Salience-Aware Lexical Trie for Long-Context Compression
- One Model, Many Graphs: Learning over Attributed Graphs across Heterogeneous Modalities with Vision-Language Models
- Point Ladder Tuning: Parameter-Efficient Hierarchical Adaptation for 3D Point Cloud Understanding
- From Blind Search to Memory-Aware Evolution: Efficient DBMS Tuning via Collaborative Diagnosis and Utility-Aware Retrieval
- Tokenizing Crosslingual Homographs
- (A)iSpy: Parasitic Trojans for Machine Learning Infrastructure
- Sparse Evidence Can Suffice: Agentic Evidence Seeking for Multimodal Video Misinformation Detection
- Patch Policy: Efficient Embodied Control via Dense Visual Representations
- O-VAD: Industrial Video Anomaly Detection through Object-Centric Tracking and Reasoning
- Entity-Relation Extraction as Multi-Turn Question Answering
- DiffAxE: Diffusion-driven Hardware Accelerator Generation and Design Space Exploration
- Dice Loss for Data-imbalanced NLP Tasks
- A Unified MRC Framework for Named Entity Recognition
- RRAM-DP: Device-Calibrated Differential Privacy for In-Memory Edge Learning
- MotifRole-Diff: Risk-Optimal Role-Aware Corruption for Masked Molecular Graph Diffusion
- Beyond Parallel Tracking: Interactive Multi-Feature Fusion Drives Semantic Reconstruction from Non-invasive Brain Recordings
- BanClickThumb: A Multimodal Dataset and Transformer Fusion Benchmarks for Clickbait Detection in Bengali YouTube Videos
- Artificially intelligent agents in the social and behavioral sciences: A history and outlook
- Learning Sparse Representations of Multimodal Content for Enhanced Cold Item Recommendation
- Explaining and Tuning Transformer-based LLMs in Arithmetic Tasks with Human Strategies
- Dynamic Interaction-Aware and Causality-Disentangled Framework for Multimodal Sentiment Analysis
- FAME: Failure-Aware Mixture-of-Experts for Message-Level Log Anomaly Detection
- Do Speech Tokens Leak Voiceprints? Speaker Inversion Attacks Against End-to-End Speech Language Models
- Cognitive-YOLO: LLM-Driven Architecture Synthesis from First Principles of Data for Object Detection
- Gradient-Free Privacy Leakage in Federated Language Models through Selective Weight Tampering
- Measuring and Evaluating the Performance of Generative AI Models for Scam Detection
- Transition-Aware Backend Dispatch for Edge LLM Inference
- Taxonomy-Targeted Error Generation for Quantitative Reasoning
- Lost in the Tower of Babel: The Adverse Effects of Incidental Multilingualism in LLMs
- Evolution Strategy-Based Calibration for Low-Bit Quantization of Speech Models
- GlyRAG: Context-Aware Retrieval-Augmented Framework for Blood Glucose Forecasting
- Phantom Transitions in Language Model Fine-Tuning: A Density-Matrix Analysis
- BADTV: Unveiling Backdoor Threats in Third-Party Task Vectors
- Graph-Embedded Intuitionistic Fuzzy Broad Learning System: A Multi-view Framework
- Cascading versus Joint Modeling for Hierarchical Offensive Language Detection
- Prompt-Guided Foundation Model Tuning for Pathology Image Classification
- Brain-CLIPLM: Semantic Compression for EEG-to-Text Decoding
- Unsupervised Multimodal Clustering for Semantics Discovery in Multimodal Utterances
- Is One Score Enough? Assessing Singing Quality of Songs with Temporal Score Curves
- Dropout and Random Gradient Masking Are Asymptotically Equivalent in Large ResNets
- Dataset Distillation by Influence Matching
- GRASP: Grounded CoT Reasoning with Dual-Stage Optimization for Multimodal Sarcasm Target Identification
- Interactive Training 2: Auditable Control Plane for Live Model Training
- Controlling Implicit Shortcut Reliance in L2 Spoken English Auto-markers
- Posts of Peril: Detecting Information About Hazards in Text
- CAMMAR: Culture-Aware Matryoshka for Metaphorical Arabic Representations
- An expressivity analysis of hierarchical modelling in deep transformers via bounded-depth grammars
- Candidate Attended Dialogue State Tracking Using BERT
- Cost-efficient generative AI summarization for scalable automated essay scoring in educational assessment
- HCIG: A Hierarchical Cross-Modal Incongruity Graph Network for Multimodal Sarcasm and Cyberbullying Detection
- Not All Tokens Matter: Data-Centric Optimization for Efficient Code Summarization
- Speculative Decoding with a Speculative Vocabulary
- Conditional Reliability of Toxicity Signals for Multilingual and Code-Mixed Abuse Detection
- Knowledge-Centric Agents for Workflow Generation in ComfyUI
- Rate-Utility Frontiers for Language Encodings: Comparing Tokens, Bytes, and Pixels Under Controlled Linguistic Content
- Multimodality as Supervision: Self-Supervised Specialization to the Test Environment via Multimodality
- MxGPS: Multiplex Graph Transformers for a Power Grid Foundation Model
- Expanding the Lexicon of Ge'ez Based African Languages: A Comparative Study of Amharic and Tigrinya
- Bifocal Attention: Harmonizing Geometric and Spectral Positional Embeddings for Algorithmic Generalization
- CF-Net: Conflict Fusion with Speaker Normalisation and Certainty Weighting for Ambivalence/Hesitancy Recognition
- Dialogue Summarization with Emotion Dynamics Using Topic- and Participant-Centric Decomposition
- Robust Explanations for User Trust in Enterprise NLP Systems
- Pretrained Event Classification Model for High Energy Physics Analysis
- CrimeNER Demo: Named-Entity Recognition in the Crime Domain
- Does generative AI supersede supervised XMLC? A Benchmark Study on Automated Subject Indexing with German Scientific Literature
- PERL: Pinyin Enhanced Rephrasing Language Model for Chinese ASR N-best Error Correction
- ME-IQA: Memory-Enhanced Image Quality Assessment via Re-Ranking
- Advancing bioinformatics with language models: components, applications, and perspectives
- Token-Level Cross-Modal Transformer with Contrastive Multi-Task Learning for Breast Cancer Subtype Classification and Survival Prediction
- Leveraging Large Language Models for Generating Research Topic Ontologies: A Multi-Disciplinary Study
- LLM Unlearning for Cyber Defense: A Survey on Methods, Challenges, and Emerging Threats
- From Errors to Rules: Iterative Prompt Optimization for Text Classification
- Explaining Attention with Program Synthesis
- The Market in the Model: Latent Diffusion as Neural Economy
- AlphaOracle: Oracle bone script decipherment via human-workflow-inspired deep learning
- FBLayout: Optimizing Memory Layout for Efficient LLM Finetuning on Mobile GPUs
- Household Movement Detection in Mixed-Format Occupancy Data Using LLM-Based Entity Resolution
- What Makes Linguistic Representations Good Models of High-Level Visual Perception in the Human Brain?
- Enhancing next token prediction based pre-training for jet foundation models
- Multi-Stream LLMs: Unblocking Language Models with Parallel Streams of Thoughts, Inputs and Outputs
- LBA: Textual Hard-Label Adversarial Attack under Low Query Budgets
- A Survey on <scp>GNN</scp> ‐Based Link Prediction: Techniques, Applications, and Challenges
- Semantic Field Theory: Historical Origin, Higher-Order Interaction, and Stabilized Semantic Inference
- thaulab@EEUCA 2026: Who Said What to Whom? A Targeting-Aware Neural-Symbolic Pipeline for Gaming Toxicity Detection
- Shapley Context Pruning: A Cooperative Game Perspective for Context Reranking and Pruning
- Toward manifest relationality in transformers via symmetry reduction
- FineServe: A Fine-Grained Dataset and Characterization of Global LLM Serving Workloads
- How Open Must Language Models be to Enable Reliable Scientific Inference?
- PR-CAD: Progressive Refinement for Unified Controllable and Faithful Text-to-CAD Generation with Large Language Models
- Why AI systems don't learn and what to do about it: Lessons on autonomous learning from cognitive science
- From transcription to semantic corpus analysis: unsupervised learning of sentence representations for ancient languages
- Semantic Novelty Trajectories in 80,000 Books: A Cross-Corpus Embedding Analysis
- Generative AI & Fictionality: How Novels Power Large Language Models
- A Human-Centric Framework for Data Attribution in Large Language Models
- Strategies for Span Labeling with Large Language Models
- Who is transitioning to green? Introducing a text-based indicator to measure green skill transferability
- Remapping and navigation of an embedding space via error minimization: a fundamental organizational principle of cognition in natural and artificial systems
- Efficient Black-Box Fault Localization for System-Level Test Code Using Large Language Models
- On Success and Simplicity: A Second Look at Transferable Vision-Language Attack Pipeline
- Leveraging External Knowledge for Historical Document Restoration via Retrieval-Augmented Large Language Models
- LLM-Based Agentic Systems for Software Engineering: Challenges and Opportunities
- Compressing What Matters: Neuron Importance Meets Data-Aware Low Rank Approximation for Language Model Compression
- A common framework for semantic memory and semantic composition
- CIG-MAE: Cross-Modal Information-Guided Masked Autoencoder for Self-Supervised WiFi Sensing
- LLM Essay Scoring Under Holistic and Analytic Rubrics: Prompt Effects and Bias
- In-Context Probing for Membership Inference in Fine-Tuned Language Models
- The Mean-Field Dynamics of Transformers
- Epistemological Fault Lines Between Human and Artificial Intelligence
- MuM: Multi-View Masked Image Modeling for 3D Vision
- The ‘design features’ of language revisited
- Mind captioning: Evolving descriptive text of mental content from human brain activity
- Effectiveness of LLMs in Temporal User Profiling for Recommendation
- Jasmine: A Simple, Performant and Scalable JAX-based World Modeling Codebase
- Platonic Transformers: A Solid Choice For Equivariance
- Investigating the Representation of Backchannels and Fillers in Fine-tuned Language Models
- Bootstrapping Task Spaces for Self-Improvement
- Towards a Physics Foundation Model
- Understanding Reinforcement Learning for Model Training, and future directions with GRAPE
- World Modeling with Probabilistic Structure Integration
- The aesthetics of climate misinformation: computational multimodal framing analysis with BERTopic and CLIP
- Constructions are Revealed in Word Distributions
- Arnold: a generalist muscle transformer policy
- Sketch-of-Thought: Efficient LLM Reasoning with Adaptive Cognitive-Inspired Sketching
- Every FLOP Counts: Scaling a 300B Mixture-of-Experts LING LLM without Premium GPUs
- A Survey on Diffusion Language Models
- Fast weight programming and linear transformers: from machine learning to neurobiology
- Assessing LLM Text Detection in Educational Contexts: Does Human Contribution Affect Detection?
- Cameras as Relative Positional Encoding
- Use as Directed? A Comparison of Software Tools Intended to Check Rigor and Transparency of Published Work
- Token Bottleneck: One Token to Remember Dynamics
- Position: Text Embeddings Should Capture Implicit Semantics, Not Just Surface Meaning
- QiMeng: Fully Automated Hardware and Software Design for Processor Chip
- Forgetful by design? A critical audit of YouTube’s search API for academic research
- JavelinGuard: Low-Cost Transformer Architectures for LLM Security
- Leaner Transformers: More Heads, Less Depth
- Breaking Quadratic Barriers: A Non-Attention LLM for Ultra-Long Context Horizons
- From Tokens to Thoughts: How LLMs and Humans Trade Compression for Meaning
- Diffusion vs. Autoregressive Language Models: A Text Embedding Perspective
- El Agente: An autonomous agent for quantum chemistry
- Do Language Models Know Who Did What to Whom?
- Neurosymbolic Diffusion Models
- ModernBERT or DeBERTaV3? Examining Architecture and Data Influence on Transformer Encoder Models Performance
- Optimizing Biophysical Large-Scale Brain Circuit Models With Deep Neural Networks
- MultiBLiMP 1.0: A Massively Multilingual Benchmark of Linguistic Minimal Pairs
- From superposition to sparse codes: interpretable representations in neural networks
- VGGT: Visual Geometry Grounded Transformer
- Probing LLMs for Multilingual Discourse Generalization Through a Unified Label Set
- MatplotAlt: A Python Library for Adding Alt Text to Matplotlib Figures in Computational Notebooks
- Foundation neural-networks quantum states as a unified Ansatz for multiple hamiltonians
- Smaller But Better: Unifying Layout Generation with Smaller Large Language Models
- Linguistic Generalizations are not Rules: Impacts on Evaluation of LMs
- Pragmatics in the Era of Large Language Models: A Survey on Datasets, Evaluation, Opportunities and Challenges
- LM2: Large Memory Models
- Fast3R: Towards 3D Reconstruction of 1000+ Images in One Forward Pass
- Vision-Enhanced Large Language Models for High-Resolution Image Synthesis and Multimodal Data Interpretation
- Context-Selective State Space Models: Feedback is All You Need
- Revisiting Theory of Contrastive Learning for Domain Generalization
- High-Performance Self-Supervised Learning by Joint Training of Flow Matching
- It's LIT! Reliability-Optimized LLMs with Inspectable Tools
- Fake News Classification in Urdu: A Domain Adaptation Approach for a Low-Resource Language
- VL-RouterBench: A Benchmark for Vision-Language Model Routing
- Targeted Linguistic Analysis of Sign Language Models with Minimal Translation Pairs
- MetFuse: Figurative Fusion between Metonymy and Metaphor
- Securing the AI Supply Chain: What Can We Learn From Developer-Reported Security Issues and Solutions of AI Projects?
- CountGD++: Generalized Prompting for Open-World Counting
- RobustMask: Certified Robustness against Adversarial Neural Ranking Attack via Randomized Masking
- Agentic Physical AI toward a Domain-Specific Foundation Model for Energy Systems: A Case Study on Nuclear Reactor Control
- Chinese Morph Resolution in E-commerce Live Streaming Scenarios
- Text-Routed Sparse Mixture-of-Experts Model with Explanation and Temporal Alignment for Multi-Modal Sentiment Analysis
- Diffusion-based Decentralized Federated Multi-Task Representation Learning
- Graph Neural Networks with Transformer Fusion of Brain Connectivity Dynamics and Tabular Data for Forecasting Future Tobacco Use
- Rethinking Fine-Tuning: Unlocking Hidden Capabilities in Vision-Language Models
- Enhancing knowledge graph interactions: A comprehensive Text-to-Cypher pipeline with large language models
- Fusion or Confusion? Multimodal Complexity Is Not All You Need
- Embodied Robot Manipulation in the Era of Foundation Models: Planning and Learning Perspectives
- Let Samples Speak: Mitigating Spurious Correlation by Exploiting the Clusterness of Samples
- Bridging Global Intent with Local Details: A Hierarchical Representation Approach for Semantic Validation in Text-to-SQL
- Discovering Transmission Dynamics of COVID-19 in China
- GHaLIB: A Multilingual Framework for Hope Speech Detection in Low-Resource Languages
- Beyond Centralization: Provable Communication Efficient Decentralized Multi-Task Learning
- Chain-of-thought Reviewing and Correction for Time Series Question Answering
- Raven: Mining Defensive Patterns in Ethereum via Semantic Transaction Revert Invariants Categories
- Learning When Not to Attend Globally
- TimePerceiver: An Encoder-Decoder Framework for Generalized Time-Series Forecasting
- SPECTRE: Spectral Pre-training Embeddings with Cylindrical Temporal Rotary Position Encoding for Fine-Grained sEMG-Based Movement Decoding
- LLM-Based vs. Lexicon-Based Sentiment Signals for Tail-Risk Detection in Meme Stocks
- HalluMat: Detecting Hallucinations in LLM-Generated Materials Science Content Through Multi-Stage Verification
- Meta-information Guided Cross-domain Synergistic Diffusion Model for Low-dose PET Reconstruction
- VideoScaffold: Elastic-Scale Visual Hierarchies for Streaming Video Understanding in MLLMs
- Greedy dynamical meta-learning
- Towards High-Level Semantic Intelligence
- RayRoPE: Projective Ray Positional Encoding for Multi-view Attention
- Choosing a Text Embedding Model: A Practical Benchmarking and Decision Framework
- ATLAS: Automated Approximation of Transformers for Efficient Homomorphic Inference in One Hour
- seqLens: Optimizing Language Models for Genomic Predictions
- The Cross-Domain Generalization Cost of Offensive Language Detection
- BERT-based Models vs. Large Language Models for Low-Resource Named Entity Recognition: A Comparative Study on Marathi
- The JEPA Paradox in Language: The Geometry of Linguistic Alternatives
- LA-RL: Label-Aware Self-Reflection for Reinforcement Learning in Information Extraction
- HiTMS: A High-Throughput Multi-Stream Linguistic Steganography Framework
- Enhancing Code Understanding for Impact Analysis by Combining Transformers and Program Dependence Graphs
- Controllable Diversity in Normalization-Based Implicit Ensembles via Softmax-Temperature Modulation
- TriShieldRAG: A Three-Ring Defense-in-Depth Framework Against Knowledge Corruption in Retrieval-Augmented Generation
- From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement
- Asleep at the Keyboard? Assessing the Security of GitHub Copilot's Code Contributions
- Generative AI for Requirements Engineering: A Systematic Literature Review
- A Scoping Review of Machine Learning Applications in Power System Protection and Disturbance Management
- Decoding Phone Pairs from MEG Signals Across Speech Modalities
- Resource-sensitive but language-blind: Community size and not grammatical complexity better predicts the accuracy of Large Language Models in a novel Wug Test
- SpecFormer: Mitigating Embedding and Attention Collapse via Spectral-Aware Transformer for Recommendation
- The Tokenizer Tax: Quantifying and Explaining the Cross-Lingual Cost of Subword Tokenization for Indian Languages
- A Coulomb Particle Model for Learning Kernel Attention in Transformers
- BHARATI: Morphology-Aware Tokenizers for Classical Indian Languages with Subword Fertility Analysis
- Speech Signals Complement LLMs for Predicting Interpersonal Attraction in Speed Dating
- Where meaning lives: Layer-wise accessibility of psycholinguistic features in encoder and decoder language models
- A Cross-lingual Comparison of Human and Classification Model Entrainment Behavior in Code-switched Speech Settings
- Automated Numerical Stability Analysis of Deep Learning Operators
- PreDiff-LM: Pretrained Discrete Masked Diffusion Language Modeling with Hybrid Attention
- Decompose and Reorganize: Planning with Primitives and Visuomotor Policies Learned from Demonstrations
- Structural Analysis of Journal Columns Using Ordinal Patterns and Information-Theoretic Measures
- Text-Prompted CLAP: Learning Query-Conditioned Audio Representations via Contrastive Learning
- Inspect India Evals: An Open Benchmarking Framework for Evaluating Large Language Models in the Indian Linguistic and Cultural Context
- Room-Mediated Co-occurrence for Zero-Shot Object-Centric Semantic Navigation via Frontier Scoring
- Bekko Embedding: Parameter-Efficient Multilingual Retrieval with Ultra-Compact Encoders
- DisasterTD: Disaster Toponym Disambiguation Using Multimodal LLMs and Cross-View Geolocalization
- Beyond Predictive Accuracy: A Reliability-Aware Audit of Molecular Representations for Human Olfaction
- Beam-Response Contrastive Learning for Transmitter-Side MIMO CSI Representation
- LongFly: Long-Horizon UAV Vision-and-Language Navigation with Spatiotemporal Context Integration
- Teaching Tiny VLA Models Where to Look and How to Move
- Multimodal Domain Generalization for Depression Detection: An Attention-Based BiLSTM Network With Domain-Adversarial Training
- LC-SEPLM: long-range contact-supervised adaptation for sequence-only protein representation learning
- Optimizing Transformer Neural Network for Real-Time Outlier Detection on FPGAs
- Real-time Reconstruction of Human Visual Perception from fMRI
- Open Your Model’s Eyes: Video and Context-Aware Multimodal Backchannel Prediction
- Imprompt: A Language Framework for Prompt Programming
- DOSA: A Tree-Guided, Self-Regressive Framework for Long Document Structure Analysis
- Reading Without a Reader: Large Language Models Collapse Reading and Writing into a Single Entangled Code
- Unified Semantic Modeling Framework for Large-Scale Job Understanding at LinkedIn
- GrocLM: Grocery Category Recommendation in E-Commerce with Large Language Models
- Evolving from Lessons: Skill-Augmented Table Graph Reasoning for Operation-wise Table Question Answering
- CHS-SQL: A Text-to-SQL approach based on Confidence-Guided Heuristic Search Schema Linking process
- Opti-Q: A Constraint-Based Optimization Framework for Multi-LLM Question Planning
- Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy
- DocAnnot -- Accelerating the Creation of Key Information Extraction Datasets with GenAI-Powered Auto-annotation
- TimeCapsule: Generative Hallucination as a Method for Historical Sensemaking
- Bridging MARL to SARL: An Order-Independent Multi-Agent Transformer via Latent Consensus
- Exact and Asymptotically Complete Robust Verifications of Neural Networks via Ising Solvers
- Self-Distillation of Hidden Layers for Self-Supervised Representation Learning
- SkipOPU: An FPGA-based Overlay Processor for Large Language Models with Dynamically Allocated Computation
- Asymmetric Hierarchical Anchoring for Robust Audio-Visual Cross-Modal Generalization
- PRIMA: Pre-Training with Risk-Integrated Image--Metadata Alignment for Medical Diagnosis with LLM-Based Feature Aggregation
- SERM: Self-Evolving Relevance Model with Agent-Driven Learning from Massive Query Streams
- Reconstructing Item Characteristic Curves using Fine-Tuned Large Language Models
- Explainable Statute Prediction via Attention-based Model and LLM Prompting
- Self-attention vector output similarities reveal how machines pay attention
- MMCTOP: A Multimodal Textualization and Mixture-of-Experts Framework for Clinical Trial Outcome Prediction
- DPAR: Dynamic Patchification for Efficient Autoregressive Visual Generation
- TimeBill: Time-Budgeted Inference for Large Language Models
- Approximation Capabilities of Feedforward Neural Networks with GELU Activations
- Detecting AI-Generated Paraphrases in Bengali: A Comparative Study of Zero-Shot and Fine-Tuned Transformers
- Cross-Semantic Transfer Learning for High-Dimensional Linear Regression
- Rethinking Output Alignment For 1-bit Post-Training Quantization of Large Language Models
- animal2vec and MeerKAT: A self-supervised transformer for rare-event raw audio input and a large-scale reference dataset for bioacoustics
- CATCH: A Controllable Theme Detection Framework with Contextualized Clustering and Hierarchical Generation
- Dynamic Attention (DynAttn): Interpretable High-Dimensional Spatio-Temporal Forecasting (with Application to Conflict Fatalities)
- GraviBERT: Transformer-based inference for gravitational-wave time series
- Fast SAM2 with Text-Driven Token Pruning
- Towards Practical Automatic Piano Reduction using BERT with Semi-supervised Learning
- ReaSeq: Unleashing World Knowledge via Reasoning for Sequential Modeling
- Beyond Context: Large Language Models Failure to Grasp Users Intent
- Semi-Supervised Learning for Large Language Models Safety and Content Moderation
- Uncovering Hierarchical Structure in LLM Embeddings with δ-Hyperbolicity, Ultrametricity, and Neighbor Joining
- Blurb-Refined Inference from Crowdsourced Book Reviews using Hierarchical Genre Mining with Dual-Path Graph Convolutions
- Language-Guided Grasp Detection with Coarse-to-Fine Learning for Robotic Manipulation
- When LLMs fall short in Deductive Coding: Model Comparison and Human AI Collaboration Workflow Design
- Learning from Next-Frame Prediction: Autoregressive Video Modeling Encodes Effective Representations
- CoSeNet: A Novel Approach for Optimal Segmentation of Correlation Matrices
- AegisAgent: An Autonomous Defense Agent Against Prompt Injection Attacks in LLM-HARs
- Diving into 3D Parallelism with Heterogeneous Spot Instance GPUs: Design and Implications
- ETP-R1: Evolving Topological Planning with Reinforcement Fine-tuning for Vision-Language Navigation in Continuous Environments
- PanoGrounder: Bridging 2D and 3D with Panoramic Scene Representations for VLM-based 3D Visual Grounding
- SegMo: Segment-aligned Text to 3D Human Motion Generation
- TokSuite: Measuring the Impact of Tokenizer Choice on Language Model Behavior
- Advancing Multimodal Teacher Sentiment Analysis:The Large-Scale T-MED Dataset & The Effective AAM-TSA Model
- Benchmarking LLMs for Predictive Applications in the Intensive Care Units
- Toward Explaining Large Language Models in Software Engineering Tasks
- SpidR: Learning Fast and Stable Linguistic Units for Spoken Language Models Without Supervision
- Designing Spatial Architectures for Sparse Attention: STAR Accelerator via Cross-Stage Tiling
- A Novel Graph-Sequence Learning Model for Inductive Text Classification
- Jensen-Shannon Divergence Message-Passing for Rich-Text Graph Representation Learning
- Adaptive Financial Sentiment Analysis for NIFTY 50 via Instruction-Tuned LLMs , RAG and Reinforcement Learning Approaches
- Reason2Decide: Rationale-Driven Multi-Task Learning
- Beyond Vision: Contextually Enriched Image Captioning with Multi-Modal Retrieval
- TAVID: Text-Driven Audio-Visual Interactive Dialogue Generation
- Making Large Language Models Efficient Dense Retrievers
- LADLE-MM: Limited Annotation based Detector with Learned Ensembles for Multimodal Misinformation
- Vehicle-centric Perception via Multimodal Structured Pre-training
- Attention Is Not What You Need
- Multimodal LLMs for Historical Dataset Construction from Archival Image Scans: German Patents (1877-1918)
- The Epistemological Consequences of Large Language Models: Rethinking collective intelligence and institutional knowledge
- Event Extraction in Large Language Model
- Multi-Modal Soccer Scene Analysis with Masked Pre-Training
- R-GenIMA: Integrating Neuroimaging and Genetics with Interpretable Multimodal AI for Alzheimer's Disease Progression
- CienaLLM: Generative Climate-Impact Extraction from News Articles with Autoregressive LLMs
- MAGIC: Achieving Superior Model Merging via Magnitude Calibration
- HARMON-E: Hierarchical Agentic Reasoning for Multimodal Oncology Notes to Extract Structured Data
- Can LLMs Estimate Student Struggles? Human-AI Difficulty Alignment with Proficiency Simulation for Item Difficulty Prediction
- How Many Experts Are Enough? Towards Optimal Semantic Specialization for Mixture-of-Experts
- Overcoming Spectral Bias via Cross-Attention
- A Comparative Study of Light-weight Language Models for PII Masking and their Deployment for Real Conversational Texts
- The Interaction Bottleneck of Deep Neural Networks: Discovery, Proof, and Modulation
- Revisiting the Learning Objectives of Vision-Language Reward Models
- Object-Centric Framework for Video Moment Retrieval
- InstructNet: A Novel Approach for Multi-Label Instruction Classification through Advanced Deep Learning
- Probabilistic Digital Twins of Users: Latent Representation Learning with Statistically Validated Semantics
- Adversarial Robustness of Vision in Open Foundation Models
- Disentangling Fact from Sentiment: A Dynamic Conflict-Consensus Framework for Multimodal Fake News Detection
- Towards Deeper Emotional Reflection: Crafting Affective Image Filters with Generative Priors
- Physics of Language Models: Part 4.1, Architecture Design and the Magic of Canon Layers
- Toward Ethical AI Through Bayesian Uncertainty in Neural Question Answering
- Attention Distance: A Novel Metric for Directed Fuzzing with Large Language Models
- Confidence-Credibility Aware Weighted Ensembles of Small LLMs Outperform Large LLMs in Emotion Detection
- DNAMotifTokenizer: Towards Biologically Informed Tokenization of Genomic Sequences
- Digitizing Nepal's Written Heritage: A Comprehensive HTR Pipeline for Old Nepali Manuscripts
- Next-Embedding Prediction Makes Strong Vision Learners
- Posterior Behavioral Cloning: Pretraining BC Policies for Efficient RL Finetuning
- TOGGLE: Temporal Logic-Guided Large Language Model Compression for Edge
- LLMCache: Layer-Wise Caching Strategies for Accelerated Reuse in Transformer Inference
- What Do Prosody and Text Convey? Characterizing How Meaningful Information is Distributed Across Multiple Channels
- The Colombian legislative process, 2014-2025: networks, topics, and polarization
- GinSign: Grounding Natural Language Into System Signatures for Temporal Logic Translation
- Abacus: Self-Supervised Event Counting-Aligned Distributional Pretraining for Sequential User Modeling
- From Essence to Defense: Adaptive Semantic-aware Watermarking for Embedding-as-a-Service Copyright Protection
- Hypernetworks That Evolve Themselves
- DualGuard: Dual-stream Large Language Model Watermarking Defense against Paraphrase and Spoofing Attack
- Decoding Fake Narratives in Spreading Hateful Stories: A Dual-Head RoBERTa Model with Multi-Task Learning
- ModelTables: A Corpus of Tables about Models
- Vision-Language-Action Models for Autonomous Driving: Past, Present, and Future
- See It Before You Grab It: Deep Learning-based Action Anticipation in Basketball
- From Words to Wavelengths: VLMs for Few-Shot Multispectral Object Detection
- In-Context Semi-Supervised Learning
- In Pursuit of Pixel Supervision for Visual Pre-training
- ArcBERT: An LLM-based Search Engine for Exploring Integrated Multi-Omics Metadata
- The Deleuzian Representation Hypothesis
- Leveraging Foundational Models and Simple Fusion for Multi-modal Physiological Signal Analysis
- RFKG-CoT: Relation-Driven Adaptive Hop-count Selection and Few-Shot Path Guidance for Knowledge-Aware QA
- No More Hidden Pitfalls? Exposing Smart Contract Bad Practices with LLM-Powered Hybrid Analysis
- Explainable Action Form Assessment by Exploiting Multimodal Chain-of-Thoughts Reasoning
- Embedding Software Intent: Lightweight Java Module Recovery
- Beyond Fast and Slow: Cognitive-Inspired Elastic Reasoning for Large Language Models
- DASH: Dialogue-Aware Similarity and Handshake Recognition for Topic Segmentation in Public-Channel Conversations
- SeBERTis: A Framework for Producing Classifiers of Security-Related Issue Reports
- Spatia: Video Generation with Updatable Spatial Memory
- On the Use of Self-Supervised Representation Learning for Speaker Diarization and Separation
- Parameter Efficient Multimodal Instruction Tuning for Romanian Vision Language Models
- Task Matrices: Linear Maps for Cross-Model Finetuning Transfer
- Focus: A Streaming Concentration Architecture for Efficient Vision-Language Models
- TiME: Tiny Monolingual Encoders for Efficient NLP Pipelines
- ParaFormer: A Generalized PageRank Graph Transformer for Graph Representation Learning
- Towards Nepali-language LLMs: Efficient GPT training with a Nepali BPE tokenizer
- Agreement Between Large Language Models and Human Raters in Essay Scoring: A Research Synthesis
- Dual-objective Language Models: Training Efficiency Without Overfitting
- Explainable Ethical Assessment on Human Behaviors by Generating Conflicting Social Norms
- Optimizing the Adversarial Perturbation with a Momentum-based Adaptive Matrix
- On Improving Deep Active Learning with Formal Verification
- IntentMiner: Intent Inversion Attack via Tool Call Analysis in the Model Context Protocol
- Citation importance-aware document representation learning for large-scale science mapping
- Unleashing the Power of Image-Tabular Self-Supervised Learning via Breaking Cross-Tabular Barriers
- Large-Language Memorization During the Classification of United States Supreme Court Cases
- Verifying Rumors via Stance-Aware Structural Modeling
- FiNERweb: Datasets and Artifacts for Scalable Multilingual Named Entity Recognition
- Beyond surface form: A pipeline for semantic analysis in Alzheimer's Disease detection from spontaneous speech
- ReFusion: A Diffusion Large Language Model with Parallel Autoregressive Decoding
- From Zipf's Law to Neural Scaling through Heaps' Law and Hilberg's Hypothesis
- Non-Resolution Reasoning (NRR): A Computational Framework for Contextual Identity and Ambiguity Preservation
- SocialNav-MoE: A Mixture-of-Experts Vision Language Model for Socially Compliant Navigation with Reinforcement Fine-Tuning
- Learning to Retrieve with Weakened Labels: Robust Training under Label Noise
- A Simple and Effective Framework for Symmetric Consistent Indexing in Large-Scale Dense Retrieval
- Forging a Dynamic Memory: Retrieval-Guided Continual Learning for Generalist Medical Foundation Models
- Alada: Alternating Adaptation of Momentum Method for Memory-Efficient Matrix Optimization
- Are Large Language Models Really Effective for Training-Free Cold-Start Recommendation?
- Let the Model Learn to Feel: Mode-Guided Tonality Injection for Symbolic Music Emotion Recognition
- Investigating Data Pruning for Pretraining Biological Foundation Models at Scale
- SPAR: Session-based Pipeline for Adaptive Retrieval on Legacy File Systems
- PrahokBART: A Pre-trained Sequence-to-Sequence Model for Khmer Natural Language Generation
- GenieDrive: Towards Physics-Aware Driving World Model with 4D Occupancy Guided Video Generation
- Fine-Tuning Causal LLMs for Text Classification: Embedding-Based vs. Instruction-Based Approaches
- One Leak Away: How Pretrained Model Exposure Amplifies Jailbreak Risks in Finetuned LLMs
- StegaVAR: Privacy-Preserving Video Action Recognition via Steganographic Domain Analysis
- Detecting Prompt Injection Attacks Against Application Using Classifiers
- Supervised Contrastive Frame Aggregation for Video Representation Learning
- Effective Fine-Tuning with Eigenvector Centrality Based Pruning
- NagaNLP: Bootstrapping NLP for Low-Resource Nagamese Creole with Human-in-the-Loop Synthetic Data
- SafeGen: Embedding Ethical Safeguards in Text-to-Image Generation
- Taint-Based Code Slicing for LLMs-based Malicious NPM Package Detection
- Semantic Distance Measurement based on Multi-Kernel Gaussian Processes
- Rethinking Label Consistency of In-Context Learning: An Implicit Transductive Label Propagation Perspective
- BaRISTA: Brain Scale Informed Spatiotemporal Representation of Human Intracranial Neural Activity
- Citation-Grounded Code Comprehension: Preventing LLM Hallucination Through Hybrid Retrieval and Graph-Augmented Context
- SCIR: A Self-Correcting Iterative Refinement Framework for Enhanced Information Extraction Based on Schema
- Epistemoverse: Toward an AI-Driven Knowledge Metaverse for Intellectual Heritage Preservation
- EditMGT: Unleashing Potentials of Masked Generative Transformers in Image Editing
- DREAM-B3P: Dual-Stream Transformer Network Enhanced by Feedback Diffusion Model for Blood-Brain Barrier Penetrating Peptide Prediction
- PhraseVAE and PhraseLDM: Latent Diffusion for Full-Song Multitrack Symbolic Music Generation
- A Scalable Multi-GPU Framework for Encrypted Large-Model Inference
- Multi-Intent Spoken Language Understanding: Methods, Trends, and Challenges
- CADKnitter: Compositional CAD Generation from Text and Geometry Guidance
- SciLaD: A Large-Scale, Transparent, Reproducible Dataset for Natural Scientific Language Processing
- Cross-modal Context-aware Learning for Visual Prompt Guided Multimodal Image Understanding in Remote Sensing
- An Anatomy of Vision-Language-Action Models: From Modules to Milestones and Challenges
- Automating Historical Insight Extraction from Large-Scale Newspaper Archives via Neural Topic Modeling
- Explanation Bias is a Product: Revealing the Hidden Lexical and Position Preferences in Post-Hoc Feature Attribution
- TAO-Net: Two-stage Adaptive OOD Classification Network for Fine-grained Encrypted Traffic Classification
- E-RayZer: Self-supervised 3D Reconstruction as Spatial Visual Pre-training
- Exposing Pink Slime Journalism: Linguistic Signatures and Robust Detection Against LLM-Generated Threats
- ASK: Adaptive Self-improving Knowledge Framework for Audio Text Retrieval
- Causal Reasoning Favors Encoders: On The Limits of Decoder-Only Models
- Error-Propagation-Free Learned Video Compression With Dual-Domain Progressive Temporal Alignment
- Semantic Reconstruction of Adversarial Plagiarism: A Context-Aware Framework for Detecting and Restoring "Tortured Phrases" in Scientific Literature
- Local LLM Ensembles for Zero-shot Portuguese Named Entity Recognition
- SemanticBBV: A Semantic Signature for Cross-Program Knowledge Reuse in Microarchitecture Simulation
- Grounding Everything in Tokens for Multimodal Large Language Models
- Decoding Student Minds: Leveraging Conversational Agents for Psychological and Learning Analysis
- Token Sample Complexity of Attention
- SEMDICE: Off-policy State Entropy Maximization via Stationary Distribution Correction Estimation
- SCOUT: A Defense Against Data Poisoning Attacks in Fine-Tuned Language Models
- Defining Cost Function of Steganography with Large Language Models
- Interpreto: An Explainability Library for Transformers
- Straggler Tolerant and Resilient DL Training on Homogeneous GPUs
- CLARGA: Multimodal Graph Representation Learning over Arbitrary Sets of Modalities
- Content-Adaptive Image Retouching Guided by Attribute-Based Text Representation
- Hands-on Evaluation of Visual Transformers for Object Recognition and Detection
- Representation Invariance and Allocation: When Subgroup Balance Matters
- Semantic Geometry for policy-constrained interpretation
- ODMA: On-Demand Memory Allocation Framework for LLM Serving on LPDDR-Class Accelerators
- Belief Is All You Need: Modeling Narrative Archetypes in Conspiratorial Discourse
- Stanford Sleep Bench: Evaluating Polysomnography Pre-training Methods for Sleep Foundation Models
- HPM-KD: Hierarchical Progressive Multi-Teacher Framework for Knowledge Distillation and Efficient Model Compression
- Learning Unmasking Policies for Diffusion Language Models
- Graph Deep Learning for Intracranial Aneurysm Blood Flow Simulation and Risk Assessment
- Ask, Answer, and Detect: Role-Playing LLMs for Personality Detection with Question-Conditioned Mixture-of-Experts
- Quantum Decision Transformers (QDT): Synergistic Entanglement and Interference for Offline Reinforcement Learning
- Adaptive Regularized Newton Method with Inexact Hessian
- The Erosion of LLM Signatures: Can We Still Distinguish Human and LLM-Generated Scientific Ideas After Iterative Paraphrasing?
- Beyond Detection: A Comprehensive Benchmark and Study on Representation Learning for Fine-Grained Webshell Family Classification
- Enhancing Clinical Note Generation with ICD-10, Clinical Ontology Knowledge Graphs, and Chain-of-Thought Prompting Using GPT-4
- PointDico: Contrastive 3D Representation Learning Guided by Diffusion Models
- Hybrid Attribution Priors for Explainable and Robust Model Training
- Is GPT-OSS All You Need? Benchmarking Large Language Models for Financial Intelligence and the Surprising Efficiency Paradox
- Beyond Traditional Diagnostics: Transforming Patient-Side Information into Predictive Insights with Knowledge Graphs and Prototypes
- VisKnow: Constructing Visual Knowledge Base for Object Understanding
- A scalable and real-time neural decoder for topological quantum codes
- No Labels, No Problem: Training Visual Reasoners with Multimodal Verifiers
- Luxical: High-Speed Lexical-Dense Text Embeddings
- Segment, Embed, and Align: A Universal Recipe for Aligning Subtitles to Signing
- CAMO: Causality-Guided Adversarial Multimodal Domain Generalization for Crisis Classification
- Bridging Code Graphs and Large Language Models for Better Code Understanding
- HOLE: Homological Observation of Latent Embeddings for Neural Network Interpretability
- An Empirical Framework for Evaluating Semantic Preservation Using Hugging Face
- In-Context and Few-Shots Learning for Forecasting Time Series Data based on Large Language Models
- MoCoRP: Modeling Consistent Relations between Persona and Response for Persona-based Dialogue
- MIDG: Mixture of Invariant Experts with knowledge injection for Domain Generalization in Multimodal Sentiment Analysis
- Amulet: Fast TEE-Shielded Inference for On-Device Model Protection
- Persian-Phi: Efficient Cross-Lingual Adaptation of Compact LLMs via Curriculum Learning
- How Far are Modern Trackers from UAV-Anti-UAV? A Million-Scale Benchmark and New Baseline
- PrivORL: Differentially Private Synthetic Dataset for Offline Reinforcement Learning
- Generalized Referring Expression Segmentation on Aerial Photos
- Generating Storytelling Images with Rich Chains-of-Reasoning
- Attention as Binding: A Vector-Symbolic Perspective on Transformer Reasoning
- Self-Supervised Learning on Molecular Graphs: A Systematic Investigation of Masking Design
- Cross-platform Product Matching Based on Entity Alignment of Knowledge Graph with RAEA model
- Pay Less Attention to Function Words for Free Robustness of Vision-Language Models
- Progress Ratio Embeddings: An Impatience Signal for Robust Length Control in Neural Text Generation
- A Patient-Doctor-NLP-System to contest inequality for less privileged
- ProAgent: Harnessing On-Demand Sensory Contexts for Proactive LLM Agent Systems
- Distribution-Aware Exploration for Adaptive HNSW Search
- Selective Masking based Self-Supervised Learning for Image Semantic Segmentation
- Patronus: Identifying and Mitigating Transferable Backdoors in Pre-trained Language Models
- Always Keep Your Promises: A Model-Agnostic Attribution Algorithm for Neural Networks
- Transferring Clinical Knowledge into ECGs Representation
- CMV-Fuse: Cross Modal-View Fusion of AMR, Syntax, and Knowledge Representations for Aspect Based Sentiment Analysis
- TopiCLEAR: Topic extraction by CLustering Embeddings with Adaptive dimensional Reduction
- Hybrid Quantum-Classical Ensemble Learning for S&P 500 Directional Prediction
- Classifying German Language Proficiency Levels Using Large Language Models
- BERTO: an Adaptive BERT-based Network Time Series Predictor with Operator Preferences in Natural Language
- AgenticCyber: A GenAI-Powered Multi-Agent System for Multimodal Threat Detection and Adaptive Response in Cybersecurity
- Heard or Halted? Gender, Interruptions, and Emotional Tone in U.S. Supreme Court Oral Arguments
- EgoEdit: Dataset, Real-Time Streaming Model, and Benchmark for Egocentric Video Editing
- From Text to Returns: Using Large Language Models for Mutual Fund Portfolio Optimization and Risk-Adjusted Allocation
- Retrieving Semantically Similar Decisions under Noisy Institutional Labels: Robust Comparison of Embedding Methods
- Ontology Learning with LLMs: A Benchmark Study on Axiom Identification
- The Road of Adaptive AI for Precision in Cybersecurity
- Lyrics Matter: Exploiting the Power of Learnt Representations for Music Popularity Prediction
- Text Rationalization for Robust Causal Effect Estimation
- Transformer-Enabled Diachronic Analysis of Vedic Sanskrit: Neural Methods for Quantifying Types of Language Change
- AI & Human Co-Improvement for Safer Co-Superintelligence
- Gradient Descent with Provably Tuned Learning-rate Schedules
- Fine-Tuning BERT for Domain-Specific Question Answering: Toward Educational NLP Resources at University Scale
- GeoPE:A Unified Geometric Positional Embedding for Structured Tensors
- LLMs Know More Than Words: A Genre Study with Syntax, Metaphor & Phonetics
- Are LLMs Truly Multilingual? Exploring Zero-Shot Multilingual Capability of LLMs for Information Retrieval: An Italian Healthcare Use Case
- Order Matters: 3D Shape Generation from Sequential VR Sketches
- OsmT: Bridging OpenStreetMap Queries and Natural Language with Open-source Tag-aware Language Models
- E3AD: An Emotion-Aware Vision-Language-Action Model for Human-Centric End-to-End Autonomous Driving
- Efficient Safety Verification of Autonomous Vehicles with Neural Network Operator
- LLM-SrcLog: Towards Proactive and Unified Log Template Extraction via Large Language Models
- Automating Complex Document Workflows via Stepwise and Rollback-Enabled Operation Orchestration
- Self-Paced and Self-Corrective Masked Prediction for Movie Trailer Generation
- Solving LLM Repetition Problem in Production: A Comprehensive Study of Multiple Solutions
- MASE: Interpretable NLP Models via Model-Agnostic Saliency Estimation
- RapidUn: Influence-Driven Parameter Reweighting for Efficient Large Language Model Unlearning
- On the Temporality for Sketch Representation Learning
- State Space Models for Bioacoustics: A comparative Evaluation with Transformers
- Diminishing Returns in Self-Supervised Learning
- Empirical Prompt Engineering for Construct Identification with Large Language Models
- The enshittification of online search? Privacy and quality of Google, Bing and Apple in coding advice
- In-Context Representation Hijacking
- MANTRA: a Framework for Multi-stage Adaptive Noise TReAtment During Training
- Tutorial on Large Language Model-Enhanced Reinforcement Learning for Wireless Networks
- Investigating the originality of scientific papers across time and domain: A quantitative analysis
- FFTrainer: Fast Failover in Large-Language Model Training with Almost-Free State Management
- AaPE: Aliasing-aware Patch Embedding for Self-Supervised Audio Representation Learning
- Fine-grained Narrative Classification in Biased News Articles
- Think Before You Drive: World Model-Inspired Multimodal Grounding for Autonomous Vehicles
- PretrainZero: Reinforcement Active Pretraining
- SHRP: Specialized Head Routing and Pruning for Efficient Encoder Compression
- Learning to Comparison-Shop
- Action Anticipation at a Glimpse: To What Extent Can Multimodal Cues Replace Video?
- Bangla Hate Speech Classification with Fine-tuned Transformer Models
- Adapting Large Language Models to Low-Resource Tibetan: A Two-Stage Continual and Supervised Fine-Tuning Study
- Enhancing Job Matching: Occupation, Skill and Qualification Linking with the ESCO and EQF taxonomies
- TokenPowerBench: Benchmarking the Power Consumption of LLM Inference
- Flexible Gravitational-Wave Parameter Estimation with Transformers
- Reasoning-Aware Multimodal Fusion for Hateful Video Detection
- CryptoQA: A Large-scale Question-answering Dataset for AI-assisted Cryptography
- CREST: Universal Safety Guardrails Through Cluster-Guided Cross-Lingual Transfer
- Learning What to Attend First: Modality-Importance-Guided Reasoning for Reliable Multimodal Emotion Understanding
- ADORE: Autonomous Domain-Oriented Relevance Engine for E-commerce
- ESACT: An End-to-End Sparse Accelerator for Compute-Intensive Transformers via Local Similarity
- Q-BERT4Rec: Quantized Semantic-ID Representation Learning for Multimodal Recommendation
- Contrastive Deep Learning for Variant Detection in Wastewater Genomic Sequencing
- Identifying attributions of causality in political text
- Efficiency and Effectiveness of SPLADE Models on Billion-Scale Web Document Title
- Empathy Level Prediction in Multi-Modal Scenario with Supervisory Documentation Assistance
- PULSE-ICU: A Pretrained Unified Long-Sequence Encoder for Multi-task Prediction in Intensive Care Units
- WhAM: Towards A Translative Model of Sperm Whale Vocalization
- Feature Selection Empowered BERT for Detection of Hate Speech with Vocabulary Augmentation
- Low-Rank Prehab: Preparing Neural Networks for SVD Compression
- Ensemble Privacy Defense for Knowledge-Intensive LLMs against Membership Inference Attacks
- Learned-Rule-Augmented Large Language Model Evaluators
- Testing Transformer Learnability on the Arithmetic Sequence of Rooted Trees
- Evaluating SAM2 for Video Semantic Segmentation
- On the Unreasonable Effectiveness of Last-layer Retraining
- Reasoning About the Unsaid: Misinformation Detection with Omission-Aware Graph Inference
- Scaling and context steer LLMs along the same computational path as the human brain
- Reconstructing Multi-Scale Physical Fields from Extremely Sparse Measurements with an Autoencoder-Diffusion Cascade
- Label Forensics: Interpreting Hard Labels in Black-Box Text Classifier
- Code Comments for Quantum Software Development Kits: An Empirical Study on Qiskit
- DyFuLM: An Advanced Multimodal Framework for Sentiment Analysis
- MARSAD: A Multi-Functional Tool for Real-Time Social Media Analysis
- Handwritten Text Recognition for Low Resource Languages
- Data-Driven Learnability Transition of Measurement-Induced Entanglement
- AI-Enabled grading with near-domain data for scaling feedback with human-level accuracy
- MDiff4STR: Mask Diffusion Model for Scene Text Recognition
- Masked Symbol Modeling for Demodulation of Oversampled Baseband Communication Signals in Impulsive Noise-Dominated Channels
- M4-BLIP: Advancing Multi-Modal Media Manipulation Detection through Face-Enhanced Local Analysis
- Enhancing BERT Fine-Tuning for Sentiment Analysis in Lower-Resourced Languages
- A Knowledge-Based Language Model: Deducing Grammatical Knowledge in a Multi-Agent Language Acquisition Simulation
- Advancing Academic Chatbots: Evaluation of Non Traditional Outputs
- AFRAgent : An Adaptive Feature Renormalization Based High Resolution Aware GUI agent
- Melody or Machine: Detecting Synthetic Music with Dual-Stream Contrastive Learning
- Large Language Model based Smart Contract Auditing with LLMBugScanner
- Accelerating Bangla NLP Tasks with Automatic Mixed Precision: Resource-Efficient Training Preserving Model Efficacy
- FastPOS: Language-Agnostic Scalable POS Tagging Framework Low-Resource Use Case
- LAP: Fast LAtent Diffusion Planner for Autonomous Driving
- Financial Text Classification Based On rLoRA Finetuning On Qwen3-8B model
- Simplex-Optimized Hybrid Ensemble for Large Language Model Text Detection Under Generative Distribution Drif
- Exploring Health Misinformation Detection with Multi-Agent Debate
- Vision Transformer for Classification of UAV and Helicopters Using Micro-Doppler Spectrograms in Surveillance Radar
- FiCoTS: Fine-to-Coarse LLM-Enhanced Hierarchical Cross-Modality Interaction for Time Series Forecasting
- BioArc: Discovering Optimal Neural Architectures for Biological Foundation Models
- DialBench: Towards Accurate Reading Recognition of Pointer Meter using Large Foundation Models
- CodeFlowLM: Incremental Just-In-Time Defect Prediction with Pretrained Language Models and Exploratory Insights into Defect Localization
- Tree Matching Networks for Natural Language Inference: Parameter-Efficient Semantic Understanding via Dependency Parse Trees
- Video-R2: Reinforcing Consistent and Grounded Reasoning in Multimodal Language Models
- Predicting Startup-VC Fund Matches with Structural Embeddings and Temporal Investment Data
- Resolving Conflicts in Lifelong Learning via Aligning Updates in Subspaces
- Tourism Question Answer System in Indian Language using Domain-Adapted Foundation Models
- Transformer-Driven Triple Fusion Framework for Enhanced Multimodal Author Intent Classification in Low-Resource Bangla
- Data-Efficient Motor Condition Monitoring with Time Series Foundation Models
- Machine learning for violence prediction: a systematic review and critical appraisal
- LUMOS: Large User MOdels for User Behavior Prediction
- Standard Occupation Classifier -- A Natural Language Processing Approach
- Decoding the Past: Explainable Machine Learning Models for Dating Historical Texts
- A transfer learning approach for automatic conflicts detection in software requirement sentence pairs based on dual encoders
- Efficient Asynchronous Federated Evaluation with Strategy Similarity Awareness for Intent-Based Networking in Industrial Internet of Things
- Pooling Attention: Evaluating Pretrained Transformer Embeddings for Deception Classification
- RadDiff: Retrieval-Augmented Denoising Diffusion for Protein Inverse Folding
- Bridging Modalities via Progressive Re-alignment for Multimodal Test-Time Adaptation
- BanglaSentNet: An Explainable Hybrid Deep Learning Framework for Multi-Aspect Sentiment Analysis with Cross-Domain Transfer Learning
- A Trainable Centrality Framework for Modern Data
- Towards Improving Interpretability of Language Model Generation through a Structured Knowledge Discovery Approach
- Leveraging Textual Compositional Reasoning for Robust Change Captioning
- Closed-Loop Transformers: Autoregressive Modeling as Iterative Latent Equilibrium
- Beyond Real versus Fake Towards Intent-Aware Video Analysis
- PISA: Prioritized Invariant Subgraph Aggregation
- Semantic-Aware Caching for Efficient Image Generation in Edge Computing
- Economies of Open Intelligence: Tracing Power & Participation in the Model Ecosystem
- HandyLabel: Towards Post-Processing to Real-Time Annotation Using Skeleton Based Hand Gesture Recognition
- Structure is Supervision: Multiview Masked Autoencoders for Radiology
- The Collapse of Patches
- Contextual Gating within the Transformer Stack: Synergistic Feature Modulation for Enhanced Lyrical Classification and Calibration
- Shoe Style-Invariant and Ground-Aware Learning for Dense Foot Contact Estimation
- Harmonic-Percussive Disentangled Neural Audio Codec for Bandwidth Extension
- Odin: Oriented Dual-module Integration for Text-rich Network Representation Learning
- CanKD: Cross-Attention-based Non-local operation for Feature-based Knowledge Distillation
- Using Text-Based Life Trajectories from Swedish Register Data to Predict Residential Mobility with Pretrained Transformers
- Going with the Speed of Sound: Pushing Neural Surrogates into Highly-turbulent Transonic Regimes
- A Systematic Study of In-the-Wild Model Merging for Large Language Models
- FITRep: Attention-Guided Item Representation via MLLMs
- BanglaMM-Disaster: A Multimodal Transformer-Based Deep Learning Framework for Multiclass Disaster Classification in Bangla
- Infinite-Story: A Training-Free Consistent Text-to-Image Generation
- HTTM: Head-wise Temporal Token Merging for Faster VGGT
- Exploring Automated Recognition of Instructional Activity and Discourse from Multimodal Classroom Data
- PEFT-Bench: A Parameter-Efficient Fine-Tuning Methods Benchmark
- Enhancing Burmese News Classification with Kolmogorov-Arnold Network Head Fine-tuning
- Context-Aware Pragmatic Metacognitive Prompting for Sarcasm Detection
- A Probabilistic Framework for Temporal Distribution Generalization in Industry-Scale Recommender Systems
- Towards Audio Token Compression in Large Audio Language Models
- Chatty-KG: A Multi-Agent AI System for On-Demand Conversational Question Answering over Knowledge Graphs
- On the Origin of Algorithmic Progress in AI
- LAPA: Log-Domain Prediction-Driven Dynamic Sparsity Accelerator for Transformer Model
- Winning with Less for Low Resource Languages: Advantage of Cross-Lingual EnglishPersian Argument Mining Model over LLM Augmentation
- MixLM: High-Throughput and Effective LLM Ranking via Text-Embedding Mix-Interaction
- Unsupervised Memorability Modeling from Tip-of-the-Tongue Retrieval Queries
- Physics Steering: Causal Control of Cross-Domain Concepts in a Physics Foundation Model
- Dynamical Properties of Tokens in Self-Attention and Effects of Positional Encoding
- Revisiting KRISP: A Lightweight Reproduction and Analysis of Knowledge-Enhanced Vision-Language Models
- BrowseSafe: Understanding and Preventing Prompt Injection Within AI Browser Agents
- From Words to Wisdom: Discourse Annotation and Baseline Models for Student Dialogue Understanding
- CrossEarth-Gate: Fisher-Guided Adaptive Tuning Engine for Efficient Adaptation of Cross-Domain Remote Sensing Semantic Segmentation
- LLM-Driven Transient Stability Assessment: From Automated Simulation to Neural Architecture Design
- ScenarioCLIP: Pretrained Transferable Visual Language Models and Action-Genome Dataset for Natural Scene Analysis
- DUO-TOK: Dual-Track Semantic Music Tokenizer for Vocal-Accompaniment Generation
- GHR-VQA: Graph-guided Hierarchical Relational Reasoning for Video Question Answering
- KyrgyzBERT: A Compact, Efficient Language Model for Kyrgyz NLP
- Spatio-Temporal Trajectory Foundation Model - Recent Advances and Future Directions
- FINE: Factorized multimodal sentiment analysis via mutual INformation Estimation
- A Compact Hybrid Convolution--Frequency State Space Network for Learned Image Compression
- "When Data is Scarce, Prompt Smarter"... Approaches to Grammatical Error Correction in Low-Resource Settings
- DinoLizer: Separating VAE and Diffusion Artifacts in Generative Inpainting Localization
- Foundry: Distilling 3D Foundation Models for the Edge
- A Machine Learning Approach for Detection of Mental Health Conditions and Cyberbullying from Social Media
- Stragglers Can Contribute More: Uncertainty-Aware Distillation for Asynchronous Federated Learning
- Distilling Cross-Modal Knowledge via Feature Disentanglement
- Language-Independent Sentiment Labelling with Distant Supervision: A Case Study for English, Sepedi and Setswana
- Learning to Clean: Reinforcement Learning for Noisy Label Correction
- HKRAG: Holistic Knowledge Retrieval-Augmented Generation Over Visually-Rich Documents
- Schema Matching on Graph: Iterative Graph Exploration for Efficient and Explainable Data Integration
- CAMformer: Associative Memory is All You Need
- fMRI-LM: Towards a Universal Foundation Model for Language-Aligned fMRI Understanding
- Pretraining Transformer-Based Models on Diffusion-Generated Synthetic Graphs for Alzheimer's Disease Prediction
- On the Utility of Foundation Models for Fast MRI: Vision-Language-Guided Image Reconstruction
- HunyuanOCR Technical Report
- Neural surrogates for designing gravitational wave detectors
- What Drives Cross-lingual Ranking? Retrieval Approaches with Multilingual Language Models
- Unboxing the Black Box: Mechanistic Interpretability for Algorithmic Understanding of Neural Networks
- Solar-GECO: Perovskite Solar Cell Property Prediction with Geometric-Aware Co-Attention
- ABM-LoRA: Activation Boundary Matching for Fast Convergence in Low-Rank Adaptation
- Emotion-Enhanced Multi-Task Learning with LLMs for Aspect Category Sentiment Analysis
- Extracting Robust Register Automata from Neural Networks over Data Sequences
- A Multi-Agent LLM Framework for Multi-Domain Low-Resource In-Context NER via Knowledge Retrieval, Disambiguation and Reflective Analysis
- A Longitudinal Measurement of Privacy Policy Evolution for Large Language Models
- On the Role of Hidden States of Modern Hopfield Network in Transformer
- FineXtrol: Controllable Motion Generation via Fine-Grained Text
- Large Language Models for the Summarization of Czech Documents: From History to the Present
- FVAR: Visual Autoregressive Modeling via Next Focus Prediction
- VideoPerceiver: Enhancing Fine-Grained Temporal Perception in Video Multimodal Large Language Models
- Scale What Counts, Mask What Matters: Evaluating Foundation Models for Zero-Shot Cross-Domain Wi-Fi Sensing
- Robust Multimodal Sentiment Analysis of Image-Text Pairs by Distribution-Based Feature Recovery and Fusion
- Multimodal Large Language Models with Adaptive Preference Optimization for Sequential Recommendation
- Yo'City: Personalized and Boundless 3D Realistic City Scene Generation via Self-Critic Expansion
- Large-Scale In-Game Outcome Forecasting for Match, Team and Players in Football using an Axial Transformer Neural Network
- Extracting Disaster Impacts and Impact Related Locations in Social Media Posts Using Large Language Models
- OpenGloss: A Synthetic Encyclopedic Dictionary and Semantic Knowledge Graph
- Leveraging Language Models for Interpretable Analysis of Narratives in a Large Corpus
- Foundations of Artificial Intelligence Frameworks: Notion and Limits of AGI
- CrossJEPA: Cross-Modal Joint-Embedding Predictive Architecture for Efficient 3D Representation Learning from 2D Images
- Proactive Defense: Compound AI for Detecting Persuasion Attacks and Measuring Inoculation Effectiveness
- "AGI" team at SHROOM-CAP: Data-Centric Approach to Multilingual Hallucination Detection using XLM-RoBERTa
- Efficient Covariance Estimation for Sparsified Functional Data
- Muskie: Multi-view Masked Image Modeling for 3D Vision Pre-training
- Paper2SysArch: Structure-Constrained System Architecture Generation from Scientific Papers
- Multi-speaker Attention Alignment for Multimodal Social Interaction
- MobileVLA-R1: Reinforcing Vision-Language-Action for Mobile Robots
- Cost-Sensitive Conformal Training with Provably Controllable Learning Bounds
- Narratives to Numbers: Large Language Models and Economic Policy Uncertainty
- Controllability Analysis of State Space-based Language Model
- Towards Efficient LLM-aware Heterogeneous Graph Learning
- A Lightweight Approach to Detection of AI-Generated Texts Using Stylometric Features
- Blu-WERP (Web Extraction and Refinement Pipeline): A Scalable Pipeline for Preprocessing Large Language Model Datasets
- MultiGA: Leveraging Multi-Source Seeding in Genetic Algorithms
- Enhancing Quranic Learning: A Multimodal Deep Learning Approach for Arabic Phoneme Recognition
- Masked-and-Reordered Self-Supervision for Reinforcement Learning from Verifiable Rewards
- SMILE: A Composite Lexical-Semantic Metric for Question-Answering Evaluation
- DSeq-JEPA: Discriminative Sequential Joint-Embedding Predictive Architecture
- DeepCoT: Deep Continual Transformers for Real-Time Inference on Data Streams
- Large Language Models for Sentiment Analysis to Detect Social Challenges: A Use Case with South African Languages
- NX-CGRA: A Programmable Hardware Accelerator for Core Transformer Algorithms on Edge Devices
- A Hybrid Classical-Quantum Fine Tuned BERT for Text Classification
- AutoLink: Autonomous Schema Exploration and Expansion for Scalable Schema Linking in Text-to-SQL at Scale
- Attention-Guided Feature Fusion (AGFF) Model for Integrating Statistical and Semantic Features in News Text Classification
- A Cross-Cultural Assessment of Human Ability to Detect LLM-Generated Fake News about South Africa
- Learning to Compress: Unlocking the Potential of Large Language Models for Text Representation
- Bridging Visual Affective Gap: Borrowing Textual Knowledge by Learning from Noisy Image-Text Pairs
- Don't Learn, Ground: A Case for Natural Language Inference with Visual Grounding
- CREST: Improving Interpretability and Effectiveness of Troubleshooting at Ericsson through Criterion-Specific Trouble Report Retrieval
- Principled Design of Interpretable Automated Scoring for Large-Scale Educational Assessments
- Senti-iFusion: An Integrity-centered Hierarchical Fusion Framework for Multimodal Sentiment Analysis under Uncertain Modality Missingness
- CroTad: A Contrastive Reinforcement Learning Framework for Online Trajectory Anomaly Detection
- Predicting Talent Breakout Rate using Twitter and TV data
- Exploring Scientific Debt: Harnessing AI for SATD Identification in Scientific Software
- A Little More Like This: Text-to-Image Retrieval with Vision-Language Models Using Relevance Feedback
- Robot Confirmation Generation and Action Planning Using Long-context Q-Former Integrated with Multimodal LLM
- Innovation by Displacement
- Membership Inference Attacks Beyond Overfitting
- When Structure Doesn't Help: LLMs Do Not Read Text-Attributed Graphs as Effectively as We Expected
- Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation
- Learning from Sufficient Rationales: Analysing the Relationship Between Explanation Faithfulness and Token-level Regularisation Strategies
- SurvAgent: Hierarchical CoT-Enhanced Case Banking and Dichotomy-Based Multi-Agent System for Multimodal Survival Prediction
- Toward Artificial Palpation: Representation Learning of Touch on Soft Bodies
- Integrating Symbolic Natural Language Understanding and Language Models for Word Sense Disambiguation
- Progressive Supernet Training for Efficient Visual Autoregressive Modeling
- Pharos-ESG: A Framework for Multimodal Parsing, Contextual Narration, and Hierarchical Labeling of ESG Report
- Walrus: A Cross-Domain Foundation Model for Continuum Dynamics
- Green Distributed AI Training: Orchestrating Compute Across Renewable-Powered Micro Datacenters
- Benchmarking Table Extraction from Heterogeneous Scientific Extraction Documents
- Multi-Faceted Attack: Exposing Cross-Model Vulnerabilities in Defense-Equipped Vision-Language Models
- ILoRA: Federated Learning with Low-Rank Adaptation for Heterogeneous Client Aggregation
- LiSTAR: Ray-Centric World Models for 4D LiDAR Sequences in Autonomous Driving
- Synergizing Deconfounding and Temporal Generalization For Time-series Counterfactual Outcome Estimation
- Boosting Medical Visual Understanding From Multi-Granular Language Learning
- Incorporating Token Importance in Multi-Vector Retrieval
- SpectralTrain: A Universal Framework for Hyperspectral Image Classification
- AssayMatch: Learning to Select Data for Molecular Activity Models
- Unsupervised Image Classification with Adaptive Nearest Neighbor Selection and Cluster Ensembles
- CoS: Towards Optimal Event Scheduling via Chain-of-Scheduling
- TB or Not TB: Coverage-Driven Direct Preference Optimization for Verilog Stimulus Generation
- Standardising the NLP Workflow: A Framework for Reproducible Linguistic Analysis
- A Review of Machine Learning for Cavitation Intensity Recognition in Complex Industrial Systems
- Judicial Sentencing Prediction Based on Hybrid Models and Two-Stage Learning Algorithms
- C2F-Space: Coarse-to-Fine Space Grounding for Spatial Instructions using Vision-Language Models
- An Optimized Machine Learning Classifier for Detecting Fake Reviews Using Extracted Features
- Eq.Bot: Enhance Robotic Manipulation Learning via Group Equivariant Canonicalization
- OEMA: Ontology-Enhanced Multi-Agent Collaboration Framework for Zero-Shot Clinical Named Entity Recognition
- Effective Code Membership Inference for Code Completion Models via Adversarial Prompts
- TiCAL:Typicality-Based Consistency-Aware Learning for Multimodal Emotion Recognition
- HEAD-QA v2: Expanding a Healthcare Benchmark for Reasoning
- ProPL: Universal Semi-Supervised Ultrasound Image Segmentation via Prompt-Guided Pseudo-Labeling
- Opinion Dynamics Models for Sentiment Evolution in Weibo Blogs
- Hierarchical Token Prepending: Enhancing Information Flow in Decoder-based LLM Embeddings
- Bias in, Bias out: Annotation Bias in Multilingual Large Language Models
- AutoTool: Efficient Tool Selection for Large Language Model Agents
- Gradient-descent methods for quantum detector tomography
- Steganographic Backdoor Attacks in NLP: Ultra-Low Poisoning and Defense Evasion
- LogPurge: Log Data Purification for Anomaly Detection via Rule-Enhanced Filtering
- Neural Posterior Estimation with Autoregressive Tiling for Detecting Objects in Astronomical Images
- 10Cache: Heterogeneous Resource-Aware Tensor Caching and Migration for LLM Training
- TTA: Transcribe, Translate and Alignment for Cross-lingual Speech Representation
- RAG-Driven Data Quality Governance for Enterprise ERP Systems
- Generalizable and Efficient Automated Scoring with a Knowledge-Distilled Multi-Task Mixture-of-Experts
- Learning Skill-Attributes for Transferable Assessment in Video
- ActVAR: Activating Mixtures of Weights and Tokens for Efficient Visual Autoregressive Generation
- TaoSearchEmb: A Multi-Objective Reinforcement Learning Framework for Dense Retrieval in Taobao Search
- MergeDNA: Context-aware Genome Modeling with Dynamic Tokenization through Token Merging
- Hierarchical Retrieval with Out-Of-Vocabulary Queries: A Case Study on SNOMED CT
- Reconstruction-Driven Multimodal Representation Learning for Automated Media Understanding
- OlmoEarth: Stable Latent Image Modeling for Multimodal Earth Observation
- RegionMarker: A Region-Triggered Semantic Watermarking Framework for Embedding-as-a-Service Copyright Protection
- Whistledown: Combining User-Level Privacy with Conversational Coherence in LLMs
- 3DAlign-DAER: Dynamic Attention Policy and Efficient Retrieval Strategy for Fine-grained 3D-Text Alignment at Scale
- Large Language Models Meet Extreme Multi-label Classification: Scaling and Multi-modal Framework
- Zero-Shot Grammar Competency Estimation Using Large Language Model Generated Pseudo Labels
- Skeletons Speak Louder than Text: A Motion-Aware Pretraining Paradigm for Video-Based Person Re-Identification
- MedGEN-Bench: Contextually entangled benchmark for open-ended multimodal medical generation
- F.A.C.U.L.: Language-Based Interaction with AI Companions in Gaming
- BeDiscovER: The Benchmark of Discourse Understanding in the Era of Reasoning Language Models
- How Good is BLI as an Alignment Measure: A Study in Word Embedding Paradigm
- TacEleven: generative tactic discovery for football open play
- BIRD: Bronze Inscription Restoration and Dating
- ExplicitLM: Decoupling Knowledge from Parameters via Explicit Memory Banks
- Multimodal Large Language Models as Image Classifiers
- MolEdit: Knowledge Editing for Multimodal Molecule Language Models
- Backdoor Attacks on Open Vocabulary Object Detectors via Multi-Modal Prompt Tuning
- HMVLM: Human Motion-Vision-Lanuage Model via MoE LoRA
- CAT-ID2: Category-Tree Integrated Document Identifier Learning for Generative Retrieval In E-commerce
- Robust Multimodal Sentiment Analysis via Double Information Bottleneck
- UniSOT: A Unified Framework for Multi-Modality Single Object Tracking
- SEPS: Semantic-enhanced Patch Slimming Framework for fine-grained cross-modal alignment
- Evaluating Embedding Generalization: How LLMs, LoRA, and SLERP Shape Representational Geometry
- HEDGE: Hallucination Estimation via Dense Geometric Entropy for VQA with Vision-Language Models
- How Language Directions Align with Token Geometry in Multilingual LLMs
- Medical Knowledge Intervention Prompt Tuning for Medical Image Classification
- Enhancing Conversational Recommender Systems with Tree-Structured Knowledge and Pretrained Language Models
- A Content-Preserving Secure Linguistic Steganography
- MaskAnyNet: Rethinking Masked Image Regions as Valuable Information in Supervised Learning
- Probing Preference Representations: A Multi-Dimensional Evaluation and Analysis Method for Reward Models
- Personality-guided Public-Private Domain Disentangled Hypergraph-Former Network for Multimodal Depression Detection
- Concept-Based Interpretability for Toxicity Detection
- Global-Lens Transformers: Adaptive Token Mixing for Dynamic Link Prediction
- Prompt Engineering Techniques for Context-dependent Text-to-SQL in Arabic
- LOBERT: Generative AI Foundation Model for Limit Order Book Messages
- Emotion and Intention Guided Multi-Modal Learning for Sticker Response Selection
- Evaluating Latent Generative Paradigms for High-Fidelity 3D Shape Completion from a Single Depth Image
- ViConBERT: Context-Gloss Aligned Vietnamese Word Embedding for Polysemous and Sense-Aware Representations
- Mixture of States: Routing Token-Level Dynamics for Multimodal Generation
- MixAR: Mixture Autoregressive Image Generation
- Understanding InfoNCE: Transition Probability Matrix Induced Feature Clustering
- PRISM of Opinions: A Persona-Reasoned Multimodal Framework for User-centric Conversational Stance Detection
- Teaching Prompts to Coordinate: Hierarchical Layer-Grouped Prompt Tuning for Continual Learning
- A Robust and Explainable Transformer-Based Framework for Phishing Email Detection
- LM4Opt-RA: A Multi-Candidate LLM Framework with Structured Ranking for Automating Network Resource Allocation
- MME-RAG: Multi-Manager-Expert Retrieval-Augmented Generation for Fine-Grained Entity Recognition in Task-Oriented Dialogues
- Augmenting The Weather: A Hybrid Counterfactual-SMOTE Algorithm for Improving Crop Growth Prediction When Climate Changes
- Constructing Political Coordinates: Aggregating Over the Opposition for Diverse News Recommendation
- Identifying Imaging Follow-Up in Radiology Reports: A Comparative Analysis of Traditional ML and LLM Approaches
- Scaling Open-Weight Large Language Models for Hydropower Regulatory Information Extraction: A Systematic Analysis
- CURENet: Combining Unified Representations for Efficient Chronic Disease Prediction
- M-DAIGT: A Shared Task on Multi-Domain Detection of AI-Generated Text
- Stroke Modeling Enables Vectorized Character Generation with Large Vectorized Glyph Model
- ChemFixer: Correcting Invalid Molecules to Unlock Previously Unseen Chemical Space
- LEMUR: Large scale End-to-end MUltimodal Recommendation
- Language-Guided Graph Representation Learning for Video Summarization
- Dynamic Temperature Scheduler for Knowledge Distillation
- Binary BPE: A Family of Cross-Platform Tokenizers for Binary Analysis
- From 2D to 3D Without Extra Baggage: Data-Efficient Cancer Detection in Digital Breast Tomosynthesis
- STAGE: A Symbolic Tensor grAph GEnerator for distributed AI system co-design
- Analogical Structure, Minimal Contextual Cues and Contrastive Distractors: Input Design for Sample-Efficient Linguistic Rule Induction
- Torch-Uncertainty: A Deep Learning Framework for Uncertainty Quantification
- RoboBenchMart: Benchmarking Robots in Retail Environment
- FineSkiing: A Fine-grained Benchmark for Skiing Action Quality Assessment
- Persona-Aware Alignment Framework for Personalized Dialogue Generation
- Ellipsoid-Based Decision Boundaries for Open Intent Classification
- VISTA: A Vision and Intent-Aware Social Attention Framework for Multi-Agent Trajectory Prediction
- Text2SQL-Flow: A Robust SQL-Aware Data Augmentation Framework for Text-to-SQL
- Generalizing to Unseen Disaster Events: A Causal View
- Adaptive Hyperbolic Kernels: Modulated Embedding in de Branges-Rovnyak Spaces
- Boosting In-Silicon Directed Evolution with Fine-Tuned Protein Language Model and Tree Search
- TermGPT: Multi-Level Contrastive Fine-Tuning for Terminology Adaptation in Legal and Financial Domain
- DELICATE: Diachronic Entity LInking using Classes And Temporal Evidence
- ACT as Human: Multimodal Large Language Model Data Annotation with Critical Thinking
- SynClaimEval: A Framework for Evaluating the Utility of Synthetic Data in Long-Context Claim Verification
- MCAD: Multimodal Context-Aware Audio Description Generation For Soccer
- C3TG: Conflict-aware, Composite, and Collaborative Controlled Text Generation
- End-to-end Contrastive Language-Speech Pretraining Model For Long-form Spoken Question Answering
- Generative AI as a Linguistic Equalizer in Global Science
- Blurred Encoding for Trajectory Representation Learning
- EVADE: LLM-Based Explanation Generation and Validation for Error Detection in NLI
- TransactionGPT
- Improve Contrastive Clustering Performance by Multiple Fusing-Augmenting ViT Blocks
- Hierarchical Memorization in Large Language Models: Evidence from Citation Generation
- A centroid based framework for text classification in itsm environments
- PriVi: Towards A General-Purpose Video Model For Primate Behavior In The Wild
- Synergistic Feature Fusion for Latent Lyrical Classification: A Gated Deep Learning Architecture
- Machine Learning-Guided Memory Optimization for DLRM Inference on Tiered Memory
- CleverBirds: A Multiple-Choice Benchmark for Fine-grained Human Knowledge Tracing
- Advancing Scientific Knowledge Retrieval and Reuse with a Novel Digital Library for Machine-Readable Knowledge
- Automated sign detection across the Electronic Babylonian Library: A large-scale dataset and end-to-end cuneiform OCR pipeline
- Leveraging unlabelled data for generalizable neural population decoding
- One Model for All: Universal Pre-training for EEG based Emotion Recognition across Heterogeneous Datasets and Paradigms
- EMAformer: Enhancing Transformer through Embedding Armor for Time Series Forecasting
- TurkEmbed: Turkish Embedding Model on NLI & STS Tasks
- Why does weak-OOD help? A Further Step Towards Understanding Jailbreaking VLMs
- JobSphere: An AI-Powered Multilingual Career Copilot for Government Employment Platforms
- H-Model: Dynamic Neural Architectures for Adaptive Processing
- Concentration bounds on response-based vector embeddings of black-box generative models
- X-IONet: Cross-Platform Inertial Odometry Network with Dual-Stage Attention
- PE-TSFM: Self-Supervised Time-Series Learning for Generalizable Power Converter Health Monitoring under Unseen Conditions
- Evaluating Gemini LLM in Food Image-Based Recipe and Nutrition Description with EfficientNet-B4 Visual Backbone
- Encoder Fine-tuning with Stochastic Sampling Outperforms Open-weight GPT in Astronomy Knowledge Extraction
- Multi-Granularity Mutual Refinement Network for Zero-Shot Learning
- Boomda: Balanced Multi-objective Optimization for Multimodal Domain Adaptation
- DiffuGR: Generative Document Retrieval with Diffusion Language Models
- On the Interplay between Positional Encodings, Morphological Complexity, and Word Order Flexibility
- A Small Leak Sinks All: Exploring the Transferable Vulnerability of Source Code Models
- Estranged Predictions: Measuring Semantic Category Disruption with Masked Language Modelling
- HipKittens: Fast and Furious AMD Kernels
- MSCR: Exploring the Vulnerability of LLMs' Mathematical Reasoning Abilities Using Multi-Source Candidate Replacement
- Generalizable Insights for Graph Transformers in Theory and Practice
- State of the Art in Text Classification for South Slavic Languages: Fine-Tuning or Prompting?
- Unified Work Embeddings: Contrastive Learning of a Bidirectional Multi-task Ranker
- Breaking the Adversarial Robustness-Performance Trade-off in Text Classification via Manifold Purification
- Parallel Sampling via Autospeculation
- VEDA: 3D Molecular Generation via Variance-Exploding Diffusion with Annealing
- VectorSynth: Fine-Grained Satellite Image Synthesis with Structured Semantics
- Cross Modal Fine-Grained Alignment via Granularity-Aware and Region-Uncertain Modeling
- LLM-GROP: Visually Grounded Robot Task and Motion Planning with Large Language Models
- Harmonic Token Projection (HTP): A Vocabulary-Free, Training-Free, Deterministic, and Reversible Embedding Methodology
- TurkEmbed4Retrieval: Turkish Embedding Model for Retrieval Task
- Beyond Fact Retrieval: Episodic Memory for RAG with Generative Semantic Workspaces
- Contact Map Transfer with Conditional Diffusion Model for Generalizable Dexterous Grasp Generation
- SPOT: An Annotated French Corpus and Benchmark for Detecting Critical Interventions in Online Conversations
- Evaluating Online Moderation Via LLM-Powered Counterfactual Simulations
- LeCoT: revisiting network architecture for two-view correspondence pruning
- EmoBang: Detecting Emotion From Bengali Texts
- Do LLMs Feel? Teaching Emotion Recognition with Prompts, Retrieval, and Curriculum Learning
- Evaluating Large Language Models for Anxiety, Depression, and Stress Detection: Insights into Prompting Strategies and Synthetic Data
- GoCkpt: Gradient-Assisted Multi-Step overlapped Checkpointing for Efficient LLM Training
- Learning to Focus: Prioritizing Informative Histories with Structured Attention Mechanisms in Partially Observable Reinforcement Learning
- Sampling and Loss Weights in Multi-Domain Training
- QUARK: Quantization-Enabled Circuit Sharing for Transformer Acceleration by Exploiting Common Patterns in Nonlinear Operations
- Argus: Quality-Aware High-Throughput Text-to-Image Inference Serving System
- Multi-Modal Continual Learning via Cross-Modality Adapters and Representation Alignment with Knowledge Preservation
- Evaluating Language Model Applications for Identifying Solution-Related Content in Issue Report Discussions
- Rethinking what Matters: Effective and Robust Multilingual Realignment for Low-Resource Languages
- CAE: Character-Level Autoencoder for Non-Semantic Relational Data Grouping
- S-DAG: A Subject-Based Directed Acyclic Graph for Multi-Agent Heterogeneous Reasoning
- Rep2Text: Decoding Full Text from a Single LLM Token Representation
- SR-KI: Scalable and Real-Time Knowledge Integration into LLMs via Supervised Attention
- Comparing Reconstruction Attacks on Pretrained Versus Full Fine-tuned Large Language Model Embeddings on Homo Sapiens Splice Sites Genomic Data
- HatePrototypes: Interpretable and Transferable Representations for Implicit and Explicit Hate Speech Detection
- Vocabulary In-Context Learning in Transformers: Benefits of Positional Encoding
- Reaction Prediction via Interaction Modeling of Symmetric Difference Shingle Sets
- EcoSpa: Efficient Transformer Training with Coupled Sparsity
- RAG-targeted Adversarial Attack on LLM-based Threat Detection and Mitigation Framework
- Interaction-Centric Knowledge Infusion and Transfer for Open-Vocabulary Scene Graph Generation
- Enhancing Multimodal Misinformation Detection by Replaying the Whole Story from Image Modality Perspective
- FLEX: Continuous Agent Evolution via Forward Learning from Experience
- LLaDA-Rec: Discrete Diffusion for Parallel Semantic ID Generation in Generative Recommendation
- ReMoD: Rethinking Modality Contribution in Multimodal Stance Detection via Dual Reasoning
- Towards Implicit Aggregation: Robust Image Representation for Place Recognition in the Transformer Era
- RF-Behavior: A Multimodal Radio-Frequency Dataset for Human Behavior and Emotion Analysis
- Kunlun Anomaly Troubleshooter: Enabling Kernel-Level Anomaly Detection and Causal Reasoning for Large Model Distributed Inference
- CoMA: Complementary Masking and Hierarchical Dynamic Multi-Window Self-Attention in a Unified Pre-training Framework
- L2T-Hyena: Enhancing State-Space Models with an Adaptive Learn-to-Teach Framework
- IDALC: A Semi-Supervised Framework for Intent Detection and Active Learning based Correction
- Open-World 3D Scene Graph Generation for Retrieval-Augmented Reasoning
- DiagnoLLM: A Hybrid Bayesian Neural Language Framework for Interpretable Disease Diagnosis
- Multi-Scale Feature Fusion and Graph Neural Network Integration for Text Classification with Large Language Models
- TabDistill: Distilling Transformers into Neural Nets for Few-Shot Tabular Classification
- A Metamorphic Testing Perspective on Knowledge Distillation for Language Models of Code: Does the Student Deeply Mimic the Teacher?
- Semantic-Guided Natural Language and Visual Fusion for Cross-Modal Interaction Based on Tiny Object Detection
- How Many Tokens Do 3D Point Cloud Transformer Architectures Really Need?
- Large Language Models for Explainable Threat Intelligence
- Evaluating Subword Tokenization Techniques for Bengali: A Benchmark Study with BengaliBPE
- Effectiveness of Chain-of-Thought in Distilling Reasoning Capability from Large Language Models
- ManufactuBERT: Efficient Continual Pretraining for Manufacturing
- Iterative Layer-wise Distillation for Efficient Compression of Large Language Models
- Deep Progressive Training: scaling up depth capacity of zero/one-layer models
- LoPT: Lossless Parallel Tokenization Acceleration for Long Context Inference of Large Language Model
- Listening Between the Lines: Decoding Podcast Narratives with Language Modeling
- Multimodal Diffusion Forcing for Forceful Manipulation
- Q3R: Quadratic Reweighted Rank Regularizer for Effective Low-Rank Training
- On the Brittleness of CLIP Text Encoders
- The Strong Lottery Ticket Hypothesis for Multi-Head Attention Mechanisms
- ScaleDL: Towards Scalable and Efficient Runtime Prediction for Distributed Deep Learning Workloads
- KGFR: A Foundation Retriever for Generalized Knowledge Graph Question Answering
- E-CARE: An Efficient LLM-based Commonsense-Augmented Framework for E-Commerce
- Plan of Knowledge: Retrieval-Augmented Large Language Models for Temporal Knowledge Graph Question Answering
- DartQuant: Efficient Rotational Distribution Calibration for LLM Quantization
- Memory- and Latency-Constrained Inference of Large Language Models via Adaptive Split Computing
- MoSa: Motion Generation with Scalable Autoregressive Modeling
- Improving the Performance of Radiology Report De-identification with Large-Scale Training and Benchmarking Against Cloud Vendor Methods
- Promoting Sustainable Web Agents: Benchmarking and Estimating Energy Consumption through Empirical and Theoretical Analysis
- Evaluating Machine Translation Datasets for Low-Web Data Languages: A Gendered Lens
- OMPILOT: Harnessing Transformer Models for Auto Parallelization to Shared Memory Computing Paradigms
- A Super-Learner with Large Language Models for Medical Emergency Advising
- Towards Transparent Stance Detection: A Zero-Shot Approach Using Implicit and Explicit Interpretability
- A systematic review of relation extraction task since the emergence of Transformers
- Towards Scalable Web Accessibility Audit with MLLMs as Copilots
- Light over Heavy: Automated Performance Requirements Quantification with Linguistic Inducement
- Knowledge-Augmented Question Error Correction for Chinese Question Answer System with QuestionRAG
- Segmentation Beyond Defaults: Asymmetrical Byte Pair Encoding for Optimal Machine Translation Performance
- GEMMA-SQL: A Novel Text-to-SQL Model Based on Large Language Models
- Depth-induced NTK: Bridging Over-parameterized Neural Networks and Deep Neural Kernels
- How to Evaluate Speech Translation with Source-Aware Neural MT Metrics
- Laugh, Relate, Engage: Stylized Comment Generation for Short Videos
- An Efficient Classification Model for Cyber Text
- A Computational Approach to Analyzing Disrupted Language in Schizophrenia: Integrating Surprisal and Coherence Measures
- An Interdisciplinary and Cross-Task Review on Missing Data Imputation
- Divide, Cache, Conquer: Dichotomic Prompting for Efficient Multi-Label LLM-Based Classification
- KnowThyself: An Agentic Assistant for LLM Interpretability
- How Different Tokenization Algorithms Impact LLMs and Transformer Models for Binary Code Analysis
- A Large Scale Study of AI-based Binary Function Similarity Detection Techniques for Security Researchers and Practitioners
- The Curved Spacetime of Transformer Architectures
- Fine-Tuning Vision-Language Models for Multimodal Polymer Property Prediction
- Data-Efficient Adaptation and a Novel Evaluation Method for Aspect-based Sentiment Analysis
- SCALE-VLP: Soft-Weighted Contrastive Volumetric Vision-Language Pre-training with Spatial-Knowledge Semantics
- Automatic Machine Translation Detection Using a Surrogate Multilingual Translation Model
- ProM3E: Probabilistic Masked MultiModal Embedding Model for Ecology
- Cache Mechanism for Agent RAG Systems
- Unified Diffusion VLA: Vision-Language-Action Model via Joint Discrete Denoising Diffusion Process
- From emotional data to decisions: A systematic review on how airlines use sentiments and emotions to stay ahead
- Modeling Hawkish-Dovish Latent Beliefs in Multi-Agent Debate-Based LLMs for Monetary Policy Decision Classification
- Next Token Knowledge Tracing: Exploiting Pretrained LLM Representations to Decode Student Behaviour
- Smart-Hiring: An Explainable end-to-end Pipeline for CV Information Extraction and Job Matching
- From the Laboratory to Real-World Application: Evaluating Zero-Shot Scene Interpretation on Edge Devices for Mobile Robotics
- Chronic Kidney Disease Prognosis Prediction Using Transformer
- Personalized Decision Modeling: Utility Optimization or Textualized-Symbolic Reasoning
- Metamorphic Testing of Large Language Models for Natural Language Processing
- Vibe Learning: Education in the age of AI
- DEER: Disentangled Mixture of Experts with Instance-Adaptive Routing for Generalizable Machine-Generated Text Detection
- Detecting Vulnerabilities from Issue Reports for Internet-of-Things
- Rescuing the Unpoisoned: Efficient Defense against Knowledge Corruption Attacks on RAG Systems
- Improving Romanian LLM Pretraining Data using Diversity and Quality Filtering
- HPLT 3.0: Very Large-Scale Multilingual Resources for LLMs and MT. Mono- and Bi-lingual Data, Multilingual Evaluation, and Pre-Trained Models
- Energy-Efficient Deep Learning Without Backpropagation: A Rigorous Evaluation of Forward-Only Algorithms
- OceanAI: A Conversational Platform for Accurate, Transparent, Near-Real-Time Oceanographic Insights
- VesSAM: Efficient Multi-Prompting for Segmenting Complex Vessel
- The Riddle of Reflection: Evaluating Reasoning and Self-Awareness in Multilingual LLMs using Indian Riddles
- G2rammar: Bilingual Grammar Modeling for Enhanced Text-attributed Graph Learning
- Transformers as Intrinsic Optimizers: Forward Inference through the Energy Principle
- ColMate: Contrastive Late Interaction and Masked Text for Multimodal Document Retrieval
- Training with Fewer Bits: Unlocking Edge LLMs Training with Stochastic Rounding
- All-in-one Graph-based Indexing for Hybrid Search on GPUs
- TriCon-Fair: Triplet Contrastive Learning for Mitigating Social Bias in Pre-trained Language Models
- Parameter Interpolation Adversarial Training for Robust Image Classification
- Reconstruction of Black Hole Ringdown Signals with Data Gaps using a Deep-Learning Framework
- Deciphering Scientific Collaboration in Biomedical LLM Research: Dynamics, Institutional Participation, and Resource Disparities
- Automated Invoice Data Extraction: Using LLM and OCR
- Do You Know About My Nation? Investigating Multilingual Language Models' Cultural Literacy Through Factual Knowledge
- CueBench: Advancing Unified Understanding of Context-Aware Video Anomalies in Real-World
- EPARA: Parallelizing Categorized AI Inference in Edge Clouds
- Listwise Preference Diffusion Optimization for User Behavior Trajectories Prediction
- Exploring and Mitigating Gender Bias in Encoder-Based Transformer Models
- ToxicTextCLIP: Text-Based Poisoning and Backdoor Attacks on CLIP Pre-training
- Penetrating the Hostile: Detecting DeFi Protocol Exploits through Cross-Contract Analysis
- Reversal Invariance in Autoregressive Language Models
- A Multimodal Framework for Depression Detection during Covid-19 via Harvesting Social Media: A Novel Dataset and Method
- Foundation Models for Trajectory Planning in Autonomous Driving: A Review of Progress and Open Challenges
- Sketch-to-Layout: Sketch-Guided Multimodal Layout Generation
- Multilingual BERT language model for medical tasks: Evaluation on domain-specific adaptation and cross-linguality
- BiSparse-AAS: Bilinear Sparse Attention and Adaptive Spans Framework for Scalable and Efficient Text Summarization
- From the Rock Floor to the Cloud: A Systematic Survey of State-of-the-Art NLP in Battery Life Cycle
- FPS: Feedforward-based Parameter Selection For Efficient Fine-Tuning
- A Survey on Deep Text Hashing: Efficient Semantic Text Retrieval with Binary Representation
- Exploring Landscapes for Better Minima along Valleys
- Probability Distributions Computed by Hard-Attention Transformers
- QiNN-QJ: A Quantum-inspired Neural Network with Quantum Jump for Multimodal Sentiment Analysis
- A Retrospect to Multi-prompt Learning across Vision and Language
- Cognitive Alignment in Personality Reasoning: Leveraging Prototype Theory for MBTI Inference
- Addressing Longstanding Challenges in Cognitive Science with Language Models
- Exploring the Utilities of the Rationales from Large Language Models to Enhance Automated Essay Scoring
- Spatial-SSRL: Enhancing Spatial Understanding via Self-Supervised Reinforcement Learning
- POSESTITCH-SLT: Linguistically Inspired Pose-Stitching for End-to-End Sign Language Translation
- Calibration Across Layers: Understanding Calibration Evolution in LLMs
- NaturalVoices: A Large-Scale, Spontaneous and Emotional Podcast Dataset for Voice Conversion
- The aftermath of compounds: Investigating Compounds and their Semantic Representations
- Dataset Creation and Baseline Models for Sexism Detection in Hausa
- Elastic Architecture Search for Efficient Language Models
- Enhancing Sentiment Classification with Machine Learning and Combinatorial Fusion
- Mixture-of-Transformers Learn Faster: A Theoretical Study on Classification Problems
- Semantic Frame Aggregation-based Transformer for Live Video Comment Generation
- Overview of the MEDIQA-OE 2025 Shared Task on Medical Order Extraction from Doctor-Patient Consultations
- Masked Diffusion Captioning for Visual Feature Learning
- Encoder-Decoder or Decoder-Only? Revisiting Encoder-Decoder Large Language Model
- The Structure of Relation Decoding Linear Operators in Large Language Models
- Hebrew Diacritics Restoration using Visual Representation
- CyberNER: A Harmonized STIX Corpus for Cybersecurity Named Entity Recognition
- SynBullying: A Multi LLM Synthetic Conversational Dataset for Cyberbullying Detection
- Don't Let It Fade: Preserving Edits in Diffusion Language Models via Token Timestep Allocation
- ReaKase-8B: Legal Case Retrieval via Knowledge and Reasoning Representations with LLMs
- Linking Heterogeneous Data with Coordinated Agent Flows for Social Media Analysis
- Bridging the Gap Between Molecule and Textual Descriptions via Substructure-aware Alignment
- StructLayoutFormer:Conditional Structured Layout Generation via Structure Serialization and Disentanglement
- Do Students Debias Like Teachers? On the Distillability of Bias Mitigation Methods
- Detecting Anomalies in Machine Learning Infrastructure via Hardware Telemetry
- SPEAR: A Unified SSL Framework for Learning Speech and Audio Representations
- Revisiting Multilingual Data Mixtures in Language Model Pretraining
- Beyond Long Context: When Semantics Matter More than Tokens
- Generalized Pseudo-Relevance Feedback
- Beyond Data Scarcity Optimizing R3GAN for Medical Image Generation from Small Datasets
- Semantic-Aware Temporal Adaptation for UAV Anti-UAV Tracking
- From Interface to Inference: Eliciting Any-Order Inference from Any-Order Models
- ClockRoPE: Random Fourier Rotations for Temporal Routine Modeling
- Diagnosing Fine-Grained Inconsistency Classification in Financial Disclosure Text
- ASARL: Autonomous Social-Aware Relevance Learning for QQ Search
- Decoupled Visual Processing: Efficient Multimodal Adaptation via Modality-Specific Transformer Substitution
- Where Detectors Fail: Closing the Tail-Domain Gap with Expert-Guided Mutual Distillation
- Speech2Grasp: Data-Efficient Transfer of Text-Conditioned Grasp Detection to Speech in Humanoid Robots
- RAGuard: A Layered Defense Framework for Retrieval-Augmented Generation Systems Against Data Poisoning
- Shape-Based Inductive Bias for Glioma Grading from Tumor Contours
- Guess Where You Go: Generative Next Point-of-Interest Recommendation in Amap
- How to make the most of your masked language model for protein engineering
- Semantic Similarity in Radiology Reports via LLMs and NER
- ExplainRec: Towards Explainable Multi-Modal Zero-Shot Recommendation with Preference Attribution and Large Language Models
- Enhancing Generative Information Extraction with Two-step Validation: A Product Attribute Use Case
- When Knowledge Changes: Metamorphic Testing of RAG Systems with Mutations
- Truth-Aware Decoding: A Program-Logic Approach to Factual Language Generation
- SCALPEL: Semantic Cross-modal Alignment via LLM-Powered Encoder Learning for Medical Vision-Language Representation
- Ideas Have Genomes: Benchmarking Scientific Lineage Reasoning and Lineage-Grounded Idea Generation
- Grammar as a Behavioral Biometric: Using Cognitively Motivated Grammar Models for Authorship Verification
- TweetyBERT: Automated parsing of birdsong through self-supervised machine learning
- <i>Making a Name for Myself</i> : On Academic Naming Policies and their Impact
- MedFusionT5: Cross-Modal Attention Boosts Semantic Quality and Reduces Hallucinations in Dental AI
- The Scaling Properties of Implicit Deductive Reasoning in Transformers
- Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini
- Understanding Wacky Weights: A Dissection of SPLADE's Learned Term Importance
- Seeking the Unfamiliar but Memorable: Conceptual Creativity as Meta-Learning
- Quantifying Concentration Phenomena of Mean-Field Transformers in the Low-Temperature Regime
- Lattice Deduction Transformers
- Exploring the Intersection of AI, Language, and Law: A Bibliometric Analysis
- Hallucination Benchmark for Speech Foundation Models
- NeurIPT: Foundation Model for Neural Interfaces
- MIN-Merging: Merge the Important Neurons for Model Merging
- On the Use of Large Language Models for Qualitative Synthesis
- Machine learning based screening of potential paper mill publications in cancer research: methodological and cross sectional study
- TrajSelector: Harnessing Latent Representations for Efficient and Effective Best-of-N in Large Reasoning Model
- Bots into the Fediverse
- Digitising Death: Benchmarking Genealogical Data and Recovering Women’s Histories in Early Modern Ireland
- DeepTaxa: a hybrid CNN-BERT framework for 16S rRNA taxonomic classification
- FrugalPrompt: Reducing Contextual Overhead in Large Language Models via Token Attribution
- Enhancing Linguistic Competence of Language Models through Pre-training with Language Learning Tasks
- RL makes MLLMs see better than SFT
- Stay Tuned: Improving Sentiment Analysis and Stance Detection Using Large Language Models
- Probing the Hidden Talent of ASR Foundation Models for L2 English Oral Assessment
- Self-Supervised Text-Vision Alignment for Automated Brain MRI Abnormality Detection: A Multicenter Study (ALIGN Study)
- Artificial intelligence in bioinformatics: a survey
- End-to-End Argument Mining through Autoregressive Argumentative Structure Prediction
- BSFA: Leveraging the Subspace Dichotomy to Accelerate Neural Network Training
- Beyond One-Size-Fits-All: Personalized Harmful Content Detection with In-Context Learning
- SynHLMA:Synthesizing Hand Language Manipulation for Articulated Object with Discrete Human Object Interaction Representation
- IBNorm: Information-Bottleneck Inspired Normalization for Representation Learning
- Testing Cross-Lingual Text Comprehension In LLMs Using Next Sentence Prediction
- MemEIC: A Step Toward Continual and Compositional Knowledge Editing
- Bridging the Divide: End-to-End Sequence-Graph Learning
- NanoVLA: Routing Decoupled Vision-Language Understanding for Nano-sized Generalist Robotic Policies
- A Survey on Unlearning in Large Language Models
- Can LLMs Estimate Cognitive Complexity of Reading Comprehension Items?
- Retrieval-Augmented Multimodal Depression Detection
- Exponential Dynamic Energy Network for High Capacity Sequence Memory
- RiddleBench: A New Generative Reasoning Benchmark for LLMs
- Aggregation Hides Out-of-Distribution Generalization Failures from Spurious Correlations
- SARC: Sentiment-Augmented Deep Role Clustering for Fake News Detection
- Eigenfunction Extraction for Ordered Representation Learning
- Optimizing Retrieval for RAG via Reinforced Contrastive Learning
- MuseTok: Symbolic Music Tokenization for Generation and Semantic Understanding
- "Mm, Wat?" Detecting Other-initiated Repair Requests in Dialogue
- BEST-RQ-Based Self-Supervised Learning for Whisper Domain Adaptation
- Generative AI for Healthcare: Fundamentals, Challenges, and Perspectives
- Talk2Ref: A Dataset for Reference Prediction from Scientific Talks
- Mitigating Hallucination in Large Language Models (LLMs): An Application-Oriented Survey on RAG, Reasoning, and Agentic Systems
- DiffusionX: Efficient Edge-Cloud Collaborative Image Generation with Multi-Round Prompt Evolution
- HiMAE: Hierarchical Masked Autoencoders Discover Resolution-Specific Structure in Wearable Time Series
- Human-Level Reasoning: A Comparative Study of Large Language Models on Logical and Abstract Reasoning
- Uncovering Gaps Between RFC Updates and TCP/IP Implementations: LLM-Facilitated Differential Checks on Intermediate Representations
- Metadata-Driven Retrieval-Augmented Generation for Financial Question Answering
- A Unified Geometric Space Bridging AI Models and the Human Brain
- LASTIST: LArge-Scale Target-Independent STance dataset
- MERGE: Minimal Expression-Replacement GEneralization Test for Natural Language Inference
- Abjad AI at NADI 2025: CATT-Whisper: Multimodal Diacritic Restoration Using Text and Speech Representations
- SCOPE: Saliency-Coverage Oriented Token Pruning for Efficient Multimodel LLMs
- BMGQ: A Bottom-up Method for Generating Complex Multi-hop Reasoning Questions from Semi-structured Data
- Leveraging LLMs for Early Alzheimer's Prediction
- Squrve: A Unified and Modular Framework for Complex Real-World Text-to-SQL Tasks
- COMMUNITYNOTES: A Dataset for Exploring the Helpfulness of Fact-Checking Explanations
- Mitigating Negative Transfer via Reducing Environmental Disagreement
- Blending Learning to Rank and Dense Representations for Efficient and Effective Cascades
- DynBERG: Dynamic BERT-based Graph neural network for financial fraud detection
- Endowing GPT-4 with a Humanoid Body: Building the Bridge Between Off-the-Shelf VLMs and the Physical World
- Inferring Group Intent as a Cooperative Game. An NLP-based Framework for Trajectory Analysis using Graph Transformer Neural Network
- Language Models for Longitudinal Clinical Prediction
- ReviewGuard: Enhancing Deficient Peer Review Detection via LLM-Driven Data Augmentation
- SSL4RL: Revisiting Self-supervised Learning as Intrinsic Reward for Visual-Language Reasoning
- RefAtomNet++: Advancing Referring Atomic Video Action Recognition using Semantic Retrieval based Multi-Trajectory Mamba
- Modeling Expert Interactions in Sparse Mixture of Experts via Graph Structures
- How Pragmatics Shape Articulation: A Computational Case Study in STEM ASL Discourse
- VIPAMIN: Visual Prompt Initialization via Embedding Selection and Subspace Expansion
- Debiasing Reward Models by Representation Learning with Guarantees
- Sentiment and Volatility in Financial Markets: A Review of BERT and GARCH Applications during Geopolitical Crises
- Variational Masked Diffusion Models
- Minimizing Human Intervention in Online Classification
- FRONTIER-RevRec: A Large-scale Dataset for Reviewer Recommendation
- COOPERA: Continual Open-Ended Human-Robot Assistance
- Education Paradigm Shift To Maintain Human Competitive Advantage Over AI
- Code Contribution and Credit in Science
- Machine Learning for Climate Policy: Understanding Policy Progression in the European Green Deal
- Large language model-based task planning for service robots: A review
- SwiftEmbed: Ultra-Fast Text Embeddings via Static Token Lookup for Real-Time Applications
- Network Intrusion Detection: Evolution from Conventional Approaches to LLM Collaboration and Emerging Risks
- Provable test-time adaptivity and distributional robustness of in-context learning
- A high-capacity linguistic steganography based on entropy-driven rank-token mapping
- Capsule Network-Based Multimodal Fusion for Mortgage Risk Assessment from Unstructured Data Sources
- How Muon's Spectral Design Benefits Generalization: A Study on Imbalanced Data
- Survey of Multimodal Geospatial Foundation Models: Techniques, Applications, and Challenges
- Encoder-Decoder Diffusion Language Models for Efficient Training and Inference
- Exploration of Summarization by Generative Language Models for Automated Scoring of Long Essays
- In Generative AI We (Dis)Trust? Computational Analysis of Trust and Distrust in Reddit Discussions
- Agentic Meta-Orchestrator for Multi-task Copilots
- Multi-Modal Fact-Verification Framework for Reducing Hallucinations in Large Language Models
- REVISION:Reflective Intent Mining and Online Reasoning Auxiliary for E-commerce Visual Search System Optimization
- Cross-view Localization and Synthesis -- Datasets, Challenges and Opportunities
- SALSA: Single-pass Autoregressive LLM Structured Classification
- Do Stop Me Now: Detecting Boilerplate Responses with a Single Iteration
- FastVLM: Self-Speculative Decoding for Fast Vision-Language Model Inference
- JiuTian Chuanliu: A Large Spatiotemporal Model for General-purpose Dynamic Urban Sensing
- A Framework for Quantifying How Pre-Training and Context Benefit In-Context Learning
- Text to Trust: Evaluating Fine-Tuning and LoRA Trade-offs in Language Models for Unfair Terms of Service Detection
- Frustratingly Easy Task-aware Pruning for Large Language Models
- Contextual Tokenization for Graph Inverted Indices
- E2Rank: Your Text Embedding can Also be an Effective and Efficient Listwise Reranker
- Chitchat with AI: Understand the supply chain carbon disclosure of companies worldwide through Large Language Model
- Expert Merging in Sparse Mixture of Experts with Nash Bargaining
- PromptReverb: Multimodal Room Impulse Response Generation Through Latent Rectified Flow Matching
- GigaEmbeddings: Efficient Russian Language Embedding Model
- Multilingual Target-Stance Extraction
- PatenTEB: A Comprehensive Benchmark and Model Family for Patent Text Embedding
- Understanding Self-Admitted Technical Debt in Test Code: An Empirical Study
- RoGBot: Relationship-Oblivious Graph-based Neural Network with Contextual Knowledge for Bot Detection
- SentiMaithili: A Benchmark Dataset for Sentiment and Reason Generation for the Low-Resource Maithili Language
- Learning 3D Anisotropic Noise Distributions Improves Molecular Force Field Modeling
- Gradual Forgetting: Logarithmic Compression for Extending Transformer Context Windows
- Energy-Efficient Domain-Specific Artificial Intelligence Models and Agents: Pathways and Paradigms
- Foundation of Intelligence: Review of Math Word Problems from Human Cognition Perspective
- From Black-box to Causal-box: Towards Building More Interpretable Models
- ArchISMiner: A Framework for Automatic Mining of Architectural Issue-Solution Pairs from Online Developer Communities
- Spatially Aware Linear Transformer (SAL-T) for Particle Jet Tagging
- Efficient semantic uncertainty quantification in language models via diversity-steered sampling
- PARL: Prompt-based Agents for Reinforcement Learning
- Flight Delay Prediction via Cross-Modality Adaptation of Large Language Models and Aircraft Trajectory Representation
- HalleluBERT: Let every token that has meaning bear its weight
- SindBERT, the Sailor: Charting the Seas of Turkish NLP
- Typoglycemia under the Hood: Investigating Language Models' Understanding of Scrambled Words
- Evaluating Prompting Strategies and Large Language Models in Systematic Literature Review Screening: Relevance and Task-Stage Classification
- Understanding AI Trustworthiness: A Scoping Review of AIES & FAccT Articles
- LLMComp: A Language Modeling Paradigm for Error-Bounded Scientific Data Compression (Technical Report)
- Large Language Models Meet Text-Attributed Graphs: A Survey of Integration Frameworks and Applications
- R2ComSync: Improving Code-Comment Synchronization with In-Context Learning and Reranking
- DictPFL: Efficient and Private Federated Learning on Encrypted Gradients
- Bridging Language Gaps with Adaptive RAG: Improving Indonesian Language Question Answering
- The Universal Landscape of Human Reasoning
- FSRF: Factorization-guided Semantic Recovery for Incomplete Multimodal Sentiment Analysis
- Multimodal Detection of Fake Reviews using BERT and ResNet-50
- Leverage Unlearning to Sanitize LLMs
- Fair Representation Learning with Controllable High Confidence Guarantees via Adversarial Inference
- Learning Grouped Lattice Vector Quantizers for Low-Bit LLM Compression
- Memory Constrained Dynamic Subnetwork Update for Transfer Learning
- CantoNLU: A benchmark for Cantonese natural language understanding
- TernaryCLIP: Efficiently Compressing Vision-Language Models with Ternary Weights and Distilled Knowledge
- BUSTED at AraGenEval Shared Task: A Comparative Study of Transformer-Based Models for Arabic AI-Generated Text Detection
- Automated Extraction of Fluoropyrimidine Treatment and Treatment-Related Toxicities from Clinical Notes Using Natural Language Processing
- Multimodal Negative Learning
- IKnow: Instruction-Knowledge-Aware Continual Pretraining for Effective Domain Adaptation
- GRATING: Low-Latency and Memory-Efficient Semantic Selection on Device
- ComProScanner: A multi-agent based framework for composition-property structured data extraction from scientific literature
- Calibrating Multimodal Consensus for Emotion Recognition
- Context-level Language Modeling by Learning Predictive Context Embeddings
- KCM: KAN-Based Collaboration Models Enhance Pretrained Large Models
- DSSmoothing: Toward Certified Dataset Ownership Verification for Pre-trained Language Models via Dual-Space Smoothing
- Hierarchical Dual-Head Model for Suicide Risk Assessment via MentalRoBERTa
- DyPE: Dynamic Position Extrapolation for Ultra High Resolution Diffusion
- MS-BART: Unified Modeling of Mass Spectra and Molecules for Structure Elucidation
- RECALL: REpresentation-aligned Catastrophic-forgetting ALLeviation via Hierarchical Model Merging
- Classical Feature Embeddings Help in BERT-Based Human Mobility Prediction
- From Masks to Worlds: A Hitchhiker's Guide to World Models
- Meta-Learning for Cross-Task Generalization in Protein Mutation Property Prediction
- On the Detectability of LLM-Generated Text: What Exactly Is LLM-Generated Text?
- Not-a-Bandit: Provably No-Regret Drafter Selection in Speculative Decoding for LLMs
- Forging GEMs: Advancing Greek NLP through Quality-Based Corpus Curation
- Adapting Multilingual Models to Code-Mixed Tasks via Model Merging
- Top-P Masking for Cross Language Information Retrieval
- KARIPAP: Quantum-Inspired Tensor Network Compression of Large Language Models Using Infinite Projected Entangled Pair States and Tensor Renormalization Group
- Study of Training Dynamics for Memory-Constrained Fine-Tuning
- Automated HIV Screening on Dutch Electronic Health Records with Large Language Models
- What is the Best Sequence Length for BABYLM?
- Learning Noise-Resilient and Transferable Graph-Text Alignment via Dynamic Quality Assessment
- AMAuT: A Flexible and Efficient Multiview Audio Transformer Framework Trained from Scratch
- Algorithmic Fairness in NLP: Persona-Infused LLMs for Human-Centric Hate Speech Detection
- Loopholing Discrete Diffusion: Deterministic Bypass of the Sampling Wall
- Vision-Based Mistake Analysis in Procedural Activities: A Review of Advances and Challenges
- Difficulty-Controllable Multiple-Choice Question Generation Using Large Language Models and Direct Preference Optimization
- Selecting and Combining Large Language Models for Scalable Code Clone Detection
- No Intelligence Without Statistics: The Invisible Backbone of Artificial Intelligence
- Aligning Multilingual News for Stock Return Prediction
- Transfer Learning Beyond the Standard Model
- CrossNews-UA: A Cross-lingual News Semantic Similarity Benchmark for Ukrainian, Polish, Russian, and English
- Automated Concern Extraction from Textual Requirements of Cyber-Physical Systems: A Multi-solution Study
- No Compute Left Behind: Rethinking Reasoning and Sampling with Masked Diffusion Models
- A Derandomization Framework for Structure Discovery: Applications in Neural Networks and Beyond
- LAPRAD: LLM-Assisted PRotocol Attack Discovery
- RailS: Load Balancing for All-to-All Communication in Distributed Mixture-of-Experts Training
- Modality Matching Matters: Calibrating Language Distances for Cross-Lingual Transfer in URIEL+
- Training-Free Spectral Fingerprints of Voice Processing in Transformers
- That's Deprecated! Understanding, Detecting, and Steering Knowledge Conflicts in Language Models for Code Generation
- See the Text: From Tokenization to Visual Reading
- When LRP Diverges from Leave-One-Out in Transformers
- Decoding Funded Research: Comparative Analysis of Topic Models and Uncovering the Effect of Gender and Geographic Location
- FeClustRE: Hierarchical Clustering and Semantic Tagging of App Features from User Reviews
- Topoformer: brain-like topographic organization in Transformer language models through spatial querying and reweighting
- SITS-DECO: A Generative Decoder Is All You Need For Multitask Satellite Image Time Series Modelling
- Reasoning Language Model Inference Serving Unveiled: An Empirical Study
- CLASP: Cost-Optimized LLM-based Agentic System for Phishing Detection
- CovMatch: Cross-Covariance Guided Multimodal Dataset Distillation with Trainable Text Encoder
- Beyond the Explicit: A Bilingual Dataset for Dehumanization Detection in Social Media
- Large language models for folktale type automation based on motifs: Cinderella case study
- Extracting alignment data in open models
- SLICE: SLO-Driven Scheduling for LLM Inference on Edge Computing Devices
- LLMs as Sparse Retrievers:A Framework for First-Stage Product Search
- CEFR-Annotated WordNet: LLM-Based Proficiency-Guided Semantic Database for Language Learning
- Embodied Navigation with Auxiliary Task of Action Description Prediction
- Benchmarking On-Device Machine Learning on Apple Silicon with MLX
- Learning to Flow from Generative Pretext Tasks for Neural Architecture Encoding
- GPTFace: Generative Pre-training of Facial-Linguistic Transformer by Span Masking and Weakly Correlated Text-image Data
- Combining Distantly Supervised Models with In Context Learning for Monolingual and Cross-Lingual Relation Extraction
- From Retrieval to Generation: Unifying External and Parametric Knowledge for Medical Question Answering
- BrailleLLM: Braille Instruction Tuning with Large Language Models for Braille Domain Tasks
- OpenInsGaussian: Open-vocabulary Instance Gaussian Segmentation with Context-aware Cross-view Fusion
- DiffGRM: Diffusion-based Generative Recommendation Model
- DuoLens: A Framework for Robust Detection of Machine-Generated Multilingual Text and Code
- Unifying and Enhancing Graph Transformers via a Hierarchical Mask Framework
- Ensembling Pruned Attention Heads For Uncertainty-Aware Efficient Transformers
- ScaleNet: Scaling up Pretrained Neural Networks with Incremental Parameters
- Engagement Undermines Safety: How Stereotypes and Toxicity Shape Humor in Language Models
- AgentChangeBench: A Multi-Dimensional Evaluation Framework for Goal-Shift Robustness in Conversational AI
- Any-Depth Alignment: Unlocking Innate Safety Alignment of LLMs to Any-Depth
- Exemplar-Guided Planing: Enhanced LLM Agent for KGQA
- DynaQuery: A Self-Adapting Framework for Querying Structured and Multimodal Data
- Unbiased Gradient Low-Rank Projection
- AcademicEval: Live Long-Context LLM Benchmark
- PANER: A Paraphrase-Augmented Framework for Low-Resource Named Entity Recognition
- An Enhanced Dual Transformer Contrastive Network for Multimodal Sentiment Analysis
- Multilingual Text-to-Image Person Retrieval via Bidirectional Relation Reasoning and Aligning
- DETree: DEtecting Human-AI Collaborative Texts via Tree-Structured Hierarchical Representation Learning
- Disparities in Multilingual LLM-Based Healthcare Q&A
- Multilingual Clinical NER for Diseases and Medications Recognition in Cardiology Texts using BERT Embeddings
- AFRICAPTION: Establishing a New Paradigm for Image Captioning in African Languages
- MILES: Modality-Informed Learning Rate Scheduler for Balancing Multimodal Learning
- Exploration via Feature Perturbation in Contextual Bandits
- Graph Attention-Guided Search for Dense Multi-Agent Pathfinding
- RubiSCoT: A Framework for AI-Supported Academic Assessment
- Efficient Toxicity Detection in Gaming Chats: A Comparative Study of Embeddings, Fine-Tuned Transformers and LLMs
- Explainability of Large Language Models: Opportunities and Challenges toward Generating Trustworthy Explanations
- SOLE: Hardware-Software Co-design of Softmax and LayerNorm for Efficient Transformer Inference
- ParaVul: A Parallel Large Language Model and Retrieval-Augmented Framework for Smart Contract Vulnerability Detection
- DVAGen: Dynamic Vocabulary Augmented Generation
- HGAdapter: Hypergraph-based Adapters in Language Models for Code Summarization and Clone Detection
- Generation then Reconstruction: Accelerating Masked Autoregressive Models via Two-Stage Sampling
- MemoryBench: A Benchmark for Memory and Continual Learning in LLM Systems
- Conveying Meaning through Gestures: An Investigation into Semantic Co-Speech Gesture Generation
- Reasoning Distillation and Structural Alignment for Improved Code Generation
- OG-Rank: Learning to Rank Fast and Slow with Uncertainty and Reward-Trend Guided Adaptive Exploration
- Extended LSTM: Adaptive Feature Gating for Toxic Comment Classification
- An Efficient Framework for Whole-Page Reranking via Single-Modal Supervision
- Connecting Domains and Contrasting Samples: A Ladder for Domain Generalization
- Does Visual Grounding Enhance the Understanding of Embodied Knowledge in Large Language Models?
- Uncovering Brain-Like Hierarchical Patterns in Vision-Language Models through fMRI-Based Neural Encoding
- Efficient High-Accuracy PDEs Solver with the Linear Attention Neural Operator
- ProtoMol: Enhancing Molecular Property Prediction via Prototype-Guided Multimodal Learning
- Renaissance of RNNs in Streaming Clinical Time Series: Compact Recurrence Remains Competitive with Transformers
- Domain-Contextualized Concept Graphs: A Computable Framework for Knowledge Representation
- Closing the Curvature Gap: Full Transformer Hessians and Their Implications for Scaling Laws
- All You Need is One: Capsule Prompt Tuning with a Single Vector
- MOSAIC: Masked Objective with Selective Adaptation for In-domain Contrastive Learning
- Mixed-Precision Quantization for Language Models: Techniques and Prospects
- Advances in Pre-trained Language Models for Domain-Specific Text Classification: A Systematic Review
- TACL: Threshold-Adaptive Curriculum Learning Strategy for Enhancing Medical Text Understanding
- Scaling Beyond Context: A Survey of Multimodal Retrieval-Augmented Generation for Document Understanding
- Enhance Large Language Models as Recommendation Systems with Collaborative Filtering
- The Hidden Cost of Modeling P(X): Vulnerability to Membership Inference Attacks in Generative Text Classifiers
- Extending Audio Context for Long-Form Understanding in Large Audio-Language Models
- Mixture of Experts Approaches in Dense Retrieval Tasks
- Leveraging LLMs for Context-Aware Implicit Textual and Multimodal Hate Speech Detection
- Policy Transfer for Continuous-Time Reinforcement Learning: A (Rough) Differential Equation Approach
- Hyperparameter Optimization and Reproducibility in Deep Learning Model Training
- FarsiMCQGen: a Persian Multiple-choice Question Generation Framework
- Measuring the Effect of Disfluency in Multilingual Knowledge Probing Benchmarks
- DMRetriever: A Family of Models for Improved Text Retrieval in Disaster Management
- MOBIUS: Big-to-Mobile Universal Instance Segmentation via Multi-modal Bottleneck Fusion and Calibrated Decoder Pruning
- TokDrift: When LLM Speaks in Subwords but Code Speaks in Grammar
- LoRAverse: A Submodular Framework to Retrieve Diverse Adapters for Diffusion Models
- ChangingGrounding: 3D Visual Grounding in Changing Scenes
- TRI-DEP: A Trimodal Comparative Study for Depression Detection Using Speech, Text, and EEG
- A Hard-Label Black-Box Evasion Attack against ML-based Malicious Traffic Detection Systems
- Detecting Early and Implicit Suicidal Ideation via Longitudinal and Information Environment Signals on Social Media
- Fantastic (small) Retrievers and How to Train Them: mxbai-edge-colbert-v0 Tech Report
- TOUCH: Text-guided Controllable Generation of Free-Form Hand-Object Interactions
- Multi-modal video data-pipelines for machine learning with minimal human supervision
- Capturing Context-Aware Route Choice Semantics for Trajectory Representation Learning
- ToolPRM: Fine-Grained Inference Scaling of Structured Outputs for Function Calling
- Purifying Task Vectors in Knowledge-Aware Subspace for Model Merging
- First Attentions Last: Better Exploiting First Attentions for Efficient Transformer Training
- Eyes Wide Open: Ego Proactive Video-LLM for Streaming Video
- ToolTweak: An Attack on Tool Selection in LLM-based Agents
- Hierarchical Semantic Retrieval with Cobweb
- Hi-Agent: Hierarchical Vision-Language Agents for Mobile Device Control
- Large Reasoning Embedding Models: Towards Next-Generation Dense Retrieval Paradigm
- Retrofitting Small Multilingual Models for Retrieval: Matching 7B Performance with 300M Parameters
- MorphoBench: A Benchmark with Difficulty Adaptive to Model Reasoning
- Generalist vs Specialist Time Series Foundation Models: Investigating Potential Emergent Behaviors in Assessing Human Health Using PPG Signals
- Large Scale Retrieval for the LinkedIn Feed using Causal Language Models
- Joint Modeling of Big Five and HEXACO for Multimodal Apparent Personality-trait Recognition
- From Language to Locomotion: Retargeting-free Humanoid Control via Motion Latent Guidance
- Inferred global dense residue transition graphs from primary structure sequences enable protein interaction prediction via directed graph convolutional neural networks
- DROID: Dual Representation for Out-of-Scope Intent Detection
- Nondeterminism-Aware Optimistic Verification for Floating-Point Neural Networks
- FedHFT: Efficient Federated Finetuning with Heterogeneous Edge Clients
- Vgent: Graph-based Retrieval-Reasoning-Augmented Generation For Long Video Understanding
- Signature in Code Backdoor Detection, how far are we?
- Scaling Vision Transformers for Functional MRI with Flat Maps
- Unlocking Public Catalogues: Instruction-Tuning LLMs for ICD Coding of German Tumor Diagnoses
- MemoTime: Memory-Augmented Temporal Knowledge Graph Enhanced Large Language Model Reasoning
- LLM one-shot style transfer for Authorship Attribution and Verification
- DistilCLIP-EEG: Enhancing Epileptic Seizure Detection Through Multi-modal Learning and Knowledge Distillation
- A New Perspective on Transformers in Online Reinforcement Learning for Continuous Control
- Document Intelligence in the Era of Large Language Models: A Survey
- Generalist++: A Meta-learning Framework for Mitigating Trade-off in Adversarial Training
- Transformer-based Scalable Beamforming Optimization via Deep Residual Learning
- Higher Satisfaction, Lower Cost: A Technical Report on How LLMs Revolutionize Meituan's Intelligent Interaction Systems
- Convergence, design and training of continuous-time dropout as a random batch method
- A Conformation-Centric Generative Foundation Model for Linear Polymer Modeling and Design
- Universal Image Restoration Pre-training via Masked Degradation Classification
- OS-HGAdapter: Open Semantic Hypergraph Adapter for Large Language Models Assisted Entropy-Enhanced Image-Text Alignment
- Reinforcement Learning Meets Masked Generative Models: Mask-GRPO for Text-to-Image Generation
- Text Anomaly Detection with Simplified Isolation Kernel
- On the Reasoning Abilities of Masked Diffusion Language Models
- ProtoTopic: Prototypical Network for Few-Shot Medical Topic Modeling
- What "Not" to Detect: Negation-Aware VLMs via Structured Reasoning and Token Merging
- Chinese ModernBERT with Whole-Word Masking
- BoN Appetit Team at LeWiDi-2025: Best-of-N Test-time Scaling Can Not Stomach Annotation Disagreements (Yet)
- Detect Anything via Next Point Prediction
- Multitask finetuning and acceleration of chemical pretrained models for small molecule drug property prediction
- ProtoSiTex: Learning Semi-Interpretable Prototypes for Multi-label Text Classification
- Learning to Recognize Correctly Completed Procedure Steps in Egocentric Assembly Videos through Spatio-Temporal Modeling
- Efficient Adaptive Transformer: An Empirical Study and Reproducible Framework
- IP-Augmented Multi-Modal Malicious URL Detection Via Token-Contrastive Representation Enhancement and Multi-Granularity Fusion
- A Hierarchical Quantized Tokenization Framework for Task-Adaptive Graph Representation Learning
- Pretraining in Actor-Critic Reinforcement Learning for Robot Locomotion
- Traveling Salesman-Based Token Ordering Improves Stability in Homomorphically Encrypted Language Models
- Simple Projection Variants Improve ColBERT Performance
- Metronome: Efficient Scheduling for Periodic Traffic Jobs with Network and Priority Awareness
- FedMMKT:Co-Enhancing a Server Text-to-Image Model and Client Task Models in Multi-Modal Federated Learning
- Unveiling the Vulnerability of Graph-LLMs: An Interpretable Multi-Dimensional Adversarial Attack on TAGs
- BIGFix: Bidirectional Image Generation with Token Fixing
- Structure-aware Propagation Generation with Large Language Models for Fake News Detection
- Chimera: State Space Models Beyond Sequences
- Order from Chaos: Comparative Study of Ten Leading LLMs on Unstructured Data Categorization
- SAIL-Embedding Technical Report: Omni-modal Embedding Foundation Model
- EReLiFM: Evidential Reliability-Aware Residual Flow Meta-Learning for Open-Set Domain Generalization under Noisy Labels
- Ethic-BERT: An Enhanced Deep Learning Model for Ethical and Non-Ethical Content Classification
- SMEC: Rethinking Matryoshka Representation Learning for Retrieval Embedding Compression
- Attribution Graphs and Causal Probing for Mechanistic Discovery and Bias Repair in Multimodal Generative Learning
- Enhanced Pre-training of Graph Neural Networks for Million-Scale Heterogeneous Graphs
- Toward LLM-Supported Automated Assessment of Critical Thinking Subskills
- FlexPipe: Adapting Dynamic LLM Serving Through Inflight Pipeline Refactoring in Fragmented Serverless Clusters
- Integrating Sequential and Relational Modeling for User Events: Datasets and Prediction Tasks
- Dimension-Free Minimax Rates for Learning Pairwise Interactions in Attention-Style Models
- REGENT: Relevance-Guided Attention for Entity-Aware Multi-Vector Neural Re-Ranking
- QDER: Query-Specific Document and Entity Representations for Multi-Vector Document Re-Ranking
- FinVet: A Collaborative Framework of RAG and External Fact-Checking Agents for Financial Misinformation Detection
- Deconstructing Attention: Investigating Design Principles for Effective Language Modeling
- MeTA-LoRA: Data-Efficient Multi-Task Fine-Tuning for Large Language Models
- An Encoder-Integrated PhoBERT with Graph Attention for Vietnamese Token-Level Classification
- Towards Real-Time Fake News Detection under Evidence Scarcity
- Exploring and Leveraging Class Vectors for Classifier Editing
- CNSocialDepress: A Chinese Social Media Dataset for Depression Risk Detection and Structured Analysis
- A Theorem-Proving-Based Evaluation of Neural Semantic Parsing
- Fairness Metric Design Exploration in Multi-Domain Moral Sentiment Classification using Transformer-Based Models
- An Explorative Study on Distributed Computing Techniques in Training and Inference of Large Language Models
- Class Prototypes based Contrastive Learning for Classifying Multi-Label and Fine-Grained Educational Videos
- Evaluating Line-level Localization Ability of Learning-based Code Vulnerability Detection Models
- Efficient In-Memory Acceleration of Sparse Block Diagonal LLMs
- Reliable Cross-modal Alignment via Prototype Iterative Construction
- CoSPED: Consistent Soft Prompt Targeted Data Extraction and Defense
- DyKnow-RAG: Dynamic Knowledge Utilization Reinforcement Framework for Noisy Retrieval-Augmented Generation in E-commerce Search Relevance
- Deep Edge Filter: Return of the Human-Crafted Layer in Deep Learning
- Therapeutic AI and the Hidden Risks of Over-Disclosure: An Embedded AI-Literacy Framework for Mental Health Privacy
- Toward Human-Centered Readability Evaluation
- Rethinking deep learning: linear regression remains a key benchmark in predicting terrestrial water storage
- From Reasoning LLMs to BERT: A Two-Stage Distillation Framework for Search Relevance
- Catch-Only-One: Non-Transferable Examples for Model-Specific Authorization
- RV-HATE: Reinforced Multi-Module Voting for Implicit Hate Speech Detection
- End-to-end Sequence Labeling via Bi-directional LSTM-CNNs-CRF: A Reproducibility Study
- GapDNER: A Gap-Aware Grid Tagging Model for Discontinuous Named Entity Recognition
- Comparative Explanations via Counterfactual Reasoning in Recommendations
- QLENS: Towards A Quantum Perspective of Language Transformers
- PAGE: Prompt Augmentation for text Generation Enhancement
- Visual Odometry with Transformers
- Early Detection and Reduction of Memorisation for Domain Adaptation and Instruction Tuning
- CLASP: Training-Free LLM-Assisted Source Code Watermarking via Semantic-Preserving Transformations
- BioOSS: A Bio-Inspired Oscillatory State System with Spatio-Temporal Dynamics
- Structured Spectral Graph Representation Learning for Multi-label Abnormality Analysis from 3D CT Scans
- HiligayNER: A Baseline Named Entity Recognition Model for Hiligaynon
- Predict Training Data Quality via Its Geometry in Metric Space
- ADiP: Adaptive Precision Systolic Array for Matrix Multiplication Acceleration
- Self-Supervised Representation Learning with ID-Content Modality Alignment for Sequential Recommendation
- MCE: Towards a General Framework for Handling Missing Modalities under Imbalanced Missing Rates
- VOLTAGE: A Versatile Contrastive Learning based OCR Methodology for ultra low-resource scripts through Auto Glyph Feature Extraction
- Head-wise Adaptive Rotary Positional Encoding for Fine-Grained Image Generation
- SASER: Stego attacks on open-source LLMs
- Multimodal Retrieval-Augmented Generation with Large Language Models for Medical VQA
- NIM: Neuro-symbolic Ideographic Metalanguage for Inclusive Communication
- On the Problem of Consistent Anomalies in Zero-Shot Industrial Anomaly Detection
- Knowing Unknowns in an Age of Information Overload
- FLAMMABLE: A Multi-Model Federated Learning Framework with Multi-Model Engagement and Adaptive Batch Sizes
- LLM-Based Multi-Task Bangla Hate Speech Detection: Type, Severity, and Target
- MetaBreak: Jailbreaking Online LLM Services via Special Token Manipulation
- B2N3D: Progressive Learning from Binary to N-ary Relationships for 3D Object Grounding
- Meronymic Ontology Extraction via Large Language Models
- A Survey of Inductive Reasoning for Large Language Models
- LLMs are All You Need? Improving Fuzz Testing for MOJO with Large Language Models
- SaFiRe: Saccade-Fixation Reiteration with Mamba for Referring Image Segmentation
- ReMix: Towards a Unified View of Consistent Character Generation and Editing
- Hybrid OCR-LLM Framework for Enterprise-Scale Document Information Extraction Under Copy-heavy Task
- CLMN: Concept based Language Models via Neural Symbolic Reasoning
- Translution: Unifying Self-attention and Convolution for Adaptive and Relative Modeling
- Collaborative Learning of Semantic-Aware Feature Learning and Label Recovery for Multi-Label Image Recognition with Incomplete Labels
- PANTHER: Generative Pretraining Beyond Language for Sequential User Behavior Modeling
- Improving Speech Emotion Recognition with Mutual Information Regularized Generative Model
- Probabilistic Hyper-Graphs using Multiple Randomly Masked Autoencoders for Semi-supervised Multi-modal Multi-task Learning
- Lightweight Baselines for Medical Abstract Classification: DistilBERT with Cross-Entropy as a Strong Default
- Unpacking Hateful Memes: Presupposed Context and False Claims
- ACT: Automatically Generating Compiler Backends from Tensor Accelerator ISA Descriptions
- Seeing Hate Differently: Hate Subspace Modeling for Culture-Aware Hate Speech Detection
- Interpretable Graph-Language Modeling for Detecting Youth Illicit Drug Use
- Serialized EHR make for good text representations
- Tool Use Enables Undetectable Steganography in Multi-Agent LLM Systems
- Exploring Large Language Models for Financial Applications: Techniques, Performance, and Challenges with FinMA
- Arab Spring’s impact on science through the lens of scholarly attention, funding, and migration
- Towards Speeding up Program Repair with Non-Autoregressive Model
- Sparse Query Attention (SQA): A Computationally Efficient Attention Mechanism with Query Heads Reduction
- Detecting LLM-Generated Spam Reviews by Integrating Language Model Embeddings and Graph Neural Network
- Comparison of Unsupervised Metrics for Evaluating Judicial Decision Extraction
- OntoLogX: Ontology-Guided Knowledge Graph Extraction from Cybersecurity Logs with Large Language Models
- Adversarial News and Lost Profits: Manipulating Headlines in LLM-Driven Algorithmic Trading
- Text2Token: Unsupervised Text Representation Learning with Token Target Prediction
- Lost in the Middle: An Emergent Property from Information Retrieval Demands in LLMs
- HUME: Measuring the Human-Model Performance Gap in Text Embedding Tasks
- Myopic Bayesian Decision Theory for Batch Active Learning with Partial Batch Label Sampling
- Cluster-Aware Prompt Ensemble Learning for Few-Shot Vision-Language Model Adaptation
- Steering Embedding Models with Geometric Rotation: Mapping Semantic Relationships Across Languages and Models
- PromptGuard at BLP-2025 Task 1: A Few-Shot Classification Framework Using Majority Voting and Keyword Similarity for Bengali Hate Speech Detection
- Patentformer: A demonstration of AI-assisted automated patent drafting
- Titans Revisited: A Lightweight Reimplementation and Critical Analysis of a Test-Time Memory Model
- Evaluating Robustness of Large Language Models Against Multilingual Typographical Errors
- On the Representations of Entities in Auto-regressive Large Language Models
- Entropy Meets Importance: A Unified Head Importance-Entropy Score for Stable and Efficient Transformer Pruning
- CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts
- MSDM: Generating Task-Specific Pathology Images with a Multimodal Conditioned Diffusion Model for Cell and Nuclei Segmentation
- Training Models to Detect Successive Robot Errors from Human Reactions
- Exploring Cross-Lingual Knowledge Transfer via Transliteration-Based MLM Fine-Tuning for Critically Low-resource Chakma Language
- Web Crawler Restrictions, AI Training Datasets & Political Biases
- Value-State Gated Attention for Mitigating Extreme-Token Phenomena in Transformers
- Promptimizer: User-Led Prompt Optimization for Personal Content Classification
- SQS: Bayesian DNN Compression through Sparse Quantized Sub-distributions
- The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections
- Bridging the Semantic Gap: Contrastive Rewards for Multilingual Text-to-SQL with GRPO
- Neural Codecs as Biosignal Tokenizers
- Co-Authoring the Self: A Human-AI Interface for Interest Reflection in Recommenders
- Training Feature Attribution for Vision Models
- SEER: Sustainability Enhanced Engineering of Software Requirements
- Toward Autonomous Soft Robotic Endovascular Navigation via Imitation Learning
- One Sentence, Two Embeddings: Contrastive Learning of Explicit and Implicit Semantic Representations
- Geodesic Calculus on Implicitly Defined Latent Manifolds
- SeCon-RAG: A Two-Stage Semantic Filtering and Conflict-Free Framework for Trustworthy RAG
- PairSem: LLM-Guided Pairwise Semantic Matching for Scientific Document Retrieval
- Diffusion-Inspired Masked Fine-Tuning for Knowledge Injection in Autoregressive LLMs
- Edu-EmotionNet: Cross-Modality Attention Alignment with Temporal Feedback Loops
- Q-Router: Agentic Video Quality Assessment with Expert Model Routing and Artifact Localization
- From Noise to Signal to Selbstzweck: Reframing Human Label Variation in the Era of Post-training in NLP
- Forecasting the Buzz: Enriching Hashtag Popularity Prediction with LLM Reasoning
- Kelp: A Streaming Safeguard for Large Models via Latent Dynamics-Guided Risk Detection
- Neuron-Level Analysis of Cultural Understanding in Large Language Models
- FedDTRE: Federated Dialogue Generation Models Powered by Trustworthiness Evaluation
- TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance
- Past, Present, and Future of Bug Tracking in the Generative AI Era
- SHE: Stepwise Hybrid Examination Reinforcement Learning Framework for E-commerce Search Relevance
- DISCO: Diversifying Sample Condensation for Efficient Model Evaluation
- Towards Human-Like Grading: A Unified LLM-Enhanced Framework for Subjective Question Evaluation
- Standard-to-Dialect Transfer Trends Differ across Text and Speech: A Case Study on Intent and Topic Classification in German Dialects
- Multilingual Generative Retrieval via Cross-lingual Semantic Compression
- Textual Entailment and Token Probability as Bias Evaluation Metrics
- Instance Relation Learning Network with Label Knowledge Propagation for Few-shot Multi-label Intent Detection
- Causality Guided Representation Learning for Cross-Style Hate Speech Detection
- Learning to Look at the Other Side: A Semantic Probing Study of Word Embeddings in LLMs with Enabled Bidirectional Attention
- AutoQual: An LLM Agent for Automated Discovery of Interpretable Features for Review Quality Assessment
- Vanishing Contributions: A Unified Framework for Smooth and Iterative Model Compression
- Improving Reasoning for Diffusion Language Models via Group Diffusion Policy Optimization
- Struc-EMB: The Potential of Structure-Aware Encoding in Language Embeddings
- Multi-modal Foundation Model for Cosmological Simulation Data
- Two-Stage Voting for Robust and Efficient Suicide Risk Detection on Social Media
- Learning What to Remember: Adaptive Probabilistic Memory Retention for Memory-Efficient Language Models
- FlyLoRA: Boosting Task Decoupling and Parameter Efficiency via Implicit Rank-Wise Mixture-of-Experts
- ZeroCard: Cardinality Estimation with Zero Dependence on Target Databases -- No Data, No Query, No Retraining
- DACIP-RC: Domain Adaptive Continual Instruction Pre-Training via Reading Comprehension on Business Conversations
- CommandSans: Securing AI Agents with Surgical Precision Prompt Sanitization
- Counterfactually Fair Conformal Prediction
- CAT: Curvature-Adaptive Transformers for Geometry-Aware Learning
- Multi-Task Pre-Finetuning of Lightweight Transformer Encoders for Text Classification and NER
- Investigating Thematic Patterns and User Preferences in LLM Interactions using BERTopic
- Label Semantics for Robust Hyperspectral Image Classification
- Phase Diagram of Dropout for Two-Layer Neural Networks in the Mean-Field Regime
- Lemma Dilemma: On Lemma Generation Without Domain- or Language-Specific Training Data
- Learning to Decide with Just Enough: Information-Theoretic Context Summarization for CMDPs
- Artificial Hippocampus Networks for Efficient Long-Context Modeling
- Evolutionary Profiles for Protein Fitness Prediction
- Language Lives in Sparse Dimensions: Toward Interpretable and Efficient Multilingual Control for Large Language Models
- Sunflower: A New Approach To Expanding Coverage of African Languages in Large Language Models
- Biasless Language Models Learn Unnaturally: How LLMs Fail to Distinguish the Possible from the Impossible
- Reasoning for Hierarchical Text Classification: The Case of Patents
- Revisiting Metric Reliability for Fine-grained Evaluation of Machine Translation and Summarization in Indian Languages
- Search-R3: Unifying Reasoning and Embedding in Large Language Models
- Unlocking Latent Discourse Translation in LLMs Through Quality-Aware Decoding
- SoftMatcha 2: A Fast and Soft Pattern Matcher for Trillion-Scale Corpora
- End-to-End Test-Time Training for Long Context
- BOTANIC-0: a series of foundation models for plant genomic data
- Iterative design of a NAND hybrid riboswitch by deep batch Bayesian optimization
- Pool Me Wisely: On the Effect of Pooling in Transformer-Based Models
- Mid-Training of Large Language Models: A Survey
- Efficient Discriminative Joint Encoders for Large Scale Vision-Language Reranking
- Study on LLMs for Promptagator-Style Dense Retriever Training
- TWIST: Training-free and Label-free Short Text Clustering through Iterative Vector Updating with LLMs
- AWM: Accurate Weight-Matrix Fingerprint for Large Language Models
- Heptapod: Language Modeling on Visual Signals
- The Effect of Attention Head Count on Transformer Approximation
- Modeling Sarcastic Speech: Semantic and Prosodic Cues in a Speech Synthesis Framework
- Learning to Route LLMs from Bandit Feedback: One Policy, Many Trade-offs
- TGM: a Modular and Efficient Library for Machine Learning on Temporal Graphs
- A Framework for Measuring How News Topics Drive Stock Movement
- HEMERA: A Human-Explainable Transformer Model for Estimating Lung Cancer Risk using GWAS Data
- Relational Transformer: Toward Zero-Shot Foundation Models for Relational Data
- LLM Bias Detection and Mitigation through the Lens of Desired Distributions
- Multi-Task Reinforcement Learning with Language-Encoded Gated Policy Networks
- Parallel Tokenizers: Rethinking Encoder Models' Vocabulary Design in Cross-Lingual Transfer of Low-Resource Languages
- lm-Meter: Unveiling Runtime Inference Latency for On-Device Language Models
- Shaken or Stirred? An Analysis of MetaFormer's Token Mixing for Medical Imaging
- Gradient-Sign Masking for Task Vector Transport Across Pre-Trained Models
- Multimodal Trajectory Representation Learning for Travel Time Estimation
- FoleyGRAM: Video-to-Audio Generation with GRAM-Aligned Multimodal Encoders
- Mixture of Neuron Experts
- Early Multimodal Prediction of Cross-Lingual Meme Virality on Reddit: A Time-Window Analysis
- Vipera: Blending Visual and LLM-Driven Guidance for Systematic Auditing of Text-to-Image Generative AI
- Membership Inference Attacks on Tokenizers of Large Language Models
- Oracle-Guided Masked Contrastive Reinforcement Learning for Visuomotor Policies
- Fine-Tuning Masked Diffusion for Provable Self-Correction
- Paraplume: A fast and accurate paratope prediction method provides insights into repertoire-scale binding dynamics
- MetaVLA: Unified Meta Co-training For Efficient Embodied Adaption
- BuilderBench -- A benchmark for generalist agents
- Automated Research Article Classification and Recommendation Using NLP and ML
- InforME: Improving Informativeness of Abstractive Text Summarization With Informative Attention Guided by Named Entity Salience
- MASA: Rethinking the Representational Bottleneck in LoRA with Multi-A Shared Adaptation
- The New Quant: A Survey of Large Language Models in Financial Prediction and Trading
- Boomerang Distillation Enables Zero-Shot Model Size Interpolation
- Decoding Partial Differential Equations: Cross-Modal Adaptation of Decoder-only Models to PDEs
- Stratum: System-Hardware Co-Design with Tiered Monolithic 3D-Stackable DRAM for Efficient MoE Serving
- KEEP: Integrating Medical Ontologies with Clinical Data for Robust Code Embeddings
- Revealing Interconnections between Diseases: from Statistical Methods to Large Language Models
- Beyond Random: Automatic Inner-loop Optimization in Dataset Distillation
- ModernBERT + ColBERT: Enhancing biomedical RAG through an advanced re-ranking retriever
- Contrastive Learning Using Graph Embeddings for Domain Adaptation of Language Models in the Process Industry
- The Role of Feature Interactions in Graph-based Tabular Deep Learning
- Transformers Discover Molecular Structure Without Graph Priors
- Language models for longitudinal analysis of abusive content in Billboard Music Charts
- Domain Generalization Under Posterior Drift
- Residualized Similarity for Faithfully Explainable Authorship Verification
- On the Limitations and Capabilities of Position Embeddings for Length Generalization
- GRACE: Generative Representation Learning via Contrastive Policy Optimization
- How Different from the Past? Spatio-Temporal Time Series Forecasting with Self-Supervised Deviation Learning
- Time Is Effort: Estimating Human Post-Editing Time for Grammar Error Correction Tool Evaluation
- Machine Learning for Detection and Analysis of Novel LLM Jailbreaks
- Representation Potentials of Foundation Models for Multimodal Alignment: A Survey
- HoRA: Cross-Head Low-Rank Adaptation with Joint Hypernetworks
- Rounding-Guided Backdoor Injection in Deep Learning Model Quantization
- Enhancing Talent Search Ranking with Role-Aware Expert Mixtures and LLM-based Fine-Grained Job Descriptions
- Thai Semantic End-of-Turn Detection for Real-Time Voice Agents
- Critical appraisal of artificial intelligence for rare-event recognition: principles and pharmacovigilance case studies
- Large Language Models Hallucination: A Comprehensive Survey
- PABSA: Hybrid Framework for Persian Aspect-Based Sentiment Analysis
- Beyond Simple Fusion: Adaptive Gated Fusion for Robust Multimodal Sentiment Analysis
- Refactoring with LLMs: Bridging Human Expertise and Machine Understanding
- Multimodal Learning with Augmentation Techniques for Natural Disaster Assessment
- Annotate Rhetorical Relations with INCEpTION: A Comparison with Automatic Approaches
- Allocation of Parameters in Transformers
- Optimizing Fine-Tuning through Advanced Initialization Strategies for Low-Rank Adaptation
- Evolutionary Computation as Natural Generative AI
- What Is The Performance Ceiling of My Classifier? Utilizing Category-Wise Influence Functions for Pareto Frontier Analysis
- Fine-Tuning Large Language Models with QLoRA for Offensive Language Detection in Roman Urdu-English Code-Mixed Text
- Towards Sampling Data Structures for Tensor Products in Turnstile Streams
- Efficient Test-Time Scaling for Small Vision-Language Models
- Model-Based Ranking of Source Languages for Zero-Shot Cross-Lingual Transfer
- Neural Correlates of Language Models Are Specific to Human Language
- Grounding Large Language Models in Clinical Evidence: A Retrieval-Augmented Generation System for Querying UK NICE Clinical Guidelines
- Evaluating Embedding Frameworks for Scientific Domain
- Improving Adversarial Robustness of Zero-Shot CLIP with Confidence-Aware Weighting
- C2|Q>: A Robust Framework for Bridging Classical and Quantum Software Development
- Bayesian Test-time Adaptation for Object Recognition and Detection with Vision-language Models
- IndiCASA: A Dataset and Bias Evaluation Framework in LLMs Using Contrastive Embedding Similarity in the Indian Context
- Team Xiaomi EV-AD VLA: Caption-Guided Retrieval System for Cross-Modal Drone Navigation -- Technical Report for IROS 2025 RoboSense Challenge Track 4
- PGMEL: Policy Gradient-based Generative Adversarial Network for Multimodal Entity Linking
- Hyperparameter Loss Surfaces Are Simple Near their Optima
- AutoMaAS: Self-Evolving Multi-Agent Architecture Search for Large Language Models
- AgenticRAG: Tool-Augmented Foundation Models for Zero-Shot Explainable Recommender Systems
- HyperAdaLoRA: Accelerating LoRA Rank Allocation During Training via Hypernetworks without Sacrificing Performance
- Dissecting Transformers: A CLEAR Perspective towards Green AI
- Semantic Differentiation in Speech Emotion Recognition: Insights from Descriptive and Expressive Speech Roles
- Visual Language Model as a Judge for Object Detection in Industrial Diagrams
- VisitHGNN: Heterogeneous Graph Neural Networks for Modeling Point-of-Interest Visit Patterns
- Fine-Tuning on Noisy Instructions: Effects on Generalization and Performance
- Coevolutionary Continuous Discrete Diffusion: Make Your Diffusion Language Model a Latent Reasoner
- Brain-Language Model Alignment: Insights into the Platonic Hypothesis and Intermediate-Layer Advantage
- Geometric Properties of Neural Multivariate Regression
- In-Situ Tweedie Discrete Diffusion Models
- GLAI: GreenLightningAI for Accelerated Training through Knowledge Decoupling
- HalluGuard: Evidence-Grounded Small Reasoning Models to Mitigate Hallucinations in Retrieval-Augmented Generation
- What You See is What You Ask: Evaluating Audio Descriptions
- Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs
- Composer: A Search Framework for Hybrid Neural Architecture Design
- Sentry: Authenticating Machine Learning Artifacts on the Fly
- A Deep Learning Pipeline for Epilepsy Genomic Analysis Using GPT-2 XL and NVIDIA H100
- The Transformer Cookbook
- Neu-RadBERT for Enhanced Diagnosis of Brain Injuries and Conditions
- Foundation vs. Specialized Models: Evaluating Catastrophic Forgetting in Continual Time Series Forecasting
- Multi-Category Materials Information Extraction (Composition, Processing, Microstructure, Properties)
- SynCED-EnDe 2025: A Synthetic and Curated English - German Dataset for Critical Error Detection in Machine Translation
- Scrolling Through Chaos: The Implications of TikTok for Crisis Sensemaking
- How Foundational are Foundation Models for Time Series Forecasting?
- Social Processes in the Intensification of Online Hate: The Effects of Verbal Replies to Anti-Muslim and Anti-Jewish Posts Following 7 October 2023
- From Factoid Questions to Data Product Requests: Benchmarking Data Product Discovery over Tables and Text
- Automated Alignment of Math Items to Content Standards in Large-Scale Assessments Using Language Models
- Efficient Layer-wise LLM Fine-tuning for Revision Intention Prediction
- MultiFair: Multimodal Balanced Fairness-Aware Medical Classification with Dual-Level Gradient Modulation
- SecureBERT 2.0: Advanced Language Model for Cybersecurity Intelligence
- BiasFreeBench: a Benchmark for Mitigating Bias in Large Language Model Responses
- LoRAFusion: Efficient LoRA Fine-Tuning for LLMs
- Secure and Robust Watermarking for AI-generated Images: A Comprehensive Survey
- Optimizing What Matters: AUC-Driven Learning for Robust Neural Retrieval
- Probing the Critical Point (CritPt) of AI Reasoning: a Frontier Physics Research Benchmark
- TVS Sidekick: Challenges and Practical Insights from Deploying Large Language Models in the Enterprise
- Text-Based Approaches to Item Alignment to Content Standards in Large-Scale Reading & Writing Tests
- TimeScope: Towards Task-Oriented Temporal Grounding In Long Videos
- Memory-Driven Self-Improvement for Decision Making with Large Language Models
- NeuroTTT: Bridging Pretraining-Downstream Task Misalignment in EEG Foundation Models via Test-Time Training
- Type-Less yet Type-Aware Inductive Link Prediction with Pretrained Language Models
- The silence of the weights: an investigation of structural pruning strategies for attention-based audio signal architectures
- FITS: Towards an AI-Driven Fashion Information Tool for Sustainability
- MUVLA: Learning to Explore Object Navigation via Map Understanding
- Self-Supervised Anatomical Consistency Learning for Vision-Grounded Medical Report Generation
- Reliability Crisis of Reference-free Metrics for Grammatical Error Correction
- Accelerating LLM Inference with Precomputed Query Storage
- Understanding the Mixture-of-Experts with Nadaraya-Watson Kernel
- HiStyle: Hierarchical Style Embedding Predictor for Text-Prompt-Guided Controllable Speech Synthesis
- Adapting SAM with Dynamic Similarity Graphs for Few-Shot Parameter-Efficient Small Dense Object Detection: A Case Study of Chickpea Pods in Field Conditions
- Better with Less: Small Proprietary Models Surpass Large Language Models in Financial Transaction Understanding
- EchoingECG: An Electrocardiogram Cross-Modal Model for Echocardiogram Tasks
- From Cheap Geometry to Expensive Physics: Elevating Neural Operators via Latent Shape Pretraining
- ProbMed: A Probabilistic Framework for Medical Multimodal Binding
- QFrBLiMP: a Quebec-French Benchmark of Linguistic Minimal Pairs
- LMOD+: A Comprehensive Multimodal Dataset and Benchmark for Developing and Evaluating Multimodal Large Language Models in Ophthalmology
- CustomIR: Unsupervised Fine-Tuning of Dense Embeddings for Known Document Corpora
- Effective Model Pruning
- FlowMoE: A Scalable Pipeline Scheduling Framework for Distributed Mixture-of-Experts Training
- Hierarchical Reasoning Models: Perspectives and Misconceptions
- Uncovering Zero-Shot Generalization Gaps in Time-Series Foundation Models Using Real-World Videos
- SafePassage: High-Fidelity Information Extraction with Black Box LLMs
- RFG: Test-Time Scaling for Diffusion Large Language Model Reasoning with Reward-Free Guidance
- Building the EHR Foundation Model via Next Event Prediction
- FishNet++: Analyzing the capabilities of Multimodal Large Language Models in marine biology
- A Cartography of Open Collaboration in Open Source AI: Mapping Practices, Motivations, and Governance in 14 Open Large Language Model Projects
- Spontaneous High-Order Generalization in Neural Theory-of-Mind Networks
- GateMABSA: Aspect-Image Gated Fusion for Multimodal Aspect-based Sentiment Analysis
- SecInfer: Preventing Prompt Injection via Inference-time Scaling
- VIVALDy: A Hybrid Generative Reduced-Order Model for Turbulent Flows, Applied to Vortex-Induced Vibrations
- Inductive Bias and Spectral Properties of Single-Head Attention in High Dimensions
- Environment-Aware Satellite Image Generation with Diffusion Models
- Efficient Sketching and Nearest Neighbor Search Algorithms for Sparse Vector Sets
- Reference-Free Rating of LLM Responses via Latent Information
- "Stop replacing salt with sugar!'': Towards Intuitive Human-Agent Teaching
- Can you SPLICE it together? A Human Curated Benchmark for Probing Visual Reasoning in VLMs
- UniDex: Rethinking Search Inverted Indexing with Unified Semantic Modeling
- LEAF: A Robust Expert-Based Framework for Few-Shot Continual Event Detection
- AstroMMBench: A Benchmark for Evaluating Multimodal Large Language Models Capabilities in Astronomy
- Multi-Item-Query Attention for Stable Sequential Recommendation
- Dynamic Orchestration of Multi-Agent System for Real-World Multi-Image Agricultural VQA
- PEARL: Performance-Enhanced Aggregated Representation Learning
- ELASTIQ: EEG-Language Alignment with Semantic Task Instruction and Querying
- Graph Optimization Foundation Model: Tokenizing Graph via A Language-Model Paradigm
- Extracting the Structure of Press Releases for Predicting Earnings Announcement Returns
- Uni-NTFM: A Unified Foundation Model for EEG Signal Representation Learning
- Metamorphic Testing for Audio Content Moderation Software
- FM-FoG: A Real-Time Foundation Model-based Wearable System for Freezing-of-Gait Mitigation
- ASTROCO: Self-Supervised Conformer-Style Transformers for Light-Curve Embeddings
- Watermarking Diffusion Language Models
- AGNOMIN -- Architecture Agnostic Multi-Label Function Name Prediction
- Ultra-Fast Language Generation via Discrete Diffusion Divergence Instruct
- Meta-Router: Bridging Gold-standard and Preference-based Evaluations in Large Language Model Routing
- Navigating prokaryotic viral genome analysis from metagenomic data
- The limits of AI for authoritarian control
- G-reasoner: Foundation Models for Unified Reasoning over Graph-structured Knowledge
- Who invented deep residual learning?
- Evaluating Spatiotemporal Consistency in Automatically Generated Sewing Instructions
- Hype or not? Formalizing Automatic Promotional Language Detection in Biomedical Research
- The Rise of AfricaNLP: A Survey of Contributions, Contributors, Community Impact, and Bibliometric Analysis
- Learning to Parallel: Accelerating Diffusion Large Language Models via Learnable Parallel Decoding
- Reinforcement Mid-Training
- BTC-SAM: Leveraging LLMs for Generation of Bias Test Cases for Sentiment Analysis Models
- ResFormer: All-Time Reservoir Memory for Long Sequence Classification
- A Second-Order Perspective on Pruning at Initialization and Knowledge Transfer
- The Hidden Costs of Translation Accuracy: Distillation, Quantization, and Environmental Impact
- LLM/Agent-as-Data-Analyst: A Survey
- Detecting and Rectifying Noisy Labels: A Similarity-based Approach
- From Neural Networks to Logical Theories: The Correspondence between Fibring Modal Logics and Fibring Neural Networks
- Disentangling Score Content and Performance Style for Joint Piano Rendering and Transcription
- Investigating Multi-layer Representations for Dense Passage Retrieval
- Collaboration of Fusion and Independence: Hypercomplex-driven Robust Multi-Modal Knowledge Graph Completion
- Merge Now, Regret Later: The Hidden Cost of Model Merging is Adversarial Transferability
- Aligning LLMs for Multilingual Consistency in Enterprise Applications
- MotionVerse: A Unified Multimodal Framework for Motion Comprehension, Generation and Editing
- Fusing Sequence Motifs and Pan-Genomic Features: Antimicrobial Resistance Prediction using an Explainable Lightweight 1D CNN-XGBoost Ensemble
- FUSAR-KLIP: Towards Multimodal Foundation Models for Remote Sensing
- GPS-MTM: Capturing Pattern of Normalcy in GPS-Trajectories with self-supervised learning
- QuadEnhancer: Leveraging Quadratic Transformations to Enhance Deep Neural Networks
- Automated extraction of fungal trophic modes from literature using BioBERT: an open pilot workflow
- AraS2P: Arabic Speech-to-Phonemes System
- Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
- Modeling the language cortex with form-independent and enriched representations of sentence meaning reveals remarkable semantic abstractness
- An Senegalese Legal Texts Structuration Using LLM-augmented Knowledge Graph
- MimiTalk: Revolutionizing Qualitative Research with Dual-Agent AI
- CCD: Mitigating Hallucinations in Radiology MLLMs via Clinical Contrastive Decoding
- Language, Culture, and Ideology: Personalizing Offensiveness Detection in Political Tweets with Reasoning LLMs
- Transforming Opioid Poisoning Surveillance Through Novel Technologies: Rationale and Methodological Protocol for Applying Natural Language Processing to Emergency Department Data
- Robust Fine-Tuning from Non-Robust Pretrained Models: Mitigating Suboptimal Transfer With Adversarial Scheduling
- Scaling medical imaging report generation with multimodal reinforcement learning
- Fin-ExBERT: User Intent based Text Extraction in Financial Context using Graph-Augmented BERT and trainable Plugin
- VeriGRAG: Enhancing LLM-Based Verilog Code Generation with Structure-Aware Soft Prompts
- Knowledge distillation through geometry-aware representational alignment
- Understanding and Enhancing the Planning Capability of Language Models via Multi-Token Prediction
- Limit Analysis for Symbolic Multi-step Reasoning Tasks with Information Propagation Rules Based on Transformers
- WARBERT: A Hierarchical BERT-based Model for Web API Recommendation
- Dense associative memory on the Bures-Wasserstein space
- Mask What Matters: Controllable Text-Guided Masking for Self-Supervised Medical Image Analysis
- CoDA: Coding LM via Diffusion Adaptation
- Beyond Embeddings: Interpretable Feature Extraction for Binary Code Similarity
- Steering Prepositional Phrases in Language Models: A Case of with-headed Adjectival and Adverbial Complements in Gemma-2
- T-TAMER: Provably Taming Trade-offs in ML Serving
- Functional Critic Modeling for Provably Convergent Off-Policy Actor-Critic
- MindCraft: How Concept Trees Take Shape In Deep Models
- What If Moderation Didn't Mean Suppression? A Case for Personalized Content Transformation
- Learning to Detect Relevant Contexts and Knowledge for Response Selection in Retrieval-based Dialogue Systems
- What Do They Fix? LLM-Aided Categorization of Security Patches for Critical Memory Bugs
- Transformers for single-cell RNA sequencing: a survey
- Beyond statistical significance: Quantifying uncertainty and statistical variability in multilingual and multitask NLP evaluation
- Category Discovery: An Open-World Perspective
- REMA: A Unified Reasoning Manifold Framework for Interpreting Large Language Model
- Representing LLMs in Prompt Semantic Task Space
- Your RAG is Unfair: Exposing Fairness Vulnerabilities in Retrieval-Augmented Generation via Backdoor Attacks
- Learning the Neighborhood: Contrast-Free Multimodal Self-Supervised Molecular Graph Pretraining
- Stochastic activations
- Decoding quantum low density parity check codes with diffusion
- The InviTE Corpus: Annotating Invectives in Tudor English Texts for Computational Modeling
- SoDaDE: Solvent Data-Driven Embeddings with Small Transformer Models
- Aurora: Towards Universal Generative Multimodal Time Series Forecasting
- Bridging Fairness and Explainability: Can Input-Based Explanations Promote Fairness in Hate Speech Detection?
- Context Parametrization with Compositional Adapters
- Mixture of Detectors: A Compact View of Machine-Generated Text Detection
- Multi-Agent Path Finding via Offline RL and LLM Collaboration
- Multilingual Vision-Language Models, A Survey
- COSPADI: Compressing LLMs via Calibration-Guided Sparse Dictionary Learning
- RedNote-Vibe: A Dataset for Capturing Temporal Dynamics of AI-Generated Text in Social Media
- Fuzzy Reasoning Chain (FRC): An Innovative Reasoning Framework from Fuzziness to Clarity
- BrainPro: Towards Large-scale Brain State-aware EEG Representation Learning
- Taxonomy of Comprehensive Safety for Clinical Agents
- Task-Adaptive Parameter-Efficient Fine-Tuning for Weather Foundation Models
- EgoInstruct: An Egocentric Video Dataset of Face-to-face Instructional Interactions with Multi-modal LLM Benchmarking
- WAVE: Learning Unified & Versatile Audio-Visual Embeddings with Multimodal LLM
- Effect of Model Merging in Domain-Specific Ad-hoc Retrieval
- AutoSCORE: Enhancing Automated Scoring with Multi-Agent Large Language Models via Structured Component Recognition
- Graph of Agents: Principled Long Context Modeling by Emergent Multi-Agent Collaboration
- Semantic Agreement Enables Efficient Open-Ended LLM Cascades
- Towards Minimal Causal Representations for Human Multimodal Language Understanding
- SelfJudge: Faster Speculative Decoding via Self-Supervised Judge Verification
- Bridging Kolmogorov Complexity and Deep Learning: Asymptotically Optimal Description Length Objectives for Transformers
- Triple-BERT: Do We Really Need MARL for Order Dispatch on Ride-Sharing Platforms?
- KurdSTS: The Kurdish Semantic Textual Similarity
- RLP: Reinforcement as a Pretraining Objective
- Blockwise Hadamard high-Rank Adaptation for Parameter-Efficient LLM Fine-Tuning
- Semantic F1 Scores: Fair Evaluation Under Fuzzy Class Boundaries
- Towards Transparent AI: A Survey on Explainable Language Models
- C-QUERI: Congressional Questions, Exchanges, and Responses in Institutions Dataset
- GraphPFN: A Prior-Data Fitted Graph Foundation Model
- Dual-Head Reasoning Distillation: Improving Classifier Accuracy with Train-Time-Only Reasoning
- Filtering with Confidence: When Data Augmentation Meets Conformal Prediction
- Are Hallucinations Bad Estimations?
- Talking Trees: Reasoning-Assisted Induction of Decision Trees for Tabular Data
- Diagnosing the Performance Trade-off in Moral Alignment: A Case Study on Gender Stereotypes
- Decipher-MR: A Vision-Language Foundation Model for 3D MRI Representations
- AutoIntent: AutoML for Text Classification
- EvoMail: Self-Evolving Cognitive Agents for Adaptive Spam and Phishing Email Defense
- An Improved Quantum Software Challenges Classification Approach using Transfer Learning and Explainable AI
- Rejuvenating Cross-Entropy Loss in Knowledge Distillation for Recommender Systems
- A short survey on almost orthogonal vectors in a few specific large dimensions
- ArchGPT: Understanding the World's Architectures with Large Multimodal Models
- Prompt-Aware Scheduling for Low-Latency LLM Serving
- Explaining Grokking and Information Bottleneck through Neural Collapse Emergence
- Measuring LLM Sensitivity in Transformer-based Tabular Data Synthesis
- RAG Security and Privacy: Formalizing the Threat Model and Attack Surface
- Retrieval over Classification: Integrating Relation Semantics for Multimodal Relation Extraction
- RedHerring Attack: Testing the Reliability of Attack Detection
- WDformer: A Wavelet-based Differential Transformer Model for Time Series Forecasting
- Incorporating LLM Embeddings for Variation Across the Human Genome
- Alignment Unlocks Complementarity: A Framework for Multiview Circuit Representation Learning
- Multi-Modal Sentiment Analysis with Dynamic Attention Fusion
- Who's Laughing Now? An Overview of Computational Humour Generation and Explanation
- Unlocking Financial Insights: An advanced Multimodal Summarization with Multimodal Output Framework for Financial Advisory Videos
- GEP: A GCG-Based method for extracting personally identifiable information from chatbots built on small language models
- Understanding and Enhancing Mask-Based Pretraining towards Universal Representations
- Performance Consistency of Learning Methods for Information Retrieval Tasks
- Enhancing Python Programming Education with an AI-Powered Code Helper: Design, Implementation, and Impact
- Document Summarization with Conformal Importance Guarantees
- FAST: Foreground-aware Diffusion with Accelerated Sampling Trajectory for Segmentation-oriented Anomaly Synthesis
- Fulcrum: Optimizing Concurrent DNN Training and Inferencing on Edge Accelerators
- Pagoda: An Energy and Time Roofline Study for DNN Workloads on Edge Accelerators
- Synergistic Enhancement of Requirement-to-Code Traceability: A Framework Combining Large Language Model based Data Augmentation and an Advanced Encoder
- Probability Signature: Bridging Data Semantics and Embedding Structure in Language Models
- Tokenization and Representation Biases in Multilingual Models on Dialectal NLP Tasks
- Embodied AI: From LLMs to World Models
- An effective control of large systems of active particles: An application to evacuation problem
- Multimodal-enhanced Federated Recommendation: A Group-wise Fusion Approach
- Towards Self-Supervised Foundation Models for Critical Care Time Series
- CollaPipe: Adaptive Segment-Optimized Pipeline Parallelism for Collaborative LLM Training in Heterogeneous Edge Networks
- Faster, Smaller, and Smarter: Task-Aware Expert Merging for Online MoE Inference
- WEE-Therapy: A Mixture of Weak Encoders Framework for Psychological Counseling Dialogue Analysis
- Rectified Decoupled Dataset Distillation: A Closer Look for Fair and Comprehensive Evaluation
- A Unified Noise-Curvature View of Loss of Trainability
- Revisiting Performance Claims for Chest X-Ray Models Using Clinical Context
- Learning Contextual Retrieval for Robust Conversational Search
- Let's Play Across Cultures: A Large Multilingual, Multicultural Benchmark for Assessing Language Models' Understanding of Sports
- Every Character Counts: From Vulnerability to Defense in Phishing Detection
- Efficiently Attacking Memorization Scores
- SeMob: Semantic Synthesis for Dynamic Urban Mobility Prediction
- Feasibility of Structuring Stress Documentation Using an Ontology-Guided Large Language Model
- Anatomy of a Feeling: Narrating Embodied Emotions via Large Vision-Language Models
- ExPe: Exact Positional Encodings for Generative Transformer Models with Extrapolating Capabilities
- Uncertainty in Semantic Language Modeling with PIXELS
- Confidence Calibration in Large Language Model-Based Entity Matching
- Building a User Foundation Model for the Open Web
- HyperClaim: Fine-Grained Cross-Modal Hypergraph Reasoning for Video Misinformation Detection
- GyRot: Leveraging Hidden Synergy between Rotation and Fine-grained Group Quantization for Low-bit LLM Inference
- AWARE-FX: An Auditable Knowledge-Guided AI System for Measuring Corporate Foreign-Exchange Hedging Disclosure
- Drawing-Recode: Annotation Grounding for Parametric CAD Code Generation from Raster 2D CAD Drawings
- CDAE: Enhancing Perturbation Robustness in Pretrained Language Models with Contrastive Denoising
- Semantic-Aligned Structural Abstraction for Multimodal Sentiment Analysis
- Improving Mental Health Screening and Early Risk Detection in Spanish
- Gradient-free Task-Conditioned Retrieval for On-Device In-Context Learning
- Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer
- Does EEG Foundation Models Transfer to Speech? A Benchmark on Overt and Imagined Speech Decoding
- Leveraging Trajectory Graphs for Pre-Execution Error Diagnosis in Agentic LLM Systems
- AHA-Memes: A Fine-Grained Multimodal Benchmark for Understanding Hate in Arabic Memes
- Latent-Kernel Discrete Flow Maps for Few-Step Generation
- Learning to Select, Not Relearn: Hard-Routed Mixtures of Reasoning LoRAs
- Benchmarking Open-Ended Multi-Agent Coordination in Language Agents
- Quadratic Objective Perturbation: Curvature-Based Differential Privacy
- NorBERTo: A ModernBERT Model Trained for Portuguese with 331 Billion Tokens Corpus
- Home sharing as affordable housing for all? Revealing the exclusionary language of shared rental listings through AI
- From Large Language Model Predicates to Logic Tensor Networks: Neurosymbolic Offer Validation in Regulated Procurement
- How Can We Synthesize High-Quality Pretraining Data? A Systematic Study of Prompt Design, Generator Model, and Source Data
- A transformer-based multi-task deep learning model for urban livability evaluation by fusing remote sensing and textual geospatial data
- RELISH: LLM REgression with a Latent Iterative State Head
- Just Use XML: Revisiting Joint Translation and Label Projection
- KuaiSearch: An E-Commerce Search Dataset with Authentic Queries and Product Texts for Recall, Ranking, and Relevance
- GCT: A Granger-Causal Transformer for Multivariate Traffic Analysis in Smart Villages
- epiGPTope: A Machine Learning-Based Epitope Generator and Classifier
- Demographically-Inspired Query Variants Using an LLM
- Large Language Model Automated Extraction of Clinical Signs and Symptoms From Emergency Department Reports for Machine Learning Prediction Models: Development and Validation Study
- Bridging gaps in hate speech detection: meta-collections and benchmarks for low-resource Iberian languages
- scGMB: A scRNA‐seq Cell Classification Method Combining GCN and Mamba
- LexSemBridge: Fine-Grained Dense Representation Enhancement through Token-Aware Embedding Augmentation
- PLM-interact: extending protein language models to predict protein-protein interactions
- LLMulator: Generalizable Cost Modeling for Dataflow Accelerators with Input-Adaptive Control Flow
- Multiple agroecological practices use and climate change mitigation. A review
- ISACL: Internal State Analyzer for Copyrighted Training Data Leakage
- CEIDM: A Controlled Entity and Interaction Diffusion Model for Enhanced Text-to-Image Generation
- Theoretical Foundations of Representation Learning using Unlabeled Data: Statistics and Optimization
- From Global to Local: Social Bias Transfer in CLIP
- Randomly Removing 50% of Dimensions in Text Embeddings has Minimal Impact on Retrieval and Classification Tasks
- DRISHTIKON: A Multimodal Multilingual Benchmark for Testing Language Models' Understanding on Indian Culture
- SloPalSpeech: A 2,8000-Hour Slovak Speech Corpus from Parliamentary Data
- Reinforcement Learning on Pre-Training Data
- CompLLM: Compression for Long Context Q&A
- Systematic Comparative Analysis of Large Pretrained Language Models on Contextualized Medication Event Extraction
- Charting a Decade of Computational Linguistics in Italy: The CLiC-it Corpus
- Pure Vision Language Action (VLA) Models: A Comprehensive Survey
- Otters: An Energy-Efficient SpikingTransformer via Optical Time-to-First-Spike Encoding
- Enhancing the Effectiveness and Durability of Backdoor Attacks in Federated Learning through Maximizing Task Distinction
- Trigger Where It Hurts: Unveiling Hidden Backdoors through Sensitivity with Sensitron
- Semantic Search for Information Retrieval
- HyperAdapt: Simple High-Rank Adaptation
- TsqLoRA: Towards Sensitivity and Quality Low-Rank Adaptation for Efficient Fine-Tuning
- LCMF: Lightweight Cross-Modality Mambaformer for Embodied Robotics VQA
- Human-Annotated NER Dataset for the Kyrgyz Language
- Financial Risk Relation Identification through Dual-view Adaptation
- Trace Is In Sentences: Unbiased Lightweight ChatGPT-Generated Text Detector
- M4SER: Multimodal, Multirepresentation, Multitask, and Multistrategy Learning for Speech Emotion Recognition
- CCQA: Generating Question from Solution Can Improve Inference-Time Reasoning in SLMs
- Multi-Hierarchical Feature Detection for Large Language Model Generated Text
- An overview of neural architectures for self-supervised audio representation learning from masked spectrograms
- Advances in Large Language Models for Medicine
- Extracting Conceptual Spaces from LLMs Using Prototype Embeddings
- Reading Images Like Texts: Sequential Image Understanding in Vision-Language Models
- WolBanking77: Wolof Banking Speech Intent Classification Dataset
- Enhancing Transformer-Based Rerankers with Synthetic Data and LLM-Based Supervision
- Bounded PCTL Model Checking of Large Language Model Outputs
- The Narcissus Hypothesis: Descending to the Rung of Illusion
- Fine-Grained Detection of AI-Generated Text Using Sentence-Level Segmentation
- Enhancing Semantic Segmentation with Continual Self-Supervised Pre-training
- Accurate and Efficient Low-Rank Model Merging in Core Space
- A Generative Framework for Personalized Sticker Retrieval
- Breaking Token Into Concepts: Exploring Extreme Compression in Token Representation Via Compositional Shared Semantics
- Enhancing Cluster Scheduling in HPC: A Continuous Transfer Learning for Real-Time Optimization
- Turk-LettuceDetect: A Hallucination Detection Models for Turkish RAG Applications
- Overview of PlantCLEF 2022: Image-based plant identification at global scale
- COLA: Context-aware Language-driven Test-time Adaptation
- Training-Free Label Space Alignment for Universal Domain Adaptation
- Pre-Trained CNN Architecture for Transformer-Based Image Caption Generation Model
- LLaVul: A Multimodal LLM for Interpretable Vulnerability Reasoning about Source Code
- Towards Open-Ended Discovery for Low-Resource NLP
- Chat-CBM: Towards Interactive Concept Bottleneck Models with Frozen Large Language Models
- ReDepress: A Cognitive Framework for Detecting Depression Relapse from Social Media
- Modeling the Attack: Detecting AI-Generated Text by Quantifying Adversarial Perturbations
- Make Every Letter Count: Building Dialect Variation Dictionaries from Monolingual Corpora
- Transformer-Gather, Fuzzy-Reconsider: A Scalable Hybrid Framework for Entity Resolution
- Individualized non-uniform quantization for vector search
- SingLEM: Single-Channel Large EEG Model
- RadEval: A framework for radiology text evaluation
- TraceHiding: Scalable Machine Unlearning for Mobility Data
- nDNA -- the Semantic Helix of Artificial Cognition
- STAR: Speech-to-Audio Generation via Representation Learning
- MaskVCT: Masked Voice Codec Transformer for Zero-Shot Voice Conversion With Increased Controllability via Multiple Guidances
- SnipSnap: A Joint Compression Format and Dataflow Co-Optimization Framework for Efficient Sparse LLM Accelerator Design
- CUTE: A Multilingual Dataset for Enhancing Cross-Lingual Knowledge Transfer in Low-Resource Languages
- Are you sure? Measuring models bias in content moderation through uncertainty
- Influence Guided Context Selection for Effective Retrieval-Augmented Generation
- DRES: Fake news detection by dynamic representation and ensemble selection
- Multi-task Pretraining for Enhancing Interpretable L2 Pronunciation Assessment
- Towards Interpretable and Efficient Attention: Compressing All by Contracting a Few
- From Prediction to Understanding: Will AI Foundation Models Transform Brain Science?
- VidCLearn: A Continual Learning Approach for Text-to-Video Generation
- Causal Representation Learning from Multimodal Clinical Records under Non-Random Modality Missingness
- Evolution of Concepts in Language Model Pre-Training
- Domain-Adaptive Pre-Training for Arabic Aspect-Based Sentiment Analysis: A Comparative Study of Domain Adaptation and Fine-Tuning Strategies
- "Digital Camouflage": The LLVM Challenge in LLM-Based Malware Detection
- Unlocking Hidden Potential in Point Cloud Networks with Attention-Guided Grouping-Feature Coordination
- Self-Supervised Learning of Graph Representations for Network Intrusion Detection
- Learn to Rank Risky Investors: A Case Study of Predicting Retail Traders' Behaviour and Profitability
- A Novel Differential Feature Learning for Effective Hallucination Detection and Classification
- Less Is More? Examining Fairness in Pruned Large Language Models for Summarising Opinions
- Mental Multi-class Classification on Social Media: Benchmarking Transformer Architectures against LSTM Models
- CommonForms: A Large, Diverse Dataset for Form Field Detection
- When Big Models Train Small Ones: Label-Free Model Parity Alignment for Efficient Visual Question Answering using Small VLMs
- On the de-duplication of the Lakh MIDI dataset
- KuBERT: Central Kurdish BERT Model and Its Application for Sentiment Analysis
- Leveraging Multilingual Training for Authorship Representation: Enhancing Generalization across Languages and Domains
- AHA -- Predicting What Matters Next: Online Highlight Detection Without Looking Ahead
- Causality-Induced Positional Encoding for Transformer-Based Representation Learning of Non-Sequential Features
- Redefining Experts: Interpretable Decomposition of Language Models for Toxicity Mitigation
- A Multi-Level Benchmark for Causal Language Understanding in Social Media Discourse
- MPCG: Multi-Round Persona-Conditioned Generation for Modeling the Evolution of Misinformation with LLMs
- The Role of Vocabularies in Learning Sparse Representations for Ranking
- LLM-Guided Co-Training for Text Classification
- Enhancing OHLC Data with Timing Features: A Machine Learning Evaluation
- DAFTED: Decoupled Asymmetric Fusion of Tabular and Echocardiographic Data for Cardiac Hypertension Diagnosis
- BEFT: Bias-Efficient Fine-Tuning of Language Models
- RACap: Relation-Aware Prompting for Lightweight Retrieval-Augmented Image Captioning
- ENSAM: an efficient foundation model for interactive segmentation of 3D medical images
- The Curious Case of Visual Grounding: Different Effects for Speech- and Text-based Language Encoders
- RMT-KD: Random Matrix Theoretic Causal Knowledge Distillation
- KITE: Kernelized and Information Theoretic Exemplars for In-Context Learning
- Psychology's Questionable Research Fundamentals (QRFs): Key problems in quantitative psychology and psychological measurement beyond Questionable Research Practices (QRPs)
- Bridging Graph and State-Space Modeling for Intensive Care Unit Length of Stay Prediction
- Adversarially Robust Assembly Language Model for Packed Executables Detection
- How do Language Models Generate Slang: A Systematic Comparison between Human and Machine-Generated Slang Usages
- Using machine learning to automate data annotation in corpus linguistics
- Dynamic Classifier-Free Diffusion Guidance via Online Feedback
- Chunk Knowledge Generation Model for Enhanced Information Retrieval: A Multi-task Learning Approach
- Optimizing Product Deduplication in E-Commerce with Multimodal Embeddings
- Latent Zoning Network: A Unified Principle for Generative Modeling, Representation Learning, and Classification
- Predicting the descent into extremism and terrorism
- HARE: an entity and relation centric evaluation framework for histopathology reports
- Evaluating the Effectiveness and Scalability of LLM-Based Data Augmentation for Retrieval
- Defining and Monitoring Complex Robot Activities via LLMs and Symbolic Reasoning
- Efficient Extractive Text Summarization for Online News Articles Using Machine Learning
- Efficient Multimodal Dataset Distillation via Generative Models
- BiRQ: Bi-Level Self-Labeling Random Quantization for Self-Supervised Speech Recognition
- SPH-Net: A Co-Attention Hybrid Model for Accurate Stock Price Prediction
- Efficient and Versatile Model for Multilingual Information Retrieval of Islamic Text: Development and Deployment in Real-World Scenarios
- Real, Fake, or Manipulated? Detecting Machine-Influenced Text
- Understand Before You Generate: Self-Guided Training for Autoregressive Image Generation
- SMARTER: A Data-efficient Framework to Improve Toxicity Detection with Explanation via Self-augmenting Large Language Models
- OmniMRI: A Unified Vision--Language Foundation Model for Generalist MRI Interpretation
- LLM-OREF: An Open Relation Extraction Framework Based on Large Language Models
- Improving French Synthetic Speech Quality via SSML Prosody Control
- Efficient Zero-Shot Long Document Classification by Reducing Context Through Sentence Ranking
- Attention Beyond Neighborhoods: Reviving Transformer for Graph Clustering
- Copycat vs. Original: Multi-modal Pretraining and Variable Importance in Box-office Prediction
- SINAI at eRisk@CLEF 2022: Approaching Early Detection of Gambling and Eating Disorders with Natural Language Processing
- Generative AI Meets Wireless Sensing: Towards Wireless Foundation Model
- Understanding the Thinking Process of Reasoning Models: A Perspective from Schoenfeld's Episode Theory
- MahaParaphrase: A Marathi Paraphrase Detection Corpus and BERT-based Models
- DeCoP: Enhancing Self-Supervised Time Series Representation with Dependency Controlled Pre-training
- Automating Modelica Module Generation Using Large Language Models: A Case Study on Building Control Description Language
- Leveraging IndoBERT and DistilBERT for Indonesian Emotion Classification in E-Commerce Reviews
- VisMoDAl: Visual Analytics for Evaluating and Improving Corruption Robustness of Vision-Language Models
- TriSPrompt: A Hierarchical Soft Prompt Model for Multimodal Rumor Detection with Incomplete Modalities
- MaiBERT: A Pre-training Corpus and Language Model for Low-Resourced Maithili Language
- Evaluating Multimodal Large Language Models on Spoken Sarcasm Understanding
- Adaptive LoRA Experts Allocation and Selection for Federated Fine-Tuning
- DeKeyNLU: Enhancing Natural Language to SQL Generation through Task Decomposition and Keyword Extraction
- Learning to Retrieve for Environmental Knowledge Discovery: An Augmentation-Adaptive Self-Supervised Learning Framework
- Patent Language Model Pretraining with ModernBERT
- Deep learning and abstractive summarisation for radiological reports: an empirical study for adapting the PEGASUS models' family with scarce data
- PCCL: Photonic circuit-switched collective communication for distributed ML
- Hierarchical Self-Attention: Generalizing Neural Attention Mechanics to Multi-Scale Problems
- Fast and Fluent Diffusion Language Models via Convolutional Decoding and Rejective Fine-tuning
- Not What the Doctor Ordered: Surveying LLM-based De-identification and Quantifying Clinical Information Loss
- Synthetic bootstrapped pretraining
- A Framework for Generating Artificial Datasets to Validate Absolute and Relative Position Concepts
- Exploring the Capabilities of LLM Encoders for Image-Text Retrieval in Chest X-rays
- MOCHA: Multi-modal Objects-aware Cross-arcHitecture Alignment
- Long-context Reference-based MT Quality Estimation
- MetricNet: Recovering Metric Scale in Generative Navigation Policies
- Combining Evidence and Reasoning for Biomedical Fact-Checking
- Detecting Struggling Student Programmers using Proficiency Taxonomies
- Speech-Based Cognitive Screening: A Systematic Evaluation of LLM Adaptation Strategies
- Masked Feature Modeling Enhances Adaptive Segmentation
- Self Identity Mapping
- Cross-modal Full-mode Fine-grained Alignment for Text-to-Image Person Retrieval
- ST-LINK: Spatially-Aware Large Language Models for Spatio-Temporal Forecasting
- State Space Models over Directed Graphs
- Re-purposing SAM into Efficient Visual Projectors for MLLM-Based Referring Image Segmentation
- Deep Lookup Network
- Is GPT-4o mini Blinded by its Own Safety Filters? Exposing the Multimodal-to-Unimodal Bottleneck in Hate Speech Detection
- Deep Learning-Driven Peptide Classification in Biological Nanopores
- CS-FLEURS: A Massively Multilingual and Code-Switched Speech Dataset
- An LLM-based multi-agent framework for agile effort estimation
- Translate, then Detect: Leveraging Machine Translation for Cross-Lingual Toxicity Classification
- Overview of the TREC 2024 NeuCLIR Track
- CodeLSI: Leveraging Foundation Models for Automated Code Generation with Low-Rank Optimization and Domain-Specific Instruction Tuning
- Teaching According to Talents! Instruction Tuning LLMs with Competence-Aware Curriculum Learning
- Risk Assessment and Security Analysis of Large Language Models
- Annotating Satellite Images of Forests with Keywords from a Specialized Corpus in the Context of Change Detection
- Event Causality Identification with Synthetic Control
- Image Realness Assessment and Localization with Multimodal Features
- RIS-FUSION: Rethinking Text-Driven Infrared and Visible Image Fusion from the Perspective of Referring Image Segmentation
- From Embeddings to Equations: Genetic-Programming Surrogates for Interpretable Transformer Classification
- The Few-shot Dilemma: Over-prompting Large Language Models
- Conan-Embedding-v2: Training an LLM from Scratch for Text Embeddings
- Membership Inference Attack against Large Language Model-based Recommendation Systems: A New Distillation-based Paradigm
- Harnessing the Power of AI in Qualitative Research: Role Assignment, Engagement, and User Perceptions of AI-Generated Follow-Up Questions in Semi-Structured Interviews
- Towards Inclusive Toxic Content Moderation: Addressing Vulnerabilities to Adversarial Attacks in Toxicity Classifiers Tackling LLM-generated Content
- Positional Encoding via Token-Aware Phase Attention
- LEAF: Knowledge Distillation of Text Embedding Models with Teacher-Aligned Representations
- Multimodal Hate Detection Using Dual-Stream Graph Neural Networks
- Empowering LLMs with Parameterized Skills for Adversarial Long-Horizon Planning
- TFANet: Three-Stage Image-Text Feature Alignment Network for Robust Referring Image Segmentation
- Module-Aware Parameter-Efficient Machine Unlearning on Transformers
- Retrieve-and-Verify: A Table Context Selection Framework for Accurate Column Annotations
- MMCIG: Multimodal Cover Image Generation for Text-only Documents and Its Dataset Construction via Pseudo-labeling
- FusionMAE: large-scale pretrained model to optimize and simplify diagnostic and control of fusion plasma
- Match Chat: Real Time Generative AI and Generative Computing for Tennis
- Automated Generation of Research Workflows from Academic Papers: A Full-text Mining Framework
- Contextualized Representation Learning for Effective Human-Object Interaction Detection
- Active Domain Knowledge Acquisition with 100-Dollar Budget: Enhancing LLMs via Cost-Efficient, Expert-Involved Interaction in Sensitive Domains
- Quantifying Language Disparities in Multilingual Large Language Models
- Spatio-Temporal Pruning for Compressed Spiking Large Language Models
- Natural Language Satisfiability: Exploring the Problem Distribution and Evaluating Transformer-based Language Models
- Post-Hoc Split-Point Self-Consistency Verification for Efficient, Unified Quantification of Aleatoric and Epistemic Uncertainty in Deep Learning
- Data Augmentation for Maltese NLP using Transliterated and Machine Translated Arabic Data
- A comparison of pipelines for the translation of a low resource language based on transformers
- Multi Anatomy X-Ray Foundation Model
- SENTRA: Selected-Next-Token Transformer for LLM Text Detection
- GhostNetV3-Small: A Tailored Architecture and Comparative Study of Distillation Strategies for Tiny Images
- MTEB-NL and E5-NL: Embedding Benchmark and Models for Dutch
- Dynamic Relational Priming Improves Transformer in Multivariate Time Series
- More Similar than Dissimilar: Modeling Annotators for Cross-Corpus Speech Emotion Recognition
- Interpreting Public Sentiment in Diplomacy Events: A Counterfactual Analysis Framework Using Large Language Models
- GTA: Supervised-Guided Reinforcement Learning for Text Classification with Large Language Models
- MusicSwarm: Biologically Inspired Intelligence for Music Composition
- LEGO: Spatial Accelerator Generation and Optimization for Tensor Applications
- RAM++: Robust Representation Learning via Adaptive Mask for All-in-One Image Restoration
- Query-Focused Extractive Summarization for Sentiment Explanation
- Uncertainty in Authorship: Why Perfect AI Detection Is Mathematically Impossible
- EgoMem: Lifelong Memory Agent for Full-duplex Omnimodal Models
- Do It Yourself (DIY): Modifying Images for Poems in a Zero-Shot Setting Using Weighted Prompt Manipulation
- Enhancing Knowledge Tracing through Leakage-Free and Recency-Aware Embeddings
- Biomedical Hypothesis Explainability with Graph-Based Context Retrieval
- AssemMate: Graph-Based LLM for Robotic Assembly Assistance
- Dynamic Span Interaction and Graph-Aware Memory for Entity-Level Sentiment Classification
- ProtoEHR: Hierarchical Prototype Learning for EHR-based Healthcare Predictions
- Proximal Vision Transformer: Enhancing Feature Representation through Two-Stage Manifold Geometry
- Bhaasha, Bhasa, Zaban: A Survey for Low-Resourced Languages in South Asia -- Current Stage and Challenges
- Graph-Enhanced Retrieval-Augmented Question Answering for E-Commerce Customer Support
- HiChunk: Evaluating and Enhancing Retrieval-Augmented Generation with Hierarchical Chunking
- Know What You Don't Know: Selective Prediction for Early Exit DNNs
- Unsupervised Candidate Ranking for Lexical Substitution via Holistic Sentence Semantics
- LOKI: Proactively Discovering Online Scam Websites by Mining Toxic Search Queries
- Is 'Hope' a person or an idea? A pilot benchmark for NER: comparing traditional NLP tools and large language models on ambiguous entities
- SAQ: Pushing the Limits of Vector Quantization through Code Adjustment and Dimension Segmentation
- Designing LLMs for cultural sensitivity: Evidence from English-Japanese translation
- Disentangling Content from Style to Overcome Shortcut Learning: A Hybrid Generative-Discriminative Learning Framework
- Efficient Hate Speech Detection: Evaluating 38 Models from Traditional Methods to Transformers
- FuseCodec: Semantic-Contextual Fusion and Supervision for Neural Codecs
- Quantifying Compositionality of Classic and State-of-the-Art Embeddings
- Transformer Enhanced Relation Classification: A Comparative Analysis of Contextuality, Data Efficiency and Sequence Complexity
- !MSA at AraHealthQA 2025 Shared Task: Enhancing LLM Performance for Arabic Clinical Question Answering through Prompt Engineering and Ensemble Learning
- RanAT4BIE: Random Adversarial Training for Biomedical Information Extraction
- Auto-Slides: An Interactive Multi-Agent System for Creating and Customizing Research Presentations
- BERT4beam: Large AI Model Enabled Generalized Beamforming Optimization
- Human Activity Recognition Based on Electrocardiogram Data Only
- We Argue to Agree: Towards Personality-Driven Argumentation-Based Negotiation Dialogue Systems for Tourism
- UserTrace: User-Level Requirements Generation and Traceability Recovery from Software Project Repositories
- AI-Generated Content in Cross-Domain Applications: Research Trends, Challenges and Propositions
- Feature Space Topology Control via Hopkins Loss
- Improving Table Understanding with LLMs and Entity-Oriented Search
- ReFineG: Synergizing Small Supervised Models and LLMs for Low-Resource Grounded Multimodal NER
- Refining Syntactic Distinctions Using Decision Trees: A Paper on Postnominal 'That' in Complement vs. Relative Clauses
- Quantifier Scope Interpretation in Language Learners and LLMs
- Towards Automated Error Discovery: A Study in Conversational AI
- Why Bonds Fail Differently? Explainable Multimodal Learning for Multi-Class Default Prediction
- Long Context Automated Essay Scoring with Language Models
- Reasoning Under Uncertainty: Exploring Probabilistic Reasoning Capabilities of LLMs
- PolyTruth: Multilingual Disinformation Detection using Transformer-Based Language Models
- SegSLR: Promptable Video Segmentation for Isolated Sign Language Recognition
- A Survey on Retrieval And Structuring Augmented Generation with Large Language Models
- An Efficient Dual-Line Decoder Network with Multi-Scale Convolutional Attention for Multi-organ Segmentation
- Safety and Security Analysis of Large Language Models: Benchmarking Risk Profile and Harm Potential
- Planning for Success: Exploring LLM Long-term Planning Capabilities in Table Understanding
- DeAR: Dual-Stage Document Reranking with Reasoning Agents via LLM Distillation
- WebSight: A Vision-First Architecture for Robust Web Agents
- Improving Audio Event Recognition with Consistency Regularization
- Opening the Black Box: Interpretable LLMs via Semantic Resonance Architecture
- Decoding Alignment: A Critical Survey of LLM Development Initiatives through Value-setting and Data-centric Lens
- Explaining Black-box Language Models with Knowledge Probing Systems: A Post-hoc Explanation Perspective
- Multi-Intent Recognition in Dialogue Understanding: A Comparison Between Smaller Open-Source LLMs
- Development of Automated Software Design Document Review Methods Using Large Language Models
- Byte by Byte: Unmasking Browser Fingerprinting at the Function Level Using V8 Bytecode Transformers
- Can we use automated approaches to measure the quality of online political discussion? How to (not) measure interactivity, diversity, rationality, and incivility in online comments to the news
- DyKen-Hyena: Dynamic Kernel Generation via Cross-Modal Attention for Multimodal Intent Recognition
- Investigating red packet fraud in Android applications: Insights from user reviews
- Scalable Training for Vector-Quantized Networks with 100% Codebook Utilization
- DOSA: Differentiable Model-Based One-Loop Search for DNN Accelerators
- ZapGPT: Free-form Language Prompting for Simulated Cellular Control
- Conditioning on PDE Parameters to Generalise Deep Learning Emulation of Stochastic and Chaotic Dynamics
- Fluent but Unfeeling: The Emotional Blind Spots of Language Models
- Personality-Enhanced Social Recommendations in SAMI: Exploring the Role of Personality Detection in Matchmaking
- Hierarchical Bracketing Encodings Work for Dependency Graphs
- Boosting Embodied AI Agents through Perception-Generation Disaggregation and Asynchronous Pipeline Execution
- CoPE: A Lightweight Complex Positional Encoding
- TextOnly: A Unified Function Portal for Text-Related Functions on Smartphones
- Boosting Data Utilization for Multilingual Dense Retrieval
- Cross-Domain Evaluation of Transformer-Based Vulnerability Detection on Open & Industry Data
- Target-oriented Multimodal Sentiment Classification with Counterfactual-enhanced Debiasing
- ViRanker: A BGE-M3 & Blockwise Parallel Transformer Cross-Encoder for Vietnamese Reranking
- Character-Level Perturbations Disrupt LLM Watermarks
- Towards Confidential and Efficient LLM Inference with Dual Privacy Protection
- Enhancing 3D Medical Image Understanding with Pretraining Aided by 2D Multimodal Large Language Models
- Quality Assessment of Tabular Data using Large Language Models and Code Generation
- LLMs as Agentic Cooperative Players in Multiplayer UNO
- IMDMR: An Intelligent Multi-Dimensional Memory Retrieval System for Enhanced Conversational AI
- CoSwin: Convolution Enhanced Hierarchical Shifted Window Attention For Small-Scale Vision
- Documents Are People and Words Are Items: A Psychometric Approach to Textual Data with Contextual Embeddings
- Diffusion-Based Action Recognition Generalizes to Untrained Domains
- Machine learning the effects of many quantum measurements
- Advancing Conversational AI with Shona Slang: A Dataset and Hybrid Model for Digital Inclusion
- BcQLM: Efficient Vision-Language Understanding with Distilled Q-Gated Cross-Modal Fusion
- Tokenizing Loops of Antibodies
- Discrimination by LLMs: Cross-lingual Bias Assessment and Mitigation in Decision-Making and Summarisation
- OTESGN: Optimal Transport-Enhanced Syntactic-Semantic Graph Networks for Aspect-Based Sentiment Analysis
- QFrCoLA: a Quebec-French Corpus of Linguistic Acceptability Judgments
- TriagerX: Dual Transformers for Bug Triaging Tasks with Content and Interaction Based Rankings
- Adversarial Attacks Against Automated Fact-Checking: A Survey
- Retrieval-Augmented VLMs for Multimodal Melanoma Diagnosis
- DiTTO-LLM: Framework for Discovering Topic-based Technology Opportunities via Large Language Model
- Foundation Models for Autonomous Driving Perception: A Survey Through Core Capabilities
- Algorithmic Tradeoffs, Applied NLP, and the State-of-the-Art Fallacy
- ReProCon: Scalable and Resource-Efficient Few-Shot Biomedical Named Entity Recognition
- Label Smoothing++: Enhanced Label Regularization for Training Neural Networks
- Un cadre paraconsistant pour l'évaluation de similarité dans les bases de connaissances
- SurgLaVi: Large-Scale Hierarchical Dataset for Surgical Vision-Language Representation Learning
- Bias after Prompting: Persistent Discrimination in Large Language Models
- Attribute-based Object Grounding and Robot Grasp Detection with Spatial Reasoning
- Bringing Multi-Modal Multi-Task Federated Foundation Models to Education Domain: Prospects and Challenges
- Towards Generalized Routing: Model and Agent Orchestration for Adaptive and Efficient Inference
- Query Expansion in the Age of Pre-trained and Large Language Models: A Comprehensive Survey
- A Survey of Long-Document Retrieval in the PLM and LLM Era
- ELEC: Efficient Large Language Model-Empowered Click-Through Rate Prediction
- BALI: Enhancing Biomedical Language Representations through Knowledge Graph and Language Model Alignment
- FLeW: Facet-Level and Adaptive Weighted Representation Learning of Scientific Documents
- AIxcellent Vibes at GermEval 2025 Shared Task on Candy Speech Detection: Improving Model Performance by Span-Level Training
- PersonaFuse: A Personality Activation-Driven Framework for Enhancing Human-LLM Interactions
- Assess and Prompt: A Generative RL Framework for Improving Engagement in Online Mental Health Communities
- Interpreting the Effects of Quantization on LLMs
- ACE and Diverse Generalization via Selective Disagreement
- M-BRe: Discovering Training Samples for Relation Extraction from Unlabeled Texts with Large Language Models
- How Good are LLM-based Rerankers? An Empirical Analysis of State-of-the-Art Reranking Models
- Transferable Direct Prompt Injection via Activation-Guided MCMC Sampling
- NOWJ@COLIEE 2025: A Multi-stage Framework Integrating Embedding Models and Large Language Models for Legal Retrieval and Entailment
- In-Context Learning Enhanced Credibility Transformer
- CP-Model-Zoo: A Natural Language Query System for Constraint Programming Models
- From Detection to Mitigation: Addressing Gender Bias in Chinese Texts via Efficient Tuning and Voting-Based Rebalancing
- Attention Maps in 3D Shape Classification for Dental Stage Estimation with Class Node Graph Attention Networks
- Towards EnergyGPT: A Large Language Model Specialized for the Energy Sector
- UniSearch: Rethinking Search System with a Unified Generative Architecture
- MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security
- Will Annotators Disagree? Identifying Subjectivity in Value-Laden Arguments
- AI-driven Remote Facial Skin Hydration and TEWL Assessment from Selfie Images: A Systematic Solution
- D-HUMOR: Dark Humor Understanding via Multimodal Open-ended Reasoning -- A Benchmark Dataset and Method
- mmBERT: A Modern Multilingual Encoder with Annealed Language Learning
- Investigating Location-Regularised Self-Supervised Feature Learning for Seafloor Visual Imagery
- PLRV-O: Advancing Differentially Private Deep Learning via Privacy Loss Random Variable Optimization
- PathoHR: Hierarchical Reasoning for Vision-Language Models in Pathology
- Benchmarking Information Retrieval Models on Complex Retrieval Tasks
- Language Bias in Information Retrieval: The Nature of the Beast and Mitigation Methods
- An Explainable Deep Neural Network with Frequency-Aware Channel and Spatial Refinement for Flood Prediction in Sustainable Cities
- Semantic-Aware Edge Intelligence for UAV Handover in 6G Networks
- ALPHA: LLM-Enabled Active Learning for Human-Free Network Anomaly Detection
- Augmented Fine-Tuned LLMs for Enhanced Recruitment Automation
- Augmenting Human-Centered Racial Covenant Detection and Georeferencing with Plug-and-Play NLP Pipelines
- Hyperbolic Large Language Models
- Automating API Documentation with LLMs: A BERTopic Approach
- InterAct: A Large-Scale Dataset of Dynamic, Expressive and Interactive Activities between Two People in Daily Scenarios
- From Joy to Fear: A Benchmark of Emotion Estimation in Pop Song Lyrics
- New Insights into Optimal Alignment of Acoustic and Linguistic Representations for Knowledge Transfer in ASR
- QCSE: A Pretrained Quantum Context-Sensitive Word Embedding for Natural Language Processing
- Few-Shot Query Intent Detection via Relation-Aware Prompt Learning
- A Survey of the State-of-the-Art in Conversational Question Answering Systems
- Learning to Route: Per-Sample Adaptive Routing for Multimodal Multitask Prediction
- Self-supervised Learning for Hyperspectral Images of Trees
- On the Contribution of Lexical Features to Speech Emotion Recognition
- Differential Robustness in Transformer Language Models: Empirical Evaluation Under Adversarial Text Attacks
- CURE: Controlled Unlearning for Robust Embeddings -- Mitigating Conceptual Shortcuts in Pre-Trained Language Models
- Scaling Performance of Large Language Model Pretraining
- Recomposer: Event-roll-guided generative audio editing
- HoPE: Hyperbolic Rotary Positional Encoding for Stable Long-Range Dependency Modeling in Large Language Models
- MEAN-RIR: Multi-Modal Environment-Aware Network for Robust Room Impulse Response Estimation
- Masked Diffusion Language Models with Frequency-Informed Training
- Towards Efficient Pixel Labeling for Industrial Anomaly Detection and Localization
- Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts
- Do Large Language Models Need Intent? Revisiting Response Generation Strategies for Service Assistant
- A Study of Large Language Models for Patient Information Extraction: Model Architecture, Fine-Tuning Strategy, and Multi-task Instruction Tuning
- Legislators’ sentiment analysis supervised by legislators
- RINSER: Accurate API Prediction Using Masked Language Models
- PLanTS: Periodicity-aware Latent-state Representation Learning for Multivariate Time Series
- Rethinking Reasoning in LLMs: Neuro-Symbolic Local RetoMaton Beyond ICL and CoT
- Towards Open World Detection: A Survey
- Optimizing Small Transformer-Based Language Models for Multi-Label Sentiment Analysis in Short Texts
- Classification of kinetic-related injury in hospital triage data using NLP
- Comparative Analysis of Transformer Models in Disaster Tweet Classification for Public Safety
- Action Chunking with Transformers for Image-Based Spacecraft Guidance and Control
- Manipulating Transformer-Based Models: Controllability, Steerability, and Robust Interventions
- Contextualized Token Discrimination for Speech Search Query Correction
- Parking Availability Prediction via Fusing Multi-Source Data with A Self-Supervised Learning Enhanced Spatio-Temporal Inverted Transformer
- Enhancing Technical Documents Retrieval for RAG
- A RoBERTa-Based Functional Syntax Annotation Model for Chinese Texts
- Rethinking the long-range dependency in Mamba/SSM and transformer models
- Sailing Towards Zero-Shot State Estimation using Foundation Models Combined with a UKF
- Explicit and Implicit Data Augmentation for Social Event Detection
- The changing role of cited papers over time: An analysis of highly cited papers based on a large full-text dataset
- SafeSpace: An Integrated Web Application for Digital Safety and Emotional Well-being
- Improving Narrative Classification and Explanation via Fine Tuned Language Models
- SMooGPT: Stylized Motion Generation using Large Language Models
- RTQA : Recursive Thinking for Complex Temporal Knowledge Graph Question Answering with Large Language Models
- Anti-establishment sentiment on TikTok: Implications for understanding influence(rs) and expertise on social media
- Beyond Interpretability: Exploring the Comprehensibility of Adaptive Video Streaming through Large Language Models
- Optimizing Frequent Checkpointing via Low-Cost Differential for Distributed Training Systems
- Code Like Humans: A Multi-Agent Solution for Medical Coding
- Spoken in Jest, Detected in Earnest: A Systematic Review of Sarcasm Recognition -- Multimodal Fusion, Challenges, and Future Prospects
- Estudio de la eficiencia en la escalabilidad de GPUs para el entrenamiento de Inteligencia Artificial
- Efficient Item ID Generation for Large-Scale LLM-based Recommendation
- Differentiable Entropy Regularization: A Complexity-Aware Approach for Neural Optimization
- MLSD: A Novel Few-Shot Learning Approach to Enhance Cross-Target and Cross-Domain Stance Detection
- OPERA: A Reinforcement Learning--Enhanced Orchestrated Planner-Executor Architecture for Reasoning-Oriented Multi-Hop Retrieval
- Cetvel: A Unified Benchmark for Evaluating Language Understanding, Generation and Cultural Capacity of LLMs for Turkish
- NoteBar: An AI-Assisted Note-Taking System for Personal Knowledge Management
- Scalable and Loosely-Coupled Multimodal Deep Learning for Breast Cancer Subtyping
- Learning Mechanism Underlying NLP Pre-Training and Fine-Tuning
- RecBase: Generative Foundation Model Pretraining for Zero-Shot Recommendation
- VQualA 2025 Challenge on Engagement Prediction for Short Videos: Methods and Results
- Advancing Minority Stress Detection with Transformers: Insights from the Social Media Datasets
- GLARE: Agentic Reasoning for Legal Judgment Prediction
- Training LLMs to be Better Text Embedders through Bidirectional Reconstruction
- DPQuant: Efficient and Differentially-Private Model Training via Dynamic Quantization Scheduling
- AIVA: An AI-based Virtual Companion for Emotion-aware Interaction
- Hierarchical Section Matching Prediction (HSMP) BERT for Fine-Grained Extraction of Structured Data from Hebrew Free-Text Radiology Reports in Crohn's Disease
- Sparse and Dense Retrievers Learn Better Together: Joint Sparse-Dense Optimization for Text-Image Retrieval
- OPRA-Vis: Visual Analytics System to Assist Organization-Public Relationship Assessment with Large Language Models
- Decoding the Rule Book: Extracting Hidden Moderation Criteria from Reddit Communities
- IDEAlign: Comparing Large Language Models to Human Experts in Open-ended Interpretive Annotations
- Scaling behavior of large language models in emotional safety classification across sizes and tasks
- Predicting Movie Success with Multi-Task Learning: A Hybrid Framework Combining GPT-Based Sentiment Analysis and SIR Propagation
- DrDiff: Dynamic Routing Diffusion with Hierarchical Attention for Breaking the Efficiency-Quality Trade-off
- Generative AI for Crystal Structures: A Review
- Probabilistic Pretraining for Neural Regression
- VASSO: Variance Suppression for Sharpness-Aware Minimization
- Vision encoders should be image size agnostic and task driven
- Exploiting Information Redundancy in Attention Maps for Extreme Quantization of Vision Transformers
- A Multimodal-Multitask Framework with Cross-modal Relation and Hierarchical Interactive Attention for Semantic Comprehension
- LLMs that Understand Processes: Instruction-tuning for Semantics-Aware Process Mining
- StructCoh: Structured Contrastive Learning for Context-Aware Text Semantic Matching
- Structure-aware Contrastive Learning for Diagram Understanding of Multimodal Models
- Weakly Supervised Medical Entity Extraction and Linking for Chief Complaints
- An Ensemble Classification Approach in A Multi-Layered Large Language Model Framework for Disease Prediction
- CellPainTR: Generalizable Representation Learning for Cross-Dataset Cell Painting Analysis
- Computational Social Science and Critical Studies of Education and Technology: An Improbable Combination?
- Abex-rat: Synergizing Abstractive Augmentation and Adversarial Training for Classification of Occupational Accident Reports
- chDzDT: Word-level morphology-aware language model for Algerian social media text
- A Narrative-Driven Computational Framework for Clinician Burnout Surveillance
- Entropy-Driven Curriculum for Multi-Task Training in Human Mobility Prediction
- Cross-Modal Prototype Augmentation and Dual-Grained Prompt Learning for Social Media Popularity Prediction
- Text Takes Over: A Study of Modality Bias in Multimodal Intent Detection
- The AudioMOS Challenge 2025
- GradeSQL: Test-Time Inference with Outcome Reward Models for Text-to-SQL Generation from Large Language Models
- DaMoC: Efficiently Selecting the Optimal Large Language Model for Fine-tuning Domain Tasks Based on Data and Model Compression
- TransGAT: Transformer-Based Graph Neural Networks for Multi-Dimensional Automated Essay Scoring
- Re3: Learning to Balance Relevance & Recency for Temporal Information Retrieval
- Bridging Thoughts and Words: Graph-Based Intent-Semantic Joint Learning for Fake News Detection
- AttnBoost: Retail Supply Chain Sales Insights via Gradient Boosting Perspective
- Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation
- A Multi-target Bayesian Transformer Framework for Predicting Cardiovascular Disease Biomarkers during Pandemics
- WATCHED: A Web AI Agent Tool for Combating Hate Speech by Expanding Data
- Serialized Output Prompting for Large Language Model-based Multi-Talker Speech Recognition
- Efficient and High-Accuracy Secure Two-Party Protocols for a Class of Functions with Real-number Inputs
- A case study of forensic psychiatry experts' reports analysis through large language models
- MEPT: Mixture of Expert Prompt Tuning as a Manifold Mapper
- LLM-Guided Semantic Relational Reasoning for Multimodal Intent Recognition
- Testing the assumptions about the geometry of sentence embedding spaces: the cosine measure need not apply
- Supervised In-Context Fine-Tuning for Generative Sequence Labeling
- TinyMusician: On-Device Music Generation with Knowledge Distillation and Mixed Precision Quantization
- Exploring Over-stationarization in Deep Learning-based Bus/Tram Arrival Time Prediction: Analysis and Non-stationary Effect Recovery
- Efficient Graph Understanding with LLMs via Structured Context Injection
- LLM Encoder vs. Decoder: Robust Detection of Chinese AI-Generated Text with LoRA
- CascadeFormer: A Family of Two-stage Cascading Transformers for Skeleton-based Human Action Recognition
- Speaker-Conditioned Phrase Break Prediction for Text-to-Speech with Phoneme-Level Pre-trained Language Model
- Political Ideology Shifts in Large Language Models
- Predicting Multi-Type Talented Students in Secondary School Using Semi-Supervised Machine Learning
- Neural Models and Language Model Prompting for the Multidimensional Evaluation of Open-Ended Conversations
- TMT: A Simple Way to Translate Topic Models Using Dictionaries
- LegalChainReasoner: A Legal Chain-guided Framework for Criminal Judicial Opinion Generation
- HePGA: A Heterogeneous Processing-in-Memory based GNN Training Accelerator
- A Deep Learning Framework for Joint Channel Acquisition and Communication Optimization in Movable Antenna Systems
- Memory Limitations of Prompt Tuning in Transformers
- Backdoor Samples Detection Based on Perturbation Discrepancy Consistency in Pre-trained Language Models
- Adaptive Point-Prompt Tuning: Fine-Tuning Heterogeneous Foundation Models for 3D Point Cloud Analysis
- SABR: A Stable Adaptive Bitrate Framework Using Behavior Cloning Pretraining and Reinforcement Learning Fine-Tuning
- Exploring Reasoning-Infused Text Embedding with Large Language Models for Zero-Shot Dense Retrieval
- Configuration Bugs Classification using LLMs and Encoders
- Not All Parameters Are Created Equal: Smart Isolation Boosts Fine-Tuning Performance
- DriveQA: Passing the Driving Knowledge Test
- VoCap: Video Object Captioning and Segmentation from Any Prompt
- Reasoning-Intensive Regression
- EPIC: Generative AI Platform for Accelerating HPC Operational Data Analytics
- Introducing LCOAI: A Standardized Economic Metric for Evaluating AI Deployment Costs
- QZhou-Embedding Technical Report
- A Survey on Current Trends and Recent Advances in Text Anonymization
- L3Cube-MahaSTS: A Marathi Sentence Similarity Dataset and Models
- Spiking Decision Transformers: Local Plasticity, Phase-Coding, and Dendritic Routing for Low-Power Sequence Control
- HSFN: Hierarchical Selection for Fake News Detection building Heterogeneous Ensemble
- Diffusion-based Multi-modal Synergy Interest Network for Click-through Rate Prediction
- MM-SeR: Multimodal Self-Refinement for Lightweight Image Captioning
- Evaluating Recabilities of Foundation Models: A Multi-Domain, Multi-Dataset Benchmark
- Quantum-Enhanced Natural Language Generation: A Multi-Model Framework with Hybrid Quantum-Classical Architectures
- Scalable Equilibrium Propagation via Intermediate Error Signals for Deep Convolutional CRNNs
- Guess-and-Learn (G&L): Measuring the Cumulative Error Cost of Cold-Start Adaptation
- Generative AI for Industrial Contour Detection: A Language-Guided Vision System
- AI Compute Architecture and Evolution Trends
- Efficient Code Embeddings from Code Generation Models
- Decoding Memories: An Efficient Pipeline for Self-Consistency Hallucination Detection
- Synthetic CVs To Build and Test Fairness-Aware Hiring Tools
- Strategic Sample Selection for Improved Clean-Label Backdoor Attacks in Text Classification
- GSTBench: A Benchmark Study on the Transferability of Graph Self-Supervised Learning
- Re-Representation in Sentential Relation Extraction with Sequence Routing Algorithm
- Native Logical and Hierarchical Representations with Subspace Embeddings
- LeMat-Traj: A Scalable and Unified Dataset of Materials Trajectories for Atomistic Modeling
- Looking Beyond the Obvious: A Survey on Abstract Concept Recognition for Video Understanding
- Specializing General-purpose LLM Embeddings for Implicit Hate Speech Detection across Datasets
- Can News Predict the Direction of Oil Price Volatility? A Language Model Approach with SHAP Explanations
- EEGDM: Learning EEG Representation with Latent Diffusion Model
- Bridging Minds and Machines: Toward an Integration of AI and Cognitive Science
- Masked Autoencoders for Ultrasound Signals: Robust Representation Learning for Downstream Applications
- SemSR: Semantics aware robust Session-based Recommendations
- KCS: Diversify Multi-hop Question Generation with Knowledge Composition Sampling
- Overview of BioASQ 2025: The Thirteenth BioASQ Challenge on Large-Scale Biomedical Semantic Indexing and Question Answering
- Overview of BioASQ 2024: The twelfth BioASQ challenge on Large-Scale Biomedical Semantic Indexing and Question Answering
- MoTAS: MoE-Guided Feature Selection from TTS-Augmented Speech for Enhanced Multimodal Alzheimer's Early Screening
- CoFormer: Collaborating with Heterogeneous Edge Devices for Scalable Transformer Inference
- LLM Chatbot-Creation Approaches
- Latent Factor Point Processes for Patient Representation in Electronic Health Records
- Generative Annotation for ASR Named Entity Correction
- End-to-End Analysis of Charge Stability Diagrams with Transformers
- OneRec-V2 Technical Report
- Automated Bug Triaging using Instruction-Tuned Large Language Models
- SEAL: Structure and Element Aware Learning to Improve Long Structured Document Retrieval
- Enhancing Health Fact-Checking with LLM-Generated Synthetic Data
- Tree-like Pairwise Interaction Networks
- Turning Tabular Foundation Models into Graph Foundation Models
- Speech Emotion Recognition via Entropy-Aware Score Selection
- MobileCLIP2: Improving Multi-Modal Reinforced Training
- Tokenization Strategies for Low-Resource Agglutinative Languages in Word2Vec: Case Study on Turkish and Finnish
- A Systematic Review on the Generative AI Applications in Human Medical Genomics
- Robustness Assessment and Enhancement of Text Watermarking for Google's SynthID
- Smart Contract Intent Detection with Pre-trained Programming Language Model
- ProMSC-MIS: Prompt-based Multimodal Semantic Communication for Multi-Spectral Image Segmentation
- Forewarned is Forearmed: Pre-Synthesizing Jailbreak-like Instructions to Enhance LLM Safety Guardrail to Potential Attacks
- Pruning Strategies for Backdoor Defense in LLMs
- Uncertainty-Aware Collaborative System of Large and Small Models for Multimodal Sentiment Analysis
- Automatic Pronunciation Error Detection and Correction of the Holy Quran's Learners Using Deep Learning
- Dhati+: Fine-tuned Large Language Models for Arabic Subjectivity Evaluation
- Self-supervised structured object representation learning
- TComQA: Extracting Temporal Commonsense from Text
- The Art of Hide and Seek: Making Pickle-Based Model Supply Chain Poisoning Stealthy Again
- An Effective Strategy for Modeling Score Ordinality and Non-uniform Intervals in Automated Speaking Assessment
- Skill-based Explanations for Serendipitous Course Recommendation
- Bi-LoRA: Efficient Sharpness-Aware Minimization for Fine-Tuning Large-Scale Models
- MobText-SISA: Efficient Machine Unlearning for Mobility Logs with Spatio-Temporal and Natural-Language Data
- Continual Neural Topic Model
- UNIFORM: Unifying Knowledge from Large-scale and Diverse Pre-trained Models
- Color Bind: Exploring Color Perception in Text-to-Image Models
- LegiScout: A Visual Tool for Understanding Complex Legislation
- The Enemy from Within: A Study of Political Delegitimization Discourse in Israeli Political Speech
- Ensemble Debates with Local Large Language Models for AI Alignment
- Breaking the Layer Barrier: Remodeling Private Transformer Inference with Hybrid CKKS and MPC
- Inference Gap in Domain Expertise and Machine Intelligence in Named Entity Recognition: Creation of and Insights from a Substance Use-related Dataset
- Stack Trace-Based Crash Deduplication with Transformer Adaptation
- Reflective Agreement: Combining Self-Mixture of Agents with a Sequence Tagger for Robust Event Extraction
- JEDI-linear: Fast and Efficient Graph Neural Networks for Jet Tagging on FPGAs
- DeepMEL: A Multi-Agent Collaboration Framework for Multimodal Entity Linking
- Enhancing Document VQA Models via Retrieval-Augmented Generation
- Novel Approaches to Artificial Intelligence Development Based on the Nearest Neighbor Method
- The GINN framework: a stochastic QED correspondence for stability and chaos in deep neural networks
- Embedding Font Impression Word Tags Based on Co-occurrence
- Alljoined-1.6M: A Million-Trial EEG-Image Dataset for Evaluating Affordable Brain-Computer Interfaces
- CoBA: Counterbias Text Augmentation for Mitigating Various Spurious Correlations via Semantic Triples
- Beyond Tokens: Enhancing RTL Quality Estimation via Structural Graph Learning
- From Linearity to Non-Linearity: How Masked Autoencoders Capture Spatial Correlations
- Beyond Benchmark: LLMs Evaluation with an Anthropomorphic and Value-oriented Roadmap
- What do language models model? Transformers, automata, and the format of thought
- Granite Embedding R2 Models
- EvoFormer: Learning Dynamic Graph-Level Representations with Structural and Temporal Bias Correction
- Mimicking associative learning of rats via a neuromorphic robot in open field maze using spatial cell models
- How Reliable are LLMs for Reasoning on the Re-ranking task?
- ExBigBang: A Dynamic Approach for Explainable Persona Classification through Contextualized Hybrid Transformer Analysis
- ANO : Faster is Better in Noisy Landscape
- From BERT to LLMs: Comparing and Understanding Chinese Classifier Prediction in Language Models
- Better Language Model-Based Judging Reward Modeling through Scaling Comprehension Boundaries
- Improving End-to-End Training of Retrieval-Augmented Generation Models via Joint Stochastic Approximation
- Learning from Few Samples: A Novel Approach for High-Quality Malcode Generation
- How Quantization Shapes Bias in Large Language Models
- Training Transformers for Mesh-Based Simulations
- VocabTailor: Dynamic Vocabulary Selection for Downstream Tasks in Small Language Models
- Self-Guided Function Calling in Large Language Models via Stepwise Experience Recall
- ReviseMate: Exploring Contextual Support for Digesting STEM Paper Reviews
- KG-EDAS: A Meta-Metric Framework for Evaluating Knowledge Graph Completion Models
- WangchanThaiInstruct: An instruction-following Dataset for Culture-Aware, Multitask, and Multi-domain Evaluation in Thai
- Dream 7B: Diffusion Large Language Models
- An Empirical Study of Knowledge Distillation for Code Understanding Tasks
- HebID: Detecting Social Identities in Hebrew-language Political Text
- Select to Know: An Internal-External Knowledge Self-Selection Framework for Domain-Specific Question Answering
- Frequency-adaptive tensor neural networks for high-dimensional multi-scale problems
- A Robust BERT-Based Deep Learning Model for Automated Cancer Type Extraction from Unstructured Pathology Reports
- Sensing, Social, and Motion Intelligence in Embodied Navigation: A Comprehensive Survey
- Multilingual Datasets for Custom Input Extraction and Explanation Requests Parsing in Conversational XAI Systems
- Quantization Meets dLLMs: A Systematic Study of Post-training Quantization for Diffusion LLMs
- Synthetic Adaptive Guided Embeddings (SAGE): A Novel Knowledge Distillation Method
- Continuous sentiment scores for literary and multilingual contexts
- PGF-Net: A Progressive Gated-Fusion Framework for Efficient Multimodal Sentiment Analysis
- SATURN: Autoregressive Image Generation Guided by Scene Graphs
- DGenCTR: Towards a Universal Generative Paradigm for Click-Through Rate Prediction via Discrete Diffusion
- Beyond Semantic Similarity: Reducing Unnecessary API Calls via Behavior-Aligned Retriever
- Repeating Words for Video-Language Retrieval with Coarse-to-Fine Objectives
- Tokens with Meaning: A Hybrid Tokenization Approach for NLP
- Two Birds with One Stone: Multi-Task Detection and Attribution of LLM-Generated Text
- Bites of Tomorrow: Personalized Recommendations for a Healthier and Greener Plate
- The Promise of Large Language Models in Digital Health: Evidence from Sentiment Analysis in Online Health Communities
- ChronoLLM: Customizing Language Models for Physics-Based Simulation Code Generation
- ReviewGraph: A Knowledge Graph Embedding Based Framework for Review Rating Prediction with Sentiment Features
- CARE: Contextual Adaptation of Recommenders for LLM-based Conversational Recommendation
- Exit Stories: Using Reddit Self-Disclosures to Understand Disengagement from Problematic Communities
- The illusion of a perfect metric: Why evaluating AI's words is harder than it looks
- PENGUIN: Enhancing Transformer with Periodic-Nested Group Attention for Long-term Time Series Forecasting
- Prediction is not Explanation: Revisiting the Explanatory Capacity of Mapping Embeddings
- Interactive Query Answering on Knowledge Graphs with Soft Entity Constraints
- `My Dataset of Love': A Preliminary Mixed-Method Exploration of Human-AI Romantic Relationships
- Equinox: Holistic Fair Scheduling in Serving Large Language Models
- Compressed Models are NOT Trust-equivalent to Their Large Counterparts
- Evaluating Open-Source Vision Language Models for Facial Emotion Recognition against Traditional Deep Learning Models
- Scalable Scientific Interest Profiling Using Large Language Models
- Extracting Structured Requirements from Unstructured Building Technical Specifications for Building Information Modeling
- CCFC: Core & Core-Full-Core Dual-Track Defense for LLM Jailbreak Protection
- DPad: Efficient Diffusion Language Models with Suffix Dropout
- MHSNet:An MoE-based Hierarchical Semantic Representation Network for Accurate Duplicate Resume Detection with Large Language Model
- Graph Concept Bottleneck Models
- FLAIR: Feedback Learning for Adaptive Information Retrieval
- Integrating Feedback Loss from Bi-modal Sarcasm Detector for Sarcastic Speech Synthesis
- MaskSem: Semantic-Guided Masking for Learning 3D Hybrid High-Order Motion Representation
- RUM: Rule+LLM-Based Comprehensive Assessment on Testing Skills
- Explicit v.s. Implicit Memory: Exploring Multi-hop Complex Reasoning Over Personalized Information
- Hallucinations in medical devices
- REACH: Reinforcement Learning for Efficient Allocation in Community and Heterogeneous Networks
- Wavy Transformer
- Prompt-Induced Linguistic Fingerprints for LLM-Generated Fake News Detection
- SDEC: Semantic Deep Embedded Clustering
- The Application of Transformer-Based Models for Predicting Consequences of Cyber Attacks
- Leveraging Large Language Models for Predictive Analysis of Human Misery
- Evaluating ASR robustness to spontaneous speech errors: A study of WhisperX using a Speech Error Database
- MuDRiC: Multi-Dialect Reasoning for Arabic Commonsense Validation
- MDPO: Overcoming the Training-Inference Divide of Masked Diffusion Language Models
- Beyond Modality Limitations: A Unified MLLM Approach to Automated Speaking Assessment with Effective Curriculum Learning
- CASPER: Concept-integrated Sparse Representation for Scientific Retrieval
- FLARE: Fast Low-rank Attention Routing Engine
- Feature Request Analysis and Processing: Tasks, Techniques, and Trends
- TaoSR1: The Thinking Model for E-commerce Relevance Search
- HuBERT-VIC: Improving Noise-Robust Automatic Speech Recognition of Speech Foundation Model via Variance-Invariance-Covariance Regularization
- Incorporating Legal Logic into Deep Learning: An Intelligent Approach to Probation Prediction
- SEA-BED: Southeast Asia Embedding Benchmark
- LinkAnchor: An Autonomous LLM-Based Agent for Issue-to-Commit Link Recovery
- Exploring Self-Supervised Audio Models for Generalized Anomalous Sound Detection
- Towards Generalizable Human Activity Recognition: A Survey
- VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine
- AI Models for Depressive Disorder Detection and Diagnosis: A Review
- MOON: Generative MLLM-based Multimodal Representation Learning for E-commerce Product Understanding
- A Comprehensive Review of AI Agents: Transforming Possibilities in Technology and Beyond
- What Matters for Bioacoustic Encoding
- The Rise of Generative AI for Metal-Organic Framework Design and Synthesis
- Representing Speech Through Autoregressive Prediction of Cochlear Tokens
- TrajSV: A Trajectory-based Model for Sports Video Representations and Applications
- Handwritten Text Recognition of Historical Manuscripts Using Transformer-Based Models
- Reference Points in LLM Sentiment Analysis: The Role of Structured Context
- CoDiEmb: A Collaborative yet Distinct Framework for Unified Representation Learning in Information Retrieval and Semantic Textual Similarity
- NucEL: Single-Nucleotide ELECTRA-Style Genomic Pre-training for Efficient and Interpretable Representations
- Model Interpretability and Rationale Extraction by Input Mask Optimization
- A Global Dataset of Location Data Integrity-Assessed Reforestation Efforts
- NeMo: A Neuron-Level Modularizing-While-Training Approach for Decomposing DNN Models
- Semantically Guided Adversarial Testing of Vision Models Using Language Models
- MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts
- Delving into Dynamic Scene Cue-Consistency for Robust 3D Multi-Object Tracking
- Group Fairness Meets the Black Box: Enabling Fair Algorithms on Closed LLMs via Post-Processing
- E-CaTCH: Event-Centric Cross-Modal Attention with Temporal Consistency and Class-Imbalance Handling for Misinformation Detection
- Representation Quantization for Collaborative Filtering Augmentation
- Expressive Speech Retrieval using Natural Language Descriptions of Speaking Style
- Overcoming Low-Resource Barriers in Tulu: Neural Models and Corpus Creation for OffensiveLanguage Identification
- Towards the Next-generation Bayesian Network Classifiers
- ADMIRE-BayesOpt: Accelerated Data MIxture RE-weighting for Language Models with Bayesian Optimization
- Forecasting Clicks in Digital Advertising: Multimodal Inputs and Interpretable Outputs
- Labels or Input? Rethinking Augmentation in Multimodal Hate Detection
- Scalable Geospatial Data Generation Using AlphaEarth Foundations Model
- Copyright Protection for Large Language Models: A Survey of Methods, Challenges, and Trends
- Hybrid-Hierarchical Fashion Graph Attention Network for Compatibility-Oriented and Personalized Outfit Recommendation
- Are Large Pre-trained Vision Language Models Effective Construction Safety Inspectors?
- Improving Text Style Transfer using Masked Diffusion Language Models with Inference-time Scaling
- Human-in-Context: Unified Cross-Domain 3D Human Motion Modeling via In-Context Learning
- MAESTRO: Masked AutoEncoders for Multimodal, Multitemporal, and Multispectral Earth Observation Data
- STream3R: Scalable Sequential 3D Reconstruction with Causal Transformer
- A Multimodal Neural Network for Recognizing Subjective Self-Disclosure Towards Social Robots
- IBEX: Information-Bottleneck-EXplored Coarse-to-Fine Molecular Generation under Limited Data
- Thinking Inside the Mask: In-Place Prompting in Diffusion LLMs
- GenOM: Ontology Matching with Description Generation and Large Language Model
- Med-GLIP: Advancing Medical Language-Image Pre-training with Large-scale Grounded Dataset
- A Retrieval Augmented Spatio-Temporal Framework for Traffic Prediction
- Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning
- Improving Generative Cross-lingual Aspect-Based Sentiment Analysis with Constrained Decoding
- Large Language Models for Summarizing Czech Historical Documents and Beyond
- Advancing Cross-lingual Aspect-Based Sentiment Analysis with LLMs and Constrained Decoding for Sequence-to-Sequence Models
- "Here Comes the Makeup Tutorial You Asked For!": Exploring Communication Strategies and Viewer Engagement in Beauty Videos on Rednote
- Flexible Personalized Split Federated Learning for On-Device Fine-Tuning of Foundation Models
- A Curriculum Learning Approach to Reinforcement Learning: Leveraging RAG for Multimodal Question Answering
- From Surface to Semantics: Semantic Structure Parsing for Table-Centric Document Analysis
- BERTector: An Intrusion Detection Framework Constructed via Joint-dataset Learning Based on Language Model
- Meta-Metrics and Best Practices for System-Level Inference Performance Benchmarking
- Pruning and Malicious Injection: A Retraining-Free Backdoor Attack on Transformer Models
- On the synchronization between Hugging Face pre-trained language models and their upstream GitHub repository
- Bridging Modality Gaps in e-Commerce Products via Vision-Language Alignment
- Interpretable Oracle Bone Script Decipherment through Radical and Pictographic Analysis with LVLMs
- Data-Driven Discovery of Interpretable Kalman Filter Variants through Large Language Models and Genetic Programming
- GBC: Generalized Behavior-Cloning Framework for Whole-Body Humanoid Imitation
- A BERT-based Hierarchical Classification Model with Applications in Chinese Commodity Classification
- SpeechForensics: Audio-Visual Speech Representation Learning for Face Forgery Detection
- A Survey of Cognitive Distortion Detection and Classification in NLP
- Can LLM-Generated Textual Explanations Enhance Model Classification Performance? An Empirical Study
- Social-Sensor Identity Cloning Detection Using Weakly Supervised Deep Forest and Cryptographic Authentication
- Personalized Product Search Ranking: A Multi-Task Learning Approach with Tabular and Non-Tabular Data
- Artificial Intelligence, Domain AI Readiness, and Firm Productivity
- Towards Self-cognitive Exploration: Metacognitive Knowledge Graph Retrieval Augmented Generation
- A Chain of Diagnosis Framework for Accurate and Explainable Radiology Report Generation
- WeatherPrompt: Multi-modality Representation Learning for All-Weather Drone Visual Geo-Localization
- Exploring the Equivalence of Closed-Set Generative and Real Data Augmentation in Image Classification
- SYNAPSE-G: Bridging Large Language Models and Graph Learning for Rare Event Classification
- Improving Dense Passage Retrieval with Multiple Positive Passages
- Cross-lingual Aspect-Based Sentiment Analysis: A Survey on Tasks, Approaches, and Challenges
- LACA: Improving Cross-lingual Aspect-Based Sentiment Analysis with LLM Data Augmentation
- MPT: Motion Prompt Tuning for Micro-Expression Recognition
- Harnessing Input-Adaptive Inference for Efficient VLN
- Link Prediction for Event Logs in the Process Industry
- Utilizing Multilingual Encoders to Improve Large Language Models for Low-Resource Languages
- Text-conditioned State Space Model For Domain-generalized Change Detection Visual Question Answering
- InteChar: A Unified Oracle Bone Character List for Ancient Chinese Language Modeling
- M2LLM: Multi-view Molecular Representation Learning with Large Language Models
- Prompt-Based Approach for Czech Sentiment Analysis
- UWB at WASSA-2024 Shared Task 2: Cross-lingual Emotion Detection
- Agentic Graph Neural Networks for Wireless Communications and Networking Towards Edge General Intelligence: A Survey
- Privacy Preserving Inference of Personalized Content for Out of Matrix Users
- LyS at SemEval 2025 Task 8: Zero-Shot Code Generation for Tabular QA
- SegDAC: Visual Generalization in Reinforcement Learning via Dynamic Object Tokens
- A Guide to Robust Generalization: The Impact of Architecture, Pre-training, and Optimization Strategy
- Explainable Graph Spectral Clustering For GloVe-like Text Embeddings
- DevNous: An LLM-Based Multi-Agent System for Grounding IT Project Management in Unstructured Conversation
- Weakly Supervised Fine-grained Span-Level Framework for Chinese Radiology Report Quality Assurance
- Integrating attention into explanation frameworks for language and vision transformers
- Deep Neural Network Calibration by Reducing Classifier Shift with Stochastic Masking
- DeCAL Tokenwise Compression
- MuGa-VTON: Multi-Garment Virtual Try-On via Diffusion Transformers with Prompt Customization
- Rethinking Tokenization for Rich Morphology: The Dominance of Unigram over BPE and Morphological Alignment
- Generating Query-Relevant Document Summaries via Reinforcement Learning
- From Source to Target: Leveraging Transfer Learning for Predictive Process Monitoring in Organizations
- Mining the Social Fabric: Unveiling Communities for Fake News Detection in Short Videos
- LLMs for Law: Evaluating Legal-Specific LLMs on Contract Understanding
- UniSVG: A Unified Dataset for Vector Graphic Understanding and Generation with Multimodal Large Language Models
- LoSemB: Logic-Guided Semantic Bridging for Inductive Tool Retrieval
- Attribution Explanations for Deep Neural Networks: A Theoretical Perspective
- Towards Comprehensible Recommendation with Large Language Model Fine-tuning
- Detecting Mislabeled and Corrupted Data via Pointwise Mutual Information
- Large Language Models for Subjective Language Understanding: A Survey
- Augmenting Bias Detection in LLMs Using Topological Data Analysis
- Temporal User Profiling with LLMs: Balancing Short-Term and Long-Term Preferences for Recommendations
- Using LLMs to Capture Users' Temporal Context for Recommendation
- Pareto Multi-Objective Alignment for Language Models
- Czech Dataset for Complex Aspect-Based Sentiment Analysis Tasks
- HGMF: A Hierarchical Gaussian Mixture Framework for Scalable Tool Invocation within the Model Context Protocol
- Heterogeneity in Entity Matching: A Survey and Experimental Analysis
- Semantic-Enhanced Time-Series Forecasting via Large Language Models
- Iterative refinement, not training objective, makes HuBERT behave differently from wav2vec 2.0
- Steerable Pluralism: Pluralistic Alignment via Few-Shot Comparative Regression
- Modeling and Detecting Company Risks from News: A Case Study in Bloomberg News
- An Embodied AR Navigation Agent: Integrating BIM with Retrieval-Augmented Generation for Language Guidance
- Préface
- A Survey on Non-Intrusive ASR Refinement: From Output-Level Correction to Full-Model Distillation
- Prompt Tuning for Few-Shot Continual Learning Named Entity Recognition
- Selection and Exploitation of High-Quality Knowledge from Large Language Models for Recommendation
- Bridging Semantic Logic Gaps: A Cognition Inspired Multimodal Boundary Preserving Network for Image Manipulation Localization
- Propagation Tree Is Not Deep: Adaptive Graph Contrastive Learning Approach for Rumor Detection
- Improved Personalized Headline Generation via Denoising Fake Interests from Implicit Feedback
- Arce: Augmented Roberta with Contextualized Elucidations for Ner in Automated Rule Checking
- Enhancing Rumor Detection Methods with Propagation Structure Infused Language Model
- AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning
- PrLM: Learning Explicit Reasoning for Personalized RAG via Contrastive Reward Optimization
- Towards Real-World Rumor Detection: Anomaly Detection Framework with Graph Supervised Contrastive Learning
- Towards Safer AI Moderation: Evaluating LLM Moderators Through a Unified Benchmark Dataset and Advocating a Human-First Approach
- ESNERA: Empirical and semantic named entity alignment for named entity dataset merging
- Natural Language-Driven Viewpoint Navigation for Volume Exploration via Semantic Block Representation
- Anatomy of a Machine Learning Ecosystem: 2 Million Models on Hugging Face
- SEVADE: Self-Evolving Multi-Agent Analysis with Decoupled Evaluation for Hallucination-Resistant Irony Detection
- Fed MobiLLM: Efficient Federated LLM Fine-Tuning over Heterogeneous Mobile Devices via Server Assisted Side-Tuning
- Adversarial Video Promotion Against Text-to-Video Retrieval
- Structure-Preserving Digital Twins via Conditional Neural Whitney Forms
- Generalizing Scaling Laws for Dense and Sparse Large Language Models
- TRUST: Leveraging Text Robustness for Unsupervised Domain Adaptation
- Learning the Topic, Not the Language: How LLMs Classify Online Immigration Discourse Across Languages
- Tree-Based Deep Learning for Ranking Symbolic Integration Algorithms
- Text as Any-Modality for Zero-Shot Classification by Consistent Prompt Tuning
- Aligning Effective Tokens with Video Anomaly in Large Language Models
- Deep Language Geometry: Constructing a Metric Space from LLM Weights
- Multi-Objective Instruction-Aware Representation Learning in Procedural Content Generation RL
- Text-guided Visual Prompt DINO for Generic Segmentation
- MAHL: Multi-Agent LLM-Guided Hierarchical Chiplet Design with Adaptive Debugging
- Learning 3D Texture-Aware Representations for Parsing Diverse Human Clothing and Body Parts
- Fine-Grained Safety Neurons with Training-Free Continual Projection to Reduce LLM Fine Tuning Risks
- Bifrost-1: Bridging Multimodal LLMs and Diffusion Models with Patch-level CLIP Latents
- Efficient Multimodal Streaming Recommendation via Expandable Side Mixture-of-Experts
- Large Language Model Data Generation for Enhanced Intent Recognition in German Speech
- Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge
- Omni Geometry Representation Learning vs Large Language Models for Geospatial Entity Resolution
- Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM
- LATTE: Learning Aligned Transactions and Textual Embeddings for Bank Clients
- PRvL: Quantifying the Capabilities and Risks of Large Language Models for PII Redaction
- Streamlining Admission with LOR Insights: AI-Based Leadership Assessment in Online Master's Program
- Task complexity shapes internal representations and robustness in neural networks
- Large Language Models Transform Organic Synthesis From Reaction Prediction to Automation
- Categorising SME Bank Transactions with Machine Learning and Synthetic Data Generation
- UNCAGE: Contrastive Attention Guidance for Masked Generative Transformers in Text-to-Image Generation
- Optimal Corpus Aware Training for Neural Machine Translation
- Textual Inversion for Efficient Adaptation of Open-Vocabulary Object Detectors Without Forgetting
- B4DL: A Benchmark for 4D LiDAR LLM in Spatio-Temporal Understanding
- CodeBoost: Boosting Code LLMs by Squeezing Knowledge from Code Snippets with RL
- ReasoningTrack: Chain-of-Thought Reasoning for Long-term Vision-Language Tracking
- RRRA: Resampling and Reranking through a Retriever Adapter
- A Survey on Video Temporal Grounding with Multimodal Large Language Model
- Tool Graph Retriever: Exploring Dependency Graph-based Tool Retrieval for Large Language Models
- Navigating Through Paper Flood: Advancing LLM-based Paper Evaluation through Domain-Aware Retrieval and Latent Reasoning
- Cognitive Duality for Adaptive Web Agents
- FedCoT: Communication-Efficient Federated Reasoning Enhancement for Large Language Models
- A Context-aware Attention and Graph Neural Network-based Multimodal Framework for Misogyny Detection
- A Study of the Framework and Real-World Applications of Language Embedding for 3D Scene Understanding
- ImpliHateVid: A Benchmark Dataset and Two-stage Contrastive Learning Framework for Implicit Hate Speech Detection in Videos
- Disentangling Bias by Modeling Intra- and Inter-modal Causal Attention for Multimodal Sentiment Analysis
- Tesserae: Scalable Placement Policies for Deep Learning Workloads
- Self-Error Adjustment: Theory and Practice of Balancing Individual Performance and Diversity in Ensemble Learning
- An Effective Approach for Node Classification in Textual Graphs
- Extending Foundational Monocular Depth Estimators to Fisheye Cameras with Calibration Tokens
- Why Attention Fails: A Taxonomy of Faults in Attention-Based Neural Networks
- Advancing Hate Speech Detection with Transformers: Insights from the MetaHate
- Revealing Temporal Label Noise in Multimodal Hateful Video Classification
- Benchmarking Sociolinguistic Diversity in Swahili NLP: A Taxonomy-Guided Approach
- Fine-Tuning Small Language Models (SLMs) for Autonomous Web-based Geographical Information Systems (AWebGIS)
- Uncertainty-aware Predict-Then-Optimize Framework for Equitable Post-Disaster Power Restoration
- A Scalable Pretraining Framework for Link Prediction with Efficient Adaptation
- Do Recommender Systems Really Leverage Multimodal Content? A Comprehensive Analysis on Multimodal Representations for Recommendation
- One Model for All: Unified Try-On and Try-Off in Any Pose via LLM-Inspired Bidirectional Tweedie Diffusion
- Privacy Risk Predictions Based on Fundamental Understanding of Personal Data and an Evolving Threat Landscape
- InceptoFormer: A Multi-Signal Neural Framework for Parkinson's Disease Severity Evaluation from Gait
- CALE : Concept-Aligned Embeddings for Both Within-Lemma and Inter-Lemma Sense Differentiation
- GFocal: A Global-Focal Neural Operator for Solving PDEs on Arbitrary Geometries
- Boosting Visual Knowledge-Intensive Training for LVLMs Through Causality-Driven Visual Object Completion
- Dialogue Response Prefetching Based on Semantic Similarity and Prediction Confidence of Language Model
- VisionTS++: Cross-Modal Time Series Foundation Model with Continual Pre-trained Vision Backbones
- Modelling and Classifying the Components of a Literature Review
- Leveraging large language models for SQL behavior-based database intrusion detection
- KG-Augmented Executable CoT for Mathematical Coding
- Segment Any Vehicle: Semantic and Visual Context Driven SAM and A Benchmark
- Graph Representation Learning with Massive Unlabeled Data for Rumor Detection
- I3-MRec: Invariant Learning with Information Bottleneck for Incomplete Modality Recommendation
- DP-DocLDM: Differentially Private Document Image Generation using Latent Diffusion Models
- Evaluating Selective Encryption Against Gradient Inversion Attacks
- STARE: Predicting Decision Making Based on Spatio-Temporal Eye Movements
- Benefit from Rich: Tackling Search Interaction Sparsity in Search Enhanced Recommendation
- Slice or the Whole Pie? Utility Control for AI Models
- A Comparative Survey of PyTorch vs TensorFlow for Deep Learning: Usability, Performance, and Deployment Trade-offs
- Enhancing Serendipity Recommendation System by Constructing Dynamic User Knowledge Graphs with Large Language Models
- Uncertainty-Aware GUI Agent: Adaptive Perception through Component Recommendation and Human-in-the-Loop Refinement
- Factor Augmented Supervised Learning with Text Embeddings
- LLM4ES: Learning User Embeddings from Event Sequences via Large Language Models
- Model Compression vs. Adversarial Robustness: An Empirical Study on Language Models for Code
- FairLangProc: A Python package for fairness in NLP
- LLMDistill4Ads: Using Cross-Encoders to Distill from LLM Signals for Advertiser Keyphrase Recommendations
- AttZoom: Attention Zoom for Better Visual Features
- OpenLifelogQA: An Open-Ended Multi-Modal Lifelog Question-Answering Dataset
- Speech-to-LaTeX: New Models and Datasets for Converting Spoken Equations and Sentences
- Cropping outperforms dropout as an augmentation strategy for training self-supervised text embeddings
- Neighborhood-Preserving Voronoi Treemaps
- User Perception of Attention Visualizations: Effects on Interpretability Across Evidence-Based Medical Documents
- Variety Is the Spice of Life: Detecting Misinformation with Dynamic Environmental Representations
- Efficient Multi-Slide Visual-Language Feature Fusion for Placental Disease Classification
- LECTOR: LLM-Enhanced Concept-based Test-Oriented Repetition for Adaptive Spaced Learning
- Exploring Stability-Plasticity Trade-offs for Continual Named Entity Recognition
- RooseBERT: A New Deal For Political Language Modelling
- BadBlocks: Low-Cost and Stealthy Backdoor Attacks Tailored for Text-to-Image Diffusion Models
- Adaptive Sparse Softmax: An Effective and Efficient Softmax Variant
- ContractEval: Benchmarking LLMs for Clause-Level Legal Risk Identification in Commercial Contracts
- AGENTiGraph: A Multi-Agent Knowledge Graph Framework for Interactive, Domain-Specific LLM Chatbots
- Separating Shared and Domain-Specific LoRAs for Multi-Domain Learning
- Probing Syntax in Large Language Models: Successes and Remaining Challenges
- Domain-Specific Fine-Tuning and Prompt-Based Learning: A Comparative Study for developing Natural Language-Based BIM Information Retrieval Systems
- CTR-Sink: Attention Sink for Language Models in Click-Through Rate Prediction
- From literature to biodiversity data: mining arthropod organismal traits with machine learning
- PyLate: Flexible Training and Retrieval for Late Interaction Models
- FedPromo: Federated Lightweight Proxy Models at the Edge Bring New Domains to Foundation Models
- HiTeC: Hierarchical Contrastive Learning on Text-Attributed Hypergraph with Semantic-Aware Augmentation
- CoughViT: A Self-Supervised Vision Transformer for Cough Audio Representation Learning
- LLM-based IR-system for Bank Supervisors
- Realizing Scaling Laws in Recommender Systems: A Foundation-Expert Paradigm for Hyperscale Model Deployment
- Tricks and Plug-ins for Gradient Boosting with Transformers
- SLIM-LLMs: Modeling of Style-Sensory Language RelationshipsThrough Low-Dimensional Representations
- Beyond Least Squares: Robust Regression Transformer (R2T)
- Evaluation and Analysis of Deep Neural Transformers and Convolutional Neural Networks on Modern Remote Sensing Datasets
- Modeling Annotator Disagreement with Demographic-Aware Experts and Synthetic Perspectives
- TransAM: Transformer-Based Agent Modeling for Multi-Agent Systems via Local Trajectory Encoding
- LaMPE: Length-aware Multi-grained Positional Encoding for Adaptive Long-context Scaling Without Training
- AirTrafficGen: Configurable Air Traffic Scenario Generation with Large Language Models
- I2CR: Intra- and Inter-modal Collaborative Reflections for Multimodal Entity Linking
- GaussianCross: Cross-modal Self-supervised 3D Representation Learning via Gaussian Splatting
- Towards Immersive Human-X Interaction: A Real-Time Framework for Physically Plausible Motion Synthesis
- VLM4D: Towards Spatiotemporal Awareness in Vision Language Models
- Toward Efficient Spiking Transformers: Synapse Pruning Meets Synergistic Learning-Based Compensation
- Charting 15 years of progress in deep learning for speech emotion recognition: A replication study
- Contextual Graph Transformer: A Small Language Model for Enhanced Engineering Document Information Extraction
- User Trajectory Prediction Unifying Global and Local Temporal Information
- Versatile yet Efficient Network Traffic Analysis: Offloading Network Foundation Model to SmartNIC
- HGTS-Former: Hierarchical HyperGraph Transformer for Multivariate Time Series Analysis
- Alleviating Attention Hacking in Discriminative Reward Modeling through Interaction Distillation
- DINA: A Dual Defense Framework Against Internal Noise and External Attacks in Natural Language Processing
- ReMoMask: Retrieval-Augmented Masked Motion Generation
- Parameter-Efficient Routed Fine-Tuning: Mixture-of-Experts Demands Mixture of Adaptation Modules
- LMAR: Language Model Augmented Retriever for Domain-specific Knowledge Indexing
- MicroMix: Efficient Mixed-Precision Quantization with Microscaling Formats for Large Language Models
- Zero-shot Compositional Action Recognition with Neural Logic Constraints
- GlaBoost: A multimodal Structured Framework for Glaucoma Risk Stratification
- Proactive Disentangled Modeling of Trigger-Object Pairings for Backdoor Defense
- Quantum-RAG and PunGPT2: Advancing Low-Resource Language Generation and Retrieval for the Punjabi Language
- Context-Adaptive Multi-Prompt Embedding with Large Language Models for Vision-Language Alignment
- Context Guided Transformer Entropy Modeling for Video Compression
- HateClipSeg: A Segment-Level Annotated Dataset for Fine-Grained Hate Video Detection
- The Bidirectional Process Reward Model
- DUP: Detection-guided Unlearning for Backdoor Purification in Language Models
- Empowering Tabular Data Preparation with Language Models: Why and How?
- Social Media Information Operations
- HeQ: a Large and Diverse Hebrew Reading Comprehension Benchmark
- CLASS: Contrastive Learning via Action Sequence Supervision for Robot Manipulation
- ChEmbed: Enhancing Chemical Literature Search Through Domain-Specific Text Embeddings
- OpenMed NER: Open-Source, Domain-Adapted State-of-the-Art Transformers for Biomedical NER Across 12 Public Datasets
- TeSent: A Benchmark Dataset for Fairness-aware Explainable Sentiment Classification in Telugu
- HT-Transformer: Event Sequences Classification by Accumulating Prefix Information with History Tokens
- ConfGuard: A Simple and Effective Backdoor Detection for Large Language Models
- Referring Remote Sensing Image Segmentation with Cross-view Semantics Interaction Network
- Aligning Language Models with Real-time Knowledge Editing
- How Far Are LLMs from Symbolic Planners? An NLP-Based Perspective
- Multi-Operator Few-Shot Learning for Generalization Across PDE Families
- CreditARF: A Framework for Corporate Credit Rating with Annual Report and Financial Feature Integration
- Benchmarking and Bridging Emotion Conflicts for Multimodal Emotion Reasoning
- Recovering Individual-Level Activity Sequences from Location-Based Service Data Using a Novel Transformer-Based Model
- FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models
- R2-CoD: Understanding Text-Graph Complementarity in Relational Reasoning via Knowledge Co-Distillation
- NATLM: Detecting Defects in NFT Smart Contracts Leveraging LLM
- T2S: Tokenized Skill Scaling for Lifelong Imitation Learning
- DisTaC: Conditioning Task Vectors via Distillation for Robust Model Merging
- Capturing More: Learning Multi-Domain Representations for Robust Online Handwriting Verification
- Graph-based Interaction Augmentation Network for Robust Multimodal Sentiment Analysis
- Teaching at Scale: Leveraging AI to Evaluate and Elevate Engineering Education
- UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu
- Interpreting Performance Profiles with Deep Learning
- Diffusion-Scheduled Denoising Autoencoders for Anomaly Detection in Tabular Data
- How LLMs are Shaping the Future of Virtual Reality
- Forecasting NCAA Basketball Outcomes with Deep Learning: A Comparative Study of LSTM and Transformer Models
- Can User Feedback Help Issue Detection? An Empirical Study on a One-billion-user Online Service System
- Court of LLMs: Evidence-Augmented Generation via Multi-LLM Collaboration for Text-Attributed Graph Anomaly Detection
- Learning Unified User Quantized Tokenizers for User Representation
- Decouple before Align: Visual Disentanglement Enhances Prompt Tuning
- From Generator to Embedder: Harnessing Innate Abilities of Multimodal LLMs via Building Zero-Shot Discriminative Embedding Model
- Bidirectional Action Sequence Learning for Long-term Action Anticipation with Large Language Models
- Multimodal Referring Segmentation: A Survey
- DACTYL: Diverse Adversarial Corpus of Texts Yielded from Large Language Models
- Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts
- SPENCER: Self-Adaptive Model Distillation for Efficient Code Retrieval
- NyayaRAG: Realistic Legal Judgment Prediction with RAG under the Indian Common Law System
- Addressing Cold Start For next-article Recommendation
- Improving Multimodal Contrastive Learning of Sentence Embeddings with Object-Phrase Alignment
- From Individuals to Crowds: Dual-Level Public Response Prediction in Social Media
- Semantic Compression for Word and Sentence Embeddings using Discrete Wavelet Transform
- TriP-LLM: A Tri-Branch Patch-wise Large Language Model Framework for Time-Series Anomaly Detection
- Deep Learning-based Prediction of Clinical Trial Enrollment with Uncertainty Estimates
- Role-Aware Language Models for Secure and Contextualized Access Control in Organizations
- Mitigating Resolution-Drift in Federated Learning: Case of Keypoint Detection
- Enhanced Arabic Text Retrieval with Attentive Relevance Scoring
- Causal2Vec: Improving Decoder-only LLMs as Versatile Embedding Models
- Multi-Prompt Progressive Alignment for Multi-Source Unsupervised Domain Adaptation
- Foundations and Models in Modern Computer Vision: Key Building Blocks in Landmark Architectures
- MUST-RAG: MUSical Text Question Answering with Retrieval Augmented Generation
- What's Taboo for You? - An Empirical Evaluation of LLMs Behavior Toward Sensitive Content
- Unveiling Super Experts in Mixture-of-Experts Large Language Models
- A Bayesian Hybrid Parameter-Efficient Fine-Tuning Method for Large Language Models
- From Image Captioning to Visual Storytelling
- Improved Algorithms for Kernel Matrix-Vector Multiplication Under Sparsity Assumptions
- Your Spending Needs Attention: Modeling Financial Habits with Transformers
- Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space
- AGA: An adaptive group alignment framework for structured medical cross-modal representation learning
- Reinitializing weights vs units for maintaining plasticity in neural networks
- Learning Like Humans: Resource-Efficient Federated Fine-Tuning through Cognitive Developmental Stages
- Distilling Knowledge from Large Language Models: A Concept Bottleneck Model for Hate and Counter Speech Recognition
- Uncovering the Fragility of Trustworthy LLMs through Chinese Textual Ambiguity
- TR-PTS: Task-Relevant Parameter and Token Selection for Efficient Tuning
- Real-time News Story Identification
- Quantifying surprise in clinical care: Detecting highly informative events in electronic health records with foundation models
- Opportunities and Challenges of LLMs in Education: An NLP Perspective
- Resource-Efficient Adaptation of Large Language Models for Text Embeddings via Prompt Engineering and Contrastive Fine-tuning
- OFCnetLLM: Large Language Model for Network Monitoring and Alertness
- Multilingual Political Views of Large Language Models: Identification and Steering
- RePaCA: Leveraging Reasoning Large Language Models for Static Automated Patch Correctness Assessment
- XAutoLM: Efficient Fine-Tuning of Language Models via Meta-Learning and AutoML
- Pre-trained Models Perform the Best When Token Distributions Follow Zipf's Law
- IFEvalCode: Controlled Code Generation
- Traits Run Deep: Enhancing Personality Assessment via Psychology-Guided LLM Representations and Multimodal Apparent Behaviors
- A Comprehensive Taxonomy of Negation for NLP and Neural Retrievers
- Towards Experiment Execution in Support of Community Benchmark Workflows for HPC
- Context-aware Rotary Position Embedding
- Intent Recognition and Out-of-Scope Detection using LLMs in Multi-party Conversations
- Multi-modal Relational Item Representation Learning for Inferring Substitutable and Complementary Items
- Generative Recommendation with Semantic IDs: A Practitioner's Handbook
- TESPEC: Temporally-Enhanced Self-Supervised Pretraining for Event Cameras
- Ashish Vaswani [wikipedia]
- Attention Is All You Need [wikipedia]
- BERT (language model) [wikipedia]
- BookCorpus [wikipedia]
- Geospatial foundation model [wikipedia]
- Information retrieval [wikipedia]
- Language model [wikipedia]
- List of large language models [wikipedia]
- Neuro-symbolic AI [wikipedia]
- Speech recognition [wikipedia]
- Transformer (deep learning) [wikipedia]
Discussions
- BERT: Pre-Training of Deep Bidirectional Transformers for Language Understanding [hn, 78 points, 5 comments]
- Bert: Pre-Training of Deep Bidirectional Transformers for Language Understanding [hn, 3 points, 0 comments]
- BERT-like (arxiv.org/abs/1810.04805) models works by doing "masked" language modeling, so instead of predicting the next token like a causal LLM, you predict an arbitrary set of tokens. In particular, [bsky, 2 points, 0 comments]
- One of the important innovations of BERT arxiv.org/abs/1810.04805, now 6 years old, was that the team released mutlilingual variants trained on 100 languages github.com/google-resea.... While the perf [bsky, 1 points, 1 comments]
- 📝 BERT arxiv.org/abs/1810.04805 [bsky, 0 points, 0 comments]
- 🎧 EP005: How BERT Mastered Language by Hiding Words 📄 BERT 🔗 https://arxiv.org/abs/1810.04805 🟢 https://podcasters.spotify.com/pod/show/yun-wu/episodes/EP005-How-BERT-Mastered-Language-by-Hiding-W [bsky, 0 points, 0 comments]
- BERT: Encoder is all you need. Also, left-to-right language modeling is NOT all you need. (Also, pre-training + finetuning 📈) https://arxiv.org/abs/1810.04805 [bsky, 0 points, 1 comments]
Related