Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
2019/08/27 by Nils Reimers, Iryna Gurevych, Reimers, Nils +1 · 2,203 citations
Computer Science · #Natural Language Processing Techniques #Sentiment Analysis and Opinion Mining #Topic Modeling #cs.CL
paper · pdf · doi:10.48550/arxiv.1908.10084
Published at EMNLP 2019
arxiv created 2019/08/27 · arxiv updated 2019/08/28
Abstract
BERT (Devlin et al., 2018) and RoBERTa (Liu et al., 2019) has set a new state-of-the-art performance on sentence-pair regression tasks like semantic textual similarity (STS). However, it requires that both sentences are fed into the network, which causes a massive computational overhead: Finding the most similar pair in a collection of 10,000 sentences requires about 50 million inference computations (~65 hours) with BERT. The construction of BERT makes it unsuitable for semantic similarity search as well as for unsupervised tasks like clustering. In this publication, we present Sentence-BERT (SBERT), a modification of the pretrained BERT network that use siamese and triplet network structures to derive semantically meaningful sentence embeddings that can be compared using cosine-similarity. This reduces the effort for finding the most similar pair from 65 hours with BERT / RoBERTa to about 5 seconds with SBERT, while maintaining the accuracy from BERT. We evaluate SBERT and SRoBERTa on common STS tasks and transfer learning tasks, where it outperforms other state-of-the-art sentence embeddings methods.
Cited by
- Evaluating the Effect of Linguistic Relatedness on Cross-Lingual Transfer in Large Multilingual Automatic Speech Recognition
- NDCG-Consistent Softmax Approximation with Accelerated Convergence
- From transcription to semantic corpus analysis: unsupervised learning of sentence representations for ancient languages
- AGRO-SQL: Agentic Group-Relative Optimization with High-Fidelity Data Synthesis
- Multi Agents Semantic Emotion Aligned Music to Image Generation with Music Derived Captions
- Debugging Tabular Log as Dynamic Graphs
- MUSON: A Reasoning-oriented Multimodal Dataset for Socially Compliant Navigation in Urban Environments
- Self-Rewarded Multimodal Coherent Reasoning Across Diverse Visual Domains
- Keep It InMind: Benchmarking the Implicit-Association Blind Spot in Agent Memory
- HalluMat: Detecting Hallucinations in LLM-Generated Materials Science Content Through Multi-Stage Verification
- Valori: A Deterministic Memory Substrate for AI Systems
- Hierarchical Geometry of Cognitive States in Transformer Embedding Spaces
- iOS as Acceleration
- Recursive Governance: A Graph-Theoretic Framework for Risk Propagation and Drift Detection in Agentic AI Systems
- From Unstructured Recall to Schema-Grounded Memory: Reliable AI Memory via Iterative, Schema-Aware Extraction
- Advocating the potential of artificial intelligence for syndrome discovery in syndromic surveillance systems: A scoping review
- PlanCraft: Sketch, Refine, and Furnish for Architect-Inspired Progressive 3D Residential Scene Generation
- Choosing a Text Embedding Model: A Practical Benchmarking and Decision Framework
- LinkRank: A Learning-to-Rank Framework for One-to-Many Issue-Commit Traceability
- seqLens: Optimizing Language Models for Genomic Predictions
- Do LLM Debates Repeat Arguments Differently Across Languages?
- LiFT-MPC: Language-in-the-Loop Feedback Tuning of Cost Previews for MPC
- TriShieldRAG: A Three-Ring Defense-in-Depth Framework Against Knowledge Corruption in Retrieval-Augmented Generation
- The Dynamics of Human and AI-Generated Language: How Semantic Content Fluctuates Across Different Timescales
- Where Quality Breaks in Compressed Short-Text Generation: Staged Bottleneck Localization
- StanceFlip: A Comprehensive Multi-Dimensional Benchmark for Multimodal Conversational Stance Flipping Forecasting
- Beyond Direct Answering: Aligning Educational LLMs as Socratic Guides via Heuristic Reinforcement Learning
- From Signals to Behaviors: Evidence-Based Android Malware Detection
- ADAGE: A Language-Agnostic Pipeline for Analogical Reasoning Evaluation
- Interpretable Column Annotation with LLM-Symbolized Decision Process Materialization
- Bekko Embedding: Parameter-Efficient Multilingual Retrieval with Ultra-Compact Encoders
- The Case Against Generation for Retrieval: Discriminative Language Models as Effective Retrievers
- From Semantics to Readout: Mechanistic Understanding of Audio Tokens after Fine-Tuning for Temporal Audio Grounding
- HVM-GraphRAG: Holistic-View Multimodal Graph Retrieval-Augmented Generation on Complex Document
- Traceable LLM Reasoning for Fake-Order Fraud Detection
- Frustratingly Simple Black-Box Adaptation of Language Models via Logit Bias
- Disentangling the Interpretive and Predictive Roles of LIWC: Controlled Substitution in Depression-Related Classification
- An Interactive Vision Language Platform for Cognitive Remediation in Schizophrenia
- MPR-CiteG: Enhancing RAG with Multi-Portfolio Retrieval and Citation-Grounded Generation
- SciClaimSeekers at CheckThat! 2026: Retrieving Scientific Sources for Social Media Claims with LLM Reranking
- SourceMinds at CheckThat! 2026: NLI-Grounded Citation Auditing in a Multi-Agent Pipeline for Full Fact-Checking Article Generation
- Multimodal Hybrid Retrieval-Augmented Generation for Scientific Document Understanding using Open-Source SLMs
- AI-Assisted Knowledge Access for Legacy Enterprise Asset Management in Energy Operations: A Practical Retrieval System
- Agent-based simulation of online social networks and disinformation
- Source-Aware Reranking for Retrieval-Augmented Generation: A Reliability Prior Approach
- Measuring and Improving Behavioral Consistency in Large Language Models through Fact-Heuristic-Emotion State Enforcement
- GrocLM: Grocery Category Recommendation in E-Commerce with Large Language Models
- HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
- JKO-RAG: Distributional Retrieval as Wasserstein Free-Energy Gradient Flow
- SeT-Diff: Towards Semantic Foundation Models for HPC Telemetry and Time-Series
- JobMatchAI-An Intelligent Job Matching Platform Using Knowledge Graphs, Semantic Search and Explainable AI
- To update or to separate: Neural signatures and consequences of latent cause inference in episodic memory
- Superposition Without Interference? Towards Isolated Interventions via Almost Orthogonal Features in Language Models
- Intuitive therapist robot patient physical interaction is worth a thousand words
- Explainable Statute Prediction via Attention-based Model and LLM Prompting
- MMCTOP: A Multimodal Textualization and Mixture-of-Experts Framework for Clinical Trial Outcome Prediction
- Knowledge Reasoning of Large Language Models Integrating Graph-Structured Information for Pest and Disease Control in Tobacco
- HELP: Hierarchical Embodied Language Planner for Household Tasks
- Enabling Conversational Behavior Reasoning Capabilities in Full-Duplex Speech
- ALETHEIA: Combating Social Media Influence Campaigns with Graph Neural Networks
- SENTINEL: A Multi-Modal Early Detection Framework for Emerging Cyber Threats using Telegram
- T2AV-Compass: Towards Unified Evaluation for Text-to-Audio-Video Generation
- Chorus: Multi-Teacher Pretraining for Holistic 3D Gaussian Scene Encoding
- AegisAgent: An Autonomous Defense Agent Against Prompt Injection Attacks in LLM-HARs
- How important is Recall for Measuring Retrieval Quality?
- Leveraging Lightweight Entity Extraction for Scalable Event-Based Image Retrieval
- X-GridAgent: An LLM-Powered Agentic AI System for Assisting Power Grid Analysis
- Structured Visualization Design Knowledge for Grounding Generative Reasoning and Situated Feedback
- Beyond Vision: Contextually Enriched Image Captioning with Multi-Modal Retrieval
- Interpolative Decoding: Exploring the Spectrum of Personality Traits in LLMs
- MemR3: Memory Retrieval via Reflective Reasoning for LLM Agents
- EchoTrail-GUI: Building Actionable Memory for GUI Agents via Critic-Guided Self-Exploration
- From Pixels to Predicates Structuring urban perception with scene graphs
- Understanding Chain-of-Thought in Large Language Models via Topological Data Analysis
- FC-MIR: A Mobile Screen Awareness Framework for Intent-Aware Recommendation based on Frame-Compressed Multimodal Trajectory Reasoning
- Learning Hierarchical Procedural Memory for LLM Agents through Bayesian Selection and Contrastive Refinement
- Cross-modal Counterfactual Explanations: Uncovering Decision Factors and Dataset Biases in Subjective Classification
- Explainable and Fine-Grained Safeguarding of LLM Multi-Agent Systems via Bi-Level Graph Anomaly Detection
- Probabilistic Digital Twins of Users: Latent Representation Learning with Statistically Validated Semantics
- External Hippocampus: Topological Cognitive Maps for Guiding Large Language Model Reasoning
- AnyTask: an Automated Task and Data Generation Framework for Advancing Sim-to-Real Policy Learning
- Incorporating Error Level Noise Embedding for Improving LLM-Assisted Robustness in Persian Speech Recognition
- Atom: Efficient On-Device Video-Language Pipelines Through Modular Reuse
- Reinforcement Learning for Self-Improving Agent with Skill Library
- Dynamic Tool Dependency Retrieval for Efficient Function Calling
- NRGPT: An Energy-based Alternative for GPT
- Persistent Multiscale Density-based Clustering
- From Essence to Defense: Adaptive Semantic-aware Watermarking for Embedding-as-a-Service Copyright Protection
- BrepLLM: Enabling Large Language Models to Understand Boundary Representations
- Design and Evaluation of Cost-Aware PoQ for Decentralized LLM Inference
- Coarse-to-Fine Open-Set Graph Node Classification with Large Language Models
- ModelTables: A Corpus of Tables about Models
- Convolutional Lie Operator for Sentence Classification
- ORACLE: Time-Dependent Recursive Summary Graphs for Foresight on News Data Using LLMs
- Social Story Frames: Contextual Reasoning about Narrative Intent and Reception
- ArcBERT: An LLM-based Search Engine for Exploring Integrated Multi-Omics Metadata
- SynGP500: A Clinically-Grounded Synthetic Dataset of Australian General Practice Medical Notes
- FAME: Fictional Actors for Multilingual Erasure
- The Semantic Illusion: Certified Limits of Embedding-Based Hallucination Detection in RAG Systems
- Examining the Utility of Self-disclosure Types for Modeling Annotators of Social Norms
- EVICPRESS: Joint KV-Cache Compression and Eviction for Efficient LLM Serving
- DP-Bench: A Benchmark for Evaluating Data Product Creation Systems
- IaC Generation with LLMs: An Error Taxonomy and A Study on Configuration Knowledge Injection
- SPARQL-LLM: Real-Time SPARQL Query Generation from Natural Language Questions
- Explainable Ethical Assessment on Human Behaviors by Generating Conflicting Social Norms
- Improving Semantic Uncertainty Quantification in LVLMs with Semantic Gaussian Processes
- FiNERweb: Datasets and Artifacts for Scalable Multilingual Named Entity Recognition
- Beyond surface form: A pipeline for semantic analysis in Alzheimer's Disease detection from spontaneous speech
- Semantic Grounding Index: Geometric Bounds on Context Engagement in RAG Systems
- Integrating Causal Reasoning into Automated Fact-Checking
- UCRBench: Benchmarking LLMs on Use Case Recovery
- Learning to Retrieve with Weakened Labels: Robust Training under Label Noise
- A Relational Model of Neighborhood Mobility: The Role of Amenities and Cultural Alignment
- Lemon: A Unified and Scalable 3D Multimodal Model for Universal Spatial Understanding
- CoRe3D: Collaborative Reasoning as a Foundation for 3D Intelligence
- Intelligent Scientific Literature Explorer using Machine Learning (ISLE)
- Tacit Understanding Game (TUG): Predicting Interpersonal Compatibility
- HyperEdit: Unlocking Instruction-based Text Editing in LLMs via Hypernetworks
- Semantic Distance Measurement based on Multi-Kernel Gaussian Processes
- Rethinking Label Consistency of In-Context Learning: An Implicit Transductive Label Propagation Perspective
- Rethinking Jailbreak Detection of Large Vision Language Models with Representational Contrastive Scoring
- PhraseVAE and PhraseLDM: Latent Diffusion for Full-Song Multitrack Symbolic Music Generation
- Extending a Parliamentary Corpus with MPs' Tweets: Automatic Annotation and Evaluation Using MultiParTweet
- Automating Historical Insight Extraction from Large-Scale Newspaper Archives via Neural Topic Modeling
- MultiScript30k: Leveraging Multilingual Embeddings to Extend Cross Script Parallel Data
- Interpretable and Steerable Concept Bottleneck Sparse Autoencoders
- GLOW: Graph-Language Co-Reasoning for Agentic Workflow Performance Prediction
- Semantic Reconstruction of Adversarial Plagiarism: A Context-Aware Framework for Detecting and Restoring "Tortured Phrases" in Scientific Literature
- AgriRegion: Region-Aware Retrieval for High-Fidelity Agricultural Advice
- Interpretation as Linear Transformation: A Cognitive-Geometric Model of Belief and Meaning
- Ontology-Based Knowledge Graph Framework for Industrial Standard Documents via Hierarchical and Propositional Structuring
- Semantic-Aware Cooperative Communication and Computation Framework in Vehicular Networks
- Refining Diffusion Models for Motion Synthesis with an Acceleration Loss to Generate Realistic IMU Data
- Beyond Traditional Diagnostics: Transforming Patient-Side Information into Predictive Insights with Knowledge Graphs and Prototypes
- How a Bit Becomes a Story: Semantic Steering via Differentiable Fault Injection
- Luxical: High-Speed Lexical-Dense Text Embeddings
- Bridging Code Graphs and Large Language Models for Better Code Understanding
- SkipKV: Selective Skipping of KV Generation and Storage for Efficient Inference with Large Reasoning Models
- Relational Visual Similarity
- When Large Language Models Do Not Work: Online Incivility Prediction through Graph Neural Networks
- MoCoRP: Modeling Consistent Relations between Persona and Response for Persona-based Dialogue
- MASim: Multilingual Agent-Based Simulation for Social Science
- PICKT: Practical Interlinked Concept Knowledge Tracing for Personalized Learning using Knowledge Map Concept Relations
- Cross-platform Product Matching Based on Entity Alignment of Knowledge Graph with RAEA model
- See Once, Then Act: Vision-Language-Action Model with Task Learning from One-Shot Video Demonstrations
- STAR-GO: Improving Protein Function Prediction by Learning to Hierarchically Integrate Ontology-Informed Semantic Embeddings
- LLM-Upgraded Graph Reinforcement Learning for Carbon-Aware Job Scheduling in Smart Manufacturing
- One Word Is Not Enough: Simple Prompts Improve Word Embeddings
- Distribution-Aware Exploration for Adaptive HNSW Search
- Prompting-in-a-Series: Psychology-Informed Contents and Embeddings for Personality Recognition With Decoder-Only Models
- Graph-Regularized Sparse Autoencoders for LLM Safety Steering
- TopiCLEAR: Topic extraction by CLustering Embeddings with Adaptive dimensional Reduction
- GradientSpace: Unsupervised Data Clustering for Improved Instruction Tuning
- MedGRPO: Multi-Task Reinforcement Learning for Heterogeneous Medical Video Understanding
- Knowing What's Missing: Assessing Information Sufficiency in Question Answering
- CommentScope: A Comment-Embedded Assisted Reading System for a Long Text
- Heard or Halted? Gender, Interruptions, and Emotional Tone in U.S. Supreme Court Oral Arguments
- The Road of Adaptive AI for Precision in Cybersecurity
- ResearchArcade: Graph Interface for Academic Tasks
- AfriStereo: A Culturally Grounded Dataset for Evaluating Stereotypical Bias in Large Language Models
- Mapping CVEs to MITRE ATT&CK Techniques: A Curated Gold-Set Classifier and the Limits of LLM-Assisted Label Expansion
- Verbalizing LLMs' assumptions to explain and control sycophancy
- OsmT: Bridging OpenStreetMap Queries and Natural Language with Open-source Tag-aware Language Models
- AdmTree: Compressing Lengthy Context with Adaptive Semantic Trees
- Distance Is All You Need: Radial Dispersion for Uncertainty Estimation in Large Language Models
- ClusterFusion: Hybrid Clustering with Embedding Guidance and LLM Adaptation
- ConsentDiff at Scale: Longitudinal Audits of Web Privacy Policy Changes and UI Frictions
- Peek-a-Boo Reasoning: Contrastive Region Masking in MLLMs
- LLM-Driven Data Generation and a Novel Soft Metric for Evaluating Text-to-SQL in Aviation MRO
- Overcoming State Inertia: Minimally Invasive Temporal Alignment for Evolving Contexts
- ProcObject-10K: Benchmarking Object-Centric Procedural Understanding in Instructional Videos
- Semantic Nutrition Estimation: Predicting Food Healthfulness from Text Descriptions
- Enhancing Job Matching: Occupation, Skill and Qualification Linking with the ESCO and EQF taxonomies
- GraphMatch: Fusing Language and Graph Representations in a Dynamic Two-Sided Work Marketplace
- SR-GRPO: Stable Rank as an Intrinsic Geometric Reward for Large Language Model Alignment
- Flowchart2Mermaid: A Vision-Language Model Powered System for Converting Flowcharts into Editable Diagram Code
- Ensemble Privacy Defense for Knowledge-Intensive LLMs against Membership Inference Attacks
- GPTrace: Effective Crash Deduplication Using LLM Embeddings
- LEC: Linear Expectation Constraints for Selection-Conditioned Risk Control in Selective Prediction and Routing Systems
- Label Forensics: Interpreting Hard Labels in Black-Box Text Classifier
- Code Comments for Quantum Software Development Kits: An Empirical Study on Qiskit
- Language-Guided Open-World Anomaly Segmentation
- MARSAD: A Multi-Functional Tool for Real-Time Social Media Analysis
- Patient Safety Risks from AI Scribes: Signals from End-User Feedback
- Advancing Academic Chatbots: Evaluation of Non Traditional Outputs
- Graph Queries from Natural Language using Constrained Language Models and Visual Editing
- WaterSearch: A Quality-Aware Search-based Watermarking Framework for Large Language Models
- Bias Injection Attacks on RAG Databases and Sanitization Defenses
- SAGE: Semantic-Aware Gray-Box Game Regression Testing with Large Language Models
- CourseTimeQA: A Lecture-Video Benchmark and a Latency-Constrained Cross-Modal Fusion Method for Timestamped QA
- Toward Automated and Trustworthy Scientific Analysis and Visualization with LLM-Generated Code
- Tracing How Annotators Think: Augmenting Preference Judgments with Reading Processes
- Learning to Prioritize IT Tickets: A Comparative Evaluation of Embedding-based Approaches and Fine-Tuned Transformer Models
- Action-guided generation of 3D functionality segmentation data
- Listwise Preference Optimization with Element-wise Confusions for Aspect Sentiment Quad Prediction
- Learning to Refuse: Refusal-Aware Reinforcement Fine-Tuning for Hard-Irrelevant Queries in Video Temporal Grounding
- A transfer learning approach for automatic conflicts detection in software requirement sentence pairs based on dual encoders
- Pooling Attention: Evaluating Pretrained Transformer Embeddings for Deception Classification
- Experts are all you need: A Composable Framework for Large Language Model Inference
- A Customer Journey in the Land of Oz: Leveraging the Wizard of Oz Technique to Model Emotions in Customer Service Interactions
- A Taxonomy-Driven Case Study of Australian Web Resources Against Technology-Facilitated Abuse
- Bandit Guided Submodular Curriculum for Adaptive Subset Selection
- Tackling a Challenging Corpus for Early Detection of Gambling Disorder: UNSL at MentalRiskES 2025
- Breaking the Illusion: Consensus-Based Generative Mitigation of Adversarial Illusions in Multi-Modal Embeddings
- Learning Programming in Informal Spaces: Using Emotion as a Lens to Understand Novice Struggles on r/learnprogramming
- Economies of Open Intelligence: Tracing Power & Participation in the Model Ecosystem
- Beyond Membership: Limitations of Add/Remove Adjacency in Differential Privacy
- From Compound Figures to Composite Understanding: Developing a Multi-Modal LLM from Biomedical Literature with Medical Multiple-Image Benchmarking and Validation
- Real-PGDN: A Two-level Classification Method for Full-Process Recognition of Newly Registered Pornographic and Gambling Domain Names
- PROMPTMINER: Black-Box Prompt Stealing against Text-to-Image Generative Models via Reinforcement Learning and Fuzz Optimization
- BAMAS: Structuring Budget-Aware Multi-Agent Systems
- Unsupervised Multimodal Graph-based Model for Geo-social Analysis
- MetaRank: Task-Aware Metric Selection for Model Transferability Estimation
- Subgoal Graph-Augmented Planning for LLM-Guided Open-World Reinforcement Learning
- Generating Querying Code from Text for Multi-Modal Electronic Health Record
- Unsupervised Memorability Modeling from Tip-of-the-Tongue Retrieval Queries
- SPHINX: A Synthetic Environment for Visual Perception and Reasoning
- Memories Retrieved from Many Paths: A Multi-Prefix Framework for Robust Detection of Training Data Leakage in Large Language Models
- Diverse Video Generation with Determinantal Point Process-Guided Policy Optimization
- Copyright Detection in Large Language Models: An Ethical Approach to Generative AI Development
- DesignPref: Capturing Personal Preferences in Visual Design Generation
- NormEval: A Unified Multi-Metric Framework for Evaluating Semantic Fidelity in Text Normalization
- A Systematic Analysis of Large Language Models with RAG-enabled Dynamic Prompting for Medical Error Detection and Correction
- Beyond Relational: Semantic-Aware Multi-Modal Analytics with LLM-Native Query Optimization
- Language-Independent Sentiment Labelling with Distant Supervision: A Case Study for English, Sepedi and Setswana
- Online-PVLM: Advancing Personalized VLMs with Online Concept Learning
- Interactive AI NPCs Powered by LLMs: Technical Report for the CPDC Challenge 2025
- Are Neuro-Inspired Multi-Modal Vision-Language Models Resilient to Membership Inference Privacy Leakage?
- Accuracy and Efficiency Trade-Offs in LLM-Based Malware Detection and Explanation: A Comparative Study of Parameter Tuning vs. Full Fine-Tuning
- Beyond Protein Language Models: An Agentic LLM Framework for Mechanistic Enzyme Design
- UISearch: Graph-Based Embeddings for Multimodal Enterprise UI Screenshots Retrieval
- FilmSceneDesigner: Chaining Set Design for Procedural Film Scene Generation
- ReEXplore: Improving MLLMs for Embodied Exploration with Contextualized Retrospective Experience Replay
- Benchmarking Corruption Robustness of LVLMs: A Discriminative Benchmark and Robustness Alignment Metric
- A Reproducible Framework for Neural Topic Modeling in Focus Group Analysis
- Yo'City: Personalized and Boundless 3D Realistic City Scene Generation via Self-Critic Expansion
- Comparative Analysis of LoRA-Adapted Embedding Models for Clinical Cardiology Text Representation
- Scalable Parameter-Light Spectral Method for Clustering Short Text Embeddings with a Cohesion-Based Evaluation Metric
- What Helps Language Models Predict Human Beliefs: Demographics or Prior Stances?
- From Reviewers' Lens: Understanding Bug Bounty Report Invalid Reasons with LLMs
- LLM Assisted Coding with Metamorphic Specification Mutation Agent
- Consolidating Diffusion-Generated Video Detection with Unified Multimodal Forgery Learning
- Paper2SysArch: Structure-Constrained System Architecture Generation from Scientific Papers
- GeeSanBhava: Sentiment Tagged Sinhala Music Video Comment Data Set
- Blu-WERP (Web Extraction and Refinement Pipeline): A Scalable Pipeline for Preprocessing Large Language Model Datasets
- EduMod-LLM: A Modular Approach for Designing Flexible and Transparent Educational Assistants
- M3-Bench: Multi-Modal, Multi-Hop, Multi-Threaded Tool-Using MLLM Agent Benchmark
- REMSA: Foundation Model Selection for Remote Sensing via a Constraint-Aware Agent
- SMILE: A Composite Lexical-Semantic Metric for Question-Answering Evaluation
- Don't Learn, Ground: A Case for Natural Language Inference with Visual Grounding
- CREST: Improving Interpretability and Effectiveness of Troubleshooting at Ericsson through Criterion-Specific Trouble Report Retrieval
- A Benchmark for Procedural Memory Retrieval in Language Agents
- Monte Carlo Expected Threat (MOCET) Scoring
- Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation
- Learning from Sufficient Rationales: Analysing the Relationship Between Explanation Faithfulness and Token-level Regularisation Strategies
- WER is Unaware: Assessing How ASR Errors Distort Clinical Understanding in Patient Facing Dialogue
- When Alignment Fails: Multimodal Adversarial Attacks on Vision-Language-Action Models
- Detecting Sleeper Agents in Large Language Models via Semantic Drift Analysis
- TurkColBERT: A Benchmark of Dense and Late-Interaction Models for Turkish Information Retrieval
- Two-Faced Social Agents: Context Collapse in Role-Conditioned Large Language Models
- Quality-Controlled Multimodal Emotion Recognition in Conversations with Identity-Based Transfer Learning and MAMBA Fusion
- Hierarchical Token Prepending: Enhancing Information Flow in Decoder-based LLM Embeddings
- Leveraging Digitized Newspapers to Collect Summarization Data in Low-Resource Languages
- Mitigating Label Length Bias in Large Language Models
- Technique to Baseline QE Artefact Generation Aligned to Quality Metrics
- RoboTidy : A 3D Gaussian Splatting Household Tidying Benchmark for Embodied Navigation and Action
- SweeperBot: Making 3D Browsing Accessible through View Analysis and Visual Question Answering
- Beyond Surface-Level Similarity: Hierarchical Contamination Detection for Synthetic Training Data in Foundation Models
- Retrieval-GRPO: A Multi-Objective Reinforcement Learning Framework for Dense Retrieval in Taobao Search
- Hierarchical Retrieval with Out-Of-Vocabulary Queries: A Case Study on SNOMED CT
- PolicyBot - Reliable Question Answering over Policy Documents
- Attention Grounded Enhancement for Visual Document Retrieval
- Mem-PAL: Towards Memory-based Personalized Dialogue Assistants for Long-term User-Agent Interaction
- Unifying points of interest taxonomies: mapping OpenStreetMap tags to the Foursquare category system
- Dual-LoRA and Quality-Enhanced Pseudo Replay for Multimodal Continual Food Learning
- EL3DD: Extended Latent 3D Diffusion for Language Conditioned Multitask Manipulation
- Uncovering and Mitigating Transient Blindness in Multimodal Model Editing
- Assessing Large Language Models in Generating RTL Design Specifications
- Synthetic Eggs in Many Baskets: The Impact of Synthetic Data Diversity on LLM Fine-Tuning
- Analyzing Sustainability Messaging in Large-Scale Corporate Social Media
- Multimodal Large Language Models as Image Classifiers
- Difficulty-Controllable Cloze Question Distractor Generation
- MedSumGraph: enhancing GraphRAG for medical QA with summarization and optimized prompts
- RLHF May Not Reflect Genuine Preferences
- CAT-ID2: Category-Tree Integrated Document Identifier Learning for Generative Retrieval In E-commerce
- Entailment as Few-Shot Learner
- ConneX: Automatically Resolving Transaction Opacity of Cross-Chain Bridges for Security Analysis
- RAGSmith: A Framework for Finding the Optimal Composition of Retrieval-Augmented Generation Methods Across Datasets
- Exploring question answering: metric analysis and evaluation framework for enhanced interpretability
- Evaluating Embedding Generalization: How LLMs, LoRA, and SLERP Shape Representational Geometry
- HEDGE: Hallucination Estimation via Dense Geometric Entropy for VQA with Vision-Language Models
- Hi-Reco: High-Fidelity Real-Time Conversational Digital Humans
- LLM-Powered Text-Attributed Graph Anomaly Detection via Retrieval-Augmented Reasoning
- Evolving Prompts for Toxicity Search in Large Language Models
- Prompt Engineering Techniques for Context-dependent Text-to-SQL in Arabic
- A Systematic Study of Model Extraction Attacks on Graph Foundation Models
- Correcting Mean Bias in Text Embeddings: A Refined Renormalization with Training-Free Improvements on MMTEB
- Draft and Refine with Visual Experts
- Generative Caching for Structurally Similar Prompts and Responses
- Cost Transparency of Enterprise AI Adoption
- Beyond Elicitation: Provision-based Prompt Optimization for Knowledge-Intensive Tasks
- Reasoning about Intent for Ambiguous Requests
- Local Hybrid Retrieval-Augmented Document QA
- OutSafe-Bench: A Benchmark for Multimodal Offensive Content Detection in Large Language Models
- ProgRAG: Hallucination-Resistant Progressive Retrieval and Reasoning over Knowledge Graphs
- A general framework for adaptive nonparametric dimensionality reduction
- PustakAI: Curriculum-Aligned and Interactive Textbooks Using Large Language Models
- TermGPT: Multi-Level Contrastive Fine-Tuning for Terminology Adaptation in Legal and Financial Domain
- Does Scientific Writing Converge to U.S. English? Evidence from Generative AI-Assisted Publications
- Contextual Graph Embeddings: Accounting for Data Characteristics in Heterogeneous Data Integration
- Improve Contrastive Clustering Performance by Multiple Fusing-Augmenting ViT Blocks
- Hierarchical Memorization in Large Language Models: Evidence from Citation Generation
- A centroid based framework for text classification in itsm environments
- Synergistic Feature Fusion for Latent Lyrical Classification: A Gated Deep Learning Architecture
- TurkEmbed: Turkish Embedding Model on NLI & STS Tasks
- DPRM: A Dual Implicit Process Reward Model in Multi-Hop Question Answering
- Prompt Tuning for Natural Language to SQL with Embedding Fine-Tuning and RAG
- Evaluating Gemini LLM in Food Image-Based Recipe and Nutrition Description with EfficientNet-B4 Visual Backbone
- Self-Correction Distillation for Structured Data Question Answering
- VSPO: Validating Semantic Pitfalls in Ontology via LLM-Based CQ Generation
- Unified Work Embeddings: Contrastive Learning of a Bidirectional Multi-task Ranker
- Last Layer Logits to Logic: Empowering LLMs with Logic-Consistent Structured Knowledge Reasoning
- Sparse3DPR: Training-Free 3D Hierarchical Scene Parsing and Task-Adaptive Subgraph Reasoning from Sparse RGB Views
- Stress Testing Factual Consistency Metrics for Long-Document Summarization
- Harmonic Token Projection (HTP): A Vocabulary-Free, Training-Free, Deterministic, and Reversible Embedding Methodology
- TurkEmbed4Retrieval: Turkish Embedding Model for Retrieval Task
- Beyond Fact Retrieval: Episodic Memory for RAG with Generative Semantic Workspaces
- Do LLMs Feel? Teaching Emotion Recognition with Prompts, Retrieval, and Curriculum Learning
- Oh That Looks Familiar: A Novel Similarity Measure for Spreadsheet Template Discovery
- Rethinking Retrieval-Augmented Generation for Medicine: A Large-Scale, Systematic Expert Evaluation and Practical Insights
- Characterizing AI Manipulation Risks in Brazilian YouTube Climate Discourse
- Optimizing Chain-of-Thought Confidence via Topological and Dirichlet Risk Analysis
- BookAsSumQA: An Evaluation Framework for Aspect-Based Book Summarization via Question Answering
- FLEX: Continuous Agent Evolution via Forward Learning from Experience
- NILC: Discovering New Intents with LLM-assisted Clustering
- Building Specialized Software-Assistant ChatBot with Graph-Based Retrieval-Augmented Generation
- Search Is Not Retrieval: Decoupling Semantic Matching from Contextual Assembly in RAG
- Leak@k: Unlearning Does Not Make LLMs Forget Under Probabilistic Decoding
- Long Grounded Thoughts: Synthesizing Visual Problems and Reasoning Chains at Scale
- REFLEX: Reference-Free Evaluation of Log Summarization via Large Language Model Judgment
- ReGen: Generative Robot Simulation via Inverse Design
- Probabilistic Textual Time Series Depression Detection
- Dynamic Jointly Batch Selection for Data Efficient Machine Translation Fine-Tuning
- Advancing Equitable AI: Evaluating Cultural Expressiveness in LLMs for Latin American Contexts
- E-CARE: An Efficient LLM-based Commonsense-Augmented Framework for E-Commerce
- ForeRobo: Unlocking Infinite Simulation Data for 3D Goal-driven Robotic Manipulation
- Watermarking Large Language Models in Europe: Interpreting the AI Act in Light of Technology
- Towards Transparent Stance Detection: A Zero-Shot Approach Using Implicit and Explicit Interpretability
- Multi-Object Tracking Retrieval with LLaVA-Video: A Training-Free Solution to MOT25-StAG Challenge
- Leveraging LLM-based agents for social science research: insights from citation network simulations
- Sparse, self-organizing ensembles of local kernels detect rare statistical anomalies
- Do Androids Dream of Unseen Puppeteers? Probing for a Conspiracy Mindset in Large Language Models
- ZoFia: Zero-Shot Fake News Detection with Entity-Guided Retrieval and Multi-LLM Interaction
- The Curved Spacetime of Transformer Architectures
- ROBoto2: An Interactive System and Dataset for LLM-assisted Clinical Trial Risk of Bias Assessment
- Cache Mechanism for Agent RAG Systems
- Smart-Hiring: An Explainable end-to-end Pipeline for CV Information Extraction and Job Matching
- Large Lemma Miners: Can LLMs do Induction Proofs for Hardware?
- Keeping it Local, Tiny and Real: Automated Report Generation on Edge Computing Devices for Mechatronic-Based Cognitive Systems
- ReAcTree: Hierarchical LLM Agent Trees with Control Flow for Long-Horizon Task Planning
- IG-Pruning: Input-Guided Block Pruning for Large Language Models
- NSYNC: Negative Synthetic Image Generation for Contrastive Training to Improve Stylized Text-To-Image Translation
- Trove: A Flexible Toolkit for Dense Retrieval
- Towards LLM-Powered Task-Aware Retrieval of Scientific Workflows for Galaxy
- Rescuing the Unpoisoned: Efficient Defense against Knowledge Corruption Attacks on RAG Systems
- OceanAI: A Conversational Platform for Accurate, Transparent, Near-Real-Time Oceanographic Insights
- SpEx: A Spectral Approach to Explainable Clustering
- OmniBrainBench: A Comprehensive Multimodal Benchmark for Brain Imaging Analysis Across Multi-stage Clinical Tasks
- Teaching LLMs to See and Guide: Context-Aware Real-Time Assistance in Augmented Reality
- Do Methods to Jailbreak and Defend LLMs Generalize Across Languages?
- Separate the Wheat from the Chaff: Winnowing Down Divergent Views in Retrieval Augmented Generation
- Issue-Oriented Agent-Based Framework for Automated Review Comment Generation
- G2: Guided Generation for Enhanced Output Diversity in LLMs
- PreferThinker: Reasoning-based Personalized Image Preference Assessment
- LIR: The First Workshop on Late Interaction and Multi Vector Retrieval @ ECIR 2026
- LingGym: How Far Are LLMs from Thinking Like Field Linguists?
- IL-PCSR: Legal Corpus for Prior Case and Statute Retrieval
- Culture Cartography: Mapping the Landscape of Cultural Knowledge
- Effect of Domain Generalization Techniques in Low Resource Systems
- Thought Branches: Interpreting LLM Reasoning Requires Resampling
- Traceable Drug Recommendation over Medical Knowledge Graphs
- Relation-Aware Bayesian Optimization of DBMS Configurations Guided by Affinity Scores
- A Memory-Efficient Retrieval Architecture for RAG-Enabled Wearable Medical LLMs-Agents
- Addressing Longstanding Challenges in Cognitive Science with Language Models
- Independent Clinical Evaluation of General-Purpose LLM Responses to Signals of Suicide Risk
- How Similar Are Grokipedia and Wikipedia? A Multi-Dimensional Textual and Structural Comparison
- Cross-Platform Evaluation of Reasoning Capabilities in Foundation Models
- Value Drifts: Tracing Value Alignment During LLM Post-Training
- Hebrew Diacritics Restoration using Visual Representation
- SCRIBE: Structured Chain Reasoning for Interactive Behaviour Explanations using Tool Calling
- Graph-Enhanced Policy Optimization in LLM Agent Training
- Similarity-Distance-Magnitude Language Models
- Reasoning Path Divergence: A New Metric and Curation Strategy to Unlock LLM Diverse Thinking
- Agentic Economic Modeling
- Counterfactual-based Agent Influence Ranker for Agentic AI Workflows
- Hallucinations in Bibliographic Recommendation: Citation Frequency as a Proxy for Training Data Redundancy
- LLM-as-a-Judge for Evaluating System Responses in Conversational Music Recommendation
- Nudging Sustainable Choices through LLM-Generated Recommendation Explanations
- Language Through a Prism: A Spectral Approach for Multiscale Language Representations
- Diagnosing Fine-Grained Inconsistency Classification in Financial Disclosure Text
- Voice Memory for Agentic Speech Recognition
- Continuous Online Evaluation of Recommendation Strategies in Social Science Academic Search
- CaIRec: Calibrated Modality Imputation for Incomplete Multimodal Recommendation
- RAGuard: A Layered Defense Framework for Retrieval-Augmented Generation Systems Against Data Poisoning
- Constitutional Midtraining: Content Presence Drives Alignment Gains
- CrisisBERT: a Robust Transformer for Crisis Classification and Contextual Crisis Embedding
- IR-BERT: Leveraging BERT for Semantic Search in Background Linking for News Articles
- Supporting Workflow Reproducibility by Linking Bioinformatics Tools across Papers and Executable Code
- Tools Are Not Islands: Set-Level Tool Retrieval for LLM Agents via Query-Conditioned Hyperedge Prediction
- RedCodeAgent: Automatic Red-teaming Agent against Diverse Code Agents
- IMFuse: Instance-Aware Multi-Layer Fusion for LLM-Enhanced Sequential Recommendation
- Ideas Have Genomes: Benchmarking Scientific Lineage Reasoning and Lineage-Grounded Idea Generation
- Improving Human-Robot Teamwork in Urban Search and Rescue Through Episodic Memory of Prior Collaboration
- Do Methods Support the Claims? Intra-Paper Verification for Peer Review
- Can neurons speak? Semantic narration of vision at single-cell resolution
- Grammar as a behavioral biometric: using cognitively motivated grammar models for authorship verification
- EUDAIMONIA: Evaluating Undesirable Dynamics in AI
- Enhancing Multi-Agent Communication through Attention Steering with Context Relevance
- Is Dimensionality a Barrier for Retrieval Models?
- Platonic Representations in the Human Brain: Unsupervised Recovery of Universal Geometry
- GraphLAMA: Enabling Efficient Adaptation of Graph Language Models with Limited Annotations
- Can Knowledge-Graph-based Retrieval Augmented Generation Really Retrieve What You Need?
- Hallucination Benchmark for Speech Foundation Models
- Does Socialization Emerge in AI Agent Society? A Case Study of Moltbook
- The Rise of AI Agent Communities: Large-Scale Analysis of Discourse and Interaction on Moltbook
- From Reviews to Actionable Insights: An LLM-Based Approach for Attribute and Feature Extraction
- Readability-Robust Code Summarization via Meta Curriculum Learning
- Better Call Grep: Evaluating and Improving Grep-Like Lexical Retrieval for Repository-Level Code Completion
- ReviewSense: Transforming Customer Review Dynamics into Actionable Business Insights
- Depression Status Estimation by Deep Learning based Hybrid Multi-Modal Fusion Model
- Probing the Hidden Talent of ASR Foundation Models for L2 English Oral Assessment
- Self-Supervised Text-Vision Alignment for Automated Brain MRI Abnormality Detection: A Multicenter Study (ALIGN Study)
- Beyond One-Size-Fits-All: Personalized Harmful Content Detection with In-Context Learning
- DIRC-RAG: Accelerating Edge RAG with Robust High-Density and High-Loading-Bandwidth Digital In-ReRAM Computation
- Optimizing Knowledge Utilization for Multi-Intent Comment Generation with Large Language Models
- Testing Cross-Lingual Text Comprehension In LLMs Using Next Sentence Prediction
- SemCoT: Accelerating Chain-of-Thought Reasoning through Semantically-Aligned Implicit Tokens
- Iterative Critique-Refine Framework for Enhancing LLM Personalization
- Evaluating Joinable Column Discovery Approaches for Context-Aware Search
- Detecting the Use of Generative AI in Crowdsourced Surveys: Implications for Data Integrity
- Politically Speaking: LLMs on Changing International Affairs
- From Cross-Task Examples to In-Task Prompts: A Graph-Based Pseudo-Labeling Framework for In-context Learning
- Talk2Ref: A Dataset for Reference Prediction from Scientific Talks
- Metadata-Driven Retrieval-Augmented Generation for Financial Question Answering
- Text Simplification with Sentence Embeddings
- From Observability Data to Diagnosis: An Evolving Multi-agent System for Incident Management in Cloud Systems
- Utilising Large Language Models for Generating Effective Counter Arguments to Anti-Vaccine Tweets
- DynaStride: Dynamic Stride Windowing with MMCoT for Instructional Multi-Scene Captioning
- AfriMTEB and AfriE5: Benchmarking and Adapting Text Embedding Models for African Languages
- Beyond Understanding: Evaluating the Pragmatic Gap in LLMs' Cultural Processing of Figurative Language
- Small Language Models Offer Significant Potential for Science Community
- PixelRefer: A Unified Framework for Spatio-Temporal Object Referring with Arbitrary Granularity
- Minimizing Human Intervention in Online Classification
- COOPERA: Continual Open-Ended Human-Robot Assistance
- Evaluating Large Language Models for Stance Detection on Financial Targets from SEC Filing Reports and Earnings Call Transcripts
- Code Contribution and Credit in Science
- SwiftEmbed: Ultra-Fast Text Embeddings via Static Token Lookup for Real-Time Applications
- LangLingual: A Personalised, Exercise-oriented English Language Learning Tool Leveraging Large Language Models
- Modeling Political Discourse with Sentence-BERT and BERTopic
- Seeing the Unseen: Towards Training-Free Inspection for Wind Turbine Blades Using Knowledge-Augmented Vision Language Models
- Agentic Meta-Orchestrator for Multi-task Copilots
- Iterative Layer Pruning for Efficient Translation Inference
- Multi-Modal Fact-Verification Framework for Reducing Hallucinations in Large Language Models
- REVISION:Reflective Intent Mining and Online Reasoning Auxiliary for E-commerce Visual Search System Optimization
- RaCoT: Plug-and-Play Contrastive Example Generation Mechanism for Enhanced LLM Reasoning Reliability
- CLIN-LLM: A Safety-Constrained Hybrid Framework for Clinical Diagnosis and Treatment Generation
- Benchmarking Egocentric Multimodal Goal Inference for Assistive Wearable Agents
- Knowledge-guided Continual Learning for Behavioral Analytics Systems
- Multimodal Item Scoring for Natural Language Recommendation via Gaussian Process Regression with LLM Relevance Judgments
- TagRuler: Interactive Tool for Span-Level Data Programming by Demonstration
- LLM-Generated Negative News Headlines Dataset: Creation and Benchmarking Against Real Journalism
- Gaze-VLM:Bridging Gaze and VLMs through Attention Regularization for Egocentric Understanding
- Large Language Models Meet Text-Attributed Graphs: A Survey of Integration Frameworks and Applications
- NoisyGRPO: Incentivizing Multimodal CoT Reasoning via Noise Injection and Bayesian Estimation
- Designing and Evaluating Hint Generation Systems for Science Education
- Dynamic Retriever for In-Context Knowledge Editing via Policy Optimization
- Preventing Catastrophic Forgetting: Behavior-Aware Sampling for Safer Language Model Fine-Tuning
- FicSim: A Dataset for Multi-Faceted Semantic Similarity in Long-Form Fiction
- Thought Communication in Multiagent Collaboration
- Structure-Conditional Minimum Bayes Risk Decoding
- Systematic Evaluation of Uncertainty Estimation Methods in Large Language Models
- Citation Failure: Definition, Analysis and Efficient Mitigation
- Rank-GRPO: Training LLM-based Conversational Recommender Systems with Reinforcement Learning
- Tri-Modal Severity Fused Diagnosis across Depression and Post-traumatic Stress Disorders
- RECALL: REpresentation-aligned Catastrophic-forgetting ALLeviation via Hierarchical Model Merging
- Re-evaluating Minimum Bayes Risk Decoding for Automatic Speech Recognition
- Dialogue Is Not Enough to Make a Communicative BabyLM (But Neither Is Developmentally Inspired Reinforcement Learning)
- Do Prompts Reshape Representations? An Empirical Study of Prompting Effects on Embeddings
- CircuitGuard: Mitigating LLM Memorization in RTL Code Generation Against IP Leakage
- Style Attack Disguise: When Fonts Become a Camouflage for Adversarial Intent
- NeSyPr: Neurosymbolic Proceduralization For Efficient Embodied Reasoning
- Learning Noise-Resilient and Transferable Graph-Text Alignment via Dynamic Quality Assessment
- Sign Language Translation with Sentence Embedding Supervision
- From Script to Stage: Automating Experimental Design for Social Simulations with LLMs
- JointCQ: Improving Factual Hallucination Detection with Joint Claim and Query Generation
- Selecting and Combining Large Language Models for Scalable Code Clone Detection
- Human-Agent Collaborative Paper-to-Page Crafting
- LLM-Augmented Symbolic NLU System for More Reliable Continuous Causal Statement Interpretation
- C2T-ID: Converting Semantic Codebooks to Textual Document Identifiers for Generative Search
- FlexiDataGen: An Adaptive LLM Framework for Dynamic Semantic Dataset Generation in Sensitive Domains
- SBAN: A Framework & Multi-Dimensional Dataset for Large Language Model Pre-Training and Software Code Mining
- FeClustRE: Hierarchical Clustering and Semantic Tagging of App Features from User Reviews
- Topoformer: brain-like topographic organization in Transformer language models through spatial querying and reweighting
- Beyond the Explicit: A Bilingual Dataset for Dehumanization Detection in Social Media
- One Size Fits All? A Modular Adaptive Sanitization Kit (MASK) for Customizable Privacy-Preserving Phone Scam Detection
- Unifying Inductive, Cross-Domain, and Multimodal Learning for Robust and Generalizable Recommendation
- IMB: An Italian Medical Benchmark for Question Answering
- PP3D: An In-Browser Vision-Based Defense Against Web Behavior Manipulation Attacks
- Med-VRAgent: A Framework for Medical Visual Reasoning-Enhanced Agents
- Evaluating LLM-Based Mobile App Recommendations: An Empirical Study
- KrishokBondhu: A Retrieval-Augmented Voice-Based Agricultural Advisory Call Center for Bengali Farmers
- PoSh: Using Scene Graphs To Guide LLMs-as-a-Judge For Detailed Image Descriptions
- Improving Topic Modeling of Social Media Short Texts with Rephrasing: A Case Study of COVID-19 Related Tweets
- ECKO: Explainable Clinical Knowledge for Oncology
- Learning from Generalization Patterns: An Evaluation-Driven Approach to Enhanced Data Augmentation for Fine-Tuning Small Language Models
- Exemplar-Guided Planing: Enhanced LLM Agent for KGQA
- CrossGuard: Safeguarding MLLMs against Joint-Modal Implicit Malicious Attacks
- AtlasKV: Augmenting LLMs with Billion-Scale Knowledge Graphs in 20GB VRAM
- AFRICAPTION: Establishing a New Paradigm for Image Captioning in African Languages
- Efficient Toxicity Detection in Gaming Chats: A Comparative Study of Embeddings, Fine-Tuned Transformers and LLMs
- TaxoAlign: Scholarly Taxonomy Generation Using Language Models
- StreamingThinker: Large Language Models Can Think While Reading
- OncoReason: Structuring Clinical Reasoning in LLMs for Robust and Interpretable Survival Prediction
- Who's Asking? Simulating Role-Based Questions for Conversational AI Evaluation
- Exact Nearest-Neighbor Search on Energy-Efficient FPGA Devices
- MOSAIC: Masked Objective with Selective Adaptation for In-domain Contrastive Learning
- LLM-based In-situ Thought Exchanges for Critical Paper Reading
- BiMax: Bidirectional MaxSim Score for Document-Level Alignment
- GraphMind: Interactive Novelty Assessment System for Accelerating Scientific Discovery
- ProofBridge: Auto-Formalization of Natural Language Proofs in Lean via Joint Embeddings
- Mixture of Experts Approaches in Dense Retrieval Tasks
- Leveraging LLMs for Context-Aware Implicit Textual and Multimodal Hate Speech Detection
- Iterative Topic Taxonomy Induction with LLMs: A Case Study of Electoral Advertising
- DMRetriever: A Family of Models for Improved Text Retrieval in Disaster Management
- LoRAverse: A Submodular Framework to Retrieve Diverse Adapters for Diffusion Models
- LLMs as Scalable, General-Purpose Simulators For Evolving Digital Agent Training
- AI-Powered Early Diagnosis of Mental Health Disorders from Real-World Clinical Conversations
- TRI-DEP: A Trimodal Comparative Study for Depression Detection Using Speech, Text, and EEG
- Harmonizing Diverse Models: A Layer-wise Merging Strategy for Consistent Generation
- Leveraging Multimodal LLM Descriptions of Activity for Explainable Semi-Supervised Video Anomaly Detection
- DEXTER: Diffusion-Guided EXplanations with TExtual Reasoning for Vision Models
- Intent Clustering with Shared Pseudo-Labels
- Multimodal RAG for Unstructured Data:Leveraging Modality-Aware Knowledge Graphs with Hybrid Retrieval
- Stealthy Dual-Trigger Backdoors: Attacking Prompt Tuning in LM-Empowered Graph Foundation Models
- Holdout-Loss-Based Data Selection for LLM Finetuning via In-Context Learning
- Hierarchical Semantic Retrieval with Cobweb
- DPRF: A Generalizable Dynamic Persona Refinement Framework for Optimizing Behavior Alignment Between Personalized LLM Role-Playing Agents and Humans
- JEDA: Query-Free Clinical Order Search from Ambient Dialogues
- DROID: Dual Representation for Out-of-Scope Intent Detection
- When Embedding Models Meet: Procrustes Bounds and Applications
- ReMindRAG: Low-Cost LLM-Guided Knowledge Graph Traversal for Efficient RAG
- Putting on the Thinking Hats: A Survey on Chain of Thought Fine-tuning from the Perspective of Human Reasoning Mechanism
- Stable LLM Ensemble: Interaction between Example Representativeness and Diversity
- Revisiting Query Variants: The Advantage of Retrieval Over Generation of Query Variants for Effective QPP
- Program of Thoughts for Financial Reasoning: Leveraging Dynamic In-Context Examples and Generative Retrieval
- When Personalization Tricks Detectors: The Feature-Inversion Trap in Machine-Generated Text Detection
- Beating Harmful Stereotypes Through Facts: RAG-based Counter-speech Generation
- PromptLocate: Localizing Prompt Injection Attacks
- Unveiling the Vulnerability of Graph-LLMs: An Interpretable Multi-Dimensional Adversarial Attack on TAGs
- GOAT: A Training Framework for Goal-Oriented Agent with Tools
- Encapsulating Textual Contents into a MOC data Structure for Advanced Applications
- GRAVITY: A Framework for Personalized Text Generation via Profile-Grounded Synthetic Preferences
- Scaling Language-Centric Omnimodal Representation Learning
- REGENT: Relevance-Guided Attention for Entity-Aware Multi-Vector Neural Re-Ranking
- QDER: Query-Specific Document and Entity Representations for Multi-Vector Document Re-Ranking
- FinVet: A Collaborative Framework of RAG and External Fact-Checking Agents for Financial Misinformation Detection
- Situat3DChange: Situated 3D Change Understanding Dataset for Multimodal Large Language Model
- Automated Skill Decomposition Meets Expert Ontologies: Bridging the Granularity Gap with LLMs
- FOSSIL: Harnessing Feedback on Suboptimal Samples for Data-Efficient Generalisation with Imitation Learning for Embodied Vision-and-Language Tasks
- Do Psychometric Tests Work for Large Language Models? Evaluation of Tests on Sexism, Racism, and Morality
- RAG-Pull: Imperceptible Attacks on RAG Systems for Code Generation
- Secret-Protected Evolution for Differentially Private Synthetic Text Generation
- Chart-RVR: Reinforcement Learning with Verifiable Rewards for Explainable Chart Reasoning
- Learning to Watermark: A Selective Watermarking Framework for Large Language Models via Multi-Objective Optimization
- Scalable and Explainable Enterprise Knowledge Discovery Using Graph-Centric Hybrid Retrieval
- Quantum NLP models on Natural Language Inference
- Detecting Hallucinations in Authentic LLM-Human Interactions
- Testing and Enhancing Multi-Agent Systems for Robust Code Generation
- NIM: Neuro-symbolic Ideographic Metalanguage for Inclusive Communication
- Steering Over-refusals Towards Safety in Retrieval Augmented Generation
- Knowing Unknowns in an Age of Information Overload
- PrediQL: Automated Testing of GraphQL APIs with LLMs
- SimKey: A Semantically Aware Key Module for Watermarking Language Models
- Are LLMs Empathetic to All? Investigating the Influence of Multi-Demographic Personas on a Model's Empathy
- Diversity Augmentation of Dynamic User Preference Data for Boosting Personalized Text Summarizers
- Are LLMs Better GNN Helpers? Rethinking Robust Graph Learning under Deficiencies with Iterative Refinement
- Comparison of Unsupervised Metrics for Evaluating Judicial Decision Extraction
- SkillOS: Learning Skill Curation for Self-Evolving Agents
- The Geometry of Forgetting
- Text2Token: Unsupervised Text Representation Learning with Token Target Prediction
- Evolution of wartime discourse on Telegram: A comparative study of Ukrainian and Russian policymakers' communication before and after Russia's full-scale invasion of Ukraine
- HUME: Measuring the Human-Model Performance Gap in Text Embedding Tasks
- iBERT: Interpretable Style Embeddings via Sense Decomposition
- Mapping Semantic & Syntactic Relationships with Geometric Rotation
- From Birdwatch to Community Notes, from Twitter to X: four years of community-based content moderation
- SIMBA UQ: Similarity-Based Aggregation for Uncertainty Quantification in Large Language Models
- Doc2Query++: Topic-Coverage based Document Expansion and its Application to Dense Retrieval via Dual-Index Fusion
- Can We Reliably Rank Model Performance across Domains without Labeled Data?
- A Living Review Pipeline for AI/ML Applications in Accelerator Physics
- CFVBench: A Comprehensive Video Benchmark for Fine-grained Multimodal Retrieval-Augmented Generation
- Maple: A Multi-agent System for Portable Deep Learning across Clusters
- When Retrieval Succeeds and Fails: Rethinking Retrieval-Augmented Generation for LLMs
- FrameEOL: Semantic Frame Induction using Causal Language Models
- Semantic-Condition Tuning: Fusing Graph Context with Large Language Models for Knowledge Graph Completion
- A Human Behavioral Baseline for Collective Governance in Software Projects
- When LLM Agents Meet Graph Optimization: An Automated Data Quality Improvement Approach
- GRETEL: A Goal-driven Retrieval and Execution-based Trial Framework for LLM Tool Selection Enhancing
- ProxRouter: Proximity-Weighted LLM Query Routing for Improved Robustness to Outliers
- One Sentence, Two Embeddings: Contrastive Learning of Explicit and Implicit Semantic Representations
- DeepPrune: Parallel Scaling without Inter-trace Redundancy
- Leveraging Whisper Embeddings for Audio-based Lyrics Matching
- HySim-LLM: Embedding-Weighted Fine-Tuning Bounds and Manifold Denoising for Domain-Adapted LLMs
- From Keywords to Clusters: AI-Driven Analysis of YouTube Comments to Reveal Election Issue Salience in 2024
- Multilingual Generative Retrieval via Cross-lingual Semantic Compression
- FedBook: A Unified Federated Graph Foundation Codebook with Intra-domain and Inter-domain Knowledge Modeling
- MemWeaver: A Hierarchical Memory from Textual Interactive Behaviors for Personalized Generation
- Safely Exploring Novel Actions in Recommender Systems via Deployment-Efficient Policy Learning
- Struc-EMB: The Potential of Structure-Aware Encoding in Language Embeddings
- RAG4Tickets: AI-Powered Ticket Resolution via Retrieval-Augmented Generation on JIRA and GitHub Data
- Self-Improving LLM Agents at Test-Time
- ZeroCard: Cardinality Estimation with Zero Dependence on Target Databases -- No Data, No Query, No Retraining
- Measuring the Hidden Cost of Data Valuation through Collective Disclosure
- On the Alignment Between Supervised and Self-Supervised Contrastive Learning
- Efficient Prompt Optimisation for Legal Text Classification with Proxy Prompt Evaluator
- Multi-Task Pre-Finetuning of Lightweight Transformer Encoders for Text Classification and NER
- Investigating Thematic Patterns and User Preferences in LLM Interactions using BERTopic
- When Thoughts Meet Facts: Reusable Reasoning for Long-Context LMs
- Bridged Clustering: Semi-Supervised Sparse Bridging
- Reasoning for Hierarchical Text Classification: The Case of Patents
- Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications
- Search-R3: Unifying Reasoning and Embedding in Large Language Models
- SoftMatcha 2: A Fast and Soft Pattern Matcher for Trillion-Scale Corpora
- A Comparison of Independent and Joint Fine-tuning Strategies for Retrieval-Augmented Generation
- Mapping global bee research with traits and plant-pollinator interaction networks
- ImageNet-Think-250K: A Large-Scale Synthetic Dataset for Multimodal Reasoning for Vision Language Models
- Iterative design of a NAND hybrid riboswitch by deep batch Bayesian optimization
- Exposing Citation Vulnerabilities in Generative Engines
- Text2Stories: Evaluating the Alignment Between Stakeholder Interviews and Generated User Stories
- Study on LLMs for Promptagator-Style Dense Retriever Training
- TWIST: Training-free and Label-free Short Text Clustering through Iterative Vector Updating with LLMs
- PTEB: Towards Robust Text Embedding Evaluation via Stochastic Paraphrasing at Evaluation Time with LLMs
- Differentially Private Synthetic Text Generation for Retrieval-Augmented Generation (RAG)
- Auto-Stega: An Agent-Driven System for Lifelong Strategy Evolution in LLM-Based Text Steganography
- How Confident are Video Models? Empowering Video Models to Express their Uncertainty
- A Framework for Measuring How News Topics Drive Stock Movement
- Controllable Stylistic Text Generation with Train-Time Attribute-Regularized Diffusion
- Relational Transformer: Toward Zero-Shot Foundation Models for Relational Data
- RoSE: Round-robin Synthetic Data Evaluation for Selecting LLM Generators without Human Test Sets
- Mnemosyne: An Unsupervised, Human-Inspired Long-Term Memory Architecture for Edge-Based LLMs
- ARRC: Advanced Reasoning Robot Control - Knowledge-Driven Autonomous Manipulation Using Retrieval-Augmented Generation
- Automated Research Article Classification and Recommendation Using NLP and ML
- Redefining Cost Estimation in Database Systems: The Role of Execution Plan Features and Machine Learning
- Scalable In-context Ranking with Generative Models
- DeepV: A Model-Agnostic Retrieval-Augmented Framework for Verilog Code Generation with a High-Quality Knowledge Base
- Guided Query Refinement: Multimodal Hybrid Retrieval with Test-Time Optimization
- ModernBERT + ColBERT: Enhancing biomedical RAG through an advanced re-ranking retriever
- Topic-Specific Classifiers are Better Relevance Judges than Prompted LLMs
- Contrastive Learning Using Graph Embeddings for Domain Adaptation of Language Models in the Process Industry
- Fine-grained auxiliary learning for real-world product recommendation
- Residualized Similarity for Faithfully Explainable Authorship Verification
- AWARE, Beyond Sentence Boundaries: A Contextual Transformer Framework for Identifying Cultural Capital in STEM Narratives
- GRACE: Generative Representation Learning via Contrastive Policy Optimization
- Challenge on Optimization of Context Collection for Code Completion
- Learning Representations Through Contrastive Neural Model Checking
- Chronological Thinking in Full-Duplex Spoken Dialogue Language Models
- RLRF: Competitive Search Agent Design via Reinforcement Learning from Ranker Feedback
- LLM Microscope: What Model Internals Reveal About Answer Correctness and Context Utilization
- SECA: Semantically Equivalent and Coherent Attacks for Eliciting LLM Hallucinations
- Systematic Diagnosis of Brittle Reasoning in Large Language Models
- How Catastrophic is Your LLM? Certifying Risk in Conversation
- MetaMuse: Algorithm Generation via Creative Ideation
- TreePrompt: Leveraging Hierarchical Few-Shot Example Selection for Improved English-Persian and English-German Translation
- Generating High-Level Test Cases from Requirements using LLM: An Industry Study
- Triplet-Structured Knowledge Integration for Multi-Turn Medical Reasoning
- Consistent Kernel Change-Point Detection under m-Dependence for Text Segmentation
- External Data Extraction Attacks against Retrieval-Augmented Large Language Models
- Transparent Reference-free Automated Evaluation of Open-Ended User Survey Responses
- Time-To-Inconsistency: A Survival Analysis of Large Language Model Robustness to Adversarial Attacks
- LegalWiz: A Multi-Agent Generation Framework for Contradiction Detection in Legal Documents
- SEER: The Span-based Emotion Evidence Retrieval Benchmark
- Learning Efficient Guardrails for Compliance
- ModernVBERT: Towards Smaller Visual Document Retrievers
- LLM Routing with Dueling Feedback
- MultiPhysio-HRC: Multimodal Physiological Signals Dataset for industrial Human-Robot Collaboration
- PolyLink: A Blockchain Based Decentralized Edge AI Platform for LLM Inference
- JoyAgent-JDGenie: Technical Report on the GAIA
- Retrieval and Augmentation of Domain Knowledge for Text-to-SQL Semantic Parsing
- TokMem: Tokenized Procedural Memory for Large Language Models
- Learning Compact Representations of LLM Abilities via Item Response Theory
- RealClass: A Framework for Classroom Speech Simulation with Public Datasets and Game Engines
- Milco: Learned Sparse Retrieval Across Languages via a Multilingual Connector
- Stochastic Self-Organization in Multi-Agent Systems
- Learning to Route: A Rule-Driven Agent Framework for Hybrid-Source Retrieval-Augmented Generation
- Automated Alignment of Math Items to Content Standards in Large-Scale Assessments Using Language Models
- PrimeX: A Dataset of Worldview, Opinion, and Explanation
- Automatic Fact-checking in English and Telugu
- MotionRAG: Motion Retrieval-Augmented Image-to-Video Generation
- An Experimental Study on Generating Plausible Textual Explanations for Video Summarization
- CliniBench: A Clinical Outcome Prediction Benchmark for Generative and Encoder-Based Language Models
- RAE: A Neural Network Dimensionality Reduction Method for Nearest Neighbors Preservation in Vector Search
- Better with Less: Small Proprietary Models Surpass Large Language Models in Financial Transaction Understanding
- V-HUB: A Visual-Centric Humor Understanding Benchmark for Video LLMs
- Think Less, Label Better: Multi-Stage Domain-Grounded Synthetic Data Generation for Fine-Tuning Large Language Models in Telecommunications
- CustomIR: Unsupervised Fine-Tuning of Dense Embeddings for Known Document Corpora
- SafePassage: High-Fidelity Information Extraction with Black Box LLMs
- Investigating Language and Retrieval Bias in Multilingual Previously Fact-Checked Claim Detection
- How Well Do LLMs Imitate Human Writing Style?
- Of-SemWat: High-payload text embedding for semantic watermarking of AI-generated images with arbitrary size
- HarmMetric Eval: Benchmarking Metrics and Judges for LLM Harmfulness Assessment
- Fin-Ally: Pioneering the Development of an Advanced, Commonsense-Embedded Conversational AI for Money Matters
- ELASTIQ: EEG-Language Alignment with Semantic Task Instruction and Querying
- Let LLMs Speak Embedding Languages: Generative Text Embeddings via Iterative Contrastive Refinement
- ViReSkill: Vision-Grounded Replanning with Skill Memory for LLM-Based Planning in Lifelong Robot Learning
- Model Correlation Detection via Random Selection Probing
- Memory Transfer Planning: LLM-driven Context-Aware Code Adaptation for Robot Manipulation
- GeoVLM-R1: Reinforcement Fine-Tuning for Improved Remote Sensing Reasoning
- Pretraining with hierarchical memories: separating long-tail and common knowledge
- Assessing Large Language Models in Updating Their Forecasts with New Information
- AnveshanaAI: A Multimodal Platform for Adaptive AI/ML Education through Automated Question Generation and Interactive Assessment
- Semantic Representation of Processes with Ontology Design Patterns
- Enhancing LLM Steering through Sparse Autoencoder-Based Vector Refinement
- An Senegalese Legal Texts Structuration Using LLM-augmented Knowledge Graph
- BioArtlas: Computational Clustering of Multi-Dimensional Complexity in Bioart
- Detecting Escalation Level from Speech with Transfer Learning and Acoustic-Lexical Information Fusion
- LLM Watermark Evasion via Bias Inversion
- Open-Vocabulary Spatio-Temporal Scene Graph for Robot Perception and Teleoperation Planning
- Semantic Voting: A Self-Evaluation-Free Approach for Efficient LLM Self-Improvement on Unverifiable Open-ended Tasks
- The Geometry of Creative Variability: How Credal Sets Expose Calibration Gaps in Language Models
- Comparison of Scoring Rationales Between Large Language Models and Human Raters
- From Deferral to Learning: Online In-Context Knowledge Distillation for LLM Cascades
- JE-IRT: A Geometric Lens on LLM Abilities through Joint Embedding Item Response Theory
- "I Don't Think RAI Applies to My Model'' -- Engaging Non-champions with Sticky Stories for Responsible AI Work
- REMA: A Unified Reasoning Manifold Framework for Interpreting Large Language Model
- Representing LLMs in Prompt Semantic Task Space
- Jailbreaking on Text-to-Video Models via Scene Splitting Strategy
- Question-Driven Analysis and Synthesis: Building Interpretable Thematic Trees with LLMs for Text Clustering and Controllable Generation
- Library Hallucinations in LLMs: Risk Analysis Grounded in Developer Queries
- Context Parametrization with Compositional Adapters
- Think Right, Not More: Test-Time Scaling for Numerical Claim Verification
- Goal-Guided Efficient Exploration via Large Language Model in Reinforcement Learning
- MotivGraph-SoIQ: Integrating Motivational Knowledge Graphs and Socratic Dialogue for Enhanced LLM Ideation
- Semantic Agreement Enables Efficient Open-Ended LLM Cascades
- RobustFlow: Towards Robust Agentic Workflow Generation
- KurdSTS: The Kurdish Semantic Textual Similarity
- Does AI Coaching Prepare us for Workplace Negotiations?
- GRAB: A Risk Taxonomy--Grounded Benchmark for Unsupervised Topic Discovery in Financial Disclosures
- What Should I Cite? A RAG Benchmark for Academic Citation Prediction
- The GDN-CC Dataset: Automatic Corpus Clarification for AI-enhanced Democratic Citizen Consultations
- MMPlanner: Zero-Shot Multimodal Procedural Planning with Chain-of-Thought Object State Reasoning
- QuantMind: A Context-Engineering Based Knowledge Framework for Quantitative Finance
- Filtering with Confidence: When Data Augmentation Meets Conformal Prediction
- One Model, Many Morals: Uncovering Cross-Linguistic Misalignments in Computational Moral Reasoning
- Position: Human Factors Reshape Adversarial Analysis in Human-AI Decision-Making Systems
- Interactive Recommendation Agent with Active User Commands
- Semantic Clustering of Civic Proposals: A Case Study on Brazil's National Participation Platform
- Query-Centric Graph Retrieval Augmented Generation
- SGMem: Sentence Graph Memory for Long-Term Conversational Agents
- AutoIntent: AutoML for Text Classification
- Acoustic-based Gender Differentiation in Speech-aware Language Models
- PMark: Towards Robust and Distortion-free Semantic-level Watermarking with Channel Constraints
- Extracting Conceptual Knowledge to Locate Software Issues
- Rejuvenating Cross-Entropy Loss in Knowledge Distillation for Recommender Systems
- PseudoBridge: Pseudo Code as the Bridge for Better Semantic and Logic Alignment in Code Retrieval
- Distilling Many-Shot In-Context Learning into a Cheat Sheet
- Enrich-on-Graph: Query-Graph Alignment for Complex Reasoning with LLM Enriching
- CLIMATE-FEVER: A Dataset for Verification of Real-World Climate Claims
- Stability of In-Context Learning: A Spectral Coverage Perspective
- Human Semantic Representations of Social Interactions from Moving Shapes
- Document Summarization with Conformal Importance Guarantees
- Synergistic Enhancement of Requirement-to-Code Traceability: A Framework Combining Large Language Model based Data Augmentation and an Advanced Encoder
- MIXRAG : Mixture-of-Experts Retrieval-Augmented Generation for Textual Graph Understanding and Question Answering
- Play by the Type Rules: Inferring Constraints for LLM Functions in Declarative Programs
- AutoSpec: An Agentic Framework for Automatically Drafting Patent Specification
- SteinerSQL: Graph-Guided Mathematical Reasoning for Text-to-SQL Generation
- GuessingGame: Measuring the Informativeness of Open-Ended Questions in Large Language Models
- Creative Transformation in Literary Texts: Modelling Change Across Representational Levels
- Cognitive Load Limits in Large Language Models: Benchmarking Multi-Hop Reasoning
- AI-based scoring systematically underestimates conceptual understanding of linguistically weak students' explanations in physics
- GLM-RAG: Graph Language Models for Graph-Based Retrieval-Augmented Generation
- Chaos in reason: How chain-of-thought LLMs can look for an answer
- Rehearse: Stepping Back from the Confidence Cliff in Self-Improving Autoresearch
- CDAE: Enhancing Perturbation Robustness in Pretrained Language Models with Contrastive Denoising
- Diversifying Personalized Research Ideation against AI-Induced Homogenization
- GGC: Selective Query Correction for Reliable Text-to-SPARQL Generation
- Reasoning Consensus: Structural Ensembling of LLM Reasoning via Weighted DAG Aggregation
- Cocktail-Talker: Multi-Speaker Dialog Modeling in Noisy Social Environments with Turn Action GRPO
- Gradient-free Task-Conditioned Retrieval for On-Device In-Context Learning
- From Backlog Items to Security Guidance: Towards Continuous Security Compliance
- MedLLM: An Open Medical Language Model at the Sub-Billion Scale
- ThreatForest: Multi-Agent Attack Tree Generation with Pluggable TTP Framework Mapping
- CAIT: A Syntactic Parsing Toolkit for Child-Adult InTeractions
- RELISH: LLM REgression with a Latent Iterative State Head
- LLM2Vec-Gen: Generative Embeddings from Large Language Models
- GCT: A Granger-Causal Transformer for Multivariate Traffic Analysis in Smart Villages
- Previously on... Automating Code Review
- MERMAID: Metaphor Generation with Symbolism and Discriminative Decoding
- Text Similarity Using Word Embeddings to Classify Misinformation
- WOER suchet, der findet nicht! Identifikation von thematisch verwandten OER-Materialien
- Enhancing Diversity in News Recommendations Increases Click-Through Rates: Insights from an Online Experiment and User Study
- A study of word embedding models for measuring topic coherence
- PhantomLint: Principled Detection of Hidden LLM Prompts in Structured Documents
- LexSemBridge: Fine-Grained Dense Representation Enhancement through Token-Aware Embedding Augmentation
- PLM-interact: extending protein language models to predict protein-protein interactions
- Randomly Removing 50% of Dimensions in Text Embeddings has Minimal Impact on Retrieval and Classification Tasks
- ConViS-Bench: Estimating Video Similarity Through Semantic Concepts
- AgentInit: Initializing LLM-based Multi-Agent Systems via Diversity and Expertise Orchestration for Effective and Efficient Collaboration
- Text Meets Topology: Rethinking Out-of-distribution Detection in Text-Rich Networks
- Single-Branch Network Architectures to Close the Modality Gap in Multimodal Recommendation
- Agentic AutoSurvey: Let LLMs Survey LLMs
- CCQA: Generating Question from Solution Can Improve Inference-Time Reasoning in SLMs
- Towards Synthesizing Normative Data for Cognitive Assessments Using Generative Multimodal Large Language Models
- Investigating Traffic Accident Detection Using Multimodal Large Language Models
- Geometric Structures and Patterns of Meaning: A PHATE Manifold Analysis of Chinese Character Embeddings
- Confidence-Aware Routing for Large Language Model Reliability Enhancement: A Multi-Signal Approach to Pre-Generation Hallucination Mitigation
- Extracting Conceptual Spaces from LLMs Using Prototype Embeddings
- Are Smaller Open-Weight LLMs Closing the Gap to Proprietary Models for Biomedical Question Answering?
- AIRwaves at CheckThat! 2025: Retrieving Scientific Sources for Implicit Claims on Social Media with Dual Encoders and Neural Re-Ranking
- When Meaning Stays the Same, but Models Drift: Evaluating Quality of Service under Token-Level Behavioral Instability in LLMs
- Mind Your Ps and Qs: Supporting Positive Reinforcement in Moderation Through a Positive Queue
- PEEK: Guiding and Minimal Image Representations for Zero-Shot Generalization of Robot Manipulation Policies
- Semantic Reformulation Entropy for Robust Hallucination Detection in QA Tasks
- Scale-free Characteristics of Multilingual Legal Texts and the Limitations of LLMs
- How Persuasive is Your Context?
- Transformer-Gather, Fuzzy-Reconsider: A Scalable Hybrid Framework for Entity Resolution
- SilentStriker:Toward Stealthy Bit-Flip Attacks on Large Language Models
- Localizing Malicious Outputs from CodeLLM
- Scaling Law for Recommendation Models: Towards General-purpose User Representations
- Quantum Abduction: A New Paradigm for Reasoning under Uncertainty
- VidCLearn: A Continual Learning Approach for Text-to-Video Generation
- Semantic-Driven Topic Modeling for Analyzing Creativity in Virtual Brainstorming
- Cognitive Linguistic Identity Fusion Score (CLIFS): A Scalable Cognition-Informed Approach to Quantifying Identity Fusion from Text
- Learn to Rank Risky Investors: A Case Study of Predicting Retail Traders' Behaviour and Profitability
- Long document summarization using page specific target text alignment and distilling page importance
- mmExpert: Integrating Large Language Models for Comprehensive mmWave Data Synthesis and Understanding
- GRIL: Knowledge Graph Retrieval-Integrated Learning with Large Language Models
- MPCG: Multi-Round Persona-Conditioned Generation for Modeling the Evolution of Misinformation with LLMs
- The Role of Vocabularies in Learning Sparse Representations for Ranking
- Patterns in the Transition From Founder-Leadership to Community Governance of Open Source
- Reward Hacking Mitigation using Verifiable Composite Rewards
- How do Language Models Generate Slang: A Systematic Comparison between Human and Machine-Generated Slang Usages
- LibriTTS-VI: A Public Corpus and Novel Methods for Efficient Voice Impression Control
- Latent Zoning Network: A Unified Principle for Generative Modeling, Representation Learning, and Classification
- Evaluating the Effectiveness and Scalability of LLM-Based Data Augmentation for Retrieval
- Enhancing Financial RAG with Agentic AI and Multi-HyDE: A Novel Approach to Knowledge Retrieval and Hallucination Reduction
- SERVAL: Surprisingly Effective Zero-Shot Visual Document Retrieval Powered by Large Vision and Language Models
- LLM-Assisted Topic Reduction for BERTopic on Social Media Data
- Efficient and Versatile Model for Multilingual Information Retrieval of Islamic Text: Development and Deployment in Real-World Scenarios
- Quantifying Self-Awareness of Knowledge in Large Language Models
- An Artificial Intelligence Driven Semantic Similarity-Based Pipeline for Rapid Literature
- Spatial-CLAP: Learning Spatially-Aware audio--text Embeddings for Multi-Source Conditions
- Reveal and Release: Iterative LLM Unlearning with Self-generated Data
- TableDART: Dynamic Adaptive Multi-Modal Routing for Table Understanding
- Catch Me If You Can? Not Yet: LLMs Still Struggle to Imitate the Implicit Writing Styles of Everyday Authors
- Retrieval Capabilities of Large Language Models Scale with Pretraining FLOPs
- Words to Waves: Emotion-Adaptive Music Recommendation System
- PhenoGnet: A Graph-Based Contrastive Learning Framework for Disease Similarity Prediction
- MICA: Multi-Agent Industrial Coordination Assistant
- DSpAST: Disentangled Representations for Spatial Audio Reasoning with Large Language Models
- Who Taught the Lie? Responsibility Attribution for Poisoned Knowledge in Retrieval-Augmented Generation
- Controllable Pareto Trade-off between Fairness and Accuracy
- GeoAware-VLA: Implicit Geometry Aware Vision-Language-Action Model
- Evaluating LLM Alignment on Personality Inference from Real-World Interview Data
- The Few-shot Dilemma: Over-prompting Large Language Models
- Rethinking the Evaluation of Alignment Methods: Insights into Diversity, Generalisation, and Safety
- Conan-Embedding-v2: Training an LLM from Scratch for Text Embeddings
- FedMentor: Domain-Aware Differential Privacy for Heterogeneous Federated LLMs in Mental Health
- Are You Sure You're Positive? Consolidating Chain-of-Thought Agents with Uncertainty Quantification for Aspect-Category Sentiment Analysis
- Don't Change My View: Ideological Bias Auditing in Large Language Models
- ScaleDoc: Scaling LLM-based Predicates over Large Document Collections
- LEAF: Knowledge Distillation of Text Embedding Models with Teacher-Aligned Representations
- Retrieve-and-Verify: A Table Context Selection Framework for Accurate Column Annotations
- Smoothed Contrastive Learning for Unsupervised Sentence Embedding
- MA-DPR: Manifold-aware Distance Metrics for Dense Passage Retrieval
- Optimizing Agricultural Research: A RAG-Based Approach to Mycorrhizal Fungi Information
- Text Adaptation to Plain Language and Easy Read via Automatic Post-Editing Cycles
- MTEB-NL and E5-NL: Embedding Benchmark and Models for Dutch
- The Power of Framing: How News Headlines Guide Search Behavior
- Query-Focused Extractive Summarization for Sentiment Explanation
- Digital Voices of Survival: From Social Media Disclosures to Support Provisions for Domestic Violence Victims
- Zero-shot Multimodal Document Retrieval via Cross-modal Question Generation
- Graph-Enhanced Retrieval-Augmented Question Answering for E-Commerce Customer Support
- Pluralistic Off-policy Evaluation and Alignment
- Context-Aware Language Models for Forecasting Market Impact from Sequences of Financial News
- GRAID: Synthetic Data Generation with Geometric Constraints and Multi-Agentic Reflection for Harmful Content Detection
- AKCIT-FN at CheckThat! 2025: Switching Fine-Tuned SLMs and LLM Prompting for Multilingual Claim Normalization
- Topic Coverage-based Demonstration Retrieval for In-Context Learning
- Understanding Prompt Management in GitHub Repositories: A Call for Best Practices
- CEMTM: Contextual Embedding-based Multimodal Topic Modeling
- Decoding Plastic Toxicity: An Intelligent Framework for Conflict-Aware Relational Metapath Extraction from Scientific Abstracts
- Difficulty-Aware Agentic Orchestration for Query-Specific Multi-Agent Workflows
- A Biosecurity Agent for Lifecycle LLM Biosecurity Alignment
- Text2Sign Diffusion: A Generative Approach for Gloss-Free Sign Language Production
- GAPrune: Gradient-Alignment Pruning for Domain-Aware Embeddings
- Evaluating Large Language Models for Evidence-Based Clinical Question Answering
- Developer-LLM Conversations: An Empirical Study of Interactions and Generated Code Quality
- The Language of Approval: Identifying the Drivers of Positive Feedback Online
- Immunizing Images from Text to Image Editing via Adversarial Cross-Attention
- JU-NLP at Touché: Covert Advertisement in Conversational AI-Generation and Detection Strategies
- The anatomy of Green AI technologies: structure, evolution, and impact
- Established Psychometric vs. Ecologically Valid Questionnaires: Rethinking Psychological Assessments in Large Language Models
- Beyond the Silence: How Men Navigate Infertility Through Digital Communities and Data Sharing
- Investigating red packet fraud in Android applications: Insights from user reviews
- ZapGPT: Free-form Language Prompting for Simulated Cellular Control
- Topic-Guided Reinforcement Learning with LLMs for Enhancing Multi-Document Summarization
- Synthetic Homes: A Multimodal Generative AI Pipeline for Residential Building Data Generation under Data Scarcity
- Retrieval-Augmented Generation for Reliable Interpretation of Radio Regulations
- Gene-R1: Reasoning with Data-Augmented Lightweight LLMs for Gene Set Analysis
- Graph Alignment via Dual-Pass Spectral Encoding and Latent Space Communication
- Towards Explainable Job Title Matching: Leveraging Semantic Textual Relatedness and Knowledge Graphs
- SEDM: Scalable Self-Evolving Distributed Memory for Agents
- Modelling Analogies and Analogical Reasoning: Connecting Cognitive Science Theory and NLP Research
- From scratch to silver: Creating trustworthy training data for patent-SDG classification using Large Language Models
- Chat-Driven Reconfiguration of Model Predictive Control
- Modality Alignment with Multi-scale Bilateral Attention for Multimodal Recommendation
- InterAct: Advancing Large-Scale Versatile 3D Human-Object Interaction Generation
- LLM-JEPA: Large Language Models Meet Joint Embedding Predictive Architectures
- SALMAN: Stability Analysis of Language Models Through the Maps Between Graph-based Manifolds
- Do All Autoregressive Transformers Remember Facts the Same Way? A Cross-Architecture Analysis of Recall Mechanisms
- JUDGEBERT: Assessing Legal Meaning Preservation Between Sentences
- TriagerX: Dual Transformers for Bug Triaging Tasks with Content and Interaction Based Rankings
- LLM Ensemble for RAG: Role of Context Length in Zero-Shot Question Answering for BioASQ Challenge
- Explainable Semantic Text Relations: A Question-Answering Framework for Comparing Document Content
- ALIGNS: Unlocking nomological networks in psychological measurement through a large language model
- Handling Open-Vocabulary Constructs in Formalizing Specifications: Retrieval-Augmented Parsing with Expert Knowledge
- Two Facets of the Same Optimization Coin: Model Degradation and Representation Collapse in Graph Foundation Models
- ImportSnare: Directed "Code Manual" Hijacking in Retrieval-Augmented Code Generation
- Are LLMs Enough for Hyperpartisan, Fake, Polarized and Harmful Content Detection? Evaluating In-Context Learning vs. Fine-Tuning
- The Role of Exploration Modules in Small Language Models for Knowledge Graph Question Answering
- Guarding Your Conversations: Privacy Gatekeepers for Secure Interactions with Cloud-Based AI Models
- ALLabel: Three-stage Active Learning for LLM-based Entity Recognition using Demonstration Retrieval
- GLEAM: Learning to Match and Explain in Cross-View Geo-Localization
- Biased Tales: Cultural and Topic Bias in Generating Children's Stories
- OmniMap: A General Mapping Framework Integrating Optics, Geometry, and Semantics
- NOWJ@COLIEE 2025: A Multi-stage Framework Integrating Embedding Models and Large Language Models for Legal Retrieval and Entailment
- LLM Analysis of 150+ years of German Parliamentary Debates on Migration Reveals Shift from Post-War Solidarity to Anti-Solidarity in the Last Decade
- Towards EnergyGPT: A Large Language Model Specialized for the Energy Sector
- On the Evaluation of Conditional GANs
- How Small Transformation Expose the Weakness of Semantic Similarity Measures
- mmBERT: A Modern Multilingual Encoder with Annealed Language Learning
- SVGauge: Towards Human-Aligned Evaluation for SVG Generation
- Benchmarking Information Retrieval Models on Complex Retrieval Tasks
- Analysis of Blood Report Images Using General Purpose Vision-Language Models
- Augmented Fine-Tuned LLMs for Enhanced Recruitment Automation
- Multimodal Fine-grained Context Interaction Graph Modeling for Conversational Speech Synthesis
- Learning to Route: Per-Sample Adaptive Routing for Multimodal Multitask Prediction
- An Optimized Pipeline for Automatic Educational Knowledge Graph Construction
- Ontology-Aligned Embeddings for Data-Driven Labour Market Analytics
- KGRAG-SC: Knowledge Graph RAG-Assisted Semantic Communication
- Evaluating Cognitive-Behavioral Fixation via Multimodal User Viewing Patterns on Social Media
- Finding your MUSE: Mining Unexpected Solutions Engine
- ThumbnailTruth: A Multi-Modal LLM Approach for Detecting Misleading YouTube Thumbnails Across Diverse Cultural Settings
- Conceptual Schema Inference for Tabular Datasets using Large Language Models
- Enhancing Technical Documents Retrieval for RAG
- MAGneT: Coordinated Multi-Agent Generation of Synthetic Multi-Turn Mental Health Counseling Sessions
- NER Retriever: Zero-Shot Named Entity Retrieval with Type-Aware Embeddings
- Anti-establishment sentiment on TikTok: Implications for understanding influence(rs) and expertise on social media
- AI-in-the-Loop: Privacy Preserving Real-Time Scam Detection and Conversational Scambaiting by Leveraging LLMs and Federated Learning
- PersonaTeaming: Exploring How Introducing Personas Can Improve Automated AI Red-Teaming
- MLSD: A Novel Few-Shot Learning Approach to Enhance Cross-Target and Cross-Domain Stance Detection
- Explainable Knowledge Graph Retrieval-Augmented Generation (KG-RAG) with KG-SMILE
- The Impact of Critique on LLM-Based Model Generation from Natural Language: The Case of Activity Diagrams
- VQualA 2025 Challenge on Engagement Prediction for Short Videos: Methods and Results
- Grocery to General Merchandise: A Cross-Pollination Recommender using LLMs and Real-Time Cart Context
- IDEAlign: Comparing Large Language Models to Human Experts in Open-ended Interpretive Annotations
- Scaling behavior of large language models in emotional safety classification across sizes and tasks
- L3Cube-IndicHeadline-ID: A Dataset for Headline Identification and Semantic Evaluation in Low-Resource Indian Languages
- Benchmarking Large Language Models for Personalized Guidance in AI-Enhanced Learning
- CLEAR: Contrastive Learning for Sentence Representation
- An Epidemiological Knowledge Graph extracted from the World Health Organization's Disease Outbreak News
- Attributes as Textual Genes: Leveraging LLMs as Genetic Algorithm Simulators for Conditional Synthetic Data Generation
- StructCoh: Structured Contrastive Learning for Context-Aware Text Semantic Matching
- Avoidance Decoding for Diverse Multi-Branch Story Generation
- Towards Agents That Know When They Don't Know: Uncertainty as a Control Signal for Structured Reasoning
- Take That for Me: Multimodal Exophora Resolution with Interactive Questioning for Ambiguous Out-of-View Instructions
- KoBLEX: Open Legal Question Answering with Multi-hop Reasoning
- Towards Open-World Retrieval-Augmented Generation on Knowledge Graph: A Multi-Agent Collaboration Framework
- Natural Context Drift Undermines the Natural Language Understanding of Large Language Models
- Hierarchical Motion Captioning Utilizing External Text Data Source
- Enhancing Large Language Model for Knowledge Graph Completion via Structure-Aware Alignment-Tuning
- AMAZe: A Multi-Agent Zero-shot Index Advisor for Relational Databases
- Testing the assumptions about the geometry of sentence embedding spaces: the cosine measure need not apply
- Dissecting Atomic Facts: Visual Analytics for Improving Fact Annotations in Language Model Evaluation
- Understanding Fanchuan in Livestreaming Platforms: A New Form of Online Antisocial Behavior
- Decomposing and Revising What Language Models Generate
- MLLMRec: Exploring the Potential of Multimodal Large Language Models in Recommender Systems
- OpinioRAG: Towards Generating User-Centric Opinion Highlights from Large-scale Online Reviews
- Modeling Motivated Reasoning in Law: Evaluating Strategic Role Conditioning in LLM Summarization
- Standard vs. Modular Sampling: Best Practices for Reliable LLM Unlearning
- Data Auctions for Retrieval Augmented Generation
- The Demon is in Ambiguity: Revisiting Situation Recognition with Single Positive Multi-Label Learning
- T-Retrievability: A Topic-Focused Approach to Measure Fair Document Exposure in Information Retrieval
- QZhou-Embedding Technical Report
- L3Cube-MahaSTS: A Marathi Sentence Similarity Dataset and Models
- Efficient Code Embeddings from Code Generation Models
- Synthetic CVs To Build and Test Fairness-Aware Hiring Tools
- GSTBench: A Benchmark Study on the Transferability of Graph Self-Supervised Learning
- InSQuAD: In-Context Learning for Efficient Retrieval via Submodular Mutual Information to Enforce Quality and Diversity
- Graph-Based Feature Augmentation for Predictive Tasks on Relational Datasets
- Native Logical and Hierarchical Representations with Subspace Embeddings
- ConspirED: A Dataset for Cognitive Traits of Conspiracy Theories and Large Language Model Safety
- From Post To Personality: Harnessing LLMs for MBTI Prediction in Social Media
- SciTopic: Enhancing Topic Discovery in Scientific Literature through Advanced LLM
- Leveraging Semantic Triples for Private Document Generation with Local Differential Privacy Guarantees
- Exploring Selective Retrieval-Augmentation for Long-Tail Legal Text Classification
- NLKI: A lightweight Natural Language Knowledge Integration Framework for Improving Small VLMs in Commonsense VQA Tasks
- SPELUNKER: Item Similarity Search Using Large Language Models and Custom K-Nearest Neighbors
- Functional Consistency of LLM Code Embeddings: A Self-Evolving Data Synthesis Framework for Benchmarking
- Continual Neural Topic Model
- LegiScout: A Visual Tool for Understanding Complex Legislation
- SafetyFlow: An Agent-Flow System for Automated LLM Safety Benchmarking
- Network-Level Prompt and Trait Leakage in Local Research Agents
- SIExVulTS: Sensitive Information Exposure Vulnerability Detection System using Transformer Models and Static Analysis
- Inference Gap in Domain Expertise and Machine Intelligence in Named Entity Recognition: Creation of and Insights from a Substance Use-related Dataset
- Stack Trace-Based Crash Deduplication with Transformer Adaptation
- DeepMEL: A Multi-Agent Collaboration Framework for Multimodal Entity Linking
- The Double-edged Sword of LLM-based Data Reconstruction: Understanding and Mitigating Contextual Vulnerability in Word-level Differential Privacy Text Sanitization
- Membership Inference Attacks on LLM-based Recommender Systems
- Controllable Conversational Theme Detection Track at DSTC 12
- Flexible metadata harvesting for ecology using large language models
- Constraint Matters: Multi-Modal Representation for Reducing Mixed-Integer Linear programming
- Rethinking Caching for LLM Serving Systems: Beyond Traditional Heuristics
- Granite Embedding R2 Models
- Uncovering Intervention Opportunities for Suicide Prevention with Language Model Assistants
- DenseRec: Revisiting Dense Content Embeddings for Sequential Transformer-based Recommendation
- Leveraging Large Language Models for Accurate Sign Language Translation in Low-Resource Scenarios
- InReAcTable: LLM-Powered Interactive Visual Data Story Construction from Tabular Data
- Named Entity Recognition of Historical Text via Large Language Model
- LLM-Guided Genetic Improvement: Envisioning Semantic Aware Automated Software Evolution
- Reference and Document Aware Semantic Evaluation Methods for Korean Language Summarization
- Subjective Behaviors and Preferences in LLM: Language of Browsing
- Continuous sentiment scores for literary and multilingual contexts
- Towards Skeletal and Signer Noise Reduction in Sign Language Production via Quaternion-Based Pose Encoding and Contrastive Learning
- Inv-Entropy: A Fully Probabilistic Framework for Uncertainty Quantification in Language Models
- Supporting Clustering with Contrastive Learning
- Democratizing News Recommenders: Modeling Multiple Perspectives for News Candidate Generation with VQ-VAE
- InPars+: Supercharging Synthetic Data Generation for Information Retrieval Systems
- The illusion of a perfect metric: Why evaluating AI's words is harder than it looks
- Interactive Query Answering on Knowledge Graphs with Soft Entity Constraints
- AdaDocVQA: Adaptive Framework for Long Document Visual Question Answering in Low-Resource Settings
- Statistical Comparative Analysis of Semantic Similarities and Model Transferability Across Datasets for Short Answer Grading
- Driving Style Recognition Like an Expert Using Semantic Privileged Information from Large Language Models
- A Functionality-Grounded Benchmark for Evaluating Web Agents in E-commerce Domains
- XAMT: Cross-Framework API Matching for Testing Deep Learning Libraries
- LumiMAS: A Comprehensive Framework for Real-Time Monitoring and Enhanced Observability in Multi-Agent Systems
- Structuring the Unstructured: A Systematic Review of Text-to-Structure Generation for Agentic AI with a Universal Evaluation Framework
- SEA-BED: Southeast Asia Embedding Benchmark
- Cost-Aware Contrastive Routing for LLMs
- Scalable RF Simulation in Generative 4D Worlds
- Multi-Modal Drift Forecasting of Leeway Objects via Navier-Stokes-Guided CNN and Sequence-to-Sequence Attention-Based Models
- LLM-as-a-Judge for Privacy Evaluation? Exploring the Alignment of Human and LLM Perceptions of Privacy in Textual Data
- In-Context Examples Matter: Improving Emotion Recognition in Conversation with Instruction Tuning
- Controlling Multimodal LLMs via Reward-guided Decoding
- CoDiEmb: A Collaborative yet Distinct Framework for Unified Representation Learning in Information Retrieval and Semantic Textual Similarity
- Retrieval-augmented reasoning with lean language models
- ORFuzz: Fuzzing the "Other Side" of LLM Safety -- Testing Over-Refusal
- RAG for Geoscience: What We Expect, Gaps and Opportunities
- From Feedback to Failure: Automated Android Performance Issue Reproduction
- Copyright Protection for Large Language Models: A Survey of Methods, Challenges, and Trends
- GenOM: Ontology Matching with Description Generation and Large Language Model
- IADGPT: Unified LVLM for Few-Shot Industrial Anomaly Detection, Localization, and Reasoning via In-Context Learning
- ChatENV: An Interactive Vision-Language Model for Sensor-Guided Environmental Monitoring and Scenario Simulation
- A Study of Commonsense Reasoning over Visual Object Properties
- Semantic IDs for Joint Generative Search and Recommendation
- Dataset Construction for Training LLM to Learn Analog Circuit Knowledge
- Empowering Multimodal LLMs with External Tools: A Comprehensive Survey
- Prompt-Response Semantic Divergence Metrics for Faithfulness Hallucination and Misalignment Detection in Large Language Models
- DS4RS: Community-Driven and Explainable Dataset Search Engine for Recommender System Research
- Estimating Machine Translation Difficulty
- Amazon Nova AI Challenge -- Trusted AI: Advancing secure, AI-assisted software development
- January Food Benchmark (JFB): A Public Benchmark Dataset and Evaluation Suite for Multimodal Food Analysis
- Social-Sensor Identity Cloning Detection Using Weakly Supervised Deep Forest and Cryptographic Authentication
- Towards Self-cognitive Exploration: Metacognitive Knowledge Graph Retrieval Augmented Generation
- UWBa at SemEval-2025 Task 7: Multilingual and Crosslingual Fact-Checked Claim Retrieval
- IAG: Input-aware Backdoor Attack on VLM-based Visual Grounding
- Semantic Distance: A New Metric for ASR Performance Analysis Towards Spoken Language Understanding
- Link Prediction for Event Logs in the Process Industry
- BiasGym: A Simple and Generalizable Framework for Analyzing and Removing Biases through Elicitation
- DiffPose-Animal: A Language-Conditioned Diffusion Framework for Animal Pose Estimation
- Exploring Palette based Color Guidance in Diffusion Models
- GreenTEA: Gradient Descent with Topic-modeling and Evolutionary Auto-prompting
- E3-Rewrite: Learning to Rewrite SQL for Executability, Equivalence,and Efficiency
- Jointly Generating and Attributing Answers using Logits of Document-Identifier Tokens
- Mitigating Popularity Bias in Counterfactual Explanations using Large Language Models
- BlindGuard: Safeguarding LLM-based Multi-Agent Systems under Unknown Attacks
- AdaptFlow: Adaptive Workflow Optimization via Meta-Learning
- DIVER: A Multi-Stage Approach for Reasoning-intensive Information Retrieval
- UniSVG: A Unified Dataset for Vector Graphic Understanding and Generation with Multimodal Large Language Models
- Semantic Caching for Low-Cost LLM Serving: From Offline Learning to Online Adaptation
- Retrieval-Augmented Multi-Agent System for Rapid Statement of Work Generation
- In-situ Value-aligned Human-Robot Interactions with Physical Constraints
- GLiClass: Generalist Lightweight Model for Sequence Classification Tasks
- Temporal User Profiling with LLMs: Balancing Short-Term and Long-Term Preferences for Recommendations
- Using LLMs to Capture Users' Temporal Context for Recommendation
- Improving Document Retrieval Coherence for Semantically Equivalent Queries
- What am I missing here?: Evaluating Large Language Models for Masked Sentence Prediction
- VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding
- Grounding Natural Language for Multi-agent Decision-Making with Multi-agentic LLMs
- Are Multimodal Embeddings Truly Beneficial for Recommendation? A Deep Dive into Whole vs. Individual Modalities
- An Embodied AR Navigation Agent: Integrating BIM with Retrieval-Augmented Generation for Language Guidance
- A Survey on Non-Intrusive ASR Refinement: From Output-Level Correction to Full-Model Distillation
- Incorporating Contextual Paralinguistic Understanding in Large Speech-Language Models
- Graph Neural Network for Product Recommendation on the Amazon Co-purchase Graph
- Canvas3D: Empowering Precise Spatial Control for Image Generation with Constraints from a 3D Virtual Canvas
- Multi-Modal Requirements Data-based Acceptance Criteria Generation using LLMs
- ESNERA: Empirical and semantic named entity alignment for named entity dataset merging
- Towards Effective Prompt Stealing Attack against Text-to-Image Diffusion Models
- BiXSE: Improving Dense Retrieval via Probabilistic Graded Relevance Distillation
- Vec2Summ: Text Summarization via Probabilistic Sentence Embeddings
- Position: Ideas Should be the Center of Machine Learning Research
- A Computational Approach to Modeling Conversational Systems: Analyzing Large-Scale Quasi-Patterned Dialogue Flows
- Linking Multi-Site Sex Ad Data at the Individual Level to Aid Counter-Trafficking Efforts
- LLM Unlearning using Gradient Ratio-Based Influence Estimation and Noise Injection
- MIRA: Leveraging Multi-Intention Co-click Information in Web-scale Document Retrieval using Deep Neural Networks
- EICAP: Deep Dive in Assessment and Enhancement of Large Language Models in Emotional Intelligence through Multi-Turn Conversations
- Multi-Objective Instruction-Aware Representation Learning in Procedural Content Generation RL
- LinguaFluid: Language Guided Fluid Control via Semantic Rewards in Reinforcement Learning
- Prosocial Behavior Detection in Player Game Chat: From Aligning Human-AI Definitions to Efficient Annotation at Scale
- Scaling Personality Control in LLMs with Big Five Scaler Prompts
- Automatic Semantic Alignment of Flow Pattern Representations for Exploration with Large Language Models
- Omni Geometry Representation Learning vs Large Language Models for Geospatial Entity Resolution
- Leveraging LLMs for Privacy-Aware Predictions in Participatory Budgeting
- Does Multimodality Improve Recommender Systems as Expected? A Critical Analysis and Future Directions
- Resource-Limited Joint Multimodal Sentiment Reasoning and Classification via Chain-of-Thought Enhancement and Distillation
- A Survey on Video Temporal Grounding with Multimodal Large Language Model
- Tool Graph Retriever: Exploring Dependency Graph-based Tool Retrieval for Large Language Models
- MultiCheck: Strengthening Web Trust with Unified Multimodal Fact Verification
- A multi-stage interactive writing task for the assessment of English language writing proficiency
- SPaRFT: Self-Paced Reinforcement Fine-Tuning for Large Language Models
- A Scalable Pretraining Framework for Link Prediction with Efficient Adaptation
- GraphProp: Training the Graph Foundation Models using Graph Properties
- UniTalker: Conversational Speech-Visual Synthesis
- Do Recommender Systems Really Leverage Multimodal Content? A Comprehensive Analysis on Multimodal Representations for Recommendation
- CALE : Concept-Aligned Embeddings for Both Within-Lemma and Inter-Lemma Sense Differentiation
- PersonaEval: Are LLM Evaluators Human Enough to Judge Role-Play?
- Dialogue Response Prefetching Based on Semantic Similarity and Prediction Confidence of Language Model
- What Do Humans Hear When Interacting? Experiments on Selective Listening for Evaluating ASR of Spoken Dialogue Systems
- SSEmb: A Joint Structural and Semantic Embedding Framework for Mathematical Formula Retrieval
- Controllable Hybrid Captioner for Improved Long-form Video Understanding
- LLMDistill4Ads: Using Cross-Encoders to Distill from LLM Signals for Advertiser Keyphrase Recommendations at eBay
- CF-RAG: A Dataset and Method for Carbon Footprint QA Using Retrieval-Augmented Generation
- fact check AI at SemEval-2025 Task 7: Multilingual and Crosslingual Fact-checked Claim Retrieval
- Cropping outperforms dropout as an augmentation strategy for training self-supervised text embeddings
- Data Overdose? Time for a Quadruple Shot: Knowledge Graph Construction using Enhanced Triple Extraction
- ReDSM5: A Reddit Dataset for DSM-5 Depression Detection
- SmartLLMs Scheduler: A Framework for Cost-Effective LLMs Utilization
- Token-Level Precise Attack on RAG: Searching for the Best Alternatives to Mislead Generation
- PyLate: Flexible Training and Retrieval for Late Interaction Models
- LLM-based IR-system for Bank Supervisors
- Vision Language Model-based Testing of Industrial Autonomous Mobile Robots
- CAAD: Context-Aware Adaptive Decoding for Truthful Text Generation
- SpeechRole: A Large-Scale Dataset and Benchmark for Evaluating Speech Role-Playing Agents
- Improving Hospital Risk Prediction with Knowledge-Augmented Multimodal EHR Modeling
- MArgE: Meshing Argumentative Evidence from Multiple Large Language Models for Justifiable Claim Verification
- Harnessing Temporal Databases for Systematic Evaluation of Factual Time-Sensitive Question-Answering in Large Language Models
- Contextually Aware E-Commerce Product Question Answering using RAG
- LMAR: Language Model Augmented Retriever for Domain-specific Knowledge Indexing
- Intention-Guided Cognitive Reasoning for Egocentric Long-Term Action Anticipation
- Empowering Tabular Data Preparation with Language Models: Why and How?
- ChEmbed: Enhancing Chemical Literature Search Through Domain-Specific Text Embeddings
- Am I Blue or Is My Hobby Counting Teardrops? Expression Leakage in Large Language Models as a Symptom of Irrelevancy Disruption
- TCDiff: Triplex Cascaded Diffusion for High-fidelity Multimodal EHRs Generation with Incomplete Clinical Data
- Instruction-based Time Series Editing
- Balancing the Blend: An Experimental Analysis of Trade-offs in Hybrid Search
- Aligning Language Models with Real-time Knowledge Editing
- Towards Bridging Review Sparsity in Recommendation with Textual Edge Graph Representation
- Towards Efficient Medical Reasoning with Minimal Fine-Tuning Data
- Disaggregated Health Data in LLMs: Evaluating Data Equity in the Context of Asian American Representation
- Team "bettercallclaude": Style Change Detection using a Sequential Sentence Pair Classifier
- Experimental Evaluation of Dynamic Topic Modeling Algorithms
- A Pilot Study on LLM-Based Agentic Translation from Android to iOS: Pitfalls and Insights
- Activation-Guided Local Editing for Jailbreaking Attacks
- The Prosody of Emojis
- The Missing Parts: Augmenting Fact Verification with Half-Truth Detection
- CyGATE: Game-Theoretic Cyber Attack-Defense Engine for Patch Strategy Optimization
- Cued-Agent: A Collaborative Multi-Agent System for Automatic Cued Speech Recognition
- Accurate and Consistent Graph Model Generation from Text with Large Language Models
- Improving Multimodal Contrastive Learning of Sentence Embeddings with Object-Phrase Alignment
- Automating AI Failure Tracking: Semantic Association of Reports in AI Incident Database
- From Static to Dynamic: A Streaming RAG Approach to Real-time Knowledge Base
- MECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks
- Role-Aware Language Models for Secure and Contextualized Access Control in Organizations
- Self-Foveate: Enhancing Diversity and Difficulty of Synthesized Instructions from Unsupervised Text via Multi-Level Foveation
- Causal2Vec: Improving Decoder-only LLMs as Versatile Embedding Models
- Text-to-SQL Task-oriented Dialogue Ontology Construction
- Accessibility Scout: Personalized Accessibility Scans of Built Environments
- Failures Are the Stepping Stones to Success: Enhancing Few-Shot In-Context Learning by Leveraging Negative Samples
- Social and Political Framing in Search Engine Results
- Real-time News Story Identification
- A Framework for Institutional Risk Identification using Knowledge Graphs and Automated News Profiling
- Resource-Efficient Adaptation of Large Language Models for Text Embeddings via Prompt Engineering and Contrastive Fine-tuning
- Leveraging Context for Multimodal Fallacy Classification in Political Debates
- PersonaTwin: A Multi-Tier Prompt Conditioning Framework for Generating and Evaluating Personalized Digital Twins
- From personalized news curation to shared issue concerns in fragmentation era: a dynamic network approach by levels of issue involvement
- RASL: Retrieval Augmented Schema Linking for Massive Database Text-to-SQL
- IdeaBlocks: Expressing and Reusing Divergent Intents for Graphic Design Exploration using Generative AI
- Culinary Crossroads: A RAG Framework for Enhancing Diversity in Cross-Cultural Recipe Adaptation
- Exploration on Demand: From Algorithmic Control to User Empowerment
- Modelling Adjectival Modification Effects on Semantic Plausibility
- HRIPBench: Benchmarking LLMs in Harm Reduction Information Provision to Support People Who Use Drugs
- Helping or Homogenizing? GenAI as a Design Partner to Pre-Service SLPs for Just-in-Time Programming of AAC
- Identification of Design Recommendations for Augmented Reality Authors in Corporate Training
- Predicting Abandonment of Open Source Software Projects with An Integrated Feature Framework
- Multilingual JobBERT for Cross-Lingual Job Title Matching
- Comparison of Information Retrieval Techniques Applied to IT Support Tickets
- MemShare: Memory Efficient Inference for Large Reasoning Models through KV Cache Reuse
- Does Editing Improve Answer Quality on Stack Overflow? A Data-Driven Investigation
- Prescriptive Agents based on RAG for Automated Maintenance (PARAM)
- On The Role of Pretrained Language Models in General-Purpose Text Embeddings: A Survey
- Uncertainty-driven Embedding Convolution
- SAND-Math: Using LLMs to Generate Novel, Difficult and Useful Mathematics Questions and Answers
- Security Challenges in AI Agent Deployment: Insights from a Large Scale Public Competition
- SDD: Self-Degraded Defense against Malicious Fine-tuning
- Modeling Insider Filing Delays in Financial Markets with an Interpretable XGBoost Framework
- K4: Online Log Anomaly Detection Via Unsupervised Typicality Learning
- VLQA: The First Comprehensive, Large, and High-Quality Vietnamese Dataset for Legal Question Answering
- RoD-TAL: A Benchmark for Answering Questions in Romanian Driving License Exams
- Towards Domain Specification of Embedding Models in Medicine
- Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization
- A Similarity Measure for Comparing Conversational Dynamics
- Smooth Reading: Bridging the Gap of Recurrent LLM to Self-Attention LLM on Long-Context Tasks
- Retrieval augmented generation based dynamic prompting for few-shot biomedical named entity recognition using large language models
- DxHF: Providing High-Quality Human Feedback for LLM Alignment via Interactive Decomposition
- From Roots to Rewards: Dynamic Tree Reasoning with Reinforcement Learning
- AccessGuru: Leveraging LLMs to Detect and Correct Web Accessibility Violations in HTML Code
- 3D Software Synthesis Guided by Constraint-Expressive Intermediate Representation
- How Well Do LLMs Predict Prerequisite Skills? Zero-Shot Comparison to Expert-Defined Concepts
- LLM-based Embedders for Prior Case Retrieval
- Actively evaluating and learning the distinctions that matter: Vaccine safety signal detection from emergency triage notes
- Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction
- EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs
- SynC: Synthetic Image Caption Dataset Refinement with One-to-many Mapping for Zero-shot Image Captioning
- DR.EHR: Dense Retrieval for Electronic Health Record with Knowledge Injection and Synthetic Data
- TDR: Task-Decoupled Retrieval with Fine-Grained LLM Feedback for In-Context Learning
- Transform Before You Query: A Privacy-Preserving Approach for Vector Retrieval with Embedding Space Alignment
- SCOPE: Stochastic and Counterbiased Option Placement for Evaluating Large Language Models
- E.A.R.T.H.: Structuring Creative Evolution through Model Error in Generative AI
- Stealthy LLM-Driven Data Poisoning Attacks Against Embedding-Based Retrieval-Augmented Recommender Systems
- MEKiT: Multi-source Heterogeneous Knowledge Injection Method via Instruction Tuning for Emotion-Cause Pair Extraction
- Quantizing Text-attributed Graphs for Semantic-Structural Integration
- Doc2Chart: Intent-Driven Zero-Shot Chart Generation from Documents
- Descrip3D: Enhancing Large Language Model-based 3D Scene Understanding with Object-Level Text Descriptions
- SemCSE: Semantic Contrastive Sentence Embeddings Using LLM-Generated Summaries For Scientific Abstracts
- GRID: Scalable Task-Agnostic Prompt-Based Continual Learning for Language Models
- Optimizing Legal Document Retrieval in Vietnamese with Semi-Hard Negative Mining
- XL-DURel: Finetuning Sentence Transformers for Ordinal Word-in-Context Classification
- FullStack-Agent: Enhancing Agentic Full-Stack Web Coding via Development-Oriented Testing and Repository Back-Translation
- How AI Coding Agents Modify Code: A Large-Scale Study of GitHub Pull Requests
- PAT++: a cautionary tale about generative visual augmentation for Object Re-identification
- PARK: Personalized academic retrieval with knowledge-graphs
- CarD-T: an automated pipeline for the nomination and analysis of potential human carcinogens
- Linguistic and Embedding-Based Profiling of Texts generated by Humans and Large Language Models
- From Extraction to Synthesis: Entangled Heuristics for Agent-Augmented Strategic Reasoning
- KiC: Keyword-inspired Cascade for Cost-Efficient Text Generation with LLMs
- Confident RAG: Enhancing the Performance of LLMs for Mathematics Question Answering through Multi-Embedding and Confidence Scoring
- Bob's Confetti: Phonetic Memorization Attacks in Music and Video Generation
- InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation
- Marcel: A Lightweight and Open-Source Conversational Agent for University Student Support
- Combining Language and Topic Models for Hierarchical Text Classification
- ReasonVQA: A Multi-hop Reasoning Benchmark with Structural Knowledge for Visual Question Answering
- Enhancing patent retrieval using automated patent summarization
- BDIViz: An Interactive Visualization System for Biomedical Schema Matching with LLM-Powered Validation
- A Computational Framework to Identify Self-Aspects in Text
- DEMONSTRATE: Zero-shot Language to Robotic Control via Multi-task Demonstration Learning
- SE-VLN: A Self-Evolving Vision-Language Navigation Framework Based on Multimodal Large Language Models
- GenBreak: Red Teaming Text-to-Image Generators Using Large Language Models
- S2WTM: Spherical Sliced-Wasserstein Autoencoder for Topic Modeling
- Advancing Retrieval-Augmented Generation for Structured Enterprise and Internal Data
- Fair Representation in Parliamentary Summaries: Measuring and Mitigating Inclusion Bias
- Text-ADBench: Text Anomaly Detection Benchmark based on LLMs Embedding
- Simplifications are Absolutists: How Simplified Language Reduces Word Sense Awareness in LLM-Generated Definitions
- Overview of the Sensemaking Task at the ELOQUENT 2025 Lab: LLMs as Teachers, Students and Evaluators
- Real-World Summarization: When Evaluation Reaches Its Limits
- TOPJoin: A Context-Aware Multi-Criteria Approach for Joinable Column Search
- CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios
- A study of search result aggregation approaches for the digital humanities
- Towards Better Code Generation: Adaptive Decoding with Uncertainty Guidance
- Dense Retrievers Can Fail on Simple Queries: Revealing The Granularity Dilemma of Embeddings
- From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge
- Reinforcement Fine-Tuning for Reasoning towards Multi-Step Multi-Source Search in Large Language Models
- H2GFM: Towards unifying Homogeneity and Heterogeneity on Text-Attributed Graphs
- Self-Anchored Attention Model for Sample-Efficient Classification of Prosocial Text Chat
- Opus: A Prompt Intention Framework for Complex Workflow Generation
- An Empirical Study of Multi-Agent RAG for Real-World University Admissions Counseling
- Paraphrase Generation as Unsupervised Machine Translation
- Aligned Query Expansion: Efficient Query Expansion for Information Retrieval through LLM Alignment
- Personalized Exercise Recommendation with Semantically-Grounded Knowledge Tracing
- Tracing the Path to Grokking: Embeddings, Dropout, and Network Activation
- LLMATCH: A Unified Schema Matching Framework with Large Language Models
- Evaluating Generated Commit Messages with Large Language Models
- Journalism-Guided Agentic In-Context Learning for News Stance Detection
- EtiCor++: Towards Understanding Etiquettical Bias in LLMs
- Language Models for Adult Service Website Text Analysis
- Can You Detect the Difference?
- Multiple Choice Learning of Low-Rank Adapters for Language Modeling
- CLOSP: A Unified Semantic Space for SAR, MSI, and Text in Remote Sensing
- Automating SPARQL Query Translations between DBpedia and Wikidata
- (Almost) Free Modality Stitching of Foundation Models
- Protective Factor-Aware Dynamic Influence Learning for Suicide Risk Prediction on Social Media
- Evaluating the Performance and Efficiency of Sentence-BERT for Code Comment Classification
- Hierarchical Job Classification with Similarity Graph Integration
- TolerantECG: A Foundation Model for Imperfect Electrocardiogram
- TRACE: Grounding Time Series in Context for Multimodal Embedding and Retrieval
- SciSummPip: An Unsupervised Scientific Paper Summarization Pipeline
- (RSA)2: A Rhetorical-Strategy-Aware Rational Speech Act Framework for Figurative Language Understanding
- CoMuMDR: Code-mixed Multi-modal Multi-domain corpus for Discourse paRsing in conversations
- Extracting Cause-Effect Pairs from a Sentence with a Dependency-Aware Transformer Model
- Visually grounded emotion regulation via diffusion models and user-driven reappraisal
- Extracting Important Tokens in E-Commerce Queries with a Tag Interaction-Aware Transformer Model
- SLIF-MR: Self-loop Iterative Fusion of Heterogeneous Auxiliary Information for Multimodal Recommendation
- A Scalable and Efficient Signal Integration System for Job Matching
- How Important is `Perfect' English for Machine Translation Prompts?
- NMIXX: Domain-Adapted Neural Embeddings for Cross-Lingual eXploration of Finance
- Adversarial Demonstration Learning for Low-resource NER Using Dual Similarity
- Back to the Basics: Rethinking Issue-Commit Linking with LLM-Assisted Retrieval
- KV Cache Steering for Controlling Frozen LLMs
- One Token to Fool LLM-as-a-Judge
- Structure-Augmented Reasoning Generation
- LLMCup: Ranking-Enhanced Comment Updating with LLMs
- The Impact of Automatic Speech Transcription on Speaker Attribution
- AutoRAG-LoRA: Hallucination-Triggered Knowledge Retuning via Lightweight Adapters
- PromotionGo at SemEval-2025 Task 11: A Feature-Centric Framework for Cross-Lingual Multi-Emotion Detection in Short Texts
- Improving Korean-English Cross-Lingual Retrieval: A Data-Centric Study of Language Composition and Model Merging
- Transforming Sensitive Documents into Quantitative Data: An AI-Based Preprocessing Toolchain for Structured and Privacy-Conscious Analysis
- CRMAgent: A Multi-Agent LLM System for E-Commerce CRM Message Template Generation
- Towards Efficient Quantity Retrieval from Text:An Approach via Description Parsing and Weak Supervision
- VIP: Visual Information Protection through Adversarial Attacks on Vision-Language Models
- Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning
- CMER: A Context-Aware Approach for Mining Ethical Concern-related App Reviews
- Societal AI Research Has Become Less Interdisciplinary
- Revisiting Graph Projections for Effective Complementary Product Recommendation
- GRASP: Generic Reasoning And SPARQL Generation across Knowledge Graphs
- Beyond the Linear Separability Ceiling: Aligning Representations in VLMs
- ODIA: Oriented Distillation for Inline Acceleration of LLM-based Function Calling
- Corvid: Improving Multimodal Large Language Models Towards Chain-of-Thought Reasoning
- Shuffling for Semantic Secrecy
- Measuring a Texts Fairness Dimensions Using Machine Learning Based on Social Psychological Factors
- Improving Clustering on Occupational Text Data through Dimensionality Reduction
- SAGE: A Visual Language Model for Anomaly Detection via Fact Enhancement and Entropy-aware Alignment
- ConsNoTrainLoRA: Data-driven Weight Initialization of Low-rank Adapters using Constraints
- Bridging the Question-Answer Gap in Retrieval-Augmented Generation: Hypothetical Prompt Embeddings
- RAG Safety: Exploring Knowledge Poisoning Attacks to Retrieval-Augmented Generation
- CLI-RAG: A Retrieval-Augmented Framework for Clinically Structured and Context Aware Text Generation with LLMs
- Attention-Aware GNN-based Input Defense against Multi-Turn LLM Jailbreak
- FuDoBa: Fusing Document and Knowledge Graph-based Representations with Bayesian Optimisation
- TextCloak: Thwarting Unauthorized LLM Exploitation via RL-Driven Unlearnable Text
- Enhancing Food-Domain Question Answering with a Multimodal Knowledge Graph: Hybrid QA Generation and Diversity Analysis
- DS@GT at CheckThat! 2025: Exploring Retrieval and Reranking Pipelines for Scientific Claim Source Retrieval on Social Media Discourse
- Towards Theme Detection in Personal Finance Questions
- Self-Consistency in Vision-Language Models for Precision Agriculture: Multi-Response Consensus for Crop Disease Management
- Data-Semantics-Aware Recommendation of Diverse Pivot Tables
- An Ensemble Embedding Approach for Improving Semantic Caching Performance in LLM-based Systems
- DocIE@XLLM25: In-Context Learning for Information Extraction using Fully Synthetic Demonstrations
- Fair Domain Generalization: An Information-Theoretic View
- SARA: Selective and Adaptive Retrieval-augmented Generation with Context Compression
- Detecting Proxy Gaming in RL and LLM Alignment via Evaluator Stress Tests
- Beyond Retrieval: Ensembling Cross-Encoders and GPT Rerankers with LLMs for Biomedical QA
- DS@GT at CheckThat! 2025: Detecting Subjectivity via Transfer-Learning and Corrective Data Augmentation
- Conditional Multi-Stage Failure Recovery for Embodied Agents
- Enhancing Test-Time Scaling of Large Language Models with Hierarchical Retrieval-Augmented MCTS
- SERUM: State Extraction and Refinement for User Modeling
- Detecting Experiential Intertextuality Across Migration Routes: Beyond Surface Similarity in French Narratives
- Semantic Certainty Assessment in Vector Retrieval Systems: A Novel Framework for Embedding Quality Evaluation
- Conversational Education at Scale: A Multi-LLM Agent Workflow for Procedural Learning and Pedagogic Quality Assessment
- Spatio-Temporal LLM: Reasoning about Environments and Actions
- IDAGC: Adaptive Generalized Human-Robot Collaboration via Human Intent Estimation and Multimodal Policy Learning
- Reproducing LightMem: Naive RAG Is Just as Good for Memory Management
- News Source Citing Patterns in AI Search Systems
- An Ontology-Guided, Deduplication-Aware Extraction Layer for Knowledge Graph Construction from Heterogeneous Documents
- Representation learning with a transformer by contrastive learning for money laundering detection
- SMART: Simulated Students Aligned with Item Response Theory for Question Difficulty Prediction
- SpiritRAG: A Q&A System for Religion and Spirituality in the United Nations Archive
- M3-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding
- Does online sustainability communication shape public discourse? Insights from six years of tenant-housing provider interactions
- A Modular Unsupervised Framework for Attribute Recognition from Unstructured Text
- Demystifying ChatGPT: How It Masters Genre Recognition
- Self-citation Analysis using Sentence Embeddings
- SMCLM: Semantically Meaningful Causal Language Modeling for Autoregressive Paraphrase Generation
- KinyaColBERT: A Lexically Grounded Retrieval Model for Low-Resource Retrieval-Augmented Generation
- ASBERT: Siamese and Triplet network embedding for open question answering
- Beyond Weaponization: NLP Security for Medium and Lower-Resourced Languages in Their Own Right
- MemOS: A Memory OS for AI System
- RCA Copilot: Transforming Network Data into Actionable Insights via Large Language Models
- Multimodal Mathematical Reasoning with Diverse Solving Perspective
- Multi-Agent Reinforcement Learning for Dynamic Pricing in Supply Chains: Benchmarking Strategic Agent Behaviours under Realistically Simulated Market Conditions
- Automated Grading of Students' Handwritten Graphs: A Comparison of Meta-Learning and Vision-Large Language Models
- CyberRAG: An Agentic RAG cyber attack classification and reporting tool
- Efficient Code LLM Training via Distribution-Consistent and Diversity-Aware Data Selection
- Coling-UniA at SciVQA 2025: Few-Shot Example Retrieval and Confidence-Informed Ensembling for Multimodal Large Language Models
- SciGA: A Comprehensive Dataset for Designing Graphical Abstracts in Academic Papers
- Improving Constrained Language Generation via Self-Distilled Twisted Sequential Monte Carlo
- Enhancing COBOL Code Explanations: A Multi-Agents Approach Using Large Language Models
- MoIRA: Modular Instruction Routing Architecture for Multi-Task Robotics
- Evaluating the Effectiveness of Direct Preference Optimization for Personalizing German Automatic Text Simplifications for Persons with Intellectual Disabilities
- Using multi-agent architecture to mitigate the risk of LLM hallucinations
- Step-Level Visual Grounding Faithfulness Predicts Out-of-Distribution Generalization in Long-Horizon Vision-Language Models
- Disentangling Similarity and Relatedness in Topic Models
- ELISA: An Interpretable Hybrid Generative AI Agent for Expression-Grounded Discovery in Single-Cell Genomics
- OPERA: Online Data Pruning for Efficient Retrieval Model Adaptation
- Shaping Scientific Explanations to Expert Perspectives with Persona-Conditioned Reinforcement Learning
- Data Agent: A Holistic Architecture for Orchestrating Data+AI Ecosystems
- Long-Tailed Distribution-Aware Router For Mixture-of-Experts in Large Vision-Language Model
- The Medium Is Not the Message: Deconfounding Document Embeddings via Linear Concept Erasure
- Matching and Linking Entries in Historical Swedish Encyclopedias
- GaussianVLM: Scene-centric 3D Vision-Language Models using Language-aligned Gaussian Splats for Embodied Reasoning and Beyond
- A Comparative Study of Competency Question Elicitation Methods from Ontology Requirements
- FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale
- Zero-Shot Skeleton-Based Action Recognition With Prototype-Guided Feature Alignment
- MassTool: A Multi-Task Search-Based Tool Retrieval Framework for Large Language Models
- Question Decomposition for Retrieval-Augmented Generation
- InvisibleInk: High-Utility and Low-Cost Text Generation with Differential Privacy
- Unveiling Decision-Making in LLMs for Text Classification : Extraction of influential and interpretable concepts with Sparse Autoencoders
- Zero-Shot Contextual Embeddings via Offline Synthetic Corpus Generation
- No Stupid Questions: An Analysis of Question Query Generation for Citation Recommendation
- GradEscape: A Gradient-Based Evader Against AI-Generated Text Detectors
- Agentic Surgical AI: Surgeon Style Fingerprinting and Privacy Risk Quantification via Discrete Diffusion in a Vision-Language-Action Framework
- Cognitive Weave: Synthesizing Abstracted Knowledge with a Spatio-Temporal Resonance Graph
- Holistic Artificial Intelligence in Medicine; improved performance and explainability
- What to Keep and What to Drop: Adaptive Table Filtering Framework
- Hierarchical Memory Organization for Wikipedia Generation
- StreamBridge: Turning Your Offline Video Large Language Model into a Proactive Streaming Assistant
- Machine Assistant with Reliable Knowledge: Enhancing Student Learning via RAG-based Retrieval
- AlignEvoSkill: Towards Knowledge-Aware and Task-Aligned Agent Skill Evolution
- MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings
- Text2VectorSQL: Towards a Unified Interface for Vector Search and SQL Queries
- Density, asymmetry and citation dynamics in scientific literature
- Generating Privacy Stories From Software Documentation
- Video Unlearning via Low-Rank Refusal Vector
- ContextCache: Context-Aware Semantic Cache for Multi-Turn Queries in Large Language Models
- Snap, Segment, Deploy: A Visual Data and Detection Pipeline for Wearable Industrial Assistants
- Persistence Paradox in Dynamic Science
- Test-Time Consistency in Vision Language Models
- Towards Fair Rankings: Leveraging LLMs for Gender Bias Detection and Measurement
- Lost at the Beginning of Reasoning
- Literature-Grounded Novelty Assessment of Scientific Ideas
- Grokking in LLM Pretraining? Monitor Memorization-to-Generalization without Test
- How Good Are Synthetic Requirements ? Evaluating LLM-Generated Datasets for AI4RE
- Premise Selection for a Lean Hammer
- FingerTip 20K: A Benchmark for Proactive and Personalized Mobile LLM Agents
- Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models
- KaLM-Embedding-V2: Superior Training Techniques and Data Inspire A Versatile Embedding Model
- "What's Up, Doc?": Analyzing How Users Seek Health Information in Large-Scale Conversational AI Datasets
- When Style Breaks Safety: Defending LLMs Against Superficial Style Alignment
- Cohort Retrieval using Dense Passage Retrieval
- TableMoE: Neuro-Symbolic Routing for Structured Expert Reasoning in Multimodal Table Understanding
- Brain2Model Transfer: Training sensory and decision models with human neural activity as a teacher
- Fine-Tuning and Prompt Engineering of LLMs, for the Creation of Multi-Agent AI for Addressing Sustainable Protein Production Challenges
- How to Retrieve Examples in In-context Learning to Improve Conversational Emotion Recognition using Large Language Models?
- OAK -- Onboarding with Actionable Knowledge
- Multimodal Information Retrieval for Open World with Edit Distance Weak Supervision
- Towards Probabilistic Question Answering Over Tabular Data
- HERCULES: Hierarchical Embedding-based Recursive Clustering Using LLMs for Efficient Summarization
- Scaling Speculative Decoding with Lookahead Reasoning
- ZeroVO: Visual Odometry with Minimal Assumptions
- QUITE: A Query Rewrite System Beyond Rules with LLM Agents
- Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models
- Synthetic Visual Genome
- Health Sentinel: An AI Pipeline For Real-time Disease Outbreak Detection
- Thought Anchors: Which LLM Reasoning Steps Matter?
- Argument-Based Consistency in Toxicity Explanations of LLMs
- Quantifying Fairness in LLMs Beyond Tokens: A Semantic and Statistical Perspective
- jina-embeddings-v4: Universal Embeddings for Multimodal Multilingual Retrieval
- When Fine-Tuning Fails: Lessons from MS MARCO Passage Ranking
- Standard Applicability Judgment and Cross-jurisdictional Reasoning: A RAG-based Framework for Medical Device Compliance
- AI-Generated Song Detection via Lyrics Transcripts
- Beyond the Sentence: A Survey on Context-Aware Machine Translation with Large Language Models
- OpenEvents V1: Large-Scale Benchmark Dataset for Multimodal Event Grounding
- Comparative Analysis of Lion and AdamW Optimizers for Cross-Encoder Reranking with MiniLM, GTE, and ModernBERT
- Memory-Augmented Architecture for Long-Term Context Handling in Large Language Models
- Spiritual-LLM : Gita Inspired Mental Health Therapy In the Era of LLMs
- How Large Language Models play humans in online conversations: a simulated study of the 2016 US politics on Reddit
- Shrinking the Generation-Verification Gap with Weak Verifiers
- GaRAGe: A Benchmark with Grounding Annotations for RAG Evaluation
- A GenAI System for Improved FAIR Independent Biological Database Integration
- QueueEDIT: Structural Self-Correction for Sequential Model Editing in LLMs
- TreeReview: A Dynamic Tree of Questions Framework for Deep and Efficient LLM-based Scientific Peer Review
- MEMOIR: Lifelong Model Editing with Minimal Overwrite and Informed Retention for LLMs
- Towards a Unified Textual Graph Framework for Spectral Reasoning via Physical and Chemical Information Fusion
- HIDE and Seek: Detecting Hallucinations in Language Models via Decoupled Representations
- Federated In-Context Learning: Iterative Refinement for Improved Answer Quality
- From Raw Corpora to Domain Benchmarks: Automated Evaluation of LLM Domain Expertise
- V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos
- Differentiation-Based Extraction of Proprietary Data from Fine-Tuned LLMs
- Open World Scene Graph Generation using Vision Language Models
- Semantic Outlier Removal with Embedding Models and LLMs
- Do We Talk to Robots Like Therapists, and Do They Respond Accordingly? Language Alignment in AI Emotional Support
- Advancing Automated Speaking Assessment Leveraging Multifaceted Relevance and Grammar Information
- SGIC: A Self-Guided Iterative Calibration Framework for RAG
- Double Entendre: Robust Audio-Based AI-Generated Lyrics Detection via Multi-View Fusion
- Drag-and-Drop LLMs: Zero-Shot Prompt-to-Weights
- Hierarchical Patch Compression for ColPali: Efficient Multi-Vector Document Retrieval with Dynamic Pruning and Quantization
- ProtoTransformer: A Meta-Learning Approach to Providing Student Feedback
- Vision-Guided Chunking Is All You Need: Enhancing RAG with Multimodal Document Understanding
- Revela: Dense Retriever Learning via Language Modeling
- REIS: A High-Performance and Energy-Efficient Retrieval System with In-Storage Processing
- Privacy-Preserving in Connected and Autonomous Vehicles Through Vision to Text Transformation
- DiscoSG: Towards Discourse-Level Text Scene Graph Parsing through Iterative Graph Refinement
- SafeMimic: Towards Safe and Autonomous Human-to-Robot Imitation for Mobile Manipulation
- MoR: Better Handling Diverse Queries with a Mixture of Sparse, Dense, and Human Retrievers
- COSMMIC: Comment-Sensitive Multimodal Multilingual Indian Corpus for Summarization and Headline Generation
- Understanding Verbatim Memorization in LLMs Through Circuit Discovery
- Issue Retrieval and Verification Enhanced Supplementary Code Comment Generation
- Exploring MLLMs Perception of Network Visualization Principles
- Can Vision Language Models Understand Mimed Actions?
- LexiMark: Robust Watermarking via Lexical Substitutions to Enhance Membership Verification of an LLM's Textual Training Data
- InsertRank: LLMs can reason over BM25 scores to Improve Listwise Reranking
- CrEst: Credibility Estimation for Contexts in LLMs via Weak Supervision
- Expectation Confirmation Preference Optimization for Multi-Turn Conversational Recommendation Agent
- Reasoning with Exploration: An Entropy Perspective
- Into the Unknown: Applying Inductive Spatial-Semantic Location Embeddings for Predicting Individuals' Mobility Beyond Visited Places
- SCISSOR: Mitigating Semantic Bias through Cluster-Aware Siamese Networks for Robust Classification
- GenerationPrograms: Fine-grained Attribution with Executable Programs
- K/DA: Automated Data Generation Pipeline for Detoxifying Implicitly Offensive Language in Korean
- SPOT: Bridging Natural Language and Geospatial Search for Investigative Journalists
- A Game-Theoretic Negotiation Framework for Cross-Cultural Consensus in LLMs
- EmoNews: A Spoken Dialogue System for Expressive News Conversations
- Bridging Unsupervised and Semi-Supervised Anomaly Detection: A Theoretically-Grounded and Practical Framework with Synthetic Anomalies
- ASMR: Augmenting Life Scenario using Large Generative Models for Robotic Action Reflection
- FinLMM-R1: Enhancing Financial Reasoning in LMM through Scalable Data and Reward Design
- Hone as You Read: A Practical Type of Interactive Summarization
- Adversarial Paraphrasing: A Universal Attack for Humanizing AI-Generated Text
- eLog analysis for accelerators: status and future outlook
- Semantic-preserved Augmentation with Confidence-weighted Fine-tuning for Aspect Category Sentiment Analysis
- Assessing the Performance Gap Between Lexical and Semantic Models for Information Retrieval With Formulaic Legal Language
- Assessing the Role of Data Quality in Training Bilingual Language Models
- CORONA: A Coarse-to-Fine Framework for Graph-based Recommendation with Large Language Models
- Enabling Precise Topic Alignment in Large Language Models Via Sparse Autoencoders
- Language Surgery in Multilingual Large Language Models
- CLIP the Landscape: Automated Tagging of Crowdsourced Landscape Images
- Generative Representational Learning of Foundation Models for Recommendation
- Post Persona Alignment for Multi-Session Dialogue Generation
- Unsupervised Document and Template Clustering using Multimodal Embeddings
- AgentSense: Virtual Sensor Data Generation Using LLM Agents in Simulated Home Environments
- Chunk Twice, Embed Once: A Systematic Study of Segmentation and Representation Trade-offs in Chemistry-Aware Retrieval-Augmented Generation
- Improving Large Language Model Safety with Contrastive Representation Learning
- Large Language Models for History, Philosophy, and Sociology of Science: Interpretive Uses, Methodological Challenges, and Critical Perspectives
- DiscoSum: Discourse-aware News Summarization
- RETUYT-INCO at BEA 2025 Shared Task: How Far Can Lightweight Models Go in AI-powered Tutor Evaluation?
- Graph Neural Networks for Automatic Addition of Optimizing Components in Printed Circuit Board Schematics
- Reliable Reasoning Path: Distilling Effective Guidance for LLM Reasoning with Knowledge Graphs
- Flick: Few Labels Text Classification using K-Aware Intermediate Learning in Multi-Task Low-Resource Languages
- Combining Log Data and Collaborative Dialogue Features to Predict Project Quality in Middle School AI Education
- Reinforcement learning fine-tuning of language model for instruction following and math reasoning
- Q2E: Query-to-Event Decomposition for Zero-Shot Multilingual Text-to-Video Retrieval
- KI4Demokratie: An AI-Based Platform for Monitoring and Fostering Democratic Discourse
- Query-Focused Retrieval Heads Improve Long-Context Reasoning and Re-ranking
- Enhancing Traffic Accident Classifications: Application of NLP Methods for City Safety
- Label-semantics Aware Generative Approach for Domain-Agnostic Multilabel Classification
- Improving LLM-Powered EDA Assistants with RAFT
- Revealing Weaknesses in Text Watermarking Through Self-Information Rewrite Attacks
- Tau-Eval: A Unified Evaluation Framework for Useful and Private Text Anonymization
- Object Navigation with Structure-Semantic Reasoning-Based Multi-level Map and Multimodal Decision-Making LLM
- Large Language Models are Demonstration Pre-Selectors for Themselves
- Towards an Explainable Comparison and Alignment of Feature Embeddings
- Tokenization Multiplicity Leads to Arbitrary Price Variation in LLM-as-a-service
- Can LLMs Express Personality Across Cultures? Introducing CulturalPersonas for Evaluating Trait Alignment
- CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval
- DynamicMind: A Tri-Mode Thinking System for Large Language Models
- Conformal Prediction Adaptive to Unknown Subpopulation Shifts
- Urania: Differentially Private Insights into AI Use
- Static Word Embeddings for Sentence Semantic Representation
- Joint Evaluation of Answer and Reasoning Consistency for Hallucination Detection in Large Reasoning Models
- Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models
- Mechanistic Decomposition of Sentence Representations
- Red Teaming the Mind of the Machine: A Systematic Evaluation of Prompt Injection and Jailbreak Vulnerabilities in LLMs
- SweRank: Software Issue Localization with Code Ranking
- Benchmarking LLMs' Swarm intelligence
- Agentic Graph Token Reasoning
- Fine-Tuning Large Language Models and Evaluating Retrieval Methods for Improved Question Answering on Building Codes
- GASCADE: Grouped Summarization of Adverse Drug Event for Enhanced Cancer Pharmacovigilance
- Theoretical Guarantees for LT-TTD: A Unified Transformer-based Architecture for Two-Level Ranking Systems
- Steerable Chatbots: Exploring Personalization Control Interfaces via LLM Activation Steering
- Video-Skill-CoT: Skill-based Chain-of-Thoughts for Domain-Adaptive Video Reasoning
- APT: Improving Specialist LLM Performance with Weakness Case Acquisition and Iterative Preference Training
- Leveraging Reward Models for Guiding Code Review Comment Generation
- The Impact of COVID-19 on Twitter Ego Networks: Structure, Sentiment, and Topics
- Enhancing Text Comprehension for Dyslexic Readers: A 3D Semantic Visualization Approach Using Transformer Mode
- Magic Mushroom: A Customizable Benchmark for Fine-grained Analysis of Retrieval Noise Erosion in RAG Systems
- LeanExplore: A search engine for Lean 4 declarations
- Compositional Generalisation for Explainable Hate Speech Detection
- Self-Supervised Contrastive Learning is Approximately Supervised Contrastive Learning
- ReXVQA: A Large-scale Visual Question Answering Benchmark for Generalist Chest X-ray Understanding
- Say It Another Way: Auditing LLMs with a User-Grounded Automated Paraphrasing Framework
- ChineseBERT: Chinese Pretraining Enhanced by Glyph and Pinyin Information
- DRE: An Effective Dual-Refined Method for Integrating Small and Large Language Models in Open-Domain Dialogue Evaluation
- Go Beyond Earth: Understanding Human Actions and Scenes in Microgravity Environments
- In-context Clustering-based Entity Resolution with Large Language Models: A Design Space Exploration
- OS-W2S: An Automatic Labeling Engine for Language-Guided Open-Set Aerial Object Detection
- Building a Recommendation System Using Amazon Product Co-Purchasing Network
- Abstract Counterfactuals for Language Model Agents
- A Chaos Driven Metric for Backdoor Attack Detection
- RAVU: Retrieval Augmented Video Understanding with Compositional Reasoning over Graph
- AUTOCIRCUIT-RL: Reinforcement Learning-Driven LLM for Automated Circuit Topology Generation
- Leaps Beyond the Seen: Reinforced Reasoning Augmented Generation for Clinical Notes
- Enriching Location Representation with Detailed Semantic Information
- INESC-ID @ eRisk 2025: Exploring Fine-Tuned, Similarity-Based, and Prompt-Based Approaches to Depression Symptom Identification
- Cycle Consistency as Reward: Learning Image-Text Alignment without Human Preferences
- LLM in the Loop: Creating the ParaDeHate Dataset for Hate Speech Detoxification
- Agentic Episodic Control
- Redundancy, Isotropy, and Intrinsic Dimensionality of Prompt-based Text Embeddings
- GLoSS: Generative Language Models with Semantic Search for Sequential Recommendation
- Quantifying Misattribution Unfairness in Authorship Attribution
- Building Entity Association Mining Framework for Knowledge Discovery
- When LLMs Team Up: The Emergence of Collaborative Affective Computing
- Integration of Large Language Models and Traditional Deep Learning for Social Determinants of Health Prediction
- DRAG: Distilling RAG for SLMs from LLMs to Transfer Knowledge and Mitigate Hallucination via Evidence and Graph-based Distillation
- SVQA-R1: Reinforcing Spatial Reasoning in MLLMs via View-Consistent Reward Optimization
- Hybrid AI for Responsive Multi-Turn Online Conversations with Novel Dynamic Routing and Feedback Adaptation
- Memory-Efficient FastText: A Comprehensive Approach Using Double-Array Trie Structures and Mark-Compact Memory Management
- Synthline: A Product Line Approach for Synthetic Requirements Engineering Data Generation using Large Language Models
- WoMAP: World Models For Embodied Open-Vocabulary Object Localization
- ChemAU: Harness the Reasoning of LLMs in Chemical Research with Adaptive Uncertainty Estimation
- Talking to Data: Designing Smart Assistants for Humanities Databases
- Multimodal Fusion with Semi-Supervised Learning Minimizes Annotation Quantity for Modeling Videoconference Conversation Experience
- Chandomitra: Towards Generating Structured Sanskrit Poetry from Natural Language Inputs
- EEG2TEXT-CN: An Exploratory Study of Open-Vocabulary Chinese Text-EEG Alignment via Large Language Model and Contrastive Learning on ChineseEEG
- NBF at SemEval-2025 Task 5: Light-Burst Attention Enhanced System for Multilingual Subject Recommendation
- Dynamic Chunking and Selection for Reading Comprehension of Ultra-Long Context in Large Language Models
- Adapting General-Purpose Embedding Models to Private Datasets Using Keyword-based Retrieval
- Survey of Abstract Meaning Representation: Then, Now, Future
- Efficient Latent Semantic Clustering for Scaling Test-Time Computation of LLMs
- ArtiScene: Language-Driven Artistic 3D Scene Generation Through Image Intermediary
- Improving Dialogue State Tracking through Combinatorial Search for In-Context Examples
- Test-time Vocabulary Adaptation for Language-driven Object Detection
- Uncertainty-Aware Large Language Models for Explainable Disease Diagnosis
- LegalEval-Q: A New Benchmark for The Quality Evaluation of LLM-Generated Legal Text
- Don't Reinvent the Wheel: Efficient Instruction-Following Text Embedding based on Guided Space Transformation
- Reducing Annotation Burden in Physical Activity Research Using Vision-Language Models
- Multi-Domain ABSA Conversation Dataset Generation via LLMs for Real-World Evaluation and Model Comparison
- PRISM: A Framework for Producing Interpretable Political Bias Embeddings with Political-Aware Cross-Encoder
- Harnessing Large Language Models for Scientific Novelty Detection
- Bench4KE: Benchmarking Automated Competency Question Generation
- Domain Pre-training Impact on Representations
- ERU-KG: Efficient Reference-aligned Unsupervised Keyphrase Generation
- LKD-KGC: Domain-Specific KG Construction via LLM-driven Knowledge Dependency Parsing
- GATE: General Arabic Text Embedding for Enhanced Semantic Textual Similarity with Matryoshka Representation Learning and Hybrid Loss Training
- MIR: Methodology Inspiration Retrieval for Scientific Research Problems
- Context is Gold to find the Gold Passage: Evaluating and Training Contextual Document Embeddings
- Hierarchical Level-Wise News Article Clustering via Multilingual Matryoshka Embeddings
- Hi-Dyna Graph: Hierarchical Dynamic Scene Graph for Robotic Autonomy in Human-Centric Environments
- Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models
- CoRet: Improved Retriever for Code Editing
- TCM-Ladder: A Benchmark for Multimodal Question Answering on Traditional Chinese Medicine
- Hidden Persuasion: Detecting Manipulative Narratives on Social Media During the 2022 Russian Invasion of Ukraine
- MSQA: Benchmarking LLMs on Graduate-Level Materials Science Reasoning and Knowledge
- SLOT: Structuring the Output of Large Language Models
- Mis-prompt: Benchmarking Large Language Models for Proactive Error Handling
- From Parameters to Prompts: Understanding and Mitigating the Factuality Gap between Fine-Tuned LLMs
- Prompt Engineer: Analyzing Skill Requirements in the AI Job Market
- WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch
- Augment or Not? A Comparative Study of Pure and Augmented Large Language Model Recommenders
- EmotionTalk: An Interactive Chinese Multimodal Emotion Dataset With Rich Annotations
- TailorSQL: An NL2SQL System Tailored to Your Query Workload
- ToolHaystack: Stress-Testing Tool-Augmented Language Models in Realistic Long-Term Interactions
- Automatic Construction of Multiple Classification Dimensions for Managing Approaches in Scientific Papers
- Threading the Needle: Reweaving Chain-of-Thought Reasoning to Explain Human Label Variation
- EVOREFUSE: Evolutionary Prompt Optimization for Evaluation and Mitigation of LLM Over-Refusal to Pseudo-Malicious Instructions
- GeNRe: A French Gender-Neutral Rewriting System Using Collective Nouns
- Event-aware analysis of cross-city visitor flows using large language models and social media data
- RAGRouter: Learning to Route Queries to Multiple Retrieval-Augmented Language Models
- DeepTheorem: Advancing LLM Reasoning for Theorem Proving Through Natural Language and Reinforcement Learning
- LoVeC: Reinforcement Learning for Better Verbalized Confidence in Long-Form Generations
- ZeroGUI: Automating Online GUI Learning at Zero Human Cost
- Self-Critique and Refinement for Faithful Natural Language Explanations
- Speaking images. A novel framework for the automated self-description of artworks
- Multilingual vs Crosslingual Retrieval of Fact-Checked Claims: A Tale of Two Approaches
- Retweets, Receipts, and Resistance: Discourse, Sentiment, and Credibility in Public Health Crisis Twitter
- Latent Reasoning via Sentence Embedding Prediction
- Rethinking Hybrid Retrieval: When Small Embeddings and LLM Re-ranking Beat Bigger Models
- Pre-Training Curriculum for Multi-Token Prediction in Language Models
- Say What You Mean: Natural Language Access Control with Large Language Models for Internet of Things
- GuessArena: Guess Who I Am? A Self-Adaptive Framework for Evaluating LLMs in Domain-Specific Knowledge and Reasoning
- Beyond Completion: A Foundation Model for General Knowledge Graph Reasoning
- Sentence Analogies: Exploring Linguistic Relationships and Regularities in Sentence Embeddings
- Precise In-Parameter Concept Erasure in Large Language Models
- Tevatron 2.0: Unified Document Retrieval Toolkit across Scale, Language, and Modality
- NOCL: Node-Oriented Conceptualization LLM for Graph Tasks without Message Passing
- Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions
- LASER: Stratified Selective Sampling for Instruction Tuning with Dedicated Scoring Strategy
- LiTEx: A Linguistic Taxonomy of Explanations for Understanding Within-Label Variation in Natural Language Inference
- StreamLink: Large-Language-Model Driven Distributed Data Engineering System
- PARTONOMY: Large Multimodal Models with Part-Level Visual Understanding
- Enhancing Transformation from Natural Language to Signal Temporal Logic Using LLMs with Diverse External Knowledge
- STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models
- A Stereotype Content Analysis on Color-related Social Bias in Large Vision Language Models
- Creativity in LLM-based Multi-Agent Systems: A Survey
- Test-Time Learning for Large Language Models
- Aligning Proteins and Language: A Foundation Model for Protein Retrieval
- How does Misinformation Affect Large Language Model Behaviors and Preferences?
- Disentangling Locality and Entropy in Ranking Distillation
- Something's Fishy In The Data Lake: A Critical Re-evaluation of Table Union Search Benchmarks
- Requirements Coverage-Guided Minimization for Natural Language Test Cases
- ALAS: An Automatic Latent Alignment Score for Audio Language Models
- When Prompts Control Robots: Prompt Injection Attacks in Multi-Agent Robotic Systems
- Towards Effective Federated Multimodal Graph Learning via Navigating Multifaceted Heterogeneity
- Benchmarking and Enhancing LLM Agents in Localizing Linux Kernel Bugs
- The Role of Diversity in In-Context Learning for Large Language Models
- Learning to Select In-Context Demonstration Preferred by Large Language Model
- Self-Reflective Planning with Knowledge Graphs: Enhancing LLM Reasoning Reliability for Question Answering
- Simple and Effective Baselines for Code Summarisation Evaluation
- Multimodal Emotion Recognition in Conversations: A Survey of Methods, Trends, Challenges and Prospects
- Multi-View Encoders for Performance Prediction in LLM-Based Agentic Workflows
- SEMMA: A Semantic Aware Knowledge Graph Foundation Model
- REARANK: Reasoning Re-ranking Agent via Reinforcement Learning
- MA-RAG: Multi-Agent Retrieval-Augmented Generation via Collaborative Chain-of-Thought Reasoning
- Gatsby Without the 'E': Crafting Lipograms with LLMs
- FoodTaxo: Generating Food Taxonomies with Large Language Models
- An Automated LLM-based Pipeline for Asset-Level Database Creation to Assess Deforestation Impact
- On Robustness of Neural Semantic Parsers
- Invoke Interfaces Only When Needed: Adaptive Invocation for Large Language Models in Question Answering
- Learning Extrapolative Sequence Transformations from Markov Chains
- CoTGuard: Using Chain-of-Thought Triggering for Copyright Protection in Multi-Agent LLM Systems
- PatentMind: A Multi-Aspect Reasoning Graph for Patent Similarity Evaluation
- ActiveDPO: Active Direct Preference Optimization for Sample-Efficient Alignment
- LLLMs: A Data-Driven Survey of Evolving Research on Limitations of Large Language Models
- From Course to Skill: Evaluating LLM Performance in Curricular Analytics
- Evaluating Text Creativity across Diverse Domains: A Dataset and Large Language Model Evaluator
- Learning to Explain: Prototype-Based Surrogate Models for LLM Classification
- Unveiling Dual Quality in Product Reviews: An NLP-Based Approach
- POQD: Performance-Oriented Query Decomposer for Multi-vector retrieval
- Optimized Text Embedding Models and Benchmarks for Amharic Passage Retrieval
- Knoll: Creating a Knowledge Ecosystem for Large Language Models
- Sci-LoRA: Mixture of Scientific LoRAs for Cross-Domain Lay Paraphrasing
- Reasoning Segmentation for Images and Videos: A Survey
- Enhancing Training Data Attribution with Representational Optimization
- Distinctive Feature Codec: An Adaptive Efficient Speech Representation for Depression Detection
- TAG-INSTRUCT: Controlled Instruction Complexity Enhancement through Structure-based Augmentation
- Hard Negative Mining for Domain-Specific Retrieval in Enterprise Systems
- Guided by Gut: Efficient Test-Time Scaling with Reinforced Intrinsic Confidence
- Frankentext: Stitching random text fragments into long-form narratives
- ProgRM: Build Better GUI Agents with Progress Rewards
- MetaGen Blended RAG: Unlocking Zero-Shot Precision for Specialized Domain Question-Answering
- Locality-Sensitive Hashing for Efficient Hard Negative Sampling in Contrastive Learning
- Multimodal Embeddings for 3D Similarity Search in Semantic Web-of-Things Digital-Twin Platforms
- Hybrid Encoder: Towards Efficient and Precise Native AdsRecommendation via Hybrid Transformer Encoding Networks
- DocNavRAG: Document-Structured Graph RAG with Stateful Evidence Construction for Complex Document Question Answering
- Auditing Data Provenance in LLM Fine-tuning via Intrinsic Distributional Fingerprints
- SemSketches-2021: experimenting with the machine processing of the pilot semantic sketches corpus
- Surfacing Semantic Orthogonality Across Model Safety Benchmarks: A Multi-Dimensional Analysis
- Evidence-Grounded Multimodal Misinformation Detection with Attention-Based GNNs
- Reverse-Speech-Finder: A Neural Network Backtracking Architecture for Generating Alzheimer's Disease Speech Samples and Improving Diagnosis Performance
- HydraRAG: Structured Cross-Source Enhanced Large Language Model Reasoning
- Taming LLMs with Negative Samples: A Reference-Free Framework to Evaluate Presentation Content with Actionable Feedback
- Dynamic Bundling with Large Language Models for Zero-Shot Inference on Text-Attributed Graphs
- Learning Shared Representations from Unpaired Data
- Evaluating NLP Embedding Models for Handling Science-Specific Symbolic Expressions in Student Texts
- Refusal Direction is Universal Across Safety-Aligned Languages
- BP-Seg: A graphical model approach to unsupervised and non-contiguous text segmentation using belief propagation
- CAIN: Hijacking LLM-Humans Conversations via Malicious System Prompts
- SPEAR: Selection-aware Personalized End-to-end Adaptive Rewriting and Retrieval for Community Search
- Lost in Permissions: Exploring the Microsoft 365 App Ecosystem
- Causal-Invariant Cross-Domain Out-of-Distribution Recommendation
- EquivPruner: Boosting Efficiency and Quality in LLM-Based Search via Action Pruning
- Cog-TiPRO: Iterative Prompt Refinement with LLMs to Detect Cognitive Decline via Longitudinal Voice Assistant Commands
- Foundation Models for Geospatial Reasoning: Assessing Capabilities of Large Language Models in Understanding Geometries and Topological Spatial Relations
- MuseScorer: Idea Originality Scoring At Scale
- Relative Bias: A Comparative Framework for Quantifying Bias in LLMs
- LLM-Based Emulation of the Radio Resource Control Layer: Towards AI-Native RAN Protocols
- MAPLE: Many-Shot Adaptive Pseudo-Labeling for In-Context Learning
- Syntax Meets Semantics: Understanding Scientific Formulae
- Silent Leaks: Implicit Knowledge Extraction Attack on RAG Systems through Benign Queries
- SCENIR: Visual Semantic Clarity through Unsupervised Scene Graph Retrieval
- Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models
- Nek Minit: Harnessing Pragmatic Metacognitive Prompting for Explainable Sarcasm Detection of Australian and Indian English
- Are the confidence scores of reviewers consistent with the review content? Evidence from top conference proceedings in AI
- CRAFT: Training-Free Cascaded Retrieval for Tabular QA
- VERDI: VLM-Embedded Reasoning for Autonomous Driving
- RePPL: Recalibrating Perplexity by Uncertainty in Semantic Propagation and Language Generation for Explainable QA Hallucination Detection
- Ranking Free RAG: Replacing Re-ranking with Selection in RAG for Sensitive Domains
- Transformer-Based Models for Question Answering on COVID19
- On the Generalization vs Fidelity Paradox in Knowledge Distillation
- EcomScriptBench: A Multi-task Benchmark for E-commerce Script Planning via Step-wise Intention-Driven Product Association
- Multilingual Test-Time Scaling via Initial Thought Transfer
- Blind Spot Navigation: Evolutionary Discovery of Sensitive Semantic Concepts for LVLMs
- Adaptive Plan-Execute Framework for Smart Contract Security Auditing
- R&D-Agent-Quant: A Multi-Agent Framework for Data-Centric Factors and Model Joint Optimization
- R-TOFU: Unlearning in Large Reasoning Models
- MaxPoolBERT: Enhancing BERT Classification via Layer- and Token-Wise Aggregation
- SciCUEval: A Comprehensive Dataset for Evaluating Scientific Context Understanding in Large Language Models
- SLMEval: Entropy-Based Calibration for Human-Aligned Evaluation of Large Language Models
- Auditing Semantic Gains in Sequential Recommendation: A Lightweight Recovery Test
- Guarded Query Routing for Large Language Models
- The Limits of Graph Samplers for Training Inductive Recommender Systems: Extended results
- Enhancing Abstractive Summarization of Scientific Papers Using Structure Information
- Local Mixtures of Experts: Essentially Free Test-Time Training via Model Merging
- A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations
- Table Foundation Models: on knowledge pre-training for tabular learning
- Studying the Role of Input-Neighbor Overlap in Retrieval-Augmented Language Models Training Efficiency
- MoMoE: Mixture of Moderation Experts Framework for AI-Assisted Online Governance
- From Templates to Natural Language: Generalization Challenges in Instruction-Tuned LLMs for Spatial Reasoning
- Graph-Guided Passage Retrieval for Author-Centric Structured Feedback
- MultiHal: Multilingual Dataset for Knowledge-Graph Grounded Evaluation of LLM Hallucinations
- Invisible Entropy: Towards Safe and Efficient Low-Entropy LLM Watermarking
- Reasoning Path Compression: Compressing Generation Trajectories for Efficient LLM Reasoning
- Towards Omnidirectional Reasoning with 360-R1: A Dataset, Benchmark, and GRPO-based Method
- Strategic Planning and Rationalizing on Trees Make LLMs Better Debaters
- U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding
- SEPS: A Separability Measure for Robust Unlearning in LLMs
- MAFA: A multi-agent framework for annotation
- Cross-Lingual Representation Alignment Through Contrastive Image-Caption Tuning
- Scalable Bayesian Monte Carlo: fast uncertainty estimation beyond deep ensembles
- New Encoders for German Trained from Scratch: Comparing ModernGBERT with Converted LLM2Vec Models
- GuRE:Generative Query REwriter for Legal Passage Retrieval
- CPRet: A Dataset, Benchmark, and Model for Retrieval in Competitive Programming
- Prompt Stability Matters: Evaluating and Optimizing Auto-Generated Prompt in General-Purpose Systems
- A Case Study of Cross-Lingual Zero-Shot Generalization for Classical Languages in LLMs
- Predicting Reaction Time to Comprehend Scenes with Foveated Scene Understanding Maps
- CS-Sum: A Benchmark for Code-Switching Dialogue Summarization and the Limits of Large Language Models
- CoIn: Counting the Invisible Reasoning Tokens in Commercial Opaque LLM APIs
- Towards Effective Federated Graph Foundation Model via Mitigating Knowledge Entanglement
- Real-Time Hybrid Retrieval in Hyperbolic Space for Retrieval-Augmented Generation on Edge Devices
- Retrieval Augmented Biomedical Question Answering with Weak Question Recovery and Neural Reranking for BioASQ Task 14b
- CoRank: LLM-Based Compact Reranking with Document Features for Scientific Retrieval
- Adapting to LLMs: How Insiders and Outsiders Reshape Scientific Knowledge Production
- Conceptual Modeling: Topics, Themes, and Technology Trends
- AutoMathKG: The automated mathematical knowledge graph based on LLM and vector database
- JIR-Arena: The First Benchmark Dataset for Just-in-time Information Recommendation
- GUARD: Generation-time LLM Unlearning via Adaptive Restriction and Detection
- CoT-Kinetics: A Theoretical Modeling Assessing LRM Reasoning Process
- topicwizard -- a Modern, Model-agnostic Framework for Topic Model Visualization and Interpretation
- V-Mem: Modality-Routed Retrieval for Long-Term Multimodal Agentic Memory
- Towards Budget-Friendly Model-Agnostic Explanation Generation for Large Language Models
- Cost-effective Deployment of BERT Models in Serverless Environment
- Doc2CI: A Multi-Service Study of CI Configuration Generation Using Large Language Models
- No Free Lunch in Active Learning: LLM Embedding Quality Dictates Query Strategy Success
- ParticleGen: A Multi-Agent System for Particle Effects Generation
- LAMeTA: Intent-Aware Agentic Network Optimization via a Large AI Model-Empowered Two-Stage Approach
- Model alignment using inter-modal bridges
- SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization
- The AI Gap: How Socioeconomic Status Affects Language Technology Interactions
- InteracVid: Building a Real Interactive Audio-Visual Response Dataset from Live-Chat Videos
- Class Distillation with Mahalanobis Contrast: An Efficient Training Paradigm for Pragmatic Language Understanding Tasks
- Retrospex: Language Agent Meets Offline Reinforcement Learning Critic
- Counterspeech the ultimate shield! Multi-Conditioned Counterspeech Generation through Attributed Prefix Learning
- TechniqueRAG: Retrieval Augmented Generation for Adversarial Technique Annotation in Cyber Threat Intelligence Text
- Relation-Aware Graph Foundation Model
- ELITE: Embedding-Less retrieval with Iterative Text Exploration
- Are vision language models robust to uncertain inputs?
- AgenTag: Attribution of AI Coding Agents from Behavioral Fingerprints
- Towards Universal Semantics With Large Language Models
- ChartEdit: How Far Are MLLMs From Automating Chart Analysis? Evaluating MLLMs' Capability via Chart Editing
- Demystifying and Enhancing the Efficiency of Large Language Model Based Search Agents
- No Gold Standard, No Problem: Reference-Free Evaluation of Taxonomies
- ProxyPrompt: Securing System Prompts against Prompt Extraction Attacks
- TI-StegoAlign: Channel-Guided Post-Training for Generative Text Steganography under Tokenization Inconsistency
- ACSE-Eval: Can LLMs threat model real-world cloud infrastructure?
- CeQe: Grounding Lexical Retrieval in Semantic Evidence
- MTRE: Multi-Token Reliability Estimation for Hallucination Detection in VLMs
- Can Global XAI Methods Reveal Injected Behaviours in LLMs? SHAP vs Rule Extraction vs RuleSHAP
- TAIJI: MCP-based Multi-Modal Data Analytics on Data Lakes
- SubGCache: Accelerating Graph-based RAG with Subgraph-level KV Cache
- Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models
- Benchmarking Critical Questions Generation: A Challenging Reasoning Task for Large Language Models
- ReWiND: Language-Guided Rewards Teach Robot Policies without New Demonstrations
- The Future is Sparse: Embedding Compression for Scalable Retrieval in Recommender Systems
- Structure-Aware Semantic Chunking with Title-Chain Prefixes: A 1600-Query Evaluation and the Measurement Trap in Text-Transform Ablations
- Temporally-Grounded Language Generation: A Benchmark for Real-Time Vision-Language Models
- Cochain: Balancing Insufficient and Excessive Collaboration in LLM Agent Workflows
- Model Performance-Guided Evaluation Data Selection for Effective Prompt Optimization
- GeoGrid-Bench: Can Foundation Models Understand Multimodal Gridded Geo-Spatial Data?
- Real-Time Out-of-Distribution Failure Prevention via Multi-Modal Reasoning
- LDIR: Low-Dimensional Dense and Interpretable Text Embeddings with Relative Representations
- Hierarchical Document Refinement for Long-context Retrieval-augmented Generation
- Top-Down vs. Bottom-Up Approaches for Automatic Educational Knowledge Graph Construction in CourseMapper
- Evaluation and Explainability of Unsupervised Scholarly Collaboration Recommendations
- SemEval-2025 Task 7: Multilingual and Crosslingual Fact-Checked Claim Retrieval
- SafePath: Conformal Prediction for Safe LLM-Based Autonomous Navigation
- Evidence-Unit Fairness and the Limits of Query-Adaptive Sparse-Dense Fusion in Financial Document Retrieval
- Ornithologist: Towards Trustworthy "Reasoning" about Central Bank Communications
- An Efficient deep learning model to Predict Stock Price Movement Based on Limit Order Book
- LLM4CD: Leveraging Large Language Models for Open-World Knowledge Augmented Cognitive Diagnosis
- LoRA-MME: Multi-Model Ensemble of LoRA-Tuned Encoders for Code Comment Classification
- H+ Embedding: Harmonizing Global and Token-Level Retrieval with Context-Dependent Phrases
- RAG-TESTER: Automated End-to-End Testing of Retrieval-Augmented Large Language Models
- Automatic Task Detection and Heterogeneous LLM Speculative Decoding
- Enhancing Cache-Augmented Generation (CAG) with Adaptive Contextual Compression for Scalable Knowledge Integration
- DSADF: Thinking Fast and Slow for Decision Making
- Optimization and Constraint Modeling using LLMs with a Retrieval Augmented Generation Process
- Hakim: Farsi Text Embedding Model
- Assessing and Mitigating Medical Knowledge Drift and Conflicts in Large Language Models
- Reproducibility, Replicability, and Insights into Visual Document Retrieval with Late Interaction
- A Reproduction Study: The Kernel PCA Interpretation of Self-Attention Fails Under Scrutiny
- Multi-Domain Audio Question Answering Toward Acoustic Content Reasoning in The DCASE 2025 Challenge
- HAMLET: Healthcare-focused Adaptive Multilingual Learning Embedding-based Topic Modeling
- Pre-training vs. Fine-tuning: A Reproducibility Study on Dense Retrieval Knowledge Acquisition
- MedEIR: A Specialized Medical Embedding Model for Enhanced Information Retrieval
- Chronocept: Instilling a Sense of Time in Machines
- Matching Tasks with Industry Groups for Augmenting Commonsense Knowledge
- Towards Multi-Agent Reasoning Systems for Collaborative Expertise Delegation: An Exploratory Design Study
- GRADA: Graph-based Reranking against Adversarial Documents Attack
- KOKKAI DOC: An LLM-driven framework for scaling parliamentary representatives
- Optimizing Recommendations using Fine-Tuned LLMs
- Bridging AI and Carbon Capture: A Dataset for LLMs in Ionic Liquids and CBE Research
- People Are Highly Cooperative with Large Language Models, Especially When Communication Is Possible or Following Human Interaction
- Behind the Byline: A Large-Scale Study of Scientific Author Contributions
- Enhancing BERTopic with Intermediate Layer Representations
- An Analysis Focused on Womens Safety: Can VAD Models Be Enhanced by a Multi-modal Dataset?
- References Indeed Matter? Reference-Free Preference Optimization for Conversational Query Reformulation
- Weakly Supervised Temporal Sentence Grounding via Positive Sample Mining
- KCluster: An LLM-based Clustering Approach to Knowledge Component Discovery
- LOG-Nav: Efficient Layout-Aware Object-Goal Navigation with Hierarchical Planning
- Estimating Quality in Therapeutic Conversations: A Multi-Dimensional Natural Language Processing Framework
- Query-driven Document-level Scientific Evidence Extraction from Biomedical Studies
- PromptIQ: Who Cares About Prompts? Let System Handle It -- A Component-Aware Framework for T2I Generation
- Exploring the Feasibility of Multilingual Grammatical Error Correction with a Single LLM up to 9B parameters: A Comparative Study of 17 Models
- Embedding Atlas: Low-Friction, Interactive Embedding Visualization
- Courtroom-Style Multi-Agent Debate with Progressive RAG and Role-Switching for Controversial Claim Verification
- MAPLE: Metadata Augmented Private Language Evolution
- Unpacking Robustness in Inflectional Languages: Adversarial Evaluation and Mechanistic Insights
- Scalable Multi-Stage Influence Function for Large Language Models via Eigenvalue-Corrected Kronecker-Factored Parameterization
- The Rise of Language Models in Mining Software Repositories: A Survey
- Generalization Analysis for Supervised Contrastive Representation Learning under Non-IID Settings
- An Open-Source Dual-Loss Embedding Model for Semantic Retrieval in Higher Education
- QBR: A Question-Bank-Based Approach to Fine-Grained Legal Knowledge Retrieval for the General Public
- CORE: Collaborative Reasoning via Cross Teaching
- When Silence Is Golden: Can LLMs Learn to Abstain in Temporal QA and Beyond?
- Embedding based retrieval for long tail search queries in ecommerce
- Exploring the Role of Diversity in Example Selection for In-Context Learning
- Enhancing the Learning Experience: Using Vision-Language Models to Generate Questions for Educational Videos
- DepreSym: A Depression Symptom Annotated Corpus and the Role of Large Language Models as Assessors of Psychological Markers
- CAMOUFLAGE: Exploiting Misinformation Detection Systems Through LLM-driven Adversarial Claim Transformation
- CO-Search: COVID-19 Information Retrieval with Semantic Search, Question Answering, and Abstractive Summarization
- Pingala: Prosody-Aware Decoding for Sanskrit Poetry Generation
- Mapping the Climate Change Landscape on TikTok
- MateICL: Mitigating Attention Dispersion in Large-Scale In-Context Learning
- Complex totopapa: predicting the successor to pope Francis
- Harnessing Structured Knowledge: A Concept Map-Based Approach for High-Quality Multiple Choice Question Generation with Effective Distractors
- Efficient Vocabulary-Free Fine-Grained Visual Recognition in the Age of Multimodal LLMs
- Scalable Unit Harmonization in Medical Informatics via Bayesian-Optimized Retrieval and Transformer-Based Re-ranking
- Steering Large Language Models with Register Analysis for Arbitrary Style Transfer
- The Price of Meaning: Why Every Semantic Memory System Forgets
- Sustainable Hybrid Document-Routed Retrieval for Financial RAG: Resolving the Robustness-Precision Trade-off
- Language-Based Digital Twins for Elderly Cognitive Assistance
- Scalable and Interpretable Representation Alignment with Ordinal Similarity
- StepCache: Step-Level Reuse with Lightweight Verification and Selective Patching for LLM Serving
- Disentangling conviction and conformity: a Bayesian ideal point model of voting behaviour in online debates
- IdeaForge: A Knowledge Graph-Grounded Multi-Agent Framework for Cross-Methodology Innovation Analysis and Patent Claim Generation
- Artificial Intelligence in Science: Returns, Reallocation, and Reorganization
- The Rise of AI Search: Implications for Information Markets and Human Judgement at Scale
- GreenServ: Energy-Efficient Context-Aware Dynamic Routing for Multi-Model LLM Inference
- Interpreting Black-Box Large Language Models with Sentence-Level Energy Landscapes
- fog: Expressing Motion and Emotion through Function Composition of AI-Generated Code
- Attractor States Emerge in Multi-Turn LLM Conversations
- CSE-SFP: Enabling Unsupervised Sentence Representation Learning via a Single Forward Pass
- Semantic Invariance in Agentic AI
- Generative Engine Optimization: A VLM and Agent Framework for Pinterest Acquisition Growth
- Hybrid privacy-aware semantic search: SVD-truncated document geometry and CKKS-encrypted query reranking under a restricted threat model
- Measuring Curriculum Alignment across Topical Coverage, Competency, and Cognitive Depth: A Longitudinal Framework Applied to CS2013 and CS2023
- Can We Unmask the Underground? Detecting and Predicting Hidden Forum Interactions
- Fast LLM-Based Semantic Filtering: From a Unified Framework to an Adaptive Two-Phase Method
- RACT: Retrieval Augmented Column-Table Learning and Prediction for Multi-Table Schema Matching
- Unlearning Sensitive Information in Multimodal LLMs: Benchmark and Attack-Defense Evaluation
- Towards Explainable Temporal User Profiling with LLMs
- Re-Centering Humans in LLM Personalization
- Relational Semantic Reasoning on 3D Scene Graphs for Open World Interactive Object Search
- Financial Bond Similarity Search Using Representation Learning
- Structural Hallucination in Large Language Models: A Network-Based Evaluation of Knowledge Organization and Citation Integrity
- TraceLLM: Leveraging Large Language Models with Prompt Engineering for Enhanced Requirements Traceability
- LLM Ethics Benchmark: A Three-Dimensional Assessment System for Evaluating Moral Reasoning in Large Language Models
- Learning Retrieval Models with Sparse Autoencoders
- Synthetic Data Powers Product Retrieval for Long-tail Knowledge-Intensive Queries in E-commerce Search
- Self-Generated In-Context Examples Improve LLM Agents for Sequential Decision-Making Tasks
- Optimization of embeddings storage for RAG systems using quantization and dimensionality reduction techniques
- Power and Limitations of Aggregation in Compound AI Systems
- The TCF doesn't really A(A)ID -- Automatic Privacy Analysis and Legal Compliance of TCF-based Android Applications
- Sub-City Real Estate Price Index Forecasting at Weekly Horizons Using Satellite Radar and News Sentiment
- Improving Topic Modeling by Distilling Soft Labels from Language Models
- Lyapunov Spectral Analysis of Speech Embedding Trajectories in Psychosis
- Behavioral Integrity Verification for AI Agent Skills
- Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning
- Intent Laundering: AI Safety Datasets Are Not What They Seem
- Semantics-Aware Denoising: A PLM-Guided Sample Reweighting Strategy for Robust Recommendation
- Fast approximate Bayesian multidimensional scaling with consistency guarantees
- Generalistic or Specific Embeddings, Which is Better? An Empirical Study on Search for Clinical Coding in Non-English Languages
- The Model Is Not the Product: A Dual-Pillar Architecture for Local-First Psychological Coaching
- Enhancing Security and Strengthening Defenses in Automated Short-Answer Grading Systems
- Homa at SemEval-2025 Task 5: Aligning Librarian Records with OntoAligner for Subject Tagging
- TartuNLP at SemEval-2025 Task 5: Subject Tagging as Two-Stage Information Retrieval
- 20min-XD: A Comparable Corpus of Swiss News Articles
- Detecting Synthetic Political Narratives in Cross-Platform Social Media Discourse
- Birdie: Natural Language-Driven Table Discovery Using Differentiable Search Index
- Enhancing New-item Fairness in Dynamic Recommender Systems
- Overreliance in Writing Tasks: Exploring Similarity-Based Measures of AI Influence on Writing and Proposing a Reflective Writing Interface Intervention
- Embeddings for Preferences, Not Semantics
- An Automated Framework for Cybersecurity Policy Compliance Assessment Against Security Control Standards
- Classifier-to-Bias: Toward Unsupervised Automatic Bias Detection for Visual Classifiers
- Improving Phishing Email Detection Performance of Small Large Language Models
- Are Information Retrieval Approaches Good at Harmonising Longitudinal Survey Questions in Social Science?
- Kernel Affine Hull Machines as Compute-Efficient Encoders for Frozen Semantic Spaces
- Graph Synthetic Out-of-Distribution Exposure with Large Language Models
- GLIP-OOD: Zero-Shot Graph OOD Detection with Graph Foundation Model
- Skill Discovery for Software Scripting Automation via Offline Simulations with LLMs
- Accurate and Diverse LLM Mathematical Reasoning via Automated PRM-Guided GFlowNets
- CRC-Screen: Certified DNA-Synthesis Hazard Screening Under Taxonomic Shift
- Prompt Injection Attack to Tool Selection in LLM Agents
- Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets
- ELSA3D: Elastic Semantic Anchoring for Unified 3D Understanding and Generation
- An LLM-Powered Semantic Alignment Framework for Journal Recommendation
- Breaking the Filter Bubble: A Semantic Pareto-DQN Framework for Multi-Objective Recommendation
- Deployment-Time Memorization in Foundation-Model Agents
- Tool Attention Is All You Need: Dynamic Tool Gating and Lazy Schema Loading for Eliminating the MCP/Tools Tax in Scalable Agentic Workflows
- A Cloud-Native Architecture for Human-in-Control LLM-Assisted OpenSearch in Investigative Settings
- Text Steganography with Dynamic Codebook and Multimodal Large Language Model
- MATCHA: Matching Text via Contrastive Semantic Alignment
- Reasoning-Based Refinement of Unsupervised Text Clusters with LLMs
- Private Seeds, Public LLMs: Realistic and Privacy-Preserving Synthetic Data Generation
- PROTOCOL: Late Interaction Retrieval for Protein Homolog Search
- Learning is Forgetting: LLM Training As Lossy Compression
- How Vulnerable Are AI Agents to Indirect Prompt Injections? Insights from a Large-Scale Public Competition
- Structured Distillation for Personalized Agent Memory: 11x Token Reduction with Retrieval Preservation
- Chatbot Conversations in Physics Education: Using Artificial Intelligence to Analyze Student Reasoning through Computational Grounded Theory
- An Expert Schema for Evaluating Large Language Model Errors in Scholarly Question-Answering Systems
- Letting Tutor Personas Speak Up for LLMs: Learning Steering Vectors from Dialogue via Preference Optimization
- GraphAgents: Knowledge Graph-Guided Agentic AI for Cross-Domain Materials Design
- BibAgent: An Agentic Framework for Traceable Miscitation Detection in Scientific Literature
- TreeHop: Generate and Filter Next Query Embeddings Efficiently for Multi-hop Question Answering
- MICE for CATs: Model-Internal Confidence Estimation for Calibrating Agents with Tools
- Large Language Models are Qualified Benchmark Builders: Rebuilding Pre-Training Datasets for Advancing Code Intelligence Tasks
- Do Automatic Comment Generation Techniques Fall Short? Exploring the Influence of Method Dependencies on Code Understanding
- Uncertainty Quantification for Language Models: A Suite of Black-Box, White-Box, LLM Judge, and Ensemble Scorers
- Detecting Hallucinations in SpeechLLMs at Inference Time Using Attention Maps
- Assessing Capabilities of Large Language Models in Social Media Analytics: A Multi-task Quest
- Context-Aware Search and Retrieval Under Token Erasure
- Reasoning Models Know What's Important, and Encode It in Their Activations
- KIRA: Knowledge-Intensive Image Retrieval and Reasoning Architecture for Specialized Visual Domains
- FACTors: A New Dataset for Studying the Fact-checking Ecosystem
- Neuroevolution of Self-Attention Over Proto-Objects
- The Influence of Text Variation on User Engagement in Cross-Platform Content Sharing
- Can We Enhance Bug Report Quality Using LLMs?: An Empirical Study of LLM-Based Bug Report Generation
- ReSIM: Re-ranking Binary Similarity Embeddings to Improve Function Search Performance
- Bandit on the Hunt: Dynamic Crawling for Cyber Threat Intelligence
- Escaping the BLEU Trap: A Signal-Grounded Framework with Decoupled Semantic Guidance for EEG-to-Text Decoding
- NovBench: Evaluating Large Language Models on Academic Paper Novelty Assessment
- SemEnrich: Self-Supervised Semantic Enrichment of Radiology Reports for Vision-Language Learning
- Reverse-Engineering Model Editing on Language Models
- Semantic Search over 9 Million Mathematical Theorems
- EnsembleLink: Accurate Record Linkage Without Training Data
- Are Whitepaper Claims Reflected in Market Structure? A Contamination-Aware Pipeline and a Power-Limited Null
- RPO-RAG: Aligning Small LLMs with Relation-aware Preference Optimization for Knowledge Graph Question Answering
- LLM-based Semantic Search for Conversational Queries in E-commerce
- SemanticALLI: Caching Reasoning, Not Just Responses, in Agentic Systems
- Is Peer Review Really in Decline? Analyzing Review Quality across Venues and Time
- Human Values in a Single Sentence: Moral Presence, Hierarchies, and Transformer Ensembles on the Schwartz Continuum
- Guardrails for trust, safety, and ethical development and deployment of Large Language Models (LLM)
- LIME-LLM: Probing Models with Fluent Counterfactuals, Not Broken Text
- Actors, Frames and Arguments: A Multi-Decade Computational Analysis of Climate Discourse in Financial News using Large Language Models
- LLM Review: Enhancing Creative Writing via Blind Peer Review Feedback
- Lightweight Chunk Selection for Mobile Retrieval-Augmented Generation
- CLEAR: Causal Context-Based Agentic Reasoning for Vulnerability Detection
- SciRet: A Compute-Aware Empirical Study of Retrieval and Reranking for Scientific RAG
- AI-Assisted Peer Review Across Research Communities: From Reviewer AI Policies to LLM Review Quality
- CRS-Triage: Confidence- and Reliability-Aware Selective Triage under Incomplete Clinical Evidence
- LEAP: Layer-wise Exit-Aware Pretraining for Efficient Transformer Inference
- TAG-HGT: A Scalable and Cost-Effective Framework for Inductive Cold-Start Academic Recommendation
- HumanMCP: A Human-Like Query Dataset for Evaluating MCP Tool Retrieval Performance
- PicPersona-TOD : A Dataset for Personalizing Utterance Style in Task-Oriented Dialogue with Image Persona
- HierSum: A Global and Local Attention Mechanism for Video Summarization
- Explaining and Improving Model Behavior with k Nearest Neighbor Representations
- On the Diversity of Analogy Making in Large Language Models
- A Comprehensive Survey of Knowledge-Based Vision Question Answering Systems: The Lifecycle of Knowledge in Visual Reasoning Task
- DataScout: Automatic Data Fact Retrieval for Statement Augmentation with an LLM-Based Agent
- The Ultimate Cookbook for Invisible Poison: Crafting Subtle Clean-Label Text Backdoors with Style Attributes
- CoheMark: A Novel Sentence-Level Watermark for Enhanced Text Quality
- Cracking the Code of Action: a Generative Approach to Affordances for Reinforcement Learning
- FashionM3: Multimodal, Multitask, and Multiround Fashion Assistant based on Unified Vision-Language Model
- Can large language models recognize complex language errors such as zeugma?
- A RAG-Based Multi-Agent LLM System for Natural Hazard Resilience and Adaptation
- DataRx: Missingness-Aware Sampling for Safer Large Language Model Task-Specific Fine-Tuning
- ODRA: Synthesizing Cognitive Behavioral Therapy Sessions with Structured Chain-Of-Thought and Dynamic Patient Resistance
- EgoAfford: Task-Oriented Affordance Grounding via Egocentric Referring Segmentation
- CLIP-CC-Bench: Evaluating Paragraph-Level Video Descriptions in Video-Language Models
- Checked-In Secret Detection: Strings Are All You Need
- DAC-Pose: Dual-Agent Collaborative Framework for Pose-Guided Human Generation
- SurgNarrator: A Generative Retrieval Framework for Surgical Video Understanding
- Consistency-Driven Co-Evolution for Self-Supervised Cross-Representation Learning
- Consensus Measures for Unstructured Biomedical Text Annotations
- RAG-Stack: Co-Optimizing RAG Serving Performance and Quality
- SONAR: Task-Aware Code Summary Evaluation for LLM Consumers Without References
- Efficient Evaluation of Large Language Models via Collaborative Filtering
- Distillation and Refinement of Reasoning in Small Language Models for Document Re-ranking
- Recall Is Not Enough: A Reader-Context Diagnostic for Budget-Constrained Retrieval-Augmented Generation
- When Better Codebooks Are Not Enough: Predictive Performance and Behavioral Reliability in LLM Political Event Coding
- ECHO-PPI: Evidence-Bundled Overlapping Protein Module Detection with Hierarchical Assignment Confidence for Network Biology
- Tree of Thoughts as a Classical Heuristic Search Problem: Formal Foundations and Design Patterns
- Document Optimization for Black-Box Retrieval via Reinforcement Learning
- LUCid: Redefining Relevance For Lifelong Personalization
- How Effective Are NPM Malicious Package Detectors? A Large-Scale Empirical Study
- A Unified Retrieval Framework with Document Ranking and EDU Filtering for Multi-document Summarization
- Modality Reliability Guided Multimodal Recommendation
- Taxonomy-Aware Evaluation of Vision-Language Models
- Information Leakage of Sentence Embeddings via Generative Embedding Inversion Attacks
- Disentangling and Generating Modalities for Recommendation in Missing Modality Scenarios
- T-VEC: A Telecom-Specific Vectorization Model with Enhanced Semantic Understanding via Deep Triplet Loss Fine-Tuning
- Mining Software Repositories for Expert Recommendation
- Emo Pillars: Knowledge Distillation to Support Fine-Grained Context-Aware and Context-Less Emotion Classification
- Texture: Structured Exploration of Text Datasets
- Private Federated Learning using Preference-Optimized Synthetic Data
- Out-of-the-Box Conditional Text Embeddings from Large Language Models
- Capturing Symmetry and Antisymmetry in Language Models through Symmetry-Aware Training Objectives
- Automated Extraction and Analysis of Developer's Rationale in Open Source Software
- Describe Anything: Detailed Localized Image and Video Captioning
- Language Models to Support Multi-Label Classification of Industrial Data
- Advancing Egocentric Video Question Answering with Multimodal Large Language Models
- Balancing Complexity and Informativeness in LLM-Based Clustering: Finding the Goldilocks Zone
- DR.FIX: Automatically Fixing Data Races at Industry Scale
- llm-jp-modernbert: A ModernBERT Model Trained on a Large-Scale Japanese Corpus with Long Context Length
- Using digital traces to analyze software work: skills, careers and programming languages
- Vision-Language Models Are Not Pragmatically Competent in Referring Expression Generation
- Cequel: Cost-Effective Querying of Large Language Models for Text Clustering
- Leveraging Language Models for Automated Patient Record Linkage
- Fully Bayesian Approaches to Topics over Time
- LLMs as Data Annotators: How Close Are We to Human Performance
- On Self-improving Token Embeddings
- TVR: Automotive System Requirement Traceability Validation and Recovery Through Retrieval-Augmented Generation
- Tell Me What You Know About Sexism: Expert-LLM Interaction Strategies and Co-Created Definitions for Zero-Shot Sexism Detection
- Evaluating BERTopic on Open-Ended Data: A Case Study with Belgian Dutch Daily Narratives
- Could AI Trace and Explain the Origins of AI-Generated Images and Text?
- HLSTester: Efficient Testing of Behavioral Discrepancies with LLMs for High-Level Synthesis
- Learning from Reasoning Failures via Synthetic Data Generation
- UFO2: The Desktop AgentOS
- Don't Retrieve, Generate: Prompting LLMs for Synthetic Training Data in Dense Retrieval
- Probing the Subtle Ideological Manipulation of Large Language Models
- Learning over von Mises-Fisher Distributions via a Wasserstein-like Geometry
- Exploring Language Patterns of Prompts in Text-to-Image Generation and Their Impact on Visual Diversity
- Mind the Language Gap: Automated and Augmented Evaluation of Bias in LLMs for High- and Low-Resource Languages
- Enhancing Math Learning in an LMS Using AI-Driven Question Recommendations
- Scaling LLaNA: Advancing NeRF-Language Understanding Through Large-Scale Training
- MIG: Automatic Data Selection for Instruction Tuning by Maximizing Information Gain in Semantic Space
- Controlled Territory and Conflict Tracking (CONTACT): (Geo-)Mapping Occupied Territory from Open Source Intelligence
- Revisiting Uncertainty Quantification Evaluation in Language Models: Spurious Interactions with Response Length Bias Results
- Long-context Non-factoid Question Answering in Indic Languages
- RAG Without the Lag: Interactive Debugging for Retrieval-Augmented Generation Pipelines
- Accuracy is Not Agreement: Expert-Aligned Evaluation of Crash Narrative Classification Models
- CDF-RAG: Causal Dynamic Feedback for Adaptive Retrieval-Augmented Generation
- An Explicit Syllogistic Legal Reasoning Framework for Large Language Models
- ConExion: Concept Extraction with Large Language Models
- Data-efficient LLM Fine-tuning for Code Generation
- ARCeR: an Agentic RAG for the Automated Definition of Cyber Ranges
- Diffusion Generative Recommendation with Continuous Tokens
- Collaboration and Controversy Among Experts: Rumor Early Detection by Tuning a Comment Generator
- The Digital Cybersecurity Expert: How Far Have We Come?
- Evaluating the Diversity and Quality of LLM Generated Content
- Refining Over Resampling: Test-Time Self-Correction for LLM Reasoning
- Task-Conditional Flow Matching for Balanced Multilingual Text Embedding Adaptation
- Where Models Converge and Humans Diverge: A Coverage Framework for Distributional Pluralism in Open-Ended Generation
- OneEmo: A Unified Multimodal Reasoning Model for Emotion Perception, Understanding, and Interaction
- Convergent Evolution in Algorithmic Space
- Mapping Similarity Spaces across Embedding Models with Synthetic Query Probing
- Personalized Deep Research Query Refinement with Graph-Scaffolded Evidence Grounding
- CourseGraph: Finding overlaps and differences in Computer Science courses across universities
- CyberBridge: Bridging the Gap Between Cybersecurity Education and Industry
- A Mechanistic Analysis of Gender Sensitivity in Dense Retrieval Models
- DocSAM: Unified Document Image Segmentation via Query Decomposition and Heterogeneous Mixed Learning
- Abstract Event Causal Rules: Induction and Application
- PEIRCE: Unifying Material and Formal Reasoning via LLM-Driven Neuro-Symbolic Refinement
- Simulator-Grounded Large Language Models for Industrial Causal Reasoning: Tool-Use, Structured Injection, and Plant-Portable Retrieval for Wastewater Treatment Decision Support
- MARS: Margin and Semantic-Aware Data Augmentation for Reward Modeling
- GraphicBench: A Planning Benchmark for Graphic Design with Language Agents
- DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
- Video Summarization with Large Language Models
- Ai2 Scholar QA: Organized Literature Synthesis with Attribution
- Can Vision-Language Models Understand and Interpret Dynamic Gestures from Pedestrians? Pilot Datasets and Exploration Towards Instructive Nonverbal Commands for Cooperative Autonomous Vehicles
- From Misleading Queries to Accurate Answers: A Three-Stage Fine-Tuning Method for LLMs
- MIEB: Massive Image Embedding Benchmark
- Can LLMs Generate Tabular Summaries of Science Papers? Rethinking the Evaluation Protocol
- Fact-Checking with Contextual Narratives: Leveraging Retrieval-Augmented LLMs for Social Media Analysis
- Auto-Test: Learning Semantic-Domain Constraints for Unsupervised Error Detection in Tables
- A Survey of Personalization: From RAG to Agent
- The Structural Safety Generalization Problem
- Metropolis-Hastings Captioning Game: Knowledge Fusion of Vision Language Models via Decentralized Bayesian Inference
- TextSplat: Text-Guided Semantic Fusion for Generalizable Gaussian Splatting
- Visual moral inference and communication
- Quantifying the Spread of Online Incivility in Brazilian Politics
- VectorLiteRAG: Latency-Aware and Fine-Grained Resource Partitioning for Efficient RAG
- LLMTaxo: Leveraging Large Language Models for Constructing Taxonomy of Factual Claims from Social Media
- Quantum Large Language Model Fine-Tuning
- Embodied Image Captioning: Self-supervised Learning Agents for Spatially Coherent Image Descriptions
- The Ultimate Configuration Management Tool? Lessons from a Mixed Methods Study of Ansible's Challenges
- PCA-RAG: Principal Component Analysis for Efficient Retrieval-Augmented Generation
- A Strategic Coordination Framework of Small LLMs Matches Large LLMs in Data Synthesis
- Out of Style: RAG's Fragility to Linguistic Variation
- VLMT: Vision-Language Multimodal Transformer for Multimodal Multi-hop Question Answering
- Towards Sustainable Creativity Support: An Exploratory Study on Prompt Based Image Generation
- MOSAIC: Modeling Social AI for Content Dissemination and Regulation in Multi-Agent Simulations
- The Platform Is Mostly Not a Platform: Token Economies and Agent Discourse on Moltbook
- Learning Long Short-Term Intention within Human Daily Behaviors
- Enhanced Question-Answering for Skill-based learning using Knowledge-based AI and Generative AI
- Leveraging LLMs for Multimodal Retrieval-Augmented Radiology Report Generation via Key Phrase Extraction
- AttentionDefense: Leveraging System Prompt Attention for Explainable Defense Against Novel Jailbreaks
- Impact of Language Guidance: A Reproducibility Study
- Missing Premise exacerbates Overthinking: Are Reasoning Models losing Critical Thinking Skill?
- Poly-Vector Retrieval: Reference and Content Embeddings for Legal Documents
- Identifying Aspects in Peer Reviews
- SemEval-2025 Task 5: LLMs4Subjects -- LLM-based Automated Subject Tagging for a National Technical Library's Open-Access Catalog
- SkillFlow: Efficient Skill and Code Transfer Through Communication in Adapting AI Agents
- Decentralizing AI Memory: SHIMI, a Semantic Hierarchical Memory Index for Scalable Agent Reasoning
- Graph-based Approaches and Functionalities in Retrieval-Augmented Generation: A Comprehensive Survey
- ShadowCoT: Cognitive Hijacking for Stealthy Reasoning Backdoors in LLMs
- Comparing Self-Disclosure Themes and Semantics to a Human, a Robot, and a Disembodied Agent
- REVEAL: Relation-based Video Representation Learning for Video-Question-Answering
- Prεεmpt: Sanitizing Sensitive Prompts for LLMs
- Few Dimensions are Enough: Fine-tuning BERT with Selected Dimensions Revealed Its Redundant Nature
- ChartQAPro: A More Diverse and Challenging Benchmark for Chart Question Answering
- Beyond the Next Token: Towards Prompt-Robust Zero-Shot Classification via Efficient Multi-Token Prediction
- SARLANG-1M: A Benchmark for Vision-Language Modeling in SAR Image Understanding
- Recursive Training Loops in LLMs: How training data properties modulate distribution shift in generated data?
Related