Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
2019/08/27 by Nils Reimers, Iryna Gurevych, Reimers, Nils +1 · 1182 citations
Computer Science · #Topic Modeling #Natural Language Processing Techniques #Sentiment Analysis and Opinion Mining
paper · doi:10.48550/arxiv.1908.10084
Abstract
BERT (Devlin et al., 2018) and RoBERTa (Liu et al., 2019) has set a new state-of-the-art performance on sentence-pair regression tasks like semantic textual similarity (STS). However, it requires that both sentences are fed into the network, which causes a massive computational overhead: Finding the most similar pair in a collection of 10,000 sentences requires about 50 million inference computations (~65 hours) with BERT. The construction of BERT makes it unsuitable for semantic similarity search as well as for unsupervised tasks like clustering. In this publication, we present Sentence-BERT (SBERT), a modification of the pretrained BERT network that use siamese and triplet network structures to derive semantically meaningful sentence embeddings that can be compared using cosine-similarity. This reduces the effort for finding the most similar pair from 65 hours with BERT / RoBERTa to about 5 seconds with SBERT, while maintaining the accuracy from BERT. We evaluate SBERT and SRoBERTa on common STS tasks and transfer learning tasks, where it outperforms other state-of-the-art sentence embeddings methods.
Cited by
- Evaluating the Effect of Linguistic Relatedness on Cross-Lingual Transfer in Large Multilingual Automatic Speech Recognition
- From transcription to semantic corpus analysis: unsupervised learning of sentence representations for ancient languages
- AGRO-SQL: Agentic Group-Relative Optimization with High-Fidelity Data Synthesis
- Multi Agents Semantic Emotion Aligned Music to Image Generation with Music Derived Captions
- Debugging Tabular Log as Dynamic Graphs
- MUSON: A Reasoning-oriented Multimodal Dataset for Socially Compliant Navigation in Urban Environments
- Self-Rewarded Multimodal Coherent Reasoning Across Diverse Visual Domains
- Keep It InMind: Benchmarking the Implicit-Association Blind Spot in Agent Memory
- HalluMat: Detecting Hallucinations in LLM-Generated Materials Science Content Through Multi-Stage Verification
- Valori: A Deterministic Memory Substrate for AI Systems
- Hierarchical Geometry of Cognitive States in Transformer Embedding Spaces
- iOS as Acceleration
- Recursive Governance: A Graph-Theoretic Framework for Risk Propagation and Drift Detection in Agentic AI Systems
- From Unstructured Recall to Schema-Grounded Memory: Reliable AI Memory via Iterative, Schema-Aware Extraction
- Advocating the potential of artificial intelligence for syndrome discovery in syndromic surveillance systems: A scoping review
- PlanCraft: Sketch, Refine, and Furnish for Architect-Inspired Progressive 3D Residential Scene Generation
- Choosing a Text Embedding Model: A Practical Benchmarking and Decision Framework
- LinkRank: A Learning-to-Rank Framework for One-to-Many Issue-Commit Traceability
- seqLens: Optimizing Language Models for Genomic Predictions
- Do LLM Debates Repeat Arguments Differently Across Languages?
- LiFT-MPC: Language-in-the-Loop Feedback Tuning of Cost Previews for MPC
- TriShieldRAG: A Three-Ring Defense-in-Depth Framework Against Knowledge Corruption in Retrieval-Augmented Generation
- The Dynamics of Human and AI-Generated Language: How Semantic Content Fluctuates Across Different Timescales
- Where Quality Breaks in Compressed Short-Text Generation: Staged Bottleneck Localization
- StanceFlip: A Comprehensive Multi-Dimensional Benchmark for Multimodal Conversational Stance Flipping Forecasting
- Beyond Direct Answering: Aligning Educational LLMs as Socratic Guides via Heuristic Reinforcement Learning
- From Signals to Behaviors: Evidence-Based Android Malware Detection
- ADAGE: A Language-Agnostic Pipeline for Analogical Reasoning Evaluation
- Interpretable Column Annotation with LLM-Symbolized Decision Process Materialization
- Bekko Embedding: Parameter-Efficient Multilingual Retrieval with Ultra-Compact Encoders
- The Case Against Generation for Retrieval: Discriminative Language Models as Effective Retrievers
- From Semantics to Readout: Mechanistic Understanding of Audio Tokens after Fine-Tuning for Temporal Audio Grounding
- HVM-GraphRAG: Holistic-View Multimodal Graph Retrieval-Augmented Generation on Complex Document
- Traceable LLM Reasoning for Fake-Order Fraud Detection
- Frustratingly Simple Black-Box Adaptation of Language Models via Logit Bias
- Disentangling the Interpretive and Predictive Roles of LIWC: Controlled Substitution in Depression-Related Classification
- An Interactive Vision Language Platform for Cognitive Remediation in Schizophrenia
- MPR-CiteG: Enhancing RAG with Multi-Portfolio Retrieval and Citation-Grounded Generation
- SciClaimSeekers at CheckThat! 2026: Retrieving Scientific Sources for Social Media Claims with LLM Reranking
- SourceMinds at CheckThat! 2026: NLI-Grounded Citation Auditing in a Multi-Agent Pipeline for Full Fact-Checking Article Generation
- Multimodal Hybrid Retrieval-Augmented Generation for Scientific Document Understanding using Open-Source SLMs
- AI-Assisted Knowledge Access for Legacy Enterprise Asset Management in Energy Operations: A Practical Retrieval System
- Agent-based simulation of online social networks and disinformation
- Source-Aware Reranking for Retrieval-Augmented Generation: A Reliability Prior Approach
- Measuring and Improving Behavioral Consistency in Large Language Models through Fact-Heuristic-Emotion State Enforcement
- GrocLM: Grocery Category Recommendation in E-Commerce with Large Language Models
- HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
- JKO-RAG: Distributional Retrieval as Wasserstein Free-Energy Gradient Flow
- SeT-Diff: Towards Semantic Foundation Models for HPC Telemetry and Time-Series
- JobMatchAI-An Intelligent Job Matching Platform Using Knowledge Graphs, Semantic Search and Explainable AI
- To update or to separate: Neural signatures and consequences of latent cause inference in episodic memory
- Towards Isolated Interventions via Almost Orthogonal Features in Language Models
- Intuitive therapist robot patient physical interaction is worth a thousand words
- Explainable Statute Prediction via Attention-based Model and LLM Prompting
- MMCTOP: A Multimodal Textualization and Mixture-of-Experts Framework for Clinical Trial Outcome Prediction
- Knowledge Reasoning of Large Language Models Integrating Graph-Structured Information for Pest and Disease Control in Tobacco
- HELP: Hierarchical Embodied Language Planner for Household Tasks
- Enabling Conversational Behavior Reasoning Capabilities in Full-Duplex Speech
- ALETHEIA: Combating Social Media Influence Campaigns with Graph Neural Networks
- SENTINEL: A Multi-Modal Early Detection Framework for Emerging Cyber Threats using Telegram
- T2AV-Compass: Towards Unified Evaluation for Text-to-Audio-Video Generation
- Chorus: Multi-Teacher Pretraining for Holistic 3D Gaussian Scene Encoding
- AegisAgent: An Autonomous Defense Agent Against Prompt Injection Attacks in LLM-HARs
- How important is Recall for Measuring Retrieval Quality?
- Leveraging Lightweight Entity Extraction for Scalable Event-Based Image Retrieval
- X-GridAgent: An LLM-Powered Agentic AI System for Assisting Power Grid Analysis
- Structured Visualization Design Knowledge for Grounding Generative Reasoning and Situated Feedback
- Beyond Vision: Contextually Enriched Image Captioning with Multi-Modal Retrieval
- Interpolative Decoding: Exploring the Spectrum of Personality Traits in LLMs
- MemR3: Memory Retrieval via Reflective Reasoning for LLM Agents
- EchoTrail-GUI: Building Actionable Memory for GUI Agents via Critic-Guided Self-Exploration
- From Pixels to Predicates Structuring urban perception with scene graphs
- Understanding Chain-of-Thought in Large Language Models via Topological Data Analysis
- FC-MIR: A Mobile Screen Awareness Framework for Intent-Aware Recommendation based on Frame-Compressed Multimodal Trajectory Reasoning
- Learning Hierarchical Procedural Memory for LLM Agents through Bayesian Selection and Contrastive Refinement
- Cross-modal Counterfactual Explanations: Uncovering Decision Factors and Dataset Biases in Subjective Classification
- Explainable and Fine-Grained Safeguarding of LLM Multi-Agent Systems via Bi-Level Graph Anomaly Detection
- Probabilistic Digital Twins of Users: Latent Representation Learning with Statistically Validated Semantics
- External Hippocampus: Topological Cognitive Maps for Guiding Large Language Model Reasoning
- AnyTask: an Automated Task and Data Generation Framework for Advancing Sim-to-Real Policy Learning
- Incorporating Error Level Noise Embedding for Improving LLM-Assisted Robustness in Persian Speech Recognition
- Atom: Efficient On-Device Video-Language Pipelines Through Modular Reuse
- Reinforcement Learning for Self-Improving Agent with Skill Library
- Dynamic Tool Dependency Retrieval for Efficient Function Calling
- NRGPT: An Energy-based Alternative for GPT
- Persistent Multiscale Density-based Clustering
- From Essence to Defense: Adaptive Semantic-aware Watermarking for Embedding-as-a-Service Copyright Protection
- BrepLLM: Native Boundary Representation Understanding with Large Language Models
- Design and Evaluation of Cost-Aware PoQ for Decentralized LLM Inference
- Coarse-to-Fine Open-Set Graph Node Classification with Large Language Models
- ModelTables: A Corpus of Tables about Models
- Convolutional Lie Operator for Sentence Classification
- ORACLE: Time-Dependent Recursive Summary Graphs for Foresight on News Data Using LLMs
- Social Story Frames: Contextual Reasoning about Narrative Intent and Reception
- ArcBERT: An LLM-based Search Engine for Exploring Integrated Multi-Omics Metadata
- SynGP500: A Clinically-Grounded Synthetic Dataset of Australian General Practice Medical Notes
- FAME: Fictional Actors for Multilingual Erasure
- The Semantic Illusion: Certified Limits of Embedding-Based Hallucination Detection in RAG Systems
- Examining the Utility of Self-disclosure Types for Modeling Annotators of Social Norms
- EVICPRESS: Joint KV-Cache Compression and Eviction for Efficient LLM Serving
- DP-Bench: A Benchmark for Evaluating Data Product Creation Systems
- IaC Generation with LLMs: An Error Taxonomy and A Study on Configuration Knowledge Injection
- SPARQL-LLM: Real-Time SPARQL Query Generation from Natural Language Questions
- Explainable Ethical Assessment on Human Behaviors by Generating Conflicting Social Norms
- Improving Semantic Uncertainty Quantification in LVLMs with Semantic Gaussian Processes
- FiNERweb: Datasets and Artifacts for Scalable Multilingual Named Entity Recognition
- Beyond surface form: A pipeline for semantic analysis in Alzheimer's Disease detection from spontaneous speech
- Semantic Grounding Index: Geometric Bounds on Context Engagement in RAG Systems
- Integrating Causal Reasoning into Automated Fact-Checking
- UCRBench: Benchmarking LLMs on Use Case Recovery
- Learning to Retrieve with Weakened Labels: Robust Training under Label Noise
- A Relational Model of Neighborhood Mobility: The Role of Amenities and Cultural Alignment
- Lemon: A Unified and Scalable 3D Multimodal Model for Universal Spatial Understanding
- CoRe3D: Collaborative Reasoning as a Foundation for 3D Intelligence
- Intelligent Scientific Literature Explorer using Machine Learning (ISLE)
- Tacit Understanding Game (TUG): Predicting Interpersonal Compatibility
- HyperEdit: Unlocking Instruction-based Text Editing in LLMs via Hypernetworks
- Semantic Distance Measurement based on Multi-Kernel Gaussian Processes
- Rethinking Label Consistency of In-Context Learning: An Implicit Transductive Label Propagation Perspective
- Rethinking Jailbreak Detection of Large Vision Language Models with Representational Contrastive Scoring
- PhraseVAE and PhraseLDM: Latent Diffusion for Full-Song Multitrack Symbolic Music Generation
- Extending a Parliamentary Corpus with MPs' Tweets: Automatic Annotation and Evaluation Using MultiParTweet
- Automating Historical Insight Extraction from Large-Scale Newspaper Archives via Neural Topic Modeling
- MultiScript30k: Leveraging Multilingual Embeddings to Extend Cross Script Parallel Data
- Interpretable and Steerable Concept Bottleneck Sparse Autoencoders
- GLOW: Graph-Language Co-Reasoning for Agentic Workflow Performance Prediction
- Semantic Reconstruction of Adversarial Plagiarism: A Context-Aware Framework for Detecting and Restoring "Tortured Phrases" in Scientific Literature
- AgriRegion: Region-Aware Retrieval for High-Fidelity Agricultural Advice
- Interpretation as Linear Transformation: A Cognitive-Geometric Model of Belief and Meaning
- Ontology-Based Knowledge Graph Framework for Industrial Standard Documents via Hierarchical and Propositional Structuring
- Semantic-Aware Cooperative Communication and Computation Framework in Vehicular Networks
- Refining Diffusion Models for Motion Synthesis with an Acceleration Loss to Generate Realistic IMU Data
- Beyond Traditional Diagnostics: Transforming Patient-Side Information into Predictive Insights with Knowledge Graphs and Prototypes
- How a Bit Becomes a Story: Semantic Steering via Differentiable Fault Injection
- Luxical: High-Speed Lexical-Dense Text Embeddings
- Bridging Code Graphs and Large Language Models for Better Code Understanding
- SkipKV: Selective Skipping of KV Generation and Storage for Efficient Inference with Large Reasoning Models
- Relational Visual Similarity
- When Large Language Models Do Not Work: Online Incivility Prediction through Graph Neural Networks
- MoCoRP: Modeling Consistent Relations between Persona and Response for Persona-based Dialogue
- MASim: Multilingual Agent-Based Simulation for Social Science
- PICKT: Practical Interlinked Concept Knowledge Tracing for Personalized Learning using Knowledge Map Concept Relations
- Cross-platform Product Matching Based on Entity Alignment of Knowledge Graph with RAEA model
- See Once, Then Act: Vision-Language-Action Model with Task Learning from One-Shot Video Demonstrations
- STAR-GO: Improving Protein Function Prediction by Learning to Hierarchically Integrate Ontology-Informed Semantic Embeddings
- LLM-Upgraded Graph Reinforcement Learning for Carbon-Aware Job Scheduling in Smart Manufacturing
- One Word Is Not Enough: Simple Prompts Improve Word Embeddings
- Distribution-Aware Exploration for Adaptive HNSW Search
- Prompting-in-a-Series: Psychology-Informed Contents and Embeddings for Personality Recognition With Decoder-Only Models
- Graph-Regularized Sparse Autoencoders for LLM Safety Steering
- TopiCLEAR: Topic extraction by CLustering Embeddings with Adaptive dimensional Reduction
- GradientSpace: Unsupervised Data Clustering for Improved Instruction Tuning
- MedGRPO: Multi-Task Reinforcement Learning for Heterogeneous Medical Video Understanding
- Knowing What's Missing: Assessing Information Sufficiency in Question Answering
- CommentScope: A Comment-Embedded Assisted Reading System for a Long Text
- Heard or Halted? Gender, Interruptions, and Emotional Tone in U.S. Supreme Court Oral Arguments
- The Road of Adaptive AI for Precision in Cybersecurity
- ResearchArcade: Graph Interface for Academic Tasks
- AfriStereo: A Culturally Grounded Dataset for Evaluating Stereotypical Bias in Large Language Models
- Mapping CVEs to MITRE ATT&CK Techniques: A Curated Gold-Set Classifier and the Limits of LLM-Assisted Label Expansion
- Verbalizing LLMs' assumptions to explain and control sycophancy
- OsmT: Bridging OpenStreetMap Queries and Natural Language with Open-source Tag-aware Language Models
- AdmTree: Compressing Lengthy Context with Adaptive Semantic Trees
- Distance Is All You Need: Radial Dispersion for Uncertainty Estimation in Large Language Models
- ClusterFusion: Hybrid Clustering with Embedding Guidance and LLM Adaptation
- ConsentDiff at Scale: Longitudinal Audits of Web Privacy Policy Changes and UI Frictions
- Peek-a-Boo Reasoning: Contrastive Region Masking in MLLMs
- Overcoming State Inertia: Minimally Invasive Temporal Alignment for Evolving Contexts
- ProcObject-10K: Benchmarking Object-Centric Procedural Understanding in Instructional Videos
- Semantic Nutrition Estimation: Predicting Food Healthfulness from Text Descriptions
- Enhancing Job Matching: Occupation, Skill and Qualification Linking with the ESCO and EQF taxonomies
- GraphMatch: Fusing Language and Graph Representations in a Dynamic Two-Sided Work Marketplace
- SR-GRPO: Stable Rank as an Intrinsic Geometric Reward for Large Language Model Alignment
- Flowchart2Mermaid: A Vision-Language Model Powered System for Converting Flowcharts into Editable Diagram Code
- Ensemble Privacy Defense for Knowledge-Intensive LLMs against Membership Inference Attacks
- GPTrace: Effective Crash Deduplication Using LLM Embeddings
- LEC: Linear Expectation Constraints for Selection-Conditioned Risk Control in Selective Prediction and Routing Systems
- Label Forensics: Interpreting Hard Labels in Black-Box Text Classifier
- Code Comments for Quantum Software Development Kits: An Empirical Study on Qiskit
- Language-Guided Open-World Anomaly Segmentation
- MARSAD: A Multi-Functional Tool for Real-Time Social Media Analysis
- Patient Safety Risks from AI Scribes: Signals from End-User Feedback
- Advancing Academic Chatbots: Evaluation of Non Traditional Outputs
- Graph Queries from Natural Language using Constrained Language Models and Visual Editing
- WaterSearch: A Quality-Aware Search-based Watermarking Framework for Large Language Models
- Bias Injection Attacks on RAG Databases and Sanitization Defenses
- SAGE: Semantic-Aware Gray-Box Game Regression Testing with Large Language Models
- CourseTimeQA: A Lecture-Video Benchmark and a Latency-Constrained Cross-Modal Fusion Method for Timestamped QA
- Toward Automated and Trustworthy Scientific Analysis and Visualization with LLM-Generated Code
- Tracing How Annotators Think: Augmenting Preference Judgments with Reading Processes
- Learning to Prioritize IT Tickets: A Comparative Evaluation of Embedding-based Approaches and Fine-Tuned Transformer Models
- Language-guided 3D scene synthesis for fine-grained functionality understanding
- Listwise Preference Optimization with Element-wise Confusions for Aspect Sentiment Quad Prediction
- Learning to Refuse: Refusal-Aware Reinforcement Fine-Tuning for Hard-Irrelevant Queries in Video Temporal Grounding
- A transfer learning approach for automatic conflicts detection in software requirement sentence pairs based on dual encoders
- Pooling Attention: Evaluating Pretrained Transformer Embeddings for Deception Classification
- Experts are all you need: A Composable Framework for Large Language Model Inference
- A Customer Journey in the Land of Oz: Leveraging the Wizard of Oz Technique to Model Emotions in Customer Service Interactions
- A Taxonomy-Driven Case Study of Australian Web Resources Against Technology-Facilitated Abuse
- Bandit Guided Submodular Curriculum for Adaptive Subset Selection
- Tackling a Challenging Corpus for Early Detection of Gambling Disorder: UNSL at MentalRiskES 2025
- Breaking the Illusion: Consensus-Based Generative Mitigation of Adversarial Illusions in Multi-Modal Embeddings
- Learning Programming in Informal Spaces: Using Emotion as a Lens to Understand Novice Struggles on r/learnprogramming
- Economies of Open Intelligence: Tracing Power & Participation in the Model Ecosystem
- Beyond Membership: Limitations of Add/Remove Adjacency in Differential Privacy
- From Compound Figures to Composite Understanding: Developing a Multi-Modal LLM from Biomedical Literature with Medical Multiple-Image Benchmarking and Validation
- Real-PGDN: A Two-level Classification Method for Full-Process Recognition of Newly Registered Pornographic and Gambling Domain Names
- PROMPTMINER: Black-Box Prompt Stealing against Text-to-Image Generative Models via Reinforcement Learning and Fuzz Optimization
- BAMAS: Structuring Budget-Aware Multi-Agent Systems
- Unsupervised Multimodal Graph-based Model for Geo-social Analysis
- MetaRank: Task-Aware Metric Selection for Model Transferability Estimation
- Subgoal Graph-Augmented Planning for LLM-Guided Open-World Reinforcement Learning
- Generating Querying Code from Text for Multi-Modal Electronic Health Record
- Unsupervised Memorability Modeling from Tip-of-the-Tongue Retrieval Queries
- SPHINX: A Synthetic Environment for Visual Perception and Reasoning
- Memories Retrieved from Many Paths: A Multi-Prefix Framework for Robust Detection of Training Data Leakage in Large Language Models
- Diverse Video Generation with Determinantal Point Process-Guided Policy Optimization
- Copyright Detection in Large Language Models: An Ethical Approach to Generative AI Development
- DesignPref: Capturing Personal Preferences in Visual Design Generation
- NormEval: A Unified Multi-Metric Framework for Evaluating Semantic Fidelity in Text Normalization
- A Systematic Analysis of Large Language Models with RAG-enabled Dynamic Prompting for Medical Error Detection and Correction
- Beyond Relational: Semantic-Aware Multi-Modal Analytics with LLM-Native Query Optimization
- Language-Independent Sentiment Labelling with Distant Supervision: A Case Study for English, Sepedi and Setswana
- Online-PVLM: Advancing Personalized VLMs with Online Concept Learning
- Interactive AI NPCs Powered by LLMs: Technical Report for the CPDC Challenge 2025
- Are Neuro-Inspired Multi-Modal Vision-Language Models Resilient to Membership Inference Privacy Leakage?
- Accuracy and Efficiency Trade-Offs in LLM-Based Malware Detection and Explanation: A Comparative Study of Parameter Tuning vs. Full Fine-Tuning
- Beyond Protein Language Models: An Agentic LLM Framework for Mechanistic Enzyme Design
- UISearch: Graph-Based Embeddings for Multimodal Enterprise UI Screenshots Retrieval
- FilmSceneDesigner: Chaining Set Design for Procedural Film Scene Generation
- ReEXplore: Improving MLLMs for Embodied Exploration with Contextualized Retrospective Experience Replay
- Benchmarking Corruption Robustness of LVLMs: A Discriminative Benchmark and Robustness Alignment Metric
- A Reproducible Framework for Neural Topic Modeling in Focus Group Analysis
- Yo'City: Personalized and Boundless 3D Realistic City Scene Generation via Self-Critic Expansion
- Comparative Analysis of LoRA-Adapted Embedding Models for Clinical Cardiology Text Representation
- Scalable Parameter-Light Spectral Method for Clustering Short Text Embeddings with a Cohesion-Based Evaluation Metric
- What Helps Language Models Predict Human Beliefs: Demographics or Prior Stances?
- From Reviewers' Lens: Understanding Bug Bounty Report Invalid Reasons with LLMs
- LLM Assisted Coding with Metamorphic Specification Mutation Agent
- Consolidating Diffusion-Generated Video Detection with Unified Multimodal Forgery Learning
- Paper2SysArch: Structure-Constrained System Architecture Generation from Scientific Papers
- GeeSanBhava: Sentiment Tagged Sinhala Music Video Comment Data Set
- Blu-WERP (Web Extraction and Refinement Pipeline): A Scalable Pipeline for Preprocessing Large Language Model Datasets
- EduMod-LLM: A Modular Approach for Designing Flexible and Transparent Educational Assistants
- M3-Bench: Multi-Modal, Multi-Hop, Multi-Threaded Tool-Using MLLM Agent Benchmark
- REMSA: An LLM Agent for Foundation Model Selection in Remote Sensing
- SMILE: A Composite Lexical-Semantic Metric for Question-Answering Evaluation
- Don't Learn, Ground: A Case for Natural Language Inference with Visual Grounding
- CREST: Improving Interpretability and Effectiveness of Troubleshooting at Ericsson through Criterion-Specific Trouble Report Retrieval
- A Benchmark for Procedural Memory Retrieval in Language Agents
- Monte Carlo Expected Threat (MOCET) Scoring
- Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation
- Learning from Sufficient Rationales: Analysing the Relationship Between Explanation Faithfulness and Token-level Regularisation Strategies
- WER is Unaware: Assessing How ASR Errors Distort Clinical Understanding in Patient Facing Dialogue
- When Alignment Fails: Multimodal Adversarial Attacks on Vision-Language-Action Models
- Detecting Sleeper Agents in Large Language Models via Semantic Drift Analysis
- TurkColBERT: A Benchmark of Dense and Late-Interaction Models for Turkish Information Retrieval
- Two-Faced Social Agents: Context Collapse in Role-Conditioned Large Language Models
- Quality-Controlled Multimodal Emotion Recognition in Conversations with Identity-Based Transfer Learning and MAMBA Fusion
- Hierarchical Token Prepending: Enhancing Information Flow in Decoder-based LLM Embeddings
- Leveraging Digitized Newspapers to Collect Summarization Data in Low-Resource Languages
- Mitigating Label Length Bias in Large Language Models
- Technique to Baseline QE Artefact Generation Aligned to Quality Metrics
- RoboTidy : A 3D Gaussian Splatting Household Tidying Benchmark for Embodied Navigation and Action
- SweeperBot: Making 3D Browsing Accessible through View Analysis and Visual Question Answering
- Beyond Surface-Level Similarity: Hierarchical Contamination Detection for Synthetic Training Data in Foundation Models
- TaoSearchEmb: A Multi-Objective Reinforcement Learning Framework for Dense Retrieval in Taobao Search
- Hierarchical Retrieval with Out-Of-Vocabulary Queries: A Case Study on SNOMED CT
- PolicyBot - Reliable Question Answering over Policy Documents
- Attention Grounded Enhancement for Visual Document Retrieval
- Mem-PAL: Towards Memory-based Personalized Dialogue Assistants for Long-term User-Agent Interaction
- Unifying points of interest taxonomies: mapping OpenStreetMap tags to the Foursquare category system
- Dual-LoRA and Quality-Enhanced Pseudo Replay for Multimodal Continual Food Learning
- EL3DD: Extended Latent 3D Diffusion for Language Conditioned Multitask Manipulation
- Uncovering and Mitigating Transient Blindness in Multimodal Model Editing
- Assessing Large Language Models in Generating RTL Design Specifications
- Synthetic Eggs in Many Baskets: The Impact of Synthetic Data Diversity on LLM Fine-Tuning
- Analyzing Sustainability Messaging in Large-Scale Corporate Social Media
- Multimodal Large Language Models as Image Classifiers
- Difficulty-Controllable Cloze Question Distractor Generation
- MedSumGraph: enhancing GraphRAG for medical QA with summarization and optimized prompts
- RLHF May Not Reflect Genuine Preferences
- CAT-ID2: Category-Tree Integrated Document Identifier Learning for Generative Retrieval In E-commerce
- Entailment as Few-Shot Learner
- ConneX: Automatically Resolving Transaction Opacity of Cross-Chain Bridges for Security Analysis
- RAGSmith: A Framework for Finding the Optimal Composition of Retrieval-Augmented Generation Methods Across Datasets
- Exploring question answering: metric analysis and evaluation framework for enhanced interpretability
- Evaluating Embedding Generalization: How LLMs, LoRA, and SLERP Shape Representational Geometry
- HEDGE: Hallucination Estimation via Dense Geometric Entropy for VQA with Vision-Language Models
- Hi-Reco: High-Fidelity Real-Time Conversational Digital Humans
- LLM-Powered Text-Attributed Graph Anomaly Detection via Retrieval-Augmented Reasoning
- Evolving Prompts for Toxicity Search in Large Language Models
- Prompt Engineering Techniques for Context-dependent Text-to-SQL in Arabic
- A Systematic Study of Model Extraction Attacks on Graph Foundation Models
- Correcting Mean Bias in Text Embeddings: A Refined Renormalization with Training-Free Improvements on MMTEB
- Draft and Refine with Visual Experts
- Generative Caching for Structurally Similar Prompts and Responses
- Cost Transparency of Enterprise AI Adoption
- Beyond Elicitation: Provision-based Prompt Optimization for Knowledge-Intensive Tasks
- Reasoning About Intent for Ambiguous Requests
- Local Hybrid Retrieval-Augmented Document QA
- OutSafe-Bench: A Benchmark for Multimodal Offensive Content Detection in Large Language Models
- ProgRAG: Hallucination-Resistant Progressive Retrieval and Reasoning over Knowledge Graphs
- A general framework for adaptive nonparametric dimensionality reduction
- PustakAI: Curriculum-Aligned and Interactive Textbooks Using Large Language Models
- TermGPT: Multi-Level Contrastive Fine-Tuning for Terminology Adaptation in Legal and Financial Domain
- Generative AI as a Linguistic Equalizer in Global Science
- Contextual Graph Embeddings: Accounting for Data Characteristics in Heterogeneous Data Integration
- Improve Contrastive Clustering Performance by Multiple Fusing-Augmenting ViT Blocks
- Hierarchical Memorization in Large Language Models: Evidence from Citation Generation
- A centroid based framework for text classification in itsm environments
- Synergistic Feature Fusion for Latent Lyrical Classification: A Gated Deep Learning Architecture
- TurkEmbed: Turkish Embedding Model on NLI & STS Tasks
- DPRM: A Dual Implicit Process Reward Model in Multi-Hop Question Answering
- Prompt Tuning for Natural Language to SQL with Embedding Fine-Tuning and RAG
- Evaluating Gemini LLM in Food Image-Based Recipe and Nutrition Description with EfficientNet-B4 Visual Backbone
- Self-Correction Distillation for Structured Data Question Answering
- VSPO: Validating Semantic Pitfalls in Ontology via LLM-Based CQ Generation
- Unified Work Embeddings: Contrastive Learning of a Bidirectional Multi-task Ranker
- Last Layer Logits to Logic: Empowering LLMs with Logic-Consistent Structured Knowledge Reasoning
- Sparse3DPR: Training-Free 3D Hierarchical Scene Parsing and Task-Adaptive Subgraph Reasoning from Sparse RGB Views
- Stress Testing Factual Consistency Metrics for Long-Document Summarization
- Harmonic Token Projection (HTP): A Vocabulary-Free, Training-Free, Deterministic, and Reversible Embedding Methodology
- TurkEmbed4Retrieval: Turkish Embedding Model for Retrieval Task
- Beyond Fact Retrieval: Episodic Memory for RAG with Generative Semantic Workspaces
- Do LLMs Feel? Teaching Emotion Recognition with Prompts, Retrieval, and Curriculum Learning
- Oh That Looks Familiar: A Novel Similarity Measure for Spreadsheet Template Discovery
- Rethinking Retrieval-Augmented Generation for Medicine: A Large-Scale, Systematic Expert Evaluation and Practical Insights
- Characterizing AI Manipulation Risks in Brazilian YouTube Climate Discourse
- Optimizing Chain-of-Thought Confidence via Topological and Dirichlet Risk Analysis
- BookAsSumQA: An Evaluation Framework for Aspect-Based Book Summarization via Question Answering
- FLEX: Continuous Agent Evolution via Forward Learning from Experience
- NILC: Discovering New Intents with LLM-assisted Clustering
- Building Specialized Software-Assistant ChatBot with Graph-Based Retrieval-Augmented Generation
- Search Is Not Retrieval: Decoupling Semantic Matching from Contextual Assembly in RAG
- Leak@k: Unlearning Does Not Make LLMs Forget Under Probabilistic Decoding
- Long Grounded Thoughts: Synthesizing Visual Problems and Reasoning Chains at Scale
- REFLEX: Reference-Free Evaluation of Log Summarization via Large Language Model Judgment
- ReGen: Generative Robot Simulation via Inverse Design
- Probabilistic Textual Time Series Depression Detection
- Dynamic Jointly Batch Selection for Data Efficient Machine Translation Fine-Tuning
- Advancing Equitable AI: Evaluating Cultural Expressiveness in LLMs for Latin American Contexts
- E-CARE: An Efficient LLM-based Commonsense-Augmented Framework for E-Commerce
- ForeRobo: Unlocking Infinite Simulation Data for 3D Goal-driven Robotic Manipulation
- Watermarking Large Language Models in Europe: Interpreting the AI Act in Light of Technology
- Towards Transparent Stance Detection: A Zero-Shot Approach Using Implicit and Explicit Interpretability
- Multi-Object Tracking Retrieval with LLaVA-Video: A Training-Free Solution to MOT25-StAG Challenge
- Leveraging LLM-based agents for social science research: insights from citation network simulations
- Sparse, self-organizing ensembles of local kernels detect rare statistical anomalies
- Do Androids Dream of Unseen Puppeteers? Probing for a Conspiracy Mindset in Large Language Models
- ZoFia: Zero-Shot Fake News Detection with Entity-Guided Retrieval and Multi-LLM Interaction
- The Curved Spacetime of Transformer Architectures
- ROBoto2: An Interactive System and Dataset for LLM-assisted Clinical Trial Risk of Bias Assessment
- Cache Mechanism for Agent RAG Systems
- Smart-Hiring: An Explainable end-to-end Pipeline for CV Information Extraction and Job Matching
- Large Lemma Miners: Can LLMs do Induction Proofs for Hardware?
- Keeping it Local, Tiny and Real: Automated Report Generation on Edge Computing Devices for Mechatronic-Based Cognitive Systems
- ReAcTree: Hierarchical LLM Agent Trees with Control Flow for Long-Horizon Task Planning
- IG-Pruning: Input-Guided Block Pruning for Large Language Models
- NSYNC: Negative Synthetic Image Generation for Contrastive Training to Improve Stylized Text-To-Image Translation
- Trove: A Flexible Toolkit for Dense Retrieval
- Towards LLM-Powered Task-Aware Retrieval of Scientific Workflows for Galaxy
- Rescuing the Unpoisoned: Efficient Defense against Knowledge Corruption Attacks on RAG Systems
- OceanAI: A Conversational Platform for Accurate, Transparent, Near-Real-Time Oceanographic Insights
- SpEx: A Spectral Approach to Explainable Clustering
- OmniBrainBench: A Comprehensive Multimodal Benchmark for Brain Imaging Analysis Across Multi-stage Clinical Tasks
- Teaching LLMs to See and Guide: Context-Aware Real-Time Assistance in Augmented Reality
- Do Methods to Jailbreak and Defend LLMs Generalize Across Languages?
- Separate the Wheat from the Chaff: Winnowing Down Divergent Views in Retrieval Augmented Generation
- Issue-Oriented Agent-Based Framework for Automated Review Comment Generation
- G2: Guided Generation for Enhanced Output Diversity in LLMs
- PreferThinker: Reasoning-based Personalized Image Preference Assessment
- LIR: The First Workshop on Late Interaction and Multi Vector Retrieval @ ECIR 2026
- LingGym: How Far Are LLMs from Thinking Like Field Linguists?
- IL-PCSR: Legal Corpus for Prior Case and Statute Retrieval
- Culture Cartography: Mapping the Landscape of Cultural Knowledge
- Effect of Domain Generalization Techniques in Low Resource Systems
- Thought Branches: Interpreting LLM Reasoning Requires Resampling
- Traceable Drug Recommendation over Medical Knowledge Graphs
- Relation-Aware Bayesian Optimization of DBMS Configurations Guided by Affinity Scores
- A Memory-Efficient Retrieval Architecture for RAG-Enabled Wearable Medical LLMs-Agents
- Addressing Longstanding Challenges in Cognitive Science with Language Models
- Independent Clinical Evaluation of General-Purpose LLM Responses to Signals of Suicide Risk
- How Similar Are Grokipedia and Wikipedia? A Multi-Dimensional Textual and Structural Comparison
- Cross-Platform Evaluation of Reasoning Capabilities in Foundation Models
- Value Drifts: Tracing Value Alignment During LLM Post-Training
- Hebrew Diacritics Restoration using Visual Representation
- SCRIBE: Structured Chain Reasoning for Interactive Behaviour Explanations using Tool Calling
- Graph-Enhanced Policy Optimization in LLM Agent Training
- Similarity-Distance-Magnitude Language Models
- Reasoning Path Divergence: A New Metric and Curation Strategy to Unlock LLM Diverse Thinking
- Agentic Economic Modeling
- Counterfactual-based Agent Influence Ranker for Agentic AI Workflows
- Hallucinations in Bibliographic Recommendation: Citation Frequency as a Proxy for Training Data Redundancy
- LLM-as-a-Judge for Evaluating System Responses in Conversational Music Recommendation
- Nudging Sustainable Choices through LLM-Generated Recommendation Explanations
- Language Through a Prism: A Spectral Approach for Multiscale Language Representations
- Diagnosing Fine-Grained Inconsistency Classification in Financial Disclosure Text
- Voice Memory for Agentic Speech Recognition
- Continuous Online Evaluation of Recommendation Strategies in Social Science Academic Search
- CaIRec: Calibrated Modality Imputation for Incomplete Multimodal Recommendation
- RAGuard: A Layered Defense Framework for Retrieval-Augmented Generation Systems Against Data Poisoning
- Constitutional Midtraining: Content Presence Drives Alignment Gains
- CrisisBERT: a Robust Transformer for Crisis Classification and Contextual Crisis Embedding
- IR-BERT: Leveraging BERT for Semantic Search in Background Linking for News Articles
- Supporting Workflow Reproducibility by Linking Bioinformatics Tools across Papers and Executable Code
- Tools Are Not Islands: Set-Level Tool Retrieval for LLM Agents via Query-Conditioned Hyperedge Prediction
- RedCodeAgent: Automatic Red-teaming Agent against Diverse Code Agents
- IMFuse: Instance-Aware Multi-Layer Fusion for LLM-Enhanced Sequential Recommendation
- Ideas Have Genomes: Benchmarking Scientific Lineage Reasoning and Lineage-Grounded Idea Generation
- Improving Human-Robot Teamwork in Urban Search and Rescue Through Episodic Memory of Prior Collaboration
- Do Methods Support the Claims? Intra-Paper Verification for Peer Review
- Can neurons speak? Semantic narration of vision at single-cell resolution
- Grammar as a Behavioral Biometric: Using Cognitively Motivated Grammar Models for Authorship Verification
- EUDAIMONIA: Evaluating Undesirable Dynamics in AI
- Enhancing Multi-Agent Communication through Attention Steering with Context Relevance
- Is Dimensionality a Barrier for Retrieval Models?
- Platonic Representations in the Human Brain: Unsupervised Recovery of Universal Geometry
- Can Knowledge-Graph-based Retrieval Augmented Generation Really Retrieve What You Need?
- Hallucination Benchmark for Speech Foundation Models
- Does Socialization Emerge in AI Agent Society? A Case Study of Moltbook
- The Rise of AI Agent Communities: Large-Scale Analysis of Discourse and Interaction on Moltbook
- From Reviews to Actionable Insights: An LLM-Based Approach for Attribute and Feature Extraction
- Readability-Robust Code Summarization via Meta Curriculum Learning
- Better Call Grep: Evaluating and Improving Grep-Like Lexical Retrieval for Repository-Level Code Completion
- ReviewSense: Transforming Customer Review Dynamics into Actionable Business Insights
- Depression Status Estimation by Deep Learning based Hybrid Multi-Modal Fusion Model
- Probing the Hidden Talent of ASR Foundation Models for L2 English Oral Assessment
- Self-Supervised Text-Vision Alignment for Automated Brain MRI Abnormality Detection: A Multicenter Study (ALIGN Study)
- Beyond One-Size-Fits-All: Personalized Harmful Content Detection with In-Context Learning
- DIRC-RAG: Accelerating Edge RAG with Robust High-Density and High-Loading-Bandwidth Digital In-ReRAM Computation
- Optimizing Knowledge Utilization for Multi-Intent Comment Generation with Large Language Models
- Testing Cross-Lingual Text Comprehension In LLMs Using Next Sentence Prediction
- SemCoT: Accelerating Chain-of-Thought Reasoning through Semantically-Aligned Implicit Tokens
- Iterative Critique-Refine Framework for Enhancing LLM Personalization
- Evaluating Joinable Column Discovery Approaches for Context-Aware Search
- Detecting the Use of Generative AI in Crowdsourced Surveys: Implications for Data Integrity
- Politically Speaking: LLMs on Changing International Affairs
- From Cross-Task Examples to In-Task Prompts: A Graph-Based Pseudo-Labeling Framework for In-context Learning
- Talk2Ref: A Dataset for Reference Prediction from Scientific Talks
- Metadata-Driven Retrieval-Augmented Generation for Financial Question Answering
- Text Simplification with Sentence Embeddings
- From Observability Data to Diagnosis: An Evolving Multi-agent System for Incident Management in Cloud Systems
- Utilising Large Language Models for Generating Effective Counter Arguments to Anti-Vaccine Tweets
- DynaStride: Dynamic Stride Windowing with MMCoT for Instructional Multi-Scene Captioning
- AfriMTEB and AfriE5: Benchmarking and Adapting Text Embedding Models for African Languages
- Beyond Understanding: Evaluating the Pragmatic Gap in LLMs' Cultural Processing of Figurative Language
- Small Language Models Offer Significant Potential for Science Community
- PixelRefer: A Unified Framework for Spatio-Temporal Object Referring with Arbitrary Granularity
- Minimizing Human Intervention in Online Classification
- COOPERA: Continual Open-Ended Human-Robot Assistance
- Evaluating Large Language Models for Stance Detection on Financial Targets from SEC Filing Reports and Earnings Call Transcripts
- Code Contribution and Credit in Science
- SwiftEmbed: Ultra-Fast Text Embeddings via Static Token Lookup for Real-Time Applications
- LangLingual: A Personalised, Exercise-oriented English Language Learning Tool Leveraging Large Language Models
- Modeling Political Discourse with Sentence-BERT and BERTopic
- Seeing the Unseen: Towards Zero-Shot Inspection for Wind Turbine Blades using Knowledge-Augmented Vision Language Models
- Agentic Meta-Orchestrator for Multi-task Copilots
- Iterative Layer Pruning for Efficient Translation Inference
- Multi-Modal Fact-Verification Framework for Reducing Hallucinations in Large Language Models
- REVISION:Reflective Intent Mining and Online Reasoning Auxiliary for E-commerce Visual Search System Optimization
- RaCoT: Plug-and-Play Contrastive Example Generation Mechanism for Enhanced LLM Reasoning Reliability
- CLIN-LLM: A Safety-Constrained Hybrid Framework for Clinical Diagnosis and Treatment Generation
- Benchmarking Egocentric Multimodal Goal Inference for Assistive Wearable Agents
- Knowledge-guided Continual Learning for Behavioral Analytics Systems
- Multimodal Item Scoring for Natural Language Recommendation via Gaussian Process Regression with LLM Relevance Judgments
- TagRuler: Interactive Tool for Span-Level Data Programming by\n Demonstration
- LLM-Generated Negative News Headlines Dataset: Creation and Benchmarking Against Real Journalism
- Gaze-VLM:Bridging Gaze and VLMs through Attention Regularization for Egocentric Understanding
- Large Language Models Meet Text-Attributed Graphs: A Survey of Integration Frameworks and Applications
- NoisyGRPO: Incentivizing Multimodal CoT Reasoning via Noise Injection and Bayesian Estimation
- Designing and Evaluating Hint Generation Systems for Science Education
- Dynamic Retriever for In-Context Knowledge Editing via Policy Optimization
- Preventing Catastrophic Forgetting: Behavior-Aware Sampling for Safer Language Model Fine-Tuning
- FicSim: A Dataset for Multi-Faceted Semantic Similarity in Long-Form Fiction
- Thought Communication in Multiagent Collaboration
- Structure-Conditional Minimum Bayes Risk Decoding
- Systematic Evaluation of Uncertainty Estimation Methods in Large Language Models
- Citation Failure: Definition, Analysis and Efficient Mitigation
- Rank-GRPO: Training LLM-based Conversational Recommender Systems with Reinforcement Learning
- Tri-Modal Severity Fused Diagnosis across Depression and Post-traumatic Stress Disorders
- RECALL: REpresentation-aligned Catastrophic-forgetting ALLeviation via Hierarchical Model Merging
- Re-evaluating Minimum Bayes Risk Decoding for Automatic Speech Recognition
- Dialogue Is Not Enough to Make a Communicative BabyLM (But Neither Is Developmentally Inspired Reinforcement Learning)
- Do Prompts Reshape Representations? An Empirical Study of Prompting Effects on Embeddings
- CircuitGuard: Mitigating LLM Memorization in RTL Code Generation Against IP Leakage
- Style Attack Disguise: When Fonts Become a Camouflage for Adversarial Intent
- NeSyPr: Neurosymbolic Proceduralization For Efficient Embodied Reasoning
- Learning Noise-Resilient and Transferable Graph-Text Alignment via Dynamic Quality Assessment
- Sign Language Translation with Sentence Embedding Supervision
- From Script to Stage: Automating Experimental Design for Social Simulations with LLMs
- JointCQ: Improving Factual Hallucination Detection with Joint Claim and Query Generation
- Selecting and Combining Large Language Models for Scalable Code Clone Detection
- Human-Agent Collaborative Paper-to-Page Crafting
- LLM-Augmented Symbolic NLU System for More Reliable Continuous Causal Statement Interpretation
- C2T-ID: Converting Semantic Codebooks to Textual Document Identifiers for Generative Search
- FlexiDataGen: An Adaptive LLM Framework for Dynamic Semantic Dataset Generation in Sensitive Domains
- SBAN: A Framework & Multi-Dimensional Dataset for Large Language Model Pre-Training and Software Code Mining
- FeClustRE: Hierarchical Clustering and Semantic Tagging of App Features from User Reviews
- Topoformer: brain-like topographic organization in Transformer language models through spatial querying and reweighting
- Beyond the Explicit: A Bilingual Dataset for Dehumanization Detection in Social Media
- One Size Fits All? A Modular Adaptive Sanitization Kit (MASK) for Customizable Privacy-Preserving Phone Scam Detection
- Unifying Inductive, Cross-Domain, and Multimodal Learning for Robust and Generalizable Recommendation
- IMB: An Italian Medical Benchmark for Question Answering
- PP3D: An In-Browser Vision-Based Defense Against Web Behavior Manipulation Attacks
- Med-VRAgent: A Framework for Medical Visual Reasoning-Enhanced Agents
- Evaluating LLM-Based Mobile App Recommendations: An Empirical Study
- KrishokBondhu: A Retrieval-Augmented Voice-Based Agricultural Advisory Call Center for Bengali Farmers
- PoSh: Using Scene Graphs To Guide LLMs-as-a-Judge For Detailed Image Descriptions
- Improving Topic Modeling of Social Media Short Texts with Rephrasing: A Case Study of COVID-19 Related Tweets
- ECKO: Explainable Clinical Knowledge for Oncology
- Learning from Generalization Patterns: An Evaluation-Driven Approach to Enhanced Data Augmentation for Fine-Tuning Small Language Models
- Exemplar-Guided Planing: Enhanced LLM Agent for KGQA
- CrossGuard: Safeguarding MLLMs against Joint-Modal Implicit Malicious Attacks
- AtlasKV: Augmenting LLMs with Billion-Scale Knowledge Graphs in 20GB VRAM
- AFRICAPTION: Establishing a New Paradigm for Image Captioning in African Languages
- Efficient Toxicity Detection in Gaming Chats: A Comparative Study of Embeddings, Fine-Tuned Transformers and LLMs
- TaxoAlign: Scholarly Taxonomy Generation Using Language Models
- StreamingThinker: Large Language Models Can Think While Reading
- OncoReason: Structuring Clinical Reasoning in LLMs for Robust and Interpretable Survival Prediction
- Who's Asking? Simulating Role-Based Questions for Conversational AI Evaluation
- Exact Nearest-Neighbor Search on Energy-Efficient FPGA Devices
- MOSAIC: Masked Objective with Selective Adaptation for In-domain Contrastive Learning
- LLM-based In-situ Thought Exchanges for Critical Paper Reading
- BiMax: Bidirectional MaxSim Score for Document-Level Alignment
- GraphMind: Interactive Novelty Assessment System for Accelerating Scientific Discovery
- ProofBridge: Auto-Formalization of Natural Language Proofs in Lean via Joint Embeddings
- Mixture of Experts Approaches in Dense Retrieval Tasks
- Leveraging LLMs for Context-Aware Implicit Textual and Multimodal Hate Speech Detection
- Latent Topic Synthesis: Leveraging LLMs for Electoral Ad Analysis
- DMRetriever: A Family of Models for Improved Text Retrieval in Disaster Management
- LoRAverse: A Submodular Framework to Retrieve Diverse Adapters for Diffusion Models
- LLMs as Scalable, General-Purpose Simulators For Evolving Digital Agent Training
- AI-Powered Early Diagnosis of Mental Health Disorders from Real-World Clinical Conversations
- TRI-DEP: A Trimodal Comparative Study for Depression Detection Using Speech, Text, and EEG
- Harmonizing Diverse Models: A Layer-wise Merging Strategy for Consistent Generation
- Leveraging Multimodal LLM Descriptions of Activity for Explainable Semi-Supervised Video Anomaly Detection
- DEXTER: Diffusion-Guided EXplanations with TExtual Reasoning for Vision Models
- Intent Clustering with Shared Pseudo-Labels
- Multimodal RAG for Unstructured Data:Leveraging Modality-Aware Knowledge Graphs with Hybrid Retrieval
- Stealthy Dual-Trigger Backdoors: Attacking Prompt Tuning in LM-Empowered Graph Foundation Models
- Holdout-Loss-Based Data Selection for LLM Finetuning via In-Context Learning
- Hierarchical Semantic Retrieval with Cobweb
- DPRF: A Generalizable Dynamic Persona Refinement Framework for Optimizing Behavior Alignment Between Personalized LLM Role-Playing Agents and Humans
- JEDA: Query-Free Clinical Order Search from Ambient Dialogues
- DROID: Dual Representation for Out-of-Scope Intent Detection
- When Embedding Models Meet: Procrustes Bounds and Applications
- ReMindRAG: Low-Cost LLM-Guided Knowledge Graph Traversal for Efficient RAG
- Putting on the Thinking Hats: A Survey on Chain of Thought Fine-tuning from the Perspective of Human Reasoning Mechanism
- Stable LLM Ensemble: Interaction between Example Representativeness and Diversity
- Revisiting Query Variants: The Advantage of Retrieval Over Generation of Query Variants for Effective QPP
- Program of Thoughts for Financial Reasoning: Leveraging Dynamic In-Context Examples and Generative Retrieval
- When Personalization Tricks Detectors: The Feature-Inversion Trap in Machine-Generated Text Detection
- Beating Harmful Stereotypes Through Facts: RAG-based Counter-speech Generation
- PromptLocate: Localizing Prompt Injection Attacks
- Unveiling the Vulnerability of Graph-LLMs: An Interpretable Multi-Dimensional Adversarial Attack on TAGs
- GOAT: A Training Framework for Goal-Oriented Agent with Tools
- Encapsulating Textual Contents into a MOC data Structure for Advanced Applications
- GRAVITY: A Framework for Personalized Text Generation via Profile-Grounded Synthetic Preferences
- Scaling Language-Centric Omnimodal Representation Learning
- REGENT: Relevance-Guided Attention for Entity-Aware Multi-Vector Neural Re-Ranking
- QDER: Query-Specific Document and Entity Representations for Multi-Vector Document Re-Ranking
- FinVet: A Collaborative Framework of RAG and External Fact-Checking Agents for Financial Misinformation Detection
- Situat3DChange: Situated 3D Change Understanding Dataset for Multimodal Large Language Model
- Automated Skill Decomposition Meets Expert Ontologies: Bridging the Granularity Gap with LLMs
- FOSSIL: Harnessing Feedback on Suboptimal Samples for Data-Efficient Generalisation with Imitation Learning for Embodied Vision-and-Language Tasks
- Do Psychometric Tests Work for Large Language Models? Evaluation of Tests on Sexism, Racism, and Morality
- RAG-Pull: Imperceptible Attacks on RAG Systems for Code Generation
- Secret-Protected Evolution for Differentially Private Synthetic Text Generation
- Chart-RVR: Reinforcement Learning with Verifiable Rewards for Explainable Chart Reasoning
- Learning to Watermark: A Selective Watermarking Framework for Large Language Models via Multi-Objective Optimization
- Scalable and Explainable Enterprise Knowledge Discovery Using Graph-Centric Hybrid Retrieval
- Quantum NLP models on Natural Language Inference
- Detecting Hallucinations in Authentic LLM-Human Interactions
- Testing and Enhancing Multi-Agent Systems for Robust Code Generation
- NIM: Neuro-symbolic Ideographic Metalanguage for Inclusive Communication
- Steering Over-refusals Towards Safety in Retrieval Augmented Generation
- Knowing Unknowns in an Age of Information Overload
- PrediQL: Automated Testing of GraphQL APIs with LLMs
- SimKey: A Semantically Aware Key Module for Watermarking Language Models
- Are LLMs Empathetic to All? Investigating the Influence of Multi-Demographic Personas on a Model's Empathy
- Diversity Augmentation of Dynamic User Preference Data for Boosting Personalized Text Summarizers
- Are LLMs Better GNN Helpers? Rethinking Robust Graph Learning under Deficiencies with Iterative Refinement
- Comparison of Unsupervised Metrics for Evaluating Judicial Decision Extraction
- SkillOS: Learning Skill Curation for Self-Evolving Agents
- The Geometry of Forgetting
- Text2Token: Unsupervised Text Representation Learning with Token Target Prediction
- Evolution of wartime discourse on Telegram: A comparative study of Ukrainian and Russian policymakers' communication before and after Russia's full-scale invasion of Ukraine
- HUME: Measuring the Human-Model Performance Gap in Text Embedding Tasks
- iBERT: Interpretable Style Embeddings via Sense Decomposition
- Steering Embedding Models with Geometric Rotation: Mapping Semantic Relationships Across Languages and Models
- From Birdwatch to Community Notes, from Twitter to X: four years of community-based content moderation
- SIMBA UQ: Similarity-Based Aggregation for Uncertainty Quantification in Large Language Models
- Doc2Query++: Topic-Coverage based Document Expansion and its Application to Dense Retrieval via Dual-Index Fusion
- Can We Reliably Rank Model Performance across Domains without Labeled Data?
- A Living Review Pipeline for AI/ML Applications in Accelerator Physics
- CFVBench: A Comprehensive Video Benchmark for Fine-grained Multimodal Retrieval-Augmented Generation
- Maple: A Multi-agent System for Portable Deep Learning across Clusters
- When Retrieval Succeeds and Fails: Rethinking Retrieval-Augmented Generation for LLMs
- FrameEOL: Semantic Frame Induction using Causal Language Models
- Semantic-Condition Tuning: Fusing Graph Context with Large Language Models for Knowledge Graph Completion
- A Human Behavioral Baseline for Collective Governance in Software Projects
- When LLM Agents Meet Graph Optimization: An Automated Data Quality Improvement Approach
- GRETEL: A Goal-driven Retrieval and Execution-based Trial Framework for LLM Tool Selection Enhancing
- ProxRouter: Proximity-Weighted LLM Query Routing for Improved Robustness to Outliers
- One Sentence, Two Embeddings: Contrastive Learning of Explicit and Implicit Semantic Representations
- DeepPrune: Parallel Scaling without Inter-trace Redundancy
- Leveraging Whisper Embeddings for Audio-based Lyrics Matching
- HySim-LLM: Embedding-Weighted Fine-Tuning Bounds and Manifold Denoising for Domain-Adapted LLMs
- From Keywords to Clusters: AI-Driven Analysis of YouTube Comments to Reveal Election Issue Salience in 2024
- Multilingual Generative Retrieval via Cross-lingual Semantic Compression
- FedBook: A Unified Federated Graph Foundation Codebook with Intra-domain and Inter-domain Knowledge Modeling
- MemWeaver: A Hierarchical Memory from Textual Interactive Behaviors for Personalized Generation
- Safely Exploring Novel Actions in Recommender Systems via Deployment-Efficient Policy Learning
- Struc-EMB: The Potential of Structure-Aware Encoding in Language Embeddings
- RAG4Tickets: AI-Powered Ticket Resolution via Retrieval-Augmented Generation on JIRA and GitHub Data
- Self-Improving LLM Agents at Test-Time
- ZeroCard: Cardinality Estimation with Zero Dependence on Target Databases -- No Data, No Query, No Retraining
- Measuring the Hidden Cost of Data Valuation through Collective Disclosure
- On the Alignment Between Supervised and Self-Supervised Contrastive Learning
- Efficient Prompt Optimisation for Legal Text Classification with Proxy Prompt Evaluator
- Multi-Task Pre-Finetuning of Lightweight Transformer Encoders for Text Classification and NER
- Investigating Thematic Patterns and User Preferences in LLM Interactions using BERTopic
- When Thoughts Meet Facts: Reusable Reasoning for Long-Context LMs
- Bridged Clustering: Semi-Supervised Sparse Bridging
- Reasoning for Hierarchical Text Classification: The Case of Patents
- Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications
- Search-R3: Unifying Reasoning and Embedding in Large Language Models
- SoftMatcha 2: A Fast and Soft Pattern Matcher for Trillion-Scale Corpora
- A Comparison of Independent and Joint Fine-tuning Strategies for Retrieval-Augmented Generation
- Mapping global bee research with traits and plant-pollinator interaction networks
- ImageNet-Think-250K: A Large-Scale Synthetic Dataset for Multimodal Reasoning for Vision Language Models
- Iterative design of a NAND hybrid riboswitch by deep batch Bayesian optimization
- Exposing Citation Vulnerabilities in Generative Engines
- Text2Stories: Evaluating the Alignment Between Stakeholder Interviews and Generated User Stories
- Study on LLMs for Promptagator-Style Dense Retriever Training
- TWIST: Training-free and Label-free Short Text Clustering through Iterative Vector Updating with LLMs
- PTEB: Towards Robust Text Embedding Evaluation via Stochastic Paraphrasing at Evaluation Time with LLMs
- Differentially Private Synthetic Text Generation for Retrieval-Augmented Generation (RAG)
- Auto-Stega: An Agent-Driven System for Lifelong Strategy Evolution in LLM-Based Text Steganography
- How Confident are Video Models? Empowering Video Models to Express their Uncertainty
- A Framework for Measuring How News Topics Drive Stock Movement
- Controllable Stylistic Text Generation with Train-Time Attribute-Regularized Diffusion
- Relational Transformer: Toward Zero-Shot Foundation Models for Relational Data
- RoSE: Round-robin Synthetic Data Evaluation for Selecting LLM Generators without Human Test Sets
- Mnemosyne: An Unsupervised, Human-Inspired Long-Term Memory Architecture for Edge-Based LLMs
- ARRC: Advanced Reasoning Robot Control - Knowledge-Driven Autonomous Manipulation Using Retrieval-Augmented Generation
- Automated Research Article Classification and Recommendation Using NLP and ML
- Redefining Cost Estimation in Database Systems: The Role of Execution Plan Features and Machine Learning
- Scalable In-context Ranking with Generative Models
- DeepV: A Model-Agnostic Retrieval-Augmented Framework for Verilog Code Generation with a High-Quality Knowledge Base
- Guided Query Refinement: Multimodal Hybrid Retrieval with Test-Time Optimization
- ModernBERT + ColBERT: Enhancing biomedical RAG through an advanced re-ranking retriever
- Topic-Specific Classifiers are Better Relevance Judges than Prompted LLMs
- Contrastive Learning Using Graph Embeddings for Domain Adaptation of Language Models in the Process Industry
- Fine-grained auxiliary learning for real-world product recommendation
- Residualized Similarity for Faithfully Explainable Authorship Verification
- AWARE, Beyond Sentence Boundaries: A Contextual Transformer Framework for Identifying Cultural Capital in STEM Narratives
- GRACE: Generative Representation Learning via Contrastive Policy Optimization
- Challenge on Optimization of Context Collection for Code Completion
- Learning Representations Through Contrastive Neural Model Checking
- Chronological Thinking in Full-Duplex Spoken Dialogue Language Models
- RLRF: Competitive Search Agent Design via Reinforcement Learning from Ranker Feedback
- LLM Microscope: What Model Internals Reveal About Answer Correctness and Context Utilization
- SECA: Semantically Equivalent and Coherent Attacks for Eliciting LLM Hallucinations
- Systematic Diagnosis of Brittle Reasoning in Large Language Models
- How Catastrophic is Your LLM? Certifying Risk in Conversation
- MetaMuse: Algorithm Generation via Creative Ideation
- TreePrompt: Leveraging Hierarchical Few-Shot Example Selection for Improved English-Persian and English-German Translation
- Generating High-Level Test Cases from Requirements using LLM: An Industry Study
- Triplet-Structured Knowledge Integration for Multi-Turn Medical Reasoning
- Consistent Kernel Change-Point Detection under m-Dependence for Text Segmentation
- External Data Extraction Attacks against Retrieval-Augmented Large Language Models
- Transparent Reference-free Automated Evaluation of Open-Ended User Survey Responses
- Time-To-Inconsistency: A Survival Analysis of Large Language Model Robustness to Adversarial Attacks
- LegalWiz: A Multi-Agent Generation Framework for Contradiction Detection in Legal Documents
- SEER: The Span-based Emotion Evidence Retrieval Benchmark
- Learning Efficient Guardrails for Compliance
- ModernVBERT: Towards Smaller Visual Document Retrievers
- LLM Routing with Dueling Feedback
- MultiPhysio-HRC: Multimodal Physiological Signals Dataset for industrial Human-Robot Collaboration
- PolyLink: A Blockchain Based Decentralized Edge AI Platform for LLM Inference
- JoyAgent-JDGenie: Technical Report on the GAIA
- Retrieval and Augmentation of Domain Knowledge for Text-to-SQL Semantic Parsing
- TokMem: Tokenized Procedural Memory for Large Language Models
- Learning Compact Representations of LLM Abilities via Item Response Theory
- RealClass: A Framework for Classroom Speech Simulation with Public Datasets and Game Engines
- Milco: Learned Sparse Retrieval Across Languages via a Multilingual Connector
- Stochastic Self-Organization in Multi-Agent Systems
- Learning to Route: A Rule-Driven Agent Framework for Hybrid-Source Retrieval-Augmented Generation
- Automated Alignment of Math Items to Content Standards in Large-Scale Assessments Using Language Models
- PrimeX: A Dataset of Worldview, Opinion, and Explanation
- Automatic Fact-checking in English and Telugu
- MotionRAG: Motion Retrieval-Augmented Image-to-Video Generation
- An Experimental Study on Generating Plausible Textual Explanations for Video Summarization
- CliniBench: A Clinical Outcome Prediction Benchmark for Generative and Encoder-Based Language Models
- RAE: A Neural Network Dimensionality Reduction Method for Nearest Neighbors Preservation in Vector Search
- Better with Less: Small Proprietary Models Surpass Large Language Models in Financial Transaction Understanding
- V-HUB: A Visual-Centric Humor Understanding Benchmark for Video LLMs
- Think Less, Label Better: Multi-Stage Domain-Grounded Synthetic Data Generation for Fine-Tuning Large Language Models in Telecommunications
- CustomIR: Unsupervised Fine-Tuning of Dense Embeddings for Known Document Corpora
- SafePassage: High-Fidelity Information Extraction with Black Box LLMs
- Investigating Language and Retrieval Bias in Multilingual Previously Fact-Checked Claim Detection
- How Well Do LLMs Imitate Human Writing Style?
- Of-SemWat: High-payload text embedding for semantic watermarking of AI-generated images with arbitrary size
- HarmMetric Eval: Benchmarking Metrics and Judges for LLM Harmfulness Assessment
- Fin-Ally: Pioneering the Development of an Advanced, Commonsense-Embedded Conversational AI for Money Matters
- ELASTIQ: EEG-Language Alignment with Semantic Task Instruction and Querying
- Let LLMs Speak Embedding Languages: Generative Text Embeddings via Iterative Contrastive Refinement
- ViReSkill: Vision-Grounded Replanning with Skill Memory for LLM-Based Planning in Lifelong Robot Learning
- Model Correlation Detection via Random Selection Probing
- Memory Transfer Planning: LLM-driven Context-Aware Code Adaptation for Robot Manipulation
- GeoVLM-R1: Reinforcement Fine-Tuning for Improved Remote Sensing Reasoning
- Pretraining with hierarchical memories: separating long-tail and common knowledge
- Assessing Large Language Models in Updating Their Forecasts with New Information
- AnveshanaAI: A Multimodal Platform for Adaptive AI/ML Education through Automated Question Generation and Interactive Assessment
- Semantic Representation of Processes with Ontology Design Patterns
- Enhancing LLM Steering through Sparse Autoencoder-Based Vector Refinement
- An Senegalese Legal Texts Structuration Using LLM-augmented Knowledge Graph
- BioArtlas: Computational Clustering of Multi-Dimensional Complexity in Bioart
- Detecting Escalation Level from Speech with Transfer Learning and Acoustic-Lexical Information Fusion
- LLM Watermark Evasion via Bias Inversion
- Open-Vocabulary Spatio-Temporal Scene Graph for Robot Perception and Teleoperation Planning
- Semantic Voting: A Self-Evaluation-Free Approach for Efficient LLM Self-Improvement on Unverifiable Open-ended Tasks
- The Geometry of Creative Variability: How Credal Sets Expose Calibration Gaps in Language Models
- Comparison of Scoring Rationales Between Large Language Models and Human Raters
- From Deferral to Learning: Online In-Context Knowledge Distillation for LLM Cascades
- JE-IRT: A Geometric Lens on LLM Abilities through Joint Embedding Item Response Theory
- "I Don't Think RAI Applies to My Model'' -- Engaging Non-champions with Sticky Stories for Responsible AI Work
- REMA: A Unified Reasoning Manifold Framework for Interpreting Large Language Model
- Representing LLMs in Prompt Semantic Task Space
- Jailbreaking on Text-to-Video Models via Scene Splitting Strategy
- Question-Driven Analysis and Synthesis: Building Interpretable Thematic Trees with LLMs for Text Clustering and Controllable Generation
- Library Hallucinations in LLMs: Risk Analysis Grounded in Developer Queries
- Context Parametrization with Compositional Adapters
- Think Right, Not More: Test-Time Scaling for Numerical Claim Verification
- Goal-Guided Efficient Exploration via Large Language Model in Reinforcement Learning
- MotivGraph-SoIQ: Integrating Motivational Knowledge Graphs and Socratic Dialogue for Enhanced LLM Ideation
- Semantic Agreement Enables Efficient Open-Ended LLM Cascades
- RobustFlow: Towards Robust Agentic Workflow Generation
- KurdSTS: The Kurdish Semantic Textual Similarity
- Does AI Coaching Prepare us for Workplace Negotiations?
- GRAB: A Risk Taxonomy--Grounded Benchmark for Unsupervised Topic Discovery in Financial Disclosures
- What Should I Cite? A RAG Benchmark for Academic Citation Prediction
- The GDN-CC Dataset: Automatic Corpus Clarification for AI-enhanced Democratic Citizen Consultations
- MMPlanner: Zero-Shot Multimodal Procedural Planning with Chain-of-Thought Object State Reasoning
- QuantMind: A Context-Engineering Based Knowledge Framework for Quantitative Finance
- Filtering with Confidence: When Data Augmentation Meets Conformal Prediction
- One Model, Many Morals: Uncovering Cross-Linguistic Misalignments in Computational Moral Reasoning
- Position: Human Factors Reshape Adversarial Analysis in Human-AI Decision-Making Systems
- Interactive Recommendation Agent with Active User Commands
- Semantic Clustering of Civic Proposals: A Case Study on Brazil's National Participation Platform
- Query-Centric Graph Retrieval Augmented Generation
- SGMem: Sentence Graph Memory for Long-Term Conversational Agents
- AutoIntent: AutoML for Text Classification
- Acoustic-based Gender Differentiation in Speech-aware Language Models
- PMark: Towards Robust and Distortion-free Semantic-level Watermarking with Channel Constraints
- Extracting Conceptual Knowledge to Locate Software Issues
- Rejuvenating Cross-Entropy Loss in Knowledge Distillation for Recommender Systems
- PseudoBridge: Pseudo Code as the Bridge for Better Semantic and Logic Alignment in Code Retrieval
- Distilling Many-Shot In-Context Learning into a Cheat Sheet
- Enrich-on-Graph: Query-Graph Alignment for Complex Reasoning with LLM Enriching
- CLIMATE-FEVER: A Dataset for Verification of Real-World Climate Claims
- Stability of In-Context Learning: A Spectral Coverage Perspective
- Human Semantic Representations of Social Interactions from Moving Shapes
- Document Summarization with Conformal Importance Guarantees
- Synergistic Enhancement of Requirement-to-Code Traceability: A Framework Combining Large Language Model based Data Augmentation and an Advanced Encoder
- MIXRAG : Mixture-of-Experts Retrieval-Augmented Generation for Textual Graph Understanding and Question Answering
- Play by the Type Rules: Inferring Constraints for LLM Functions in Declarative Programs
- AutoSpec: An Agentic Framework for Automatically Drafting Patent Specification
- SteinerSQL: Graph-Guided Mathematical Reasoning for Text-to-SQL Generation
- GuessingGame: Measuring the Informativeness of Open-Ended Questions in Large Language Models
- Creative Transformation in Literary Texts: Modelling Change Across Representational Levels
- Cognitive Load Limits in Large Language Models: Benchmarking Multi-Hop Reasoning
- AI-based scoring systematically underestimates conceptual understanding of linguistically weak students' explanations in physics
- GLM-RAG: Graph Language Models for Graph-Based Retrieval-Augmented Generation
- Chaos in reason: How chain-of-thought LLMs can look for an answer
- Rehearse: Stepping Back from the Confidence Cliff in Self-Improving Autoresearch
- CDAE: Enhancing Perturbation Robustness in Pretrained Language Models with Contrastive Denoising
- Diversifying Personalized Research Ideation against AI-Induced Homogenization
- GGC: Selective Query Correction for Reliable Text-to-SPARQL Generation
- Reasoning Consensus: Structural Ensembling of LLM Reasoning via Weighted DAG Aggregation
- Cocktail-Talker: Multi-Speaker Dialog Modeling in Noisy Social Environments with Turn Action GRPO
- Gradient-free Task-Conditioned Retrieval for On-Device In-Context Learning
- From Backlog Items to Security Guidance: Towards Continuous Security Compliance
- MedLLM: An Open Medical Language Model at the Sub-Billion Scale
- ThreatForest: Multi-Agent Attack Tree Generation with Pluggable TTP Framework Mapping
- CAIT: A Syntactic Parsing Toolkit for Child-Adult InTeractions
- RELISH: LLM REgression with a Latent Iterative State Head
- LLM2Vec-Gen: Generative Embeddings from Large Language Models
- GCT: A Granger-Causal Transformer for Multivariate Traffic Analysis in Smart Villages
- Previously on... Automating Code Review
- MERMAID: Metaphor Generation with Symbolism and Discriminative Decoding
- Text Similarity Using Word Embeddings to Classify Misinformation
- WOER suchet, der findet nicht! Identifikation von thematisch verwandten OER-Materialien
- Enhancing Diversity in News Recommendations Increases Click-Through Rates: Insights from an Online Experiment and User Study
- A study of word embedding models for measuring topic coherence
- PhantomLint: Principled Detection of Hidden LLM Prompts in Structured Documents
- LexSemBridge: Fine-Grained Dense Representation Enhancement through Token-Aware Embedding Augmentation
- PLM-interact: extending protein language models to predict protein-protein interactions
- Randomly Removing 50% of Dimensions in Text Embeddings has Minimal Impact on Retrieval and Classification Tasks
- ConViS-Bench: Estimating Video Similarity Through Semantic Concepts
- AgentInit: Initializing LLM-based Multi-Agent Systems via Diversity and Expertise Orchestration for Effective and Efficient Collaboration
- Text Meets Topology: Rethinking Out-of-distribution Detection in Text-Rich Networks
- Single-Branch Network Architectures to Close the Modality Gap in Multimodal Recommendation
- Agentic AutoSurvey: Let LLMs Survey LLMs
- CCQA: Generating Question from Solution Can Improve Inference-Time Reasoning in SLMs
- Towards Synthesizing Normative Data for Cognitive Assessments Using Generative Multimodal Large Language Models
- Investigating Traffic Accident Detection Using Multimodal Large Language Models
- Geometric Structures and Patterns of Meaning: A PHATE Manifold Analysis of Chinese Character Embeddings
- Confidence-Aware Routing for Large Language Model Reliability Enhancement: A Multi-Signal Approach to Pre-Generation Hallucination Mitigation
- Extracting Conceptual Spaces from LLMs Using Prototype Embeddings
- Are Smaller Open-Weight LLMs Closing the Gap to Proprietary Models for Biomedical Question Answering?
- AIRwaves at CheckThat! 2025: Retrieving Scientific Sources for Implicit Claims on Social Media with Dual Encoders and Neural Re-Ranking
- Mind Your Ps and Qs: Supporting Positive Reinforcement in Moderation Through a Positive Queue
- PEEK: Guiding and Minimal Image Representations for Zero-Shot Generalization of Robot Manipulation Policies
- Semantic Reformulation Entropy for Robust Hallucination Detection in QA Tasks
- Scale-free Characteristics of Multilingual Legal Texts and the Limitations of LLMs
- How Persuasive is Your Context?
- Transformer-Gather, Fuzzy-Reconsider: A Scalable Hybrid Framework for Entity Resolution
- SilentStriker:Toward Stealthy Bit-Flip Attacks on Large Language Models
- Localizing Malicious Outputs from CodeLLM
- Scaling Law for Recommendation Models: Towards General-purpose User Representations
- Quantum Abduction: A New Paradigm for Reasoning under Uncertainty
- VidCLearn: A Continual Learning Approach for Text-to-Video Generation
- Semantic-Driven Topic Modeling for Analyzing Creativity in Virtual Brainstorming
- Cognitive Linguistic Identity Fusion Score (CLIFS): A Scalable Cognition-Informed Approach to Quantifying Identity Fusion from Text
- Learn to Rank Risky Investors: A Case Study of Predicting Retail Traders' Behaviour and Profitability
- Long document summarization using page specific target text alignment and distilling page importance
- mmExpert: Integrating Large Language Models for Comprehensive mmWave Data Synthesis and Understanding
- GRIL: Knowledge Graph Retrieval-Integrated Learning with Large Language Models
- MPCG: Multi-Round Persona-Conditioned Generation for Modeling the Evolution of Misinformation with LLMs
- The Role of Vocabularies in Learning Sparse Representations for Ranking
- Patterns in the Transition From Founder-Leadership to Community Governance of Open Source
- Reward Hacking Mitigation using Verifiable Composite Rewards
- How do Language Models Generate Slang: A Systematic Comparison between Human and Machine-Generated Slang Usages
- LibriTTS-VI: A Public Corpus and Novel Methods for Efficient Voice Impression Control
- Latent Zoning Network: A Unified Principle for Generative Modeling, Representation Learning, and Classification
- Evaluating the Effectiveness and Scalability of LLM-Based Data Augmentation for Retrieval
- Enhancing Financial RAG with Agentic AI and Multi-HyDE: A Novel Approach to Knowledge Retrieval and Hallucination Reduction
- SERVAL: Surprisingly Effective Zero-Shot Visual Document Retrieval Powered by Large Vision and Language Models
- LLM-Assisted Topic Reduction for BERTopic on Social Media Data
- Efficient and Versatile Model for Multilingual Information Retrieval of Islamic Text: Development and Deployment in Real-World Scenarios
- Quantifying Self-Awareness of Knowledge in Large Language Models
- An Artificial Intelligence Driven Semantic Similarity-Based Pipeline for Rapid Literature
- Spatial-CLAP: Learning Spatially-Aware audio--text Embeddings for Multi-Source Conditions
- Reveal and Release: Iterative LLM Unlearning with Self-generated Data
- TableDART: Dynamic Adaptive Multi-Modal Routing for Table Understanding
- Catch Me If You Can? Not Yet: LLMs Still Struggle to Imitate the Implicit Writing Styles of Everyday Authors
- Retrieval Capabilities of Large Language Models Scale with Pretraining FLOPs
- Words to Waves: Emotion-Adaptive Music Recommendation System
- PhenoGnet: A Graph-Based Contrastive Learning Framework for Disease Similarity Prediction
- MICA: Multi-Agent Industrial Coordination Assistant
- DSpAST: Disentangled Representations for Spatial Audio Reasoning with Large Language Models
- Who Taught the Lie? Responsibility Attribution for Poisoned Knowledge in Retrieval-Augmented Generation
- Controllable Pareto Trade-off between Fairness and Accuracy
- GeoAware-VLA: Implicit Geometry Aware Vision-Language-Action Model
- Evaluating LLM Alignment on Personality Inference from Real-World Interview Data
- The Few-shot Dilemma: Over-prompting Large Language Models
- Rethinking the Evaluation of Alignment Methods: Insights into Diversity, Generalisation, and Safety
- Conan-Embedding-v2: Training an LLM from Scratch for Text Embeddings
- FedMentor: Domain-Aware Differential Privacy for Heterogeneous Federated LLMs in Mental Health
- Are You Sure You're Positive? Consolidating Chain-of-Thought Agents with Uncertainty Quantification for Aspect-Category Sentiment Analysis
- Don't Change My View: Ideological Bias Auditing in Large Language Models
- ScaleDoc: Scaling LLM-based Predicates over Large Document Collections
- LEAF: Knowledge Distillation of Text Embedding Models with Teacher-Aligned Representations
- Retrieve-and-Verify: A Table Context Selection Framework for Accurate Column Annotations
- Smoothed Contrastive Learning for Unsupervised Sentence Embedding
- MA-DPR: Manifold-aware Distance Metrics for Dense Passage Retrieval
- Optimizing Agricultural Research: A RAG-Based Approach to Mycorrhizal Fungi Information
- Text Adaptation to Plain Language and Easy Read via Automatic Post-Editing Cycles
- MTEB-NL and E5-NL: Embedding Benchmark and Models for Dutch
- The Power of Framing: How News Headlines Guide Search Behavior
- Query-Focused Extractive Summarization for Sentiment Explanation
- Digital Voices of Survival: From Social Media Disclosures to Support Provisions for Domestic Violence Victims
- Zero-shot Multimodal Document Retrieval via Cross-modal Question Generation
- Graph-Enhanced Retrieval-Augmented Question Answering for E-Commerce Customer Support
- Pluralistic Off-policy Evaluation and Alignment
- Context-Aware Language Models for Forecasting Market Impact from Sequences of Financial News
- GRAID: Synthetic Data Generation with Geometric Constraints and Multi-Agentic Reflection for Harmful Content Detection
- AKCIT-FN at CheckThat! 2025: Switching Fine-Tuned SLMs and LLM Prompting for Multilingual Claim Normalization
- Topic Coverage-based Demonstration Retrieval for In-Context Learning
- Understanding Prompt Management in GitHub Repositories: A Call for Best Practices
- CEMTM: Contextual Embedding-based Multimodal Topic Modeling
- Decoding Plastic Toxicity: An Intelligent Framework for Conflict-Aware Relational Metapath Extraction from Scientific Abstracts
- Difficulty-Aware Agentic Orchestration for Query-Specific Multi-Agent Workflows
- A Biosecurity Agent for Lifecycle LLM Biosecurity Alignment
- Text2Sign Diffusion: A Generative Approach for Gloss-Free Sign Language Production
- GAPrune: Gradient-Alignment Pruning for Domain-Aware Embeddings
- Evaluating Large Language Models for Evidence-Based Clinical Question Answering
- Developer-LLM Conversations: An Empirical Study of Interactions and Generated Code Quality
- The Language of Approval: Identifying the Drivers of Positive Feedback Online
- Immunizing Images from Text to Image Editing via Adversarial Cross-Attention
- JU-NLP at Touché: Covert Advertisement in Conversational AI-Generation and Detection Strategies
- The anatomy of Green AI technologies: structure, evolution, and impact
- Established Psychometric vs. Ecologically Valid Questionnaires: Rethinking Psychological Assessments in Large Language Models
- Beyond the Silence: How Men Navigate Infertility Through Digital Communities and Data Sharing
- Investigating red packet fraud in Android applications: Insights from user reviews
- ZapGPT: Free-form Language Prompting for Simulated Cellular Control
- Topic-Guided Reinforcement Learning with LLMs for Enhancing Multi-Document Summarization
- A Modular and Multimodal Generative AI Framework for Urban Building Energy Data: Generating Synthetic Homes
- Retrieval-Augmented Generation for Reliable Interpretation of Radio Regulations
- Gene-R1: Reasoning with Data-Augmented Lightweight LLMs for Gene Set Analysis
- Graph Alignment via Dual-Pass Spectral Encoding and Latent Space Communication
- Towards Explainable Job Title Matching: Leveraging Semantic Textual Relatedness and Knowledge Graphs
- SEDM: Scalable Self-Evolving Distributed Memory for Agents
- Modelling Analogies and Analogical Reasoning: Connecting Cognitive Science Theory and NLP Research
- From scratch to silver: Creating trustworthy training data for patent-SDG classification using Large Language Models
- Chat-Driven Reconfiguration of Model Predictive Control
- Modality Alignment with Multi-scale Bilateral Attention for Multimodal Recommendation
- InterAct: Advancing Large-Scale Versatile 3D Human-Object Interaction Generation
- LLM-JEPA: Large Language Models Meet Joint Embedding Predictive Architectures
- SALMAN: Stability Analysis of Language Models Through the Maps Between Graph-based Manifolds
- Do All Autoregressive Transformers Remember Facts the Same Way? A Cross-Architecture Analysis of Recall Mechanisms
- JUDGEBERT: Assessing Legal Meaning Preservation Between Sentences
- TriagerX: Dual Transformers for Bug Triaging Tasks with Content and Interaction Based Rankings
- LLM Ensemble for RAG: Role of Context Length in Zero-Shot Question Answering for BioASQ Challenge
- Explainable Semantic Text Relations: A Question-Answering Framework for Comparing Document Content
- ALIGNS: Unlocking nomological networks in psychological measurement through a large language model
- Handling Open-Vocabulary Constructs in Formalizing Specifications: Retrieval-Augmented Parsing with Expert Knowledge
- Two Facets of the Same Optimization Coin: Model Degradation and Representation Collapse in Graph Foundation Models
- ImportSnare: Directed "Code Manual" Hijacking in Retrieval-Augmented Code Generation
- Are LLMs Enough for Hyperpartisan, Fake, Polarized and Harmful Content Detection? Evaluating In-Context Learning vs. Fine-Tuning
- The Role of Exploration Modules in Small Language Models for Knowledge Graph Question Answering
- Guarding Your Conversations: Privacy Gatekeepers for Secure Interactions with Cloud-Based AI Models
- ALLabel: Three-stage Active Learning for LLM-based Entity Recognition using Demonstration Retrieval
- GLEAM: Learning to Match and Explain in Cross-View Geo-Localization
- Biased Tales: Cultural and Topic Bias in Generating Children's Stories
- OmniMap: A General Mapping Framework Integrating Optics, Geometry, and Semantics
- NOWJ@COLIEE 2025: A Multi-stage Framework Integrating Embedding Models and Large Language Models for Legal Retrieval and Entailment
- LLM Analysis of 150+ years of German Parliamentary Debates on Migration Reveals Shift from Post-War Solidarity to Anti-Solidarity in the Last Decade
- Towards EnergyGPT: A Large Language Model Specialized for the Energy Sector
- On the Evaluation of Conditional GANs
- How Small Transformation Expose the Weakness of Semantic Similarity Measures
- mmBERT: A Modern Multilingual Encoder with Annealed Language Learning
- SVGauge: Towards Human-Aligned Evaluation for SVG Generation
- Benchmarking Information Retrieval Models on Complex Retrieval Tasks
- Analysis of Blood Report Images Using General Purpose Vision-Language Models
- Augmented Fine-Tuned LLMs for Enhanced Recruitment Automation
- Multimodal Fine-grained Context Interaction Graph Modeling for Conversational Speech Synthesis
- Learning to Route: Per-Sample Adaptive Routing for Multimodal Multitask Prediction
- An Optimized Pipeline for Automatic Educational Knowledge Graph Construction
- Ontology-Aligned Embeddings for Data-Driven Labour Market Analytics
- KGRAG-SC: Knowledge Graph RAG-Assisted Semantic Communication
- Evaluating Cognitive-Behavioral Fixation via Multimodal User Viewing Patterns on Social Media
- Finding your MUSE: Mining Unexpected Solutions Engine
- ThumbnailTruth: A Multi-Modal LLM Approach for Detecting Misleading YouTube Thumbnails Across Diverse Cultural Settings
- Conceptual Schema Inference for Tabular Datasets using Large Language Models
- Enhancing Technical Documents Retrieval for RAG
- MAGneT: Coordinated Multi-Agent Generation of Synthetic Multi-Turn Mental Health Counseling Sessions
- NER Retriever: Zero-Shot Named Entity Retrieval with Type-Aware Embeddings
- Anti-establishment sentiment on TikTok: Implications for understanding influence(rs) and expertise on social media
- AI-in-the-Loop: Privacy Preserving Real-Time Scam Detection and Conversational Scambaiting by Leveraging LLMs and Federated Learning
- PersonaTeaming: Exploring How Introducing Personas Can Improve Automated AI Red-Teaming
- MLSD: A Novel Few-Shot Learning Approach to Enhance Cross-Target and Cross-Domain Stance Detection
- Explainable Knowledge Graph Retrieval-Augmented Generation (KG-RAG) with KG-SMILE
- The Impact of Critique on LLM-Based Model Generation from Natural Language: The Case of Activity Diagrams
- VQualA 2025 Challenge on Engagement Prediction for Short Videos: Methods and Results
- Grocery to General Merchandise: A Cross-Pollination Recommender using LLMs and Real-Time Cart Context
- IDEAlign: Comparing Large Language Models to Human Experts in Open-ended Interpretive Annotations
- Scaling behavior of large language models in emotional safety classification across sizes and tasks
- L3Cube-IndicHeadline-ID: A Dataset for Headline Identification and Semantic Evaluation in Low-Resource Indian Languages
- Benchmarking Large Language Models for Personalized Guidance in AI-Enhanced Learning
- CLEAR: Contrastive Learning for Sentence Representation
- An Epidemiological Knowledge Graph extracted from the World Health Organization's Disease Outbreak News
- Attributes as Textual Genes: Leveraging LLMs as Genetic Algorithm Simulators for Conditional Synthetic Data Generation
- StructCoh: Structured Contrastive Learning for Context-Aware Text Semantic Matching
- Avoidance Decoding for Diverse Multi-Branch Story Generation
- Towards Agents That Know When They Don't Know: Uncertainty as a Control Signal for Structured Reasoning
- Take That for Me: Multimodal Exophora Resolution with Interactive Questioning for Ambiguous Out-of-View Instructions
- KoBLEX: Open Legal Question Answering with Multi-hop Reasoning
- Towards Open-World Retrieval-Augmented Generation on Knowledge Graph: A Multi-Agent Collaboration Framework
- Natural Context Drift Undermines the Natural Language Understanding of Large Language Models
- Hierarchical Motion Captioning Utilizing External Text Data Source
- Enhancing Large Language Model for Knowledge Graph Completion via Structure-Aware Alignment-Tuning
- AMAZe: A Multi-Agent Zero-shot Index Advisor for Relational Databases
- Testing the assumptions about the geometry of sentence embedding spaces: the cosine measure need not apply
- Dissecting Atomic Facts: Visual Analytics for Improving Fact Annotations in Language Model Evaluation
- Understanding Fanchuan in Livestreaming Platforms: A New Form of Online Antisocial Behavior
- Decomposing and Revising What Language Models Generate
- MLLMRec: Exploring the Potential of Multimodal Large Language Models in Recommender Systems
- OpinioRAG: Towards Generating User-Centric Opinion Highlights from Large-scale Online Reviews
- Modeling Motivated Reasoning in Law: Evaluating Strategic Role Conditioning in LLM Summarization
- Standard vs. Modular Sampling: Best Practices for Reliable LLM Unlearning
- Data Auctions for Retrieval Augmented Generation
- The Demon is in Ambiguity: Revisiting Situation Recognition with Single Positive Multi-Label Learning
- T-Retrievability: A Topic-Focused Approach to Measure Fair Document Exposure in Information Retrieval
- QZhou-Embedding Technical Report
- L3Cube-MahaSTS: A Marathi Sentence Similarity Dataset and Models
- Efficient Code Embeddings from Code Generation Models
- Synthetic CVs To Build and Test Fairness-Aware Hiring Tools
- GSTBench: A Benchmark Study on the Transferability of Graph Self-Supervised Learning
- InSQuAD: In-Context Learning for Efficient Retrieval via Submodular Mutual Information to Enforce Quality and Diversity
- Graph-Based Feature Augmentation for Predictive Tasks on Relational Datasets
- Native Logical and Hierarchical Representations with Subspace Embeddings
- ConspirED: A Dataset for Cognitive Traits of Conspiracy Theories and Large Language Model Safety
- From Post To Personality: Harnessing LLMs for MBTI Prediction in Social Media
- SciTopic: Enhancing Topic Discovery in Scientific Literature through Advanced LLM
- Leveraging Semantic Triples for Private Document Generation with Local Differential Privacy Guarantees
- Exploring Selective Retrieval-Augmentation for Long-Tail Legal Text Classification
- NLKI: A lightweight Natural Language Knowledge Integration Framework for Improving Small VLMs in Commonsense VQA Tasks
- SPELUNKER: Item Similarity Search Using Large Language Models and Custom K-Nearest Neighbors
- Functional Consistency of LLM Code Embeddings: A Self-Evolving Data Synthesis Framework for Benchmarking
- Continual Neural Topic Model
- LegiScout: A Visual Tool for Understanding Complex Legislation
- SafetyFlow: An Agent-Flow System for Automated LLM Safety Benchmarking
- Network-Level Prompt and Trait Leakage in Local Research Agents
- SIExVulTS: Sensitive Information Exposure Vulnerability Detection System using Transformer Models and Static Analysis
- Inference Gap in Domain Expertise and Machine Intelligence in Named Entity Recognition: Creation of and Insights from a Substance Use-related Dataset
- Stack Trace-Based Crash Deduplication with Transformer Adaptation
- DeepMEL: A Multi-Agent Collaboration Framework for Multimodal Entity Linking
- The Double-edged Sword of LLM-based Data Reconstruction: Understanding and Mitigating Contextual Vulnerability in Word-level Differential Privacy Text Sanitization
- Membership Inference Attacks on LLM-based Recommender Systems
- Controllable Conversational Theme Detection Track at DSTC 12
- Flexible metadata harvesting for ecology using large language models
- Constraint Matters: Multi-Modal Representation for Reducing Mixed-Integer Linear programming
- Rethinking Caching for LLM Serving Systems: Beyond Traditional Heuristics
- Granite Embedding R2 Models
- Uncovering Intervention Opportunities for Suicide Prevention with Language Model Assistants
- DenseRec: Revisiting Dense Content Embeddings for Sequential Transformer-based Recommendation
- Leveraging Large Language Models for Accurate Sign Language Translation in Low-Resource Scenarios
- InReAcTable: LLM-Powered Interactive Visual Data Story Construction from Tabular Data
- Named Entity Recognition of Historical Text via Large Language Model
- LLM-Guided Genetic Improvement: Envisioning Semantic Aware Automated Software Evolution
- Reference and Document Aware Semantic Evaluation Methods for Korean Language Summarization
- Subjective Behaviors and Preferences in LLM: Language of Browsing
- Continuous sentiment scores for literary and multilingual contexts
- Towards Skeletal and Signer Noise Reduction in Sign Language Production via Quaternion-Based Pose Encoding and Contrastive Learning
- Supporting Clustering with Contrastive Learning
- Democratizing News Recommenders: Modeling Multiple Perspectives for News Candidate Generation with VQ-VAE
- InPars+: Supercharging Synthetic Data Generation for Information Retrieval Systems
- The illusion of a perfect metric: Why evaluating AI's words is harder than it looks
- Interactive Query Answering on Knowledge Graphs with Soft Entity Constraints
- AdaDocVQA: Adaptive Framework for Long Document Visual Question Answering in Low-Resource Settings
- Statistical Comparative Analysis of Semantic Similarities and Model Transferability Across Datasets for Short Answer Grading
- Driving Style Recognition Like an Expert Using Semantic Privileged Information from Large Language Models
- A Functionality-Grounded Benchmark for Evaluating Web Agents in E-commerce Domains
- XAMT: Cross-Framework API Matching for Testing Deep Learning Libraries
- LumiMAS: A Comprehensive Framework for Real-Time Monitoring and Enhanced Observability in Multi-Agent Systems
- Structuring the Unstructured: A Systematic Review of Text-to-Structure Generation for Agentic AI with a Universal Evaluation Framework
- SEA-BED: Southeast Asia Embedding Benchmark
- Cost-Aware Contrastive Routing for LLMs
- Scalable RF Simulation in Generative 4D Worlds
- Multi-Modal Drift Forecasting of Leeway Objects via Navier-Stokes-Guided CNN and Sequence-to-Sequence Attention-Based Models
- LLM-as-a-Judge for Privacy Evaluation? Exploring the Alignment of Human and LLM Perceptions of Privacy in Textual Data
- In-Context Examples Matter: Improving Emotion Recognition in Conversation with Instruction Tuning
- Controlling Multimodal LLMs via Reward-guided Decoding
- CoDiEmb: A Collaborative yet Distinct Framework for Unified Representation Learning in Information Retrieval and Semantic Textual Similarity
- Retrieval-augmented reasoning with lean language models
- ORFuzz: Fuzzing the "Other Side" of LLM Safety -- Testing Over-Refusal
- RAG for Geoscience: What We Expect, Gaps and Opportunities
- From Feedback to Failure: Automated Android Performance Issue Reproduction
- Copyright Protection for Large Language Models: A Survey of Methods, Challenges, and Trends
- GenOM: Ontology Matching with Description Generation and Large Language Model
- IADGPT: Unified LVLM for Few-Shot Industrial Anomaly Detection, Localization, and Reasoning via In-Context Learning
- ChatENV: An Interactive Vision-Language Model for Sensor-Guided Environmental Monitoring and Scenario Simulation
- ORBIT: An Object Property Reasoning Benchmark for Visual Inference Tasks
- Semantic IDs for Joint Generative Search and Recommendation
- Dataset Construction for Training LLM to Learn Analog Circuit Knowledge
- Empowering Multimodal LLMs with External Tools: A Comprehensive Survey
- Prompt-Response Semantic Divergence Metrics for Faithfulness Hallucination and Misalignment Detection in Large Language Models
- DS4RS: Community-Driven and Explainable Dataset Search Engine for Recommender System Research
- Estimating Machine Translation Difficulty
- Amazon Nova AI Challenge -- Trusted AI: Advancing secure, AI-assisted software development
- January Food Benchmark (JFB): A Public Benchmark Dataset and Evaluation Suite for Multimodal Food Analysis
- Social-Sensor Identity Cloning Detection Using Weakly Supervised Deep Forest and Cryptographic Authentication
- Towards Self-cognitive Exploration: Metacognitive Knowledge Graph Retrieval Augmented Generation
- UWBa at SemEval-2025 Task 7: Multilingual and Crosslingual Fact-Checked Claim Retrieval
- IAG: Input-aware Backdoor Attack on VLM-based Visual Grounding
- Semantic Distance: A New Metric for ASR Performance Analysis Towards Spoken Language Understanding
- Link Prediction for Event Logs in the Process Industry
- BiasGym: A Simple and Generalizable Framework for Analyzing and Removing Biases through Elicitation
- DiffPose-Animal: A Language-Conditioned Diffusion Framework for Animal Pose Estimation
- Exploring Palette based Color Guidance in Diffusion Models
- GreenTEA: Gradient Descent with Topic-modeling and Evolutionary Auto-prompting
- E3-Rewrite: Learning to Rewrite SQL for Executability, Equivalence,and Efficiency
- Jointly Generating and Attributing Answers using Logits of Document-Identifier Tokens
- Mitigating Popularity Bias in Counterfactual Explanations using Large Language Models
- BlindGuard: Safeguarding LLM-based Multi-Agent Systems under Unknown Attacks
- AdaptFlow: Adaptive Workflow Optimization via Meta-Learning
- DIVER: A Multi-Stage Approach for Reasoning-intensive Information Retrieval
- UniSVG: A Unified Dataset for Vector Graphic Understanding and Generation with Multimodal Large Language Models
- Semantic Caching for Low-Cost LLM Serving: From Offline Learning to Online Adaptation
- Retrieval-Augmented Multi-Agent System for Rapid Statement of Work Generation
- In-situ Value-aligned Human-Robot Interactions with Physical Constraints
- GLiClass: Generalist Lightweight Model for Sequence Classification Tasks
- Temporal User Profiling with LLMs: Balancing Short-Term and Long-Term Preferences for Recommendations
- Using LLMs to Capture Users' Temporal Context for Recommendation
- Improving Document Retrieval Coherence for Semantically Equivalent Queries
- What am I missing here?: Evaluating Large Language Models for Masked Sentence Prediction
- VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding
- Grounding Natural Language for Multi-agent Decision-Making with Multi-agentic LLMs
- Are Multimodal Embeddings Truly Beneficial for Recommendation? A Deep Dive into Whole vs. Individual Modalities
- An Embodied AR Navigation Agent: Integrating BIM with Retrieval-Augmented Generation for Language Guidance
- A Survey on Non-Intrusive ASR Refinement: From Output-Level Correction to Full-Model Distillation
- Incorporating Contextual Paralinguistic Understanding in Large Speech-Language Models
- Graph Neural Network for Product Recommendation on the Amazon Co-purchase Graph
- Canvas3D: Empowering Precise Spatial Control for Image Generation with Constraints from a 3D Virtual Canvas
- Multi-Modal Requirements Data-based Acceptance Criteria Generation using LLMs
- ESNERA: Empirical and semantic named entity alignment for named entity dataset merging
- Towards Effective Prompt Stealing Attack against Text-to-Image Diffusion Models
- BiXSE: Improving Dense Retrieval via Probabilistic Graded Relevance Distillation
- Vec2Summ: Text Summarization via Probabilistic Sentence Embeddings
- Position: Ideas Should be the Center of Machine Learning Research
- LLM Unlearning using Gradient Ratio-Based Influence Estimation and Noise Injection
- MIRA: Leveraging Multi-Intention Co-click Information in Web-scale Document Retrieval using Deep Neural Networks
- EICAP: Deep Dive in Assessment and Enhancement of Large Language Models in Emotional Intelligence through Multi-Turn Conversations
- Multi-Objective Instruction-Aware Representation Learning in Procedural Content Generation RL
- LinguaFluid: Language Guided Fluid Control via Semantic Rewards in Reinforcement Learning
- Prosocial Behavior Detection in Player Game Chat: From Aligning Human-AI Definitions to Efficient Annotation at Scale
- Scaling Personality Control in LLMs with Big Five Scaler Prompts
- Automatic Semantic Alignment of Flow Pattern Representations for Exploration with Large Language Models
- Omni Geometry Representation Learning vs Large Language Models for Geospatial Entity Resolution
- Leveraging LLMs for Privacy-Aware Predictions in Participatory Budgeting
- Does Multimodality Improve Recommender Systems as Expected? A Critical Analysis and Future Directions
- Resource-Limited Joint Multimodal Sentiment Reasoning and Classification via Chain-of-Thought Enhancement and Distillation
- A Survey on Video Temporal Grounding with Multimodal Large Language Model
- Tool Graph Retriever: Exploring Dependency Graph-based Tool Retrieval for Large Language Models
- Multimodal Fact Checking with Unified Visual, Textual, and Contextual Representations
- SPaRFT: Self-Paced Reinforcement Fine-Tuning for Large Language Models
- A Scalable Pretraining Framework for Link Prediction with Efficient Adaptation
- GraphProp: Training the Graph Foundation Models using Graph Properties
- UniTalker: Conversational Speech-Visual Synthesis
- Do Recommender Systems Really Leverage Multimodal Content? A Comprehensive Analysis on Multimodal Representations for Recommendation
- CALE : Concept-Aligned Embeddings for Both Within-Lemma and Inter-Lemma Sense Differentiation
- PersonaEval: Are LLM Evaluators Human Enough to Judge Role-Play?
- Dialogue Response Prefetching Based on Semantic Similarity and Prediction Confidence of Language Model
- What Do Humans Hear When Interacting? Experiments on Selective Listening for Evaluating ASR of Spoken Dialogue Systems
- SSEmb: A Joint Structural and Semantic Embedding Framework for Mathematical Formula Retrieval
- LLMDistill4Ads: Using Cross-Encoders to Distill from LLM Signals for Advertiser Keyphrase Recommendations
- CF-RAG: A Dataset and Method for Carbon Footprint QA Using Retrieval-Augmented Generation
- fact check AI at SemEval-2025 Task 7: Multilingual and Crosslingual Fact-checked Claim Retrieval
- Cropping outperforms dropout as an augmentation strategy for training self-supervised text embeddings
- Data Overdose? Time for a Quadruple Shot: Knowledge Graph Construction using Enhanced Triple Extraction
- ReDSM5: A Reddit Dataset for DSM-5 Depression Detection
- SmartLLMs Scheduler: A Framework for Cost-Effective LLMs Utilization
- Token-Level Precise Attack on RAG: Searching for the Best Alternatives to Mislead Generation
- PyLate: Flexible Training and Retrieval for Late Interaction Models
- LLM-based IR-system for Bank Supervisors
- Vision Language Model-based Testing of Industrial Autonomous Mobile Robots
- CAAD: Context-Aware Adaptive Decoding for Truthful Text Generation
- SpeechRole: A Large-Scale Dataset and Benchmark for Evaluating Speech Role-Playing Agents
- Improving Hospital Risk Prediction with Knowledge-Augmented Multimodal EHR Modeling
- MArgE: Meshing Argumentative Evidence from Multiple Large Language Models for Justifiable Claim Verification
- Harnessing Temporal Databases for Systematic Evaluation of Factual Time-Sensitive Question-Answering in Large Language Models
- Contextually Aware E-Commerce Product Question Answering using RAG
- LMAR: Language Model Augmented Retriever for Domain-specific Knowledge Indexing
- Intention-Guided Cognitive Reasoning for Egocentric Long-Term Action Anticipation
- Empowering Tabular Data Preparation with Language Models: Why and How?
- ChEmbed: Enhancing Chemical Literature Search Through Domain-Specific Text Embeddings
- Am I Blue or Is My Hobby Counting Teardrops? Expression Leakage in Large Language Models as a Symptom of Irrelevancy Disruption
- TCDiff: Triplex Cascaded Diffusion for High-fidelity Multimodal EHRs Generation with Incomplete Clinical Data
- Instruction-based Time Series Editing
- Balancing the Blend: An Experimental Analysis of Trade-offs in Hybrid Search
- Aligning Language Models with Real-time Knowledge Editing
- Towards Bridging Review Sparsity in Recommendation with Textual Edge Graph Representation
- Towards Efficient Medical Reasoning with Minimal Fine-Tuning Data
- Disaggregated Health Data in LLMs: Evaluating Data Equity in the Context of Asian American Representation
- Team "bettercallclaude": Style Change Detection using a Sequential Sentence Pair Classifier
- Experimental Evaluation of Dynamic Topic Modeling Algorithms
- Activation-Guided Local Editing for Jailbreaking Attacks
- The Prosody of Emojis
- The Missing Parts: Augmenting Fact Verification with Half-Truth Detection
- CyGATE: Game-Theoretic Cyber Attack-Defense Engine for Patch Strategy Optimization
- Cued-Agent: A Collaborative Multi-Agent System for Automatic Cued Speech Recognition
- Accurate and Consistent Graph Model Generation from Text with Large Language Models
- Improving Multimodal Contrastive Learning of Sentence Embeddings with Object-Phrase Alignment
- Automating AI Failure Tracking: Semantic Association of Reports in AI Incident Database
- From Static to Dynamic: A Streaming RAG Approach to Real-time Knowledge Base
- MECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks
- Role-Aware Language Models for Secure and Contextualized Access Control in Organizations
- Self-Foveate: Enhancing Diversity and Difficulty of Synthesized Instructions from Unsupervised Text via Multi-Level Foveation
- Causal2Vec: Improving Decoder-only LLMs as Versatile Embedding Models
- Text-to-SQL Task-oriented Dialogue Ontology Construction
- Accessibility Scout: Personalized Accessibility Scans of Built Environments
- Failures Are the Stepping Stones to Success: Enhancing Few-Shot In-Context Learning by Leveraging Negative Samples
- Real-time News Story Identification
- A Framework for Institutional Risk Identification using Knowledge Graphs and Automated News Profiling
- Resource-Efficient Adaptation of Large Language Models for Text Embeddings via Prompt Engineering and Contrastive Fine-tuning
- PersonaTwin: A Multi-Tier Prompt Conditioning Framework for Generating and Evaluating Personalized Digital Twins
- From personalized news curation to shared issue concerns in fragmentation era: a dynamic network approach by levels of issue involvement
- RASL: Retrieval Augmented Schema Linking for Massive Database Text-to-SQL
- IdeaBlocks: Expressing and Reusing Divergent Intents for Graphic Design Exploration using Generative AI
- Culinary Crossroads: A RAG Framework for Enhancing Diversity in Cross-Cultural Recipe Adaptation
Related