A Coefficient of Agreement for Nominal Scales
1960/04/01 by Jacob Cohen · 41,705 citations
Decision Sciences · Mathematics · #Mathematics #Multi-Criteria Decision Making #Reliability and Agreement in Measurement #Statistics
paper · doi:10.1177/001316446002000104
published in Educational and Psychological Measurement 20(1), 37-46 (SAGE Publishing)
openalex publication_date 1960/04/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05
Cited by
- A Ratio Test of Interrater Agreement With High Specificity
- Quantitative Image Analysis as an Adjunct to Manual Scoring of ER, PgR, and HER2 in Invasive Breast Carcinoma
- The Equivalence of Weighted Kappa and the Intraclass Correlation Coefficient as Measures of Reliability
- A Note on the Interpretation of Weighted Kappa and its Relations to Other Rater Agreement Statistics for Metric Scales
- Coefficient Kappa: Some Uses, Misuses, and Alternatives
- Acceptability and Efficacy of Group Behavioral Activation for Depression Among Adults: A Meta-Analysis
- The Measurement of Observer Agreement for Categorical Data
- A test of the interpersonal theory of suicide in a large, representative, retrospective and prospective study: Results from the Army Study to Assess Risk and Resilience in Servicemembers (Army STARRS)
- Visual Function Classification System for children with cerebral palsy: development and validation
- Prioritizing refuge sites for migratory geese to alleviate conflicts with agriculture
- ECG-based convolutional neural network in pediatric obstructive sleep apnea diagnosis
- Is It Good to Cooperate? Testing the Theory of Morality-as-Cooperation in 60 Societies
- An improved approach for predicting the distribution of rare and endangered species from occurrence and pseudo‐absence data
- Stability and predictors of somatic symptoms in men and women over 10 years: A real-world perspective from the prospective MONICA/KORA study
- Determining the Minimum Reliability Standard Based on a Decision Criterion
- Climate change drives range contraction and shapes species distribution in an alpine passerine: Caution required when comparing atlas data
- Back-Channel Representation: A Study of the Strategic Communication of Senators with the US Department of Labor
- Effect of active learning versus traditional lecturing on the learning achievement of college students in humanities and social sciences: a meta-analysis
- The effectiveness of an integratedSTEMcurriculum unit on middle school students' life science learning
- Population distributions of time to collision at brake application during car following from naturalistic driving data
- Assessing the accuracy of species distribution models: prevalence, kappa and the true skill statistic (TSS)
- Dimensions of Religion and Spirituality: A Longitudinal Topic Modeling Approach
- Fuzzy nearest neighbor algorithms: Taxonomy, experimental analysis and prospects
- Unilateral Posterior Crossbite is Not Associated with TMJ Clicking in Young Adolescents
- Who do they think you are? Inconsistencies in self- and proxy-reports of education within families
- Negotiator confidence: The impact of self-efficacy on tactics and outcomes
- Questionnaire-based diagnosis of benign paroxysmal positional vertigo
- A Systematic Survey on Image Description Techniques for STEM Domains
- The Effects of Spaced Practice on Second Language Learning: A Meta‐Analysis
- Evaluating migration hypotheses for the extinct Glyptotherium using ecological niche modeling
- How to map biomes: Quantitative comparison and review of biome‐mapping methods
- Tau-b or Not Tau-b: Measuring the Similarity of Foreign Policy Positions
- Chronic Obstructive Pulmonary Disease and Association With Mild Cognitive Impairment: The Mayo Clinic Study of Aging
- Interpretable EEG biomarkers with bag-of-waves: Spatial and temporal waveform dictionaries for low-data regimes
- Precision of Health-Related Quality-of-Life Data Compared With Other Clinical Measures
- EmoComicNet: A multi-task model for comic emotion recognition
- Green Governance: Boards of Directors’ Composition and Environmental Corporate Social Responsibility
- Scaffolding through prompts in digital learning: A systematic review and meta-analysis of effectiveness on learning achievement
- The Kappa Statistic in Reliability Studies: Use, Interpretation, and Sample Size Requirements
- Development and Validation of a Scale for Rating Motor Compensations Used for Reaching in Patients With Hemiparesis: The Reaching Performance Scale
- Fleiss’ kappa statistic without paradoxes
- Unfit for stranding assessment: a panel-scale multimodal-LLM audit of building-decarbonisation disclosure (BeDA)
- KaPilot: LLM-Assisted Generation of Kani Specifications for Unsafe Rust Verification
- A Latent Class Extension of Signal Detection Theory, with Applications
- ROOTCLUS: Searching for “ROOT CLUSters” in Three-Way Proximity Data
- Disentangling semantic and prosodic features of English poetry
- The Populist Style in American Politics: Presidential Campaign Discourse, 1952–1996
- Discordance between pain specialists and patients on the perception of dependence on pain medication: A multi-centre cross-sectional study
- Rape Myth Acceptance: Exploration of Its Structure and Its Measurement Using theIllinois Rape Myth Acceptance Scale
- Dynamic Capability Scoping for Enterprise AI Agents: A Synthetic Dataset and Three-Source Permission Architecture
- Identifying and Supporting Academically Low-Performing Schools in a Developing Country: An Application of a Specialized Multilevel IRT Model to PISA-D Assessment Data
- The Strengths and Difficulties Questionnaire (SDQ): the Factor Structure and Scale Validation in U.S. Adolescents
- How to Tell More is More: Quantity Discrimination in Eastern Box Turtles (Emydidae: Terrapene carolina)
- Identity Formation in Early and Middle Adolescents From Various Ethnic Groups: From Three Dimensions to Five Statuses
- The use of the Vocal Profile Analysis for speaker characterization: Methodological proposals
- Inter-Rater Reliability Methods in Qualitative Case Study Research
- Exploration, Explanation, and Parent–Child Interaction in Museums
- Evolutionary Personality Psychology
- Meta-analysis of action video game impact on perceptual, attentional, and cognitive skills.
- A meta-analytic review of experiments examining the effects of extrinsic rewards on intrinsic motivation.
- Data Annotations as Pedagogical Hints: From Subjective Labels to Critical Thinking
- MOF-Sleuth: Tool-Grounded Reward Alignment for Explainable Fine-Grained MOF CIF Auditing
- Active Listening in Integrative Negotiation
- Towards an Automated Test of LLM Security Knowledge
- How Far Can Wearable-Compatible Signals Go? A Controlled Decomposition of Non-EEG Sleep Staging
- Large Language Models for Citation Function Classification
- In-Context Learning for Wound Classification with Small Multimodal Language Models
- Evaluating medical AI under missing information: same-provider judges and human raters change apparent safety
- Visual Semantic Decoding of Electrocorticography from Video Stimuli using End-to-End Deep Learning
- SciCodePile: A 128GB Corpus and Executable Benchmark for Challenging Scientific Code Generation
- Real-World Evaluation of an AI Agent Drafting Translational Impact Summaries
- Espoonlahti mobile laser scanning tree species classification
- BanClickThumb: A Multimodal Dataset and Transformer Fusion Benchmarks for Clickbait Detection in Bengali YouTube Videos
- Safety That Does Not Transfer: Cross-Lingual Clinical Correctness Drift in Deployable Medical Language Models
- CARV: A Diagnostic Benchmark for Compositional Analogical Reasoning in Multimodal LLMs
- PriorProof: A Point-in-Time Measure of Technique Novelty for Formal Proofs
- One Modality to Forget Them All: Enhancing Cross-Modal Unlearning in Vision-Language Models
- Evaluating Open-Weight LLMs for Generating Structured Threat Information for Autonomous Vehicle Vulnerabilities
- Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning
- From Stateless to Situated: Building a Psychological World for LLM-Based Agents
- Large Language Models for Code Generation from Multilingual Prompts: A Curated Benchmark and a Study on Code Quality
- Students' Perceptions of Peer Grading
- Identity statuses and psychosocial functioning in Turkish youth: A person‐centered approach
- SWE-chat: Coding Agent Interactions From Real Users in the Wild
- Measuring the State of Open Science in Transportation Using Large Language Models
- Evaluating RAG for French immigration law: a benchmark and baseline study
- Can We Hide Machines in the Crowd? Quantifying Equivalence in LLM-in-the-loop Annotation Tasks
- Training language models to be warm and empathetic makes them less reliable and more sycophantic
- Base Models Beat Aligned Models at Randomness and Creativity
- Large language models for scientometric mapping of scientific controversy: A validated hybrid AI–Human framework
- Using Natural Language Processing to Automatically Detect Self-Admitted Technical Debt
- Empirical Evaluation of the Impact of Object-Oriented Code Refactoring on Quality Attributes: A Systematic Literature Review
- The Secret Life of Software Vulnerabilities: A Large-Scale Empirical Study
- Computing inter‐rater reliability and its variance in the presence of high agreement
- Looking at the dark and bright sides of identity formation: New insights from adolescents and emerging adults in Japan
- A Systematic Review of Collective Tactical Behaviours in Football Using Positional Data
- A multi‐dimensional measure of vocational identity status
- Factors Relating to Sprint Swimming Performance: A Systematic Review
- Language models and Automated Essay Scoring
- Using Fisher's Exact Test to Evaluate Association Measures for N-grams
- Remotely mapping gullying and incision in Maryland Piedmont headwater streams using repeat airborne lidar
- Acoustic characterization and classification of rockfish and walleye pollock in the Gulf of Alaska at bottom trawl locations
- Style Amnesia: Investigating Speaking Style Degradation and Mitigation in Multi-Turn Spoken Language Models
- Semicircular canal morphology in Rodentia and its relationship to locomotion
- Beyond the Tip of the Iceberg: Assessing Coherence of Text Classifiers
- Deep Contextualized Biomedical Abbreviation Expansion
- Can we really reduce ethnic prejudice outside the lab? A meta‐analysis of direct and indirect contact interventions
- Smart Contract Security: a Practitioners' Perspective
- Family Firms, M&A Strategies, and M&A Performance: A Meta-Analysis
- Grading exams using large language models: A comparison between human and AI grading of exams in higher education using ChatGPT
- A systematic review of economic evidence of artificial intelligence in healthcare
- Evaluation of performance measures in predictive artificial intelligence models to support medical decisions: overview and guidance
- Digital transformation: A multidisciplinary reflection and research agenda
- DICE: Discrete Interpretable Comparative Evaluation with Probabilistic Scoring for Retrieval-Augmented Generation
- Group-wise Contrastive Learning for Neural Dialogue Generation
- Are open educational resources (OER) and practices (OEP) effective in improving learning achievement? A meta-analysis and research synthesis
- Impact of environmental barriers on temnospondyl biogeography and dispersal during the Middle–Late Triassic
- ATCNet-CIAM for Multi-Session Motor Imagery EEG Signal Classification
- Global tracking of marine megafauna space use reveals how to achieve conservation targets
- Dyadic differences in empathy scores are associated with kinematic similarity during conversational question–answer pairs
- Multi-label Categorization of Accounts of Sexism using a Neural Framework
- Dynamic Context Selection for Document-level Neural Machine Translation via Reinforcement Learning
- Calibrated Tree-Neural Fusion for Fine-Grained Vegetation Community Classification
- Do Current Retrievers Cover All the Evidence? A Controlled Study of Conjunctive Cross-Page Retrieval
- Consumer Perception of Cows' Milk and Plant‐Based Milk Alternatives: Comparing Aotearoa–New Zealand and Singapore Consumers
- Leveraging Semantic Maps for City-Scale Cross-View Localization
- Sense it with your eyes: Sensation Generation and Understanding for Advertisements
- Beyond Exact Match: How Evaluation Methodology Dominates Model Choice in LLM-Based Product Attribute Extraction
- Dialectical Contradictions in Relationship Development
- Beyond resilients, undercontrollers, and overcontrollers? an extension of personality prototype research
- TrajAudit: Automated Failure Diagnosis for Agentic Coding Systems
- Refining Implicit Argument Annotation for UCCA
- Translation Quality Assessment: A Brief Survey on Manual and Automatic Methods
- Automated Modernization of Machine Learning Engineering Notebooks for Reproducibility
- Sleep Stage Classification Using Bidirectional LSTM in Wearable Multi-sensor Systems
- Co‐occurrence of depression and delinquency in personality types
- ClarifyMT-Bench: Benchmarking and Improving Multi-Turn Clarification for Conversational Large Language Models
- Assessing the Software Security Comprehension of Large Language Models
- Improving ML Training Data with Gold-Standard Quality Metrics
- Competing or Collaborating? The Role of Hackathon Formats in Shaping Team Dynamics and Project Choices
- A Large-Language-Model Framework for Automated Humanitarian Situation Reporting
- DramaBench: A Six-Dimensional Evaluation Framework for Drama Script Continuation
- Understanding Typing-Related Bugs in Solidity Compiler
- When the Gold Standard Isn't Necessarily Standard: Challenges of Evaluating the Translation of User-Generated Content
- Are Vision Language Models Cross-Cultural Theory of Mind Reasoners?
- Towards Deeper Emotional Reflection: Crafting Affective Image Filters with Generative Priors
- OLAF: Towards Robust LLM-Based Annotation Framework in Empirical Software Engineering
- Governance by Evidence: Regulated Predictors in Decision-Tree Models
- Towards Proactive Personalization through Profile Customization for Individual Users in Dialogues
- Thematic Dispersion in Arabic Applied Linguistics: A Bibliometric Analysis using Brookes' Measure
- APT-ClaritySet: A Large-Scale, High-Fidelity Labeled Dataset for APT Malware with Alias Normalization and Graph-Based Deduplication
- Emotion Recognition in Signers
- Vibe Spaces for Creatively Connecting and Expressing Visual Concepts
- LexRel: Benchmarking Legal Relation Extraction for Chinese Civil Cases
- Unveiling Malicious Logic: Towards a Statement-Level Taxonomy and Dataset for Securing Python Packages
- NagaNLP: Bootstrapping NLP for Low-Resource Nagamese Creole with Human-in-the-Loop Synthetic Data
- Journey Before Destination: On the importance of Visual Faithfulness in Slow Thinking
- The Effect of Document Summarization on LLM-Based Relevance Judgments
- Evaluating the Efficacy of Sentinel-2 versus Aerial Imagery in Serrated Tussock Classification
- Decoding Human-LLM Collaboration in Coding: An Empirical Study of Multi-Turn Conversations in the Wild
- Generate-Then-Validate: A Novel Question Generation Approach Using Small Language Models
- Towards Practical and Usable In-network Classification
- Evaluation of Text Generation: A Survey
- Machine learning for smell: Ordinal odor strength prediction of molecular perfumery components
- Towards a Science of Scaling Agent Systems
- Human– AI collaborative learning in mixed reality: Examining the cognitive and socio‐emotional interactions
- Balanced Accuracy: The Right Metric for Evaluating LLM Judges -- Explained through Youden's J statistic
- LocalSearchBench: Benchmarking Agentic Search in Real-World Local Life Services
- Desktop versus VR for collaborative sensemaking
- CAuSE: Decoding Multimodal Classifiers using Faithful Natural Language Explanation
- Configuration Defects in Kubernetes
- When Do Domain-Specific Foundation Models Justify Their Cost? A Systematic Evaluation Across Retinal Imaging Tasks
- Can gestures speak louder than words? The effect of gestural discourse markers on discourse expectations
- "Dragon Slayer Becomes the Dragon": How Players Perceive and Respond to Inequality in the Game World of Whiteout Survival
- StageGuard: Physiologically Constrained Sleep Staging
- KH-FUNSD: A Hierarchical and Fine-Grained Layout Analysis Dataset for Low-Resource Khmer Business Document
- Learn like a Pathologist: Curriculum Learning by Annotator Agreement for Histopathology Image Classification
- Executable Governance for AI: Translating Policies into Rules Using LLMs
- Systematic Review and Meta-Analysis: Adolescent Depression and Long-Term Psychosocial Outcomes
- Evaluating the Robustness of Large Language Model Safety Guardrails Against Adversarial Attacks
- LAURA: Enhancing Code Review Generation with Context-Enriched Retrieval-Augmented LLM
- CodeDistiller: Automatically Generating Code Libraries for Scientific Coding Agents
- CentaurEval: Benchmarking Human-in-the-Loop Value in Agentic Coding
- CourseTimeQA: A Lecture-Video Benchmark and a Latency-Constrained Cross-Modal Fusion Method for Timestamped QA
- Assessing method agreement for paired repeated binary measurements administered by multiple raters
- A standard protocol for reporting species distribution models
- From Words to Wisdom: Discourse Annotation and Baseline Models for Student Dialogue Understanding
- MTBBench: A Multimodal Sequential Clinical Decision-Making Benchmark in Oncology
- Large Language Models' Complicit Responses to Illicit Instructions across Socio-Legal Contexts
- Language-Independent Sentiment Labelling with Distant Supervision: A Case Study for English, Sepedi and Setswana
- Cross-replication Reliability -- An Empirical Approach to Interpreting Inter-rater Reliability
- Interobserver Agreement Among Sleep Scorers From Different Centers in a Large Dataset
- MUCH: A Multilingual Claim Hallucination Benchmark
- A Diversity-optimized Deep Ensemble Approach for Accurate Plant Leaf Disease Detection
- FinCriticalED: A Visual Benchmark for Financial Fact-Level OCR Evaluation
- Leveraging Digitized Newspapers to Collect Summarization Data in Low-Resource Languages
- A systematic review and meta-analysis of factors associated with anthelmintic resistance in sheep
- Function of language skills in preschooler's problem-solving performance: The role of self-directed speech
- Spiking Neural Networks for Early Prediction in Human Robot Collaboration
- Leveraging Medical Sentiment to Understand Patients Health on Social Media
- Towards Consistent Detection of Cognitive Distortions: LLM-Based Annotation and Dataset-Agnostic Evaluation
- Evaluating consistency of deterministic streamline tractography in\n non-linearly warped DTI data
- An empirical study of Policy-as-Code adoption in open-source software projects
- Exploring the stigma experienced by people affected by Parkinson’s disease: a systematic review
- Exception handling bugs in Python: An empirical study of root causes, fix patterns, and anti-patterns
- Drawing on Education: Using Drawings to Document Schooling and Support Change
- Analyzing Dataset Annotation Quality Management in the Wild
- PRISM of Opinions: A Persona-Reasoned Multimodal Framework for User-centric Conversational Stance Detection
- Towards a Human-in-the-Loop Framework for Reliable Patch Evaluation Using an LLM-as-a-Judge
- Perceive, Act and Correct: Confidence Is Not Enough for Hyperspectral Classification
- AI Annotation Orchestration: Evaluating LLM verifiers to Improve the Quality of LLM Annotations in Learning Analytics
- Exploringand Unleashing the Power of Large Language Models in CI/CD Configuration Translation
- On User Interfaces for Large-Scale Document-Level Human Evaluation of Machine Translation Outputs
- Predicting student outcomes using digital logs of learning behaviors: Review, current standards, and suggestions for future work
- DiagramIR: An Automatic Pipeline for Educational Math Diagram Evaluation
- Concrete images, diverse ideas: The role of pictures in learning from multimedia texts of varying abstractness
- Predictive mapping of forest composition and structure with direct gradient analysis and nearest- neighbor imputation in coastal Oregon, U.S.A.
- Evaluating Language Model Applications for Identifying Solution-Related Content in Issue Report Discussions
- Predicting Length of Stay in the Intensive Care Unit with Temporal Pointwise Convolutional Networks
- Who Is the Story About? Protagonist Entity Recognition in News
- Groundwater potential mapping using C5.0, random forest, and multivariate adaptive regression spline models in GIS
- Testing the Testers: Human-Driven Quality Assessment of Voice AI Testing Platforms
- Detecting Silent Failures in Multi-Agentic AI Trajectories
- Language-Enhanced Generative Modeling for Amyloid PET Synthesis from MRI and Blood Biomarkers
- From Pre-labeling to Production: Engineering Lessons from a Machine Learning Pipeline in the Public Sector
- Sustainability of Machine Learning-Enabled Systems: The Machine Learning Practitioner's Perspective
- The Eigenvalues Entropy as a Classifier Evaluation Measure
- Rating Roulette: Self-Inconsistency in LLM-As-A-Judge Frameworks
- TheraMind: A Strategic and Adaptive Agent for Longitudinal Psychological Counseling
- Developing and Validating a Diagnostic Checklist to Assess Argumentation in EFL Writing
- DecoEvo: Score-Decoupled Co-Evolution of Solver and Rubric-Generator Skills in Text Space
- Luxury value perceptions and consumer outcomes: A meta‐analysis
- Multimodal information density is highest in question beginnings, and early entropy is associated with fewer but longer visual signals
- Rashomon Alignment
- The Audit Committee Oversight Process*
- Opinion Mining in Online Reviews About Distance Education Programs
- Target Based Speech Act Classification in Political Campaign Text
- Turn-taking cues in task-oriented dialogue
- RICA: Evaluating Robust Inference Capabilities Based on Commonsense Axioms
- Comparison between methods of vascular calibre characterization and measurement protocols: Influence of vessels number considered
- Modeling water table trends in high-latitude peatlands: Divergent hydrological responses and fire risk implications
- Information Leakage and Performance Overestimation in EEG-Based Schizophrenia Detection: Evidence from Literature and Empirical Analyses
- Mustelid Herpesvirus-2, a Novel Herpes Infection in Northern Sea Otters (Enhydra Lutris Kenyoni)
- Interaction coding in leadership research: A critical review and best-practice recommendations to measure behavior
- Model family selection for classification using Neural Decision Trees
- Chance‐Corrected Interrater Agreement Statistics for Two‐Rater Dichotomous Responses: A Method Review With Comparative Assessment Under Possibly Correlated Decisions
- Handwriting in primary school: comparing standardized tests and evaluating impact of grapho-motor parameters
- Hysteresis in streamflow‐water table relation provides a new classification system of rainfall‐runoff events
- StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents
- GuidedRAG: Semantic Steering of Retrieval-Augmented Generation
- High agreement but low Kappa: I. the problems of two paradoxes
- Assessment of interrater and intrarater reliability of the Fahn–Tolosa–Marin Tremor Rating Scale in essential tremor
- Antibiotic Exposure and Risk of Parkinson's Disease in Finland: A Nationwide Case‐Control Study
- An Intersectional Analysis of Agentic Efforts Individuals Under Community Supervision Describe to Improve Their Lives
- Interobserver agreement in describing adnexal masses using the International Ovarian Tumor Analysis simple rules in a real-time setting and using three-dimensional ultrasound volumes and digital clips
- Intra- and interobserver agreement with regard to describing adnexal masses using International Ovarian Tumor Analysis terminology: reproducibility study involving seven observers
- How Do Therapists Experience 20‐Session Cognitive‐Behavioral Therapy for Anorexia Nervosa ( CBT ‐ AN ‐20)? A Qualitative Study
- Deep attentive spatio-temporal feature learning for automatic resting-state fMRI denoising
- The paradox of paradoxical leadership: A multi-level conceptualization
- Do risk assessment tools help manage and reduce risk of violence and reoffending? A systematic review.
- Muscle size and composition in people with articular hip pathology: a systematic review with meta-analysis
- Textual Indicators of Deliberative Dialogue: A Systematic Review of Methods for Studying the Quality of Online Dialogues
- Does Media Coverage of Partisan Polarization Affect Political Attitudes?
- The Forgotten Margins of AI Ethics
- Understanding partition comparison indices based on counting object pairs
- The motion of trees in the wind: a data synthesis
- Automated identification of hedgerows and hedgerow gaps using deep learning
- Overcoming Climate Gridlock: Perspectives of Climate Leaders on How to Achieve Social Change During Persistent Failure in Australia
- Emotion Ratings: How Intensity, Annotation Confidence and Agreements are Entangled
- Crowdfunding Success Factors: A Meta-Analytic Investigation
- Multi-scale habitat selection modeling: a review and outlook
- Short text classification with machine learning in the social sciences: The case of climate change on Twitter
- Hybrid CNN XGBoost intrusion detection approach tuned by modified sine cosine algorithm towards better cloud security
- Classifying types of gully changes with unoccupied aircraft vehicles 3D multitemporal point clouds for training of satellite data analysis in Northwest Namibia
- The Relation of Ambulatory Blood Pressure and Pulse Rate to Retinopathy in Type 1 Diabetes Mellitus
- WEC: Deriving a Large-scale Cross-document Event Coreference dataset from Wikipedia
- Ripple Effects: Social Turmoil Following Infant Kidnapping Attempts in Wild Geladas
- Ecotrends: an R package for estimating habitat suitability trends over time
- Identifying psychological distress data available in nationally representative surveys: A scoping review and case study of Australian surveys
- Effect of potentially modifiable risk factors associated with myocardial infarction in 52 countries (the INTERHEART study): case-control study
- Deep problems with neural network models of human vision
- Factors Associated with Teacher Wellbeing: A Meta-Analysis
- Predictive habitat distribution models in ecology
- Battle for Inbox and Bucks
- European Blame Games
- Classification of tree species and standing dead trees in Boreal forests using UAV‐based RGB, multispectral, and LiDAR point clouds
- When Knowledge Changes: Metamorphic Testing of RAG Systems with Mutations
- Instructional Coaching as a Tool for Professional Development: Coaches’ Roles and Considerations
- Intercoder Reliability in Qualitative Research: Debates and Practical Guidelines
- A Tertiary lymphoid structures-based pathological score predicts survival and recurrence in colorectal Cancer patients
- Couples and Reproductive Health: A Review of Couple Studies
- Does Google Scholar contain all highly cited documents (1950-2013)?
- Symbol Emergence as an Interpersonal Multimodal Categorization
- Land-cover change in the Kruger to Canyons Biosphere Reserve (1993– 2006): A first step towards creating a conservation plan for the subregion
- An Empirical Study on Deployment Faults of Deep Learning Based Mobile Applications
- Semantic Change Detection with Asymmetric Siamese Networks
- Min-Mid-Max Scaling, Limits of Agreement, and Agreement Score
- The role of neuroticism and extraversion in the stress–anxiety and stress–depression relationships
- Nomogram for sample size calculation on a straightforward basis for the kappa statistic
- Team teaching or solo teaching? Evidence from a crossover experiment on the effects of team teaching on student achievement
- Stimulating language awareness in the foreign language classroom: exploring EFL teaching practices
- Dynamics of affective states during complex learning
- Automated sleep stage identification system based on time–frequency analysis of a single EEG channel and random forest classifier
- Personal and Social Facets of Job Identity: A Person-Centered Approach
- Paraspinal muscle quality in chronic low back pain: a systematic review and meta-analysis of muscle atrophy and fat infiltration
- Software Development During COVID-19 Pandemic: an Analysis of Stack Overflow and GitHub
- Biases in the Blind Spot: Detecting What LLMs Fail to Mention
- 5-Hydroxymethylcytosine signatures in cell-free DNA provide information about tumor types and stages
- Advanced statistics: Understanding Medical Record Review (MRR) Studies
- Self-regulation in young children: Is there a role for sociodramatic play?
- Effects of maternal mentalization-related parenting on toddlers’ self-regulation
- Adaptive Data Collection for Latin-American Community-sourced Evaluation of Stereotypes (LACES)
- Global Chlorophyll-a Retrieval algorithm from Sentinel 2 Using Residual Deep Learning and Novel Machine Learning Water Classification
- Towards AI as Colleagues: Multi-Agent System Improves Structured Professional Ideation
- When Good Fit Goes Bad: Identifying and Minimising Overfitting in Ecological Niche Models
- Promoting healthy eating through nudges: Multimodal discursive representations of food and eating in weight‑loss articles in Chinese official WeChat posts
- The Paris 1976 Wine Tastings Revisited Once More: Comparing Ratings of Consistent and Inconsistent Tasters
- Analysis of hotspot areas in China's satellite internet innovation policies and research on policy evolution
- The proteome of the late Middle Pleistocene Harbin individual
- Code Contribution and Credit in Science
- Modeling Hierarchical Thinking in Large Reasoning Models
- A First Look at the Self-Admitted Technical Debt in Test Code: Taxonomy and Detection
- Measure what Matters: Psychometric Evaluation of AI with Situational Judgment Tests
- On making causal claims: A review and recommendations
- Leader humility in Singapore
- ArchISMiner: A Framework for Automatic Mining of Architectural Issue-Solution Pairs from Online Developer Communities
- A meta-analytic review of authentic and transformational leadership: A test for redundancy
- Implicit and explicit processes in phonological concept learning
- “Safer to plant corn and beans”? Navigating the challenges and opportunities of agricultural diversification in the U.S. Corn Belt
- Learning to Triage Taint Flows Reported by Dynamic Program Analysis in Node.js Packages
- RatioWaveNet: A Learnable RDWT Front-End for Robust and Interpretable EEG Motor-Imagery Classification
- Beyond the Surface: Sharenting as a Source of Family Quandaries: Mapping Parents’ Social Media Dilemmas
- The advantages of the Matthews correlation coefficient (MCC) over F1 score and accuracy in binary classification evaluation
- Resting-State Functional MRI in Dyslexia: A Systematic Review
- Patterns of Technical Variation in Chimpanzee Termite Fishing Behavior in Mbam and Djerem National Park, Cameroon
- Visual representations of energy and chemical bonding in biology and chemistry textbooks: A case study of ATP hydrolysis
- Spike detection in the wild: Screening of suspected temporal lobe epilepsy cases using a tailored 2‐channel wearable EEG
- The social studies discourse instrument: Validating an observation tool for classroom discussions
- LexChain: Modeling Legal Reasoning Chains for Chinese Tort Case Analysis
- Enhancing reliability in AI inference services: An empirical study on real production incidents
- Iterative Topic Taxonomy Induction with LLMs: A Case Study of Electoral Advertising
- Speculative Model Risk in Healthcare AI: Using Storytelling to Surface Unintended Harms
- Low-income Latino mothers’ booksharing styles and children's emergent literacy development
- GAPS: A Clinically Grounded, Automated Benchmark for Evaluating AI Clinicians
- Development and Benchmarking of a Blended Human-AI Qualitative Research Assistant
- The Child–Adult Medical Procedure Interaction Scale-Short Form (CAMPIS-SF)
- Getting more out of binary data. Segmenting markets by bagged clustering.
- Lingxi: Repository-Level Issue Resolution Framework Enhanced by Procedural Knowledge Guided Scaling
- Testing and Enhancing Multi-Agent Systems for Robust Code Generation
- A non-invasive machine learning mechanism for early disease recognition on Twitter: The case of anemia
- From literature to biodiversity data: mining arthropod organismal traits with machine learning
- Masculine Republicans and Feminine Democrats: Gender and Americans’ Explicit and Implicit Images of the Political Parties
- Brick: Spatial Capability Routing for the Mixture-of-Models (MoM) Paradigm
- PreprintToPaper dataset: connecting bioRxiv preprints with journal publications
- Deconstruct to Reconstruct a Configurable Evaluation Metric for Open-Domain Dialogue Systems
- Mapping riparian forest fragmentation along the Iori River in Georgia
- Dominance Style Among Macaca thibetana on Mt. Huangshan, China
- Constraint-Guided Unit Test Generation for Machine Learning Libraries
- Investigating the Impact of Rational Dilated Wavelet Transform on Motor Imagery EEG Decoding with Deep Learning Models
- DITING: A Multi-Agent Evaluation Framework for Benchmarking Web Novel Translation
- MigrateLib: a tool for end-to-end Python library migration
- Emotionally Charged, Logically Blurred: AI-driven Emotional Framing Impairs Human Fallacy Detection
- Linguistic Resources for Bhojpuri, Magahi and Maithili: Statistics about them, their Similarity Estimates, and Baselines for Three Applications
- Beyond Monolingual Assumptions: A Survey of Code-Switched NLP in the Era of Large Language Models
- The virtual census: representations of gender, race and age in video games
- Review: Adolescents' perspectives on and experiences with post‐primary school‐based suicide prevention as end‐users, co‐creators and peer helpers – a systematic review meta‐ethnography
- Aligning EU policies to address biological invasions: assessing invasion impacts across sectors
- Refactoring with LLMs: Bridging Human Expertise and Machine Understanding
- Person-Centric Annotations of LAION-400M: Auditing Bias and Its Transfer to Models
- Can Emulating Semantic Translation Help LLMs with Code Translation? A Study Based on Pseudocode
- Curiosity-Driven LLM-as-a-judge for Personalized Creative Judgment
- A Note on the Use of Categorical Subscores
- VitaBench: Benchmarking LLM Agents with Versatile Interactive Tasks in Real-world Applications
- An Annotation Scheme for Factuality and its Application to Parliamentary Proceedings
- Feedback Forensics: A Toolkit to Measure AI Personality
- Discovering Self-Regulated Learning Patterns in Chatbot-Powered Education Environment
- Generative Value Conflicts Reveal LLM Priorities
- Towards Reliable Generation of Executable Workflows by Foundation Models
- Identifying deep leverage points to destabilize ‘lock-in’ and empower farmers in the Midwestern agrifood system
- SimuHome: A Temporal- and Environment-Aware Benchmark for Smart Home LLM Agents
- A bot identification model and tool based on GitHub activity sequences
- Beyond kappa: A review of interrater agreement measures
- Moving towards Europe-wide freshwater restoration through model-based integration of policy objectives
- Revisiting Maturity Data: Using Oocyte Diameter and Gonadosomatic Index to Retroactively Apply a New Maturity Scale to Greenland Halibut ( Reinhardtius hippoglossoides )
- Deteção de estruturas permanentes a partir de dados de séries temporais Sentinel 1 e 2
- Building work engagement: A systematic review and meta‐analysis investigating the effectiveness of work engagement interventions
- MASH: A Multiplatform and Multimodal Annotated Dataset for Societal Impact of Hurricane
- Open-DeBias: Toward Mitigating Open-Set Bias in Language Models
- What does it mean to be European? How identity content shapes adolescent's views towards immigrants and support for the EU
- Liaozhai through the Looking-Glass: On Paratextual Explicitation of Culture-Bound Terms in Machine Translation
- Bridging science and society: Developing a citizen science biomonitoring approach for river ecosystems in Italy
- Identification of a 10-species microbial signature of inflammatory bowel disease by machine learning and external validation
- Examining the Influence of Different Inventories on Shallow Landslide Susceptibility Modeling: An Assessment Using Machine Learning and Statistical Approaches
- Comparison between Dense L-Band and C-Band Synthetic Aperture Radar (SAR) Time Series for Crop Area Mapping over a NISAR Calibration-Validation Site
- Could you teach new tricks to old dogs? An analysis of online communication of political leaders in Spanish regional elections
- Do 14–17-month-old infants use iconic speech and gesture cues to interpret word meanings?
- LLMs Behind the Scenes: Enabling Narrative Scene Illustration
- Consider the following: A pilot study of the effects of an educational television program on viewer perceptions of anthropogenic climate change and ocean acidification
- Environmental constraints shaping constituent order in emerging communication systems: Structural iconicity, interactive alignment and conventionalization
- A Machine Learning Framework for Predicting and Understanding the Canadian Drought Monitor
- Personality and prosocial behavior: A theoretical framework and meta-analysis.
- The double‐edged sword of CEO narcissism: A meta‐analysis of innovation and firm performance implications
- Addressing Failing Water Systems: A Qualitative Protocol for Intervention Research on Water Insecurity
- Different gazes and hands: Visual representations of causes and solutions of adolescent depression on Chinese state media social platforms
- Building a Pilot Software Quality-in-Use Benchmark Dataset
- TrueGradeAI: Retrieval-Augmented and Bias-Resistant AI for Transparent and Explainable Digital Assessments
- Do LLM Agents Know How to Ground, Recover, and Assess? A Benchmark for Epistemic Competence in Information-Seeking Agents
- ASSESS: A Semantic and Structural Evaluation Framework for Statement Similarity
- ReviewScore: Misinformed Peer Review Detection with Large Language Models
- Acoustic-based Gender Differentiation in Speech-aware Language Models
- Human-AI Narrative Synthesis to Foster Shared Understanding in Civic Decision-Making
- AutoSpec: An Agentic Framework for Automatically Drafting Patent Specification
- Reverse Engineering User Stories from Code using Large Language Models
- Beyond Similarity: Grounded Agentic Extraction and Expert-Adjudicated Evaluation of Intertextuality in Classical Chinese Histories
- Hallucination‐Free? Assessing the Reliability of Leading AI Legal Research Tools
- A Closer Look at Classification Evaluation Metrics and a Critical Reflection of Common Evaluation Practice
- Gender Discrimination in the Visual Representation of Athletes on the Official Instagram Accounts of Sports Federations in Indonesia
- Land Use and Land Cover Change in the Shafarood Watershed, Northern Iran (2000–2020): A Case Study Using Landsat Imagery and Support Vector Machine Classification
- A systematic review of echo chamber research: comparative analysis of conceptualizations, operationalizations, and varying outcomes
- Challenges in annotations by humans and LLMs: A case study of evaluative language
- Associations of Physician Empathy with Patient Anxiety and Ratings of Communication in Hospital Admission Encounters
- Open Security Benchmark: Towards Autonomous Enterprise Cyber Defense
- Consonant articulation accuracy in paediatric cochlear implant recipients
- GeDi: Simplifying Gene Set Distances for Enhanced Omics Interpretation in R/Bioconductor
- Large language models as first-pass filters for corpus annotation: semantic disambiguation of Galician pobo
- Continued use of retracted papers: Temporal trends in citations and (lack of) awareness of retractions shown in citation contexts in biomedicine
- Subjective assessment of adnexal masses with the use of ultrasonography: an analysis of interobserver variability and experience
- From pixels to peaks: integrating LiDAR and RGB drone imagery to map mussel spat on intertidal rocky shores
- The Chrysalis Effect
- Impact of schematic representations of road maps on pedestrian route planning
- Sentinel-2 satellite image application for establishing a coral reef distribution map in Bai Tu Long National Park, Quang Ninh Province, Vietnam
- Measuring and comparing the accuracy of species distribution models with presence–absence data
- How Sexual Consent is Portrayed in Sex Comics (Eromanga): A Content Analysis in Japan
- Assessing the Cross-Version Applicability of Java Library Vulnerability Exploits
- Comparing Clinical, Microbiological, and Genetic Definitions of Relapse in Patients With Subsequent Episodes of Methicillin-resistant Staphylococcus aureus Bone and Joint Infection
- Impact of interaction with an architectural example on design behavior in student teams
- Ictal emotional features in pediatric and young adult patients with frontal lobe epilepsy
- A custom‐built single‐channel in‐ear electroencephalography sensor for sleep phase detection: an interdependent solution for at‐home sleep studies
- The impact and return-on-investment of evidence-based practice in conservation and environmental management: A machine learning-assisted scoping review protocol
- Annotation of biological samples data to standard ontologies with support from large language models
- A systematic review and meta-analysis of interventions that target the intersection of body image and movement among girls and women
- The value of small forest fragments and urban tree canopy for Neotropical migrant birds during winter and migration seasons in Latin American countries: A systematic review
- Evaluating ChatGPT, Gemini and other Large Language Models (LLMs) in orthopaedic diagnostics: A prospective clinical study
- Integrating climate indices and land use practices for comprehensive drought monitoring in Syria: Impacts and implications
- Reusability of Bayesian Networks case studies: a survey
- The impact of rapport on intelligence yield: police source handler telephone interactions with covert human intelligence sources
- Examining item content across nine psychological (in)flexibility scales: What do they measure?
- The effect of reward value on the performance of long-tailed macaques (Macaca fascicularis) in a delay of gratification exchange task
- Demographically-Inspired Query Variants Using an LLM
- Toward Understanding Free‐Flowing Manual Object Contact: Real‐Time Interaction Between Body Position and Object Type in 9‐Month‐Olds
- Machine learning-based ensemble species distribution models to guide monitoring and survey design for offshore wind
- Evidence on the effects of flame retardant substances at ecologically relevant endpoints: a systematic map protocol
- Clustering longitudinal data: comparison of model-based and distance-based approaches using simulated and real-world data in psychiatric research
- Machine learning predictive modelling for sediment risk indices within an urbanized river channel
- Evaluation of Remotely Sensed Inundation Data Sets to Estimate Flood‐Associated Emergency Department Visits After Hurricane Harvey
- Identifying suitable habitats under climate change for non-targeted demersal fish in the Mediterranean Sea
- Doctoral Education Trends: Content Analyses of Dissertations and Job Postings
- Uncertainty-driven ensembles of deep architectures for multiclass classification. Application to COVID-19 diagnosis in chest X-ray images
- Automatic sampling and training method for wood-leaf classification based on tree terrestrial point cloud
- Semixup: In- and Out-of-Manifold Regularization for Deep Semi-Supervised Knee Osteoarthritis Severity Grading from Plain Radiographs
- Telephone versus In-Person Clinical and Health Status Assessment Interviews in Patients with Bipolar Disorder
- Longitudinal Analysis of Discussion Topics in an Online Breast Cancer Community using Convolutional Neural Networks
- Advanced, Analytic, Automated (AAA) Measurement of Engagement During Learning
- MedFact: A Large-scale Chinese Dataset for Evidence-based Medical Fact-checking of LLM Responses
- Robustness and Reliability of Gender Bias Assessment in Word Embeddings: The Role of Base Pairs
- Psychometric properties of the EQ-5D-5L: a systematic review of the literature
- Mental Multi-class Classification on Social Media: Benchmarking Transformer Architectures against LSTM Models
- SciEvent: Benchmarking Multi-domain Scientific Event Extraction
- RulER: Automated Rule-Based Semantic Error Localization and Repair for Code Translation
- Position: Thematic Analysis of Unstructured Clinical Transcripts with Large Language Models
- A Study on Thinking Patterns of Large Reasoning Models in Code Generation
- Development of a Structured Psychiatric Interview for Children: Agreement Between Child and Parent on Individual Symptoms
- Annotating Training Data for Conditional Semantic Textual Similarity Measurement using Large Language Models
- Crash Report Enhancement with Large Language Models: An Empirical Study
- When Large Language Models Meet UAV Projects: An Empirical Study from Developers' Perspective
- Accelerating Discovery: Rapid Literature Screening with LLMs
- Determinants of Entrepreneurial Intent: A Meta–Analytic Test and Integration of Competing Models
- A New Statistical Approach for Comparing Algorithms for Lexicon Based Sentiment Analysis
- A Simple and Efficient Multi-Task Learning Approach for Conditioned Dialogue Generation
- Obsessive-Compulsive Personality Disorder Co-occurring in Individuals with Obsessive-Compulsive Disorder: A Systematic Review and Meta-analysis
- Interactive Task and Concept Learning from Natural Language Instructions and GUI Demonstrations
- Few-NERD: A Few-Shot Named Entity Recognition Dataset
- Measuring Whiteness: A Systematic Review of Instruments and Call to Action
- Ensemble Pruning via Margin Maximization
- LLM-as-a-Judge: Rapid Evaluation of Legal Document Recommendation for Retrieval-Augmented Generation
- An Attention-Based Deep Learning Approach for Sleep Stage Classification With Single-Channel EEG
- TREMO: A dataset for emotion analysis in Turkish
- Auto-Slides: An Interactive Multi-Agent System for Creating and Customizing Research Presentations
- Humanizing Automated Programming Feedback: Fine-Tuning Generative Models with Student-Written Feedback
- My Favorite Streamer is an LLM: Discovering, Bonding, and Co-Creating in AI VTuber Fandom
- Venture capitalists' decision criteria in new venture evaluation
- r/Fakeddit: A New Multimodal Benchmark Dataset for Fine-grained Fake News Detection
- Examining teaching assistant pedagogies in traditional laboratories and recitations
- Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles
- Development of an Instrument to Measure the Perceptions of Adopting an Information Technology Innovation
- An Alternative Measure of Effect Size for Cochran's Q Test for Related Proportions
- RADAR: A Reasoning-Guided Attribution Framework for Explainable Visual Data Analysis
- Fear of the coronavirus (COVID-19): Predictors in an online study conducted in March 2020
- What Were You Thinking? An LLM-Driven Large-Scale Study of Refactoring Motivations in Open-Source Projects
- GitHub Copilot AI pair programmer: Asset or Liability?
- MS2: Multi-Document Summarization of Medical Studies
- From Vision to Validation: A Theory- and Data-Driven Construction of a GCC-Specific AI Adoption Index
- Grader variability and the importance of reference standards for evaluating machine learning models for diabetic retinopathy
- What if I ask in alia lingua? Measuring Functional Similarity Across Languages
- Towards generalisable hate speech detection: a review on obstacles and solutions
- The Impact of Critique on LLM-Based Model Generation from Natural Language: The Case of Activity Diagrams
- Automatic evaluation of reading aloud performance in children
- Resilients, Overcontrollers, and Undercontrollers: The replicability of the three personality prototypes across informants
- StableSleep: Source-Free Test-Time Adaptation for Sleep Staging with Lightweight Safety Rails
- Understanding Architecture Erosion: The Practitioners' Perceptive
- DynaGuard: A Dynamic Guardian Model With User-Defined Policies
- ABCD-LINK: Annotation Bootstrapping for Cross-Document Fine-Grained Links
- Understanding Code Understandability Improvements in Code Reviews
- Flaw or Artifact? Rethinking Prompt Sensitivity in Evaluating LLMs
- Towards Emotional Support Dialog Systems
- Partial success in closing the gap between human and machine vision
- A Benchmark for Modeling Violation-of-Expectation in Physical Reasoning Across Event Categories
- Using technology in special education: current practices and trends
- Automated Test Validators for Flaky Cyber-Physical System Simulators: Approach and Evaluation
- Automated Quality Assessment for LLM-Based Complex Qualitative Coding: A Confidence-Diversity Framework
- Guidelines for Empirical Studies in Software Engineering involving Large Language Models
- Inference Gap in Domain Expertise and Machine Intelligence in Named Entity Recognition: Creation of and Insights from a Substance Use-related Dataset
- Multimodal Emotion-Cause Pair Extraction in Conversations
- Understanding the role of single-board computers in engineering and computer science education: A systematic literature review
- Mapping Toxic Comments Across Demographics: A Dataset from German Public Broadcasting
- MAVEN: A Massive General Domain Event Detection Dataset
- Are Companies Taking AI Risks Seriously? A Systematic Analysis of Companies' AI Risk Disclosures in SEC 10-K forms
- M3HG: Multimodal, Multi-scale, and Multi-type Node Heterogeneous Graph for Emotion Cause Triplet Extraction in Conversations
- LiveMCP-101: Stress Testing and Diagnosing MCP-enabled Agents on Challenging Queries
- Toward Responsible ASR for African American English Speakers: A Scoping Review of Bias and Equity in Speech Technology
- Counterspeech for Mitigating the Influence of Media Bias: Comparing Human and LLM-Generated Responses
- Don't Think Twice! Over-Reasoning Impairs Confidence Calibration
- EmoTale: An Enacted Speech-emotion Dataset in Danish
- Temperature-Dependent Evolutionary Speed Shapes the Evolution of Biodiversity Patterns Across Tetrapod Radiations
- Real-Time Classification of Twitter Trends
- Improving K-12 Teachers’ Acceptance of Open Educational Resources by Open Educational Practices: A Mixed Methods Inquiry
- Towards Automatic Bot Detection in Twitter for Health-related Tasks
- Trouble when they walk in? Candidates, dark personality, and attack behavior in German televised debates
- The component structure of memory during development.
- Inter-rater Reliability of the Modified Japanese Orthopedic Association Score in Degenerative Cervical Myelopathy
- A First Look at Bugs in LLM Inference Engines
- An interpretable semi-supervised classifier using two different strategies for amended self-labeling
- Defects4Log: Benchmarking LLMs for Logging Code Defect Detection and Reasoning
- Psychedelic therapy for smoking cessation: Qualitative analysis of participant accounts
- Clustering of Social Media Messages for Humanitarian Aid Response during Crisis
- Case study: Mapping potential informal settlements areas in Tegucigalpa with machine learning to plan ground survey
- Lost in Moderation: How Commercial Content Moderation APIs Over- and Under-Moderate Group-Targeted Hate Speech and Linguistic Variations
- Sensitivity toward dark matter annihilation imprints on 21-cm signal with SKA-Low: A convolutional neural network approach
- Planner-Refiner: Dynamic Space-Time Refinement for Vision-Language Alignment in Videos
- BharatBBQ: A Multilingual Bias Benchmark for Question Answering in the Indian Context
- The dire disregard of measurement invariance testing in psychological science.
- Alcohol‐cancer risk communication on social media: A content analysis of alcohol‐related Instagram and TikTok posts
- Clinical Utility of the Automatic Phenotype Annotation in Unstructured Clinical Notes: ICU Use Cases
- Examining the Association Between Internet Use and Perceived Stress in Adults: Longitudinal Observational Study Combining Web Tracking Data With Questionnaires
- Demystifying Feature Requests: Leveraging LLMs to Refine Feature Requests in Open-Source Software
- Scaling Success: A Systematic Review of Peer Grading Strategies for Accuracy, Efficiency, and Learning in Contemporary Education
- Understanding Inconsistent State Update Vulnerabilities in Smart Contracts
- Variable selection via knockoffs in missing data settings with categorical predictors
- Do Ethical AI Principles Matter to Users? A Large-Scale Analysis of User Sentiment and Satisfaction
- FineDialFact: A benchmark for Fine-grained Dialogue Fact Verification
- How Do LLMs Persuade? Linear Probes Can Uncover Persuasion Dynamics in Multi-Turn Conversations
- An Overview of 7726 User Reports: Uncovering SMS Scams and Scammer Strategies
- Finding Needles in Images: Can Multimodal LLMs Locate Fine Details?
- An ML-based Approach to Predicting Software Change Dependencies: Insights from an Empirical Study on OpenStack
- I Think, Therefore I Am Under-Qualified? A Benchmark for Evaluating Linguistic Shibboleth Detection in LLM Hiring Evaluations
- Beneath the Surface of the Sexual Harassment Label: A Mixed Methods Study of Young Working Women
- Fostering argumentative knowledge construction through enactive role play in Second Life
- Towards Transparent AI Grading: Semantic Entropy as a Signal for Human-AI Disagreement
- Are Today's LLMs Ready to Explain Well-Being Concepts?
- Explainable Deep Neural Network for Multimodal ECG Signals: Intermediate vs Late Fusion
- From App Features to Explanation Needs: Analyzing Correlations and Predictive Potential
- Argument Quality Annotation and Gender Bias Detection in Financial Communication through Large Language Models
- Psychological safety in software workplaces: A systematic literature review
- Understanding Prediction Discrepancies in Machine Learning Classifiers
- An Effective Entropy-assisted Mind-wandering Detection System with EEG Signals based on MM-SART Database
- Cross-lingual Opinions and Emotions Mining in Comparable Documents
- Understanding environmental tweets of for-profits and nonprofits and their effects on user responses
- PunchPulse: A Physically Demanding Virtual Reality Boxing Game Designed with, for and by Blind and Low-Vision Players
- A Confidence-Diversity Framework for Calibrating AI Judgement in Accessible Qualitative Coding Tasks
- EHSAN: Leveraging ChatGPT in a Hybrid Framework for Arabic Aspect-Based Sentiment Analysis in Healthcare
- Toward Using Machine Learning as a Shape Quality Metric for Liver Point Cloud Generation
- A Methodological Framework for LLM-Based Mining of Software Repositories
- MCeT: Behavioral Model Correctness Evaluation using Large Language Models
- NyayaRAG: Realistic Legal Judgment Prediction with RAG under the Indian Common Law System
- Robust Collaborative Learning of Patch-level and Image-level Annotations for Diabetic Retinopathy Grading from Fundus Image
- BigTokDetect: A Clinically-Informed Vision-Language Modeling Framework for Detecting Pro-Bigorexia Videos on TikTok
- Stakeholder engagement and dialogic accounting
- Helping or Homogenizing? GenAI as a Design Partner to Pre-Service SLPs for Just-in-Time Programming of AAC
- AgriEval: A Comprehensive Chinese Agricultural Benchmark for Large Language Models
- How Growing Toxicity Manifests: A Topic Trajectory Analysis of U.S. Immigration Discourse on Social Media
- Towards Recognizing Phrase Translation Processes: Experiments on English-French
- STARN-GAT: A Multi-Modal Spatio-Temporal Graph Attention Network for Accident Severity Prediction
- Named Entities in Medical Case Reports: Corpus and Experiments
- An Inventory of Preposition Relations
- A mixed method study of DevOps challenges
- A study on cost behaviors of binary classification measures in class-imbalanced problems
- Social perception in negotiation
- Causal relatedness and importance of story events
- What Makes Code Generation Ethically Sourced?
- RoD-TAL: A Benchmark for Answering Questions in Romanian Driving License Exams
- Using Under-trained Deep Ensembles to Learn Under Extreme Label Noise
- Schedule for Affective Disorders and Schizophrenia for School-Age Children-Present and Lifetime Version (K-SADS-PL): Initial Reliability and Validity Data
- ShEMO -- A Large-Scale Validated Database for Persian Speech Emotion Detection
- Significance Tests for the Measure of Raw Agreement
- Dutch General Public Reaction on Governmental COVID-19 Measures and Announcements in Twitter Data
- Multimodal Behavioral Patterns Analysis with Eye-Tracking and LLM-Based Reasoning
- Who shapes crisis communication on Twitter? An analysis of influential German-language accounts during the COVID-19 pandemic
- Multi-Year Evaluations of an FTA Card–Based Detection Protocol for Four Vector-Borne Viruses Affecting Potato
- Combining participatory and modeling approaches to investigate factors and drivers of soil erosion risk in mixed crop-livestock farms
- Neural networks with divisive normalization for image segmentation
- Validating telehealth assessment of physical concussion symptoms
- Beyond Accuracy: Ecological Validity of Machine Learning Models for Assessing Vegetation Conservation Classes in South Korea
- Systematic classification differences across eye movement detection algorithms
- Prevalence and Spirometric Transitions of PRISm and Obstruction: A Population‐Based Study
- A systematic review of latent class analysis in psychology: Examining the gap between guidelines and research practice
- Screening for REM Sleep Behaviour Disorder with Minimal Sensors
- CatSIM: A Categorical Image Similarity Metric
- Snow cover dynamics: an overlooked yet important feature of winter bird occurrence and abundance across the United States
- Hierarchical method for cataract grading based on retinal images using improved Haar wavelet
- The neural correlates of pain-related fear: A meta-analysis comparing fear conditioning studies using painful and non-painful stimuli
- Friendships are flexible, not fragile: Turning points in geographically-close and long-distance friendships
- Quantity and quality of parental language input to late-talking toddlers during play
- Lexico-semantic and affective modelling of Spanish poetry: A semi-supervised learning approach
- What Should I Learn First: Introducing LectureBank for NLP Education and Prerequisite Chain Learning
- Don't forget your classics: Systematizing 45 years of Ancestry for Security API Usability Recommendations
- Review on Requirements Modeling and Analysis for Self-Adaptive Systems: A Ten-Year Perspective
- Evaluating the ability of habitat suitability models to predict species presences
- Policy-Grounded Safety Evaluation of 20 Large Language Models
- Assessing the Reliability of Large Language Models for Deductive Qualitative Coding: A Comparative Study of ChatGPT Interventions
- Assessing Post Deletion in Sina Weibo: Multi-modal Classification of Hot Topics
- From Text to Codings
- Reproducibility of Machine Learning-Based Fault Detection and Diagnosis for HVAC Systems in Buildings: An Empirical Study
- DRAFT-What you always wanted to know but could not find about block-based environments
- CPC-CMS: Cognitive Pairwise Comparison Classification Model Selection Framework for Document-level Sentiment Analysis
- Hallucination Score: Towards Mitigating Hallucinations in Generative Image Super-Resolution
- On the Effectiveness of LLM-as-a-judge for Code Generation and Summarization
- A framework for reliable traffic surrogate safety assessment based on multi-object tracking data
- A protocol for evaluating AI chatbots’ capabilities for low-resource language teachers
- Use of artificial intelligence to support the assessment of the methodological quality of systematic reviews
- Advancing Mental Disorder Detection: A Comparative Evaluation of Transformer and LSTM Architectures on Social Media
- Automatically assessing oral narratives of Afrikaans and isiXhosa children
- Towards Better Requirements from the Crowd: Developer Engagement with Feature Requests in Open Source Software
- Content Analysis in Mass Communication: Assessment and Reporting of Intercoder Reliability
- Temporal Stability of Heavy Drinking Days and Drinking Reductions Among Heavy Drinkers in the COMBINE Study
- Diagnostic reliability of the Semi-structured Assessment for Drug Dependence and Alcoholism (SSADDA)
- ANLIzing the Adversarial Natural Language Inference Dataset
- CD2CR: Co-reference Resolution Across Documents and Domains
- Meta‐analysis reveals that the effects of precipitation change on soil and litter fauna in forests depend on body size
- Novel Antischistosomal Drug Targets: Identification of Alkaloid Inhibitors of SmTGR via Integrated In Silico Methods
- Concordance between FVC and FEV 6 for identifying chronic airflow obstruction and spirometric restriction in the Burden of Obstructive Lung Disease (BOLD) study
- Macular OCT Classification Using a Multi-Scale Convolutional Neural Network Ensemble
- Security Smells in Ansible and Chef Scripts: A Replication Study
- A nonparametric Bayesian test of dependence
- Development of a qualitative data analysis codebook for peri‐ictal behavior in suspected functional seizures
- Tracking Sumatran Tiger (Panthera tigris sumatrae Pocock, 1929) distribution in Gunung Leuser National Park: The influence of prey presence and environmental variables on habitat selection
- Evaluation of Facebook as a Longitudinal Data Source for Parkinson’s Disease Insights
- A Dataset of General-Purpose Rebuttal
- Ranking Computer Vision Service Issues using Emotion
- Identifying Morality Frames in Political Tweets using Relational Learning
- Empirical studies of agile software development: A systematic review
- Mental Models of Adversarial Machine Learning
- Internal Value Alignment in Large Language Models through Controlled Value Vector Activation
- A training programme for early-stage researchers that focuses on developing personal science outreach portfolios
- Dr.Copilot: A Multi-Agent Prompt Optimized Assistant for Improving Patient-Doctor Communication in Romanian
- The Eighth Dialog System Technology Challenge
- Is Quantization a Deal-breaker? Empirical Insights from Large Code Models
- THAI Speech Emotion Recognition (THAI-SER) corpus
- From Research to Resources: Assessing Student Understanding and Skills in Quantum Computing
- ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models
- SPICE: An Automated SWE-Bench Labeling Pipeline for Issue Clarity, Test Coverage, and Effort Estimation
- Enhancing Essay Cohesion Assessment: A Novel Item Response Theory Approach
- Measuring Hypothesis Testing Errors in the Evaluation of Retrieval Systems
- Quantifying Uncertainty in Error Consistency: Towards Reliable Behavioral Comparison of Classifiers
- A proposal and assessment of an improved heuristic for the Eager Test smell detection
- EduCoder: An Open-Source Annotation System for Education Transcript Data
- Improving Label Quality by Jointly Modeling Items and Annotators
- A versatile index to characterize hysteresis between hydrological variables at the runoff event timescale
- WSCoach: Wearable Real-time Auditory Feedback for Reducing Unwanted Words in Daily Communication
- SymbolicThought: Integrating Language Models and Symbolic Reasoning for Consistent and Interpretable Human Relationship Understanding
- ScopeJudge: Cost-Aware Pre-Execution Gating for Offensive Security Agents
- A Structured Learning Approach with Neural Conditional Random Fields for Sleep Staging
- Agreement Metrics for LLM-as-Judge Evaluation: What to Report and Why
- Leveraging Information Technology Infrastructure to Facilitate a Firm's Customer Agility and Competitive Activity: An Empirical Investigation
- Intraclass Correlation Coefficient (ICC): A Framework for Monitoring and Assessing Performance of Trained Sensory Panels and Panelists
- Weakly Supervised Learning of Nuanced Frames for Analyzing Polarization in News Media
- An Investigation of Posttraumatic Stress Disorder and Depressive Symptomatology among Female Victimsof Interpersonal Trauma
- Do LLMs Hold Their Values? MANTA: A Multi-Turn Adversarial Benchmark for Animal Welfare Reasoning
- Sustainability Flags for the Identification of Sustainability Posts in Q&A Platforms
- Predicting the Transition from Short-term to Long-term Memory based on Deep Neural Network
- MedVAL: Toward Expert-Level Medical Text Validation with Language Models
- Sentinel SAR-optical fusion for crop type mapping using deep learning and Google Earth Engine
- Enhancing COBOL Code Explanations: A Multi-Agents Approach Using Large Language Models
- HCqa: Hybrid and Complex Question Answering on Textual Corpus and Knowledge Graph
- Evaluating the Effectiveness of Direct Preference Optimization for Personalizing German Automatic Text Simplifications for Persons with Intellectual Disabilities
- MRI can assess glenoid bone loss after shoulder luxation: inter- and intra-individual comparison with CT
- Recommending Variable Names for Extract Local Variable Refactorings
- Relevance of Unsupervised Metrics in Task-Oriented Dialogue for Evaluating Natural Language Generation
- Evaluating Lexicon-Based Sentiment Analysis Methods for Small Datasets on Low Compute Devices
- Development of a Framework for Youth- and Family-Specific Engagement in Research: Proposal for a Scoping Review and Qualitative Descriptive Study
- On Recipe Memorization and Creativity in Large Language Models: Is Your Model a Creative Cook, a Bad Cook, or Merely a Plagiator?
- Climate change-induced range shift of the endemic epiphytic lichenLobaria pindarensisin the Hindu Kush Himalayan region
- Learning from various labeling strategies for suicide-related messages on social media: An experimental study
- Evaluating and Improving Large Language Models for Competitive Program Generation
- On the limits of cross-domain generalization in automated X-ray prediction
- Accountability in Intervention Research: Developing a Fidelity Checklist of a Mental Health Intervention in Prisons
- Interaction Analysis by Humans and AI: A Comparative Perspective
- Domain Knowledge in Artificial Intelligence: Using Conceptual Modeling to Increase Machine Learning Accuracy and Explainability
- Semantic Caching for Improving Web Affordability
- The International Caries Detection and Assessment System (ICDAS): an integrated system for measuring dental caries
- Magnetic Resonance Classification of Lumbar Intervertebral Disc Degeneration
- The reliability of the pass/fail decision for assessments comprised of multiple components
- See-in-Pairs: Reference Image-Guided Comparative Vision-Language Models for Medical Diagnosis
- Predictive modelling of football injuries
- Leveraging Network Methods for Hub-like Microservice Detection
- Categorization of interestingness measures for knowledge extraction
- LAION-C: An Out-of-Distribution Benchmark for Web-Scale Vision Models
- Identifying Explanation Needs: Towards a Catalog of User-based Indicators
- Understanding the Challenges and Opportunities of Generative AI Apps: An Empirical Study
- PRISON: Unmasking the Criminal Potential of Large Language Models
- Specificity-Based Sentence Ordering for Multi-Document Extractive Risk Summarization
- Detecting Online Hate Speech Using Context Aware Models
- Essential-Web v1.0: 24T tokens of organized web data
- What Makes a Good Natural Language Prompt?
- Reciprocal Relationships Between Parenting Behavior and Disruptive Psychopathology from Childhood Through Adolescence
- K/DA: Automated Data Generation Pipeline for Detoxifying Implicitly Offensive Language in Korean
- Cross-project Classification of Security-related Requirements
- Konooz: Multi-domain Multi-dialect Corpus for Named Entity Recognition
- More Diverse Means Better: Multimodal Deep Learning Meets Remote Sensing Imagery Classification
- Automatic Argument Quality Assessment -- New Datasets and Methods
- Hatevolution: What Static Benchmarks Don't Tell Us
- Are You Convinced? Choosing the More Convincing Evidence with a Siamese Network
- Posted in Error: Did the CDC’s Retraction of Aerosol Guidance Undercut Its Public Reputation?
- Distribución espacio-temporal de Eichhornia crassipes (Mart.) Solms a través de teledetección en laguna La Turbina, Cuba
- What Users Value and Critique: Large-Scale Analysis of User Feedback on AI-Powered Mobile Apps
- Unsupervised patient representations from clinical notes with interpretable classification decisions
- Let ’em Talk!
- Peering Inside a Canadian Interrogation Room
- Quality of remotely-collected gaze data in autistic and nonspectrum children
- Cauda equina redundant nerve roots are associated to the degree of spinal stenosis and to spondylolisthesis
- A multi-task learning method for extraction of newly constructed areas based on bi-temporal hyperspectral images
- AI Techniques in the Microservices Life-Cycle: a Systematic Mapping Study
- Ultra-Fast, Low-Storage, Highly Effective Coarse-grained Selection in Retrieval-based Chatbot by Using Deep Semantic Hashing
- Interobserver reliability of the Tile classification system for pelvic fractures among radiologists and surgeons
- Emotion-Aware, Emotion-Agnostic, or Automatic: Corpus Creation Strategies to Obtain Cognitive Event Appraisal Annotations
- Development and validation of a measurement instrument for studying supply chain management practices
- Developing Students’ Critical Thinking Skills and Argumentation Abilities Through Augmented Reality–Based Argumentation Activities in Science Classes
- Technical-tactical evolution of women’s football: a comparativeanalysis of ball possessions in the FIFA Women’s World CupFrance 2019 and Australia & New Zealand 2023
- Adaptive Keywords Extraction with Contextual Bandits for Advertising on Parked Domains
- Open-Domain Dialogue Generation Based on Pre-trained Language Models
- From trust to accountability: Negotiating media performance in the Netherlands, 1987—2007
- An Empirical Study of OpenAI API Discussions on Stack Overflow
- Character 3-gram Mover's Distance: An Effective Method for Detecting Near-duplicate Japanese-language Recipes
- Fine-grained Human Evaluation of Transformer and Recurrent Approaches to Neural Machine Translation for English-to-Chinese
- R-VGAE: Relational-variational Graph Autoencoder for Unsupervised Prerequisite Chain Learning
- Unsupervised Paraphrasing by Simulated Annealing
- Affinity-Based Hierarchical Learning of Dependent Concepts for Human Activity Recognition
- An Evidence-Based Systematic Review on Cognitive Interventions for Individuals With Dementia
- Supportive Messages Female Offenders Receive From Probation and Parole Officers About Substance Avoidance: Message Perceptions and Effects
- Recognizing Explicit and Implicit Hate Speech Using a Weakly Supervised Two-path Bootstrapping Approach
- Matter-of-Fact: A Benchmark for Verifying the Feasibility of Literature-Supported Claims in Materials Science
- Behavior-based evaluation of session satisfaction
- Graph-Embedding Empowered Entity Retrieval
- Automated phonological analysis and treatment target selection using AutoPATT
- DRE: An Effective Dual-Refined Method for Integrating Small and Large Language Models in Open-Domain Dialogue Evaluation
- Aspect-Based Argument Mining
- Expanded understanding of student conceptions of engineers: Validation of the modified draw‐an‐engineer test (mDAET) scoring rubric
- FailureSensorIQ: A Multi-Choice QA Dataset for Understanding Sensor Relationships and Failure Modes
- Quantifying task-relevant representational similarity using decision variable correlation
- Harnessing GIS, remote sensing, and machine learning for sustainable management and carbon sequestration of non-timber forest products in Gujarat, India
- A Systematic Mapping Study on Software Architecture for AI-based Mobility Systems
- Citations versus expert opinions: Citation analysis of Featured Reviews of the American Mathematical Society
- Evaluating Discourse and Dialogue Coding Schemes
- Analyzing Web Search Behavior for Software Engineering Tasks
- Climate-driven spread of giant hogweed [Heracleum mantegazzianum (Sommier & Levier) in Turkey: assessing future invasion risks under CMIP6 climate projections
- Corpus-pragmatic perspectives on the contemporary weakening of fuck: The case of teenage British English conversation
- Reducing cardiovascular risk among people living with HIV: Rationale and design of the INcreasing Statin Prescribing in HIV Behavioral Economics REsearch (INSPIRE) randomized controlled trial
- From Chat Logs to Collective Insights: Aggregative Question Answering
- Evaluating AI capabilities in detecting conspiracy theories on YouTube
- Bridging Semantic Gaps between Natural Languages and APIs with Word Embedding
- Automated Essay Scoring Incorporating Annotations from Automated Feedback Systems
- Human footprint with machine learning identifies risks of the invasive weed Conyza sumatrensis across land-use types under climate change
- Chinese Cyberbullying Detection: Dataset, Method, and Validation
- THiNK: Can Large Language Models Think-aloud?
- CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation
- Classification of dinosaur footprints using machine learning
- Preschoolers' questions and parents' explanations: Causal thinking in everyday activity
- Rethinking Metrics and Benchmarks of Video Anomaly Detection
- Can LLMs Evaluate What They Cannot Annotate? Revisiting LLM Reliability in Hate Speech Detection
- Generalized Permutation Framework for Testing Model Variable Significance
- Comparing Human and AI Rater Effects Using the Many-Facet Rasch Model
- The Militarized Interstate Events (MIE) dataset, 1816–2014
- You Can Do Better! If You Elaborate the Reason When Making Prediction
- Reliability: on the reproducibility of assessment data
- How does Human Resource Management help service organizations to thrive in uncertainties and risks: Postcrisis as a context
- A Bayesian Ordinal Logistic Regression Model to Correct for Interobserver Measurement Error in a Geographical Oral Health Study
- Annotation of Emotion Carriers in Personal Narratives
- Evaluating NLP Embedding Models for Handling Science-Specific Symbolic Expressions in Student Texts
- Learning and Exploiting Interclass Visual Correlations for Medical Image Classification
- Introducing the Psychological Autopsy Methodology Checklist
- RAIM: Recurrent Attentive and Intensive Model of Multimodal Patient Monitoring Data
- Data Doping or True Intelligence? Evaluating the Transferability of Injected Knowledge in LLMs
- When Do LLMs Admit Their Mistakes? Understanding The Role Of Model Belief In Retraction
- An Empirical Analysis of Vulnerability Detection Tools for Solidity Smart Contracts Using Line Level Manually Annotated Vulnerabilities
- ParsiNLU: A Suite of Language Understanding Challenges for Persian
- The Corrective Commit Probability Code Quality Metric
- Understanding the Link Between Spatial Distance and Social Distance
- BugRepro: Enhancing Android Bug Reproduction with Domain-Specific Knowledge Integration
- Voice to Vision: Enhancing Civic Decision-Making through Co-Designed Data Infrastructure
- Analysis of Trending Topics and Text-based Channels of Information Delivery in Cybersecurity
- Mixed Signals: Understanding Model Disagreement in Multimodal Empathy Detection
- SDLog: A Deep Learning Framework for Detecting Sensitive Information in Software Logs
- Reliable Decision Support with LLMs: A Framework for Evaluating Consistency in Binary Text Classification Applications
- Product Fit Uncertainty in Online Markets: Nature, Effects, and Antecedents
- The Impact of Executives’ IT Expertise on Reported Data Security Breaches
- Tracking Turbulence Through Financial News During COVID-19
- Parents’ understanding of genome and exome sequencing for pediatric health conditions: a systematic review
- Noise exposure and the risk of cancer: a comprehensive systematic review
- How Accurate Is Clinician Reporting of Chemotherapy Adverse Effects? A Comparison With Patient-Reported Symptoms From the Quality-of-Life Questionnaire C30
- Agreement and reproducibility between 3DStent vs. Optical Coherence Tomography for evaluation of stent area and diameter
- Using Multilabel Neural Network to Score High‐Dimensional Assessments for Different Use Foci: An Example with College Major Preference Assessment
- Doc2CI: A Multi-Service Study of CI Configuration Generation Using Large Language Models
- A Multi-Task Benchmark for Abusive Language Detection in Low-Resource Settings
- Counterspeech the ultimate shield! Multi-Conditioned Counterspeech Generation through Attributed Prefix Learning
- RefactorAssist: Agentic Refinement for Reliable Code Refactoring
- Covariance Symmetries Classification in Multitemporal/Multipass PolSAR Images
- Feedback from nuclear RNA on transcription promotes robust RNA concentration homeostasis in human cells
- Beyond bigrams: call sequencing in the common marmoset ( Callithrix jacchus ) vocal system
- Visual Environment, Attention Allocation, and Learning in Young Children
- Smiling underwater: Exploring playful signals and rapid mimicry in bottlenose dolphins
- The iconicity toolbox: empirical approaches to measuring iconicity
- Ecological and social outcomes of urbanization on regional farming systems: a global synthesis
- Truth Is a Lie: Crowd Truth and the Seven Myths of Human Annotation
- SurgXBench: Explainable Vision-Language Model Benchmark for Surgery
- Benchmarking Critical Questions Generation: A Challenging Reasoning Task for Large Language Models
- Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation
- Equal is Not Always Fair: A New Perspective on Hyperspectral Representation Non-Uniformity
- Center for Epidemiologic Studies Depression Scale (CES-D) as a screening instrument for depression among community-residing older adults.
- Comparing LLM Text Annotation Skills: A Study on Human Rights Violations in Social Media Data
- Potential types of bias when estimating causal effects in environmental research and how to interpret them
- How white is the global elite? An analysis of race, gender and network structure
- The infection-tolerant white-footed deermouse tempers interferon responses to endotoxin in comparison to the mouse and rat
- Synoptic patterns and mesoscale precursors of Italian tornadoes
- Towards Automated Situation Awareness: A RAG-Based Framework for Peacebuilding Reports
- Predicting reptile distributions at the mesoscale: relation to climate and topography
- Structural and interactional aspects of adverbial sentences in English mother-child interactions: an analysis of two dense corpora
- RadPRISM: Schema-stratified radiology-report supervision for concept-disentangled image representations and visual grounding
- Evaluating Simplification Algorithms for Interpretability of Time Series Classification
- LCES: Zero-shot Automated Essay Scoring via Pairwise Comparisons Using Large Language Models
- A Machine Learning model of the combination of normalized SD1 and SD2\n indexes from 24h-Heart Rate Variability as a predictor of myocardial\n infarction
- A Consolidated System for Robust Multi-Document Entity Risk Extraction and Taxonomy Augmentation
- Exploring Challenges in Test Mocking: Developer Questions and Insights from StackOverflow
- LLM-Based Detection of Tangled Code Changes for Higher-Quality Method-Level Bug Datasets
- Automatic detection of abnormal clinical EEG: comparison of a finetuned foundation model with two deep learning models
- ViMRHP: A Vietnamese Benchmark Dataset for Multimodal Review Helpfulness Prediction via Human-AI Collaborative Annotation
- When are Deep Networks really better than Decision Forests at small sample sizes, and how?
- A model for spectra-based software diagnosis
- Equivariant Imaging Biomarkers for Robust Unsupervised Segmentation of Histopathology
- Frame In, Frame Out: Do LLMs Generate More Biased News Headlines than Humans?
- UKElectionNarratives: A Dataset of Misleading Narratives Surrounding Recent UK General Elections
- Comparing two SVM models through different metrics based on the confusion matrix
- POTENTIAL DISTRIBUTION OF TWO INSECTS WITH GASTRONOMIC VALUE IN MEXICO
- Decoding Open-Ended Information Seeking Goals from Eye Movements in Reading
- Gender Representations on Disney Channel, Cartoon Network, and Nickelodeon Broadcasts in the United States
- Mental disorders and suicide in Northern Ireland
- A Defect Taxonomy for Infrastructure as Code: A Replication Study
- The Social Media Disorder Scale
- Behavior of agreement measures in the presence of zero cells and biased marginal distributions
- Descriptor: C++ Self-Admitted Technical Debt Dataset (CppSATD)
- Evaluating the Impact of Data Cleaning on the Quality of Generated Pull Request Descriptions
- Where is the “theory” within the field of educational technology research?
- Summer and winter habitat suitability of Marco Polo argali in southeastern Tajikistan: A modeling approach
- Habitat assessment of Marco Polo sheep (Ovis ammon polii) in Eastern Tajikistan: Modeling the effects of climate change
- Modified Overt Aggression Scale (MOAS) for People with Intellectual Disability and Aggressive Challenging Behaviour: A Reliability Study
- Proper Correlation Coefficients for Nominal Random Variables
- Corporate social identity: an analysis of the Indian banking sector
- Search-Induced Issues in Web-Augmented LLM Code Generation: Detecting and Repairing Error-Inducing Pages
- Distributed Partial Information Puzzles: Examining Common Ground Construction Under Epistemic Asymmetry
- Attention Interpretability Across NLP Tasks
- ScienceExamCER: A High-Density Fine-Grained Science-Domain Corpus for Common Entity Recognition
- Measuring Faithfulness Depends on How You Measure: Classifier Sensitivity in LLM Chain-of-Thought Evaluation
- Prodromal Assessment With the Structured Interview for Prodromal Syndromes and the Scale of Prodromal Symptoms: Predictive Validity, Interrater Reliability, and Training to Reliability
- Technology-Enhanced Collaborative Inquiry in K–12 Classrooms: A Systematic Review of Empirical Studies
- Qualitative methods can enrich quantitative research on occupational stress: An example from one occupational group
- Transgender adults, gender-affirming hormone therapy and blood pressure: a systematic review
- Digital discourse research in linguistics (2010-2025): A bibliometric and content analysis
- The Individual-Targeting Assumption: A Systematic Review of Proactive Robots in Human Group Settings
- The effectiveness of an aerobic exercise training on patients with neck pain during a short- and long-term follow-up: a prospective double-blind randomized controlled trial
- Light-weight sleep monitoring: electrode distance matters more than placement for automatic scoring
- Clinical gait data analysis based on Spatio-Temporal features
- Machine learning for the recognition of emotion in the speech of couples in psychotherapy using the Stanford Suppes Brain Lab Psychotherapy Dataset
- “That's not My Job”: Developing Flexible Employee Work Orientations
- Stuck in the Past: Why Managers Persist with New Product Failures
- Coder Reliability and Misclassification in the Human Coding of Party Manifestos
- “Friendship” Interactions and Expression of Agitation among Residents of a Dementia Care Unit
- Students’ sociodemographic characteristics and writing performance: a systematic literature review
- A comparison of methods for defining sociometric status among children.
- Health-Care Provider Planned Responses to Patient Misunderstandings about End-of-Life Care
- Inventoried and observed stress in parent-child interactions
- Reliability and Expected Loss: A Unifying Principle
- General Estimators for the Reliability of Qualitative Data
- Inspecting the Achilles heel: a quantitative analysis of 50 years of family business definitions
- THAT'S NOT MY JOB: DEVELOPING FLEXIBLE EMPLOYEE WORK.
- Primary and secondary goals in the production of interpersonal influence messages
- Begging with a Purpose? Testing Behavioural Hallmarks of First-Order Intentionality in Free-ranging Hanuman Langurs
- Reflux Symptom Index versus Reflux Finding Score
- IslamicLegalBench: Evaluating LLMs Knowledge and Reasoning of Islamic Law Across 1,200 Years of Islamic Pluralist Legal Traditions
- Using Finite-State Machines to Automatically Scan Classical Greek Hexameter
- Antidepressants versus placebo for panic disorder in adults
- Antidepressants and benzodiazepines for panic disorder in adults
- The effectiveness of semantic feature analysis: An evidence-based systematic review
- The Sentience Readiness Index: A Preliminary Framework for Measuring National Preparedness for the Possibility of Artificial Sentience
- The drinking water contamination crisis in Flint: Modeling temporal trends of lead level since returning to Detroit water system
- Estimating the quality of landslide susceptibility models
- Optimal landslide susceptibility zonation based on multiple forecasts
- Interventions for latent autoimmune diabetes (LADA) in adults
- A Benthic Terrain Classification Scheme for American Samoa
- Effect of 3 years of SAFE (surgery, antibiotics, facial cleanliness, and environmental change) strategy for trachoma control in southern Sudan: a cross-sectional study
- The concept of major depression
- A Re-analysis of the Reliability of Psychiatric Diagnosis
- Intent Laundering: AI Safety Datasets Are Not What They Seem
- Special Paper: A Global Biome Model Based on Plant Physiology and Dominance, Soil Properties and Climate
- Why parapsychology cannot become a science
- Continuous glucose monitoring systems for type 1 diabetes mellitus
- Reliability of Scores on the Stroke Rehabilitation Assessment of Movement (STREAM) Measure
- Phosphodiesterase inhibitors for erectile dysfunction in patients with diabetes mellitus
- Mechanisms of Team-Sport-Related Brain Injuries in Children 5 to 19 Years Old: Opportunities for Prevention
- Staying Centered: Curriculum Leadership in a Turbulent Era
- Eviction and the Reproduction of Urban Poverty
- Benchmarking Reward Hack Detection in Code Environments via Contrastive Analysis
- The Invisible Hand of AI Libraries Shaping Open Source Projects and Communities
- TRUST: An LLM-Based Dialogue System for Trauma Understanding and Structured Assessments
- Current threats faced by Neotropical parrot populations
- Truecluster matching
- A snow cover climatology for the Pyrenees from MODIS snow products
- Safety and accuracy follow different scaling laws in clinical large language models
- Using LLMs in Generating Design Rationale for Software Architecture Decisions
- Modelling invasion for a habitat generalist and a specialist plant species
- An Empirical Study on Common Defects in Modern Web Browsers Using Knowledge Embedding in GPT-4o
- Measuring the Gap Between Human and LLM Research Ideas
- ClinHallu: A Benchmark for Diagnosing Stage-Wise Hallucinations in Medical MLLM Reasoning
- CATEKAPPA: An R Shiny Application for Design and Analysis of Consistency Tests Based on the Kappa Statistic for Categorical Responses
- MalwareTextDB: A Database for Annotated Malware Articles
- Multituberculate Mammals Show Evidence of a Life History Strategy Similar to That of Placentals, Not Marsupials
- Reconstructing spatial vulnerability to forest loss by fire in pre‐historic New Zealand
- The DDI corpus: An annotated corpus with pharmacological substances and drug–drug interactions
- When Cow Urine Cures Constipation on YouTube: Limits of LLMs in Detecting Culture-specific Health Misinformation
- Can LLMs Accurately Score Medical Diagnoses and Clinical Reasoning?
- Reasoning-Based Refinement of Unsupervised Text Clusters with LLMs
- AI-Mediated Feedback Improves Student Revisions: A Randomized Trial with FeedbackWriter in a Large Undergraduate Course
- Improving Code Generation via Small Language Model-as-a-judge
- Performance Smells in ML and Non-ML Python Projects: A Comparative Study
- Ordinal Regression Methods: Survey and Experimental Study
- Towards Automated Detection of Inline Code Comment Smells
- Biotic and abiotic variables show little redundancy in explaining tree species distributions
- Cuba: Exploring the History of Admixture and the Genetic Basis of Pigmentation Using Autosomal and Uniparental Markers
- Exercise therapy for multiple sclerosis
- TextTIGER: Text-based Intelligent Generation with Entity Prompt Refinement for Text-to-Image Generation
- Confirmatory bias in evaluating personality test information: Am I really that kind of person?
- Early Detection of Alzheimer's Disease Using Explainable Machine Learning on Clinical Biomarkers: A Multi-Class Classification Study Using the Alzheimer's Disease Neuroimaging Initiative (ADNI) Dataset
- Relative expertise in an everyday reasoning task: Epistemic understanding, problem representation, and reasoning competence
- Commercially Available Mobile Phone Headache Diary Apps: A Systematic Review
- Is dispersal guided by the environment? A comparison of interspecific gene flow estimates among differentiated regions of a newt hybrid zone
- Rosiglitazone for type 2 diabetes mellitus
- Low glycaemic index, or low glycaemic load, diets for diabetes mellitus
- Self-monitoring of blood glucose in patients with type 2 diabetes mellitus who are not using insulin
- Counting polyamorists who count: Prevalence and definitions of an under-researched form of consensual nonmonogamy
- Stratification by Skin Color in Contemporary Mexico
- Rewards of kindness? A meta-analysis of the link between prosociality and well-being.
- The human factor in SCM
- Truecluster: robust scalable clustering with model selection
- Green tea for weight loss and weight maintenance in overweight or obese adults
- An Examination of Interrater Reliability for Scoring the Rorschach Comprehensive System in Eight Data Sets
- An on-line interpretive Rorschach approach: Using Exner’s comprehensive system
- LogSieve: Task-Aware CI Log Reduction for Sustainable LLM-Based Analysis
- Topotecan for ovarian cancer
- The Evolution of Wikipedia’s Norm Network
- Prosodic Markers of Saliency in Humorous Narratives
- Reasoning Models Will Sometimes Lie About Their Reasoning
- An Empirical Study of Policy-as-Code Adoption in Open-Source Software Projects
- LDU-Bench: Multimodal LLM Evaluation for Lithography Defect Understanding under Layout-Varying Circuit Backgrounds
- Impaired Reception of Nonverbal Cues in Women with Premenstrual Tension Syndrome
- FACTWASH: Catching AI Rewrites That Wash Hearsay into Fact
- Depicting the Quarterback in Black and White: A Content Analysis of College and Professional Football Broadcast Commentary
- Halothane Binding Proteome in Human Brain Cortex
- Mitigating work alienation: what can we learn from employee ownership?
- Immature siblings and mother-infant relationships among free-ranging rhesus monkeys on Cayo Santiago
- Path dependence and the validation of agent‐based spatial models of land use
- Building UD Cairo for Old English in the Classroom
- From Bugs to Benchmarks: A Comprehensive Survey of Software Defect Datasets
- Recalculation of the Critical Values for Lawshe’s Content Validity Ratio
- Improving clinical practice using clinical decision support systems: a systematic review of trials to identify features critical to success
- A semi-structured clinical interview for the assessment of diagnosis and mental state in the elderly: the Geriatric Mental State Schedule: I. Development and reliability
- Modeling Communication Perception in Development Teams Using Monte Carlo Methods
- Learning under Concept Drift: A Review
- Assessing Media Influences on Middle School–Aged Children's Perceptions of Women in Science Using the Draw-A-Scientist Test (DAST)
- Data Stream Classification Guided by Clustering on Nonstationary Environments and Extreme Verification Latency
- STRIVE: Probing Reasoning Limits in Graded Plausibility Generation and Evaluation
- On Reliability of Patch Correctness Assessment
- Guideline-as-Oracle: Zero-Annotation Training of an Ophthalmic Telephone Triage Agent
- Consensus Measures for Unstructured Biomedical Text Annotations
- A Finnish news corpus for named entity recognition
- FinProBench: Evaluating Financial AI Agents with Role-Grounded Rubrics Derived from Professional Deliverables
- A scoping review of AI-mediated informal language learning: Mapping out the terrain and identifying future directions
- The Effect of Number of Rating Scale Categories on Levels of Interrater Reliability : A Monte Carlo Investigation
- Diagnostic accuracy of emergency department ECGs in hyperkalemia detection: A cross-sectional study
- Characterizing and Comparing External Measures for the Assessment of Cluster Analysis and Community Detection
- Zooming in: Identifying fine‐grained verbal dynamics that influence coachees' self‐regulation statements during copreneur coaching sessions
- TathyaNyaya and FactLegalLlama: Advancing Factual Judgment Prediction and Explanation in the Indian Legal Context
- CliME: Evaluating Multimodal Climate Discourse on Social Media and the Climate Alignment Quotient (CAQ)
- Voluntary stuttering suppresses true stuttering: A window on the speech perception-production link
- The Inhibition of Stuttering Via the Perceptions and Production of Syllable Repetitions
- PreSumm: Predicting Summarization Performance Without Summarizing
- Campaña digital en clave de género: un estudio de caso de las elecciones autonómicas andaluzas de 2022 a través de Facebook
- Do they agree? Bibliometric evaluation versus informed peer review in the Italian research assessment exercise
- A Comprehensive Evaluation of Code Language Models for Security Patch Detection
- Validity and reliability of the Movement Assessment Battery for Children‐2 Checklist for children with and without motor impairments
- When Does Diversity Help Generalization in Classification Ensembles?
- On Developers' Self-Declaration of AI-Generated Code: An Analysis of Practices
- Emo Pillars: Knowledge Distillation to Support Fine-Grained Context-Aware and Context-Less Emotion Classification
- Applying Cooperative Machine Learning to Speed Up the Annotation of Social Signals in Large Multi-modal Corpora
- Preliminary study of optical coherence tomography imaging to identify microscopic extrathyroidal extension in patients with papillary thyroid carcinoma
- Using the full-text content of academic articles to identify and evaluate algorithm entities in the domain of natural language processing
- Significativity Indices for Agreement Values
- Combating Toxic Language: A Review of LLM-Based Strategies for Software Engineering
- Development of the Patient Education Materials Assessment Tool (PEMAT): A new measure of understandability and actionability for print and audiovisual patient information
- Epidemiology of brucellosis, Q Fever and Rift Valley Fever at the human and livestock interface in northern Côte d’Ivoire
- The Stochastic Replica Approach to Machine Learning: Stability and Parameter Optimization
- Structured Legal Document Generation in India: A Model-Agnostic Wrapper Approach with VidhikDastaavej
- A Local Method for Identifying Causal Relations under Markov Equivalence
- Revisiting Uncertainty Quantification Evaluation in Language Models: Spurious Interactions with Response Length Bias Results
- The mini nutritional assessment (MNA) and its use in grading the nutritional state of elderly patients
- Exploring the reliability of social and environmental disclosures content analysis
- Photometric Classifications of Evolved Massive Stars: Preparing for the Era of Webb and Roman with Machine Learning
- Sleep staging from electrocardiography and respiration with deep learning
- Deriving suitability factors for CA-Markov land use simulation model based on local historical data
- The ‘as code’ activities: development anti-patterns for infrastructure as code
- Multi-hop assortativities for network classification
- DeepSeeNet: A Deep Learning Model for Automated Classification of Patient-based Age-related Macular Degeneration Severity from Color Fundus Photographs
- Contextual Factors Contributing to Ethnic Identity Development of Second-Generation Iranian American Adolescents
- Harnessing Generative AI to Facilitate Epistemic Understanding of Students in Scientific Inquiry
- The Role of Educational Data Mining and Artificial Intelligence Supported Learning Analytics on Conceptual Change: New Approaches to Differentiated Instruction
- An Empirical Study of Python Library Migration Using Large Language Models
- Ensemble Pruning Based on Objection Maximization With a General Distributed Framework
- Mayo clinical Text Analysis and Knowledge Extraction System (cTAKES): architecture, component evaluation and applications
- Comparing the effectiveness of process-experiential with cognitive-behavioral psychotherapy in the treatment of depression.
- The Construction of a Model of the Process of Couples’ Forgiveness in Emotion-Focused Therapy for Couples
- TriQua: Reconciling Granularity and Context in Factuality Evaluation
- Neurobiology of culturally common maternal responses to infant cry
- Aging in the Digital Age : Public Beliefs About the Potential of Virtual Reality (VR) for the Aging Population
- LLMs as Span Annotators: A Comparative Study of LLMs and Humans
- An Empirical Study of Production Incidents in Generative AI Cloud Services
- The Ultimate Configuration Management Tool? Lessons from a Mixed Methods Study of Ansible's Challenges
- Should you use LLMs to simulate opinions? Quality checks for early-stage deliberation
- BOISHOMMO: Holistic Approach for Bangla Hate Speech
- Efficient Tuning of Large Language Models for Knowledge-Grounded Dialogue Generation
- Synthesizing High-Quality Programming Tasks with LLM-based Expert and Student Agents
- Beyond LLMs: A Linguistic Approach to Causal Graph Generation from Narrative Texts
- Conditional Data Synthesis Augmentation
- Multistream Vision Transformer Fusion with Stain-Physics Priors for Reliable Colorectal Histopathology Classification
- Patterns and issues in multiaxial psychiatric diagnosis
- Spurious precision: procedural validity of diagnostic assessment in psychotic disorders
- Teaching Data Science Students to Sketch Privacy Designs through Heuristics (Extended Technical Report)
- Accuracy assessment of land cover maps [wikipedia]
- Validity/reliability of PHQ-9 and PHQ-2 depression scales among adults living with HIV/AIDS in western Kenya. [europepmc]
- Radiographic evaluation of the hip has limited reliability. [europepmc]
- Reliability of a complication classification system for orthopaedic surgery. [europepmc]
- A comparison of Cohen's Kappa and Gwet's AC1 when calculating inter-rater reliability coefficients: a study conducted with personality disorder samples. [europepmc]
- Multidimensional in vivo hazard assessment using zebrafish. [europepmc]
- A review and meta-analysis of age-based stereotype threat: negative stereotypes, not facts, do the damage. [europepmc]
- Pictorial cigarette pack warnings: a meta-analysis of experimental studies. [europepmc]
- Metrics for evaluating 3D medical image segmentation: analysis, selection, and tool. [europepmc]
- Can mental health diagnoses in administrative data be used for research? A systematic review of the accuracy of routinely collected diagnoses. [europepmc]
- Utility of spherical human liver microtissues for prediction of clinical drug-induced liver injury. [europepmc]
- The relationship between physician burnout and quality of healthcare in terms of safety and acceptability: a systematic review. [europepmc]
- Empirical Comparison of Publication Bias Tests in Meta-Analysis. [europepmc]
- Best Practices for Developing and Validating Scales for Health, Social, and Behavioral Research: A Primer. [europepmc]
- Automated Gleason grading of prostate cancer tissue microarrays via deep learning. [europepmc]
- Agreement Between Prospective and Retrospective Measures of Childhood Maltreatment: A Systematic Review and Meta-analysis. [europepmc]
- Metascape provides a biologist-oriented resource for the analysis of systems-level datasets. [europepmc]
- Multitask learning and benchmarking with clinical time series data. [europepmc]
- Development and validation of a deep learning algorithm for improving Gleason scoring of prostate cancer. [europepmc]
- Understanding the care and support needs of older people: a scoping review and categorisation using the WHO international classification of functioning, disability and health framework (ICF). [europepmc]
- The advantages of the Matthews correlation coefficient (MCC) over F1 score and accuracy in binary classification evaluation. [europepmc]
- Promoting memory consolidation during sleep: A meta-analysis of targeted memory reactivation. [europepmc]
- Automatic diagnosis of the 12-lead ECG using a deep neural network. [europepmc]
- Fear of the coronavirus (COVID-19): Predictors in an online study conducted in March 2020. [europepmc]
- The Matthews correlation coefficient (MCC) is more reliable than balanced accuracy, bookmaker informedness, and markedness in two-class confusion matrix evaluation. [europepmc]
- Towards a guideline for evaluation metrics in medical image segmentation. [europepmc]