A Coefficient of Agreement for Nominal Scales
1960/04/01 by Jacob Cohen · 741 citations
Decision Sciences · #Reliability and Agreement in Measurement #Multi-Criteria Decision Making
paper · doi:10.1177/001316446002000104
Cited by
- A Ratio Test of Interrater Agreement With High Specificity
- Quantitative Image Analysis as an Adjunct to Manual Scoring of ER, PgR, and HER2 in Invasive Breast Carcinoma
- The Equivalence of Weighted Kappa and the Intraclass Correlation Coefficient as Measures of Reliability
- A Note on the Interpretation of Weighted Kappa and its Relations to Other Rater Agreement Statistics for Metric Scales
- Coefficient Kappa: Some Uses, Misuses, and Alternatives
- Acceptability and Efficacy of Group Behavioral Activation for Depression Among Adults: A Meta-Analysis
- The Measurement of Observer Agreement for Categorical Data
- A test of the interpersonal theory of suicide in a large, representative, retrospective and prospective study: Results from the Army Study to Assess Risk and Resilience in Servicemembers (Army STARRS)
- Visual Function Classification System for children with cerebral palsy: development and validation
- Prioritizing refuge sites for migratory geese to alleviate conflicts with agriculture
- ECG-based convolutional neural network in pediatric obstructive sleep apnea diagnosis
- Is It Good to Cooperate? Testing the Theory of Morality-as-Cooperation in 60 Societies
- An improved approach for predicting the distribution of rare and endangered species from occurrence and pseudo‐absence data
- Stability and predictors of somatic symptoms in men and women over 10 years: A real-world perspective from the prospective MONICA/KORA study
- Determining the Minimum Reliability Standard Based on a Decision Criterion
- Climate change drives range contraction and shapes species distribution in an alpine passerine: Caution required when comparing atlas data
- Back-Channel Representation: A Study of the Strategic Communication of Senators with the US Department of Labor
- Effect of active learning versus traditional lecturing on the learning achievement of college students in humanities and social sciences: a meta-analysis
- The effectiveness of an integratedSTEMcurriculum unit on middle school students' life science learning
- Population distributions of time to collision at brake application during car following from naturalistic driving data
- Assessing the accuracy of species distribution models: prevalence, kappa and the true skill statistic (TSS)
- Dimensions of Religion and Spirituality: A Longitudinal Topic Modeling Approach
- Fuzzy nearest neighbor algorithms: Taxonomy, experimental analysis and prospects
- Unilateral Posterior Crossbite is Not Associated with TMJ Clicking in Young Adolescents
- Who do they think you are? Inconsistencies in self- and proxy-reports of education within families
- Negotiator confidence: The impact of self-efficacy on tactics and outcomes
- Questionnaire-based diagnosis of benign paroxysmal positional vertigo
- A Systematic Survey on Image Description Techniques for STEM Domains
- The Effects of Spaced Practice on Second Language Learning: A Meta‐Analysis
- Evaluating migration hypotheses for the extinct Glyptotherium using ecological niche modeling
- How to map biomes: Quantitative comparison and review of biome‐mapping methods
- Tau-b or Not Tau-b: Measuring the Similarity of Foreign Policy Positions
- Chronic Obstructive Pulmonary Disease and Association With Mild Cognitive Impairment: The Mayo Clinic Study of Aging
- Interpretable EEG biomarkers with bag-of-waves: Spatial and temporal waveform dictionaries for low-data regimes
- Precision of Health-Related Quality-of-Life Data Compared With Other Clinical Measures
- EmoComicNet: A multi-task model for comic emotion recognition
- Green Governance: Boards of Directors’ Composition and Environmental Corporate Social Responsibility
- Scaffolding through prompts in digital learning: A systematic review and meta-analysis of effectiveness on learning achievement
- The Kappa Statistic in Reliability Studies: Use, Interpretation, and Sample Size Requirements
- Development and Validation of a Scale for Rating Motor Compensations Used for Reaching in Patients With Hemiparesis: The Reaching Performance Scale
- Fleiss’ kappa statistic without paradoxes
- Unfit for stranding assessment: a panel-scale multimodal-LLM audit of building-decarbonisation disclosure (BeDA)
- KaPilot: LLM-Assisted Generation of Kani Specifications for Unsafe Rust Verification
- A Latent Class Extension of Signal Detection Theory, with Applications
- ROOTCLUS: Searching for “ROOT CLUSters” in Three-Way Proximity Data
- Disentangling semantic and prosodic features of English poetry
- The Populist Style in American Politics: Presidential Campaign Discourse, 1952–1996
- Discordance between pain specialists and patients on the perception of dependence on pain medication: A multi-centre cross-sectional study
- Rape Myth Acceptance: Exploration of Its Structure and Its Measurement Using theIllinois Rape Myth Acceptance Scale
- Dynamic Capability Scoping for Enterprise AI Agents: A Synthetic Dataset and Three-Source Permission Architecture
- Identifying and Supporting Academically Low-Performing Schools in a Developing Country: An Application of a Specialized Multilevel IRT Model to PISA-D Assessment Data
- The Strengths and Difficulties Questionnaire (SDQ): the Factor Structure and Scale Validation in U.S. Adolescents
- How to Tell More is More: Quantity Discrimination in Eastern Box Turtles (Emydidae: Terrapene carolina)
- Identity Formation in Early and Middle Adolescents From Various Ethnic Groups: From Three Dimensions to Five Statuses
- The use of the Vocal Profile Analysis for speaker characterization: Methodological proposals
- Inter-Rater Reliability Methods in Qualitative Case Study Research
- Exploration, Explanation, and Parent–Child Interaction in Museums
- Evolutionary Personality Psychology
- Meta-analysis of action video game impact on perceptual, attentional, and cognitive skills.
- A meta-analytic review of experiments examining the effects of extrinsic rewards on intrinsic motivation.
- Data Annotations as Pedagogical Hints: From Subjective Labels to Critical Thinking
- MOF-Sleuth: Tool-Grounded Reward Alignment for Explainable Fine-Grained MOF CIF Auditing
- Active Listening in Integrative Negotiation
- Towards an Automated Test of LLM Security Knowledge
- How Far Can Wearable-Compatible Signals Go? A Controlled Decomposition of Non-EEG Sleep Staging
- Large Language Models for Citation Function Classification
- In-Context Learning for Wound Classification with Small Multimodal Language Models
- Evaluating medical AI under missing information: same-provider judges and human raters change apparent safety
- Visual Semantic Decoding of Electrocorticography from Video Stimuli using End-to-End Deep Learning
- SciCodePile: A 128GB Corpus and Executable Benchmark for Challenging Scientific Code Generation
- Real-World Evaluation of an AI Agent Drafting Translational Impact Summaries
- Espoonlahti mobile laser scanning tree species classification
- BanClickThumb: A Multimodal Dataset and Transformer Fusion Benchmarks for Clickbait Detection in Bengali YouTube Videos
- Safety That Does Not Transfer: Cross-Lingual Clinical Correctness Drift in Deployable Medical Language Models
- CARV: A Diagnostic Benchmark for Compositional Analogical Reasoning in Multimodal LLMs
- PriorProof: A Point-in-Time Measure of Technique Novelty for Formal Proofs
- One Modality to Forget Them All: Enhancing Cross-Modal Unlearning in Vision-Language Models
- Evaluating Open-Weight LLMs for Generating Structured Threat Information for Autonomous Vehicle Vulnerabilities
- Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning
- From Stateless to Situated: Building a Psychological World for LLM-Based Agents
- Large Language Models for Code Generation from Multilingual Prompts: A Curated Benchmark and a Study on Code Quality
- Students' Perceptions of Peer Grading
- Identity statuses and psychosocial functioning in Turkish youth: A person‐centered approach
- SWE-chat: Coding Agent Interactions From Real Users in the Wild
- Measuring the State of Open Science in Transportation Using Large Language Models
- Evaluating RAG for French immigration law: a benchmark and baseline study
- Can We Hide Machines in the Crowd? Quantifying Equivalence in LLM-in-the-loop Annotation Tasks
- Training language models to be warm and empathetic makes them less reliable and more sycophantic
- Base Models Beat Aligned Models at Randomness and Creativity
- Large language models for scientometric mapping of scientific controversy: A validated hybrid AI–Human framework
- Using Natural Language Processing to Automatically Detect Self-Admitted Technical Debt
- Empirical Evaluation of the Impact of Object-Oriented Code Refactoring on Quality Attributes: A Systematic Literature Review
- The Secret Life of Software Vulnerabilities: A Large-Scale Empirical Study
- Computing inter‐rater reliability and its variance in the presence of high agreement
- Looking at the dark and bright sides of identity formation: New insights from adolescents and emerging adults in Japan
- A Systematic Review of Collective Tactical Behaviours in Football Using Positional Data
- A multi‐dimensional measure of vocational identity status
- Factors Relating to Sprint Swimming Performance: A Systematic Review
- Language models and Automated Essay Scoring
- Using Fisher's Exact Test to Evaluate Association Measures for N-grams
- Remotely mapping gullying and incision in Maryland Piedmont headwater streams using repeat airborne lidar
- Acoustic characterization and classification of rockfish and walleye pollock in the Gulf of Alaska at bottom trawl locations
- Style Amnesia: Investigating Speaking Style Degradation and Mitigation in Multi-Turn Spoken Language Models
- Semicircular canal morphology in Rodentia and its relationship to locomotion
- Beyond the Tip of the Iceberg: Assessing Coherence of Text Classifiers
- Deep Contextualized Biomedical Abbreviation Expansion
- Can we really reduce ethnic prejudice outside the lab? A meta‐analysis of direct and indirect contact interventions
- Smart Contract Security: a Practitioners' Perspective
- Family Firms, M&A Strategies, and M&A Performance: A Meta-Analysis
- Grading exams using large language models: A comparison between human and AI grading of exams in higher education using ChatGPT
- A systematic review of economic evidence of artificial intelligence in healthcare
- Evaluation of performance measures in predictive artificial intelligence models to support medical decisions: overview and guidance
- Digital transformation: A multidisciplinary reflection and research agenda
- DICE: Discrete Interpretable Comparative Evaluation with Probabilistic Scoring for Retrieval-Augmented Generation
- Group-wise Contrastive Learning for Neural Dialogue Generation
- Are open educational resources (OER) and practices (OEP) effective in improving learning achievement? A meta-analysis and research synthesis
- Impact of environmental barriers on temnospondyl biogeography and dispersal during the Middle–Late Triassic
- ATCNet-CIAM for Multi-Session Motor Imagery EEG Signal Classification
- Global tracking of marine megafauna space use reveals how to achieve conservation targets
- Dyadic differences in empathy scores are associated with kinematic similarity during conversational question–answer pairs
- Multi-label Categorization of Accounts of Sexism using a Neural\n Framework
- Dynamic Context Selection for Document-level Neural Machine Translation via Reinforcement Learning
- Calibrated Tree-Neural Fusion for Fine-Grained Vegetation Community Classification
- Do Current Retrievers Cover All the Evidence? A Controlled Study of Conjunctive Cross-Page Retrieval
- Consumer Perception of Cows' Milk and Plant‐Based Milk Alternatives: Comparing Aotearoa–New Zealand and Singapore Consumers
- Leveraging Semantic Maps for City-Scale Cross-View Localization
- Sense it with your eyes: Sensation Generation and Understanding for Advertisements
- Beyond Exact Match: How Evaluation Methodology Dominates Model Choice in LLM-Based Product Attribute Extraction
- Dialectical Contradictions in Relationship Development
- Beyond resilients, undercontrollers, and overcontrollers? an extension of personality prototype research
- TrajAudit: Automated Failure Diagnosis for Agentic Coding Systems
- Refining Implicit Argument Annotation for UCCA
- Translation Quality Assessment: A Brief Survey on Manual and Automatic Methods
- Automated Modernization of Machine Learning Engineering Notebooks for Reproducibility
- Sleep Stage Classification Using Bidirectional LSTM in Wearable Multi-sensor Systems
- Co‐occurrence of depression and delinquency in personality types
- ClarifyMT-Bench: Benchmarking and Improving Multi-Turn Clarification for Conversational Large Language Models
- Assessing the Software Security Comprehension of Large Language Models
- Improving ML Training Data with Gold-Standard Quality Metrics
- Competing or Collaborating? The Role of Hackathon Formats in Shaping Team Dynamics and Project Choices
- A Large-Language-Model Framework for Automated Humanitarian Situation Reporting
- DramaBench: A Six-Dimensional Evaluation Framework for Drama Script Continuation
- Understanding Typing-Related Bugs in Solidity Compiler
- When the Gold Standard Isn't Necessarily Standard: Challenges of Evaluating the Translation of User-Generated Content
- Are Vision Language Models Cross-Cultural Theory of Mind Reasoners?
- Towards Deeper Emotional Reflection: Crafting Affective Image Filters with Generative Priors
- OLAF: Towards Robust LLM-Based Annotation Framework in Empirical Software Engineering
- Governance by Evidence: Regulated Predictors in Decision-Tree Models
- Towards Proactive Personalization through Profile Customization for Individual Users in Dialogues
- Thematic Dispersion in Arabic Applied Linguistics: A Bibliometric Analysis using Brookes' Measure
- APT-ClaritySet: A Large-Scale, High-Fidelity Labeled Dataset for APT Malware with Alias Normalization and Graph-Based Deduplication
- Emotion Recognition in Signers
- Vibe Spaces for Creatively Connecting and Expressing Visual Concepts
- LexRel: Benchmarking Legal Relation Extraction for Chinese Civil Cases
- Unveiling Malicious Logic: Towards a Statement-Level Taxonomy and Dataset for Securing Python Packages
- NagaNLP: Bootstrapping NLP for Low-Resource Nagamese Creole with Human-in-the-Loop Synthetic Data
- Journey Before Destination: On the importance of Visual Faithfulness in Slow Thinking
- The Effect of Document Summarization on LLM-Based Relevance Judgments
- Evaluating the Efficacy of Sentinel-2 versus Aerial Imagery in Serrated Tussock Classification
- Decoding Human-LLM Collaboration in Coding: An Empirical Study of Multi-Turn Conversations in the Wild
- Generate-Then-Validate: A Novel Question Generation Approach Using Small Language Models
- Can LLMs Evaluate What They Cannot Annotate? Revisiting LLM Reliability in Hate Speech Detection
- Towards Practical and Usable In-network Classification
- Evaluation of Text Generation: A Survey
- Machine learning for smell: Ordinal odor strength prediction of molecular perfumery components
- Towards a Science of Scaling Agent Systems
- Human– AI collaborative learning in mixed reality: Examining the cognitive and socio‐emotional interactions
- Balanced Accuracy: The Right Metric for Evaluating LLM Judges -- Explained through Youden's J statistic
- LocalSearchBench: Benchmarking Agentic Search in Real-World Local Life Services
- Desktop versus VR for collaborative sensemaking
- CAuSE: Decoding Multimodal Classifiers using Faithful Natural Language Explanation
- Configuration Defects in Kubernetes
- When Do Domain-Specific Foundation Models Justify Their Cost? A Systematic Evaluation Across Retinal Imaging Tasks
- Can gestures speak louder than words? The effect of gestural discourse markers on discourse expectations
- "Dragon Slayer Becomes the Dragon": How Players Perceive and Respond to Inequality in the Game World of Whiteout Survival
- StageGuard: Physiologically Constrained Sleep Staging
- KH-FUNSD: A Hierarchical and Fine-Grained Layout Analysis Dataset for Low-Resource Khmer Business Document
- Learn like a Pathologist: Curriculum Learning by Annotator Agreement for Histopathology Image Classification
- Executable Governance for AI: Translating Policies into Rules Using LLMs
- Systematic Review and Meta-Analysis: Adolescent Depression and Long-Term Psychosocial Outcomes
- Evaluating the Robustness of Large Language Model Safety Guardrails Against Adversarial Attacks
- LAURA: Enhancing Code Review Generation with Context-Enriched Retrieval-Augmented LLM
- CodeDistiller: Automatically Generating Code Libraries for Scientific Coding Agents
- CentaurEval: Benchmarking Human-in-the-Loop Value in Agentic Coding
- CourseTimeQA: A Lecture-Video Benchmark and a Latency-Constrained Cross-Modal Fusion Method for Timestamped QA
- Assessing method agreement for paired repeated binary measurements administered by multiple raters
- A standard protocol for reporting species distribution models
- From Words to Wisdom: Discourse Annotation and Baseline Models for Student Dialogue Understanding
- MTBBench: A Multimodal Sequential Clinical Decision-Making Benchmark in Oncology
- Large Language Models' Complicit Responses to Illicit Instructions across Socio-Legal Contexts
- Language-Independent Sentiment Labelling with Distant Supervision: A Case Study for English, Sepedi and Setswana
- Cross-replication Reliability -- An Empirical Approach to Interpreting Inter-rater Reliability
- Interobserver Agreement Among Sleep Scorers From Different Centers in a Large Dataset
- MUCH: A Multilingual Claim Hallucination Benchmark
- A Diversity-optimized Deep Ensemble Approach for Accurate Plant Leaf Disease Detection
- FinCriticalED: A Visual Benchmark for Financial Fact-Level OCR Evaluation
- Leveraging Digitized Newspapers to Collect Summarization Data in Low-Resource Languages
- A systematic review and meta-analysis of factors associated with anthelmintic resistance in sheep
- Function of language skills in preschooler's problem-solving performance: The role of self-directed speech
- Spiking Neural Networks for Early Prediction in Human Robot Collaboration
- Leveraging Medical Sentiment to Understand Patients Health on Social Media
- Towards Consistent Detection of Cognitive Distortions: LLM-Based Annotation and Dataset-Agnostic Evaluation
- Evaluating consistency of deterministic streamline tractography in\n non-linearly warped DTI data
- An empirical study of Policy-as-Code adoption in open-source software projects
- Exploring the stigma experienced by people affected by Parkinson’s disease: a systematic review
- Exception handling bugs in Python: An empirical study of root causes, fix patterns, and anti-patterns
- Drawing on Education: Using Drawings to Document Schooling and Support Change
- Analyzing Dataset Annotation Quality Management in the Wild
- PRISM of Opinions: A Persona-Reasoned Multimodal Framework for User-centric Conversational Stance Detection
- Towards a Human-in-the-Loop Framework for Reliable Patch Evaluation Using an LLM-as-a-Judge
- Perceive, Act and Correct: Confidence Is Not Enough for Hyperspectral Classification
- AI Annotation Orchestration: Evaluating LLM verifiers to Improve the Quality of LLM Annotations in Learning Analytics
- Exploringand Unleashing the Power of Large Language Models in CI/CD Configuration Translation
- On User Interfaces for Large-Scale Document-Level Human Evaluation of Machine Translation Outputs
- Predicting student outcomes using digital logs of learning behaviors: Review, current standards, and suggestions for future work
- DiagramIR: An Automatic Pipeline for Educational Math Diagram Evaluation
- Concrete images, diverse ideas: The role of pictures in learning from multimedia texts of varying abstractness
- Predictive mapping of forest composition and structure with direct gradient analysis and nearest- neighbor imputation in coastal Oregon, U.S.A.
- Evaluating Language Model Applications for Identifying Solution-Related Content in Issue Report Discussions
- Predicting Length of Stay in the Intensive Care Unit with Temporal\n Pointwise Convolutional Networks
- Who Is the Story About? Protagonist Entity Recognition in News
- Groundwater potential mapping using C5.0, random forest, and multivariate adaptive regression spline models in GIS
- Testing the Testers: Human-Driven Quality Assessment of Voice AI Testing Platforms
- Detecting Silent Failures in Multi-Agentic AI Trajectories
- Language-Enhanced Generative Modeling for Amyloid PET Synthesis from MRI and Blood Biomarkers
- From Pre-labeling to Production: Engineering Lessons from a Machine Learning Pipeline in the Public Sector
- Sustainability of Machine Learning-Enabled Systems: The Machine Learning Practitioner's Perspective
- The Eigenvalues Entropy as a Classifier Evaluation Measure
- Rating Roulette: Self-Inconsistency in LLM-As-A-Judge Frameworks
- TheraMind: A Strategic and Adaptive Agent for Longitudinal Psychological Counseling
- Developing and Validating a Diagnostic Checklist to Assess Argumentation in EFL Writing
- DecoEvo: Score-Decoupled Co-Evolution of Solver and Rubric-Generator Skills in Text Space
- Luxury value perceptions and consumer outcomes: A meta‐analysis
- Multimodal information density is highest in question beginnings, and early entropy is associated with fewer but longer visual signals
- Rashomon Alignment
- The Audit Committee Oversight Process*
- Opinion Mining in Online Reviews About Distance Education Programs
- Target Based Speech Act Classification in Political Campaign Text
- Turn-taking cues in task-oriented dialogue
- RICA: Evaluating Robust Inference Capabilities Based on Commonsense Axioms
- Comparison between methods of vascular calibre characterization and measurement protocols: Influence of vessels number considered
- Modeling water table trends in high-latitude peatlands: Divergent hydrological responses and fire risk implications
- Information Leakage and Performance Overestimation in EEG-Based Schizophrenia Detection: Evidence from Literature and Empirical Analyses
- Mustelid Herpesvirus-2, a Novel Herpes Infection in Northern Sea Otters (Enhydra Lutris Kenyoni)
- Interaction coding in leadership research: A critical review and best-practice recommendations to measure behavior
- Model family selection for classification using Neural Decision Trees
- Chance‐Corrected Interrater Agreement Statistics for Two‐Rater Dichotomous Responses: A Method Review With Comparative Assessment Under Possibly Correlated Decisions
- Handwriting in primary school: comparing standardized tests and evaluating impact of grapho-motor parameters
- Hysteresis in streamflow‐water table relation provides a new classification system of rainfall‐runoff events
- StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents
- GuidedRAG: Semantic Steering of Retrieval-Augmented Generation
- High agreement but low Kappa: I. the problems of two paradoxes
- Assessment of interrater and intrarater reliability of the Fahn–Tolosa–Marin Tremor Rating Scale in essential tremor
- Antibiotic Exposure and Risk of Parkinson's Disease in Finland: A Nationwide Case‐Control Study
- An Intersectional Analysis of Agentic Efforts Individuals Under Community Supervision Describe to Improve Their Lives
- Interobserver agreement in describing adnexal masses using the International Ovarian Tumor Analysis simple rules in a real-time setting and using three-dimensional ultrasound volumes and digital clips
- Intra- and interobserver agreement with regard to describing adnexal masses using International Ovarian Tumor Analysis terminology: reproducibility study involving seven observers
- How Do Therapists Experience 20‐Session Cognitive‐Behavioral Therapy for Anorexia Nervosa ( CBT ‐ AN ‐20)? A Qualitative Study
- Deep attentive spatio-temporal feature learning for automatic resting-state fMRI denoising
- The paradox of paradoxical leadership: A multi-level conceptualization
- Do risk assessment tools help manage and reduce risk of violence and reoffending? A systematic review.
- Muscle size and composition in people with articular hip pathology: a systematic review with meta-analysis
- Textual Indicators of Deliberative Dialogue: A Systematic Review of Methods for Studying the Quality of Online Dialogues
- Does Media Coverage of Partisan Polarization Affect Political Attitudes?
- The Forgotten Margins of AI Ethics
- Understanding partition comparison indices based on counting object pairs
- The motion of trees in the wind: a data synthesis
- Automated identification of hedgerows and hedgerow gaps using deep learning
- Overcoming Climate Gridlock: Perspectives of Climate Leaders on How to Achieve Social Change During Persistent Failure in Australia
- Emotion Ratings: How Intensity, Annotation Confidence and Agreements are Entangled
- Crowdfunding Success Factors: A Meta-Analytic Investigation
- Multi-scale habitat selection modeling: a review and outlook
- Short text classification with machine learning in the social sciences: The case of climate change on Twitter
- Hybrid CNN XGBoost intrusion detection approach tuned by modified sine cosine algorithm towards better cloud security
- Classifying types of gully changes with unoccupied aircraft vehicles 3D multitemporal point clouds for training of satellite data analysis in Northwest Namibia
- The Relation of Ambulatory Blood Pressure and Pulse Rate to Retinopathy in Type 1 Diabetes Mellitus
- WEC: Deriving a Large-scale Cross-document Event Coreference dataset from Wikipedia
- Ripple Effects: Social Turmoil Following Infant Kidnapping Attempts in Wild Geladas
- Ecotrends: an R package for estimating habitat suitability trends over time
- Identifying psychological distress data available in nationally representative surveys: A scoping review and case study of Australian surveys
- Effect of potentially modifiable risk factors associated with myocardial infarction in 52 countries (the INTERHEART study): case-control study
- Deep problems with neural network models of human vision
- Factors Associated with Teacher Wellbeing: A Meta-Analysis
- Predictive habitat distribution models in ecology
- Battle for Inbox and Bucks
- European Blame Games
- Classification of tree species and standing dead trees in Boreal forests using UAV‐based RGB, multispectral, and LiDAR point clouds
- When Knowledge Changes: Metamorphic Testing of RAG Systems with Mutations
- Instructional Coaching as a Tool for Professional Development: Coaches’ Roles and Considerations
- Intercoder Reliability in Qualitative Research: Debates and Practical Guidelines
- A Tertiary lymphoid structures-based pathological score predicts survival and recurrence in colorectal Cancer patients
- Couples and Reproductive Health: A Review of Couple Studies
- Does Google Scholar contain all highly cited documents (1950-2013)?
- Symbol Emergence as an Interpersonal Multimodal Categorization
- Land-cover change in the Kruger to Canyons Biosphere Reserve (1993– 2006): A first step towards creating a conservation plan for the subregion
- An Empirical Study on Deployment Faults of Deep Learning Based Mobile Applications
- Semantic Change Detection with Asymmetric Siamese Networks
- Min-Mid-Max Scaling, Limits of Agreement, and Agreement Score
- The role of neuroticism and extraversion in the stress–anxiety and stress–depression relationships
- Nomogram for sample size calculation on a straightforward basis for the kappa statistic
- Team teaching or solo teaching? Evidence from a crossover experiment on the effects of team teaching on student achievement
- Stimulating language awareness in the foreign language classroom: exploring EFL teaching practices
- Dynamics of affective states during complex learning
- Automated sleep stage identification system based on time–frequency analysis of a single EEG channel and random forest classifier
- Personal and Social Facets of Job Identity: A Person-Centered Approach
- Paraspinal muscle quality in chronic low back pain: a systematic review and meta-analysis of muscle atrophy and fat infiltration
- Software Development During COVID-19 Pandemic: an Analysis of Stack\n Overflow and GitHub
- Biases in the Blind Spot: Detecting What LLMs Fail to Mention
- 5-Hydroxymethylcytosine signatures in cell-free DNA provide information about tumor types and stages
- Advanced statistics: Understanding Medical Record Review (MRR) Studies
- Self-regulation in young children: Is there a role for sociodramatic play?
- Effects of maternal mentalization-related parenting on toddlers’ self-regulation
- Adaptive Data Collection for Latin-American Community-sourced Evaluation of Stereotypes (LACES)
- Global Chlorophyll-a Retrieval algorithm from Sentinel 2 Using Residual Deep Learning and Novel Machine Learning Water Classification
- Towards AI as Colleagues: Multi-Agent System Improves Structured Professional Ideation
- When Good Fit Goes Bad: Identifying and Minimising Overfitting in Ecological Niche Models
- Promoting healthy eating through nudges: Multimodal discursive representations of food and eating in weight‑loss articles in Chinese official WeChat posts
- The Paris 1976 Wine Tastings Revisited Once More: Comparing Ratings of Consistent and Inconsistent Tasters
- Analysis of hotspot areas in China's satellite internet innovation policies and research on policy evolution
- The proteome of the late Middle Pleistocene Harbin individual
- Code Contribution and Credit in Science
- Modeling Hierarchical Thinking in Large Reasoning Models
- A First Look at the Self-Admitted Technical Debt in Test Code: Taxonomy and Detection
- Measure what Matters: Psychometric Evaluation of AI with Situational Judgment Tests
- On making causal claims: A review and recommendations
- Leader humility in Singapore
- ArchISMiner: A Framework for Automatic Mining of Architectural Issue-Solution Pairs from Online Developer Communities
- A meta-analytic review of authentic and transformational leadership: A test for redundancy
- Implicit and explicit processes in phonological concept learning
- “Safer to plant corn and beans”? Navigating the challenges and opportunities of agricultural diversification in the U.S. Corn Belt
- Learning to Triage Taint Flows Reported by Dynamic Program Analysis in Node.js Packages
- RatioWaveNet: A Learnable RDWT Front-End for Robust and Interpretable EEG Motor-Imagery Classification
- Beyond the Surface: Sharenting as a Source of Family Quandaries: Mapping Parents’ Social Media Dilemmas
- The advantages of the Matthews correlation coefficient (MCC) over F1 score and accuracy in binary classification evaluation
- Resting-State Functional MRI in Dyslexia: A Systematic Review
- Patterns of Technical Variation in Chimpanzee Termite Fishing Behavior in Mbam and Djerem National Park, Cameroon
- Visual representations of energy and chemical bonding in biology and chemistry textbooks: A case study of ATP hydrolysis
- Spike detection in the wild: Screening of suspected temporal lobe epilepsy cases using a tailored 2‐channel wearable EEG
- The social studies discourse instrument: Validating an observation tool for classroom discussions
- LexChain: Modeling Legal Reasoning Chains for Chinese Tort Case Analysis
- Enhancing reliability in AI inference services: An empirical study on real production incidents
- Iterative Topic Taxonomy Induction with LLMs: A Case Study of Electoral Advertising
- Speculative Model Risk in Healthcare AI: Using Storytelling to Surface Unintended Harms
- Low-income Latino mothers’ booksharing styles and children's emergent literacy development
- GAPS: A Clinically Grounded, Automated Benchmark for Evaluating AI Clinicians
- Development and Benchmarking of a Blended Human-AI Qualitative Research Assistant
- The Child–Adult Medical Procedure Interaction Scale-Short Form (CAMPIS-SF)
- Getting more out of binary data. Segmenting markets by bagged clustering.
- Lingxi: Repository-Level Issue Resolution Framework Enhanced by Procedural Knowledge Guided Scaling
- Testing and Enhancing Multi-Agent Systems for Robust Code Generation
- A non-invasive machine learning mechanism for early disease recognition on Twitter: The case of anemia
- From literature to biodiversity data: mining arthropod organismal traits with machine learning
- Masculine Republicans and Feminine Democrats: Gender and Americans’ Explicit and Implicit Images of the Political Parties
- Brick: Spatial Capability Routing for the Mixture-of-Models (MoM) Paradigm
- PreprintToPaper dataset: connecting bioRxiv preprints with journal publications
- Deconstruct to Reconstruct a Configurable Evaluation Metric for Open-Domain Dialogue Systems
- Mapping riparian forest fragmentation along the Iori River in Georgia
- Dominance Style Among Macaca thibetana on Mt. Huangshan, China
- Constraint-Guided Unit Test Generation for Machine Learning Libraries
- Investigating the Impact of Rational Dilated Wavelet Transform on Motor Imagery EEG Decoding with Deep Learning Models
- DITING: A Multi-Agent Evaluation Framework for Benchmarking Web Novel Translation
- PyMigTool: a tool for end-to-end Python library migration
- Emotionally Charged, Logically Blurred: AI-driven Emotional Framing Impairs Human Fallacy Detection
- Linguistic Resources for Bhojpuri, Magahi and Maithili: Statistics about\n them, their Similarity Estimates, and Baselines for Three Applications
- Beyond Monolingual Assumptions: A Survey of Code-Switched NLP in the Era of Large Language Models
- The virtual census: representations of gender, race and age in video games
- Review: Adolescents' perspectives on and experiences with post‐primary school‐based suicide prevention as end‐users, co‐creators and peer helpers – a systematic review meta‐ethnography
- Aligning EU policies to address biological invasions: assessing invasion impacts across sectors
- Refactoring with LLMs: Bridging Human Expertise and Machine Understanding
- Person-Centric Annotations of LAION-400M: Auditing Bias and Its Transfer to Models
- Can Emulating Semantic Translation Help LLMs with Code Translation? A Study Based on Pseudocode
- Curiosity-Driven LLM-as-a-judge for Personalized Creative Judgment
- A Note on the Use of Categorical Subscores
- VitaBench: Benchmarking LLM Agents with Versatile Interactive Tasks in Real-world Applications
- An Annotation Scheme for Factuality and its Application to Parliamentary Proceedings
- Feedback Forensics: A Toolkit to Measure AI Personality
- Discovering Self-Regulated Learning Patterns in Chatbot-Powered Education Environment
- Generative Value Conflicts Reveal LLM Priorities
- Towards Reliable Generation of Executable Workflows by Foundation Models
- Identifying deep leverage points to destabilize ‘lock-in’ and empower farmers in the Midwestern agrifood system
- SimuHome: A Temporal- and Environment-Aware Benchmark for Smart Home LLM Agents
- A bot identification model and tool based on GitHub activity sequences
- Beyond kappa: A review of interrater agreement measures
- Moving towards Europe-wide freshwater restoration through model-based integration of policy objectives
- Revisiting Maturity Data: Using Oocyte Diameter and Gonadosomatic Index to Retroactively Apply a New Maturity Scale to Greenland Halibut ( Reinhardtius hippoglossoides )
- Deteção de estruturas permanentes a partir de dados de séries temporais Sentinel 1 e 2
- Building work engagement: A systematic review and meta‐analysis investigating the effectiveness of work engagement interventions
- MASH: A Multiplatform and Multimodal Annotated Dataset for Societal Impact of Hurricane
- Open-DeBias: Toward Mitigating Open-Set Bias in Language Models
- What does it mean to be European? How identity content shapes adolescent's views towards immigrants and support for the EU
- Liaozhai through the Looking-Glass: On Paratextual Explicitation of Culture-Bound Terms in Machine Translation
- Bridging science and society: Developing a citizen science biomonitoring approach for river ecosystems in Italy
- Identification of a 10-species microbial signature of inflammatory bowel disease by machine learning and external validation
- Examining the Influence of Different Inventories on Shallow Landslide Susceptibility Modeling: An Assessment Using Machine Learning and Statistical Approaches
- Comparison between Dense L-Band and C-Band Synthetic Aperture Radar (SAR) Time Series for Crop Area Mapping over a NISAR Calibration-Validation Site
- Could you teach new tricks to old dogs? An analysis of online communication of political leaders in Spanish regional elections
- Do 14–17-month-old infants use iconic speech and gesture cues to interpret word meanings?
- LLMs Behind the Scenes: Enabling Narrative Scene Illustration
- Consider the following: A pilot study of the effects of an educational television program on viewer perceptions of anthropogenic climate change and ocean acidification
- Environmental constraints shaping constituent order in emerging communication systems: Structural iconicity, interactive alignment and conventionalization
- A Machine Learning Framework for Predicting and Understanding the Canadian Drought Monitor
- Personality and prosocial behavior: A theoretical framework and meta-analysis.
- The double‐edged sword of CEO narcissism: A meta‐analysis of innovation and firm performance implications
- Addressing Failing Water Systems: A Qualitative Protocol for Intervention Research on Water Insecurity
- Different gazes and hands: Visual representations of causes and solutions of adolescent depression on Chinese state media social platforms
- Building a Pilot Software Quality-in-Use Benchmark Dataset
- TrueGradeAI: Retrieval-Augmented and Bias-Resistant AI for Transparent and Explainable Digital Assessments
- Do LLM Agents Know How to Ground, Recover, and Assess? A Benchmark for Epistemic Competence in Information-Seeking Agents
- ASSESS: A Semantic and Structural Evaluation Framework for Statement Similarity
- ReviewScore: Misinformed Peer Review Detection with Large Language Models
- Acoustic-based Gender Differentiation in Speech-aware Language Models
- Human-AI Narrative Synthesis to Foster Shared Understanding in Civic Decision-Making
- AutoSpec: An Agentic Framework for Automatically Drafting Patent Specification
- Reverse Engineering User Stories from Code using Large Language Models
- Beyond Similarity: Grounded Agentic Extraction and Expert-Adjudicated Evaluation of Intertextuality in Classical Chinese Histories
- Hallucination‐Free? Assessing the Reliability of Leading AI Legal Research Tools
- A Closer Look at Classification Evaluation Metrics and a Critical Reflection of Common Evaluation Practice
- Gender Discrimination in the Visual Representation of Athletes on the Official Instagram Accounts of Sports Federations in Indonesia
- Land Use and Land Cover Change in the Shafarood Watershed, Northern Iran (2000–2020): A Case Study Using Landsat Imagery and Support Vector Machine Classification
- A systematic review of echo chamber research: comparative analysis of conceptualizations, operationalizations, and varying outcomes
- Challenges in annotations by humans and LLMs: A case study of evaluative language
- Associations of Physician Empathy with Patient Anxiety and Ratings of Communication in Hospital Admission Encounters
- Open Security Benchmark: Towards Autonomous Enterprise Cyber Defense
- Consonant articulation accuracy in paediatric cochlear implant recipients
- GeDi: Simplifying Gene Set Distances for Enhanced Omics Interpretation in R/Bioconductor
- Large language models as first-pass filters for corpus annotation: semantic disambiguation of Galician pobo
- Continued use of retracted papers: Temporal trends in citations and (lack of) awareness of retractions shown in citation contexts in biomedicine
- Subjective assessment of adnexal masses with the use of ultrasonography: an analysis of interobserver variability and experience
- From pixels to peaks: integrating LiDAR and RGB drone imagery to map mussel spat on intertidal rocky shores
- The Chrysalis Effect
- Impact of schematic representations of road maps on pedestrian route planning
- Sentinel-2 satellite image application for establishing a coral reef distribution map in Bai Tu Long National Park, Quang Ninh Province, Vietnam
- Measuring and comparing the accuracy of species distribution models with presence–absence data
- How Sexual Consent is Portrayed in Sex Comics (Eromanga): A Content Analysis in Japan
- Assessing the Cross-Version Applicability of Java Library Vulnerability Exploits
- Comparing Clinical, Microbiological, and Genetic Definitions of Relapse in Patients With Subsequent Episodes of Methicillin-resistant Staphylococcus aureus Bone and Joint Infection
- Impact of interaction with an architectural example on design behavior in student teams
- Ictal emotional features in pediatric and young adult patients with frontal lobe epilepsy
- A custom‐built single‐channel in‐ear electroencephalography sensor for sleep phase detection: an interdependent solution for at‐home sleep studies
- The impact and return-on-investment of evidence-based practice in conservation and environmental management: A machine learning-assisted scoping review protocol
- Annotation of biological samples data to standard ontologies with support from large language models
- A systematic review and meta-analysis of interventions that target the intersection of body image and movement among girls and women
- The value of small forest fragments and urban tree canopy for Neotropical migrant birds during winter and migration seasons in Latin American countries: A systematic review
- Evaluating ChatGPT, Gemini and other Large Language Models (LLMs) in orthopaedic diagnostics: A prospective clinical study
- Integrating climate indices and land use practices for comprehensive drought monitoring in Syria: Impacts and implications
- Reusability of Bayesian Networks case studies: a survey
- The impact of rapport on intelligence yield: police source handler telephone interactions with covert human intelligence sources
- Examining item content across nine psychological (in)flexibility scales: What do they measure?
- The effect of reward value on the performance of long-tailed macaques (Macaca fascicularis) in a delay of gratification exchange task
- Demographically-Inspired Query Variants Using an LLM
- Toward Understanding Free‐Flowing Manual Object Contact: Real‐Time Interaction Between Body Position and Object Type in 9‐Month‐Olds
- Machine learning-based ensemble species distribution models to guide monitoring and survey design for offshore wind
- Evidence on the effects of flame retardant substances at ecologically relevant endpoints: a systematic map protocol
- Clustering longitudinal data: comparison of model-based and distance-based approaches using simulated and real-world data in psychiatric research
- Machine learning predictive modelling for sediment risk indices within an urbanized river channel
- Evaluation of Remotely Sensed Inundation Data Sets to Estimate Flood‐Associated Emergency Department Visits After Hurricane Harvey
- Identifying suitable habitats under climate change for non-targeted demersal fish in the Mediterranean Sea
- Doctoral Education Trends: Content Analyses of Dissertations and Job Postings
- Uncertainty-driven ensembles of deep architectures for multiclass classification. Application to COVID-19 diagnosis in chest X-ray images
- Automatic sampling and training method for wood-leaf classification based on tree terrestrial point cloud
- Semixup: In- and Out-of-Manifold Regularization for Deep Semi-Supervised\n Knee Osteoarthritis Severity Grading from Plain Radiographs
- Telephone versus In-Person Clinical and Health Status Assessment Interviews in Patients with Bipolar Disorder
- Longitudinal Analysis of Discussion Topics in an Online Breast Cancer Community using Convolutional Neural Networks
- Advanced, Analytic, Automated (AAA) Measurement of Engagement During Learning
- MedFact: A Large-scale Chinese Dataset for Evidence-based Medical Fact-checking of LLM Responses
- Robustness and Reliability of Gender Bias Assessment in Word Embeddings: The Role of Base Pairs
- Psychometric properties of the EQ-5D-5L: a systematic review of the literature
- Mental Multi-class Classification on Social Media: Benchmarking Transformer Architectures against LSTM Models
- SciEvent: Benchmarking Multi-domain Scientific Event Extraction
- RulER: Automated Rule-Based Semantic Error Localization and Repair for Code Translation
- Position: Thematic Analysis of Unstructured Clinical Transcripts with Large Language Models
- A Study on Thinking Patterns of Large Reasoning Models in Code Generation
- Development of a Structured Psychiatric Interview for Children: Agreement Between Child and Parent on Individual Symptoms
- Annotating Training Data for Conditional Semantic Textual Similarity Measurement using Large Language Models
- Crash Report Enhancement with Large Language Models: An Empirical Study
- When Large Language Models Meet UAV Projects: An Empirical Study from Developers' Perspective
- Accelerating Discovery: Rapid Literature Screening with LLMs
- Determinants of Entrepreneurial Intent: A Meta–Analytic Test and Integration of Competing Models
- A New Statistical Approach for Comparing Algorithms for Lexicon Based Sentiment Analysis
- A Simple and Efficient Multi-Task Learning Approach for Conditioned Dialogue Generation
- Obsessive-Compulsive Personality Disorder Co-occurring in Individuals with Obsessive-Compulsive Disorder: A Systematic Review and Meta-analysis
- Interactive Task and Concept Learning from Natural Language Instructions and GUI Demonstrations
- Few-NERD: A Few-Shot Named Entity Recognition Dataset
- Measuring Whiteness: A Systematic Review of Instruments and Call to Action
- Ensemble Pruning via Margin Maximization
- LLM-as-a-Judge: Rapid Evaluation of Legal Document Recommendation for Retrieval-Augmented Generation
- An Attention-Based Deep Learning Approach for Sleep Stage Classification With Single-Channel EEG
- TREMO: A dataset for emotion analysis in Turkish
- Auto-Slides: An Interactive Multi-Agent System for Creating and Customizing Research Presentations
- Humanizing Automated Programming Feedback: Fine-Tuning Generative Models with Student-Written Feedback
- My Favorite Streamer is an LLM: Discovering, Bonding, and Co-Creating in AI VTuber Fandom
- Venture capitalists' decision criteria in new venture evaluation
- Textual misinformation on Reddit
- Examining teaching assistant pedagogies in traditional laboratories and recitations
- Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles
- Development of an Instrument to Measure the Perceptions of Adopting an Information Technology Innovation
- An Alternative Measure of Effect Size for Cochran's Q Test for Related Proportions
- RADAR: A Reasoning-Guided Attribution Framework for Explainable Visual Data Analysis
- Fear of the coronavirus (COVID-19): Predictors in an online study conducted in March 2020
- What Were You Thinking? An LLM-Driven Large-Scale Study of Refactoring Motivations in Open-Source Projects
- GitHub Copilot AI pair programmer: Asset or Liability?
- MS2: Multi-Document Summarization of Medical Studies
- From Vision to Validation: A Theory- and Data-Driven Construction of a GCC-Specific AI Adoption Index
- Grader Variability and the Importance of Reference Standards for Evaluating Machine Learning Models for Diabetic Retinopathy
- What if I ask in alia lingua? Measuring Functional Similarity Across Languages
- Towards generalisable hate speech detection: a review on obstacles and solutions
- The Impact of Critique on LLM-Based Model Generation from Natural Language: The Case of Activity Diagrams
- Automatic evaluation of reading aloud performance in children
- Resilients, Overcontrollers, and Undercontrollers: The replicability of the three personality prototypes across informants
- StableSleep: Source-Free Test-Time Adaptation for Sleep Staging with Lightweight Safety Rails
- Understanding Architecture Erosion: The Practitioners' Perceptive
- DynaGuard: A Dynamic Guardian Model With User-Defined Policies
- ABCD-LINK: Annotation Bootstrapping for Cross-Document Fine-Grained Links
- Understanding Code Understandability Improvements in Code Reviews
- Flaw or Artifact? Rethinking Prompt Sensitivity in Evaluating LLMs
- Towards Emotional Support Dialog Systems
- Partial success in closing the gap between human and machine vision
- A Benchmark for Modeling Violation-of-Expectation in Physical Reasoning Across Event Categories
- Using technology in special education: current practices and trends
- Automated Test Validators for Flaky Cyber-Physical System Simulators: Approach and Evaluation
- Automated Quality Assessment for LLM-Based Complex Qualitative Coding: A Confidence-Diversity Framework
- Guidelines for Empirical Studies in Software Engineering involving Large Language Models
- Inference Gap in Domain Expertise and Machine Intelligence in Named Entity Recognition: Creation of and Insights from a Substance Use-related Dataset
- Multimodal Emotion-Cause Pair Extraction in Conversations
- Understanding the role of single‐board computers in engineering and computer science education: A systematic literature review
- Mapping Toxic Comments Across Demographics: A Dataset from German Public Broadcasting
- MAVEN: A Massive General Domain Event Detection Dataset
- Are Companies Taking AI Risks Seriously? A Systematic Analysis of Companies' AI Risk Disclosures in SEC 10-K forms
- M3HG: Multimodal, Multi-scale, and Multi-type Node Heterogeneous Graph for Emotion Cause Triplet Extraction in Conversations
- LiveMCP-101: Stress Testing and Diagnosing MCP-enabled Agents on Challenging Queries
- Toward Responsible ASR for African American English Speakers: A Scoping Review of Bias and Equity in Speech Technology
- Counterspeech for Mitigating the Influence of Media Bias: Comparing Human and LLM-Generated Responses
- Don't Think Twice! Over-Reasoning Impairs Confidence Calibration
- EmoTale: An Enacted Speech-emotion Dataset in Danish
- Temperature-Dependent Evolutionary Speed Shapes the Evolution of Biodiversity Patterns Across Tetrapod Radiations
- Real-Time Classification of Twitter Trends
- Improving K-12 Teachers’ Acceptance of Open Educational Resources by Open Educational Practices: A Mixed Methods Inquiry
- Towards Automatic Bot Detection in Twitter for Health-related Tasks
- Trouble when they walk in? Candidates, dark personality, and attack behavior in German televised debates
- The component structure of memory during development.
- Inter-rater Reliability of the Modified Japanese Orthopedic Association Score in Degenerative Cervical Myelopathy
- An interpretable semi-supervised classifier using two different\n strategies for amended self-labeling
- Defects4Log: Benchmarking LLMs for Logging Code Defect Detection and Reasoning
- Psychedelic therapy for smoking cessation: Qualitative analysis of participant accounts
- Clustering of Social Media Messages for Humanitarian Aid Response during Crisis
- Case study: Mapping potential informal settlements areas in Tegucigalpa\n with machine learning to plan ground survey
- Lost in Moderation: How Commercial Content Moderation APIs Over- and Under-Moderate Group-Targeted Hate Speech and Linguistic Variations
- Sensitivity toward dark matter annihilation imprints on 21-cm signal with SKA-Low: A convolutional neural network approach
- Planner-Refiner: Dynamic Space-Time Refinement for Vision-Language Alignment in Videos
- BharatBBQ: A Multilingual Bias Benchmark for Question Answering in the Indian Context
- The dire disregard of measurement invariance testing in psychological science.
- Alcohol‐cancer risk communication on social media: A content analysis of alcohol‐related Instagram and TikTok posts
- Clinical Utility of the Automatic Phenotype Annotation in Unstructured Clinical Notes: ICU Use Cases
- Examining the Association Between Internet Use and Perceived Stress in Adults: Longitudinal Observational Study Combining Web Tracking Data With Questionnaires
- Demystifying Feature Requests: Leveraging LLMs to Refine Feature Requests in Open-Source Software
- Scaling Success: A Systematic Review of Peer Grading Strategies for Accuracy, Efficiency, and Learning in Contemporary Education
- Understanding Inconsistent State Update Vulnerabilities in Smart Contracts
- Variable selection via knockoffs in missing data settings with categorical predictors
- Do Ethical AI Principles Matter to Users? A Large-Scale Analysis of User Sentiment and Satisfaction
- FineDialFact: A benchmark for Fine-grained Dialogue Fact Verification
- How Do LLMs Persuade? Linear Probes Can Uncover Persuasion Dynamics in Multi-Turn Conversations
- An Overview of 7726 User Reports: Uncovering SMS Scams and Scammer Strategies
- Finding Needles in Images: Can Multimodal LLMs Locate Fine Details?
- An ML-based Approach to Predicting Software Change Dependencies: Insights from an Empirical Study on OpenStack
- I Think, Therefore I Am Under-Qualified? A Benchmark for Evaluating Linguistic Shibboleth Detection in LLM Hiring Evaluations
- Beneath the Surface of the Sexual Harassment Label: A Mixed Methods Study of Young Working Women
- Fostering argumentative knowledge construction through enactive role play in Second Life
- Towards Transparent AI Grading: Semantic Entropy as a Signal for Human-AI Disagreement
- Are Today's LLMs Ready to Explain Well-Being Concepts?
- Explainable Deep Neural Network for Multimodal ECG Signals: Intermediate vs Late Fusion
- From App Features to Explanation Needs: Analyzing Correlations and Predictive Potential
- Argument Quality Annotation and Gender Bias Detection in Financial Communication through Large Language Models
- Psychological safety in software workplaces: A systematic literature review
- Understanding Prediction Discrepancies in Machine Learning Classifiers
- An Effective Entropy-assisted Mind-wandering Detection System with EEG Signals based on MM-SART Database
- Cross-lingual Opinions and Emotions Mining in Comparable Documents
- Understanding environmental tweets of for-profits and nonprofits and their effects on user responses
- PunchPulse: A Physically Demanding Virtual Reality Boxing Game Designed with, for and by Blind and Low-Vision Players
- A Confidence-Diversity Framework for Calibrating AI Judgement in Accessible Qualitative Coding Tasks
- EHSAN: Leveraging ChatGPT in a Hybrid Framework for Arabic Aspect-Based Sentiment Analysis in Healthcare
- Toward Using Machine Learning as a Shape Quality Metric for Liver Point Cloud Generation
- A Methodological Framework for LLM-Based Mining of Software Repositories
- MCeT: Behavioral Model Correctness Evaluation using Large Language Models
- NyayaRAG: Realistic Legal Judgment Prediction with RAG under the Indian Common Law System
- Robust Collaborative Learning of Patch-level and Image-level Annotations for Diabetic Retinopathy Grading from Fundus Image
- BigTokDetect: A Clinically-Informed Vision-Language Modeling Framework for Detecting Pro-Bigorexia Videos on TikTok
- Stakeholder engagement and dialogic accounting
- Helping or Homogenizing? GenAI as a Design Partner to Pre-Service SLPs for Just-in-Time Programming of AAC
- AgriEval: A Comprehensive Chinese Agricultural Benchmark for Large Language Models
- How Growing Toxicity Manifests: A Topic Trajectory Analysis of U.S. Immigration Discourse on Social Media
- Towards Recognizing Phrase Translation Processes: Experiments on English-French
- STARN-GAT: A Multi-Modal Spatio-Temporal Graph Attention Network for Accident Severity Prediction
- Named Entities in Medical Case Reports: Corpus and Experiments
- An Inventory of Preposition Relations
- A mixed method study of DevOps challenges
- A study on cost behaviors of binary classification measures in class-imbalanced problems
- Social perception in negotiation
- Causal relatedness and importance of story events
- What Makes Code Generation Ethically Sourced?
- RoD-TAL: A Benchmark for Answering Questions in Romanian Driving License Exams
- Using Under-trained Deep Ensembles to Learn Under Extreme Label Noise
- Schedule for Affective Disorders and Schizophrenia for School-Age Children-Present and Lifetime Version (K-SADS-PL): Initial Reliability and Validity Data
- ShEMO -- A Large-Scale Validated Database for Persian Speech Emotion\n Detection
- Significance Tests for the Measure of Raw Agreement
- Dutch General Public Reaction on Governmental COVID-19 Measures and Announcements in Twitter Data
- Multimodal Behavioral Patterns Analysis with Eye-Tracking and LLM-Based Reasoning
- Who shapes crisis communication on Twitter? An analysis of influential German-language accounts during the COVID-19 pandemic
- Multi-Year Evaluations of an FTA Card–Based Detection Protocol for Four Vector-Borne Viruses Affecting Potato
- Combining participatory and modeling approaches to investigate factors and drivers of soil erosion risk in mixed crop-livestock farms
- Neural networks with divisive normalization for image segmentation
- Validating telehealth assessment of physical concussion symptoms
- Beyond Accuracy: Ecological Validity of Machine Learning Models for Assessing Vegetation Conservation Classes in South Korea
- Systematic classification differences across eye movement detection algorithms
- Prevalence and Spirometric Transitions of PRISm and Obstruction: A Population‐Based Study
- A systematic review of latent class analysis in psychology: Examining the gap between guidelines and research practice
- Screening for REM Sleep Behaviour Disorder with Minimal Sensors
- CatSIM: A Categorical Image Similarity Metric
- Snow cover dynamics: an overlooked yet important feature of winter bird occurrence and abundance across the United States
- Hierarchical method for cataract grading based on retinal images using improved Haar wavelet
- The neural correlates of pain-related fear: A meta-analysis comparing fear conditioning studies using painful and non-painful stimuli
- Lexico-semantic and affective modelling of Spanish poetry: A semi-supervised learning approach
- What Should I Learn First: Introducing LectureBank for NLP Education and Prerequisite Chain Learning
- Don't forget your classics: Systematizing 45 years of Ancestry for Security API Usability Recommendations
- Review on Requirements Modeling and Analysis for Self-Adaptive Systems: A Ten-Year Perspective
- Evaluating the ability of habitat suitability models to predict species presences
- Policy-Grounded Safety Evaluation of 20 Large Language Models
- Assessing the Reliability of Large Language Models for Deductive Qualitative Coding: A Comparative Study of ChatGPT Interventions
- Assessing Post Deletion in Sina Weibo: Multi-modal Classification of Hot Topics
- From Text to Codings
- Reproducibility of Machine Learning-Based Fault Detection and Diagnosis for HVAC Systems in Buildings: An Empirical Study
- DRAFT-What you always wanted to know but could not find about block-based environments
- CPC-CMS: Cognitive Pairwise Comparison Classification Model Selection Framework for Document-level Sentiment Analysis
- Hallucination Score: Towards Mitigating Hallucinations in Generative Image Super-Resolution
- On the Effectiveness of LLM-as-a-judge for Code Generation and Summarization
- A framework for reliable traffic surrogate safety assessment based on multi-object tracking data
- A protocol for evaluating AI chatbots’ capabilities for low-resource language teachers
- Use of artificial intelligence to support the assessment of the methodological quality of systematic reviews
- Advancing Mental Disorder Detection: A Comparative Evaluation of Transformer and LSTM Architectures on Social Media
- Automatically assessing oral narratives of Afrikaans and isiXhosa children
- Towards Better Requirements from the Crowd: Developer Engagement with Feature Requests in Open Source Software
- Content Analysis in Mass Communication: Assessment and Reporting of Intercoder Reliability
- Temporal Stability of Heavy Drinking Days and Drinking Reductions Among Heavy Drinkers in the COMBINE Study
- Diagnostic reliability of the Semi-structured Assessment for Drug Dependence and Alcoholism (SSADDA)
- ANLIzing the Adversarial Natural Language Inference Dataset
- CD2CR: Co-reference Resolution Across Documents and Domains
- Meta‐analysis reveals that the effects of precipitation change on soil and litter fauna in forests depend on body size
- Novel Antischistosomal Drug Targets: Identification of Alkaloid Inhibitors of SmTGR via Integrated In Silico Methods
- Concordance between FVC and FEV 6 for identifying chronic airflow obstruction and spirometric restriction in the Burden of Obstructive Lung Disease (BOLD) study
- Macular OCT Classification Using a Multi-Scale Convolutional Neural Network Ensemble
- Security Smells in Ansible and Chef Scripts: A Replication Study
- A nonparametric Bayesian test of dependence
- Development of a qualitative data analysis codebook for peri‐ictal behavior in suspected functional seizures
- Tracking Sumatran Tiger (Panthera tigris sumatrae Pocock, 1929) distribution in Gunung Leuser National Park: The influence of prey presence and environmental variables on habitat selection
- Evaluation of Facebook as a Longitudinal Data Source for Parkinson’s Disease Insights
- A Dataset of General-Purpose Rebuttal
- Ranking Computer Vision Service Issues using Emotion
- Identifying Morality Frames in Political Tweets using Relational Learning
- Empirical studies of agile software development: A systematic review
- Mental Models of Adversarial Machine Learning
- Internal Value Alignment in Large Language Models through Controlled Value Vector Activation
- A training programme for early-stage researchers that focuses on\n developing personal science outreach portfolios
- Dr.Copilot: A Multi-Agent Prompt Optimized Assistant for Improving Patient-Doctor Communication in Romanian
- The Eighth Dialog System Technology Challenge
- Is Quantization a Deal-breaker? Empirical Insights from Large Code Models
- THAI Speech Emotion Recognition (THAI-SER) corpus
- From Research to Resources: Assessing Student Understanding and Skills in Quantum Computing
- ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models
- SPICE: An Automated SWE-Bench Labeling Pipeline for Issue Clarity, Test Coverage, and Effort Estimation
- Enhancing Essay Cohesion Assessment: A Novel Item Response Theory Approach
- Measuring Hypothesis Testing Errors in the Evaluation of Retrieval Systems
- Quantifying Uncertainty in Error Consistency: Towards Reliable Behavioral Comparison of Classifiers
- A proposal and assessment of an improved heuristic for the Eager Test smell detection
- EduCoder: An Open-Source Annotation System for Education Transcript Data
- Improving Label Quality by Jointly Modeling Items and Annotators
- A versatile index to characterize hysteresis between hydrological variables at the runoff event timescale
- WSCoach: Wearable Real-time Auditory Feedback for Reducing Unwanted Words in Daily Communication
- SymbolicThought: Integrating Language Models and Symbolic Reasoning for Consistent and Interpretable Human Relationship Understanding
- A Structured Learning Approach with Neural Conditional Random Fields for\n Sleep Staging
- Leveraging Information Technology Infrastructure to Facilitate a Firm's Customer Agility and Competitive Activity: An Empirical Investigation
- Intraclass Correlation Coefficient (ICC): A Framework for Monitoring and Assessing Performance of Trained Sensory Panels and Panelists
- Weakly Supervised Learning of Nuanced Frames for Analyzing Polarization in News Media
- An Investigation of Posttraumatic Stress Disorder and Depressive Symptomatology among Female Victimsof Interpersonal Trauma
- Sustainability Flags for the Identification of Sustainability Posts in Q&A Platforms
- Predicting the Transition from Short-term to Long-term Memory based on Deep Neural Network
- MedVAL: Toward Expert-Level Medical Text Validation with Language Models
- Sentinel SAR-optical fusion for crop type mapping using deep learning and Google Earth Engine
- Enhancing COBOL Code Explanations: A Multi-Agents Approach Using Large Language Models
- HCqa: Hybrid and Complex Question Answering on Textual Corpus and Knowledge Graph
- Evaluating the Effectiveness of Direct Preference Optimization for Personalizing German Automatic Text Simplifications for Persons with Intellectual Disabilities
- MRI can assess glenoid bone loss after shoulder luxation: inter- and intra-individual comparison with CT
- Recommending Variable Names for Extract Local Variable Refactorings
- Relevance of Unsupervised Metrics in Task-Oriented Dialogue for Evaluating Natural Language Generation
- Evaluating Lexicon-Based Sentiment Analysis Methods for Small Datasets on Low Compute Devices
- Development of a Framework for Youth- and Family-Specific Engagement in Research: Proposal for a Scoping Review and Qualitative Descriptive Study
- On Recipe Memorization and Creativity in Large Language Models: Is Your Model a Creative Cook, a Bad Cook, or Merely a Plagiator?
- Climate change-induced range shift of the endemic epiphytic lichenLobaria pindarensisin the Hindu Kush Himalayan region
- Learning from various labeling strategies for suicide-related messages on social media: An experimental study
- Evaluating and Improving Large Language Models for Competitive Program Generation
- On the limits of cross-domain generalization in automated X-ray prediction
- Accountability in Intervention Research: Developing a Fidelity Checklist of a Mental Health Intervention in Prisons
- Domain Knowledge in Artificial Intelligence: Using Conceptual Modeling to Increase Machine Learning Accuracy and Explainability
- Semantic Caching for Improving Web Affordability
- The International Caries Detection and Assessment System (ICDAS): an integrated system for measuring dental caries
- Magnetic Resonance Classification of Lumbar Intervertebral Disc Degeneration
- The reliability of the pass/fail decision for assessments comprised of multiple components
- See-in-Pairs: Reference Image-Guided Comparative Vision-Language Models for Medical Diagnosis
- Predictive modelling of football injuries
- LAION-C: An Out-of-Distribution Benchmark for Web-Scale Vision Models
- Identifying Explanation Needs: Towards a Catalog of User-based Indicators
- Understanding the Challenges and Promises of Developing Generative AI Apps: An Empirical Study
- PRISON: Unmasking the Criminal Potential of Large Language Models
- Specificity-Based Sentence Ordering for Multi-Document Extractive Risk Summarization
- Detecting Online Hate Speech Using Context Aware Models
- Essential-Web v1.0: 24T tokens of organized web data
- Reciprocal Relationships Between Parenting Behavior and Disruptive Psychopathology from Childhood Through Adolescence
- Accuracy assessment of land cover maps [wikipedia]
- Validity/reliability of PHQ-9 and PHQ-2 depression scales among adults living with HIV/AIDS in western Kenya. [europepmc]
- Radiographic evaluation of the hip has limited reliability. [europepmc]
- Reliability of a complication classification system for orthopaedic surgery. [europepmc]
- A comparison of Cohen's Kappa and Gwet's AC1 when calculating inter-rater reliability coefficients: a study conducted with personality disorder samples. [europepmc]
- Multidimensional in vivo hazard assessment using zebrafish. [europepmc]
- A review and meta-analysis of age-based stereotype threat: negative stereotypes, not facts, do the damage. [europepmc]
- Pictorial cigarette pack warnings: a meta-analysis of experimental studies. [europepmc]
- Metrics for evaluating 3D medical image segmentation: analysis, selection, and tool. [europepmc]
- Can mental health diagnoses in administrative data be used for research? A systematic review of the accuracy of routinely collected diagnoses. [europepmc]
- Utility of spherical human liver microtissues for prediction of clinical drug-induced liver injury. [europepmc]
- The relationship between physician burnout and quality of healthcare in terms of safety and acceptability: a systematic review. [europepmc]
- Empirical Comparison of Publication Bias Tests in Meta-Analysis. [europepmc]
- Best Practices for Developing and Validating Scales for Health, Social, and Behavioral Research: A Primer. [europepmc]
- Automated Gleason grading of prostate cancer tissue microarrays via deep learning. [europepmc]
- Agreement Between Prospective and Retrospective Measures of Childhood Maltreatment: A Systematic Review and Meta-analysis. [europepmc]
- Metascape provides a biologist-oriented resource for the analysis of systems-level datasets. [europepmc]
- Multitask learning and benchmarking with clinical time series data. [europepmc]
- Development and validation of a deep learning algorithm for improving Gleason scoring of prostate cancer. [europepmc]
- Understanding the care and support needs of older people: a scoping review and categorisation using the WHO international classification of functioning, disability and health framework (ICF). [europepmc]
- The advantages of the Matthews correlation coefficient (MCC) over F1 score and accuracy in binary classification evaluation. [europepmc]
- Promoting memory consolidation during sleep: A meta-analysis of targeted memory reactivation. [europepmc]
- Automatic diagnosis of the 12-lead ECG using a deep neural network. [europepmc]
- Fear of the coronavirus (COVID-19): Predictors in an online study conducted in March 2020. [europepmc]
- The Matthews correlation coefficient (MCC) is more reliable than balanced accuracy, bookmaker informedness, and markedness in two-class confusion matrix evaluation. [europepmc]
- Towards a guideline for evaluation metrics in medical image segmentation. [europepmc]