2025/08/14 by Karl‐Patrik Kresoja, Anne Rebecca Schöber, Thomas F. Lüscher +7 · 1 voice
Computer Science · Medicine · #Artificial Intelligence in Healthcare and Education #Explainable Artificial Intelligence (XAI) #Topic Modeling
paper · pdf · doi:10.1093/eurheartj/ehaf654
openalex publication_date 2025/08/14 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Artificial intelligence (AI) is rapidly evolving and has significant potential to impact modern medicine. Recognizing its impact, international societies, including the European Society of Cardiology (ESC), have launched initiatives such as the committee on Digital Cardiology and Artificial Intelligence1 and, more recently, ESCChat, an AI-enhanced Chatbot containing all ESC Clinical Practice Guidelines. AI-based large language models, with ChatGPT being among the most well-known, are set to transform the way medical literature is written and reviewed.2 The use of ChatGPT in scientific writing has sparked debate and critical discussion,3 prompting major publishers to implement specific policies.4 However, like any innovation, AI can not only be helpful, but also be misused, for instance, to generate and disseminate misinformation or fraudulent scientific work, as seen with paper mills.5 It is excessively difficult even for human experts to distinguish between human written abstracts based on true data or large language-based models such as ChatGPT-manufactured abstracts based on imaginary data. We hypothesized that AI-generated abstracts would not receive lower ratings than human-written abstracts in a real-world setting where reviewers were unaware of their presence and had no prior knowledge of the baseline rate of fraudulent submissions. Additionally, we hypothesized that even the highest-rated AI-generated abstracts could be identified using dedicated detection software. To test this, we aimed to generate ∼10% of all submitted abstracts for the German Cardiac Society Annual Congress using AI. Historical submission data were analyzed to ensure proportional representation across subcategories. The ChatGPT-4o web interface was programmed using a hybrid of ‘instruction’ and ‘few-shot’ prompting to generate fabricated abstracts and data with predefined word counts, structures, and appropriate categories according to abstract submission instructions. Fabricated randomized controlled trials were prohibited from creation to avoid ethical concerns and prevent hypothetical results from being misinterpreted or reported as factual. The AI-generated content, including wording and data, was not modified by the authors. Additionally, ChatGPT’s integrated Python tool was used to generate figures for 50% of the AI-generated abstracts. More details on prompting and all generated abstracts can be found at https://zenodo.org/records/16489919. These abstracts were submitted alongside regular submissions, with all reviewers involved in the abstract assessment blinded to the inclusion of AI-generated content and unaware of the experimental design. After evaluation, all AI-generated abstracts were immediately retracted to prevent any influence on the congress proceedings. The primary outcome measure was the abstract rating on a scale from 1 (lowest) to 5 (highest). Furthermore, the licensed ‘essential’ version of GPTZero (https://gptzero.me), a tool designed to detect AI-generated content, was used to assess the probability of AI-generated text in the 12 highest-rated AI abstracts, as well as in 12 human-authored abstracts that were selected for a ‘Young Investigator Award’ competition during the annual congress in 2025 based on reviewer grading and 12 from 2022 as absolute controls for the same categories, as ChatGPT was not released at the time of abstract submission deadline in 2022. Across 19 categories, a total of 1348 abstracts were submitted. As part of this experiment of those 136 (10%) were deliberately generated by the research team using ChatGPT-4o. Overall, there was no significant difference in ratings between human-authored and AI-generated abstracts [human median 3.3 (IQR 3.0 to 3.7), AI median 3.4, (IQR 2.9 to 3.7); P = .85 Mann–Whitney test, Figure 1). The inclusion of fabricated figures in AI-generated abstracts did not significantly affect their ratings compared with those without figures (P = .57). Performance of artificial intelligence fabricated vs human-authored abstracts. The top panel shows design, findings and potential solution for the presented study which investigated the artificial intelligence vs human-authored abstracts. The table shows the performance of the ChatGPT generated abstracts according to subcategories, with the total number of submitted abstracts, the number and fraction of ChatGPT submitted abstracts, the highest ranked ChatGPT abstract, the lowest ranked ChatGPT abstract and whether ChatGPT or humans had significantly higher scores on overall evaluation. Lastly medals indicate whether ChatGPT received first (gold), second (silver) or third (bronze) best abstracts in the respective categories. Comparison of parametric values was done with Mann–Whitney test. Box plots represent median and corresponding interquartile range, whiskers represent respective minimal and maximal values. NA* no statistical comparison performed due to low number of observations (i.e. n = 2) Performance varied by category: AI-generated abstracts were rated lower than human-authored ones only in rhythmology (P < .001), whereas they were rated significantly higher in cardiovascular imaging (P = .011). AI-generated abstracts ranked highest in six categories, second-highest in three categories, and third-highest in four categories (Figure 1). Two AI-generated abstracts were flagged by reviewers due to suspected prior publication, but no evidence of comparable studies was found. No other AI-generated abstract was flagged as suspicious. GPTZero effectively distinguished AI-generated abstracts from human-authored ones, assigning a significantly higher probability to AI-generated texts [AI: median 100%, 95% confidence interval (CI) 98% to100% vs humans 2025: median 7%, 95% CI: 1% to 78%; P < .001]. Receiver operating characteristic analysis demonstrated excellent discrimination, with an area under the curve of 0.97 (95% CI: .91–1.0; P < .001), a sensitivity of 90% (95% CI: 74% to 100%), and a specificity of 92% (95% CI: 62% to 99%) at a Youden-Index derived GPTZero probability threshold of ≥90%. There was a trend for an increase in the AI probability prediction for human written abstracts from 2022 to 2025 (2022: median 1%, 95% CI 1%–5% vs 2025: median 7%, 95% CI: 1%–78%; P = .078). The main findings of this study indicate that human reviewers, when left uninformed, did not rate AI-generated abstracts differently from human-authored ones, demonstrating neither overall inferiority nor superiority. However, a dedicated AI detection software was effective in identifying fabricated abstracts, even among those ranked highest by human reviewers. Our results, though concerning, align with prior research showing that human reviewers struggle to identify AI-generated content, even when explicitly asked to do so.5 In contrast, detection software proved highly effective, confirming earlier findings.2,6 Using AI to generate or refine scientific texts is not inherently problematic. However, lack of disclosure,4 and risks of misuse raise serious ethical concerns.5 Generative AI depends on existing literature, and unchecked fabricated content could trigger a snowball effect: spreading misinformation, undermining reproducibility, and potentially endangering patient safety.7,8 These issues highlight the urgent need for a global AI strategy to uphold scientific integrity, as in the end submitting authors are responsible for the scientific accuracy.1 This study's limitations include the absence of a reporting option for suspected AI content and congress policies that do not yet mandate AI disclosure. Further, this study addresses the application and detection of generative AI in the context of scientific publishing and fraud detection, which may extend beyond medical domains. AI-generated abstracts achieved comparable reviewer scores and acceptance rates to human-authored submissions, underscoring the challenge of detecting fraudulent content. These findings emphasize the need for mandatory AI disclosure policies and the implementation of reliable detection tools to safeguard research integrity in the future. As a key takeaway from this study, scientific congresses, societies and journals may wish to explore clearer guidance regarding the disclosure of AI-assisted writing tools during abstract submission. The integration of AI-detection software, such as GPTZero, could be considered as a supportive measure to enhance transparency during the submission and review process. While effective, its accessibility may be limited by licensing, and the software primarily analyses syntactic and not semantic features. While such tools may help identify AI-generated content, decisions around exclusion should be approached with caution and informed by further validation studies. Reviewers, in turn, could benefit from awareness of any preliminary screening measures. Finally, larger datasets will be needed to determine whether variations in abstract evaluation are consistent across different thematic areas. None declared. Supplementary data are not available at European Heart Journal online. ChatGPT, was used for minor language editing (shortening) of this manuscript. All scientific content, analyses, discussion and interpretation were developed by the authors. K.-P.K. is a consultant to Edwards Lifesciences. T.L. does no longer accept any honoraria from industry, but outside this work received research and educational grants to the institution from Abbott, Amgen, AstraZeneca, Boehringer-Ingelheim, Daichi-Sankyo, Menarini Foundation, Novartis, Novo Nordisk, Roche Diagnostics, Sanofi and Vifor. All other authors report no conflicts of interest. Data sharing is possible based on a reasonable request (e.g. for individual data metaanalyses). All authors declare no funding for this contribution. Ethical approval was not required. None supplied.