Abstract
Artificial intelligence (AI)–based language models are increasingly explored as tools for interpreting and applying clinical guideline recommendations. In urology, the European Association of Urology (EAU) recently introduced a guideline-specific chatbot; however, its comparative performance relative to contemporary general-purpose large language models (LLMs) remains unclear. In this structured comparative study, five AI systems—the EAU Guidelines Bot, ChatGPT-5, Gemini 2.5 Pro, Copilot – Smart GPT-5, and Perplexity Pro—were evaluated using 13 clinical questions derived directly from strongly recommended statements in the EAU erectile dysfunction (ED) guidelines. Responses were independently assessed by three senior reviewers across five predefined domains: relevance, clarity, structure, clinical utility, and factual accuracy, using a 5-point Likert scale. The primary outcome of the study was the composite performance score, which was calculated as the mean of the five domain scores. Inter-rater reliability was calculated using ICC(2,k), and differences among models were analyzed with the Friedman test followed by Holm-adjusted Wilcoxon post-hoc comparisons. Significant performance differences were observed across all domains (all p < 0.001). The highest composite scores were observed for Gemini 2.5 Pro [4.60 (4.40–4.73)] and the EAU Guidelines Bot [4.53 (4.47–4.80)], followed by ChatGPT-5 [4.27 (4.07–4.47)]. Lower composite scores were observed for Copilot – Smart GPT-5 [3.73 (3.40–3.87)] and Perplexity Pro [3.60 (3.47–3.80)]. Domain-level analysis showed consistently high median scores (≥ 4) for factual accuracy among top-performing models, whereas variability was more pronounced in clarity, structure, and clinical utility. These findings suggest that both guideline-specific systems and advanced general-purpose LLMs may generate responses broadly consistent with guideline-based recommendations in structured ED scenarios. However, variability across domains—particularly in structure and clinical utility—and modest differences in composite performance suggest that these models should be interpreted as supportive tools rather than definitive clinical decision-making systems, requiring further validation in real-world settings.
This is a preview of subscription content, access
Access options
- Purchase on SpringerLink
- Instant access to the full article PDF.
Prices may be subject to local taxes which are calculated during checkout
Data availability
Data are available from the corresponding author on reasonable request.
References
-
Impotence: NIH Consensus Development Panel on Impotence. JAMA. 1993;270:83–90. https://doi.org/10.1001/JAMA.1993.03510010089036.
-
McKinlay JB. The worldwide prevalence and epidemiology of erectile dysfunction. Int J Impot Res. 2000;12:S6–11. https://doi.org/10.1038/SJ.IJIR.3900567.
-
Salonia, Capogrosso A, Boeri P, Cocci L, Corona G A, Dinkelman-Smit M, et al. European Association of Urology guidelines on male sexual and reproductive health: 2025 update on male hypogonadism, erectile dysfunction, premature ejaculation, and Peyronie’s disease. Eur Urol. 2025;88:76–102. https://doi.org/10.1016/J.EURURO.2025.04.010.
-
Hussain W, Khoriba G, Maity S, Jyoti Saikia M. Large language models in healthcare and medical applications: a review. Bioengineering. 2025;12:631. https://doi.org/10.3390/BIOENGINEERING12060631.
-
Baturu M, Solakhan M, Kazaz TG, Bayrak O. Frequently asked questions on erectile dysfunction: evaluating artificial intelligence answers with expert mentorship. Int J Impot Res. 2024;37:310–4. https://doi.org/10.1038/s41443-024-00898-3.
-
EAU Guidelines Bot. Guideline Bot – Uroweb 2025. https://uroweb.org/chat//6704965 (accessed September 15, 2025).
-
Razdan S, Siegal AR, Brewer Y, Sljivich M, Valenzuela RJ. Assessing ChatGPT’s ability to answer questions pertaining to erectile dysfunction: can our patients trust it?. Int J Impot Res. 2023;36:734–40. https://doi.org/10.1038/s41443-023-00797-z.
-
Şahin MF, Ateş H, Keleş A, Özcan R, Doğan Ç, Akgül M, et al. Responses of five different artificial intelligence chatbots to the top searched queries about erectile dysfunction: a comparative Analysis. J Med Syst. 2024;48:38. https://doi.org/10.1007/S10916-024-02056-0.
-
ChatGPT version 5.0 [Internet]. OpenAI [cited 2025 Sep 15]. Available from: https://chatgpt.com/
-
Gemini 2.5 Pro [Internet]. Google DeepMind [cited 2025 Sep 15]. Available from: https://gemini.google.com/
-
Copilot (GPT-5-based) [Internet]. GitHub [cited 2025 Sep 15]. Available from: https://copilot.microsoft.com/
-
Perplexity Pro [Internet]. Perplexity AI [cited 2025 Sep 15]. Available from: https://www.perplexity.ai/
-
Joshi A, Kale S, Chandel S, Pal D. Likert scale: explored and explained. Br J Appl Sci Technol. 2015;7:396–403. https://doi.org/10.9734/BJAST/2015/14975.
-
Koo TK, Li MY. A guideline of selecting and reporting intraclass correlation coefficients for reliability research. J Chiropr Med. 2016;15:155–63. https://doi.org/10.1016/J.JCM.2016.02.012.
-
Caglar U, Yildiz O, Meric A, Ayranci A, Gelmis M, Sarilar O, et al. Evaluating the performance of ChatGPT in answering questions related to pediatric urology. J Pediatr Urol. 2024;20:26.e1–26.e5. https://doi.org/10.1016/j.jpurol.2023.08.003.
-
Zhu L, Mou W, Chen R. Can the ChatGPT and other large language models with internet-connected database solve the questions and concerns of patient with prostate cancer and help democratize medical knowledge? J Transl Med. 2023;21:269. https://doi.org/10.1186/s12967-023-04123-5.
-
Malak A, Şahin MF. How useful are current chatbots regarding urology patient information? Comparison of the ten most popular chatbots’ responses about female urinary incontinence. J Med Syst. 2024;48:102. https://doi.org/10.1007/s10916-024-02125-4.
-
May M, Körner-Riffard K, Kollitsch L, Burger M, Brookman-May SD, Rauchenwald M, et al. Evaluating the efficacy of AI chatbots as tutors in urology: a comparative analysis of responses to the 2022 in-service assessment of the European Board of Urology. Urol Int. 2024;108:359–66. https://doi.org/10.1159/000537854.
-
Karches KE. Against the iDoctor: why artificial intelligence should not replace physician judgment. Theor Med Bioeth. 2018;39:91–110. https://doi.org/10.1007/s11017-018-9442-3.
-
Lebovitz S, Lifshitz-Assaf H, Levina N. To engage or not to engage with AI for critical judgments: how professionals deal with opacity when using AI for medical diagnosis. Organ Sci. 2022;33:126–48. https://doi.org/10.1287/ORSC.2021.1549.
-
Almada M, Petit N. The EU AI Act: between the rock of product safety and the hard place of fundamental rights. Common Mark Law Rev. 2025;62:85–120. https://doi.org/10.54648/cola2025004.
Acknowledgements
The authors thank all participants who contributed to this study.
Funding
The authors received no <a href="https://bitcomme.com/release-of-marimekkos-half-year-financial-report-1-january-30-june-2026/” title=”Release of Marimekko's Half-year Financial Report, 1 January–30 June 2026″>financial support for the research, authorship, and/or publication of this article.
Authors and Affiliations
Consortia
Contributions
GÇ conceptualized and designed the study. GÇ and AM performed the primary data analysis and drafted the manuscript. OA conducted the statistical analyses. GIR and MF contributed to data acquisition and provided clinical input. EGR, CM, OA, and RR revised the manuscript for important intellectual content. All authors reviewed and approved the final version of the manuscript.
Ethics declarations
Competing interests
Marco Falcone is a Deputy Editor of the International Journal of Impotence Research. Afonso Morgado is an Associate Editor of the International Journal of Impotence Research. The other authors declare no competing interests.
Ethical approval
Ethical approval was not required for this study as it did not involve human participants, animals, or identifiable patient data.
Additional information
Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Supplementary information
Supplementary 2- Holm-adjusted pairwise comparisons (p-values) between AI models across evaluation (download PDF )
Rights and permissions
Springer Nature or its licensor (e.g. a society or other partner) holds exclusive rights to this article under a publishing agreement with the author(s) or other rightsholder(s); author self-archiving of the accepted manuscript version of this article is solely governed by the terms of such publishing agreement and applicable law.
About this article
Cite this article
Çeker, G., Morgado, A., Russo, G.I. et al. Quality assessment of artificial intelligence responses in erectile dysfunction: a comparative study based on EAU recommendations.
Int J Impot Res (2026). https://doi.org/10.1038/s41443-026-01339-z
-
Version of record:07 August 2026
-
DOI
:https://doi.org/10.1038/s41443-026-01339-z
