Evaluating the image recognition capabilities of GPT-4V and Gemini Pro in the Japanese national dental examination

IF 3.1 3区医学 Q1 DENTISTRY, ORAL SURGERY & MEDICINE Journal of Dental Sciences Pub Date : 2025-01-01 Epub Date: 2024-07-02 DOI:10.1016/j.jds.2024.06.015

Hikaru Fukuda , Masaki Morishita , Kosuke Muraoka , Shino Yamaguchi , Taiji Nakamura , Izumi Yoshioka , Shuji Awano , Kentaro Ono

{"title":"Evaluating the image recognition capabilities of GPT-4V and Gemini Pro in the Japanese national dental examination","authors":"Hikaru Fukuda , Masaki Morishita , Kosuke Muraoka , Shino Yamaguchi , Taiji Nakamura , Izumi Yoshioka , Shuji Awano , Kentaro Ono","doi":"10.1016/j.jds.2024.06.015","DOIUrl":null,"url":null,"abstract":"<div><h3>Background/purpose</h3><div>OpenAI's GPT-4V and Google's Gemini Pro, being Large Language Models (LLMs) equipped with image recognition capabilities, have the potential to be utilized in future medical diagnosis and treatment, ands serve as valuable educational support tools for students. This study compared and evaluated the image recognition capabilities of GPT-4V and Gemini Pro using questions from the Japanese National Dental Examination (JNDE) to investigate their potential as educational support tools.</div></div><div><h3>Materials and methods</h3><div>We analyzed 160 questions from the 116th JNDE, administered in March 2023, using ChatGPT-4V, and Gemini Pro, which have image recognition functions. Standardized prompts were used for all LLMs, and statistical analysis was conducted using Fisher's exact test and the Mann–Whitney U test.</div></div><div><h3>Results</h3><div>For the 160 JNDE questions, the accuracy rates of GPT-4V and Gemini Pro were 35.0% and 28.1%, respectively, with GPT-4V being the highest, although not statistically significant. Across dental specialties, the accuracy rates of the GPT-4V were generally higher than those of the Gemini Pro, with some areas showing equal accuracy. Accuracy rates tended to decrease with an increased number of images within a question, suggesting that the number of images influenced the correctness of the responses.</div></div><div><h3>Conclusion</h3><div>The overall superior performance of GPT-4V compared to Gemini Pro may be attributed to the continuous updates in OpenAI's model. This research demonstrates the potential of LLMs as educational support tools in dentistry, while also highlighting areas that require further technological development.</div></div>","PeriodicalId":15583,"journal":{"name":"Journal of Dental Sciences","volume":"20 1","pages":"Pages 368-372"},"PeriodicalIF":3.1000,"publicationDate":"2025-01-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Journal of Dental Sciences","FirstCategoryId":"3","ListUrlMain":"https://www.sciencedirect.com/science/article/pii/S1991790224002125","RegionNum":3,"RegionCategory":"医学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"2024/7/2 0:00:00","PubModel":"Epub","JCR":"Q1","JCRName":"DENTISTRY, ORAL SURGERY & MEDICINE","Score":null,"Total":0}

引用次数: 0

Abstract

Background/purpose

OpenAI's GPT-4V and Google's Gemini Pro, being Large Language Models (LLMs) equipped with image recognition capabilities, have the potential to be utilized in future medical diagnosis and treatment, ands serve as valuable educational support tools for students. This study compared and evaluated the image recognition capabilities of GPT-4V and Gemini Pro using questions from the Japanese National Dental Examination (JNDE) to investigate their potential as educational support tools.

Materials and methods

We analyzed 160 questions from the 116th JNDE, administered in March 2023, using ChatGPT-4V, and Gemini Pro, which have image recognition functions. Standardized prompts were used for all LLMs, and statistical analysis was conducted using Fisher's exact test and the Mann–Whitney U test.

Results

For the 160 JNDE questions, the accuracy rates of GPT-4V and Gemini Pro were 35.0% and 28.1%, respectively, with GPT-4V being the highest, although not statistically significant. Across dental specialties, the accuracy rates of the GPT-4V were generally higher than those of the Gemini Pro, with some areas showing equal accuracy. Accuracy rates tended to decrease with an increased number of images within a question, suggesting that the number of images influenced the correctness of the responses.

Conclusion

The overall superior performance of GPT-4V compared to Gemini Pro may be attributed to the continuous updates in OpenAI's model. This research demonstrates the potential of LLMs as educational support tools in dentistry, while also highlighting areas that require further technological development.

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

评估 GPT-4V 和 Gemini Pro 在日本全国牙科考试中的图像识别能力

背景/目的openai的GPT-4V和谷歌的Gemini Pro是具有图像识别功能的大型语言模型（llm），具有在未来医学诊断和治疗中使用的潜力，并可作为学生宝贵的教育支持工具。本研究比较并评估了GPT-4V和Gemini Pro的图像识别能力，并使用日本国家牙科考试（JNDE）中的问题来调查它们作为教育支持工具的潜力。材料和方法我们使用ChatGPT-4V和Gemini Pro分析了2023年3月参加的第116届JNDE考试的160个问题，这些问题具有图像识别功能。所有llm均采用标准化提示，采用Fisher精确检验和Mann-Whitney U检验进行统计分析。结果在160个JNDE题目中，GPT-4V和Gemini Pro的准确率分别为35.0%和28.1%，其中GPT-4V的准确率最高，但无统计学意义。在牙科专业中，GPT-4V的准确率通常高于Gemini Pro，在某些领域显示出相同的准确性。正确率往往随着问题中图像数量的增加而降低，这表明图像数量影响了回答的正确性。结论GPT-4V的整体性能优于Gemini Pro可能与OpenAI模型的不断更新有关。这项研究证明了法学硕士作为牙科教育支持工具的潜力，同时也强调了需要进一步技术发展的领域。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文去求助

来源期刊

Journal of Dental Sciences 医学-牙科与口腔外科

CiteScore

5.10

自引率

14.30%

发文量

348

审稿时长

6 days

期刊介绍： he Journal of Dental Sciences (JDS), published quarterly, is the official and open access publication of the Association for Dental Sciences of the Republic of China (ADS-ROC). The precedent journal of the JDS is the Chinese Dental Journal (CDJ) which had already been covered by MEDLINE in 1988. As the CDJ continued to prove its importance in the region, the ADS-ROC decided to move to the international community by publishing an English journal. Hence, the birth of the JDS in 2006. The JDS is indexed in the SCI Expanded since 2008. It is also indexed in Scopus, and EMCare, ScienceDirect, SIIC Data Bases. The topics covered by the JDS include all fields of basic and clinical dentistry. Some manuscripts focusing on the study of certain endemic diseases such as dental caries and periodontal diseases in particular regions of any country as well as oral pre-cancers, oral cancers, and oral submucous fibrosis related to betel nut chewing habit are also considered for publication. Besides, the JDS also publishes articles about the efficacy of a new treatment modality on oral verrucous hyperplasia or early oral squamous cell carcinoma.