A Novel Validated Real-World Dataset for the Diagnosis of Multiclass Serous Effusion Cytology according to the International System and Ground-Truth Validation Data.

IF 1.6 4区 医学 Q3 PATHOLOGY Acta Cytologica Pub Date : 2024-01-01 Epub Date: 2024-03-24 DOI:10.1159/000538465
Esraa Abd-Almoniem, Nadia Abd-Alsabour, Samar Elsheikh, Rasha R Mostafa, Yasmine Fathy Elesawy
{"title":"A Novel Validated Real-World Dataset for the Diagnosis of Multiclass Serous Effusion Cytology according to the International System and Ground-Truth Validation Data.","authors":"Esraa Abd-Almoniem, Nadia Abd-Alsabour, Samar Elsheikh, Rasha R Mostafa, Yasmine Fathy Elesawy","doi":"10.1159/000538465","DOIUrl":null,"url":null,"abstract":"<p><strong>Introduction: </strong>The application of artificial intelligence (AI) algorithms in serous fluid cytology is lacking due to the deficiency in standardized publicly available datasets. Here, we develop a novel public serous effusion cytology dataset. Furthermore, we apply AI algorithms on it to test its diagnostic utility and safety in clinical practice.</p><p><strong>Methods: </strong>The work is divided into three phases. Phase 1 entails building the dataset based on the multitiered evidence-based classification system proposed by the International System (TIS) of serous fluid cytology along with ground-truth tissue diagnosis for malignancy. To ensure reliable results of future AI research on this dataset, we carefully consider all the steps of the preparation and staining from a real-world cytopathology perspective. In phase 2, we pay special consideration to the image acquisition pipeline to ensure image integrity. Then we utilize the power of transfer learning using the convolutional layers of the VGG16 deep learning model for feature extraction. Finally, in phase 3, we apply the random forest classifier on the constructed dataset.</p><p><strong>Results: </strong>The dataset comprises 3,731 images distributed among the four TIS diagnostic categories. The model achieves 74% accuracy in this multiclass classification problem. Using a one-versus-all classifier, the fallout rate for images that are misclassified as negative for malignancy despite being a higher risk diagnosis is 0.13. Most of these misclassified images (77%) belong to the atypia of undetermined significance category in concordance with real-life statistics.</p><p><strong>Conclusion: </strong>This is the first and largest publicly available serous fluid cytology dataset based on a standardized diagnostic system. It is also the first dataset to include various types of effusions and pericardial fluid specimens. In addition, it is the first dataset to include the diagnostically challenging atypical categories. AI algorithms applied on this novel dataset show reliable results that can be incorporated into actual clinical practice with minimal risk of missing a diagnosis of malignancy. This work provides a foundation for researchers to develop and test further AI algorithms for the diagnosis of serous effusions.</p>","PeriodicalId":6959,"journal":{"name":"Acta Cytologica","volume":" ","pages":"160-170"},"PeriodicalIF":1.6000,"publicationDate":"2024-01-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Acta Cytologica","FirstCategoryId":"3","ListUrlMain":"https://doi.org/10.1159/000538465","RegionNum":4,"RegionCategory":"医学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"2024/3/24 0:00:00","PubModel":"Epub","JCR":"Q3","JCRName":"PATHOLOGY","Score":null,"Total":0}
引用次数: 0

Abstract

Introduction: The application of artificial intelligence (AI) algorithms in serous fluid cytology is lacking due to the deficiency in standardized publicly available datasets. Here, we develop a novel public serous effusion cytology dataset. Furthermore, we apply AI algorithms on it to test its diagnostic utility and safety in clinical practice.

Methods: The work is divided into three phases. Phase 1 entails building the dataset based on the multitiered evidence-based classification system proposed by the International System (TIS) of serous fluid cytology along with ground-truth tissue diagnosis for malignancy. To ensure reliable results of future AI research on this dataset, we carefully consider all the steps of the preparation and staining from a real-world cytopathology perspective. In phase 2, we pay special consideration to the image acquisition pipeline to ensure image integrity. Then we utilize the power of transfer learning using the convolutional layers of the VGG16 deep learning model for feature extraction. Finally, in phase 3, we apply the random forest classifier on the constructed dataset.

Results: The dataset comprises 3,731 images distributed among the four TIS diagnostic categories. The model achieves 74% accuracy in this multiclass classification problem. Using a one-versus-all classifier, the fallout rate for images that are misclassified as negative for malignancy despite being a higher risk diagnosis is 0.13. Most of these misclassified images (77%) belong to the atypia of undetermined significance category in concordance with real-life statistics.

Conclusion: This is the first and largest publicly available serous fluid cytology dataset based on a standardized diagnostic system. It is also the first dataset to include various types of effusions and pericardial fluid specimens. In addition, it is the first dataset to include the diagnostically challenging atypical categories. AI algorithms applied on this novel dataset show reliable results that can be incorporated into actual clinical practice with minimal risk of missing a diagnosis of malignancy. This work provides a foundation for researchers to develop and test further AI algorithms for the diagnosis of serous effusions.

查看原文
分享 分享
微信好友 朋友圈 QQ好友 复制链接
本刊更多论文
根据 TIS 和地面实况验证数据诊断多类浆液性渗出细胞学的新型验证真实世界数据集。
简介由于缺乏标准化的公开数据集,人工智能算法在浆液细胞学中的应用十分匮乏。在此,我们开发了一个新的公共浆液细胞学数据集。此外,我们还将人工智能算法应用于该数据集,以测试其在临床实践中的诊断实用性和安全性:工作分为三个阶段。第一阶段是根据国际浆液细胞学系统(TIS)提出的多层循证分类系统以及恶性肿瘤的基本组织诊断建立数据集。为确保未来人工智能研究在该数据集上取得可靠的结果,我们从现实世界细胞病理学的角度出发,仔细考虑了制备和染色的所有步骤。在第二阶段,我们对图像采集管道进行了特别考虑,以确保图像的完整性。然后,我们利用 VGG16 深度学习模型卷积层的迁移学习能力进行特征提取。最后,在第 3 阶段,我们在构建的数据集上应用随机森林分类器:该数据集包含 3731 张图像,分布在四个 TIS 诊断类别中。该模型在这一多类分类问题上达到了 74% 的准确率。使用 "一个对所有 "分类器,尽管诊断风险较高,但被误判为阴性恶性肿瘤的图像的漏判率为 0.13。这些被误判的图像中,大部分(77%)属于意义不明的非典型,与现实生活中的统计数据相符:这是首个基于标准化诊断系统的最大的公开浆液细胞学数据集。这也是第一个包含各种类型渗出液的数据集,也是第一个包含心包积液标本的数据集。此外,它还是首个包含具有诊断挑战性的非典型类别的数据集。在这一新型数据集上应用的人工智能算法显示出可靠的结果,可用于实际临床实践,将漏诊恶性肿瘤的风险降至最低。这项工作为研究人员进一步开发和测试用于诊断浆液性渗出液的人工智能算法奠定了基础。
本文章由计算机程序翻译,如有差异,请以英文原文为准。
求助全文
约1分钟内获得全文 去求助
来源期刊
Acta Cytologica
Acta Cytologica 生物-病理学
CiteScore
3.70
自引率
11.10%
发文量
46
审稿时长
4-8 weeks
期刊介绍: With articles offering an excellent balance between clinical cytology and cytopathology, ''Acta Cytologica'' fosters the understanding of the pathogenetic mechanisms behind cytomorphology and thus facilitates the translation of frontline research into clinical practice. As the official journal of the International Academy of Cytology and affiliated to over 50 national cytology societies around the world, ''Acta Cytologica'' evaluates new and existing diagnostic applications of scientific advances as well as their clinical correlations. Original papers, review articles, meta-analyses, novel insights from clinical practice, and letters to the editor cover topics from diagnostic cytopathology, gynecologic and non-gynecologic cytopathology to fine needle aspiration, molecular techniques and their diagnostic applications. As the perfect reference for practical use, ''Acta Cytologica'' addresses a multidisciplinary audience practicing clinical cytopathology, cell biology, oncology, interventional radiology, otorhinolaryngology, gastroenterology, urology, pulmonology and preventive medicine.
期刊最新文献
Reclassification of Urinary Cytology according to the Paris System for Reporting Urinary Cytology Correlation with Histological Diagnosis. Cytopathology of a Newly Described Salivary Gland Neoplasm: A Case Report of Microsecretory Adenocarcinoma Presenting in the Parotid Gland. Does the Diagnostic Performance of the Pathologist on the Indeterminate Categories of the Bethesda System for Reporting Thyroid Cytopathology Vary between Pediatric and Adult Patients? Evaluation of a Cytology-Molecular Co-Test in Liquid-Based Cytology-Processed Urine for Defining Indeterminate Categories of the Paris System. Diagnostic and Predictive Immunocytochemistry in Lung Cancer.
×
引用
GB/T 7714-2015
复制
MLA
复制
APA
复制
导出至
BibTeX EndNote RefMan NoteFirst NoteExpress
×
×
提示
您的信息不完整,为了账户安全,请先补充。
现在去补充
×
提示
您因"违规操作"
具体请查看互助需知
我知道了
×
提示
现在去查看 取消
×
提示
确定
0
微信
客服QQ
Book学术公众号 扫码关注我们
反馈
×
意见反馈
请填写您的意见或建议
请填写您的手机或邮箱
已复制链接
已复制链接
快去分享给好友吧!
我知道了
×
扫码分享
扫码分享
Book学术官方微信
Book学术文献互助
Book学术文献互助群
群 号:481959085
Book学术
文献互助 智能选刊 最新文献 互助须知 联系我们:info@booksci.cn
Book学术提供免费学术资源搜索服务,方便国内外学者检索中英文文献。致力于提供最便捷和优质的服务体验。
Copyright © 2023 Book学术 All rights reserved.
ghs 京公网安备 11010802042870号 京ICP备2023020795号-1