{"title":"Data Science Challenges of Automated Quality Verification Process in Product Data Catalogues","authors":"Niemir Maciej","doi":"10.21741/9781644902691-45","DOIUrl":null,"url":null,"abstract":"Abstract. Product master data are an essential and key component of purchasing processes, ensuring the smooth running of business operations within companies. Unfortunately, due to the lack of a single, complete, worldwide information system storing reference data, managing the data, maintaining its quality, reliability, and timeliness, requires building quality assurance teams for such processes in most companies. There are numerous errors in product data, and identification and correction of them are time-consuming, especially for large data sets that contain many millions of products. These errors are due to the so-called human factor but are also the result of technical errors and limitations of IT systems. Therefore, in the paper, we proposed a number of solutions by category and group that can automate, simplify, and shorten the master data management process. There are also presented examples of data validation using a variety of techniques, rule-based, dictionary-based, and machine learning, that enable mass verification of both images, textual parameters, digital parameters, and classifiers, while indicating the probability of errors in specific attributes as well as in their combination, and in some cases correcting or proposing correct records. The performed tests illustrate the magnitude of problems and potential on a sample dataset.","PeriodicalId":106609,"journal":{"name":"Quality Production Improvement and System Safety","volume":"17 4 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2023-09-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Quality Production Improvement and System Safety","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.21741/9781644902691-45","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}
引用次数: 0
Abstract
Abstract. Product master data are an essential and key component of purchasing processes, ensuring the smooth running of business operations within companies. Unfortunately, due to the lack of a single, complete, worldwide information system storing reference data, managing the data, maintaining its quality, reliability, and timeliness, requires building quality assurance teams for such processes in most companies. There are numerous errors in product data, and identification and correction of them are time-consuming, especially for large data sets that contain many millions of products. These errors are due to the so-called human factor but are also the result of technical errors and limitations of IT systems. Therefore, in the paper, we proposed a number of solutions by category and group that can automate, simplify, and shorten the master data management process. There are also presented examples of data validation using a variety of techniques, rule-based, dictionary-based, and machine learning, that enable mass verification of both images, textual parameters, digital parameters, and classifiers, while indicating the probability of errors in specific attributes as well as in their combination, and in some cases correcting or proposing correct records. The performed tests illustrate the magnitude of problems and potential on a sample dataset.