{"title":"Playing with matches: An assessment of accuracy in linked historical data","authors":"Catherine G. Massey","doi":"10.1080/01615440.2017.1288598","DOIUrl":null,"url":null,"abstract":"ABSTRACT This article evaluates linkage quality achieved by various record linkage techniques used in historical demography. The author creates benchmark, or truth, data by linking the 2005 Current Population Survey Annual Social and Economic Supplement to the Social Security Administration's numeric identification system by social security number. By comparing simulated linkages to the benchmark data, she examines the value added (in terms of number and quality of links) from incorporating text-string comparators, adjusting age, and using a probabilistic matching algorithm. She finds that text-string comparators and probabilistic approaches are useful for increasing the linkage rate, but use of text-string comparators may decrease accuracy in some cases. Overall, probabilistic matching offers the best balance between linkage rates and accuracy.","PeriodicalId":154465,"journal":{"name":"Historical Methods: A Journal of Quantitative and Interdisciplinary History","volume":"85 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2017-03-13","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"30","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Historical Methods: A Journal of Quantitative and Interdisciplinary History","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1080/01615440.2017.1288598","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}
引用次数: 30
Abstract
ABSTRACT This article evaluates linkage quality achieved by various record linkage techniques used in historical demography. The author creates benchmark, or truth, data by linking the 2005 Current Population Survey Annual Social and Economic Supplement to the Social Security Administration's numeric identification system by social security number. By comparing simulated linkages to the benchmark data, she examines the value added (in terms of number and quality of links) from incorporating text-string comparators, adjusting age, and using a probabilistic matching algorithm. She finds that text-string comparators and probabilistic approaches are useful for increasing the linkage rate, but use of text-string comparators may decrease accuracy in some cases. Overall, probabilistic matching offers the best balance between linkage rates and accuracy.