词汇法定意义的语料库语言学方法：外延主义与内涵主义方法

IF 2.1 Applied Corpus Linguistics Pub Date : 2024-04-01 Epub Date: 2023-12-19 DOI:10.1016/j.acorp.2023.100079

Stefan Th. Gries, Brian G. Slocum, Kevin Tobia

{"title":"词汇法定意义的语料库语言学方法：外延主义与内涵主义方法","authors":"Stefan Th. Gries, Brian G. Slocum, Kevin Tobia","doi":"10.1016/j.acorp.2023.100079","DOIUrl":null,"url":null,"abstract":"<div>Scholars and practitioners interested in legal interpretation have become increasingly interested in corpus-linguistic methodology. Lee and Mouritsen (2018) developed and helped popularize the use of concordancing and collocate displays (of mostly COCA and COHA) to operationalize a central notion in legal interpretation, the ordinary meaning of expressions. This approach provides a good first approximation but is ultimately limited. Here, we outline an approach to ordinary meaning that is intensionalist (i.e., 'feature-based'), top-down, and informed by the notion of cue validity in prototype theory. The key advantages of this approach are that (i) it avoids the which-value-on-a-dimension problem of extensionalist approaches, (ii) it provides quantifiable prototypicality values for things whose membership status in a category is in question, and (iii) it can be extended even to cases for which no textual data are yet available. We exemplify the approach with two case studies that offer the option of utilizing survey data and/or word embeddings trained on corpora by deriving cue validities from word similarities. We exemplify this latter approach with the word vehicle on the basis of (i) an embedding model trained on 840 billion words crawled from the web, but now also with the more realistic application (in terms of corpus size and time frame) of (ii) an embedding model trained on the 1950s time slice of COHA to address the question to what degree Segways, which didn't exist in the 1950s, qualify as vehicles in this intensional approach.</div>","PeriodicalId":72254,"journal":{"name":"Applied Corpus Linguistics","volume":"4 1","pages":"Article 100079"},"PeriodicalIF":2.1000,"publicationDate":"2024-04-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.sciencedirect.com/science/article/pii/S2666799123000394/pdfft?md5=fffa64c5cf04e01a22d462ddb9e4441e&pid=1-s2.0-S2666799123000394-main.pdf","citationCount":"0","resultStr":"{\"title\":\"Corpus-linguistic approaches to lexical statutory meaning: Extensionalist vs. intensionalist approaches\",\"authors\":\"Stefan Th. Gries, Brian G. Slocum, Kevin Tobia\",\"doi\":\"10.1016/j.acorp.2023.100079\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"<div>Scholars and practitioners interested in legal interpretation have become increasingly interested in corpus-linguistic methodology. Lee and Mouritsen (2018) developed and helped popularize the use of concordancing and collocate displays (of mostly COCA and COHA) to operationalize a central notion in legal interpretation, the ordinary meaning of expressions. This approach provides a good first approximation but is ultimately limited. Here, we outline an approach to ordinary meaning that is intensionalist (i.e., 'feature-based'), top-down, and informed by the notion of cue validity in prototype theory. The key advantages of this approach are that (i) it avoids the which-value-on-a-dimension problem of extensionalist approaches, (ii) it provides quantifiable prototypicality values for things whose membership status in a category is in question, and (iii) it can be extended even to cases for which no textual data are yet available. We exemplify the approach with two case studies that offer the option of utilizing survey data and/or word embeddings trained on corpora by deriving cue validities from word similarities. We exemplify this latter approach with the word vehicle on the basis of (i) an embedding model trained on 840 billion words crawled from the web, but now also with the more realistic application (in terms of corpus size and time frame) of (ii) an embedding model trained on the 1950s time slice of COHA to address the question to what degree Segways, which didn't exist in the 1950s, qualify as vehicles in this intensional approach.</div>\",\"PeriodicalId\":72254,\"journal\":{\"name\":\"Applied Corpus Linguistics\",\"volume\":\"4 1\",\"pages\":\"Article 100079\"},\"PeriodicalIF\":2.1000,\"publicationDate\":\"2024-04-01\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"https://www.sciencedirect.com/science/article/pii/S2666799123000394/pdfft?md5=fffa64c5cf04e01a22d462ddb9e4441e&pid=1-s2.0-S2666799123000394-main.pdf\",\"citationCount\":\"0\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"Applied Corpus Linguistics\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://www.sciencedirect.com/science/article/pii/S2666799123000394\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"2023/12/19 0:00:00\",\"PubModel\":\"Epub\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"Applied Corpus Linguistics","FirstCategoryId":"1085","ListUrlMain":"https://www.sciencedirect.com/science/article/pii/S2666799123000394","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"2023/12/19 0:00:00","PubModel":"Epub","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 0

摘要

对法律解释感兴趣的学者和从业人员对语料库语言学方法越来越感兴趣。Lee 和 Mouritsen（2018 年）开发并帮助普及了使用协词和搭配显示（主要是 COCA 和 COHA）来操作法律解释中的一个核心概念--表达的普通意义。这种方法提供了一个良好的初步近似，但终究是有限的。在此，我们概述了一种普通意义的方法，这种方法是内向主义的（即 "基于特征"）、自上而下的，并借鉴了原型理论中的线索有效性概念。这种方法的主要优势在于：(i) 它避免了外延主义方法中的 "维度上的值 "问题；(ii) 它为那些在某个类别中的成员地位受到质疑的事物提供了可量化的原型性值；(iii) 它甚至可以扩展到尚无文本数据的情况。我们通过两个案例研究来说明这种方法，这两个案例研究提供了利用调查数据和/或在语料库中通过从词语相似性中推导线索有效性来训练词语嵌入的选项。我们以 "车辆 "一词为例，说明了后一种方法：(i) 基于从网络中抓取的 8400 亿个单词训练的嵌入模型，但现在也更现实地应用了（在语料库规模和时间框架方面）(ii) 基于 COHA 的 20 世纪 50 年代时间片训练的嵌入模型，以解决 20 世纪 50 年代并不存在的赛格威在多大程度上符合这种内向方法中的车辆的问题。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

Corpus-linguistic approaches to lexical statutory meaning: Extensionalist vs. intensionalist approaches

Scholars and practitioners interested in legal interpretation have become increasingly interested in corpus-linguistic methodology. Lee and Mouritsen (2018) developed and helped popularize the use of concordancing and collocate displays (of mostly COCA and COHA) to operationalize a central notion in legal interpretation, the ordinary meaning of expressions. This approach provides a good first approximation but is ultimately limited. Here, we outline an approach to ordinary meaning that is intensionalist (i.e., 'feature-based'), top-down, and informed by the notion of cue validity in prototype theory. The key advantages of this approach are that (i) it avoids the which-value-on-a-dimension problem of extensionalist approaches, (ii) it provides quantifiable prototypicality values for things whose membership status in a category is in question, and (iii) it can be extended even to cases for which no textual data are yet available. We exemplify the approach with two case studies that offer the option of utilizing survey data and/or word embeddings trained on corpora by deriving cue validities from word similarities. We exemplify this latter approach with the word vehicle on the basis of (i) an embedding model trained on 840 billion words crawled from the web, but now also with the more realistic application (in terms of corpus size and time frame) of (ii) an embedding model trained on the 1950s time slice of COHA to address the question to what degree Segways, which didn't exist in the 1950s, qualify as vehicles in this intensional approach.

求助全文

通过发布文献求助，成功后即可免费获取论文全文。去求助

来源期刊