Corpora最新文献

英文中文

Front matter 前页

IF 0.5 Q3 LINGUISTICS

Corpora

Pub Date : 2020-11-01 DOI: 10.3366/cor.2020.0198

引用次数: 0

Multilingualism in Greater Poland court records (1386–1448): tagging discourse boundaries and code-switching 大波兰法庭记录中的多语现象（1386-1448）：标记话语边界和代码转换

IF 0.5 Q3 LINGUISTICS

Corpora

Pub Date : 2020-11-01 DOI: 10.3366/cor.2020.0200

M. Włodarczyk, J. Kopaczyk, M. Kozák

This paper introduces the Electronic Repository of Greater Poland Oaths, eROThA (1386–1446), a digitisation project of a diplomatic edition of mediaeval land court oaths recorded in Latin and Old Polish, resulting in a small, lightly tagged specialised bilingual corpus. We present the background, aims, design and methodology of the project. We also discuss the problems and limitations entrenched in turning a printed diplomatic edition into a machine-readable diplomatic edition equipped with a new interpretative layer that is sensitive to the switches between Latin and Old Polish. In addition to the automatic annotation of code-switched items on the basis of typographic characteristics of the printed edition, flexible coding of recurrent language and discourse boundary phenomena has been introduced manually to account for linguistically ambiguous or neutral forms. The project offers a fully multilingual corpus, as well as customised Polish-only and Latin-only datasets, and enables filtered metadata searches in the online front-end. Overall, the report presents a methodology for constructing multilingual corpora in the context of legal cultures in medieval Central Europe that may be extrapolated to datasets originating in other periods and regions.

本文介绍了大波兰宣誓电子库eROThA（1386-1446），这是一个用拉丁语和古波兰语记录的中世纪土地法院宣誓外交版的数字化项目，产生了一个小的、标记较轻的专业双语语料库。我们介绍了该项目的背景、目标、设计和方法。我们还讨论了将印刷外交版转变为机器可读外交版的问题和局限性，该外交版配备了对拉丁语和古波兰语之间的转换敏感的新解释层。除了根据印刷版的印刷特点自动注释代码转换项目外，还手动引入了对反复出现的语言和话语边界现象的灵活编码，以解释语言歧义或中性形式。该项目提供了一个完全多语言的语料库，以及定制的仅限波兰语和仅限拉丁语的数据集，并在在线前端实现过滤元数据搜索。总的来说，该报告提出了一种在中世纪中欧法律文化背景下构建多语言语料库的方法，该方法可以外推到其他时期和地区的数据集。

{"title":"Multilingualism in Greater Poland court records (1386–1448): tagging discourse boundaries and code-switching","authors":"M. Włodarczyk, J. Kopaczyk, M. Kozák","doi":"10.3366/cor.2020.0200","DOIUrl":"https://doi.org/10.3366/cor.2020.0200","url":null,"abstract":"This paper introduces the Electronic Repository of Greater Poland Oaths, eROThA (1386–1446), a digitisation project of a diplomatic edition of mediaeval land court oaths recorded in Latin and Old Polish, resulting in a small, lightly tagged specialised bilingual corpus. We present the background, aims, design and methodology of the project. We also discuss the problems and limitations entrenched in turning a printed diplomatic edition into a machine-readable diplomatic edition equipped with a new interpretative layer that is sensitive to the switches between Latin and Old Polish. In addition to the automatic annotation of code-switched items on the basis of typographic characteristics of the printed edition, flexible coding of recurrent language and discourse boundary phenomena has been introduced manually to account for linguistically ambiguous or neutral forms. The project offers a fully multilingual corpus, as well as customised Polish-only and Latin-only datasets, and enables filtered metadata searches in the online front-end. Overall, the report presents a methodology for constructing multilingual corpora in the context of legal cultures in medieval Central Europe that may be extrapolated to datasets originating in other periods and regions.","PeriodicalId":44933,"journal":{"name":"Corpora","volume":" ","pages":""},"PeriodicalIF":0.5,"publicationDate":"2020-11-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"44457108","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}

引用次数: 1

A historical characterisation of American and Brazilian cultures based on lexical representations 基于词汇表征的美国和巴西文化的历史特征

IF 0.5 Q3 LINGUISTICS

Corpora

Pub Date : 2020-08-27 DOI: 10.3366/cor.2020.0194

Tony Berber Sardinha

The goal of this study is to detect the historical distribution of representations of the United States and Brazil formed around the use of the nationality adjectives American and Brazilian. To ach...

本研究的目的是检测美国和巴西的表征在使用国籍形容词American和Brazilian前后形成的历史分布。到ach。。。

引用次数: 2

Calculating and displaying key labels: the texts, sections, authors and neighbourhoods where words and collocations are likely to be prominent 计算和显示关键标签：单词和搭配可能突出的文本、章节、作者和社区

IF 0.5 Q3 LINGUISTICS

Corpora

Pub Date : 2020-08-27 DOI: 10.3366/cor.2020.0193

Stephen Jeaco

Corpora are usually not only made up of words, sentences and plain texts; they usually also have metadata, background information and structural features which can be used to filter searches or pro...

语料库通常不仅由单词、句子和普通文本组成;它们通常还具有元数据、背景信息和结构特征，可用于过滤搜索或改进。

引用次数: 1

Review: McIntyre and Walker. 2019. Corpus Stylistics 评论：麦金太尔和沃克。2019.语料库文体学

IF 0.5 Q3 LINGUISTICS

Corpora

Pub Date : 2020-08-27 DOI: 10.3366/cor.2020.0196

Jordan Smith

引用次数: 0

ESP corpus design: compilation of the Veterinary Nursing Medical Chart Corpus and the Veterinary Nursing Wordlist ESP语料库设计:编制兽医护理医学图语料库和兽医护理词汇表

IF 0.5 Q3 LINGUISTICS

Corpora

Pub Date : 2020-08-27 DOI: 10.3366/cor.2020.0191

Yukiko Ohashi, N. Katagiri, K. Oka, Michiko Hanada

This paper reports on two research results: (1) designing an English for Specific Purposes (esp) corpus architecture complete with annotations structured by regular expressions; and (2) a case stud...

本文报告了两个研究成果:(1)设计了一个专用英语语料库体系结构，其中包含正则表达式结构的注释;(2)一个案例……

引用次数: 7

Back matter 回到问题

IF 0.5 Q3 LINGUISTICS

Corpora

Pub Date : 2020-08-01 DOI: 10.3366/cor.2020.0197

引用次数: 0

Crowdsourcing formulaic phrases: towards a new type of spoken corpus 众包公式化短语:走向新型口语语料库

IF 0.5 Q3 LINGUISTICS

Corpora

Pub Date : 2020-08-01 DOI: 10.3366/COR.2020.0192

S. Adolphs, Dawn Knight, Catherine Smith, Dominic T. Price

Spoken corpora have traditionally been assembled through careful recording and transcription of discourse events, a process which is both labour intensive and often restrictive in terms of breadth of recording contexts available. To overcome these potential challenges in spoken corpus compilation, we explore the use of crowdsourcing of language samples that are reported by participants. We investigate the level of precision and recall of the ‘crowd’ when it comes to reporting language they have heard in certain contexts, alongside the use of a crowdsourcing toolkit to facilitate this task. As a focussing device for the selection of reported language samples, we draw on the use of formulaic phrases as an area that has received considerable attention by corpus linguists and applied linguists over the years. We argue that while studying reported language usage instead of actual language-in-use is problematic for several reasons, many of which have been highlighted in the literature on Discourse Completion Tasks ( Schauer and Adolphs, 2006 ), our suggested approach presents several advantages and opportunities for spoken corpus linguistics.

口语语料库传统上是通过仔细记录和转录话语事件来组装的，这一过程既劳动密集型，而且在记录上下文的广度方面往往受到限制。为了克服口语语料库编纂中的这些潜在挑战，我们探索了参与者报告的语言样本的众包使用。我们调查了“人群”在某些情况下听到的报告语言的准确性和记忆力，同时使用众包工具包来促进这项任务。作为选择报告语言样本的一种集中手段，我们将公式化短语的使用作为一个领域，多年来受到语料库语言学家和应用语言学家的极大关注。我们认为，虽然研究报告的语言使用而不是实际使用的语言是有问题的，原因有几个，其中许多已经在关于语篇完成任务的文献中得到了强调（Schauer和Adolphs，2006），但我们提出的方法为口语语料库语言学提供了一些优势和机会。

引用次数: 4

Front matter 前页

IF 0.5 Q3 LINGUISTICS

Corpora

Pub Date : 2020-08-01 DOI: 10.3366/cor.2020.0190

引用次数: 0

Mandative subjunctive versus should in world Englishes: a new take on an old alternation 世界英语中的强制虚拟语气与should：对旧交替的新诠释

IF 0.5 Q3 LINGUISTICS

Corpora

Pub Date : 2020-08-01 DOI: 10.3366/cor.2020.0195

Sandra C. Deshors, S. Gries

This study explores the alternation between the mandative subjunctive and its modal alternative with should across native and non-native Englishes. Methodologically, we try to improve on existing standards by investigating over 3,300 occurrences of the alternation from the Corpus of Web-based Global English and annotated for a range of linguistic factors analysed with a forest of conditional inference trees; also, we are exemplifying a new strategy for the use of random or conditional inference forests in corpus-based alternation studies. We obtain a forest with significant prediction accuracies and a good C-score and discuss the strongest predictors of the subjunctive versus should alternation across Englishes. Contrasting with existing research, our multi-factorial results: ( i) suggest that in British English the mandative subjunctive may not be dying out as much as we thought; and ( ii) individual suasive verbs influence speakers' use of the two variants more than their variety of English.

本研究探讨了在母语和非母语英语中，强制虚拟语气及其语气替换词与should之间的交替。在方法论上，我们试图通过调查基于网络的全球英语语料库中的3300多个交替事件来改进现有标准，并用条件推理树森林对一系列语言因素进行注释；此外，我们还举例说明了在基于语料库的交替研究中使用随机或条件推理森林的新策略。我们获得了一个具有显著预测精度和良好C核的森林，并讨论了英语中虚拟语气与should交替的最强预测因子。与现有研究相比，我们的多因素结果表明：（i）在英国英语中，强制虚拟语气可能并没有像我们想象的那样消亡；（ii）个别的说服动词对说话人使用这两种变体的影响大于对其英语变体的影响。

引用次数: 8

首页上一页

下一页尾页

类型

全部化学•材料生命科学医学物理工程技术环境•农林材料科学地球科学法学管理学化学环境科学与生态学计算机科学教育学经济学农林科学人文科学生物学数学物理与天体物理心理学综合性期刊其他工业工程理学历史学农学文学信息工程

数据库

全部 ACS Publications Elsevier ieeexplore Springer The Royal Society of Chemistry Wiley

期刊

Corpora

全部 Acc. Chem. Res. ACS Applied Bio Materials ACS Appl. Electron. Mater. ACS Appl. Energy Mater. ACS Appl. Mater. Interfaces ACS Appl. Nano Mater. ACS Appl. Polym. Mater. ACS BIOMATER-SCI ENG ACS Catal. ACS Cent. Sci. ACS Chem. Biol. ACS Chemical Health & Safety ACS Chem. Neurosci. ACS Comb. Sci. ACS Earth Space Chem. ACS Energy Lett. ACS Infect. Dis. ACS Macro Lett. ACS Mater. Lett. ACS Med. Chem. Lett. ACS Nano ACS Omega ACS Photonics ACS Sens. ACS Sustainable Chem. Eng. ACS Synth. Biol. Anal. Chem. BIOCHEMISTRY-US Bioconjugate Chem. BIOMACROMOLECULES Chem. Res. Toxicol. Chem. Rev. Chem. Mater. CRYST GROWTH DES ENERG FUEL Environ. Sci. Technol. Environ. Sci. Technol. Lett. Eur. J. Inorg. Chem. IND ENG CHEM RES Inorg. Chem. J. Agric. Food. Chem. J. Chem. Eng. Data J. Chem. Educ. J. Chem. Inf. Model. J. Chem. Theory Comput. J. Med. Chem. J. Nat. Prod. J PROTEOME RES J. Am. Chem. Soc. LANGMUIR MACROMOLECULES Mol. Pharmaceutics Nano Lett. Org. Lett. ORG PROCESS RES DEV ORGANOMETALLICS J. Org. Chem. J. Phys. Chem. J. Phys. Chem. A J. Phys. Chem. B J. Phys. Chem. C J. Phys. Chem. Lett. Analyst Anal. Methods Biomater. Sci. Catal. Sci. Technol. Chem. Commun. Chem. Soc. Rev. CHEM EDUC RES PRACT CRYSTENGCOMM Dalton Trans. Energy Environ. Sci. ENVIRON SCI-NANO ENVIRON SCI-PROC IMP ENVIRON SCI-WAT RES Faraday Discuss. Food Funct. Green Chem. Inorg. Chem. Front. Integr. Biol. J. Anal. At. Spectrom. J. Mater. Chem. A J. Mater. Chem. B J. Mater. Chem. C Lab Chip Mater. Chem. Front. Mater. Horiz. MEDCHEMCOMM Metallomics Mol. Biosyst. Mol. Syst. Des. Eng. Nanoscale Nanoscale Horiz. Nat. Prod. Rep. New J. Chem. Org. Biomol. Chem. Org. Chem. Front. PHOTOCH PHOTOBIO SCI PCCP Polym. Chem.

﹀