LLOD schema for Simplified Offensive Language Taxonomy in multilingual detection and applications

Q2 Arts and Humanities Lodz Papers in Pragmatics Pub Date : 2023-12-01 DOI:10.1515/lpp-2023-0016
Barbara Lewandowska-Tomaszczyk, Anna Bączkowska, Olga Dontcheva-Navrátilová, Chaya Liebeskind, Giedrė Valūnaitė Oleškevičienė, Slavko Žitnik, Marcin Trojszczak, Renata Povolná, Linas Selmistraitis, A. Utka, Dangis Gudelis
{"title":"LLOD schema for Simplified Offensive Language Taxonomy in multilingual detection and applications","authors":"Barbara Lewandowska-Tomaszczyk, Anna Bączkowska, Olga Dontcheva-Navrátilová, Chaya Liebeskind, Giedrė Valūnaitė Oleškevičienė, Slavko Žitnik, Marcin Trojszczak, Renata Povolná, Linas Selmistraitis, A. Utka, Dangis Gudelis","doi":"10.1515/lpp-2023-0016","DOIUrl":null,"url":null,"abstract":"Abstract The goal of the paper is to present a Simplified Offensive Language (SOL) Taxonomy, its application and testing in the Second Annotation Campaign conducted between March-May 2023 on four languages: English, Czech, Lithuanian, and Polish to be verified and located in LLOD. Making reference to the previous Offensive Language taxonomic models proposed mostly by the same COST Action Nexus Linguarum WG 4.1.1 team, the number and variety of the categories underwent the definitional revision, and the present typology was tested in the annotation on the publicly available offensive language datasets of each of the four languages. The results of the annotation are presented and as they are contained within the accepted statistical values on the inter-annotator agreement in the SOL categories and their aspects, we propose this taxonomy as a core ontology which represents the encoding of the supported offensive languages and justify its use on new data in terms of a more universal Linguistic Linked Open Data (LLOD) schema.","PeriodicalId":39423,"journal":{"name":"Lodz Papers in Pragmatics","volume":"104 32","pages":"301 - 324"},"PeriodicalIF":0.0000,"publicationDate":"2023-12-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Lodz Papers in Pragmatics","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1515/lpp-2023-0016","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"Q2","JCRName":"Arts and Humanities","Score":null,"Total":0}
引用次数: 0

Abstract

Abstract The goal of the paper is to present a Simplified Offensive Language (SOL) Taxonomy, its application and testing in the Second Annotation Campaign conducted between March-May 2023 on four languages: English, Czech, Lithuanian, and Polish to be verified and located in LLOD. Making reference to the previous Offensive Language taxonomic models proposed mostly by the same COST Action Nexus Linguarum WG 4.1.1 team, the number and variety of the categories underwent the definitional revision, and the present typology was tested in the annotation on the publicly available offensive language datasets of each of the four languages. The results of the annotation are presented and as they are contained within the accepted statistical values on the inter-annotator agreement in the SOL categories and their aspects, we propose this taxonomy as a core ontology which represents the encoding of the supported offensive languages and justify its use on new data in terms of a more universal Linguistic Linked Open Data (LLOD) schema.
查看原文
分享 分享
微信好友 朋友圈 QQ好友 复制链接
本刊更多论文
多语言检测和应用中的简化攻击性语言分类 LLOD 模式
本文的目标是提出一个简化的攻击性语言(SOL)分类法,并在2023年3月至5月期间对英语、捷克语、立陶宛语和波兰语四种语言进行的第二次注释活动中进行应用和测试,这些语言将被验证并位于LLOD中。参考先前主要由相同的COST Action Nexus Linguarum WG 4.1.1团队提出的攻击性语言分类模型,对类别的数量和种类进行了定义修订,并在公开的四种语言的攻击性语言数据集上的注释中对目前的类型进行了测试。我们给出了注释的结果,因为它们包含在SOL类别及其方面的注释者间协议的可接受统计值中,我们建议将该分类法作为核心本体,代表支持的攻击性语言的编码,并根据更通用的语言链接开放数据(LLOD)模式证明其在新数据上的使用。
本文章由计算机程序翻译,如有差异,请以英文原文为准。
求助全文
约1分钟内获得全文 去求助
来源期刊
Lodz Papers in Pragmatics
Lodz Papers in Pragmatics Arts and Humanities-Language and Linguistics
CiteScore
1.10
自引率
0.00%
发文量
0
期刊最新文献
Writer and participant visibility in quantitative and qualitative research: a corpus-assisted study of human agent verbs in health science publications The most common graphicons in Mexican Spanish speaking WhatsApp communities composed of school parents Ecological discourse analysis and meaning interpretation of BBC news reports on 2019 Australian bushfires from the perspective of transitivity system A multimodal contrastive analysis of regulations and instructions during the COVID-19 lockdown in the context of the Island of Madeira and the United Kingdom Apology strategies in Tashelhit: linguistic realization and religious influence
×
引用
GB/T 7714-2015
复制
MLA
复制
APA
复制
导出至
BibTeX EndNote RefMan NoteFirst NoteExpress
×
×
提示
您的信息不完整,为了账户安全,请先补充。
现在去补充
×
提示
您因"违规操作"
具体请查看互助需知
我知道了
×
提示
现在去查看 取消
×
提示
确定
0
微信
客服QQ
Book学术公众号 扫码关注我们
反馈
×
意见反馈
请填写您的意见或建议
请填写您的手机或邮箱
已复制链接
已复制链接
快去分享给好友吧!
我知道了
×
扫码分享
扫码分享
Book学术官方微信
Book学术文献互助
Book学术文献互助群
群 号:481959085
Book学术
文献互助 智能选刊 最新文献 互助须知 联系我们:info@booksci.cn
Book学术提供免费学术资源搜索服务,方便国内外学者检索中英文文献。致力于提供最便捷和优质的服务体验。
Copyright © 2023 Book学术 All rights reserved.
ghs 京公网安备 11010802042870号 京ICP备2023020795号-1