一种用于数学表达式在线手势识别的变压器结构

Irish Conference on Artificial Intelligence and Cognitive Science Pub Date : 2022-11-04 DOI:10.48550/arXiv.2211.02643

Mirco Ramo, G. Silvestre

{"title":"一种用于数学表达式在线手势识别的变压器结构","authors":"Mirco Ramo, G. Silvestre","doi":"10.48550/arXiv.2211.02643","DOIUrl":null,"url":null,"abstract":"The Transformer architecture is shown to provide a powerful framework as an end-to-end model for building expression trees from online handwritten gestures corresponding to glyph strokes. In particular, the attention mechanism was successfully used to encode, learn and enforce the underlying syntax of expressions creating latent representations that are correctly decoded to the exact mathematical expression tree, providing robustness to ablated inputs and unseen glyphs. For the first time, the encoder is fed with spatio-temporal data tokens potentially forming an infinitely large vocabulary, which finds applications beyond that of online gesture recognition. A new supervised dataset of online handwriting gestures is provided for training models on generic handwriting recognition tasks and a new metric is proposed for the evaluation of the syntactic correctness of the output expression trees. A small Transformer model suitable for edge inference was successfully trained to an average normalised Levenshtein accuracy of 94%, resulting in valid postfix RPN tree representation for 94% of predictions.","PeriodicalId":286718,"journal":{"name":"Irish Conference on Artificial Intelligence and Cognitive Science","volume":"75 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2022-11-04","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"1","resultStr":"{\"title\":\"A Transformer Architecture for Online Gesture Recognition of Mathematical Expressions\",\"authors\":\"Mirco Ramo, G. Silvestre\",\"doi\":\"10.48550/arXiv.2211.02643\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"The Transformer architecture is shown to provide a powerful framework as an end-to-end model for building expression trees from online handwritten gestures corresponding to glyph strokes. In particular, the attention mechanism was successfully used to encode, learn and enforce the underlying syntax of expressions creating latent representations that are correctly decoded to the exact mathematical expression tree, providing robustness to ablated inputs and unseen glyphs. For the first time, the encoder is fed with spatio-temporal data tokens potentially forming an infinitely large vocabulary, which finds applications beyond that of online gesture recognition. A new supervised dataset of online handwriting gestures is provided for training models on generic handwriting recognition tasks and a new metric is proposed for the evaluation of the syntactic correctness of the output expression trees. A small Transformer model suitable for edge inference was successfully trained to an average normalised Levenshtein accuracy of 94%, resulting in valid postfix RPN tree representation for 94% of predictions.\",\"PeriodicalId\":286718,\"journal\":{\"name\":\"Irish Conference on Artificial Intelligence and Cognitive Science\",\"volume\":\"75 1\",\"pages\":\"0\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2022-11-04\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"1\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"Irish Conference on Artificial Intelligence and Cognitive Science\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.48550/arXiv.2211.02643\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"Irish Conference on Artificial Intelligence and Cognitive Science","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.48550/arXiv.2211.02643","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 1

摘要

Transformer体系结构提供了一个强大的框架作为端到端模型，用于从与字形笔画相对应的在线手写手势构建表达式树。特别是，注意机制被成功地用于编码、学习和执行表达式的底层语法，从而创建了能够被正确解码为精确数学表达式树的潜在表示，为删除的输入和看不见的符号提供了鲁棒性。编码器第一次被输入了时空数据符号，有可能形成一个无限大的词汇表，这发现了在线手势识别之外的应用。为通用手写识别任务的训练模型提供了一个新的在线手写手势监督数据集，并提出了一种新的指标来评估输出表达式树的语法正确性。一个适合边缘推理的小型Transformer模型被成功训练到94%的平均归一化Levenshtein准确率，从而在94%的预测中获得有效的后缀RPN树表示。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

A Transformer Architecture for Online Gesture Recognition of Mathematical Expressions

The Transformer architecture is shown to provide a powerful framework as an end-to-end model for building expression trees from online handwritten gestures corresponding to glyph strokes. In particular, the attention mechanism was successfully used to encode, learn and enforce the underlying syntax of expressions creating latent representations that are correctly decoded to the exact mathematical expression tree, providing robustness to ablated inputs and unseen glyphs. For the first time, the encoder is fed with spatio-temporal data tokens potentially forming an infinitely large vocabulary, which finds applications beyond that of online gesture recognition. A new supervised dataset of online handwriting gestures is provided for training models on generic handwriting recognition tasks and a new metric is proposed for the evaluation of the syntactic correctness of the output expression trees. A small Transformer model suitable for edge inference was successfully trained to an average normalised Levenshtein accuracy of 94%, resulting in valid postfix RPN tree representation for 94% of predictions.

求助全文

通过发布文献求助，成功后即可免费获取论文全文。去求助

来源期刊

Irish Conference on Artificial Intelligence and Cognitive Science

自引率

0.00%

发文量