Optimized Conversion of Categorical and Numerical Features in Machine Learning Models

2021 Fifth International Conference on I-SMAC (IoT in Social, Mobile, Analytics and Cloud) (I-SMAC) Pub Date : 2021-11-11 DOI:10.1109/I-SMAC52330.2021.9640967

K. P. N. V. Satya Sree, J. Karthik, Chava Niharika, P. Srinivas, N. Ravinder, Chitturi Prasad

引用次数: 2

Abstract

While some data have an explicit, numerical form, many other data, such as gender or nationality, do not typically use numbers and are referred to as categorical data. Thus, machine learning algorithms need a way of representing categorical information numerically in order to be able to analyze them. Our project specifically focuses on optimizing the conversion of categorical features to a numerical form in order to maximize the effectiveness of various machine learning models. From the methods utilized, it has been observed that wide and deep is the most effective model for datasets that contain high-cardinality features, as opposed to learn embedding and one-hot encoding.

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

机器学习模型中分类和数值特征的优化转换

虽然有些数据有明确的数字形式，但许多其他数据，如性别或国籍，通常不使用数字，被称为分类数据。因此，机器学习算法需要一种以数字方式表示分类信息的方法，以便能够对它们进行分析。我们的项目特别侧重于优化分类特征到数值形式的转换，以最大限度地提高各种机器学习模型的有效性。从所使用的方法中，已经观察到，对于包含高基数特征的数据集，与学习嵌入和单热编码相反，宽和深是最有效的模型。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文去求助

来源期刊

2021 Fifth International Conference on I-SMAC (IoT in Social, Mobile, Analytics and Cloud) (I-SMAC)

自引率

0.00%

发文量