Feature Extraction and Feature Selection: Reducing Data Complexity with Apache Spark

Authors: Dimitrios Sisiaridis, Olivier Markowitch | Published: 2017-12-11

2017.12.112025.04.03

Authors: Dimitrios Sisiaridis, Olivier Markowitch
Published: 2017-12-11

Source: https://arxiv.org/abs/1712.08618

PDF: https://arxiv.org/pdf/1712.08618

AIにより推定されたラベル

特徴抽出手法データ前処理クラスタリング手法

※ こちらのラベルはAIによって自動的に追加されました。そのため、正確でないことがあります。
詳細は文献データベースについてをご覧ください。

Abstract

Feature extraction and feature selection are the first tasks in pre-processing of input logs in order to detect cyber security threats and attacks while utilizing machine learning. When it comes to the analysis of heterogeneous data derived from different sources, these tasks are found to be time-consuming and difficult to be managed efficiently. In this paper, we present an approach for handling feature extraction and feature selection for security analytics of heterogeneous data derived from different network sensors. The approach is implemented in Apache Spark, using its python API, named pyspark.