Abstract

Feature selection has been proven to be efficient in preparing high dimensional data for data mining and machine learning. As most data is unlabeled, unsupervised feature selection has attracted more and more attention in recent years. Discriminant analysis has been proven to be a powerful technique to select discriminative features for supervised feature selection. To apply discriminant analysis, we usually need label information which is absent for unlabeled data. This gap makes it challenging to apply discriminant analysis for unsupervised feature selection. In this paper, we investigate how to exploit discriminant analysis in unsupervised scenarios to select discriminative features. We introduce the concept of pseudo labels, which enable discriminant analysis on unlabeled data, propose a novel unsupervised feature selection framework DisUFS which incorporates learning discriminative features with generating pseudo labels, and develop an effective algorithm for DisUFS. Experimental results on different types of real-world data demonstrate the effectiveness of the proposed framework DisUFS.

Original languageEnglish (US)
Title of host publicationSIAM International Conference on Data Mining 2014, SDM 2014
PublisherSociety for Industrial and Applied Mathematics Publications
Pages938-946
Number of pages9
Volume2
ISBN (Print)9781510811515
DOIs
StatePublished - 2014
Event14th SIAM International Conference on Data Mining, SDM 2014 - Philadelphia, United States
Duration: Apr 24 2014Apr 26 2014

Other

Other14th SIAM International Conference on Data Mining, SDM 2014
CountryUnited States
CityPhiladelphia
Period4/24/144/26/14

Fingerprint

Discriminant analysis
Feature extraction
Labels
Data mining
Learning systems

ASJC Scopus subject areas

  • Computer Science Applications
  • Software

Cite this

Tang, J., Hu, X., Gao, H., & Liu, H. (2014). Discriminant analysis for unsupervised feature selection. In SIAM International Conference on Data Mining 2014, SDM 2014 (Vol. 2, pp. 938-946). Society for Industrial and Applied Mathematics Publications. https://doi.org/10.1137/1.9781611973440.107

Discriminant analysis for unsupervised feature selection. / Tang, Jiliang; Hu, Xia; Gao, Huiji; Liu, Huan.

SIAM International Conference on Data Mining 2014, SDM 2014. Vol. 2 Society for Industrial and Applied Mathematics Publications, 2014. p. 938-946.

Research output: Chapter in Book/Report/Conference proceedingConference contribution

Tang, J, Hu, X, Gao, H & Liu, H 2014, Discriminant analysis for unsupervised feature selection. in SIAM International Conference on Data Mining 2014, SDM 2014. vol. 2, Society for Industrial and Applied Mathematics Publications, pp. 938-946, 14th SIAM International Conference on Data Mining, SDM 2014, Philadelphia, United States, 4/24/14. https://doi.org/10.1137/1.9781611973440.107
Tang J, Hu X, Gao H, Liu H. Discriminant analysis for unsupervised feature selection. In SIAM International Conference on Data Mining 2014, SDM 2014. Vol. 2. Society for Industrial and Applied Mathematics Publications. 2014. p. 938-946 https://doi.org/10.1137/1.9781611973440.107
Tang, Jiliang ; Hu, Xia ; Gao, Huiji ; Liu, Huan. / Discriminant analysis for unsupervised feature selection. SIAM International Conference on Data Mining 2014, SDM 2014. Vol. 2 Society for Industrial and Applied Mathematics Publications, 2014. pp. 938-946
@inproceedings{2b626cca7cad471495994860cc858021,
title = "Discriminant analysis for unsupervised feature selection",
abstract = "Feature selection has been proven to be efficient in preparing high dimensional data for data mining and machine learning. As most data is unlabeled, unsupervised feature selection has attracted more and more attention in recent years. Discriminant analysis has been proven to be a powerful technique to select discriminative features for supervised feature selection. To apply discriminant analysis, we usually need label information which is absent for unlabeled data. This gap makes it challenging to apply discriminant analysis for unsupervised feature selection. In this paper, we investigate how to exploit discriminant analysis in unsupervised scenarios to select discriminative features. We introduce the concept of pseudo labels, which enable discriminant analysis on unlabeled data, propose a novel unsupervised feature selection framework DisUFS which incorporates learning discriminative features with generating pseudo labels, and develop an effective algorithm for DisUFS. Experimental results on different types of real-world data demonstrate the effectiveness of the proposed framework DisUFS.",
author = "Jiliang Tang and Xia Hu and Huiji Gao and Huan Liu",
year = "2014",
doi = "10.1137/1.9781611973440.107",
language = "English (US)",
isbn = "9781510811515",
volume = "2",
pages = "938--946",
booktitle = "SIAM International Conference on Data Mining 2014, SDM 2014",
publisher = "Society for Industrial and Applied Mathematics Publications",

}

TY - GEN

T1 - Discriminant analysis for unsupervised feature selection

AU - Tang, Jiliang

AU - Hu, Xia

AU - Gao, Huiji

AU - Liu, Huan

PY - 2014

Y1 - 2014

N2 - Feature selection has been proven to be efficient in preparing high dimensional data for data mining and machine learning. As most data is unlabeled, unsupervised feature selection has attracted more and more attention in recent years. Discriminant analysis has been proven to be a powerful technique to select discriminative features for supervised feature selection. To apply discriminant analysis, we usually need label information which is absent for unlabeled data. This gap makes it challenging to apply discriminant analysis for unsupervised feature selection. In this paper, we investigate how to exploit discriminant analysis in unsupervised scenarios to select discriminative features. We introduce the concept of pseudo labels, which enable discriminant analysis on unlabeled data, propose a novel unsupervised feature selection framework DisUFS which incorporates learning discriminative features with generating pseudo labels, and develop an effective algorithm for DisUFS. Experimental results on different types of real-world data demonstrate the effectiveness of the proposed framework DisUFS.

AB - Feature selection has been proven to be efficient in preparing high dimensional data for data mining and machine learning. As most data is unlabeled, unsupervised feature selection has attracted more and more attention in recent years. Discriminant analysis has been proven to be a powerful technique to select discriminative features for supervised feature selection. To apply discriminant analysis, we usually need label information which is absent for unlabeled data. This gap makes it challenging to apply discriminant analysis for unsupervised feature selection. In this paper, we investigate how to exploit discriminant analysis in unsupervised scenarios to select discriminative features. We introduce the concept of pseudo labels, which enable discriminant analysis on unlabeled data, propose a novel unsupervised feature selection framework DisUFS which incorporates learning discriminative features with generating pseudo labels, and develop an effective algorithm for DisUFS. Experimental results on different types of real-world data demonstrate the effectiveness of the proposed framework DisUFS.

UR - http://www.scopus.com/inward/record.url?scp=84959879602&partnerID=8YFLogxK

UR - http://www.scopus.com/inward/citedby.url?scp=84959879602&partnerID=8YFLogxK

U2 - 10.1137/1.9781611973440.107

DO - 10.1137/1.9781611973440.107

M3 - Conference contribution

SN - 9781510811515

VL - 2

SP - 938

EP - 946

BT - SIAM International Conference on Data Mining 2014, SDM 2014

PB - Society for Industrial and Applied Mathematics Publications

ER -