103/06/20(五) race-Norm-Based PCA and DCA for High-Dimensional Big Data Analysis 主講人: Prof. Sun-Yuan Kung(Princeton University)
國立清華大學 資訊工程學系
Department Of Computer Science
National Tsing Hua University
專題演講
SEMINAR
主 講 人: Prof. Sun-Yuan Kung
SPEAKER (Princeton University)
題 目: Trace-Norm-Based PCA and DCA
TOPIC for High-Dimensional Big Data
Analysis
時 間: 103年6月20日(五) 10AM-11:30AM
DATE
地 點: 台達館613室
PLACE
摘 要:
Abstract
Big data analysis presents two fronts of research challenges: (1) large data size and (2) high feature dimensionality. As to the latter, the major concerns involve computation and over-training, both solvable via effective dimension reduction . Principal Component Analysis (PCA) has been the prevailing solution for unsupervised learning applications,. This talk, however, will explore PCA’s counterparts in supervised learning scenarios.
Our proposed SNR metric for dimension reduction stems from the classic Fisher Discriminant Analysis (FDA). Recall that FDA aims at finding a single optimal projection component. To extend it to multiple components, one must (1) extend SNR to SoSNR (Sum of SNRs) and (2) ensure an orthogonality between the projection components (so as to avoid the wasteful redundancy). For example, Successively Orthogonal Discriminant Analysis (SODA ) aims at maximizing SoSNR while enforcing the orthogonality by a computationally simple deflation method .
A conceptually illuminating derivation of FDA involves a pre-whitened space in which the within-class perturbation may be modeled as an isotropic noise. The good news is that there exists a closed-form solution for the optimal multiple components (orthogonal in the whitened space ) which maximize SoSNR . Its theoretical foundation hinges upon a trace-optimization theorem stating that optimal solution for a SoSNR-type cost function may be derived from the generalized eigenvectors of a pair of matrices associated with the signal and noise powers, respectively. The theorem is a key formulation unifying four subspace projection methods: PCA, MD-FDA, PC-DCA, MD-PDA, and DCA (Discriminant Component Analysis).
Just like KRR, DCA incorporates a ridge parameter to mitigates the effect of data perturbation . (KRR stands for kernel ridge regressor which is equivalent to Perturbational Discriminant Analysis (PDA ). The ridge parameter plays an additional role in extending the single-projection PDA to its multi-component variants: MD-PDA and DCA. By some real-world application examples, we shall demonstrate the resilience property of DCA (w.r.t. SODA). Moreover, we shall show that DCA and SODA may jointly boost the prediction accuracies as they naturally complement each other .
連絡人: 賴尚宏教授
