103/06/19(四) Kernel Approach to Incomplete Data Analysis (KAIDA) 主講人: Prof. Sun-Yuan Kung and Mr. Peiyuan Wu(Princeton University)
國立清華大學 資訊工程學系
Department Of Computer Science
National Tsing Hua University
專題演講
SEMINAR
主 講 人: Prof. Sun-Yuan Kung and Mr. Peiyuan Wu
SPEAKER (Princeton University)
題 目: Kernel Approach to Incomplete
TOPIC Data Analysis (KAIDA)
時 間: 103 年 6月19日 (四) 15:00-16:30
DATE
地 點: 台達601室
PLACE
摘 要:
Abstract
In big data analysis (BDA), the chance of missing data will become increasingly likely, prompting a renewed interest on Incomplete Data Analysis (IDA). In IDA, we shall assume that the missing entries may vary from one feature vector to another. Moreover, there is no assumption of any particular sparsity structure, i.e. the locations of missing entries are sparse in a totally random fashion. Such assumptions are natural to the following application scenarios:
1. In disaster and emergency, sensors may become temporarily disrupted or permanently out of order.
2. Any large array of sensors or measuring devices will inevitably suffer from some extent of partial failures or disorders.
3. The topics covered in, say, numerous tweets from set of a heterogeneous network users are presumably diversified and thus highly sparse.
4. Some personal or private data are often deliberately hidden or masked for security reason.
In this talk, we shall explore a kernel approach to incomplete data analysis (KAIDA), with potential applications to both supervised and unsupervised problems. In KAIDA, two data-masking schemes are adopted to convert partially-specified vectors into fully-specified vectors, which in turn lead to five well-defined kernels. We have applied KAIDA to supervised learning applications, including both kernel ridge regressor (KRR) and support vector machine (SVM). Our findings suggest that, except for the linear kernel, each of the remaining kernels has its own claim of winning sparsity regions. First, there is no surprise that two conventional kernels (SM-RBF, and M-Poly2) yield higher accuracies than the linear kernel, with M-Poly2 exhibiting higher resilience than SM-RBF. Overall, two partial-cosine (PC) kernels tailor-designed for IDA appear to outperform SM-RBF and M-Poly2 in resilience. In the case of KRR, the two PCs also delivers higher accuracies.
A particularly worth-mentioning kernel for IDA is a Doubly-Masked Partial-Cosine (DM-PC) , defined as the normalized inner-product of the pairwise co-existing features. Note that, the DM-vector associated with a vector will change when its partner changes, rendering a doubly-masked data matrix undefinable. (Note that a fully specified data matrix is usually expected in conventional learning models.) Fortunately, the kernel method basically circumvents the need of a data matrix, making KAIDA feasible. It is, however, accompanied with a price that the DM-PC kernel function fails the vital Mercer condition.
This would imply a very poor performance if the kernel methods were directly applied. On the other hand, once a proper ``Mercerization" measure is attained, the performance tends to improve significantly. Indeed, as confirmed by our simulations that highly resilient performance may be attained by KRR and/or SVM using the proposed kernels.
連絡人: 賴尚宏教授
