跳到主要內容區

103/06/19(四) Kernel Approach to Incomplete Data Analysis (KAIDA) 主講人: Prof. Sun-Yuan Kung and Mr. Peiyuan Wu(Princeton University)

國立清華大學  資訊工程學系

Department Of Computer Science

National  Tsing  Hua  University

專題演講

SEMINAR

 

主 講 人: Prof. Sun-Yuan Kung and Mr. Peiyuan Wu

SPEAKER    (Princeton University)

題    目: Kernel Approach to Incomplete

TOPIC     Data Analysis (KAIDA)

 

時    間: 103 年 6月19日 (四) 15:00-16:30

DATE

 

地    點: 台達601

PLACE                       

 

摘    要:

Abstract

In big data analysis (BDA), the chance of missing data will become  increasingly likely, prompting a renewed interest on Incomplete Data Analysis (IDA).  In IDA, we shall assume that  the missing entries may vary from one feature vector to another.  Moreover, there is no assumption of any particular sparsity structure,  i.e. the locations of missing entries are sparse in a totally random fashion.  Such assumptions are natural to the following  application scenarios:

1. In disaster and emergency, sensors may become temporarily disrupted or permanently out of order.

2. Any  large array of sensors or measuring devices will inevitably suffer from some extent of partial failures or disorders.

3. The  topics covered in, say, numerous tweets from set of a  heterogeneous network users are  presumably diversified and thus highly sparse.

4.  Some personal or private data are often deliberately hidden or masked for security reason.

In this talk, we shall explore a kernel approach to  incomplete data analysis (KAIDA), with potential applications to both supervised and unsupervised problems.  In KAIDA,  two data-masking schemes  are adopted to convert partially-specified vectors into fully-specified vectors, which in turn lead to   five well-defined kernels.   We have applied  KAIDA to supervised learning applications, including both kernel ridge regressor (KRR) and support vector machine (SVM).  Our findings suggest that,  except for  the linear kernel, each of the remaining kernels has its own claim of winning sparsity regions. First, there is no surprise that two conventional kernels (SM-RBF, and M-Poly2) yield higher accuracies  than the linear kernel, with M-Poly2 exhibiting higher resilience than SM-RBF. Overall, two partial-cosine (PC)  kernels tailor-designed for IDA appear to outperform SM-RBF and M-Poly2 in resilience.  In the case of KRR, the two PCs  also delivers higher accuracies.

A particularly worth-mentioning  kernel for IDA is  a Doubly-Masked Partial-Cosine (DM-PC) , defined as the normalized inner-product of the pairwise co-existing features.  Note that, the DM-vector associated with a vector will change when its partner changes,  rendering a doubly-masked  data matrix undefinable.  (Note that a  fully specified  data matrix  is usually expected in  conventional learning models.)  Fortunately,  the kernel method basically circumvents the need of a   data matrix, making KAIDA feasible.  It is, however,  accompanied with a  price that the DM-PC kernel function fails the vital Mercer condition.

This would imply a very  poor performance if the kernel methods were directly applied.  On the other hand, once a proper ``Mercerization" measure is attained, the performance tends to improve significantly. Indeed, as confirmed by our simulations that highly resilient performance may be attained by KRR and/or SVM using the proposed kernels.

 

連絡人:  賴尚宏教授

瀏覽數:
登入成功