跳到主要內容區

10/22(五) OCR, Visual Linguistic Pretraining, and Document Understanding 主講人:Cha Zhang, Ph.D.

國立清華大學資訊工程學系

Department of Computer Science

National Tsing Hua University

專題演講

SEMINAR

主講人: Cha Zhang, Ph.D.

SPEAKER  Microsoft Cloud & AI

題 目:OCR, Visual Linguistic Pretraining, and Document Understanding

TOPIC     

時  間:1101022()上午1010分至12

DATE  

地 點:線上會議室

PLACE

(meeting link: https://teams.microsoft.com/l/meetup-join/19%3afkE_SDD1vdIkP4vY8kqmtlL5o3uDwVlMfQxiWsjw_UI1%40thread.tacv2/1632317678550?context=%7b%22Tid%22%3a%226c3bc511-43c7-4596-baeb-2335c69c41f1%22%2c%22Oid%22%3a%22821516e4-93a6-42fb-bdf6-31b22f1d8c2c%22%7d )

Abstract:

Recent progress in AI has brought Optical Character Recognition (OCR), visual linguistics and document understanding to a whole new level. In this talk, we will first provide an overview of Microsoft’s latest OCR engine (aka OneOCR), which applies the latest deep learning techniques to recognize mixed printed and handwritten text in over 100 languages, with text lines along arbitrary orientations (even flipped), and with varying degrees of quality and distortion. We then introduce text-aware visual-linguistic pretraining and some applications in text-VQA and text captioning. And lastly, we present a novel technology for document understanding: LayoutLM. LayoutLM bridges computer vision and language, producing state-of-the art results on a number of tasks, including document segmentation, classification, TextVQA, and others.

About the speaker:

Cha Zhang is a Partner Engineering Manager at Microsoft Cloud & AI. He received the B.S. and M.S. degrees from Tsinghua University, Beijing, China in 1998 and 2000, respectively, both in Electronic Engineering, and the Ph.D. degree in Electrical and Computer Engineering from Carnegie Mellon University, in 2004. After graduation, he worked at Microsoft Research for 12 years investigating research topics including multimedia signal processing, computer vision and machine learning. He has published more than 150 technical papers and hold more than 50 U.S. patents. He served as Program Co-Chair for VCIP 2012 and MMSP 2018, and General Co-Chair for ICME 2016. He is a Fellow of the IEEE. Since joining Cloud & AI, he has led teams to ship industry-leading technologies in Microsoft Cognitive Services such as emotion recognition, optical character recognition and document understanding.

 

Host: 郭昱廷

 

瀏覽數:
登入成功