Video-based fall detection by multiview 3D spatial features(本文)

Huu Hung Dao · Institutional Repositories DataBase (IRDB) · 2014

In this dissertation, we address the problem of detecting fall incidents by using vision technology.This problem is critical to ensure the safety of the elderly who increasingly prefer to live alone at home but are prone to suffer from accidental falls.The aim of detecting falls instantly is to offer immediate help to fallen elderly, in turn, not worsening their injuries or even saving their lives.Even though a large body of literature have been dedicating to fall detection, many challenges still remain for further investigation.One of the major challenges is to discriminate carefully falls from various activities of daily living (ADL), especially like-fall ones.e.g., crouch on the ground and sit down brutally, etc.Secondly, fall detection seems to be meaningless without real-time performance.Other challenges include low image quality, cluttered background, illumination variations, appearance variations, camera viewpoints, and occlusion by furniture, etc.We realize that falls are associated with fast body movements to change postures from upright to almost lengthened, followed by a sufficient duration of staying almost motionlessly on the ground.This is contrary to slow manners of doing ADL of the elderly.Hence, we propose using 3D spatial features which are efficiently estimated from multiple views and are highly discriminative to classify human states into standing, sitting, and lying.Once a sequence of human states is given, fall events can be reliably inferred by analyzing human state transition.Firstly, we describe in this dissertation a combination of heights and occupied areas, extracted from 3D cuboids of the person of interest for human state classification.Lying people take larger areas than sitting and standing people.Standing people are higher than sitting and lying people.These three states intuitively lie in three separable region of the feature space which can be classified by SVM.Falls are inferred by time-series analysis of human state transition.For efficient feature estimation, we configure two cameras whose fields of view are relatively orthogonal.Thus, 2D bounding boxes of the person, extracted from two views, serve as two orthographic projections of the 3D cuboids.The features are normalized by using Local Empirical i Okinawa with Ikeda san and thank for his support.Lastly, special thanks to Tamakisan, Honda-san, Shinozuka-san, Nakagawa-san, Nakayama-san, Kawasaki-san, and other lab members.I also need to include staffs in YISH (Yokohama International Student House) in which I spent 2 years on living.They gave me a good opportunity to live in an international environment with many social activities with local people so that I not only discover Japanese culture and traditions but also experience those of other countries.YISH is also a special place for my family in which we welcomed our first son born in Japan.

Read the paper · More papers on PaperTik