Composite-DBN for recognition of environmental contexts
Selina Chu, Shrikanth Shri Narayanan, C.‐C. Jay Kuo · Asia-Pacific Signal and Information Processing Association Annual Summit and Conference · 2012
People's behaviors are usually dictated by their surroundings. The surrounding environment affects the character and disposition of the people within it. The goal of our work is to automatically recognize the type of environments one is in. In this paper, we introduce a hierarchical structure to recognize environments using the surrounding audio. We can use this structure to discover high-level representations for different acoustic environments in a data-driven fashion. Being able to perform such function would allow us to better understand how we could utilize such information to assist in predicting a person's emotion or behavior. To accurately make an informative decision about behaviors or emotions, it is important to have the ability to differentiate between different types of environments. Environmental sound contains large variances even within a single environment and is constantly changing. These changes and events are dynamic and inconsistent. The goal is to come up with models that is robust enough to generalize to different situations. Learning a hierarchy of sound types would improve and clarify problems caused by the confusion between multiple acoustic environments with similar characteristics. We propose a framework for a composite of deep belief networks (composite-DBNs) as a way to represent various levels of representations and to recognize twelve different types of common everyday environments. Experimental results demonstrate promising performance in improving the state of art recognition for acoustic environments.