AI Vision: Smart speaker design and implementation with object detection custom skill and advanced voice interaction capability

Bharath Sudharsan, Sree Prem Kumar, Rakesh Dhakshinamurthy · 2019

Amazon has provided a cloud-based developer console to write and deploy custom skills that can integrate with Alexa Voice Service device SDK running on the host hardware. When this deployed custom skill is invoked using a specific voice request, the host hardware performs the requested task and sends back its response. Currently, developers have implemented AI algorithms as a part of their custom skills on existing Alexa devices which focuses on voice-based applications and not the real-time image or video-based applications since current Alexa devices are not camera enabled. Hence there is a need to design an advanced camera-enabled Alexa smart speaker platform and provide it to the open-source community to facilitate implementing image/video based AI algorithms as a part of their custom skills. This article describes the design and development of a state-of-the-art camera-enabled, Linux-based modern Alexa smart speaker prototype. A microphone array with on-chip advanced DSP is interfaced with the prototype to provide speech algorithms processed voice data to Alexa Voice Service for achieving a seamless, full-duplex user-Alexa interaction. Finally, the main Deep Learning & Computer Vision-based object detection custom skill is realized and tested on the developed prototype.

Read the paper · More papers on PaperTik