Mobidic - a mobile dictation and notetaking application

Markku Turunen, Aleksi Melto, Anssi Kainulainen, Jaakko Hakulinen · 2008

Abstract Mobile devices have become ubiquitous and reasonably po-werful and well connected. However, their physical size limits possibilities of interaction, especially document creation. Dictation in mobile setting provides one solution, but limited processing power requires that the actual speech recognition is distributed to a server computer. We present MobiDic, a distributed mobile dictation application. Its user interface solutions solve problems that arise from the technical limita-tion of mobile devices and mobile user context. Index Terms : mobile applications, dictation, audio editing 1. Introduction Dictation is one of the few real success stories of speech rec-ognition technology, and speech in general can be very effi-cient data entry method [1], although there are several chal-lenges for its use [2]. Traditional dictation applications are mostly used by specific professions, e.g., doctors and lawyers, in acoustically safe environments and on computationally adequate hardware. However, people are more and more on the move most of their day, and have immediate access only to less capable hardware, both computationally and interac-tion-wise. Furthermore, their schedule is often fragmented, so they do not have time for long-lasting dictation sessions to dictate in the traditional sense. People perceive mobile dicta-tion to be very useful, but there are also great demands for its accessibility and usability [3][4]. Thus, mobile solutions al-lowing incremental dictation and notetaking are necessary to meet the current demands. However, mobile platforms, and standard mobile phones in particular, are problematic for dictation applications be-cause of lack of processing power. There are only initial pro-totypes of embedded large-vocabulary speech recognition [5]. Most advanced dictation systems need more hardware capa-bilities than is available in current and near-future mobile phones. Because of these reasons we have focused on distri-buted solution, where the dictation and editing can be done with a mobile device but ASR takes place in a server ma-chine. Furthermore, the server stores user audio recordings and recognized texts, and maintains acoustic models and dis-tributes generated document, e.g., via email. The limited interaction capabilities offered by mobile de-vices form another fundamental issue. In particular, small screens and keypads make it hard to design efficient docu-ment editing interfaces. Most of the desktop applications are assuming standard sized screens and multimodal interaction, in particular when error correction is considered [6]. Further-more, most of them are targeted for single session, real-time interaction, not asynchronous and fragmented interaction as in mobile dictation applications. There are some commercial asynchronous mobile dicta-tion solutions available. Typically, the user dictates while on move, and the recognition takes place afterwards. In this case, the process is not interactive, and the user is not able to re-view, edit and distribute the results in the mobile settings. We present MobiDic, an asynchronous dictation and note-taking application for mobile phones. In this paper, we de-scribe the MobiDic application, its functionality and architec-ture, and then discuss the user interface solutions for incre-mental dictation, audio editing, recognition results editing, and user adaptation. Figure 1.

Read the paper · More papers on PaperTik