ROBUST PITCH MARKING FOR PROSODIC MODIFICATION OF SPEECH USING TD-PSOLA

Wesley Mattheyses, Werner Verhelst, Piet Verhoeve · VUBIR (Vrije Universiteit Brussel) · 2006

ABSTRACT TD-PSOLA is a well-known technique for prosodic modica-tionofspeechsignals,especiallyforpitchshiftingandtimescal-ing. The quality of the modication results can be very high, butcritically depends on the determination of the individual pitchperiods (epochs) in the speech signal. As speech is a naturallyproduced signal, the robust estimation of pitch epochs thus be-comes extremely important for TD-PSOLA applications, espe-cially where real-time operation is required. This paper pro-poses an efcient algorithm for robust pitch epoch detection thatis amenable for real-time implementations. 1. INTRODUCTIONThe time domain pitch synchronized overlap-add algo-rithm(TD-PSOLA)[1]iswellknownintheeldofspeechsynthesis as it allows for high quality pitch and time scalemodications of stored speech segments and has a verylow complexity and computational cost. However, it isalso well known that the sound quality of TD-PSOLAmodied speech is very sensitive to a proper position-ing of the pitch marks that delimit the individual pitchepochs. Therefore, TD-PSOLA has been mostly used inapplications such as text-to-speech synthesis, where thepitch marking can be done off-line and corrected manu-ally. Unfortunately, hand correcting pitch marks is notvery enjoyable, but time consuming and expensive. Fur-thermore,bylackofrobustsimplepitchmarkingmethods,many interesting real-time applications for pitch shiftingand time stretching (e.g., karaoke applications) have re-mained problematic with TD-PSOLA. For these reasons,the search for better pitch marking methods has remainedan active research area.In a recent publication [2], a simple robust pitch markingmethod was presented that could be amenable for real-time implementation. Like many modern techniques, thismethodrstselectsasetofpossiblepitchmarkcandidatesin voiced speech segments. The nal sequence of pitchmarks is obtained from these candidates as the sequencethat optimises a functional, which depends on the con-dence in the individual pitch mark candidates and on thediscrepancybetweentheinstantaneouspitch(measuredasthe difference between consecutive pitch marks) and theaverage pitch contour as obtained from a standard pitchdetection algorithm. Unfortunately, the method gives noindication as to how to place the pitch marks in unvoicedsegments,ontheborderbetweenvoicedandunvoicedseg-ments or in mixed-excitation segments. We found that, ifother than voiced only speech is used, the overall qualityof TD-PSOLA can strongly depend on a proper strategyforpitchmarkinginotherthanthepurelyvoicedsegmentsas well. In this case, using the same technique [2] for allthe segments lead to rather poor results. Additionally, [2]did not discuss how the different initial pitch mark candi-dates should be selected in each segment.In section 2 of this paper, we briey review how we per-formed the initial voiced/unvoiced classication and aver-age pitch detection for the input speech segments. Then,in section 3, we present the detailed strategy for pitchmarking the voiced segments including our strategy forselecting a proper set of initial pitch mark candidates. Insection 4, we propose an original and effective approachfor pitch marking the other speech segments (unvoicedand mixed-excitation). Finally, section 5 presents an eval-uation of the results and section 6 concludes the paper.2. PITCH DETECTION ANDVOICED-UNVOICED DECISION2.1. Voiced-unvoiced decisionWe start by segmenting the signal into frames, on whichweperformavoicetypeclassicationwithacomputation-ally simple method that is based on the local energy andthe amount of zero-crossings that occur in the consideredframe [3].

Read the paper · More papers on PaperTik