A Sound-Synthesis Technique Based on Multidimensional Scaling of Spectra
Christophe Hourdin, G. Charbonneau, Tarek M. Moussa · Computer Music Journal · 1997
Many methods for sound synthesis have been developed since Max Mathews's first digital synthesis experiments in 1958 (Mathews 1969). Generally, there are two kinds of synthesis techniques: those that have many control parameters, such as additive synthesis, and others that use only a few parameters, such as frequency modulation (FM) synthesis (Chowning 1973) or waveshaping synthesis (Arfib 1979; LeBrun 1979). It seems obvious that the more parameters one has, the more accurately the soundsynthesis method can be controlled. The main attraction of FM synthesis is its ability to create complex sounds with few control parameters, but this advantage is counterbalanced by limited control of the sound. On the other hand, using many parameters for the description of a tone does not guarantee, from a perceptual point of view, flexibility of sound control or high quality. To be useful to a musician, a synthesis method must allow the sound to be controlled by a limited number of factors that correspond to well-defined perceptual features. As David Jaffe has pointed out, the quality of a synthesis is often judged in a very intuitive manner (Jaffe 1995). All synthesis starts from an explicit or implicit physical description of sound, but such a description is often of little relevance to the musician. Ideally, a synthesis method would build the sound from a perceptual description. Common-practice Western music notation does not describe sound sufficiently, because it does not define the timbre. Thus, if the musician cannot describe the sound in either physical or musical terms, it is necessary to find an intermediate way (Ethington and Punch 1994). One possible solution to this problem is to elaborate a perceptual model of sound, including timbre, that allows one to resynthesize the sound. In another article in this issue (Hourdin, Charbonneau, and Moussa 1997), we presented a physical analysis of sound that we believe leads to a representation of timbre. We started with spectrum analysis data and further analyzed the data using a multidimensional scaling technique, yielding a space in which each axis corresponds to some acoustical property. We found that a sound's perceived timbre is directly associated with the shape of the corresponding curve in the space. Such a curve can be considered a simple representation of the timbre, one that can be easily and intuitively used by a musician to describe a sound. In this article, we present the corresponding sound-synthesis technique, based on the same representation of sound. We will first summarize the analysis of sound data, which yields a database from which we can rebuild the original sounds or synthesize new ones. We will study the influence of each axis of the database, as well as the quality of the sound reconstruction with a limited number of axes. Finally, we will discuss the intuitive and physical means of control in this new method of sound synthesis.