A DT-Neural Parametric Violin Synthesizer
Muhammad Nizami, Dessi Puji Lestari · 2021
In the field of musical sound synthesis, expressive synthesis is generally more preferred for many genres. One of highly demanded musical instrument synthesis is for bowed string instruments. Currently, there is no fully automated expressive synthesis for string instruments, only partially automated which needs expressive gestures as additional inputs. We present a model for expressive violin synthesis adapted from neural parametric singing synthesis technique. Using neural parametric technique allows the system to learn expressive gestures from dataset, without the need for expression labeling. We also modified the system to better match the domain of bowed string instruments. We use harmonic plus stochastic encoding as the output to better represent bowed instrument sound. We also use decision tree (DT) instead of neural network to generate the expressive timing deviations. We compare our system to existing RPM synthesis technique as our baseline, without expression as input, or represented by flat expression. Correlation coefficient metric shows that the proposed system is superior in learning expressive gesture patterns. However, listener's preference scores show that the baseline RPM technique is more preferred in terms of naturalness. Blind open-ended questions reveal that the proposed system is deemed more natural in terms of timing deviations, pitch deviations, dynamics variations, phrasing, and some other minor things. The listener's preference scores are highly inclined against the proposed system because the proposed system produces unrealistic timbre. Future works should be done to improve the timbre while maintaining superiority in timing, pitch, dynamics, and phrasing.