Training-Free Watermarking for Flow Matching TTS
Today's digest spotlights AudioNoisePrints, a model-free watermarking technique that embeds provenance signals directly into flow matching TTS synthesis by exploiting spatial correlations in the initial noise — no training, no quality loss.
General overview of our approach ( ) comparing to the commonly used post-hoc watermarking schemes. As shown above, the post-hoc watermarking model required altering the audio after it is generated, which creates computation overhead and alters the quality of the audio. Our method simply changes the initial noise $ _init$ during generation (without altering the quality of the generated audio), and uses either the model-free option (calculating the cosine similarity of the audio-Mel-spectrogram and the initial noise) or the more robust detector option to determine if the audio is in fact watermarked. From Independent Research.