Ask HN: Are there AI models for generating sounds based on a text and reference?

15 points | by onemiketwelve 1 day ago

2 comments

  • xg15 29 minutes ago
    Maybe a dumb idea, but how good are image models with manipulating spectrogram images? Then a workaround could be: convert input audio to spectrogram -> pass spectrogram + prompt to an image model -> convert modified spectrogram back to audio waveform.
    • Buttons840 20 minutes ago
      I wouldn't call it a dumb idea, but there's soooo much subtlety to sound that wont be visible in any reasonably sized image.
  • narrationbox 43 minutes ago
    Plenty of models take in text + audio and spits out audio. It's the format of most newer generation accent conversion/voice cloning models.

    What's your exact use case?