This is really cool work! I'm curious like what do you see as the biggest lever for speeding up TTS models or from a technical perspective that this was a promising direction in the first place to push on. If I were to guess, some distillation but I'm certain there are probably TTS model aware architectural changes that just make inference wayyyy faster?
https://github.com/loudreader/loudkit
I think real time natural tts should be possible everywhere soon
For OP the clip name is nari-nina-01a0a12f-980a-765e-8029-fa56bd23210d.wav
Is the ASR inference engine open source as well?