5 comments

  • rao-v 14 minutes ago
    This is a really neat use of the per-layer embedding trick. It's also worth noting that there viable TTS models that are ~20-30M param, so it might mean you can have a ESP32 with no network access read stuff out to you in near real time!
  • chrishynes 12 minutes ago
    Why can't this scale to run much larger models on CPU backed by flash with good access patterns?
  • althea_tx 26 minutes ago
    This is a really cool project. Thanks for sharing!
  • cr125rider 38 minutes ago
    9.7 tokens/sec actually seems like a lot! That’s fun!