Need help?
<- Back

Comments (11)

  • yjftsjthsd-h
    Couple highlights:> Complete local text-to-waveform speech synthesis under 10M parameters.In case, like me, you hoped "complete" voice might mean both stt and tts. Not to speak poorly of it, just clarifying.> English only, with one fixed male voice. This is not zero-shot voice cloning.(And then a bunch of statements on limitations that I read as 'quality can be spotty but if you play with it it should be fine') But like. In <10M params I'm not judging:)
  • modinfo
    This is amazing, the quality blow my mind for such small model! I just replaced my old onnx model with yours!here my implementation with speech dispatcher and server: https://github.com/skorotkiewicz/inflect-speechdthanks for shearing!
  • NetOpWibby
    The inflections are weird but this doesn't sound like a robot. Not bad!
  • tmaly
    This is impressive. I wish there were a voice clone option.
  • itake
    Amazing quality for small size, but definitely not that enjoyable to listen to.IMHO, its at about the same quality level of historic TTS tools.
  • jsomedon
    amazing quality for such small size!
  • mcbetz
    Alternative title: Text to speech in 9.36M, English only.