I see Claude implemented a very crude upsampling/downsampling algorithm, which is what LLMs usually do when prompted to handle such a problem. But I would suggest restraining the model from implementing DSP processing on their own and instead use battle tested libraries. You can use rubato's FFT Resampler.
Audio processing is genuinely a hard engineering problem, LLMs usually don't get it right. If you decide to get deep into it, the knowledge you'll get is very rewarding.
There's so many of these, at this point I've seen 10 clones make the front page each time as if there never existed local only options before.
Also whisper is pretty outdated vs parakeet
It's dead simple. Hold the fn key, speak and release. I use a quantized Wisper small.en model for transcription. It inserts the text into the active application. There's also a hands-free model for longer dictation. Audio transcription is kept in memory. There's no account or transcription history. Clipboard contents are restored after it inserts it. GPLv3, Mac-only, English only..still in alpha. Hope you enjoy it! Would love some feedback.
- havent seen a single one on HN yet in the last year (i read HN twice a day like brushing my teeth)
- When I input my voice into the mic, I want an AI voice as output converting my words in real time in AI voice
- Use case: gaming, I have a terrible voice and dont want to do a voice over with that but at the same time I would love to if I could
- Know any github projects capable of pulling this off? maybe direct integration as an OBS plugin would make it godtier
I think what I was more alluding to was the engineering value such a project brings. But once again, sorry about the messaging.