Projects / Open Source Package

Open Source Package

kokoro-js

Kokoro is a genuinely good text-to-speech model: small, fast, free, and open enough to run entirely on your own hardware instead of paying per character to a cloud API.

Kokoro is a genuinely good text-to-speech model: small, fast, free, and open enough to run entirely on your own hardware instead of paying per character to a cloud API. The problem is it lives in the Python ecosystem, and I build mostly in Node. Every time I wanted to give something a voice (a prototype assistant, a narrated demo, a piece of Nyxera's earlier voice experiments) I hit the same wall: a good model, no clean way to call it from JavaScript.

kokoro-js is the bridge. Not a JavaScript or WASM port of the model itself (that would be a much bigger undertaking than a solo side project justifies), but a thin orchestration layer that shells out to a Python subprocess actually running Kokoro, generates a WAV file into the system temp directory, and reads it back into a Buffer your Node app can use. pip install kokoro soundfile underneath, npm install kokoro-js on top.

The sentence-boundary trick

The feature I'm actually proud of here is synthesizeStream(). LLM output arrives as a stream of tokens, not a finished sentence, and naive text-to-speech either waits for the whole response to finish (which kills the feeling of a live conversation) or tries to speak partial, grammatically broken fragments, which sounds wrong in an obvious way. synthesizeStream is an async generator that buffers incoming text as it arrives and flushes to the TTS engine only when it detects a real sentence boundary, using a lookbehind regex that splits on . , ! , or ? without consuming the punctuation. The result is audio that starts playing meaningfully sooner than "wait for the whole response," while still speaking complete, natural sentences instead of fragments. Piping a streaming LLM response straight into this and hearing the first sentence spoken while the model is still generating the second one is one of those small technical satisfactions that's disproportionate to how much code it actually took.

How the Python bridge actually works, mechanically

buildPythonScript() constructs a literal Python script as a template string, with your text and options interpolated in, and pipes the whole thing to python -c. It's not calling into a persistent Python process or a proper IPC bridge: every call spins up a fresh interpreter, runs the script, and tears it down. That's simpler to reason about and more resilient (a crashed synthesis attempt can't corrupt a long-lived process state, because there is no long-lived process), at the cost of paying Python's startup overhead on every single call. For occasional narration or assistant responses that overhead is invisible. For anything wanting rapid-fire TTS calls back to back, it would start to matter, and I haven't optimized for that case because it hasn't been the case I've actually needed.

Duration comes back by parsing a duration_ms: line out of the subprocess's stdout with a regex, genuinely fragile in the sense that if Kokoro's Python side ever changes its print formatting, this silently stops working rather than failing loudly. It's worked every time I've used it so far, which is a different claim than "it's robust," and I want to be honest about which of those two things is actually true here.

The dependency that does nothing

axios is declared as the package's only dependency. It's never imported anywhere in the source. I must have started building this expecting to need an HTTP call somewhere (maybe an early version fetched the model differently, or I was thinking ahead to a hosted fallback that never got built) and the dependency just never got removed once the actual implementation didn't need it. It's dead weight in the package.json, harmless but genuinely pointless, and exactly the kind of thing that accumulates in a project when you're moving fast enough that cleaning up unused imports isn't the next thing on your mind after the feature works.

Where the README oversells it

The documented public API is a KokoroTTS class: new KokoroTTS({model}), then .synthesize(), .stream() as methods on an instance. That class doesn't exist. The real exports are free functions: synthesize() and synthesizeStream(), called directly, no instantiation required. I wrote the docs imagining a slightly more object-oriented shape than what I actually built, and never reconciled the two. Anyone copying the README's example verbatim would get an error on the first line (KokoroTTS is not a constructor or similar), which is a bad first experience for a package with, realistically, an audience of one.

Why it still earns a place in the toolkit

Local TTS matters to me for the same reason local LLM routing does: cost and control. Cloud TTS APIs are metered per character and require sending whatever you're narrating to a third party. Kokoro running locally, wrapped just enough to be callable from Node, means I can prototype anything voice-related (a narrated walkthrough, an early assistant experiment) without a running bill and without a network dependency. It's rough around the edges. It also works, every time I've reached for it, for exactly the thing I built it to do.

Why a fresh interpreter every call, instead of a long-lived one

I thought seriously about keeping a single Python process alive in the background: loading the Kokoro model once, then feeding it text over stdin repeatedly, instead of paying model-load time on every single call. That would be faster for sustained use. I didn't build it that way, because a long-lived process introduces a whole category of state I didn't want to manage for a small utility: what happens if it crashes mid-synthesis, how do I know it's still alive, what happens to a call that was in-flight when it died. The fresh-interpreter-per-call approach means every synthesis attempt starts from a known-clean state and can't be corrupted by whatever happened on the previous call. It's slower. It's also nearly impossible to get into a broken state that requires restarting anything, because there's nothing long-lived to get into a broken state in the first place. For occasional narration rather than a live, rapid-fire voice pipeline, I'd take that trade again.

The VIRTUAL_ENV detection, and why it exists

The package auto-detects a VIRTUAL_ENV environment variable to pick the right Python executable, rather than always shelling out to a bare python and hoping it resolves to the right interpreter. This came from a real, specific annoyance: I have multiple Python environments on my machine for different projects, and Kokoro's dependencies aren't installed globally. They're installed in a specific virtual environment set up for exactly this purpose. Without the detection, calling python -c from inside a Node process would resolve to whatever python happens to mean in that shell's PATH, which is not reliably the environment with kokoro and soundfile actually installed. Checking VIRTUAL_ENV first and using that environment's interpreter if it's set closes that gap, quietly, without requiring the caller to configure anything extra in the common case.

Full stack developer. Founder of Yashveer Labs. A good model, a thin bridge, and a sentence-boundary trick that makes the wait feel shorter than it is.

Start a conversation about this.

Whether it's kokoro-js itself or the next system worth building, the lab is reachable.