llm-stream
Every LLM provider streams responses differently, at the protocol level, in ways that have nothing to do with the model underneath and everything to do with whichever engineering team designed that provider's API.
Every LLM provider streams responses differently, at the protocol level, in ways that have nothing to do with the model underneath and everything to do with whichever engineering team designed that provider's API. OpenAI streams server-sent events with data: prefixed lines and a [DONE] sentinel to mark the end. Anthropic streams SSE too, but with a completely different event vocabulary, content_block_delta, message_stop, and expects your system prompt pulled out into its own top-level field instead of living in the message list like every other provider does it. Ollama doesn't use SSE at all; it streams newline-delimited JSON over a plain /api/chat endpoint. Three providers, three protocols, and if you want to build something that can swap between local and hosted models without a rewrite, you end up writing three parsers.
llm-stream is those three parsers, written once, hidden behind one interface: stream(opts) for the raw async generator, complete(opts) if you just want the whole response collected into a string, streamTo(opts, onChunk) for callback-style consumption when a generator doesn't fit the rest of your code. Every provider comes back through the same StreamChunk shape: text, a done flag, and the raw underlying payload if you need to reach past the abstraction for something provider-specific.
Zero dependencies, on purpose
The whole thing is two hundred and sixty-two lines, one file, and it runs on nothing but the native fetch API and hand-rolled stream parsing. No SDK from any provider. That's a deliberate constraint, not an accident of laziness. Every official SDK brings its own response shape, its own retry behavior, its own opinions about how you should structure a request, and pulling in three different SDKs just to get three different providers' text out as a stream felt like exactly the wrong trade for something conceptually this simple. Reading raw bytes off a fetch response and parsing them by hand is more work upfront. It's also the only way to guarantee the exact same output shape regardless of which provider is on the other end, because I control the entire parsing path myself instead of inheriting three different SDKs' three different ideas about what a "chunk" is.
Why local-first mattered enough to build this
This exists because I wanted to build things, early Nyxera experiments especially, that could run against a local Ollama model during development and switch to a hosted provider for anything that needed more capability, without the actual application code caring which one it was talking to. Local-first isn't just a cost decision for me, though it is that too. It's also about being able to iterate on a conversational feature at 2am without a network dependency, and about not sending every experimental prompt to a third party while I'm still figuring out whether the idea even works. llm-stream is the piece of plumbing that makes "swap the provider, keep the code" actually true instead of aspirational.
What I promised that isn't there
The README lists token counting per chunk and error handling with retries as features. Neither exists in the code. StreamChunk has no token count field anywhere. I never wired up tokenizer logic for any of the three providers, which would have meant either calling a provider-specific counting endpoint or bundling a tokenizer library, and I decided the complexity wasn't worth it for what was meant to be a lightweight utility. And there's no retry logic at all: a dropped connection or a malformed chunk just propagates as an error to whatever's consuming the generator, full stop. I wrote both of those into the README as intended features, thinking ahead to what a "real" version of this package would eventually have, and then didn't build either one before publishing. Same pattern as a couple of the other tools in this batch: the vision got written down before the code caught up to it, and the code stopped a step short of the vision.
The duplication I know about and haven't fixed
Each of the three provider branches, OpenAI, Anthropic, Ollama, has its own buffered line-reading loop, handling the specific framing quirks of that provider's stream format. They're structurally almost identical: read a chunk, split on newlines, hold back the last incomplete line for the next read, parse whatever's complete. I wrote that loop three times, with small variations, instead of extracting one shared line-buffering primitive that each provider's parser could call into. It would be a better-factored library if I had. It isn't a worse library in practice, because the three loops work correctly and independently, and modifying one to fix a bug in, say, the Anthropic event parsing doesn't risk breaking the OpenAI path the way a shared abstraction gone wrong might. Sometimes duplication is a smell. Sometimes it's just three genuinely different things that happen to look similar at a glance, and I've landed on treating this as closer to the second case, at least for now.
Where it actually lives in my stack
llm-stream sits directly underneath local-llm-router in how I actually use them together. The router figures out which local backend is running and exposes a stable OpenAI-compatible endpoint, and llm-stream is what actually speaks to that endpoint, or to a hosted provider directly, and normalizes whatever comes back. Neither package depends on the other explicitly, but they were built to be used side by side, which is a pattern across most of this batch: small, independently useful pieces that happen to click together cleanly when I need the fuller picture.
The Anthropic system-prompt quirk that cost me the most debugging time
Of the three providers, Anthropic's protocol was the one that actually surprised me while building this. OpenAI and Ollama both let you put a system message directly in the message array, same as any other turn. Anthropic doesn't. It wants the system prompt pulled out entirely and sent as its own top-level system field, separate from the messages array, and if you send it the OpenAI-style shape instead, it either errors or silently ignores your system instructions depending on exactly how you got it wrong. I lost a genuinely frustrating stretch of time early on to a case where the model just wasn't following its system prompt at all, and the actual bug was that I was building the request the same way for all three providers, which worked for two of them and quietly broke the third. stream()'s Anthropic branch now extracts the system message out of the array before constructing the request, specifically because of that one debugging session, a real scar from a real mistake, not a feature I planned for from the start.
Why raw exists on every chunk
Every StreamChunk carries a raw field alongside the normalized text and done fields, the actual, unprocessed payload from whichever provider sent it. I added that after the first time I needed something provider-specific that the normalized shape didn't expose, a finish reason, a usage count buried in a final chunk, something that varies enough between providers that baking it into the common interface would have meant either omitting it entirely or inventing a lowest-common-denominator field that fit none of the providers well. raw is the escape hatch: use the normalized shape for the common case, reach into raw for anything provider-specific without needing me to have anticipated exactly what you'd need in advance.
Full stack developer. Founder of Yashveer Labs. One shape in, no matter which provider is actually talking.
Start a conversation about this.
Whether it's llm-stream itself or the next system worth building, the lab is reachable.