Projects / Open Source Package

Open Source Package

local-llm-router

If you experiment with local models seriously, you end up running more than one local server without meaning to. Ollama for the models it handles best.

If you experiment with local models seriously, you end up running more than one local server without meaning to. Ollama for the models it handles best. LM Studio because a particular quantization was easier to load there. llama.cpp directly, for the models neither of the above wrapped well yet. Each one binds its own port, exposes its own slightly-different API surface, and every tool you point at "your local model" needs to be told, specifically, which port and which shape to expect. And that pointer breaks the moment you switch which runtime happens to be running that day.

local-llm-router exists to make that pointer stop breaking. It probes the known ports, Ollama on 11434, LM Studio on 1234, llama.cpp on 8080, Text-Generation-WebUI on 5000, Jan on 1337, figures out which ones are actually alive right now, and runs a small Express proxy that exposes one stable, OpenAI-compatible endpoint forwarding to whichever backend is currently active. Point every tool at the router's port, once, and never update that pointer again regardless of which local runtime you're actually running underneath.

How detection actually works

detectBackends() fires probes at all five known ports in parallel, via Promise.allSettled with a two-second timeout per probe, so a dead port doesn't stall the whole scan waiting for a connection that's never coming. Each probe is genuinely two-tiered, which is the one piece of this I'd call properly considered rather than quick-and-dirty: it tries /v1/models first, the OpenAI-compatible endpoint most of these runtimes expose, and only falls back to /api/tags, Ollama's native, non-OpenAI-compatible endpoint, if the first one 404s. That fallback exists because Ollama didn't always implement the OpenAI-compatible surface as completely as the others, and without the fallback, an Ollama instance running an older version would show up as "not detected" even though it's right there, listening, just not answering the request shape I checked first. Small resilience touch, but it's the difference between the tool working reliably across the actual versions people run versus only working in the one configuration I originally tested against.

The part where the README gets ahead of the code

I want to be direct about this one. The README describes a routing strategy, fastest, round-robin, capability-aware, fallback, like the router makes an intelligent choice about which backend should handle which request based on real signals. It doesn't. The actual code picks allBackends[0], whatever showed up first in the scan, as the active backend, full stop, and the only way to change that is a manual POST /router/use/:id call to switch which one is active. There's no load balancing. There's no capability matching. There's no "fastest" anything. I wrote the strategy section describing the router I wanted to eventually build, and shipped the router that actually does the one useful thing I'd built so far: detect what's running and give you a stable single endpoint to talk to it.

I'm not going to dress that up as more than it is. The gap between "intelligent multi-backend routing" and "detect-and-proxy-to-whichever-one's-first" is real, and if someone read only the README they'd expect meaningfully more sophistication than what's actually running. What's there is genuinely useful on its own, a stable endpoint that survives you switching which local runtime you're using that day is not nothing, but it's the honest floor of the idea, not the ceiling I described.

Why plain mutable state, and why that's fine here

activeBackend and allBackends are held in ordinary module-level variables, not wrapped in a class, not persisted anywhere. For a single-process CLI tool that you start, use, and stop, that's a completely reasonable choice: there's exactly one instance of this router running at a time, on one machine, for one person. Wrapping that in more formal state management would be solving a concurrency problem that doesn't exist in the actual use case. I bring this up because it's the kind of decision that looks like corner-cutting out of context and looks like appropriate scoping once you know what the tool is actually for. Not every piece of state needs a class around it. Some of it just needs to be a variable that lives as long as the process does, because that's exactly as long as it needs to live.

The CLI, and why it exists alongside the library

llm-router start -p <port> boots the proxy server directly from the terminal, the fast path for "I just want this running." llm-router scan runs detection without starting anything, useful for a quick "what do I actually have running right now" check before deciding whether to bother starting the proxy at all. Both wrap the same underlying detectBackends() and createRouter() functions the library exposes, so the CLI isn't a separate implementation. It's a thin, convenient shell around the same code any project could import directly.

Why this matters more than its size suggests

This is the piece of infrastructure that makes "build against local models" a real, sustainable practice instead of a constant configuration headache. Every time I switch which local runtime I'm experimenting with, usually because a specific model loads better in one than another, every other tool I've built that talks to "the local LLM" keeps working, unmodified, because they're all pointed at the router, not at a specific port that happens to belong to whichever runtime I was using last week.

What actually pushed me to build this instead of just remembering the port

The honest trigger wasn't some abstract appreciation for good infrastructure. It was switching from Ollama to LM Studio for one specific model that loaded more reliably there, and then spending ten minutes hunting through three different project configs to update a hardcoded localhost:11434 to localhost:1234 in each one, only to switch back a week later and do the whole thing in reverse. That's a genuinely small amount of friction on any single occasion, and a genuinely large amount of friction across a year of doing it repeatedly without a system for it. The router isn't solving a hard technical problem. It's solving my own inability to remember, reliably, which port I was supposed to be pointing at this week, which is a real problem even though it sounds like a silly one to build a tool for.

Why the manual backend switch, and not automatic failover

POST /router/use/:id requiring a deliberate action to change the active backend, rather than automatically falling back to a different one if the active backend goes quiet, was a conscious choice. Automatic failover sounds strictly better until you think about what it actually means for a local development setup: if my active backend crashes mid-request and the router silently swaps to a different one running a completely different model, the next response I get back might come from a model with different capabilities, different context length, different everything, without me knowing that happened. For local experimentation, where I usually care specifically which model I'm testing, silent failover is closer to a footgun than a feature. I'd rather have the router fail loudly when the active backend is unreachable, so I notice and fix it, than have it quietly paper over the problem by routing to whatever else happens to be listening.

Full stack developer. Founder of Yashveer Labs. One port, whichever local model actually answered first.

Start a conversation about this.

Whether it's local-llm-router itself or the next system worth building, the lab is reachable.