Janus, a single Go binary, runs GGUF models locally via Vulkan

Janus, a single Go binary, runs GGUF models locally via Vulkan

Janus is an open-source project shared on Show HN. Its README describes it as a single Go binary that runs .gguf models on your own machine, on a GPU or a CPU, and exposes an OpenAI-compatible API. The README states that no Python, no Docker and no Ollama are required.

Inference is done by llama.cpp. On the GPU it goes through Vulkan, which the README lists as covering AMD, Intel and NVIDIA cards; if there is no Vulkan, a CPU fallback is available. The API offers two endpoints, /v1/chat/completions and /v1/models, and streaming is shown in the examples. You can call it from the command line (curl, PowerShell, scripts) or wire it into Cursor, Cline or any OpenAI client, using the same local models.

The feature list also names hot-swapping models (change the .gguf file without restarting), support for thinking models (the reasoning is split into a separate reasoning_content field), and automatic chat template detection from the GGUF metadata. On Windows the README promises zero dependencies: one .exe.

Setup on Windows: clone the repository and run build.ps1, which downloads pre-built llama.cpp Vulkan DLLs and compiles janus.exe into the dist folder. Then put a .gguf file in the models folder. A bundled downloader, modelget, fetches files from Hugging Face; the README's example pulls Llama-3.2-3B-Instruct-Q8_0.gguf from meta-llama/Llama-3.2-3B-Instruct, and any GGUF from Hugging Face works too. Next, copy .env.example to .env and set INFERENCE_BACKEND=vulkan, JANUS_MODEL_PATH to the model file and, in the example, JANUS_MAX_TOKENS=4096. Running janus.exe opens http://127.0.0.1:8990 in the browser, and curl against /health checks that the server is up.

On Linux the path is a build from source: go mod tidy, go build, copy .env.example, set JANUS_MODEL_PATH, and set INFERENCE_BACKEND=cpu if there is no Vulkan. The README adds that libllama.so must sit next to the binary or on LD_LIBRARY_PATH.

In clients, the base URL is http://127.0.0.1:8990/v1 and the API key is left blank. The README also gives disk guidance: plan for the model size, often 2 to 8 GB per model, plus about 50 MB for Janus and llama.dll. The repository layout includes cmd/janus (the main server), cmd/modelget (the downloader), internal/engine (the llama.cpp Vulkan and CPU backend), internal/bridge (DLL loader and FFI bindings) and internal/singleton (a single-instance guard). The project is MIT licensed and contributions are welcome.

Key facts

  • Janus is a single Go binary that runs .gguf models on a GPU or CPU and serves an OpenAI-compatible API (/v1/chat/completions and /v1/models); no Python, Docker or Ollama needed.
  • Inference uses llama.cpp via Vulkan (AMD, Intel, NVIDIA) with a CPU fallback.
  • Features listed: hot-swapping .gguf files without a restart, thinking-model support with reasoning split into reasoning_content, and chat template auto-detection from GGUF metadata.
  • Disk budget per the README: the model (often 2 to 8 GB) plus about 50 MB for Janus and llama.dll; the local server listens on 127.0.0.1:8990.
  • MIT licensed. Windows is built with build.ps1; Linux is a build-from-source path that needs libllama.so.

Why it matters

Running open-weight models locally normally means picking a runtime and its dependencies. Janus aims to cut that down to one Go binary that speaks the OpenAI API, so tools such as Cursor and Cline can point at a local model. Its use of Vulkan means one GPU path for AMD, Intel and NVIDIA cards rather than a vendor-specific one. The source does not say how Janus differs in performance or quality from existing runtimes, so the appeal rests on simplicity of setup.

Who it affects

People who want to run GGUF models on their own Windows or Linux machine, especially those with AMD or Intel graphics who want Vulkan acceleration, and developers who want a local endpoint for OpenAI-compatible editors and scripts. Anyone who prefers to avoid a Python or Docker stack is the target of the README's pitch.

How to use it

On Windows: clone the repository, run build.ps1 (it downloads pre-built llama.cpp Vulkan DLLs and builds janus.exe), put a .gguf file in the models folder or fetch one with modelget, copy .env.example to .env, set INFERENCE_BACKEND and JANUS_MODEL_PATH, and run janus.exe. The server opens http://127.0.0.1:8990 in your browser. In an OpenAI client, set the base URL to http://127.0.0.1:8990/v1, leave the API key blank and use the model name "local" as in the README examples. On Linux, build from source with go build, set INFERENCE_BACKEND=cpu if there is no Vulkan, and place libllama.so next to the binary or on LD_LIBRARY_PATH. The project is MIT licensed.

How solid is it

The only source is the project README, which describes features but offers no evidence for them. No benchmarks, tokens-per-second figures or comparisons with llama.cpp, Ollama or LM Studio are given. No release date, version number or GitHub star count is given. No list of tested GPUs, drivers or supported models beyond a Llama-3.2-3B-Instruct example is given. No authors, company or team behind Janus are named; Vibra-Ingenn appears only in the repository URL.

Risks and caveats

Linux support is described only as a build-from-source path; no prebuilt Linux binary is mentioned, and the README says libllama.so must be supplied. There is no statement about macOS support; only Windows and Linux instructions appear. Hardware and driver compatibility is unverified beyond the README's claim of AMD, Intel and NVIDIA via Vulkan. Model files are large, often 2 to 8 GB each, so disk space matters. The server runs locally on port 8990, and the examples tell clients to leave the API key blank.

“No Python, no Docker, no Ollama required.”

— Janus README