Under the hood
BUILT FROM THE METAL UP
Hybrid attention
Gated delta-net and full-attention layers implemented to match the reference, with runtime-generalized GQA head mapping and partial RoPE.
Per-layer quantization
Every layer picks its own FP16, INT8, or INT4 mix, with 32 to 256 element Q4 blocks. The web UI edits the mix before a load.
MoE expert offloading
Routed experts are split between VRAM and host RAM pools, so the 35B-A3B model runs on an 8 GB card at around 10.7 tok/s.
Chunked prefill + KV cache
Prompts are processed in 512-token chunks, and a tiered KV block cache spans VRAM, host RAM, and disk, so multi-turn chats resume instead of re-prefilling.
On-GPU sampler
Temperature, top-k, top-p, min-p, and repetition penalty run in a compute shader. Greedy decoding uses an on-device argmax.
Agent-ready API
The same server exposes a /v1 API with streaming, tool calls, reasoning, and usage. Point an agent harness at it and the model auto-loads on the first request.
Ships with a web UI
LOCAL, PRIVATE, SELF-HOSTED
The npm package bundles a prebuilt engine binary and a local React interface: a model loader with validation, a per-layer quantization editor, sampling controls, and a streaming chat playground. The UI opens automatically at 127.0.0.1:8787.
RUN IT
LOCALLY.
Windows x64, Node.js 18+, and a Vulkan-capable GPU with a current driver. Install it globally and run it from a folder with your models.