01 / ProjectVulkan · C · GLSL

PUMICE.

A Vulkan-based LLM inference engine written from scratch in C and GLSL compute shaders. It runs Qwen3.5 and Qwen3.6 MoE checkpoints locally on consumer GPUs — no PyTorch, no CUDA, no llama.cpp.

$npm i -g @h4zel/pumice
3
Models supported
Qwen3.5 2B · 9B · Qwen3.6 35B-A3B MoE
2
GPUs validated
AMD RX 580 8 GB · NVIDIA RTX 4060
3
Quant formats
FP16 · INT8 · INT4, per layer
0
ML frameworks
No PyTorch, no CUDA, no llama.cpp

Under the hood

BUILT FROM THE METAL UP

Hybrid attention

Gated delta-net and full-attention layers implemented to match the reference, with runtime-generalized GQA head mapping and partial RoPE.

Per-layer quantization

Every layer picks its own FP16, INT8, or INT4 mix, with 32 to 256 element Q4 blocks. The web UI edits the mix before a load.

MoE expert offloading

Routed experts are split between VRAM and host RAM pools, so the 35B-A3B model runs on an 8 GB card at around 10.7 tok/s.

Chunked prefill + KV cache

Prompts are processed in 512-token chunks, and a tiered KV block cache spans VRAM, host RAM, and disk, so multi-turn chats resume instead of re-prefilling.

On-GPU sampler

Temperature, top-k, top-p, min-p, and repetition penalty run in a compute shader. Greedy decoding uses an on-device argmax.

Agent-ready API

The same server exposes a /v1 API with streaming, tool calls, reasoning, and usage. Point an agent harness at it and the model auto-loads on the first request.

Ships with a web UI

LOCAL, PRIVATE, SELF-HOSTED

The npm package bundles a prebuilt engine binary and a local React interface: a model loader with validation, a per-layer quantization editor, sampling controls, and a streaming chat playground. The UI opens automatically at 127.0.0.1:8787.

pumice — server
$pumice
uihttp://127.0.0.1:8787
modelsscan a folder, pick a checkpoint
quantfp16 / int8 / int4, per layer
api/v1 · streaming · tool calls
cachekv resume, no re-prefill
>explain the kv cache in one line

RUN IT
LOCALLY.

Windows x64, Node.js 18+, and a Vulkan-capable GPU with a current driver. Install it globally and run it from a folder with your models.