A thin Bun.js library that brings local LLM inference to JavaScript via llama.cpp. Runs GGUF-format language models locally (Llama 3, Mistral, Gemma, DeepSeek, ChatML) with streaming token generation, entirely from JavaScript/Bun.
| Layer | File | Purpose |
|---|---|---|
| FFI shim | src/ffi/shim.c |
Thin C wrapper around libllama (compiled to lib/llama_shim.dylib) |
| FFI binding | src/ffi/llama.js |
Bun dlopen binding — exposes the C shim as a JS class (LlamaFFI) |
| Inference engine | src/inference/engine.js |
InferenceEngine — loads a model, manages context/batch/sampler, yields tokens as an async generator |
| KV cache | src/cache/kv-cache.js |
LRU cache keyed by FNV-1a hash of system-prompt tokens; snapshots and restores KV state to skip re-encoding repeated system prompts |
- Uses Bun's built-in FFI (
bun:ffi) — no Node.js native addons, no Python bridge generate()is an async generator that streams text pieces token-by-token- Chat template auto-detection from model metadata or filename (supports chatml, llama3, mistral, gemma, deepseek)
- Metal/GPU acceleration supported via
nGpuLayersand compiled with-framework Metal vendor/llama.cppis a git submodule;build.shhandles cmake + clang compilation with a Homebrew fallback
Option A — vendored submodule (recommended):
git submodule add https://github.com/ggerganov/llama.cpp vendor/llama.cpp
./build.shOption B — Homebrew:
brew install cmake llama.cpp
./build.shimport { InferenceEngine, KVCache } from "llama-bun"
const engine = new InferenceEngine({
modelPath: "/path/to/model.gguf",
nGpuLayers: 32, // 0 for CPU-only
nCtx: 4096,
nThreads: 4,
flashAttn: false,
verbose: false,
})
await engine.load()
const cache = new KVCache({ maxEntries: 4 })
const messages = [
{ role: "system", content: "You are a helpful assistant." },
{ role: "user", content: "Hello!" },
]
for await (const piece of engine.generate(messages, { temperature: 0.7, topP: 0.9, maxTokens: 512, kvCache: cache })) {
process.stdout.write(piece)
}
engine.unload()| Option | Default | Description |
|---|---|---|
modelPath |
— | Path to a .gguf model file |
nGpuLayers |
0 |
Number of layers to offload to GPU |
nCtx |
4096 |
Context window size in tokens |
nThreads |
4 |
CPU threads for inference |
flashAttn |
false |
Enable Flash Attention |
verbose |
false |
Enable llama.cpp logging |
| Option | Default | Description |
|---|---|---|
temperature |
0.7 |
Sampling temperature |
topP |
0.9 |
Top-p (nucleus) sampling |
maxTokens |
2048 |
Maximum tokens to generate |
seed |
random | RNG seed for reproducibility |
kvCache |
— | KVCache instance for system-prompt caching |
import { LlamaFFI, getLlama } from "llama-bun" // raw FFI binding
import { InferenceEngine } from "llama-bun/engine" // inference engine
import { KVCache } from "llama-bun/cache" // KV state cache
import { LlamaFFI } from "llama-bun/ffi" // FFI only