Skip to content

Latest commit

 

History

3 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

llama-bun

A thin Bun.js library that brings local LLM inference to JavaScript via llama.cpp. Runs GGUF-format language models locally (Llama 3, Mistral, Gemma, DeepSeek, ChatML) with streaming token generation, entirely from JavaScript/Bun.

Architecture

Layer File Purpose
FFI shim src/ffi/shim.c Thin C wrapper around libllama (compiled to lib/llama_shim.dylib)
FFI binding src/ffi/llama.js Bun dlopen binding — exposes the C shim as a JS class (LlamaFFI)
Inference engine src/inference/engine.js InferenceEngine — loads a model, manages context/batch/sampler, yields tokens as an async generator
KV cache src/cache/kv-cache.js LRU cache keyed by FNV-1a hash of system-prompt tokens; snapshots and restores KV state to skip re-encoding repeated system prompts

Design

  • Uses Bun's built-in FFI (bun:ffi) — no Node.js native addons, no Python bridge
  • generate() is an async generator that streams text pieces token-by-token
  • Chat template auto-detection from model metadata or filename (supports chatml, llama3, mistral, gemma, deepseek)
  • Metal/GPU acceleration supported via nGpuLayers and compiled with -framework Metal
  • vendor/llama.cpp is a git submodule; build.sh handles cmake + clang compilation with a Homebrew fallback

Build

Option A — vendored submodule (recommended):

git submodule add https://github.com/ggerganov/llama.cpp vendor/llama.cpp
./build.sh

Option B — Homebrew:

brew install cmake llama.cpp
./build.sh

API

import { InferenceEngine, KVCache } from "llama-bun"

const engine = new InferenceEngine({
  modelPath: "/path/to/model.gguf",
  nGpuLayers: 32,   // 0 for CPU-only
  nCtx: 4096,
  nThreads: 4,
  flashAttn: false,
  verbose: false,
})

await engine.load()

const cache = new KVCache({ maxEntries: 4 })

const messages = [
  { role: "system", content: "You are a helpful assistant." },
  { role: "user",   content: "Hello!" },
]

for await (const piece of engine.generate(messages, { temperature: 0.7, topP: 0.9, maxTokens: 512, kvCache: cache })) {
  process.stdout.write(piece)
}

engine.unload()

InferenceEngine options

Option Default Description
modelPath — Path to a .gguf model file
nGpuLayers 0 Number of layers to offload to GPU
nCtx 4096 Context window size in tokens
nThreads 4 CPU threads for inference
flashAttn false Enable Flash Attention
verbose false Enable llama.cpp logging

generate() options

Option Default Description
temperature 0.7 Sampling temperature
topP 0.9 Top-p (nucleus) sampling
maxTokens 2048 Maximum tokens to generate
seed random RNG seed for reproducibility
kvCache — KVCache instance for system-prompt caching

Exports

import { LlamaFFI, getLlama } from "llama-bun"        // raw FFI binding
import { InferenceEngine }     from "llama-bun/engine" // inference engine
import { KVCache }             from "llama-bun/cache"  // KV state cache
import { LlamaFFI }            from "llama-bun/ffi"    // FFI only

About

A thin Bun.js library that brings local LLM inference to JavaScript via llama.cpp (https://github.com/ggerganov/llama.cpp)

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages