Skip to content

Latest commit

 

History

167 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
Shimmy Logo

Shimmy — Local Inference, OpenAI-Compatible

🔒 The 5MB alternative to Ollama — 100% Rust, zero dependencies 🚀

License: MIT CI Crates.io Downloads Rust GitHub Stars

Languages: 简体中文 · 繁體中文

Shimmy is independently maintained and free forever. Sponsorship funds certification, compatibility work, and releases.


What Is Shimmy?

Shimmy is a single-binary OpenAI-compatible inference server for GGUF models. Point your existing AI tools at Shimmy and they just work — locally, privately, and free.

Shimmy is the server. Airframe is the engine. Under the hood, Shimmy runs on Airframe (v0.4.0), a pure-Rust WebGPU (WGSL) transformer engine. No C++ toolchain, no Python runtime, no backend flags. 26 models certified across 12 families. Version history: CHANGELOG · Airframe CHANGELOG.

Why this matters:

  • No Python runtime or C++ toolchain — Rust only, top to bottom
  • F32 accumulation precision with deterministic output (same model + seed + params → same output)
  • WGSL compute shaders via WebGPU — NVIDIA, AMD, Intel, integrated GPUs, Apple Silicon
  • Model spec auto-derived from GGUF metadata — no hardcoded per-model constants
  • YaRN RoPE scaling for extended context via SHIMMY_MAX_CTX (see Extended Context)

🎯 Supported Models

12 model families · 26 certified model/quant combinations — every model below passes Shimmy's 3-box certification regimen (MATH + INFERENCE + DETERMINISM) against the certification ledger. Certification applies to the named model/quant combination; architecture recognition does not automatically mean certification. GGUF files load as-is; no recompilation, no hardcoded per-model constants.

Family Model Quants
Llama Llama-3.2-1B-Instruct Q4_K_M · Q6_K
Llama-3.2-3B-Instruct Q4_K_M
Llama-3.1-8B-Instruct Q4_K_M
TinyLlama-1.1B-Chat Q4_0 · Q5_K_M · Q6_K
Qwen3 Qwen3-0.6B Q4_K_M
Qwen3-1.7B Q4_K_M
Qwen3-4B Q4_K_M
Qwen3-4B-Thinking Q4_K_M
Qwen3-8B Q4_K_M
Qwen2 Qwen2-0.5B-Instruct Q4_K_M
Qwen2-1.5B-Instruct Q4_K_M
Qwen2-7B-Instruct Q4_K_M
Qwen3.5 Qwen3.5-9B Q4_K_M
Phi-3 Phi-3.5-mini-Instruct Q4_K_M
Phi-3-mini-4k-Instruct Q4_0
Phi-2 Phi-2 Q4_K_M
Gemma-2 Gemma-2-2B-it Q4_K_M
Gemma-2-9B-it Q4_K_M (supported; cert: see v2-roadmap)
Gemma-4 Gemma-4-12B-coder Q4_K_M
Gemma-4-E4B Q4_K_M
DeepSeek-R1 DeepSeek-R1-0528-Qwen3-8B Q4_K_M
Ministral Ministral-3-14B-Reasoning Q4_K_M
StarCoder2 StarCoder2-3B Q4_K_M

SafeTensors format (.safetensors) is supported for model loading via safetensors_native. Full Airframe-native inference for SafeTensors remains roadmap work; see docs/v2-roadmap.md.

Features

  • TurboShimmy INT4 KV Cache — About 7× lower KV-cache memory in tested configurations. Run Llama-3.2-3B on 4 GB GPUs.
  • 🚀 OpenAI SDK Compatibility — Chat completions, text completions, streaming, and model endpoints. Works with OpenAI SDKs and tools using that surface.
  • 🔧 Extended Context — YaRN RoPE scaling via SHIMMY_MAX_CTX.
  • 📦 Migrating from v1.x — llama.cpp, MLX, HuggingFace, and RustChain backends removed in v2.0+. Shimmy is now a pure Airframe product.
  • 🏆 Certification — Every model passes a 3-box certification regimen (MATH + INFERENCE + DETERMINISM). See docs/CERTIFICATION.md.

Quick Start

cargo install shimmy
shimmy serve --model-path /absolute/path/to/model.gguf --bind 127.0.0.1:11435

Then in another terminal:

shimmy list --short
curl -s http://127.0.0.1:11435/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model":"tinyllama-1.1b","messages":[{"role":"user","content":"Say hi in 5 words."}],"max_tokens":32}'

Full install, model acquisition, GPU, VRAM sizing, platform-specific builds: docs/quickstart.md


Documentation

Start here What you need
Quick Start Install, models, GPU, VRAM
Supported Models Certified models and quantization
API Compatibility Endpoints, SDKs, integration
Configuration Env vars and config options
Troubleshooting GPU errors, model failures
Complete documentation index
Section Documents
Models & Performance TurboShimmy — INT4 KV cache compression · Extended Context — YaRN RoPE scaling, VRAM math · Performance — Tuning and token/sec · Model Expansion — Onboarding protocol
API & Integration API Reference · Integration Guides · Examples · Cross-Compilation
Engine Architecture · GPU Pipeline · Quantization · Chat Templates
Certification Certification · Methodology · Regression Testing · PPT Testing · Metrics
FAQ FAQ · Features · Migration · Windows GPU

Development Testing

Shimmy maintains high code quality through comprehensive testing:

# Full test suite (default features = GPU engine)
cargo test --features airframe,huggingface

# Quick CPU-only tests (no GPU required)
cargo test --lib --no-default-features --features huggingface -- --test-threads=1

See docs/ppt-invariant-testing.md for technical details.


Community & Support

📰 As Featured On

🔥 Hacker News · IPE Newsletter


Performance

Tool Startup Memory API
Shimmy <1s ~50MB Chat, completions, streaming, models
Ollama 5-10s 200MB+ Partial

Measured on RTX 3060, Shimmy v2.6.0, TinyLlama-1.1B. Your results vary by hardware.


Sponsor Shimmy

Shimmy is independently maintained. Sponsorship funds certification, compatibility work, and releases.

  • $5/month: Coffee tier ☕ — Sponsor badge + name in SPONSORS.md
  • $25/month: Supporter 🐛 — Priority support + name in SPONSORS.md
  • $100/month: Corporate backer 🏢 — Logo placement + release recognition
  • $500/month: Infrastructure partner 🚀 — Office hours + roadmap consultation

Current sponsors: ZephyrCloudIO · gqf2008 · alistairheath

🎯 Become a Sponsor · Invoicing


License & Philosophy

MIT License — see LICENSE. Shimmy will be free forever.

Shimmy is infrastructure: it should be invisible. Reliability through comprehensive validation and property-based testing.


Maintainer: Michael A. Kuykendall · Mission: Making local model inference simple and reliable


Trans rights are human rights. Shimmy is built by and for everyone — discrimination has no place in our community or our code.

About

⚡ Pure-Rust WebGPU inference engine — OpenAI-API compatible, GGUF native, runs on any GPU. No Python. No llama.cpp. Single binary.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

5.8k stars

Watchers

40 watching

Forks

Releases

Sponsor this project

Packages

Used by

Contributors

Languages