Build production-ready AI agent systems.
Agenkit is a lightweight, cross-language toolkit for building distributed AI agents that scale from prototype to production without rewriting your code.
Building production AI agent systems is hard:
- Reliability: LLMs fail unpredictably - you need circuit breakers, retries, and timeouts
- Scale: Prototypes work locally but break in production when you need distributed deployment
- Observability: Understanding what went wrong requires distributed tracing across services
- Language Lock-in: Python is great for prototyping, but you need Go/Rust for performance
- Integration: Every agent framework has its own incompatible abstractions
Agenkit provides the production infrastructure you need:
Prototype → Production
┌─────────────────────────────────────────────┐
│ Your Agents (Python for development) │
└──────────────┬──────────────────────────────┘
│
│ Same Interface
↓
┌─────────────────────────────────────────────┐
│ Agenkit Toolkit │
│ • Automatic retries & circuit breakers │
│ • Distributed tracing & metrics │
│ • Cross-language support (Python ↔ Go) │
│ • Multiple transports (HTTP/gRPC/WS) │
└─────────────────────────────────────────────┘
Key Insight: Write your agents once in Python. Deploy them in Go for up to 18x lower framework/transport overhead in our benchmarks (see Performance below) — the LLM call itself still dominates wall-clock time. Same interface, zero rewrites.
Create a simple agent:
from agenkit import Agent, Message
class MyAgent(Agent):
@property
def name(self) -> str:
return "my-agent"
async def process(self, message: Message) -> Message:
return Message(
role="agent",
content=f"Processed: {message.content}"
)
# Use it
agent = MyAgent()
response = await agent.process(Message(role="user", content="Hello!"))from agenkit.middleware import RetryConfig, RetryDecorator, CircuitBreakerDecorator
# Wrap with resilience
production_agent = RetryDecorator(
CircuitBreakerDecorator(agent),
RetryConfig(max_retries=3),
)
# Now handles failures automatically
response = await production_agent.process(message)# Run locally
python my_agent.py
# Deploy with Docker
docker-compose up
# Scale with Kubernetes
kubectl apply -f deploy/kubernetes/class Agent:
name: str # Unique identifier
async def process(msg) -> Message # Process messagesThat's it. Everything else is optional.
from agenkit.middleware import (
CircuitBreakerConfig,
CircuitBreakerDecorator,
RetryConfig,
RetryDecorator,
TimeoutConfig,
TimeoutDecorator,
)
# Start simple
agent = MyAgent()
# Add retry logic
agent = RetryDecorator(agent, RetryConfig(max_retries=3))
# Add circuit breaker
agent = CircuitBreakerDecorator(agent, CircuitBreakerConfig(failure_threshold=5))
# Add timeouts (timeout_ms for clarity)
agent = TimeoutDecorator(agent, TimeoutConfig(timeout_ms=30000))
# Stack as many as you needEvery middleware is <100 lines. Easy to understand, modify, or replace.
Write once. Deploy anywhere. Nine language implementations (Python is the
reference; all nine share the same Agent/Message/Tool core and most of
the 18 core patterns — C#, Java, and Scala are missing AgentsAsTools; see
Status for the exact per-language spec-conformance breakdown):
# Python - Prototype quickly with ML ecosystem
class MyAgent(Agent):
async def process(self, message):
return process_with_python_libs(message)// TypeScript - Full-stack with browser support
class MyAgent implements Agent {
async process(message: Message): Promise<Message> {
return processWithTypeScriptLibs(message);
}
}// Go - Production scale (up to 18x lower transport overhead vs Python in our benchmarks)
type MyAgent struct{}
func (a *MyAgent) Process(ctx context.Context, msg *Message) (*Message, error) {
return processWithGoLibs(msg)
}// C++ - Maximum performance with zero-overhead abstractions
class MyAgent : public Agent {
public:
Message process(const Message& msg) override {
return process_with_cpp_libs(msg);
}
};// Rust - Memory safety + performance (up to 20x lower transport overhead vs Python in our benchmarks)
struct MyAgent;
impl Agent for MyAgent {
async fn process(&self, message: Message) -> Result<Message, AgentError> {
process_with_rust_libs(message).await
}
}// Zig - Systems programming with safety (up to 22x lower transport overhead vs Python in our benchmarks)
const MyAgent = struct {
pub fn process(self: *MyAgent, message: Message) !Message {
return processWithZigLibs(message);
}
};Same interface across all languages. Python, Go, and Rust are the most complete; C#, Java, and Scala are newer and still filling in advanced subsystems (skills, some adapters). Choose the right tool for each service.
from agenkit.observability import TracingMiddleware, init_tracing
# Enable tracing
init_tracing("my-service")
# Wrap agent
traced_agent = TracingMiddleware(agent)
# Every operation now traced in JaegerView traces: http://localhost:16686
# Inspect agent state during debugging or production
result = agent.introspect()
print(f"Agent: {result.agent_name}")
print(f"Capabilities: {result.capabilities}")
print(f"Internal state: {result.internal_state}")
print(f"Memory: {result.memory_state}")
# Use for monitoring
def check_agent_health(agent):
result = agent.introspect()
return result.internal_state.get("error_count", 0) < 10
# Use for testing
def test_agent_state():
result1 = agent.introspect()
await agent.process(message)
result2 = agent.introspect()
assert result2.internal_state["msg_count"] > result1.internal_state["msg_count"]Introspection ≠ Reflection:
- Introspection (this feature): Examines current state ("What do I know?")
- Reflection (pattern): Analyzes past performance ("How did I do?")
from agenkit.adapters.python import GRPCServer, LocalAgent
from agenkit.adapters.python.http_server import HTTPAgentServer
# Same agent, different transports
await HTTPAgentServer(agent, port=8080).start() # REST API
await GRPCServer(agent, "localhost:50051").start() # High performance
await LocalAgent(agent, endpoint="ws://127.0.0.1:8765").start() # Bidirectional streamingChoose the right protocol for each use case.
- Building production AI agent systems that need to scale
- Teams that prototype in Python but deploy in Go
- Systems that need resilience (retries, circuit breakers, timeouts)
- Distributed agent systems that need observability
- Integrating multiple AI agent frameworks
- Simple one-off scripts (too much overhead)
- Embedded systems (requires async runtime)
- Real-time systems (<1ms latency requirements)
# Python
pip install agenkit
# Go
go get github.com/scttfrdmn/agenkit-go
# TypeScript/Node.js
npm install @agenkit/core
# C++
# Clone and build (CMake required)
git clone https://github.com/scttfrdmn/agenkit.git
cd agenkit/agenkit-cpp && mkdir build && cd build && cmake .. && make
# Rust
cargo add agenkit
# Zig
# Clone and build (Zig 0.12+ required)
git clone https://github.com/scttfrdmn/agenkit.git
cd agenkit/agenkit-zig && zig build# Install only what you need
pip install agenkit[openai] # OpenAI (GPT-4, GPT-3.5)
pip install agenkit[anthropic] # Anthropic (Claude 3.5)
pip install agenkit[aws] # AWS Bedrock
pip install agenkit[google] # Google Gemini
pip install agenkit[ollama] # Ollama (local models)
# Multiple providers
pip install agenkit[openai,anthropic]
# All providers
pip install agenkit[all-providers]
# With Redis memory backend
pip install agenkit[redis]
# Everything (for development)
pip install agenkit[all]See INSTALLATION.md for complete installation guide.
- Minimal interfaces - Agent, Message, Tool (50 lines total)
- Orchestration patterns - Sequential, Parallel, Router, Fallback
- Type safety - Full type hints (Python), compile-time checks (Go)
- Circuit Breaker - Prevent cascading failures
- Retry - Exponential backoff with jitter
- Timeout - Request deadline enforcement
- Rate Limiter - Token bucket algorithm
- Caching - LRU cache with TTL
- Batching - Request aggregation
- HTTP - REST APIs (HTTP/1.1, HTTP/2, HTTP/3)
- gRPC - High-performance binary protocol
- WebSocket - Bidirectional streaming
- Distributed Tracing - OpenTelemetry integration
- Metrics - Prometheus endpoints
- Structured Logging - JSON logs with trace correlation
- Memory Management - Context retention with compression strategies
- Budget Tracking - Token counting and cost optimization
- Checkpointing - State persistence and recovery
- Safety - Prompt injection detection, input/output validation
- Evaluation - Quality metrics and benchmarking
Transport Overhead: <1% of total time in realistic LLM workloads
Language Performance (framework/transport overhead, not general application performance):
- Go HTTP transport: 18.5x lower latency than Python's in the same benchmark (0.055ms vs 1.02ms)
- Middleware overhead: <0.01% of request time
These figures — and the "18x"/"20x"/"22x" multipliers mentioned earlier in
this README for Go/Rust/Zig — were measured in November 2025 against Python
3.14.0 and Go 1.21–1.22 on Apple Silicon (see benchmarks/BASELINES.md).
Both the Go toolchain and several dependency versions have since moved, so
the exact multiplier should not be read as current; the qualitative
conclusion that transport/framework overhead is a small fraction of a
100–1000ms LLM call has held across every measurement to date. See
COMPATIBILITY.md for the
full caveat and how to regenerate current numbers.
Scale:
- Kubernetes autoscaling: 3-10 replicas based on load
- Message size scaling: 10,000x size = 190x latency (linear)
See benchmarks/BASELINES.md for detailed performance data.
- Cross-language integration tests (Python ↔ Go ↔ C++)
- Chaos engineering tests (network failures, crashes)
- Property-based tests (invariant validation)
- Thousands of unit and integration tests across 9 languages (Python alone: 2200+; see Test Parity below for the current per-language breakdown)
- Input validation
- Prompt injection detection
- RBAC and audit logging
- Non-root containers
- Dropped capabilities
- Jaeger distributed tracing
- Prometheus metrics
- Structured JSON logging
- Health check endpoints
- Docker images
- Kubernetes manifests with HPA
- Horizontal pod autoscaling (3-10 replicas)
- Zero-downtime rolling updates
┌─────────────────────────────────────────┐
│ Your Application (Agents + Logic) │
└──────────────┬──────────────────────────┘
│
┌──────────────┴──────────────────────────┐
│ Optional Middleware Layer │
│ (Add only what you need) │
│ • Retry • Caching │
│ • Timeout • Batching │
│ • Circuit Breaker │
└──────────────┬──────────────────────────┘
│
┌──────────────┴──────────────────────────┐
│ Optional Transport Layer │
│ • HTTP • gRPC • WebSocket │
└──────────────┬──────────────────────────┘
│
┌──────────────┴──────────────────────────┐
│ Core Interface (Required) │
│ • Agent • Message • Tool │
└─────────────────────────────────────────┘
Design Philosophy: Start simple. Add complexity only when you need it.
- Getting Started (15 min) - Choose your language and create your first agent
- Agent Patterns Book (2-3 hours) - Master the 18 core patterns and advanced architectures
- Advanced Architectures (1 hour) - Pattern compositions and multi-agent systems
- Examples (1 hour) - Learn by doing with 27+ examples
- Architecture (30 min) - Understand design principles
- Migration Guides (30 min per language) - Migrate between languages
- Production Deployment (1 hour) - Docker + Kubernetes
Total time investment: ~6 hours from zero to production deployment with pattern mastery.
- Python Guide - 15-30 min to first agent
- Go Guide - Idiomatic Go patterns
- TypeScript Guide - Modern JavaScript/TypeScript
- Rust Guide - Memory-safe agents
- C++ Guide - High-performance systems
- Zig Guide - Systems programming with safety
- Agent Patterns Book - Comprehensive guide to 18 core patterns
- Advanced Architectures - Pattern compositions and multi-agent systems
- Architecture Principles - Design philosophy
- Streaming Patterns - Language-specific streaming approaches
- Migration Guide Index - Complete migration documentation for all 6 languages
- API Reference - Complete API documentation
- Deployment Guide - Docker and Kubernetes
- Examples - 27+ comprehensive examples
- Security Policy - Vulnerability reporting and best practices
- Compatibility Matrix - Language, platform, and version support
- Testing Strategy - Test philosophy and coverage
- Memory Management - Context retention strategies
- Budget Tracking - Token and cost management
- Checkpointing - State persistence
- Safety & Security - Input validation
- Evaluation - Quality metrics
We provide 27+ examples covering common use cases:
Getting Started:
- Basic Agent - Your first agent in 20 lines
- Sequential Pipeline - Chain agents together
- Parallel Execution - Run agents concurrently
Production Features:
- Retry & Circuit Breaker - Handle failures
- Distributed Tracing - Debug with Jaeger
- Rate Limiting - Protect your APIs
Advanced:
- Remote Agents - Cross-process communication
- Streaming Responses - Server-sent events
- Website: agenkit.dev
- GitHub: github.com/scttfrdmn/agenkit
- Issues: Report bugs or request features
- Discussions: Ask questions
We welcome contributions! See our Contributing Guide.
# Clone repository
git clone https://github.com/scttfrdmn/agenkit.git
cd agenkit
# Install dependencies
pip install -e ".[dev]"
# Run tests
pytest tests/ -v
# Run type checking
mypy agenkit/Apache License 2.0 - See LICENSE for details.
v0.89.0 — 9 language implementations, Python reference (see the root VERSION
file for the current number — this line will drift again if hand-maintained)
The shared core is the Agent/Message/Tool interface plus 18 named
patterns (listed below). The "Patterns implemented" column is a
spec-presence conformance score from spec-conformance.json
(scripts/parity/spec_conformance.py, #909/#913) — for each of the 18
patterns named in specs/patterns/*.yaml, does a source file implementing
it actually exist in that language, checked by filename/class-name (with an
alias table for legitimate naming divergence, e.g. Python's
memory.py/MemoryHierarchy vs. C#'s MemoryAugmentedAgent.cs) rather than
by counting classes. This replaced an earlier raw class-count column
(feature-manifest.json's patterns count) that answered a different,
noisier question — it could not distinguish "no implementation exists" from
"the implementation has an unconventional class name," and undercounted
C#/Java/Scala for months due to a scanner bug (#918) before this metric
existed. feature-manifest.json remains as a secondary, diagnostic
class-count signal; it is no longer this table's source. The "Tests" column
is the test count from the parity report; "Depth" reflects how many advanced
subsystems (memory, skills, reasoning memory, full adapter set) are
implemented.
| Language | Patterns implemented | LLM Adapters | Tests | Depth |
|---|---|---|---|---|
| Python | 18/18 | 7 | 2229 | Reference — all subsystems |
| Go | 18/18 | 7 (+vLLM, SGLang) | 1341 | Complete — incl. reasoning memory, skills |
| TypeScript | 18/18 | 7 | 976 | Broad — no skills / reasoning memory |
| Rust | 18/18 | 6 | 1352 | Complete — incl. skills |
| C++ | 18/18 | 5 | 1133 | Broad — safety/ not yet implemented |
| Zig | 18/18 | 8 | 671 | Broad — no skills |
| C# (.NET) | 17/18 (missing AgentsAsTools) |
2 (+mock) | 272 | Newer — no skills |
| Java | 17/18 (missing AgentsAsTools) |
2 (+mock) | 358 | Newer — no skills |
| Scala | 17/18 (missing AgentsAsTools) |
mock only | 363 | Newest — LLM adapters are stubs |
18 Core Patterns documented in the Agent Patterns Book: AgentsAsTools, Autonomous, Collaborative, Conversational, Fallback, HumanInLoop, MemoryHierarchy, MultiAgent, Orchestration, Parallel, Planning, ReAct, ReasoningWithTools, Reflection, Router, Sequential, Supervisor, Task.
Six languages (Python, Go, TypeScript, Rust, C++, Zig) implement all 18; C#, Java, and Scala implement 17 of 18 (missing AgentsAsTools).
- ✅ Defect repair (v0.89) — 82 commits fixing release-gate and CI bugs found by a fleet audit (SBOM/signing, version-declaration drift, a test gate that couldn't fail)
- ✅ Typed cross-language token
Usage(v0.86–v0.87) — unifiedUsagestruct across all 9 language cores, plus Bedrock prompt-cache token counts - ✅ Agent Skills (v0.85) —
AgentSkill,SkillRegistry,SkillEnabledAgent(Python, Go, Rust)
v0.88.0 is intentionally skipped — reserved for the observability milestone (#715). See
CHANGELOG.md for the full release history.
- ✅ Core toolkit across 9 languages; 18 shared patterns in 6 of them, 17 of 18 in C#/Java/Scala (see Status)
- ✅ MCP (Model Context Protocol) client/server in every language
- ✅ Production middleware (retry, circuit breaker, timeout, rate limiting, caching, batching)
- ✅ Multiple transports (HTTP/1.1, HTTP/2, HTTP/3, gRPC, WebSocket)
- ✅ Deployment manifests (Docker + Kubernetes with HPA)
- 🚧 Advanced-subsystem parity (skills, reasoning memory, full adapter sets) still landing in the JVM/.NET tier
Ready to build production AI agents? Visit agenkit.dev or get started in 5 minutes →
Patterns are at full parity across languages; test-count parity varies — the
secondary languages have fewer tests than the Python reference. Counts are
regenerated via scripts/test-parity.sh (and surfaced by the Parity Validation
CI workflow). Relative to Python's 2229 tests:
| Language | Tests | vs Python |
|---|---|---|
| Rust | 1352 | 61% |
| Go | 1330 | 60% |
| C++ | 1133 | 51% |
| TypeScript | 976 | 44% |
| Zig | 671 | 30% |
| Java | 358 | 16% |
| Scala | 363 | 16% |
| C# | 272 | 12% |
Counts are regenerated via scripts/test-parity.sh.