Staff AI/ML Engineer · I build the system layer that makes AI products actually work.
AI runtime & agent harnesses · headless developer platforms (MCP / SDK / CLI) · applied LLMs & Document AI
Blog · LinkedIn · Email · Publications
I'm a Staff AI/ML Engineer with 8+ years turning research-grade models into production platforms that serve millions of users. I work on the unglamorous layer that most AI demos skip — the runtime, harness, evals, and developer surface that make AI reliable, safe, and easy to build on.
- 🏗️ Currently at HighLevel, where I architected the unified AI runtime + agent harness behind 6+ product surfaces (Voice AI, Conversation AI, Agent Studio, Ask AI, Ads Copilot, Email Agent).
- 📈 That platform runs 15M+ conversations / 800B+ tokens per quarter across 250K+ active locations at 99.9% success.
- 🔌 I built HighLevel's headless developer platform —
mcp,agent-sdk,cli,plugins— exposing 600+ operations over MCP so anyone can automate via Claude, ChatGPT, Cursor, or any standard agent. - 📚 Published at EMNLP 2024 and Elsevier ASOC (×2) on layout-aware Document AI and parameter-efficient fine-tuning.
I care about the boundary where ML research meets production systems: model routing, tool/function calling, long-context state, memory, eval gates, and the reliability engineering that keeps it all up.
AI Platforms agent runtime · agent harnesses · MCP · model routing · tool/function calling
hybrid RAG · memory · long-context systems · Voice AI · copilots
Model Quality evals · prompt optimization · post-training · LoRA/QLoRA/PEFT · vLLM serving
Document AI · OCR · layout-aware extraction
Engineering Python · TypeScript · PyTorch · Hugging Face · FastAPI · Node.js · SDK/CLI design
Infra & Data AWS · GCP · Azure · Docker · Kubernetes · Postgres · Redis · OpenSearch · vector search
| Where | What I did | Outcome |
|---|---|---|
| HighLevel | Unified AI runtime + agent harness for 6+ surfaces | 15M+ convos · 800B+ tokens/qtr · 99.9% success |
| HighLevel | Headless dev platform over MCP (mcp/sdk/cli/plugins) |
600+ operations, no-UI automation |
| HighLevel | Retell-based real-time voice execution path | 4.7M+ voice calls at 99.9%+ |
| Zania AI | Auto Compliance Agents (SOC2 / NIST / HIPAA / TPRA) | evidence F1@10: 86% → 97% |
| Intellect AI | Layout-aware Document AI extraction stack | replaced 23+ models · –70% Textract spend · 6M+ PDFs |
I write about building production AI systems — runtimes, agent harnesses, evals, and applied LLM engineering — at kforcode.dev.
- The system layer behind AI products: what an agent runtime actually does
- From 86% to 97%: engineering evidence retrieval for compliance agents
- One model instead of 23: consolidating a Document AI stack
- EMNLP 2024 — De-Identification of Sensitive Personal Data in IIT-CDIP-Derived Datasets — aclanthology.org
- Elsevier ASOC — Parameter-Efficient Fine-Tuning for Hospital Discharge Summarization — sciencedirect.com
- Elsevier ASOC — DoRA+: Enhancing Weight-Decomposed Low-Rank Adaptation — sciencedirect.com
Open to Staff / Principal AI Engineering roles · AI platforms & applied LLMs · remote or relocation.

