LLM Orchestrator
A working replication of NVIDIA's “Small Language Models are the Future of Agentic AI” — their six-step LLM→SLM conversion pipeline, run against my own agent telemetry. Replaying every routing decision in hindsight put 57% of local-tier spend within reach of a 35B model on my own hardware, inside the paper's predicted 40–70% band.
- Python
- SLM research replication
- llama.cpp · 35B local
- Claude API
- evals + telemetry