AI — Chispa and VON
Chispa classifies. VON generates.
Chispa: microseconds, no daemon
Section titled “Chispa: microseconds, no daemon”A ~1 MB linear model for decisions that do not deserve an LLM: classifying an event, routing a request.
kling ai chispa train -data train.jsonl -valid valid.jsonl -o events.chispakling ai chispa predict -model events.chispa -text "panic in the parser" -fields '{"service":"api"}'Output (one JSON line)
{"label":"fix","index":5,"prob":0.393,"threshold":0.517,"confident":false,"decision":"escalate", …}Every answer says confident or escalate. It answers in 1.5–6 µs, with no memory allocations, bit for bit the same on amd64 and arm64.
VON: small models that scale to zero
Section titled “VON: small models that scale to zero”SmolLM2-360M, Qwen2.5 0.5B and 1.5B, served with llama-server from a golden with the model already loaded.
kling ai model add von-smol -model smollm2-360m-instructkling run -from von-smol -name smol-1kling ai model ask smol-1 "What is a microVM?"461 MiBPSS of 4 SmolLM2 replicas (their RSS adds up to 1715)
~0.8 sfirst token from the golden on a Mac (vz)
~140 tok/sSmolLM2 in a vz guest with 2 vCPUs
The AI gateway
Section titled “The AI gateway”kling ai servekling ai test commit-type "fix crash when the cache is cold"kling ai eval commit-type -data test.jsonl -von qwenAn OpenAI-compatible API. The cascade (VON answers what Chispa is unsure about) only turns on if kling ai eval proves that it wins.