Skip to content

AI — Chispa and VON

Chispa classifies. VON generates.

A ~1 MB linear model for decisions that do not deserve an LLM: classifying an event, routing a request.

Terminal window
kling ai chispa train -data train.jsonl -valid valid.jsonl -o events.chispa
kling ai chispa predict -model events.chispa -text "panic in the parser" -fields '{"service":"api"}'

Output (one JSON line)

{"label":"fix","index":5,"prob":0.393,"threshold":0.517,"confident":false,"decision":"escalate", …}

Every answer says confident or escalate. It answers in 1.5–6 µs, with no memory allocations, bit for bit the same on amd64 and arm64.

SmolLM2-360M, Qwen2.5 0.5B and 1.5B, served with llama-server from a golden with the model already loaded.

Terminal window
kling ai model add von-smol -model smollm2-360m-instruct
kling run -from von-smol -name smol-1
kling ai model ask smol-1 "What is a microVM?"
461 MiBPSS of 4 SmolLM2 replicas (their RSS adds up to 1715)
~0.8 sfirst token from the golden on a Mac (vz)
~140 tok/sSmolLM2 in a vz guest with 2 vCPUs
Terminal window
kling ai serve
kling ai test commit-type "fix crash when the cache is cold"
kling ai eval commit-type -data test.jsonl -von qwen

An OpenAI-compatible API. The cascade (VON answers what Chispa is unsure about) only turns on if kling ai eval proves that it wins.