Blog

Notes on human-robot interaction research, product engineering, and the craft of building software.

LLMsModel SelectionEvaluationPrompt CachingProduct Engineering

Four Lessons on Making LLM Calls Fast, Cheap, and Accurate

September 3, 2026 · 5 min read

Production notes from hyzl, our AI voice agent platform: reach for a bigger model at low reasoning, use GPT-5.6 Luna for classification, benchmark against an oracle, and cache your prompt prefix.


© MMXXVI