From the AltSlate Blog
Notes from the workshop — how we think about agents, cloud engineering, and turning ambitious ideas into production reality.
Half the Memory, None of the Answers Lost
Small AI models get expensive the moment you ask them about something long, because they have to remember all of it. We tried two ways to cut that memory bill. One flopped. The other halved it for free — and shrank it to a fifth with a little tuning, without touching a single answer.
The Verifier's Blind Spot
We built a "self-improving" LLM cascade to get cheaper every round. It got quietly worse — and every dashboard said it was fine. What we measured, and why you cannot see it from inside.
Technology Finally Reached the Delivery Layer
Cloud rebuilt the front door of services firms and left the workshop untouched. Agents walk into the workshop — and that changes the economics of professional services.
Why We Build Agent-First, Not Agent-Bolted-On
Bolting an LLM onto an existing app gets you a demo. Designing the agent as the operating core — with humans in command — gets you a product. Here is how we think about the difference.
Serverless That Actually Scales (and Stays Cheap)
Serverless is not automatically cheap or scalable — it is a set of trade-offs. Here are the patterns we lean on to get fan-out compute that bills to zero when idle.
Production, Not Pilots — How We Build
Most AI work stalls at the demo. We design, build, and operate production systems end to end — applying AI where it earns its place and conventional engineering everywhere else.