We cut the inference cost by 16x
A frontier model was writing a product’s core output at 11 cents a pop. A small open model, trained on a rented GPU for $8, does the same job for less than a cent.
We audited every call the product made to a model, found one repetitive task eating the whole bill, and distilled it into a focused open model.




