Can an expensive model actually be cheaper?
We used libfx to give 11 models the same business tasks, check their answers, and compare cost per successful task. Here is what the bills showed.
We test models on synthetic tasks and production workloads, then publish what we measure: what each one costs, how fast it is, and whether the output does the job.
We used libfx to give 11 models the same business tasks, check their answers, and compare cost per successful task. Here is what the bills showed.

Ten inexpensive models, three production workloads, identical inputs. We scored quality, latency and cost, then ranked each model against its peers on the two ratios that decide a deployment.