What Actually Happened
Bloomberg reported that Harvey left OpenAI and Anthropic after gross margins collapsed from around 50 percent to minus 50 percent by June. The cause was volume: token usage jumped roughly twentyfold under usage-based pricing as customers used the product more.
Harvey launched an in-house model post-trained on Moonshot Kimi K3 in August. Margins went positive after migration. Abridge, Decagon and Ramp are reported to be pursuing the same pivot.
Claude Opus 5
Best-in-class reasoning, strongest long-horizon agentic performance, and no infrastructure to run. For most teams that combination is worth the price, and the API bill is a line item rather than a business model problem.
It becomes a problem at one specific point: when your product charges a flat subscription and your customers can drive unbounded token usage. That is the trap Harvey fell into, and it is structural rather than a pricing complaint.
Kimi K3 Post-Trained
Open weights mean you pay for compute rather than per token, and your marginal cost per customer stops rising with usage. That is the entire argument and it is decisive above a certain volume.
What it costs you instead: serving infrastructure, an ML team that can post-train and evaluate, ongoing quality monitoring, and a capability gap on the hardest reasoning tasks that you close with post-training rather than by waiting for the next frontier release.
The Break-Even Question
There is a crossover point and it is specific to your business. Below it, closed frontier models are cheaper all-in once you count engineering salaries. Above it, the per-token bill grows faster than anything else in your cost structure.
Harvey was demonstrably above it. A team serving a few hundred customers on bounded workloads is almost certainly below it.
The honest test: model your API spend at ten times current usage. If that number breaks the business, the migration conversation starts now rather than after margins invert.
Decision Framework
- Flat-rate pricing with unbounded customer usage - this is the Harvey shape. Model it at ten times volume today.
- Usage-based pricing that passes token cost through - stay on frontier models. Your economics already work.
- No ML engineering capacity - stay. Post-training is not a weekend project and a badly tuned open model costs more in support than it saves in tokens.
- Hardest-tier reasoning is the product - Claude Opus 5. The capability gap on top-end tasks is real.
- High volume on a narrow, well-defined task - the best case for post-training an open model, and where the margin swing is largest.
Verdict
Claude Opus 5 for capability, and for any team without ML infrastructure. Kimi K3 post-trained for high-volume, narrow workloads where per-token pricing has become the dominant cost line. Harvey is not evidence that open weights beat frontier models - it is evidence that flat pricing plus unbounded usage plus per-token cost is an unstable combination.