News Weekly
LV 10 XP
0% read
Your progress · 0/5 chapters
About 6 min total
MoneyISSUE #2 · STORY 8 OF 20Sep 21, 2026CONFIRMED

Harvey's margins crashed until it switched to an open model

Bloomberg reported that legal AI firm Harvey's gross margins swung from about plus 50 percent to minus 50 percent after an agent update spiked usage roughly 20 times. The fix: building its own model on Moonshot's open Kimi K3 weights.

Illustration: a dramatic abstract pendulum swinging between a tall glossy peak and a deep ravine, tracing a wide arc — an artistic impression of reported margins swinging from roughly plus-fifty to minus-fifty p…

Read it your way

CHAPTER 1 · THE 60-SECOND VERSIONPicked for Explorers

Profit, a cliff, then a fix

According to Bloomberg, Harvey ran about 50 percent gross margins at the start of 2026. By June, the report says margins hit minus 50 percent after a March agent update grew token usage roughly 20 times. The response was Harvey Tenet, the company's own model built on Moonshot's open-weight Kimi K3, and margins turned positive again after that, per people familiar cited by Bloomberg.

Why usage explodedA March 2026 update to Harvey's AI agents spiked customer usage roughly 20 times, per Bloomberg.
The per-token trapBecause frontier labs charge per token, every extra use inflated Harvey's costs faster than its revenue.
The Tenet pivotHarvey post-trained Kimi K3 with Fireworks and released Tenet as a research preview on Aug 20, 2026.
Margins came backPeople familiar say margins turned positive again, per Bloomberg, though Harvey declined to share the number.
Finish this chapter for +15 XP
Flip the switch

Inside the margin flip

YOU GETAn in-house modelHarvey Tenet, post-trained on Kimi K3, shipped as a research preview on Aug 20, 2026.
YOU GETMargins positive againPeople familiar say margins recovered, though the exact figure remains undisclosed.
YOU GETA routing playbookVolume goes to Tenet while frontier models keep the hardest tasks, a hybrid design.
Your next move · as a Explorer

Spot the margin trap

1Ask who owns the model layer behind any AI app you use.
2Notice when AI features bill per token, because usage becomes cost.
3Follow Abridge, Decagon and Ramp for copycat pivots.

Switch your reading mode at the top to see a different next move.

Tap to open

Things to keep an eye on

Pop quiz · unlock the Margin Mentor badge

Did it stick?

0/3
Why did Harvey's margins collapse, per Bloomberg?+20 XP
What is Harvey Tenet built on?+20 XP
What stays true after the pivot?+20 XP
Your call · +5 XP

Can app-layer startups survive long-term by renting frontier APIs, or must high-usage ones own their models?

Deep dive

The full research, labeled and sourced

CONFIRMED18 sources · 51 min
Story identity
FieldValue
Story IDS08
TitleBloomberg: Harvey's margins swung from ~50% to −50% until open-model pivot to Kimi K3
OrganizationsHarvey (legal AI); Moonshot AI (base-model provider); Fireworks (post-training partner); Abridge, Decagon, Ramp, Rogo (peers)
CategoryBusiness / Startup
Event date2026-09-21 (Bloomberg article publication, 4:05 PM UTC; TNW follow-on same day 5:30 PM UTC)
Window check2026-09-21 ∈ [2026-09-18, 2026-09-22] — IN WINDOW
Article dates2026-09-21 (Bloomberg, TNW); preceding evidence dated 2026-06 → 2026-08 (margin swing, Tenet launch)
Evidence statusCONFIRMED as a reporting event (the Sep 21 Bloomberg report itself); the underlying margin figures are report-level (EARLY RESEARCH) — sourced to a person familiar with the matter plus company statements, corroborated by The Next Web and byoBot but not audited or confirmed by Harvey (Harvey declined to comment on specific financials)

Evidence-status labels used below: CONFIRMED (verifiable event/fact), COMPANY CLAIM (stated by Harvey/Moonshot), INDEPENDENTLY VERIFIED (corroborated by two or more independent parties), EARLY RESEARCH (freshly reported, not yet fully corroborated), RUMOR (not used — no rumor-grade material in this story).


✓

What happened?

🎓 For Explorer

CONFIRMED (reporting event): On Monday September 21, 2026, Bloomberg (Rebecca Torrence and Natasha Mascarenhas) published "Startups Like Harvey Embrace Open Models to Cut Reliance on Anthropic, OpenAI." The report describes AI-native startups migrating from frontier-lab APIs to open-weight models to survive the economics:

  • COMPANY CLAIM / EARLY RESEARCH (per Bloomberg, citing a person familiar with the matter): Harvey's gross margins collapsed from roughly +50% at the start of 2026 to −50% by June 2026, after a March 2026 update to its AI agents spiked customer usage. Bloomberg reports token usage grew roughly 20x this year (attributed to the company, per TNW).
  • COMPANY CLAIM / EARLY RESEARCH: Harvey post-trained its own model — Harvey Tenet, built on Moonshot AI's open-weight Kimi K3 base (announced as a research preview on August 20, 2026) — and gross margins turned positive again thereafter, according to people familiar with the effort.
  • CONFIRMED (corroborated by multiple outlets + company materials): Harvey is valued around $15.5–15.6B (company-announced $15.5B on its $550M raise; Bloomberg cites $15.6B); Kimi K3 is a 2.8T-parameter open-weight MoE model released by Moonshot in mid-July 2026 with weights published July 27, 2026.
  • CONFIRMED (Bloomberg + TNW): Abridge (clinical AI; building on Nvidia's Nemotron open weights), Decagon (customer-support AI; ~80% of queries already routed through its own models per TNW/Digital Today), and Ramp (fintech; raised $750M in June, weighing in-house training for the first time) are pursuing the same path; Sequoia Capital and General Catalyst are funding the shift; Rogo and Canva are also moving toward open/smaller models (per 36Kr English / Digital Today).

Bloomberg.com itself blocks direct fetching (403), so the article text was verified via search-engine excerpts, Google News RSS syndication of the same content, and an independent TNW article that cites the same figures with the same source attribution.

Δ

What changed?

The premise that "renting frontier APIs is the standard cost of doing business for an AI-native startup" changed. Previously, being an app-layer company on top of OpenAI/Anthropic APIs was a badge; the Harvey numbers make it a gross-margin liability. Key change:

  • Before: AI-native startups resold frontier models (Harvey originally built on OpenAI GPT-4-class models) under seat/subscription pricing, with model providers charging usage fees on top of base fees; gross margin ~+50% (Harvey, start of 2026) — standard SaaS-like economics.
  • Change (March→June 2026): An agent update multiplied usage ~20x; because the underlying cost is variable per-token usage pricing from OpenAI/Anthropic, COGS swamped revenue — margin swung to −50%.
  • Response (June→August 2026): Harvey built its own post-trained model (Tenet) on Moonshot's open-weight Kimi K3, with Fireworks research; launched research preview August 20; margins positive again (people familiar, per Bloomberg/TNW).
  • After (Sep 2026): Open-weight post-training is repositioned from a cost experiment to the survival play for agentic startups; Abridge (Nemotron), Decagon (in-house), Ramp/Rogo (exploring) follow; investors (Sequoia, General Catalyst) explicitly funding the shift; Anthropic is publicly tracking the trend (investor-forum slide noting Harvey still needs Opus for hardest tasks, per TNW).
↔

Before → Change → After

🎓 For Explorer
  • BEFORE (start of 2026): Harvey, valued ~$15.5B, $350M+ annualized revenue run rate (per TNW Aug 2026), routes legal work through models from OpenAI, Anthropic (and Google). Gross margin ≈ +50%. Model calls on a customer's behalf = invoices from rival labs; OpenAI is an investor in Harvey (alongside Sequoia and a16z).
  • CHANGE (March 2026): Agent update → usage spike → token usage grows ~20x on usage-based pricing. By June, gross margin ≈ −50% — every dollar of legal-AI revenue now loses money at the model layer.
  • CHANGE (June–August 2026): Harvey pivots: research agenda ("building frontier legal intelligence using open-weight models"; "systems to allow law firms to build their own specialized models and own their intelligence"). Harvey Tenet = Kimi K3 base post-trained with Fireworks (GSPO RL, ~150 B300 GPUs, ~2 months); released as research preview Aug 20, 2026.
  • AFTER (Sep 2026): Margins positive again (per people familiar — exact figure undisclosed). Harvey still uses frontier models for the hardest tasks (hybrid routing), but the flagship is now the in-house model. Peers copy the playbook; Bloomberg documents the wave.
⚙

How it works

The mechanics, per Harvey's own technical blog (Aug 20, 2026) and Moonshot's materials:

  • Base: Kimi K3 — Moonshot's 2.8T-parameter open-weight Mixture-of-Experts model (16/896 experts active), Kimi Delta Attention + Attention Residuals, ~2.5x scaling efficiency vs Kimi K2, 1M-token context, native vision. Weights published July 27, 2026 under the "Kimi K3 License" (derivative models permitted; separate Moonshot agreement required for model-as-a-service operators above $20M revenue in any 12 months; prominent "Kimi K3" attribution required above 100M MAU).
  • Post-training: Harvey + Fireworks ran asynchronous reinforcement learning (GSPO — group sequence policy optimization) with rank-64 LoRA over the full 2.8T network (~500k expert tensors). ~1,750 agentic legal task environments (closed-universe client matters with expert rubrics, ~50 criteria/task), >10,000 rollouts, ~150 NVIDIA B300 GPUs over ~2 months. LLM-as-a-judge grading (they converged on Kimi 2.6 as judge); no customer data used in training.
  • Cost engineering: Harvey reports reward shaping that prefers token-efficient trajectories — because cost = price-per-token × tokens-used, they co-optimize both. Post-trained model roughly doubles held-out LAB task completions over base K3 (+9pp all-pass on LAB; +20% on LAB Contracts) and, on Review Table work, improves answer quality +3.6 and citation quality +12.1 at "roughly one-tenth the cost per cell." Their "intelligence-per-token" metric: 190.8 vs 129.3 for best frontier config (FACT per Harvey's blog — company-reported).
  • INTERPRETATION — private-company parameters are unobservable, but the mechanism is not: the exact price/revenue mix that produced Harvey's −50% is undisclosed, yet any realistic combination (see the sensitivity analysis in labs/S08.md) reproduces the inversion above ~3x token growth; the numbers Bloomberg reported are plausible interior points of the model space, which is why the story reads as a mechanism lesson rather than an anomaly.
  • Private-company economics: Harvey declined to comment on specific financials (Bloomberg/TNW); margin figures derive from a person familiar with the matter + company statements. Moonshot's own model card even lists a "Harvey Lab-AA 94.6" criterion pass rate for Kimi K3 — an independent marker of Harvey's benchmark appearing in the base model's eval table.
!

Why it matters

🎓 For Explorer

This is the first on-the-record account of the margin inversion hitting the AI-native layer:

  1. It converts open-weight post-training — previously a cost experiment — into the standard survival play for agentic startups with high token consumption.
  2. It reframes frontier-API dependence as a strategic liability, not a badge of quality, for high-volume workloads.
  3. It names the mechanism: OpenAI and Anthropic charge usage on top of subscriptions; 20x usage growth turns a 50% gross margin into −50% when the COGS is a per-token metered bill.
  4. It has a geopolitical flavor: the fix for a US startup was Chinese open weights (Moonshot) underneath privileged legal client work — immediately raising privilege, national-security, AI Act and licensing questions (TNW).
  5. It arrives in a pricing-war window (Grok 4.7 at $2/$6 per M tokens the same day; Xiaomi's MiMo-V2.6 open weights 2 days earlier at MIT; OpenAI cutting GPT-6 Sol/Luna prices ~50% the next day) — the front-end evidence that frontier-lab pricing power is being contested from below.
✦

What became possible?

🎓 For Explorer
  • Post-training as a moat: A well-capitalized vertical AI startup can now take a 2.8T open-weight frontier-class base, add domain RL, and get legal-specific quality that beats the base while cutting cost per task by an order of magnitude (company-reported figures).
  • "Owning your intelligence": Harvey's public positioning is that law firms (and other professional-services firms) can eventually own specialized models themselves — turning Harvey from a model intermediary into an intelligence-infrastructure provider.
  • Inference-economics arbitrage: Hosted open-weight pricing (Kimi K3 at ~$2.85–3.00 / $14.25–15.00 per M tokens via DeepInfra/Moonshot) plus self-host options (MXFP4 weights ≈1.4TB for the full 2.8T model) gives startups a third option between "rent closed frontier" and "do nothing."
  • Investor category formation: Sequoia and General Catalyst funding the open-weight pivot signals a new diligence criterion for AI-native startups: model-layer dependency and margin exposure.
◎

Implications

Technical

  • Domain RL post-training is now production-proven at the 2.8T scale for a vertical: GSPO + LoRA over a huge MoE, environment-graded RL (rubric-based LLM judging), no customer data — reproducible-ish recipe others will copy.
  • Cost = price × tokens, not price alone: Harvey's reward shaping for token efficiency (and its "intelligence-per-token" framing) is a design principle that will propagate to every agentic workload.
  • Open-weight frontier parity is here: K3 (AA 46-class on several agentic suites, DeepSWE 67.5, Terminal-Bench 2.1 88.3 per Moonshot's card) makes base-model choice a licensing/hosting question as much as a quality question.
  • License complexity is now a technical constraint: The Kimi K3 License's $20M-MaaS-revenue clause and 100M-MAU attribution clause create real compliance surface for US startups — a new class of "model-supply-chain due diligence."
  • Hybrid routing remains: even after Tenet, Harvey "still needs Opus for its hardest tasks" (Anthropic's investor-forum framing, per TNW) — multi-model routing (per-task, per-model) is the emerging architecture, not full replacement.

Developer

  • Re-evaluate every high-volume workload against hosted-open-weight and self-hosted options; token cost curves now dominate gross-margin math.
  • Adopt multi-model routing defaults (frontier for hardest tasks, post-trained open-weight for volume) — the Harvey pattern.
  • Audit open-weight licenses before building SaaS: the Kimi K3 License (MaaS >$20M/12mo requires a Moonshot agreement; 100M MAU attribution) is the template to read carefully; MIT (MiMo-V2.6, GLM-5.2) remains the permissive benchmark.
  • Learn RL post-training tooling (GSPO, LoRA-on-MoE, harness-level training like Fireworks/Baseten) — this is becoming a differentiator skill; Harvey's own hiring (Research Engineer, Post-Training) signals the demand.
  • Track provider leverage events: OpenAI ending the Cursor contract post-SpaceX-acquisition (TNW, Sep 21) shows supplier-side risk is real and asymmetric.

Enterprise

  • Legal AI procurement resets: law-firm buyers now have two questions — output quality AND whose model layer is underneath (provenance, hosting jurisdiction, privilege exposure). Harvey's "own your intelligence" pitch targets exactly this.
  • Model-supply-chain due diligence becomes enterprise process: base-model origin (Chinese weights under privileged work), license terms, data flows, and the EU AI Act documentation chain (modified general-purpose models; post-training generally below the "significant change" ~1/3-compute mark, per TNW's analysis).
  • Negotiating leverage vs frontier labs increases: enterprises with volume can now credibly threaten open-weight fallback; simultaneously, they must price in the risk that the frontier lab is also their competitor's supplier (Anthropic slide; Cursor precedent).
  • Vertical AI economics become investable/auditable: gross-margin trajectories of AI-native vendors (health, legal, support, finance) will be scrutinized the way SaaS magic-numbers are — margin inversion stories will scare diligence.

Strategic

  • Frontier labs face a new competitive flank: their own app-layer "tenants" are building exits, converting the biggest customers from API revenue into fixed-compute deployments. This undercuts the $100B run-rate narrative (Anthropic, S07 same week) at the margin.
  • Open-weight supply now has a demand-side proof point: Nvidia/Microsoft's open-weight coalition letter (July 2026, post-Kimi) gains an economic exhibit: Harvey's margin recovery.
  • China's open-weight strategy is working commercially in the US: Moonshot's K3 under Harvey's flagship is a quiet win for the "open-model diplomacy" play (with attendant political backlash risk).
  • Venture strategy shifts: Sequoia/General Catalyst funding post-training infrastructure and in-house-model programs changes the default financing ask for AI-native startups ("what's your model-layer plan?").
  • The pricing war rotates: with open weights at 46-index quality (Xiaomi) and frontier pricing collapsing (GPT-6 Sol/Luna half-price, Grok 4.7 low-cost), the only defensible position left for closed labs is the very-hardest reasoning/agentic tier — the exact tier Anthropic argues Harvey still needs.
⚠

Risks & limitations

Risks
  • Chinese base weights beneath privileged legal client data: privilege, data-residency, and national-security exposure (TNW flagged it as "who your counterparty is when the work is covered by privilege"). Some markets/procurement rules may prohibit or constrain it.
  • License non-compliance: if Harvey operates Tenet as MaaS and its revenue exceeds $20M/12mo (it runs at $350M+ annualized, per TNW), the Kimi K3 License requires a separate agreement with Moonshot — a real, unresolved compliance surface.
  • Concentration risk relocated: dependence shifts from OpenAI/Anthropic to Moonshot — a China-based single supplier; export-control or US-China policy changes could cut the supply of updates/support.
  • Quality cliff on hardest tasks: Harvey still needs Opus-class models for the hardest legal work; if frontier labs restrict/rescope API access (Cursor precedent), hybrid strategies get squeezed.
  • Report-level figures could be wrong: the −50% and "positive again" numbers come from one person familiar + company statements; Harvey declined to comment; no audited financials. Publicized soft numbers can mislead competitors and customers alike.
Limitations
  • No primary company financial disclosure: margin figures are Bloomberg-reported (person familiar with the matter; company statements), corroborated by TNW and byoBot, but not confirmed by Harvey (declined to comment) and not audited. Treated as report-level (EARLY RESEARCH), not CONFIRMED fact.
  • Definition ambiguity: "gross margin" basis (which costs included — inference only? hosting? amortized GPU?) and the exact post-Tenet margin are undisclosed; "positive" is unquantified.
  • Bloomberg URL is bot-blocked (403); article content verified via search excerpts, Google News syndication of the identical piece, and TNW's independent cite of the same figures.
  • Harvey Tenet performance claims are company-reported; some third-party benchmark runs exist (Artificial Analysis, Vals, Mercor APEX/Redline Bench entries), but the headline "10x cost reduction" and intelligence-per-token figures are Harvey's own.
  • Dates: the K3 release date is cited as July 16–17, 2026 (announcement) with weights July 27; slight date variance across sources (Bloomberg says July 17; Moonshot site lists 2026-07-16) — immaterial to this story.
  • Valuation: company-announced $15.5B (Aug 2026 raise); Bloomberg cites $15.6B — treated as approximate (~$15.5–15.6B).
?

Open questions

  1. What exactly changed in the March 2026 agent update that produced the usage spike?
  2. What is Harvey's actual gross margin now — and does Tenet serve the majority of tokens, with frontier reserved for a minority of tasks?
  3. Does Harvey's product constitute "model as a service" under the Kimi K3 License, and has it entered a separate agreement with Moonshot (required above $20M/12mo)?
  4. Will frontier labs respond with pricing/contract changes for high-volume app-layer customers, or with cut-offs (Cursor precedent)?
  5. How do Abridge/Decagon/Ramp quantitative stories land — and does the trend reach mid-market AI-native startups without $15B war chests?
  6. Will regulators (EU AI Act; US policy) create guidance on Chinese open weights under professional-services privilege?
  7. Does Harvey publish its promised Tenet technical report, and can the RL recipe be replicated at smaller scale?
↗

What happens next?

🎓 For Explorer
  • Short term (Q4 2026): more startups disclose margin-driven pivots (Ramp, Rogo, Canva per 36Kr); Harvey scales Tenet to production, may publish its technical report; frontier labs respond with price/contract moves; watch for the Moonshot–Harvey license agreement surfacing.
  • Medium term (2027): post-training becomes a standard competency in vertical AI; enterprises adopt model-supply-chain audits; the EU AI Act "significant modification" guidance gets tested on post-trained open weights; expect at least one regulator/flagging on Chinese open weights under privileged data.
  • Prediction (medium confidence): within two quarters, at least one more major vertical AI company will announce an open-weight flagship, and at least one frontier lab will introduce a "developer/app-layer" commercial program to slow the migration — mirroring Cursor-style lock-ins turned on their head.
  • Prediction (medium confidence): the "own your model layer" pitch becomes the standard enterprise-AI procurement demand in legal and health by mid-2027.
★

Editorial takeaway

🎓 For Explorer

The quiet assumption that AI-native application companies cannot exit their frontier-lab supply chains died on September 21, 2026. Harvey's swing from +50% to −50% gross margin on a 20x usage curve — and back to positive on a post-trained open-weight model — is the first fully documented instance of the metered-token business model inverting on its own tenants. The strategic lesson is blunt: when your COGS is a competitor's per-token invoice, your volume is their profit and your margin. Open-weight parity (Kimi K3, then MiMo-V2.6 days earlier) turned "we rent GPT-4" from a feature into a liability overnight, and the companies that own their model layer now hold both the cost curve and the geopolitical questions that come with it.

Illustration: frame: a pendulum swings in a wide arc between a tall polished peak and a deep ravine — an artistic impression of reported margins collapsing from roughly plus-fifty to minus-fifty percent.
⌘

Lab: COMPARE

≡

Research sources

Primary Sources (5)
Primary
Limitation note — no primary company source for the margin figures - No URL available (Bloomberg reporting is based on company statements and a person familiar with the matter; Harvey declined to comment on specific financials — recorded per schema source gate)documents that the ~50% → −50% → positive margin trajectory is report-level, not company-disclosed financial data. ---Date: 2026-09-21
URL unavailable
Primary
Moonshot AI — Kimi K3 License (GitHub)license terms — derivative models permitted; a separate agreement with Moonshot is required if the licensee or affiliates operate a Model-as-a-Service business with aggregate revenue above $20M in any consecutive 12 months; "Kimi K3" must be prominently displayed above 100M MAU or $20M monthly revenue; internal use exempt. Underpins the compliance-risk analysis (Harvey runs at $350M+ annualized per TNW). — primary / official license text (CONFIRMED)Date: 2026-07-27
Visit source ↗
Primary
Moonshot AI — Kimi K3 model card (Hugging Face)full Kimi K3 weights released under the Kimi K3 License; architecture (KDA + Attention Residuals, Stable LatentMoE, 16/896 experts active, 2.5x scaling efficiency vs K2); 1M context; official benchmark table — including a "Harvey Lab-AA 94.6" criterion pass rate row for Kimi K3 (evidence that Harvey's own legal-agent benchmark appears in the base model's eval card). — primary / official release artifact (CONFIRMED for contents of the card)Date: 2026-07-27 (weights release)
Visit source ↗
Primary
Moonshot AI — Kimi K3 official model page (moonshot.ai)Kimi K3 is Moonshot's 2.8T-parameter natively multimodal open model, 1M-token context, "world's first open 3T-class model"; positioned for long-horizon coding, knowledge work, deep reasoning. — primary / official product documentation (CONFIRMED for model identity and specs as stated by Moonshot)Date: 2026-07-16 (Kimi K3 listed as 2026-07-16 release)
Visit source ↗
Primary
Harvey — "Update on Harvey's Post-Training Effort" (Harvey Tenet research preview)Tenet is a Kimi K3 base post-trained with Fireworks research; GSPO RL + rank-64 LoRA over the full 2.8T network (~500k expert tensors), ~1,750 agentic legal task environments, >10,000 rollouts, ~150 NVIDIA B300 GPUs over ~2 months; no customer data used; cost engineering via reward shaping (cost = price × tokens); "roughly one-tenth the cost per cell" on Review Table tasks vs baselines; intelligence-per-token 190.8 vs 129.3 for best frontier config; partner stack (Fireworks, Engram, Baseten, Applied Compute, NVIDIA, Mercor, Snorkel AI); company banner confirms $550M raise at $15.5B valuation. — primary / official company documentation (COMPANY CLAIM for performance and cost figures; CONFIRMED for technical methodology facts as described by the company)Date: 2026-08-20
Visit source ↗
Independent Sources (4)
Independent
Artificial Lawyer — "Harvey Tenet, Nashville, Legal Innovators +"corroborates, from legal-tech trade media, that Harvey Tenet is Harvey's first post-trained open-weight model built on a Kimi K3 base (post-trained with partners). — independent reporting (secondary corroboration of the Tenet fact set) ---Date: 2026-08-21
Visit source ↗
Independent
The Next Web — "Harvey's first in-house model for legal work is here" (Tenet on Kimi K3)Tenet is post-trained on Moonshot's Kimi K3 with Fireworks AI; OpenAI is an investor in Harvey (with Sequoia and a16z); "every model call is an invoice from a rival, owning the engine turns a variable cost into a fixed one"; cofounder Gabe Pereyra on routing; training data manufactured with hired lawyers via Mercor and Snorkel; Kimi K3 license MaaS threshold ($20M/12mo) vs Harvey's $350M+ annualized run rate; EU AI Act "significant modification" analysis for post-trained models; valuation context (~$15.5bn). — independent reporting (supports the license-compliance and strategic-risk analysis; background for Section 5/12/13)Date: 2026-08-23 (9:25 AM UTC)
Visit source ↗
Independent
The Next Web — "AI model costs are pushing startups towards cheaper open weights"independent corroboration of every Bloomberg figure (50% → −50% by June; 20x token usage per the company; margins positive after Kimi K3-based model; $15.6B valuation; Harvey declined to comment); Abridge building clinical model on Nvidia open weights; Decagon routing 80% of queries through own models; Ramp raised $750M in June and is weighing first in-house training ("It made absolutely no sense a year ago... It's starting to make a lot more sense now." — co-CEO Karim Atiyeh); Sequoia and General Catalyst funding the shift; Anthropic investor-forum slide noting Harvey still needs Opus for its hardest tasks; OpenAI ending the Cursor contract after the SpaceX acquisition; Mistral CEO Arthur Mensch's July warnings; Legora ($100M revenue in 18 months); Menlo Ventures' Matt Kraning skepticism quote. — independent reporting (INDEPENDENT EVIDENCE — corroborates Bloomberg without duplicating its paywall)Date: 2026-09-21 (5:30 PM UTC)
Visit source ↗
Independent
Bloomberg — "Startups Like Harvey Embrace Open Models to Cut Reliance on Anthropic, OpenAI"the reporting event itself — Harvey's gross margin dropped from ~50% (start of year) to −50% by June after the March agent update spiked usage; ~20x token growth; valuation $15.6B; post-Tenet (August, on Kimi K3) margins turned positive per people familiar; Abridge, Decagon, Ramp pursuing the same open-weight path; open-weight offerings cheaper than US proprietary tech and let firms build custom models on their own data. NOTE: bloomberg.com returns 403 to automated fetches; content verified via search-engine excerpts, Google News syndication (item 10), and The Next Web's independent article (item 7) citing identical figures. — independent reporting; the definitive source of the margin figures (report-level / EARLY RESEARCH for the numbers themselves)Date: 2026-09-21 (4:05 PM UTC), by Rebecca Torrence and Natasha Mascarenhas
Visit source ↗
Secondary Sources (9)
Secondary
DeepInfra — moonshotai/Kimi-K3 hosted model page and API docshosted availability and pricing of the open-weight Kimi K3 ($2.85 / $14.25 per 1M tokens, cached $0.285), confirming the "cheaper than frontier API" economics the story relies on; model card excerpt includes the Harvey Lab-AA benchmark row for K3 (94.6). — secondary/infrastructure evidence (hosted open-weight pricing reality) ---Date: accessed 2026-09-22 (model page live)
Visit source ↗
Secondary
WSJ — "Nvidia Is Developing an AI Healthcare Model With Startup Abridge"additional confirmation of the Nvidia–Abridge open-weights healthcare-model effort (model used exclusively within Abridge's platform). — secondary/background independent reporting (Abridge open-weight path)Date: 2026-06-11
Visit source ↗
Secondary
Fierce Healthcare — "Nvidia teams up with Abridge to build AI healthcare model"predates and supports the Bloomberg claim that Abridge is building its clinical model on NVIDIA Nemotron open weights (trained on Blackwell with de-identified data; NVentures is an Abridge investor). — secondary/background independent reporting (Abridge open-weight path)Date: 2026-06-13
Visit source ↗
Secondary
Digital Today (Korea, English) — "Startups race to build in-house AI models as fintech joins in"additional corroboration of the peer moves (Abridge on open Nvidia model; Decagon 80% in-house; Ramp considering own model since its $750M June raise; Rogo; Cursor; Cognition; Canva). — secondary international corroborationDate: 2026-09-21
Visit source ↗
Secondary
36Kr English — "OpenAI & Anthropic Face Revenue Threats: Cost Pressures Push More AI Startups to Adopt Open-Source Models"non-US corroboration — Harvey margin story; Abridge building clinical foundation model on NVIDIA open-source model; Decagon 80% of queries via self-owned model; Ramp and Rogo exploring first-time in-house training; Cursor (SpaceX-acquired) and Cognition ($48B) among earliest app companies to release customized models; Canva switching to smaller/open models for cost. — secondary international corroborationDate: 2026-09-21/22
Visit source ↗
Secondary
byoBot AI Daily Newsstand — September 22, 2026 digest"Harvey's gross margins went from roughly 50% to negative 50%... token usage jumped 20x under usage-based pricing, and margins recovered only after the legal AI company post-trained its own model on Moonshot's Kimi K3" (citing the Bloomberg URL); its FAQ also asserts "Harvey is better evidence than any leaderboard" for open-weight replaceability. NOTE: AI-written content, unedited; used for corroboration only. — secondary digest (corroboration; treated with caution as AI-generated)Date: 2026-09-22
Visit source ↗
Secondary
AI Weekly — "Harvey Moves Flagship Off Frontier Labs to Moonshot's Kimi K3"same figures plus the editorial framing that "every enterprise AI startup priced against frontier-lab tokens is now staring at Harvey's minus-50% number." — secondary summary (corroboration only)Date: 2026-09-21
Visit source ↗
Secondary
AI Weekly — "Harvey, Abridge, Ramp Ditch Frontier-Lab APIs for Open-Weight Models After Margins Collapsed to -50%"industry digest restating Bloomberg's figures (margins ~50% → −50% by June; token usage +20x; margins positive after August Tenet launch on Kimi K3; Abridge/Decagon/Ramp pursuing same pivot with Sequoia and General Catalyst backing; "first serious wave of enterprise AI startups peeling away from closed-model dependence"). — secondary summary (corroboration only)Date: 2026-09-21
Visit source ↗
Secondary
Google News RSS syndication — "OpenAI and Anthropic costs push Harvey toward cheaper open models" (Bloomberg wire content)full-text syndication of the Bloomberg wire piece — margin swing ~50% → −50% by June; 20x usage; $15.6B valuation; Harvey released its own model in August using Kimi K3; "the model can perform close to Anthropic's best offerings at a fraction of the cost, while Harvey's gross margins have since turned positive, according to people familiar with the effort." Served as the accessible full text of the paywalled article. — syndicated secondary copy of the primary independent report (used to verify Bloomberg content that cannot be fetched directly)Date: 2026-09-21 (last verified 2026-09-21)
Visit source ↗