Over the past week, Harvey and Thomas Reuters have announced they’ve built their own models to be used within their orgs and in Harvey’s case, deploying and training with customers. This strikes me as an inflection point where companies are shifting from aggressive token burning at any cost to utilizing routers and building their own models to control their data and costs in their AI stack.
Going back to April 2026 at Meta, an employee-built internal dashboard called Claudeonomics turned AI use into a competition. It ranked the company’s top 250 token users, handed out titles such as “Token Legend” and “Session Immortal,” and showed that more than 85,000 employees had consumed over 60 trillion tokens in 30 days. Meta also maintained a separate official usage dashboard for software engineers, who were among the company’s heaviest users. The leaderboard became popular enough that Meta took it down after the data leaked.
The thinking behind it was spreading as an effective way to catalyze employees to be using AI in their day-to-day as much as possible. Nvidia CEO Jensen Huang said that if a company pays an engineer $500,000 a year and that person isn’t burning $250,000 in tokens, something is wrong. Meta CTO Andrew Bosworth offered an even cleaner anecdote: his best engineer was spending the equivalent of his salary on AI tokens and, Bosworth claimed, producing five to ten times more work.
For a moment, token consumption became a proxy for ambition. Companies bought ChatGPT Enterprise and Claude Enterprise seats, opened access to coding agents, and encouraged employees to use as much intelligence as the model companies could sell them. The focuse was whether people were using AI and higher usage was assumed to improve the value each employee was delivering.
Then the bills arrived. Long-running agents consumed far more than chatbots. The same task could cost radically different amounts depending on the model and the harness around it. High usage didn’t automatically mean useful work. By the summer of 2026, the tokenmaxxing fad was giving way to budgets and cheaper models.
That correction marks the start of enterprise AI’s mature phase. The goal is no longer usage for its own sake. Companies are now deploying sophisticated tokenomics into their spending models to decide how teams on a quarterly/yearly basis should be using AI. They’re asking which jobs still deserve an expensive frontier model, which requests can be routed to cheaper models, and which repeated workflows contain enough proprietary value to support a model of their own.
From model access to stack ownership
Enterprise AI began with GPT or Claude behind a chat interface or copilot. Agents moved the vendor decision from the model to the harness. Choosing Claude Code, Codex, Cursor, or Copilot means committing workflows, permissions, tools, and evaluations to the environment around the model. That choice can become more consequential than the difference between Claude and GPT on a benchmark.
Companies with enough usage are now pulling that layer inside their own built infrastructure. They own the session, policy, and production history, then route work among models beneath it. Shopify’s Aquifer platform does this across Claude Code, Codex, Copilot, Cursor, and other tools, allowing Shopify to change a model or runtime without giving up its control plane.
That control also exposes where an owned model makes sense. Shopify processes 40 million multimodal inferences a day across its catalogue. When commercial APIs became prohibitively expensive, it fine-tuned smaller open models and deployed them on its own infrastructure. With Sidekick, Shopify turns unsuccessful merchant interactions into training data. It says the resulting GraphQL model surpassed its frontier baseline and reduced estimated annual serving costs from roughly $27 million to about $1 million, a reduction of roughly 96%.
Companies are starting to build models of their own
The companies below have publicly disclosed training, continued training, or material post-training of a generative or domain foundation model.
Shopify. Shopify fine-tunes open multimodal models for catalogue work and smaller language models for Sidekick. It serves them on infrastructure it controls and continually retrains them using failures from production.
Walmart. Wallaby is a family of retail-specific language models trained on decades of Walmart data. Walmart combines those models with outside LLMs in shopping, personalization, and customer-support systems.
eBay. LiLiuM is a family of 1 billion, 7 billion, and 13 billion parameter models developed entirely in-house. eBay trained them for marketplace tasks including product titles, descriptions, attribute extraction, and pricing.
Intuit. GenOS gives Intuit teams access to commercial and open models alongside its own custom-trained models for tax, personal finance, accounting, and marketing work.
Bloomberg. BloombergGPT is a 50 billion parameter model trained on general text and Bloomberg’s financial corpus. The project showed how a company with a valuable information archive could turn that archive into model weights.
Mastercard. Mastercard is training a tabular foundation model on billions of anonymized transactions, with plans to expand into hundreds of billions. It expects the model to support fraud detection, loyalty, personalization, portfolio tools, and other payment workloads.
Thomson Reuters. The company starts with open Qwen models and adds professional content, tools, and feedback from hundreds of subject-matter experts. Its Thomson model will operate inside CoCounsel alongside models from outside providers.
Harvey. Harvey began as a defining application of GPT-4, then added models from Anthropic and Google. It has now post-trained an open-weight Kimi model inside long-running legal environments using public law, synthetic examples, expert work, and legal rubrics.
Cursor. Cursor operates a router across outside and internal models, while training its own Composer coding models inside the harness where they will be used. The product can keep difficult tasks on frontier systems and move repeatable coding work to models it controls.
C3 AI. Narwhal is a 27 billion parameter model trained further on C3 AI’s proprietary programming system. The company turns failed developer questions into evaluation rubrics and training examples.
ServiceNow. ServiceNow’s Apriel 13B was continued from an open Mistral model and trained for enterprise workflows, tool use, and Now Assist applications. Its smaller size also makes restricted and private deployments more practical.
Salesforce. xGen-Sales and xLAM are proprietary model families built for CRM work and actions. Salesforce uses synthetic-data pipelines to train models that can select tools and carry out work inside Agentforce.
IBM. Granite is IBM’s open family of enterprise models for language, code, documents, speech, time series, and safety. IBM sells the models through watsonx while allowing customers to run and customize them elsewhere.
Snowflake. Arctic was trained for enterprise work including SQL generation and instruction following. Snowflake released its weights and training information while serving it inside Cortex alongside outside models.
Databricks. DBRX is Databricks’ open mixture-of-experts model. It also serves as proof for the company’s larger pitch that enterprises can use their own data to train and govern custom models through Mosaic AI.
SAP. RPT-1 is a relational foundation model built for tables and connected business data. It performs classification and regression from examples without requiring a separate model to be trained for each prediction task.
Adobe. Firefly is a family of Adobe-trained models for images, video, audio, vectors, and design. Adobe also lets companies customize Firefly models on approved brand assets for production work.
Autodesk. Project Bernini is an experimental model trained on ten million 3D shapes. It generates functional geometry from text, images, sketches, voxels, and point clouds.
Cisco. Foundation-sec-8B starts with Llama and adds continued training on a cybersecurity corpus built inside Cisco. It is designed for security workflows that can’t send sensitive material to a hosted general-purpose API.
Siemens. Its Industrial Foundation Model is being developed for engineering drawings, 3D models, technical specifications, and manufacturing data that general LLMs rarely understand.
Recursion. Recursion trained Phenom-1, a vision foundation model built on billions of cellular images from its proprietary phenomics library. It is extending the same approach into transcriptomics, chemistry, and patient biology.
Insilico Medicine. Nach01 is a foundation model for chemical structures and molecular prediction. Insilico sells access to the model while using related systems throughout its own drug-discovery platform.
Apple. Apple trains its own on-device and server foundation models for Apple Intelligence. Small models run locally, while larger workloads can move to Apple’s Private Cloud Compute infrastructure.
Samsung. Gauss2 is Samsung’s multimodal language, code, and image model. It already powers internal coding and productivity tools and is being adapted for Samsung devices.
LG. LG runs a smaller model derived from EXAONE locally on its laptops, giving users private search, summarization, and device assistance without requiring a cloud connection.
Naver. HyperCLOVA X is a from-scratch model family built around Korean language and business needs. Naver uses it in its own consumer services and sells access and dedicated infrastructure to other enterprises.
Rakuten. Rakuten AI 7B was continually trained from Mistral on Japanese and English data using Rakuten’s own GPU cluster. Rakuten uses its model family across its businesses and has created versions that can run locally on PCs.
Sovereignty without isolation
None of this means every enterprise needs its own LLM. The economics only work when a task happens often, the company has data or expertise that improves it, and success can be measured reliably. Plenty of work will remain on frontier APIs because the best general model is difficult to reproduce and cheaper to rent than build. Training a model requires a team and dedicated resources to put together.
Sovereignty also doesn’t require a company to cut itself off from OpenAI or Anthropic. Shopify, Thomson Reuters, Harvey, Cursor, Walmart, and Intuit all use outside models while building models of their own. The frontier model can handle unfamiliar work, teach a smaller model, or judge its output. A specialized model can absorb high-volume work where cost and company context matter more.
Tokenmaxxing was about proving that a company could consume AI. The next phase is deciding what methods yield the highest ROI per token burned.


