Google bid $10 million for Spirit Airlines’ internal business data, including roughly 100 million emails, 500 million Teams messages, 30 million lines of code and operational records, according to the bankruptcy-court notice.
That number stopped me. If Spirit’s leftover emails and Teams messages are worth eight figures, every company sitting on years of its own work history should be paying attention.
Subscribe to get Trove’s reporting on the AI data market every Monday.
From Books to the Internet
In eight years, model training has moved from books, to the public internet, to human feedback and proprietary data. GPT-1 trained largely on books. GPT-2 shifted toward WebText, a collection of webpages discovered through links shared on Reddit. The bet was simple: instead of training AI for a single task like chess or Go, train one model on a broad sample of human language. GPT-3 scaled that approach with more webpages, books and Wikipedia. OpenAI details the WebText construction in its GPT-2 report.
GPT-3 became very good at continuing text, but it struggled to answer questions, follow complex instructions or admit uncertainty. Some of those problems remain. What came next was the human-feedback race. InstructGPT, the predecessor to ChatGPT’s GPT-3.5, added human demonstrations, output rankings and user prompts through reinforcement learning from human feedback. Later generations of GPT, Claude, Llama and Gemini incorporated first-party feedback and licensed third-party data into their training and post-training. The method is detailed in the InstructGPT paper.
Digital Data as a Commodity
Human feedback showed that better-targeted data could improve a model without simply making it larger. At the same time, the public internet was becoming less differentiated. Every major model company could crawl similar websites, books and code repositories.
Model companies began racing to assemble datasets their competitors did not have. Shutterstock, Reddit and major publishers began licensing their archives for model training. The prices are public enough now to sketch the market. News Corp reportedly sold OpenAI access to its archives for more than $250 million over five years. Reddit’s licensing deal with Google was reported at around $60 million a year. The AP, Axel Springer, the Financial Times, Time, all of them have cut deals in the tens of millions. Add up the announced contracts and model builders have already spent something north of a billion dollars on content they used to scrape for free.
But most of these licenses were nonexclusive, which means everyone’s model got the same books. You can see that sameness in the AI-isms developed by every model (looking at you, em dash). Shared data can’t be an edge. Which raises the question: what do you buy when everyone has already bought the internet?
The Race for Proprietary Data
As access to capable base models becomes more common, competition is shifting toward the data used to specialize them.
This is where Spirit comes in. Google bid $10 million for Spirit’s Teams messages, code and operational records, beating a $7.5 million backup bid from Mercor. As agents enter specific business functions, data showing how work actually gets done could become more valuable than isolated code, manuals or messages. The figures come from the auction result.
The value lies in connecting instructions, decisions and outcomes into a usable training set.
The race extends beyond the frontier labs. Harvey says it post-trained Tenet from Kimi K3 with Fireworks for long-horizon legal work. The company is using that effort to develop proprietary model intelligence for legal work, as described in its Tenet update.
If Tenet outperforms general models on legal work, Harvey owns an advantage that competitors cannot copy simply by calling the same model API.
From Digital Data to Physical Data
Digital agents can learn from emails, Slack messages and operational documents. Robots need an even scarcer asset: physical-world data. If we believe in a future with abundant robots, those robots will need enormous amounts of data showing how physical work is performed.
Atoms, Travis Kalanick’s new company, is assembling operating businesses around physical automation. Its portfolio includes Pronto in mining and Lab37 in food robotics. Pronto says its software learns a haul route after a human operator drives it once. As Atoms expands into more industries, access to those environments and the data they produce could become one of its biggest advantages, according to Pronto’s route-learning description.
A supplier market is already forming around that demand. MicroAGI’s Shift records physical work and turns it into training-ready data. Its New York offer exchanges free home cleaning for recorded work, as reported by Ars Technica.
If you run a business, your Slack archive might be your most underpriced asset. Worth thinking about before a bankruptcy auction decides its price for you.
Why I started Trove
Eight years ago, data was the thing companies kept in backups nobody priced. Now it’s turning up in bankruptcy auctions with eight-figure bids.
I started Trove because I kept noticing the gap between what companies think their data is worth and what someone will actually pay for it. Spirit didn’t know it was sitting on a $10 million asset until the auction notice went up. Most companies still don’t.
So that’s the beat: every Monday, one piece of this market — a deal broken down, a company profiled, or someone who sells or buys data explaining how it works. If you own data, need data, or just like watching a new asset class get built in public, stick around.
See you Monday.
— Brian
Subscribe to get Trove’s reporting on the AI data market every Monday.



