<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Trove]]></title><description><![CDATA[A weekly report on the market for AI training data — who owns it, who's buying, and what it sells for.]]></description><link>https://www.trovereport.com</link><image><url>https://substackcdn.com/image/fetch/$s_!MGqP!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff44e7cf4-e5ce-4857-b1bf-1aea0a9a318d_512x512.png</url><title>Trove</title><link>https://www.trovereport.com</link></image><generator>Substack</generator><lastBuildDate>Sat, 22 Aug 2026 07:02:19 GMT</lastBuildDate><atom:link href="https://www.trovereport.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Brian D'Erario]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[trovemedia@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[trovemedia@substack.com]]></itunes:email><itunes:name><![CDATA[Brian D'Erario]]></itunes:name></itunes:owner><itunes:author><![CDATA[Brian D'Erario]]></itunes:author><googleplay:owner><![CDATA[trovemedia@substack.com]]></googleplay:owner><googleplay:email><![CDATA[trovemedia@substack.com]]></googleplay:email><googleplay:author><![CDATA[Brian D'Erario]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[Data: The Next Asset Class]]></title><description><![CDATA[Why Google&#8217;s $10 million bid for Spirit Airlines&#8217; data matters to every company]]></description><link>https://www.trovereport.com/p/data-the-next-asset-class</link><guid isPermaLink="false">https://www.trovereport.com/p/data-the-next-asset-class</guid><dc:creator><![CDATA[Brian D'Erario]]></dc:creator><pubDate>Thu, 20 Aug 2026 22:25:02 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/1d5889aa-1916-4d50-ac7e-c11f85d40310_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Google bid $10 million for Spirit Airlines&#8217; internal business data, including roughly 100 million emails, 500 million Teams messages, 30 million lines of code and operational records, according to the <a href="https://news.bloomberglaw.com/new-york-brief/google-aims-to-boost-ai-with-purchase-of-spirit-airlines-data">bankruptcy-court notice</a>.</p><p>That number stopped me. If Spirit&#8217;s leftover emails and Teams messages are worth eight figures, every company sitting on years of its own work history should be paying attention.</p><p>Subscribe to get Trove&#8217;s reporting on the AI data market every Monday.</p><h2>From Books to the Internet</h2><p>In eight years, model training has moved from books, to the public internet, to human feedback and proprietary data. GPT-1 trained largely on books. GPT-2 shifted toward WebText, a collection of webpages discovered through links shared on Reddit. The bet was simple: instead of training AI for a single task like chess or Go, train one model on a broad sample of human language. GPT-3 scaled that approach with more webpages, books and Wikipedia. OpenAI details the WebText construction in its <a href="https://cdn.openai.com/better-language-models/language_models_are_unsupervised_multitask_learners.pdf">GPT-2 report</a>.</p><p>GPT-3 became very good at continuing text, but it struggled to answer questions, follow complex instructions or admit uncertainty. Some of those problems remain. What came next was the human-feedback race. InstructGPT, the predecessor to ChatGPT&#8217;s GPT-3.5, added human demonstrations, output rankings and user prompts through reinforcement learning from human feedback. Later generations of GPT, Claude, Llama and Gemini incorporated first-party feedback and licensed third-party data into their training and post-training. The method is detailed in the <a href="https://arxiv.org/abs/2203.02155">InstructGPT paper</a>.</p><h2>Digital Data as a Commodity</h2><p>Human feedback showed that better-targeted data could improve a model without simply making it larger. At the same time, the public internet was becoming less differentiated. Every major model company could crawl similar websites, books and code repositories.</p><p>Model companies began racing to assemble datasets their competitors did not have. Shutterstock, Reddit and major publishers began licensing their archives for model training. The prices are public enough now to sketch the market. News Corp reportedly sold OpenAI access to its archives for more than $250 million over five years. Reddit&#8217;s licensing deal with Google was reported at around $60 million a year. The AP, Axel Springer, the Financial Times, Time, all of them have cut deals in the tens of millions. Add up the announced contracts and model builders have already spent something north of a billion dollars on content they used to scrape for free.</p><p>But most of these licenses were nonexclusive, which means everyone&#8217;s model got the same books. You can see that sameness in the AI-isms developed by every model (looking at you, em dash). Shared data can&#8217;t be an edge. Which raises the question: what do you buy when everyone has already bought the internet?</p><h2>The Race for Proprietary Data</h2><p>As access to capable base models becomes more common, competition is shifting toward the data used to specialize them.</p><p>This is where Spirit comes in. Google bid $10 million for Spirit&#8217;s Teams messages, code and operational records, beating a $7.5 million backup bid from Mercor. As agents enter specific business functions, data showing how work actually gets done could become more valuable than isolated code, manuals or messages. The figures come from the <a href="https://www.axios.com/2026/08/17/google-spirit-airlines-bankruptcy">auction result</a>.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!5hIY!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65758103-64fe-4b14-ad2c-ff5b96082745_1360x936.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!5hIY!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65758103-64fe-4b14-ad2c-ff5b96082745_1360x936.png 424w, https://substackcdn.com/image/fetch/$s_!5hIY!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65758103-64fe-4b14-ad2c-ff5b96082745_1360x936.png 848w, https://substackcdn.com/image/fetch/$s_!5hIY!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65758103-64fe-4b14-ad2c-ff5b96082745_1360x936.png 1272w, https://substackcdn.com/image/fetch/$s_!5hIY!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65758103-64fe-4b14-ad2c-ff5b96082745_1360x936.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!5hIY!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65758103-64fe-4b14-ad2c-ff5b96082745_1360x936.png" width="1360" height="936" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/65758103-64fe-4b14-ad2c-ff5b96082745_1360x936.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:936,&quot;width&quot;:1360,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:94637,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:&quot;&quot;,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://trovemedia.substack.com/i/211995147?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65758103-64fe-4b14-ad2c-ff5b96082745_1360x936.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!5hIY!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65758103-64fe-4b14-ad2c-ff5b96082745_1360x936.png 424w, https://substackcdn.com/image/fetch/$s_!5hIY!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65758103-64fe-4b14-ad2c-ff5b96082745_1360x936.png 848w, https://substackcdn.com/image/fetch/$s_!5hIY!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65758103-64fe-4b14-ad2c-ff5b96082745_1360x936.png 1272w, https://substackcdn.com/image/fetch/$s_!5hIY!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65758103-64fe-4b14-ad2c-ff5b96082745_1360x936.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The value lies in connecting instructions, decisions and outcomes into a usable training set.</p><p>The race extends beyond the frontier labs. Harvey says it post-trained Tenet from Kimi K3 with Fireworks for long-horizon legal work. The company is using that effort to develop proprietary model intelligence for legal work, as described in its <a href="https://www.harvey.ai/blog/post-training-update-harvey-tenet">Tenet update</a>.</p><p>If Tenet outperforms general models on legal work, Harvey owns an advantage that competitors cannot copy simply by calling the same model API.</p><h2>From Digital Data to Physical Data</h2><p>Digital agents can learn from emails, Slack messages and operational documents. Robots need an even scarcer asset: physical-world data. If we believe in a future with abundant robots, those robots will need enormous amounts of data showing how physical work is performed.</p><p>Atoms, Travis Kalanick&#8217;s new company, is assembling operating businesses around physical automation. Its portfolio includes Pronto in mining and Lab37 in food robotics. Pronto says its software learns a haul route after a human operator drives it once. As Atoms expands into more industries, access to those environments and the data they produce could become one of its biggest advantages, according to <a href="https://pronto.ai/solution/">Pronto&#8217;s route-learning description</a>.</p><p>A supplier market is already forming around that demand. MicroAGI&#8217;s Shift records physical work and turns it into training-ready data. Its New York offer exchanges free home cleaning for recorded work, as reported by <a href="https://arstechnica.com/ai/2026/05/robot-training-startup-will-send-humans-wearing-cameras-to-clean-your-home/">Ars Technica</a>.</p><p>If you run a business, your Slack archive might be your most underpriced asset. Worth thinking about before a bankruptcy auction decides its price for you.</p><h2>Why I started Trove</h2><p>Eight years ago, data was the thing companies kept in backups nobody priced. Now it&#8217;s turning up in bankruptcy auctions with eight-figure bids.</p><p>I started Trove because I kept noticing the gap between what companies think their data is worth and what someone will actually pay for it. Spirit didn&#8217;t know it was sitting on a $10 million asset until the auction notice went up. Most companies still don&#8217;t.</p><p>So that&#8217;s the beat: every Monday, one piece of this market &#8212; a deal broken down, a company profiled, or someone who sells or buys data explaining how it works. If you own data, need data, or just like watching a new asset class get built in public, stick around.</p><p>See you Monday.</p><p>&#8212; Brian</p><p>Subscribe to get Trove&#8217;s reporting on the AI data market every Monday.</p>]]></content:encoded></item></channel></rss>