<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Trove]]></title><description><![CDATA[Trove covers how the data behind AI is sourced, owned, licensed, priced, and turned into model advantage.]]></description><link>https://www.trovereport.com</link><image><url>https://substackcdn.com/image/fetch/$s_!MGqP!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff44e7cf4-e5ce-4857-b1bf-1aea0a9a318d_512x512.png</url><title>Trove</title><link>https://www.trovereport.com</link></image><generator>Substack</generator><lastBuildDate>Tue, 06 Oct 2026 10:11:58 GMT</lastBuildDate><atom:link href="https://www.trovereport.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Brian D'Erario]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[trovemedia@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[trovemedia@substack.com]]></itunes:email><itunes:name><![CDATA[Brian D'Erario]]></itunes:name></itunes:owner><itunes:author><![CDATA[Brian D'Erario]]></itunes:author><googleplay:owner><![CDATA[trovemedia@substack.com]]></googleplay:owner><googleplay:email><![CDATA[trovemedia@substack.com]]></googleplay:email><googleplay:author><![CDATA[Brian D'Erario]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[Gemini Search Data Marketplace]]></title><description><![CDATA[About 100 publishers now get paid when Gemini and AI Overviews use their work. One is making more than $1 million a year. Some have made less than $1,000.]]></description><link>https://www.trovereport.com/p/gemini-search-data-marketplace</link><guid isPermaLink="false">https://www.trovereport.com/p/gemini-search-data-marketplace</guid><dc:creator><![CDATA[Brian D'Erario]]></dc:creator><pubDate>Mon, 05 Oct 2026 14:17:32 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/4562ed28-37a3-4eb7-a661-b1b82d76e2b0_1920x1080.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Google is starting to pay publishers when Gemini uses their work to answer questions. <a href="https://www.theinformation.com/articles/google-paying-100-digital-publishers-ai-overviews">The Information reported</a> last week that one of them is on pace to make more than $1 million a year from it. Smaller sites are making in the thousands, depending on the type of content they&#8217;re writing about.</p><p>What&#8217;s factoring into the payouts isn&#8217;t simply the volume or breathe of content on the site, but the specifics of what&#8217;s being written about. If you want to know what a page of the internet is worth to an AI model in 2026, this is the closest thing we have to a price list.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.trovereport.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Trove! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2>How it works</h2><p>Google calls it the AI contribution pilot, and <a href="https://digiday.com/media/google-rolls-out-pay-value-ai-licensing-program-to-publishers/">Digiday first reported it</a> in mid-September. Publishers who get invited see a new panel in Search Console that tracks how much they've been paid, but not how the number was worked out, and participants say it can change from month to month without explanation. "It's quite a black box," one exec said.</p><p>What Google pays for is narrower than you might expect. A model like Gemini knows a lot from training, but that knowledge goes stale, so for many questions it checks its answer against what it finds in Google Search and revises it. This is called grounding. Not every page it looks at counts. The Information's example: if the model consulted 100 websites for an answer, perhaps only five would meaningfully improve it, and those five are the ones that get paid.</p><p>That's a different kind of data purchase from most of the ones we've covered here. Reddit's reported $60 million a year from Google and News Corp's reported $250 million over five years from OpenAI were flat fees, one price for the whole archive no matter how much any single post mattered. Grounding happens every time someone asks a question, and the pages the model leans on are the ones that earn. It looks like the start of a marketplace for inference data.</p><h2>Why anime pays better than news</h2><p>The $1 million earner joined early, and the money is a meaningful slice of its revenue. A publisher that joined a few months ago has made $50,000 to $60,000, which it said doesn't represent much of its business. For several small and mid-sized sites, payouts came to less than 0.1% of their ad revenue. A site doing $100,000 a year in ads would make less than $100 a year from Google's AI.</p><p>According to The Information, topics that are less widely covered online but draw strong interest, like anime and gaming, appear to earn more. That makes sense once you think about how grounding works. If twenty sites have written up the same iPhone launch, Gemini can lean on any of them, so no single one is worth much. If one fan site is the only place with a detailed answer about a specific anime arc, the model needs that page.</p><p>We keep seeing this across the data market, from <a href="https://www.trovereport.com/p/data-the-next-asset-class">Spirit Airlines' records</a> to <a href="https://www.trovereport.com/p/recursion-turned-an-internal-model">Recursion's lab data</a>. Buyers pay for what they can't get anywhere else, and this pilot applies that logic to individual answers.</p><h2>Who sets the price</h2><p>The other ways AI companies pay for content mostly differ on who decides what it's worth. Flat deals are negotiated up front. Perplexity and ProRata publish a revenue split. TollBit and Cloudflare let publishers set their own rates for bot access. Google's pilot is the only one where the buyer decides what each contribution was worth and doesn't show the math. As one participating executive told The Information, Google appears to be testing a deal in which it "uses publishers' content and decides for itself what that content is worth."</p><h2>Why pay at all?</h2><p>The timing is a little funny. The day after The Information's story, Judge Amit Mehta <a href="https://www.forbes.com/sites/rickellis/2026/10/01/google-wins-dismissal-of-penske-media-chegg-ai-lawsuits/">dismissed the antitrust suits</a> Penske Media and Chegg had brought over AI Overviews. They argued Google broke an unwritten deal: publishers let Google crawl their sites, and Google sends them traffic. "But an expectation is not an agreement," Mehta wrote. For now, Google has a federal judge on record saying it never promised anyone clicks.</p><p>The less generous explanation for the payments is that they're cheap insurance. The European Commission has an open antitrust investigation into how Google uses publisher content for AI, and UK regulators forced Google to offer an AI opt-out in June. A hundred small payments is a useful thing to point to in Brussels and London. David Buttle, founder of the publisher coalition Spur, told Digiday it looks like "a kind of hedge" against a future where Google has to pay for actual usage.</p><p>Publishers don't have much leverage either way. Turning down the payment doesn't take your pages out of AI answers, and opting out of AI features costs whatever traffic they still send. Several larger publishers have declined to join, hoping to pressure Google to pay more, which tells you how the people with the most leverage feel about the current terms.</p><p>My guess is that it's both: a hedge, and a real experiment in learning which content its models depend on. We&#8217;ll watch whether the program grows well past 100 publishers and whether Google ever shows publishers which pages earned and why. Until then, publishers know they're getting paid, just not what for.</p><h2>Worth Reading</h2><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://x.com/mernit/status/2106751265648632240&quot;,&quot;full_text&quot;:&quot;selling data to labs is the modern version of selling to the government\n\nthere are just a handful of buyers. you hustle to get connected to a researcher, get approved as a vendor, and setup a shared slack channel  \n\nyou&#8217;ll get an endless stream of RFPs and be ruthlessly evaluated&#8230;&quot;,&quot;username&quot;:&quot;mernit&quot;,&quot;name&quot;:&quot;Eli Mernit&quot;,&quot;profile_image_url&quot;:&quot;https://pbs.substack.com/profile_images/1462180814993637381/YPonWOXz_normal.jpg&quot;,&quot;date&quot;:&quot;2026-10-04T14:20:17.000Z&quot;,&quot;photos&quot;:[],&quot;quoted_tweet&quot;:{},&quot;reply_count&quot;:14,&quot;retweet_count&quot;:3,&quot;like_count&quot;:355,&quot;impression_count&quot;:40667,&quot;expanded_url&quot;:null,&quot;video_url&quot;:null,&quot;video_preview_media_key&quot;:null,&quot;belowTheFold&quot;:true}" data-component-name="Twitter2ToDOM"></div><p>Eli provides an interesting take here, especially given my media work in Aerospace &amp; Defense before writing Trove. As with any industry, relationships will matter more than performance of a single data set. Given how early we are, there&#8217;s certainly still time fore the winners of the market, the future RTX and Lockheed Martin&#8217;s, to solidify. </p><h2>Partnerships &amp; Deals</h2><h3>Compute, Energy &amp; Infrastructure</h3><ul><li><p><strong>Broadcom</strong> agreed to lend <strong>Anthropic</strong> up to $42 billion in convertible notes to finance its leases of TPU capacity, covering roughly one-third of the $125.2 billion Anthropic has committed under a five-year compute agreement, according to Anthropic's IPO filing.</p></li><li><p><strong>Tencent</strong> reportedly signed a five-year, roughly $7 billion lease with <strong>Oracle</strong> for about 100,000 advanced AI chips in Southeast Asian data centers, with about 30% due upfront, giving Tencent access to hardware it cannot buy in China.</p></li><li><p><strong>Samsung Electronics</strong> and five Samsung affiliates committed $1 billion to <strong>Helix Digital Infrastructure</strong>, the AI-infrastructure company KKR founded in June 2026 to build data centers, power, and fiber.</p></li><li><p><strong>GMI Cloud</strong> raised $668 million, made up of $223 million in Series B equity led by ARCHIV with <strong>NVIDIA</strong> participating and a $445 million credit facility led by CTBC, to add GPU capacity in the U.S., Taiwan, and Asia-Pacific; the company says contracted annual recurring revenue exceeds $600 million.</p></li><li><p><strong>Lambda</strong> closed a $1.008 billion senior secured loan at a 6.78% fixed rate, rated A(low) by Morningstar DBRS and Baa1 by Moody's, to buy GPUs for three deployments backed by two investment-grade customers.</p></li><li><p><strong>Modal Labs</strong> is reportedly closing a $750 million round led by <strong>Accel</strong> at a $15.75 billion valuation, more than triple its mark four months ago, as demand for inference capacity grows.</p></li><li><p><strong>CScale</strong> emerged from stealth with a $145 million Series C co-led by <strong>Atreides Management</strong>, <strong>Valor Equity Partners</strong>, and <strong>Premji Invest</strong>, with <strong>NVIDIA</strong> and <strong>Intel Capital</strong> joining, to build optical interconnect for AI scale-up networks.</p></li></ul><h3>Data, Knowledge &amp; Retrieval</h3><ul><li><p><strong>Supabase</strong> raised $150 million led by <strong>GIC</strong>, with <strong>CapitalG</strong> also investing, and agreed to acquire <strong>Turso</strong>, adding SQLite-based databases built for AI agents; proceeds fund employee liquidity and new agent features.</p></li><li><p><strong>Quartermaster</strong> raised $140 million, made up of a $100 million Series B led by <strong>Insight Partners</strong> and a $40 million debt facility from <strong>Stifel</strong>, to expand its SmartMast sensor network, which streams real-time maritime data from more than 650 vessels to insurers, shippers, and governments.</p></li><li><p><strong>Halluminate</strong> raised a $30 million Series A led by <strong>Oak HC/FT</strong>, bringing total funding to $38.5 million; it builds benchmarks and reinforcement-learning environments for financial work, and says four of the five leading closed-source U.S. labs are customers.</p></li><li><p><strong>OneMedNet</strong> signed a seven-figure, multiyear agreement to supply de-identified imaging and clinical data to an AI precision-medicine company that will build a derivative dataset, with OneMedNet earning sublicensing revenue each time the dataset is licensed to an end user.</p></li></ul><h3>Models, Training &amp; Developer Tools</h3><ul><li><p><strong>SoftBank</strong> completed the third and final $10 billion tranche of its $30 billion follow-on investment in <strong>OpenAI</strong>, lifting its total to $64.6 billion and an ownership stake of about 13%.</p></li><li><p><strong>Modulate</strong> raised a $25 million Series B led by <strong>Future Ventures</strong>, bringing total funding to $60 million, to expand its audio-native models and developer tools for detecting emotion, intent, and synthetic voices.</p></li></ul><h3>Enterprise Deployment &amp; Distribution</h3><ul><li><p><strong>Baseten</strong> joined the <strong>OpenAI</strong> B2B Marketplace as one of the first open-model inference providers, letting enterprise customers count spend on open models in Codex and the Responses API against their existing OpenAI commitments.</p></li><li><p><strong>Reco</strong> raised $55 million led by <strong>AT&amp;T Ventures</strong>, bringing total funding to $140 million, for software that shows enterprises which AI agents run in their systems, what permissions they hold, and what data they reach.</p></li></ul><h3>Industry Applications &amp; Workflows</h3><ul><li><p><strong>Instinct</strong> raised $1 billion in a Series C led by <strong>Sequoia Capital</strong>, <strong>Benchmark</strong>, and <strong>Coatue</strong> at a $10 billion valuation, four times its August mark, for a consumer agent that completes tasks by phone and computer.</p></li><li><p><strong>Salesforce</strong> agreed to acquire <strong>Listen Labs</strong>, an AI customer-research platform with a network of more than 50 million potential participants, in a deal reported at about $2 billion, expected to close in its fiscal fourth quarter of 2027.</p></li><li><p><strong>EliseAI</strong> raised $350 million at a $4 billion valuation, led by <strong>Andreessen Horowitz</strong> and <strong>Bessemer Venture Partners</strong>, to extend its AI for housing and healthcare operations; the company reported $200 million in annual recurring revenue as of June.</p></li><li><p><strong>Armadin</strong>, founded by Kevin Mandia, raised a $255.5 million Series B at a $2.5 billion-plus valuation, co-led by <strong>Andreessen Horowitz</strong> and <strong>Accel</strong>, for always-on agent swarms that test enterprise security by chaining together vulnerabilities.</p></li><li><p><strong>OpenAI</strong> and <strong>Synopsys</strong> formed a multi-year partnership to build GPT-Synopsys, a model that operates chip-design software, with joint go-to-market, revenue sharing, and OpenAI licensing Synopsys EDA tools for development.</p></li></ul><h3>Physical AI &amp; Robotics</h3><ul><li><p><strong>AMD</strong> agreed to acquire <strong>World Labs</strong>, the Fei-Fei Li startup building spatial world models for robotics and simulation, for $8.2 billion; Li becomes AMD's executive vice president and chief scientist, and the deal is expected to close before year-end.</p></li><li><p><strong>General Intuition</strong> raised $220 million at a $6.2 billion valuation, led by <strong>Valor Equity Partners</strong> and <strong>Atreides Management</strong> with <strong>776</strong>, <strong>Point72</strong>, <strong>Khosla Ventures</strong>, and <strong>General Catalyst</strong>, nearly tripling its Series A mark, to build world models trained on action-labeled gameplay video for robotics and simulation customers.</p></li><li><p><strong>SiMa.ai</strong> raised $150 million in a Series C at a $1.45 billion valuation, co-led by <strong>Fidelity</strong> and <strong>Amplify Partners</strong>, to scale its chips and software for robots, drones, and vehicles.</p></li><li><p><strong>FieldAI</strong>, which builds navigation models for robots, is reportedly raising $700 million at about a $10 billion valuation, five times its August 2025 mark, with more than $135 million in customer contracts.</p></li></ul><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.trovereport.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Trove! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Recursion Turned an Internal Model into a $12M Product]]></title><description><![CDATA[Recursion spent years building AI to discover drugs. That data helped them build models useful for their industry]]></description><link>https://www.trovereport.com/p/recursion-turned-an-internal-model</link><guid isPermaLink="false">https://www.trovereport.com/p/recursion-turned-an-internal-model</guid><dc:creator><![CDATA[Brian D'Erario]]></dc:creator><pubDate>Mon, 28 Sep 2026 18:04:25 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/a177af55-050f-4827-919c-5a3de0ecd120_6144x3240.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Recursion was founded to discover drugs, not to sell foundation models.</p><p>The company built automated laboratories, ran biological experiments at industrial scale, and photographed how human cells responded to diseases, genes, and potential treatments. Machine learning helped its scientists turn those images into predictions about which compounds might work.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.trovereport.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Trove! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>Alongside drug candidates, that system produced proprietary data, software, and models that captured how Recursion worked.</p><p>Now one of those models has a customer. Under an <a href="https://www.sec.gov/Archives/edgar/data/1601830/000160183026000112/rxrx-20260915.htm">agreement signed this month</a>, Tempus will pay Recursion $12 million for a two-year license to TxFM, a foundation model for RNA-sequencing data. Tempus can use it in oncology research, diagnostics, and clinical applications.</p><p>The license points to a larger path for AI-native companies. An enterprise begins by building models to improve its own work, while its operations keep producing proprietary data. As the models become more specialized, useful, and difficult for an outside vendor to reproduce, the system built to run the company can become something it sells to its own industry.</p><p>Recursion is an early example of that progression. It remains a drug-discovery company, but its model has become a product.</p><h2>The model inside the company</h2><p>Recursion has <a href="https://www.recursion.com/news/the-spark-for-decoding-biology-an-origin-story">used machine learning since its earliest experiments</a>. Its founding idea was to replace subjective judgments about whether diseased cells looked healthier with repeatable computational measurements. It built <a href="https://www.sec.gov/Archives/edgar/data/1601830/000119312521117033/d89478ds1a.htm">automated laboratories and a continuous learning loop</a>: experiments generated standardized data, models learned from the results, and those models helped choose the next experiments. For years, the model was infrastructure. The commercial output was expected to be a medicine or pharmaceutical partnership.</p><p>Foundation models widened that possibility. TxFM grew from Recursion&#8217;s work across cellular imaging, chemistry, and gene expression. The <a href="https://arxiv.org/abs/2605.31562">public paper</a> describes a version trained on 1.4 million curated public RNA-sequencing samples; Recursion has <a href="https://ir.recursion.com/news-releases/news-release-details/recursion-reports-first-quarter-financial-results-and-provides">separately said</a> the model used public and proprietary data.</p><p>In 2023, Recursion <a href="https://www.sec.gov/Archives/edgar/data/1601830/000160183023000068/rxrx-20231109.htm">agreed to pay as much as $160 million</a> for limited access to more than 20 petabytes of Tempus oncology data. This month, it replaced two scheduled $42 million annual fees with three $14 million payments, accepted a lower cap on unique records, and surrendered its right to terminate for convenience. Tempus traded the possibility of larger payments for $42 million of longer-term commitments.</p><p>In a <a href="https://www.sec.gov/Archives/edgar/data/1601830/000160183026000112/rxrx-20260915.htm">separate agreement signed the same day</a>, Tempus licensed TxFM for $12 million and agreed to provide additional de-identified pathology records linked to clinical data. The deal captures why unique datasets are becoming so valuable. Models will become easier to build and distribute, but the proprietary records created inside a company&#8217;s daily work, together with the outcomes needed to evaluate them, are far harder to reproduce. That data can give an enterprise a durable advantage in building models tuned to the problems it understands best.</p><p>At first, those models will serve the company&#8217;s own use cases by lowering costs, improving decisions, or strengthening its core product. Once they perform well inside their workflows, the company can sell access to other organizations facing the same problems but lacking the same data, expertise, or operating history. Recursion&#8217;s agreement with Tempus is an early example of that progression. TxFM began as part of Recursion&#8217;s drug-discovery machinery; it&#8217;s now a product another healthcare company will pay to use. The long-term opportunity is to turn what proprietary data teaches into models sold beyond the walls of the company that trained them.</p><h2>Partnerships &amp; Deals</h2><h3>Compute, Energy &amp; Infrastructure</h3><ul><li><p><strong>Anthropic and Akamai</strong> signed a seven-year, $11.6 billion cloud agreement for CPU workloads. The commitment could expand by another $9 billion, while an Akamai warrant issued to Anthropic can vest into as much as roughly 5 percent of the cloud company&#8217;s common stock.</p></li><li><p><strong>NetApp</strong> agreed to acquire PEAK:AIO to add independently scalable metadata services and parallel file architecture for AI clouds operating trillions of files and multi-exabyte storage environments.</p></li><li><p><strong>Sunrun and SPAN</strong> expanded their partnership to combine residential solar and battery systems with SPAN&#8217;s distributed data-center nodes, targeting behind-the-meter AI compute deployments at gigawatt scale.</p></li></ul><h3>Data, Knowledge &amp; Retrieval</h3><ul><li><p><strong>TinyFish and 15 data providers</strong> launched the Data Partners Alliance, connecting agents to licensed market, company, identity, legal, research, and specialized data through TinyFish&#8217;s web infrastructure. Founding members include Alpha Vantage, Databento, Crunchbase, Similarweb, Tracxn, Enigma, OpenAlex, and Trellis Law.</p></li><li><p><strong>Progress Software</strong> completed its $400 million acquisition of substantially all of Domo&#8217;s AI and data platform business, adding more than 2,400 customers and technology for connecting, governing, and activating enterprise data.</p></li></ul><h3>Models, Training &amp; Developer Tools</h3><ul><li><p><strong>Snorkel AI</strong> raised a $350 million Series E at a $3.5 billion valuation to expand its training- and evaluation-data business for frontier labs, enterprises, and government agencies.</p></li><li><p><strong>Ando</strong> raised a $20 million seed round from Accel, Index Ventures, and Emergence Capital to launch a messaging platform where people and AI agents share channels, context, identities, and permissions.</p></li></ul><h3>Enterprise Deployment &amp; Distribution</h3><ul><li><p><strong>BNP Paribas and Google Cloud</strong> signed a five-year partnership covering infrastructure, Gemini models, and Gemini Enterprise. Initial plans include integrating Gemini into the bank&#8217;s internal LLM@CIB assistant and deploying agents for corporate credit memos and other investment-banking workflows.</p></li><li><p><strong>Accenture</strong> invested in Within and formed a delivery partnership around Within&#8217;s system for capturing undocumented processes, exceptions, and workarounds as context for enterprise agents.</p></li></ul><h3>Physical AI &amp; Robotics</h3><ul><li><p><strong>Cognex</strong> agreed to acquire RealSense for approximately $500 million in cash, adding 3D depth cameras and robotic-perception technology used in autonomous mobile robots, industrial automation, quadrupeds, and humanoids. The transaction is expected to close in the fourth quarter.</p></li></ul><h3>Science &amp; Healthcare</h3><ul><li><p><strong>Enveda</strong> raised a $311 million Series E led by Catalio Capital Management to advance three clinical-stage medicines, bring additional programs into trials, and expand its AI platform for finding drug candidates in natural chemistry.</p></li><li><p><strong>Harell Data and CoreWeave</strong> signed a multi-year agreement to run training, fine-tuning, and inference against proprietary biotech datasets without releasing the raw data. Harell says data owners will receive a share of each training run while model builders retain their models and can charge for their use.</p></li></ul><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.trovereport.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Trove! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[The Dead-Data Market]]></title><description><![CDATA[SpaceXAI has reportedly discussed buying the operational memory of dead startups.]]></description><link>https://www.trovereport.com/p/the-dead-data-market</link><guid isPermaLink="false">https://www.trovereport.com/p/the-dead-data-market</guid><dc:creator><![CDATA[Brian D'Erario]]></dc:creator><pubDate>Mon, 21 Sep 2026 16:03:05 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/e2755322-7d26-4bd2-9470-be605356f346_2700x1600.webp" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Bloomberg reported on September 17 that SpaceXAI has held internal discussions about buying customer and operational records from troubled or bankrupt startups to train Grok. Labs are quickly identifying that records of work, whether clean or messy, are going to be more valuable as broad internet data is commoditized and labs continue to move deeper into the enterprise.</p><p>A dead startup is leaving behind years of Slack threads where real decisions got made, Jira tickets showing how engineers triaged bugs, email chains where deals moved or died, support tickets with the messy back-and-forth of actual customers. That's the closest thing in existence to a recording of human judgment at scale, and until recently nobody priced had a use case for it.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.trovereport.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Trove! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>SimpleClosure told Forbes it facilitated nearly 100 workplace-data transactions in a year, with payouts generally running $10,000 to $100,000 per company. Cielo24 reportedly sold a bundle of Slack messages, internal emails, and Jira tickets for hundreds of thousands of dollars. Small numbers, but real ones and the scale will only get larger as enterprises identify opportunities to sell de-anonymized information that preserves their customer records but provide labs with key model training data.</p><p>The auction that made the market visible was Spirit Airlines'. <a href="https://www.trovereport.com/p/data-the-next-asset-class">I covered the details in Trove's first issue</a>: Google's $10 million winning bid for the airline's operational corpus, Mercor's $7.5 million backup bid, Micro1's reportedly late $12.5 million offer. This was the first visible example of the market taking shape, both bankrupt and active companies data is now for sale.</p><p>Biotech is running a similar play, with the OpenAI Foundation recently giving $500,000 to 1Day Sooner for an effort called CTD Commons, which buys regulatory dossiers from failed biotech companies: toxicology data, manufacturing records, FDA correspondence. The group's president figures nonexclusive copies could go for a few tens of thousands of dollars each.</p><p>The labs don&#8217;t need another scrape of the web. They needy to see how people actually decided things, and the cheapest place to buy that may be companies that no longer have a use case for their emails and slack messages.</p><p></p><h2>Partnerships &amp; Deals</h2><h3>Compute, Energy &amp; Infrastructure</h3><ul><li><p><strong>Crusoe</strong> raised $3.9 billion in a Series F at a $30.9 billion post-money valuation. The company says its vertically integrated AI infrastructure platform has more than $140 billion in total contracted value.</p></li><li><p><strong><span>Crusoe</span></strong> signed a multi-year deal to run <span>Perplexity&#8217;s</span> full model lifecycle on Crusoe Cloud, training frontier models on dedicated NVIDIA GB300 NVL72 clusters and serving them through managed inference, while Crusoe rolls out Perplexity Enterprise to its 1,800 employees.</p></li><li><p><strong>Euclyd</strong> raised more than &#8364;200 million in a Series A co-led by Samsung, Somerset Capital Partners, the Scaleup Europe Fund, and Innovation Industries to develop energy-efficient AI inference chips, memory architecture, and data-center systems.</p></li><li><p><strong>Emerald AI, Google, and NVIDIA</strong> launched the AI Energy Management Alliance with Anthropic, utilities, power producers, and infrastructure companies. The group will develop standards for data centers that adjust electricity use in response to grid conditions.</p></li></ul><h3>Data, Knowledge &amp; Retrieval</h3><ul><li><p><strong>Infillion</strong> agreed to acquire Foursquare, adding a location-data business with more than 100 million points of interest, 16 billion human-verified check-ins, and aggregated coverage of 250 million U.S. devices to its advertising platform.</p></li><li><p><strong>Village Media and OpenAI</strong> partnered to build Open Door, an AI-powered community health and care navigation service. Village Media will own the product, while OpenAI is providing funding, API credits, and technical support.</p></li><li><p><strong><span>Google</span></strong> is piloting a &#8220;pay per value&#8221; AI licensing model with publishers, paying based on how valuable their content was for generating responses across Gemini, AI Overviews, and AI Mode. Dozens of publishers have been approached; early payouts are reportedly tiny next to ad revenue, but most would rather be inside the tent than out.</p></li><li><p><strong><span>OpenText</span></strong> and <strong><span>Cohere</span></strong> announced a strategic partnership at the ALL IN AI conference in Montreal, pairing Cohere&#8217;s agentic AI platform with OpenText&#8217;s enterprise data layer for governments and regulated industries that need private or sovereign deployment options.</p></li><li><p><strong><span>Mozilla Data Collective</span></strong> raised $5 million from <strong><span>Mozilla</span></strong> to scale its consented, provenance-tracked training datasets, now spanning more than 450 languages with 350 approved contributing organizations.</p></li><li><p><strong><span>Keewano</span></strong> launched with $12 million in seed funding led by <strong><span>Hetz Ventures</span>, </strong>with<strong> <span>a16z speedrun</span> </strong>participating, for a database built for AI agent reasoning that keeps data in sequence and context.</p></li><li><p><strong><span>Antfly</span></strong> raised $2 million led by <strong><span>Heavybit</span></strong> to build a unified retrieval engine combining lexical, semantic, and graph retrieval with fine-grained access controls, deployable self-hosted or in Antfly Cloud.</p></li></ul><h3>Models, Training &amp; Developer Tools</h3><ul><li><p><strong>Cohere and Aleph Alpha</strong> signed a definitive business combination agreement, advancing a plan first announced in April. The combined sovereign-AI company will operate as Cohere, employ more than 1,000 people, and be dual-headquartered in Toronto and Berlin, subject to regulatory approval.</p></li><li><p><strong>TypeSafe AI</strong> released Jev, its first System One model, after two years in stealth. The company built a new architecture and a training method called Reinforcement Learning for Calibrated Decisions so the model returns typed, probability-scored decisions rather than generated text for uses such as classification, routing, and agent monitoring.</p></li><li><p><strong>Arcee AI</strong> raised a Series B led by Vista Equity Partners, Cambium Capital, and Emergence Capital at a valuation above $1 billion. The company will use the funding to develop open-weight foundation models and enterprise deployment tools.</p></li><li><p><strong>AIUC</strong> raised a $40 million Series A led by Ribbit Capital to expand its auditing, standards, and insurance system from AI agents to frontier models.</p></li><li><p><strong>MIND</strong> raised a $72 million Series B led by Crosspoint Capital Partners to expand its AI-native data-loss-prevention platform. The round brings its total funding to $112 million.</p></li><li><p><strong>Accenture and Anthropic</strong> will build a team of evaluators embedded inside Anthropic to red-team models, assess alignment, and test safeguards. Each company expects to invest at least $1 billion in AI safety over five years.</p></li></ul><h3>Enterprise Deployment &amp; Distribution</h3><ul><li><p><strong>Salesforce and Google Cloud</strong> expanded their partnership so agents on Agentforce and Gemini Enterprise can reason and act across a shared data foundation. Salesforce workloads will also run on Google Cloud through Hyperforce.</p></li><li><p><strong>Mozilla and Mistral</strong> partnered to bring Mistral Small 4 to Firefox Smart Window beta in the United States and Canada while expanding the product to France with French-language support.</p></li><li><p><strong>Nokia and Microsoft</strong> integrated Nokia Data Suite with Microsoft Fabric to create a governed data foundation for agent-driven telecom operations, including network assurance, root-cause analysis, and predictive maintenance.</p></li><li><p><strong>Siemens and Salesforce</strong> connected Siemens Teamcenter service data with Agentforce so industrial sales and service teams can retrieve engineering information inside customer workflows. Siemens is also using Agentforce to qualify inbound leads for its 18,000 sellers.</p></li></ul><h3>Industry Applications &amp; Workflows</h3><ul><li><p><strong>TotalEnergies and Mistral</strong> launched a three-year, &#8364;100 million-plus program to train frontier models on nearly a century of geoscience and reservoir-engineering knowledge and almost 10 petaflops of subsurface data.</p></li><li><p><strong>Magentic</strong> raised an $18 million Series A led by Felicis to expand AI agents that work through email, Microsoft Teams, and enterprise systems for manufacturers.</p></li><li><p><strong><span>EnforceShield</span></strong> raised &#8364;1.7 million led by <strong><span>Vendep Capital</span></strong> to automate the IP enforcement lifecycle across marketplaces, social media, and ad channels, already removing more than 33,500 violations a month for 20+ enterprise clients.</p></li></ul><h3>Physical AI &amp; Robotics</h3><ul><li><p><strong>Treble Technologies</strong> raised $18 million in a Series A-2 led by Paladin Capital Group to expand its acoustic simulation and synthetic-data platform for robots, wearables, and other audio-enabled systems.</p></li><li><p><strong>Arrive AI and DXC Technology</strong> expanded their partnership to deploy secure exchange points for drones, autonomous mobile robots, vehicles, and human couriers across large manufacturing campuses.</p></li><li><p><strong><span>D-Robotics</span></strong> closed a $400 million Series C reportedly led by <strong><span>Mirae Asset</span>to </strong>to<strong> </strong>own the picks-and-shovels layer of physical AI, with 8 million robot chips shipped and the S600 adopted by more than 20 robotics companies.</p></li><li><p><strong><span>Shutu Technology</span></strong> raised nearly RMB 100 million in a pre-Series A co-led by <strong><span>CDH Baifu</span></strong> and <strong><span>Kunpeng Fund</span></strong> for infrastructure that converts human behavior into structured robot-training data. The company says its data products serve more than 60% of China&#8217;s embodied-AI unicorns valued above RMB 10 billion; the figure is company-reported.</p></li><li><p><strong><span>Embra AI</span></strong> raised $1 million in pre-seed funding from investors with <span>Boston Dynamics</span>, <span>Agility Robotics</span>, and <span>NVIDIA</span> backgrounds to build data infrastructure for robotics teams: dataset discovery, evaluation, and sourcing across multimodal data.</p></li></ul><h3>Science &amp; Healthcare</h3><ul><li><p><strong>Novo Nordisk and Anthropic</strong> partnered to use Claude Science and Anthropic models for biological reasoning, drug discovery, and agentic software development inside Novo&#8217;s research organization.</p></li></ul><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.trovereport.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Trove! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Latham & Watkins Keeps Their AI Options Open]]></title><description><![CDATA[The firm wants the option to change models without surrendering their data, evaluations, and legal expertise.]]></description><link>https://www.trovereport.com/p/latham-and-watkins-keeps-its-ai-options</link><guid isPermaLink="false">https://www.trovereport.com/p/latham-and-watkins-keeps-its-ai-options</guid><dc:creator><![CDATA[Brian D'Erario]]></dc:creator><pubDate>Mon, 14 Sep 2026 16:56:32 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/f2ce9189-a7d1-40a3-88aa-88c1dd782e23_1400x933.avif" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><span>For three years, Latham &amp; Watkins has invested to stay at the technological frontier, and the strategy appears to be paying off.</span></p><p><span>The firm currently runs Harvey across the company, while also leveraging ChatGPT, Claude, and Gemini. And in locked space at a third-party data center, reachable only by Latham staff, the firm keeps several Nvidia servers of its own, where its engineers fine-tune open-weight models on the work the firm does every day.</span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.trovereport.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Trove! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p><span>&#8220;We are not hitching our wagon to one particular company,&#8221; Rene Mendoza, the firm&#8217;s CIO, told the Financial Times. &#8220;Nobody can predict where any of this is going, so we have the best optionality.&#8221; Legal Cheek called the move a first for BigLaw.</span></p><p><span>Latham doesn&#8217;t train models from scratch. It downloads open-weight models, adapts them to legal work, and wires them into its own software. The firm settled on the direction three years ago, runs Nvidia H200 chips today, and is looking at Blackwell systems next. It won&#8217;t say what it paid but Legal Cheek puts a build of this size, hardware plus the specialists to run it, in the tens of millions a year. Not a small chunk of change for an insurance policy, as Harvey still runs firm-wide for research, document analysis, and drafting. Michael Rubin, who leads the firm&#8217;s AI work, told Bloomberg Law the buildout was never meant to cut reliance on Harvey. It&#8217;s been there as a hedge to ensure they can remain on the frontier, whatever direction the AI industry takes it.</span></p><p><span>Another great reason to want the optionality came earlier this year when Anthropic launched Claude for Legal in May, with connectors into Harvey, iManage, NetDocuments, Relativity, and Thomson Reuters. Google shipped Gemini Enterprise for Legal in August, and Cleary Gottlieb, Freshfields, Weil, and Williams &amp; Connolly signed on as launch customers. OpenAI hired Ironclad&#8217;s co-founder in June to run a legal vertical and is preparing its own legal integrations.</span></p><p><span>The companies that supply Latham&#8217;s models are now selling into Latham&#8217;s market. A firm running only rented models is exposed twice: to a vendor&#8217;s price and access terms, and to that vendor competing for the same client work.</span></p><p><span>THE ASSET THAT COMPOUNDS</span></p><p><span>Legal work is a hard case for the cloud. Privileged material, deal documents, client confidences: the files a firm least wants passing through someone else&#8217;s API. Owning the hardware answers that directly, and it does something else too.</span></p><p><span>Latham has about 100 of its 900 technologists on AI, and it keeps hiring innovation attorneys and AI services attorneys. &#8220;No AI company, no technology company will ever be able to replicate what we&#8217;re doing because they don&#8217;t have our lawyers,&#8221; partner Amber Banks told Bloomberg Law.</span></p><p><span>A firm&#8217;s most useful training material is its own casework: the briefs it wrote, the clauses it negotiated, the corrections a senior lawyer made to a junior&#8217;s draft. Using a vendor&#8217;s model or app out of the box, without integrating your own data into the model training itself, is taking away part of your competitive advantage against other firms using the same providers.</span></p><p><span>A frontier model arrives knowing nothing about how a particular firm argues a particular kind of case. Latham can take an open-weight model, RLHF using data sets only they have, and improve the models capability for the work they do. Every model generation makes the set more useful, because the firm is applying a better model to better examples. I&#8217;m sure they can take their model and also apply it to Harvey (or the legal app of their choice outside of the labs), for the app&#8217;s superior workflows.</span></p><p>The technology itself is no longer an edge. For firms like Latham, their advantage will be their relationships and decades of proprietary data.</p><h2>Pacing the Frontier</h2><p>Over the weekend, three rivals who rarely agree issued the same warning and offered three different cures. Anthropic&#8217;s Dario Amodei called for pacing frontier capability gains. OpenAI&#8217;s Sam Altman backed independent evaluators embedded inside the labs. Elon Musk wants peer review by competing labs. Gavin Baker made a useful distinction: the headlines describe a broad slowdown, but the only concrete change so far is third-party evaluation at OpenAI and Anthropic.</p><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://x.com/gavinsbaker/status/2099170698887315833?s=46&quot;,&quot;full_text&quot;:&quot;Wild 24 hours for AI and lots of different proposals have been made.\n\nTLDR; the only *tangible* new fact is that OpenAI and Anthropic are going to have embedded 3rd party evaluators from unknown organizations with Dario floating METR as a possibility. Having 3rd party evaluators&#8230;&quot;,&quot;username&quot;:&quot;GavinSBaker&quot;,&quot;name&quot;:&quot;Gavin Baker&quot;,&quot;profile_image_url&quot;:&quot;https://pbs.substack.com/profile_images/1396219525754937345/5L4n5L3O_normal.jpg&quot;,&quot;date&quot;:&quot;2026-09-13T16:17:49.000Z&quot;,&quot;photos&quot;:[],&quot;quoted_tweet&quot;:{},&quot;reply_count&quot;:129,&quot;retweet_count&quot;:158,&quot;like_count&quot;:971,&quot;impression_count&quot;:128102,&quot;expanded_url&quot;:null,&quot;video_url&quot;:null,&quot;video_preview_media_key&quot;:null,&quot;belowTheFold&quot;:true}" data-component-name="Twitter2ToDOM"></div><p>While the announcement itself has taken X by storm over the weekend, I was left wondering what does this mean for enterprises deploying AI?</p><p>In an environment where yes, the frontier labs are putting out the best product today, many are starting to turn to open-weights. The models aren&#8217;t at the frontier but they&#8217;re getting very close, and you don&#8217;t have to worry about your vendor cutting off your model access or launching a direct competitor into the market. </p><p>On top of that, <a href="https://api-docs.deepseek.com/news/news260910/">DeepSeek&#8217;s own release</a> for v4.1 flash (a model many are saying is nearing or exceeding Opus 5 performance) noted that quality of data mattered just as much if not more than volume of data in this training set. If you&#8217;re a well run firm in your industry, who is going to have better data to train models set to do work in your business than you? </p><h2>Partnerships &amp; Deals</h2><h3>Compute, Energy &amp; Infrastructure</h3><p><strong>Google</strong> committed &#8364;13 billion over two years to expand AI and cloud infrastructure in Finland, including data-center capacity and related energy investments.</p><p><strong>Positron AI</strong> raised an $875 million Series C at a $5 billion valuation to bring its Asimov inference silicon and Titan systems to market.</p><p><strong>TAR</strong> raised a $120 million Series A led by Spark Capital at a $1 billion post-money valuation. The company builds modular, off-grid renewable power and battery systems for AI data centers.</p><p><strong>Qualcomm</strong> and <strong>Amazon</strong> entered a multi-generation collaboration on customized silicon for AWS AI inference and optical connectivity for data-center networks.</p><p><strong>Oracle</strong> booked more than $30 billion of new AI cloud contracts in its latest quarter, lifting remaining performance obligations to $664 billion. It also reported 850 megawatts of added capacity, more than 300,000 additional GPUs, and a completed $20 billion equity sale.</p><p><strong>Palantir</strong> named <strong>Nebius</strong> its preferred sovereign AI infrastructure partner. Nebius compute and inference endpoints will sit inside Palantir&#8217;s enterprise perimeter so eligible customers can control their compute, data, and models.</p><h3>Data, Knowledge &amp; Retrieval</h3><p><strong>OpenAI</strong> signed content partnerships with <strong>BCCL</strong>, publisher of The Times of India and The Economic Times, and the <strong>Indian Express Group</strong>. ChatGPT can surface attributed summaries and excerpts from live and archived reporting across seven languages.</p><p><strong>OpenAI</strong> launched ChatGPT for Financial Services after design work with <strong>Morgan Stanley</strong> and <strong>Evercore</strong>. Premium data from <strong>Daloopa</strong>, <strong>PitchBook</strong>, and <strong>LSEG News</strong> is indexed and hosted by OpenAI with source-level citations.</p><p><strong>Box</strong> partnered with <strong>OpenAI</strong> so users can browse, access, and work with governed Box content directly inside ChatGPT.</p><p><strong>Versos AI</strong> launched an agentic curation system built with NVIDIA NeMo and LangChain that turns natural-language requirements into structured, rights-cleared video training datasets while preserving ownership and provenance records.</p><p><strong>Universal Music Group</strong> and <strong>ElevenLabs</strong> signed a multi-year licensing and product-development agreement. Their first product will let fans create with music from participating artists, with additional AI audio tools planned for artists and songwriters.</p><h3>Models, Training &amp; Developer Tools</h3><p><strong>Mistral AI</strong> raised &#8364;3 billion in a Series D at a post-money valuation above &#8364;21 billion. <strong>Samsung</strong> led the financing and paired its investment with plans to deploy Mistral models inside semiconductor operations.</p><p><strong>Harvey</strong> raised $550 million at a $15.5 billion valuation to expand its legal and professional-services AI platform.</p><p><strong>Harvey</strong> acquired <strong>Guardrails AI</strong>, an agent-security and simulation company whose tools test where AI agents depart from intended behavior. The team will join Harvey&#8217;s product and engineering organization.</p><p><strong>Cloudera</strong> and <strong>Mistral AI</strong> partnered to run and customize Mistral models over governed enterprise data across cloud, on-premises, edge, sovereign, and air-gapped environments.</p><p><strong>Meta</strong> agreed to acquire Swedish agent startup <strong>Stilla.ai</strong> to strengthen Meta Business Agent across WhatsApp, Messenger, and Instagram.</p><h3>Enterprise Deployment &amp; Distribution</h3><p><strong>Accenture</strong> and <strong>Google Cloud</strong> created the Accenture Gemini Enterprise Business Group. The companies plan a 1,000-person forward-deployed engineering workforce to implement agentic AI and data projects.</p><p><strong>OpenAI</strong> and the <strong>U.S. General Services Administration</strong> signed a 27-month OneGov agreement offering eligible federal, state, local, and tribal organizations a zero-dollar license fee and 50 percent off usage. Eligibility extends across a public-sector workforce of approximately 23 million.</p><p><strong>Amazon Ads</strong> partnered with <strong>OpenAI</strong> to let selected U.S. advertisers extend Amazon Ads campaigns into a ChatGPT Ads pilot.</p><p><strong>Palantir</strong> and <strong>NVIDIA</strong> introduced a sovereign AI stack for critical supply chains, combining Palantir&#8217;s ontology and software with customized NVIDIA Nemotron models. The first deployment is inside NVIDIA&#8217;s own supply chain.</p><p><strong>Cisco</strong>, <strong>Palantir</strong>, and <strong>NVIDIA</strong> are integrating Palantir&#8217;s cybersecurity ontology with Cisco&#8217;s Secure AI Factory to give regulated organizations an AI architecture they can operate under their own controls.</p><p><strong>Fujitsu</strong> expanded its partnership with <strong>Palantir</strong>, signing a new agreement for AIP and Foundry and becoming a Global Forward Deployed Engineering partner for sales, use-case design, and implementation.</p><p><strong>Avid</strong> and <strong>Google Cloud</strong> expanded their partnership with a browser-based Media Composer and new Gemini Enterprise and BigQuery integrations for media search, editing, and production workflows.</p><p><strong>Morgan State University</strong> and <strong>Google Public Sector</strong> partnered on an AI research campus with high-performance GPU infrastructure, sovereign-AI governance, cybersecurity tools, and a Google-focused Center of Excellence.</p><p><strong>Meta</strong> launched Muse with <strong>Stripe</strong> Link purchase protection and one-time payment cards. Shop Pay and 1Password integrations are planned as additional commerce and credential layers.</p><p><strong>Napster</strong> and <strong>Kameha Ventures</strong> formed a strategic partnership to develop locally guided multimodal-agent deployments for government and private-sector customers in the United Arab Emirates.</p><h3>Industry Applications &amp; Workflows</h3><p><strong>Sony Semiconductor Solutions</strong> and <strong>Aramco</strong> signed a non-binding memorandum to combine Sony sensing and edge-AI technology with Aramco&#8217;s industrial data, infrastructure, and operating environments.</p><p><strong>UNDP</strong> and <strong>NEC</strong> signed an agreement to use AI and environmental data for nature conservation, climate resilience, supply chains, and assessment of digital infrastructure including data centers.</p><h3>Physical AI &amp; Robotics</h3><p><strong>Maven Robotics</strong> launched with a $100 million Series A to build general-purpose robotic systems for material handling and assembly in logistics and manufacturing.</p><p><strong>Vecna Robotics</strong> raised $31 million led by Unless to scale autonomous forklifts, tuggers, case-picking systems, and new dock automation capabilities.</p><p><strong>Swarmer</strong> signed a definitive agreement to acquire Ukrainian unmanned-ground-vehicle maker <strong>Ratel Robotics</strong> for cash and stock worth up to $224 million if earnout milestones are met.</p><p><strong>NEURA Robotics</strong> and <strong>SECO</strong> partnered on compute modules for NEURA&#8217;s cognitive robots, including the 4NE1 humanoid, and on real-world industrial data and automation for semiconductor manufacturing.</p><p><strong>JOYX</strong> and <strong>Hitch Interactive</strong> partnered to combine verified human and robot demonstration data with standardized robot interfaces and real-world deployment environments.</p><h3>Science &amp; Healthcare</h3><p><strong>Owkin</strong> licensed its K Pro AI scientist and multimodal patient data from the MOSAIC network to <strong>Servier</strong> for oncology research. This is a new September 11 agreement, separate from Owkin&#8217;s earlier Boehringer Ingelheim license.</p><p><strong>Tempus AI</strong> received an award of up to $9.5 million from <strong>ARPA-H</strong> to develop and clinically validate an autonomous AI agent for heart-failure care using clinical and consumer-generated data.</p><p><strong>IBM</strong> and <strong>NASA</strong> released an open-source Lunar Foundation Model and a machine-learning-ready dataset combining more than 30 aligned layers from nine instruments across four lunar missions.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.trovereport.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Trove! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Whose Data is it Anyway?]]></title><description><![CDATA[A mathematician, his unpublished drafts, and the one question no AI provider can answer.]]></description><link>https://www.trovereport.com/p/whose-data-is-it-anyway</link><guid isPermaLink="false">https://www.trovereport.com/p/whose-data-is-it-anyway</guid><dc:creator><![CDATA[Brian D'Erario]]></dc:creator><pubDate>Thu, 10 Sep 2026 12:57:41 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/a43159b5-837f-412d-89ea-bbe78fd21db2_1254x1254.webp" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>The dispute between OpenAI and two mathematicians turns on a practical question: when unpublished work is entered into an AI product, can anyone later determine whether it influenced the model? In this case, the answer appears to be no.</p><p>Tristan Buckmaster, a mathematics professor at NYU, and Levent Alp&#246;ge, a researcher at Anthropic, spent roughly a year on a private collaboration around the Navier&#8211;Stokes equations. They used Codex along the way, including for drafts of their work. By late August, they had three Lean-verified proofs of finite-time blowup and believed they had a path toward the larger problem.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.trovereport.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Trove! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>On September 3, Buckmaster learned that OpenAI had heard of the work. In meetings three days later, he says OpenAI representatives told him that an internal model had produced a long proof of blowup for a version of Navier&#8211;Stokes. OpenAI says its own effort began on September 1 after rumors that Millennium problems had fallen. The parties disagree about the surrounding discussion, including how much human work went into OpenAI&#8217;s result.</p><p>Buckmaster&#8217;s central question was narrower: had the model been trained on, or otherwise had access to, the drafts he and Alp&#246;ge had entered into Codex? He writes that he did not receive an answer. OpenAI says its researchers and agents did not see the pair&#8217;s work before it was public and that no specific user data was accessed to solve the problem. It also says it cannot rule out that de-identified data derived from product use helped improve its models.</p><p>The strategic issue is bigger than one training record. As intelligence becomes cheaper and more available, the value moves downstream into products, discoveries and operating systems. Model companies will follow it. They won&#8217;t remain neutral infrastructure providers.</p><p>Anthropic is an early example. It launched Claude Science while also announcing a drug-discovery program focused on neglected diseases. That isn&#8217;t evidence of customer-data misuse, but it is evidence that a company can sell a platform to drugmakers while becoming a participant in drug discovery and that changes the relationship to pharmaceutical manufacturing companies.</p><p>Frontier models may still be the best tool for a proof search, simulation or design loop. But agents use proprietary context and produce valuable work. The decision isn&#8217;t only which model gives the best answer. It is where the company&#8217;s data, agent workflow and resulting work product should live.</p><p>It continues to drive the case for Sovereign AI to ensure you&#8217;re keeping your data protected and you remain competitive with model companies who may turn into competitors in your field. Buckmaster&#8217;s case is not proof of wrongdoing. It is a warning about where the market is going. As model suppliers continue moving into more industries directly, ownership of your data and your AI infrastructure becomes more important than just choosing your favorite model.</p><h2>Model Releases</h2><p><strong>Anthropic</strong> released Claude Fable 5.1 and Claude Mythos 5.1. Fable is the broadly available version with cyber and biology safeguards; Mythos is restricted to vetted cyberdefenders and life-sciences researchers.</p><p><strong>Google</strong> released Gemini 3.8 Flash and Gemini 3.8 Flash Cyber, its new agentic and coding workhorse plus a cybersecurity variant.</p><p><strong>Alibaba</strong> updated Qwen3.8-Max-0902, a one-million-token snapshot post-trained for coding, agent orchestration and visual analysis.</p><p><strong>Meta</strong> released Muse Spark 1.3 for longer agentic and coding workflows, then launched Muse, a personal agent for coordinating tasks through connected services.</p><p><strong>OpenAI</strong> released GPT-6 Astra, its most capable broadly deployed model and first to reach its Critical cyber-capability threshold. It also launched GPT-Image-2.5 Flare and Sunburst, new API image models for faster generation and higher-precision editing.</p><p><strong>DeepSeek</strong> released V4.1 Flash today, the smallest model in its new architecture family and the company&#8217;s new default API model. It combines native image understanding with a one-million-token context window. DeepSeek says it outperforms the larger V4 Pro while running faster and at lower cost.</p><h2>Partnerships &amp; Deals</h2><h3>Compute, Energy &amp; Infrastructure</h3><ul><li><p><strong>Nscale</strong> signed a compute agreement with <strong>Figure</strong> carrying an initial $3.5 billion commitment and plans to expand beyond $6 billion. Nscale will also take an equity stake in Figure. The first GPUs are targeted for the second half of 2027 in Barstow, Texas.</p></li><li><p><strong>Crusoe</strong> raised more than $3 billion at a roughly $30 billion valuation in a round co-led by <strong>Atreides Management</strong> and <strong>Valor Equity Partners</strong>, Bloomberg reported. The company builds data centers and cloud infrastructure for AI workloads.</p></li><li><p><strong>Fluidstack</strong> raised $1.5 billion at an $18 billion valuation in a round led by <strong>Jane Street</strong>, according to Crunchbase. It supplies large-scale compute and data center infrastructure for AI.</p></li><li><p><strong>Gimlet Labs</strong> raised a $300 million Series B at a $3 billion valuation led by <strong>Andreessen Horowitz</strong>. It&#8217;s building an inference cloud that distributes workloads across GPUs and specialized accelerators.</p></li></ul><h3>Data, Knowledge &amp; Retrieval</h3><ul><li><p><strong>AfterQuery</strong> reached a reported $3.2 billion valuation, up from $300 million five months earlier, according to Forbes. The company supplies AI training data built around expert reasoning.</p></li></ul><h3>Models, Training &amp; Developer Tools</h3><ul><li><p><strong>Nvidia</strong> confirmed an agreement to acquire <strong>Hugging Face</strong> for $12.93 billion, following the unconfirmed reports in our last issue. Nvidia says the model and dataset platform will continue to support competing hardware and cloud providers.</p></li><li><p><strong>HiddenLayer</strong> raised a $100 million Series B led by <strong>Delta-v Capital</strong> to expand security tools for AI models, agents and workflows. Investors included <strong>M12</strong>, <strong>Morgan Stanley</strong> and <strong>Booz Allen Ventures</strong>.</p></li></ul><h3>Enterprise Deployment &amp; Distribution</h3><ul><li><p><strong>Wonderful</strong> raised a $550 million Series C at a $5 billion valuation led by <strong>Insight Partners</strong>, with <strong>Salesforce</strong> participating. The funding will support its platform for building and deploying enterprise agents and applications.</p></li><li><p><strong>MOZN</strong> announced partnerships with <strong>Redington</strong>, <strong>sirar by stc</strong>, <strong>HPE</strong>, <strong>Microsoft</strong> and <strong>Edarat</strong> to expand enterprise AI in Saudi Arabia and MENA. The agreements cover regional distribution, cybersecurity, private infrastructure and locally hosted deployments.</p></li></ul><h3>Industry Applications &amp; Workflows</h3><ul><li><p><strong>HiBob</strong> raised $166 million with <strong>Salesforce</strong> providing most of the capital and <strong>Farallon Capital</strong> participating. Its HR platform brings workforce and skills data together for employee management, hiring and training decisions.</p></li></ul><h3>Physical AI &amp; Robotics</h3><ul><li><p><strong>HUMAIN</strong> and <strong>Applied Intuition</strong> signed a long-term commercial partnership targeting thousands of autonomous trucks across Saudi logistics routes by 2030. Applied Intuition will provide the self-driving system and vehicle software, with other industrial applications planned later.</p></li></ul><h3>Science &amp; Healthcare</h3><ul><li><p><strong>Owkin</strong> licensed its K Pro AI research platform and multimodal oncology data to <strong>Boehringer Ingelheim</strong>. It will also generate new immunology data under the agreement, which builds on a 2025 pilot.</p></li><li><p><strong>Merge Labs</strong> signed a multi-year license for <strong>Butterfly Network&#8217;s</strong> ultrasound-on-chip technology to develop brain-computer interfaces. The agreement includes upfront and milestone payments, hardware commitments and royalties on future commercial sales.</p></li><li><p><strong>DataHow</strong> extended its Series A through a bridge round led by <strong>Momenta</strong>, with <strong>Rockwell Automation</strong> participating. The funding will expand real-time connectivity and model-based process control in its biopharmaceutical development and manufacturing platform.</p></li></ul><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.trovereport.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Trove! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Robo Robotics]]></title><description><![CDATA[Robo Robotics are selling robot labor before their machines can work alone. Their bet is that every human intervention can train the system and make the next robot-hour cheaper than the last.]]></description><link>https://www.trovereport.com/p/robo-robotics</link><guid isPermaLink="false">https://www.trovereport.com/p/robo-robotics</guid><dc:creator><![CDATA[Brian D'Erario]]></dc:creator><pubDate>Mon, 31 Aug 2026 15:31:24 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/3ad85c98-3d8e-4212-bdac-87a3ff9f20be_2017x1055.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>The primary data set the last few years of artificial intelligence was largely digital. Companies collected text, images, video and code to train models that operated on screens. The next iteration is moving into the physical world, where machines can learn from the work people perform and the environments they perform it in.</p><p><a href="https://robo.inc/">Robo Robotics</a> entered that market last week with Robo-T, a two-armed mobile robot built to work at stations designed for people.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.trovereport.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Trove! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>Robo&#8217;s thesis: deploy inexpensive, multipurpose robots before they are fully autonomous, use human operators to keep them working and capture useful training data from selected demonstrations and interventions. If the loop works, better models will gradually reduce the human oversight each robot requires, lowering the cost of physical labor across the fleet.</p><p>I interviewed co-founders Don Morton and Kyle Noble about how the company started and their plans for building out the Robo-T and their future platforms.</p><h2>From Selling the Data to Selling a Robot</h2><p>That plan grew out of Morton and Noble&#8217;s path from software to model training and, eventually, physical data.</p><p>Noble joined baby-food company Yumi as its first engineer in 2017 and brought Morton on as its second the following year. They left in 2022 to build Atlas, beginning with a browser and later developing an education and AI product.</p><p>Model tooling was still immature, so the team built its own search index, fine-tuned models and trained models internally. &#8220;We were doing all of the AI stuff that you don&#8217;t have to do today,&#8221; they told Trove, &#8220;It&#8217;s all solved now.&#8221;</p><p>That experience led them into the training-data market. They began collecting data for video-model companies and AI labs, but conversations with prospective buyers exposed a sharper need. &#8220;What we realized as we talked to labs and companies was that the robotics companies need data the most,&#8221; the team said.</p><p>Building the arms changed how the team saw the opportunity. Once they had a working version, they started talking with customers about how it could be deployed. Most wanted something that could roll up to an existing line without requiring them to redesign their processes.</p><p>That requirement shaped Robo-T. The team placed the arms on a T-shaped body and added a movable camera head, building toward the complete robot one design decision at a time. What began as an effort to understand robotics companies as customers became a robotics company itself.</p><h2>Priced in Labor</h2><p>Robo offers two ways to use the machine.</p><p>In the first, a customer hires Robo for labor. Robo deploys and manages their platform, using whatever combination of autonomy and remote operation is required to complete the task. The customer buys the labor, irregardless that it&#8217;s performed by a machine.</p><p>In the second, an operator buys a fleet of Robo-Ts with Robo providing the hardware and operating infrastructure, while the fleet owner runs the robots themselves.</p><p>&#8220;The customer should not have to care whether there&#8217;s a human operating or whether it&#8217;s autonomous,&#8221; the team said. &#8220;The output is just the labor.&#8221;</p><p>Today, humans remain essential to Robo-T&#8217;s output. A remote operator wears a headset and uses Quest controllers to demonstrate a task or take control when the model encounters something it cannot handle. The person keeps production moving through the failures that would otherwise stop the line.</p><p>This gives Robo a wider initial task range than its autonomous models could support alone. A person can attempt anything the robot&#8217;s body can physically perform.</p><p>Robo is candid that teleoperation isn&#8217;t unique, other deployment companies can put a human behind a robot. The differentiator is whether the economics of that human in the loop work.</p><p>Every hour a person spends driving one machine is an expense attached to that robot-hour. The economics improve only when the person shifts from continuous driver to occasional supervisor.</p><p>Robo&#8217;s <a href="https://robo.inc/blog/robotics-model-t-moment">Model T launch essay</a> models a $30,000 robot amortized over two years at 240 operating hours per month, plus $1 per hour in other expenses. That creates a fixed cost of roughly $6.21 per robot-hour before teleoperation. With operators earning $25 an hour, the modeled total is $81 when three operators are assigned to each robot, $8.71 when one operator supervises 10 robots and $6.83 when one supervises 40. The exercise shows how increasing the number of robots per operator can drastically change deployment economics.</p><h2>Price buys room to learn</h2><p>Low-cost hardware does more than reduce the amount of the customer&#8217;s bill, Robo believes it changes the customer&#8217;s tolerance for an early product.</p><p>During the interview, the team used a $40,000 robot as an example. At that price, every period of downtime invites the same reaction: what was all that money for?</p><p>A cheaper machine gives the company more room to learn in public. Repairs and replacement parts cost less. Customers may accept early imperfections if the work remains economical and continues to improve. More customers can deploy robots, producing more chances to find failure cases and improve the system.</p><p>&#8220;The lower your price, the more accessible it is for people who have a higher threshold for working with something that isn&#8217;t perfect,&#8221; the team said.</p><p>Henry Ford&#8217;s achievement wasn&#8217;t simply producing a car that worked. It was imposing a price constraint that forced the company to improve manufacturing while putting enough cars into the world to create a larger economy around them.</p><p>Robo is attempting the same sequence with robot labor.</p><p>The company says demand following the launch exceeded what it could fulfill. They declined to disclose customer names, deployment counts or completed operating results, although they spoke at a high-level of upcoming work in healthcare, food production, and textile manufacturing.</p><h2>Who owns what the robot learns?</h2><p>The data loop introduces a problem that better models can&#8217;t currently solve.</p><p>A hospital or manufacturer may want robotic labor without allowing video from its facility to train a system later used by a competitor. A fleet operator may want control over the data produced by customers it sourced itself.</p><p>Robo has not publicly settled those rights yet. In the interview, Morton and Noble described a possible economic bargain rather than a finished framework, something they are actively developing.</p><p>In Robo&#8217;s proposed bargain, customers that contribute data to a shared autonomy model would benefit when the model improves. Their robots would require less teleoperation, lowering the hourly cost and improving deployment margins. Robo is also considering revenue sharing if the data can be packaged and sold to another buyer.</p><p>The team offered a hypothetical example in which an hour of video had $10 of resale value and participants in the chain received a share.</p><p>Robo&#8217;s flywheel is built around a physical-world scaling law: the more work its robots perform, the more demonstrations and useful interventions it collects, and the more capable its models become. Within an industry, data from one job trains similar tasks across the fleet. Across warehouses, factories and hospitals, recurring patterns in movement, perception and recovery strengthen the models beneath every deployment.</p><p>Each deployment expands the training set for the next. As the fleet grows, skills compound across tasks and industries, reducing human intervention and driving down the cost of robot labor.</p><p>They&#8217;ll need to find a framework that is mutually beneficial to their partners but also enables them to compound the work they&#8217;re completing and data they&#8217;re collecting to generate better models and robots.</p><h2>The Manufacturing Bottleneck</h2><p>Besides the data collection issue, Robo faces another constraint common to American robotics companies: once a product finds demand, scaling becomes as much a manufacturing problem as an AI problem. The models must improve, but the company also has to source components, assemble machines and expand production without losing the cost advantage that made deployments possible.</p><p>Morton and Noble said data and model training have a clearer development path than manufacturing. For Robo, the two are linked. It needs more robots in the field to produce data and improve autonomy, but it also needs a cost-effective supply chain across the United States and allied countries to build that fleet. That&#8217;s something the company will still need to solve as it scales, but for now it&#8217;s focused on fulfilling the demand it generated from the launch and starting to build out the flywheel.</p><h2>The Path Forward</h2><p>There is a clear playbook for improving models than for manufacturing large numbers of inexpensive robots. Robo&#8217;s answer is to begin the feedback loop immediately.</p><p>The team pointed to Unitree&#8217;s own history: make robots, sell them, collect feedback and repeat. The only way through the manufacturing pain is sale, feedback and improvement. They need to build enough robots, start them working with actual customers, and work towards a product that turns from primarily human operated to mainly automated.</p><p>From their small office in the Arts District today, Robo is trying to prove that the path to autonomous labor is not to wait for it. Deploy imperfect machines, work with real customers in complex environments, and continuously iterate through real work</p><h2>Partnerships &amp; Deals</h2><p>Transactions are grouped by the primary part of the AI supply chain they support.</p><ul><li><p><strong>Compute, Energy &amp; Infrastructure:</strong> Power, data centers, chips, servers, networking, storage and cloud capacity.</p></li><li><p><strong>Data, Knowledge &amp; Retrieval:</strong> Data rights, proprietary datasets, indexing, search and grounding systems.</p></li><li><p><strong>Models, Training &amp; Developer Tools:</strong> Models, post-training systems, agent platforms, security tools and the software used to build AI.</p></li><li><p><strong>Enterprise Deployment &amp; Distribution:</strong> Integrations, channels and services that bring AI into organizations.</p></li><li><p><strong>Industry Applications &amp; Workflows:</strong> AI products built for specific business functions and vertical markets.</p></li><li><p><strong>Physical AI &amp; Robotics:</strong> Embodied systems, autonomous machines and robotics platforms.</p></li><li><p><strong>Science &amp; Healthcare:</strong> AI-enabled biology, medicine, diagnostics, agriculture and clinical systems.</p></li></ul><h3>Compute, Energy &amp; Infrastructure</h3><ul><li><p><strong>Anthropic</strong> reportedly signed a six-year, roughly $45 billion compute agreement with <strong>Nscale</strong> covering 460 megawatts at its West Virginia campus. The companies haven&#8217;t confirmed the reported terms.</p></li><li><p><strong>Cisco</strong> expanded its Secure AI Factory with <strong>Nvidia</strong> to include rack-scale systems from <strong>Supermicro</strong>, with channel availability scheduled for October.</p></li><li><p><strong>Andreessen Horowitz</strong> raised a $1.1 billion Machine Age Fund focused on AI infrastructure and hardware, including chips, memory, networking, edge systems, appliances and robots.</p></li><li><p><strong>Emerald AI</strong> raised a $150 million Series A at a $1.05 billion valuation to build software that lets AI data centers operate as flexible grid resources.</p></li><li><p><strong>Navitas</strong> agreed to acquire AI data-center power startup <strong>Claros</strong> for approximately $232.8 million in cash and stock, including milestone consideration.</p></li><li><p><strong>Lambda</strong> closed a $926 million senior secured term-loan facility to fund GPU infrastructure for a committed investment-grade customer deployment.</p></li><li><p><strong>Lambda</strong> separately secured $1 billion of short-dated private debt to buy <strong>Nvidia</strong> chips that will be leased to <strong>Microsoft</strong>, according to Bloomberg.</p></li><li><p><strong>Lancium</strong> partnered with <strong>Nvidia</strong> across a 4-gigawatt leased AI-factory portfolio and a development pipeline exceeding 15 gigawatts, while Nvidia made a strategic investment. The pipeline isn&#8217;t the same as contracted deployed capacity.</p></li><li><p><strong>SuperX</strong> signed a supply agreement with <strong>Ezisight</strong> and received an initial order for 128 Nvidia B300 AI server clusters, with delivery planned for the fourth quarter.</p></li><li><p><strong>SCX.ai</strong> selected <strong>DDN</strong> as an infrastructure partner for what the companies describe as Australia&#8217;s largest sovereign AI inference cloud.</p></li><li><p><strong>Kasm Technologies</strong> expanded its <strong>Intel</strong> partnership to run private local AI workspaces on Xeon 6 processors with AMX acceleration.</p></li></ul><h3>Data, Knowledge &amp; Retrieval</h3><ul><li><p><strong>Keenable</strong> raised a $26 million seed round to build web-indexing, search and retrieval infrastructure for AI agents.</p></li><li><p><strong>Repodo</strong> raised a EUR 8.2 million pre-seed round to build a data and development platform for physical AI.</p></li><li><p><strong>Dun &amp; Bradstreet</strong> connected its Commercial Graph to <strong>Google Cloud&#8217;s Gemini Enterprise</strong> through MCP, grounding credit, onboarding and underwriting agents in verified business data.</p></li><li><p><strong>Dun &amp; Bradstreet</strong> brought its Commercial Graph to <strong>Perplexity</strong> through MCP connectors for research, risk and growth workflows.</p></li><li><p><strong>Reuters</strong> made its multilingual news archive, dating to 1987, available through <strong>Snowflake Marketplace</strong> for enterprise AI and analytics.</p></li></ul><h3>Models, Training &amp; Developer Tools</h3><ul><li><p><strong>Nvidia</strong> reportedly agreed to acquire <strong>Hugging Face</strong> for $12.9 billion, although neither company has announced a transaction and reports conflict over whether an agreement was signed.</p></li><li><p><strong>Deep Cogito</strong> raised a $43 million Series A to expand reinforcement-learning post-training for open-weight and enterprise models.</p></li><li><p><strong>Alice</strong> raised $140 million from <strong>Apax Digital</strong>, bringing total funding for its autonomous security-testing and remediation platform to $280 million.</p></li><li><p><strong>Stability AI</strong> raised a $76 million Series B from strategic entertainment investors to support image, video, audio and 3D models for media production.</p></li><li><p><strong>Runable</strong> raised a $21 million Series A at a $65 million post-money valuation to build agents that can operate and grow businesses.</p></li><li><p><strong>Embedd</strong> raised a EUR 2.3 million pre-seed round to build chip digital twins and agents that automate embedded-software integration for physical AI systems.</p></li></ul><h3>Enterprise Deployment &amp; Distribution</h3><ul><li><p><strong>Clearlake</strong> formed a portfolio-wide partnership with <strong>Google Cloud</strong> covering AI infrastructure, enterprise data systems, Gemini Enterprise, custom models and cybersecurity.</p></li><li><p><strong>Arga Labs</strong> raised a $10 million seed round to build enterprise agents that learn how a business operates from its data and workflows.</p></li><li><p><strong>Google Cloud</strong> and <strong>Verizon</strong> formed a strategic partnership to scale enterprise AI across customer-service, network and employee workflows.</p></li><li><p><strong>Bain</strong> partnered with <strong>Anthropic</strong> to help clients deploy Claude in regulated and operationally complex enterprises.</p></li><li><p><strong>Salesforce</strong> expanded its <strong>Anthropic</strong> partnership through Claudeforce, a package of Claude integrations and 37 prebuilt sales skills.</p></li><li><p><strong>Enabled Intelligence</strong> selected <strong>Seekr</strong> for National Geospatial-Intelligence Agency and other defense and intelligence programs.</p></li><li><p><strong>Sayari</strong> partnered with <strong>DataExpert</strong> to distribute sovereign AI tools for economic-security and investigative work across European governments.</p></li></ul><h3>Industry Applications &amp; Workflows</h3><ul><li><p><strong>Owner</strong> raised $240 million at a $2.3 billion valuation to expand AI agents for ordering, marketing and back-office work at independent restaurants.</p></li><li><p><strong>Socure</strong> raised $156 million at a $5.2 billion valuation and acquired agentic-AI startup <strong>Fravity</strong>.</p></li><li><p><strong>Instinct</strong> raised a $250 million Series B at a $2.5 billion valuation, bringing the consumer AI company&#8217;s total funding to $350 million.</p></li><li><p><strong>Standard Metrics</strong> raised a $20 million Series B from customers to expand AI-assisted portfolio monitoring and fund administration.</p></li><li><p><strong>eComID</strong> raised a $17 million seed round to expand AI-assisted returns, resale and product-lifecycle infrastructure for retailers.</p></li><li><p><strong>Ringg</strong> raised a $10 million Series A to expand its voice AI platform into multimodal customer workflows.</p></li><li><p><strong>Neno</strong> raised a EUR 6.6 million seed round to build AI-native accounting, banking and finance workflows for European small businesses.</p></li><li><p><strong>Euler</strong> raised a $4.3 million seed round for its agentic partner-management platform after bootstrapping to profitability.</p></li><li><p><strong>Volve</strong> raised a $3 million seed round to expand AI systems that analyze construction bids, scopes and procurement documents.</p></li><li><p><strong>Itoflow</strong> raised a $2.5 million pre-seed round to bring AI-assisted analysis and workflows to investment managers.</p></li><li><p><strong>Descartes</strong> acquired <strong>Tai Software</strong> for approximately $100 million in cash, adding an AI-powered transportation-management platform for freight brokers.</p></li></ul><h3>Physical AI &amp; Robotics</h3><ul><li><p><strong>Generalist</strong> raised nearly $200 million in an extension that values the robotics-model company at $3 billion and brings its Series B to roughly $600 million.</p></li><li><p><strong>XPeng&#8217;s Dogotix</strong> raised more than $900 million at a post-money valuation above $6.3 billion for its embodied-intelligence and robotics business.</p></li><li><p><strong>Corvus Robotics</strong> raised $20 million to expand autonomous inventory drones that scan warehouses without fixed infrastructure.</p></li><li><p><strong>Mara</strong> raised $7 million to develop low-cost autonomous systems for detecting and defeating FPV drones.</p></li><li><p><strong>Motion</strong> raised a $2 million pre-seed round to build a humanoid robots-as-a-service platform for European industrial customers.</p></li></ul><h3>Science &amp; Healthcare</h3><ul><li><p><strong>Adaptyv Bio</strong> raised a $40 million Series A to expand automated wet-lab infrastructure for testing proteins designed by AI systems.</p></li><li><p><strong>Lupin Dental</strong> raised a EUR 15 million Series A for a supervised robotic system that prepares teeth for veneers under a dentist&#8217;s control.</p></li><li><p><strong>Legato</strong> emerged from stealth with $12 million to develop AI-powered hearing glasses combining directional audio, speech enhancement and computer vision.</p></li><li><p><strong>CropX</strong> acquired portable spectroscopy company <strong>SCIO</strong> to connect crop-quality measurements with its AI agronomy platform.</p></li><li><p><strong>Monod Bio</strong> granted <strong>SignalChem</strong> a non-exclusive license to AI-designed protein technologies for custom discovery assays.</p></li><li><p><strong>HiRO</strong> and <strong>Differentia Biotech</strong> signed an MOU to apply AI-powered simulation to clinical-trial design.</p></li><li><p><strong>PENTAX Medical</strong> expanded its AI-assisted colonoscopy partnership with <strong>MAGENTIQ EYE</strong> from the United States to EMEA, JAPAC and Latin America.</p></li></ul><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.trovereport.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Trove! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[The State of Sovereign Enterprise AI]]></title><description><![CDATA[Enterprise AI started by maximizing access to outside models. Its next phase is about deciding which intelligence is valuable enough to own.]]></description><link>https://www.trovereport.com/p/the-state-of-sovereign-enterprise</link><guid isPermaLink="false">https://www.trovereport.com/p/the-state-of-sovereign-enterprise</guid><dc:creator><![CDATA[Brian D'Erario]]></dc:creator><pubDate>Wed, 26 Aug 2026 13:43:07 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/bf0cb6e7-8f46-4d70-a12c-74ffef32fbc7_1800x1005.webp" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Over the past week, <a href="https://www.trovereport.com/p/harvey-is-becoming-a-model-company">Harvey</a> and <a href="https://www.thomsonreuters.com/en/press-releases/2026/august/thomson-reuters-leverages-its-world-class-data-assets-to-launch-its-own-frontier-model">Thomas Reuters</a> have announced they&#8217;ve built their own models to be used within their orgs and in Harvey&#8217;s case, deploying and training with customers. This strikes me as an inflection point where companies are shifting from aggressive token burning at any cost to utilizing routers and building their own models to control their data and costs in their AI stack. </p><p>Going back to April 2026 at Meta, an employee-built internal dashboard called Claudeonomics turned AI use into a competition. It ranked the company&#8217;s top 250 token users, handed out titles such as &#8220;Token Legend&#8221; and &#8220;Session Immortal,&#8221; and showed that more than 85,000 employees had consumed over 60 trillion tokens in 30 days. Meta also maintained a separate official usage dashboard for software engineers, who were among the company&#8217;s heaviest users. The leaderboard became popular enough that <a href="https://www.theinformation.com/briefings/meta-shutters-internal-ai-token-leaderboard">Meta took it down after the data leaked</a>.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.trovereport.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">If you&#8217;re interested in model creation and the future of data and AI, subscribe!</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>The thinking behind it was spreading as an effective way to catalyze employees to be using AI in their day-to-day as much as possible. Nvidia CEO Jensen Huang said that if a company pays an engineer $500,000 a year and that person isn&#8217;t burning $250,000 in tokens, something is wrong. Meta CTO Andrew Bosworth offered an even cleaner anecdote: <a href="https://www.forbes.com/sites/richardnieva/2026/03/31/the-ai-gods-spending-as-much-as-they-can-on-ai-tokens/">his best engineer was spending the equivalent of his salary on AI tokens</a> and, Bosworth claimed, producing five to ten times more work.</p><p>For a moment, token consumption became a proxy for ambition. Companies bought ChatGPT Enterprise and Claude Enterprise seats, opened access to coding agents, and encouraged employees to use as much intelligence as the model companies could sell them. The focuse was whether people were using AI and higher usage was assumed to improve the value each employee was delivering.</p><p>Then the bills arrived. Long-running agents consumed far more than chatbots. The same task could cost radically different amounts depending on the model and the harness around it. High usage didn&#8217;t automatically mean useful work. By the summer of 2026, <a href="https://apnews.com/article/31bb80ac1cd7862d05f6397177d826b1">the tokenmaxxing fad was giving way to budgets and cheaper models</a>.</p><p>That correction marks the start of enterprise AI&#8217;s mature phase. The goal is no longer usage for its own sake. Companies are now deploying sophisticated tokenomics into their spending models to decide how teams on a quarterly/yearly basis should be using AI. They&#8217;re asking which jobs still deserve an expensive frontier model, which requests can be routed to cheaper models, and which repeated workflows contain enough proprietary value to support a model of their own.</p><h2>From model access to stack ownership</h2><p>Enterprise AI began with GPT or Claude behind a chat interface or copilot. Agents moved the vendor decision from the model to the harness. Choosing Claude Code, Codex, Cursor, or Copilot means committing workflows, permissions, tools, and evaluations to the environment around the model. That choice can become more consequential than the difference between Claude and GPT on a benchmark.</p><p>Companies with enough usage are now pulling that layer inside their own built infrastructure. They own the session, policy, and production history, then route work among models beneath it. Shopify&#8217;s <a href="https://shopify.engineering/under-the-river">Aquifer platform</a> does this across Claude Code, Codex, Copilot, Cursor, and other tools, allowing Shopify to change a model or runtime without giving up its control plane.</p><p>That control also exposes where an owned model makes sense. Shopify <a href="https://shopify.engineering/leveraging-multimodal-llms">processes 40 million multimodal inferences a day</a> across its catalogue. When commercial APIs became prohibitively expensive, it fine-tuned smaller open models and deployed them on its own infrastructure. With Sidekick, Shopify turns unsuccessful merchant interactions into training data. It says the resulting GraphQL model surpassed its frontier baseline and reduced estimated annual serving costs from roughly $27 million to about $1 million, <a href="https://shopify.engineering/sidekicks-continual-learning-loop">a reduction of roughly 96%</a>.</p><h2>Companies are starting to build models of their own</h2><p>The companies below have publicly disclosed training, continued training, or material post-training of a generative or domain foundation model.</p><ul><li><p><strong>Shopify.</strong> Shopify fine-tunes open multimodal models for catalogue work and smaller language models for Sidekick. It serves them on infrastructure it controls and continually retrains them using failures from production.</p></li><li><p><strong>Walmart.</strong> <a href="https://corporate.walmart.com/content/corporate/en_us/news/2024/10/09/walmart-reveals-plan-for-scaling-artificial-intelligence-generative-ai-augmented-reality-and-immersive-commerce-experiences.html">Wallaby</a> is a family of retail-specific language models trained on decades of Walmart data. Walmart combines those models with outside LLMs in shopping, personalization, and customer-support systems.</p></li><li><p><strong>eBay.</strong> <a href="https://arxiv.org/abs/2406.12023">LiLiuM</a> is a family of 1 billion, 7 billion, and 13 billion parameter models developed entirely in-house. eBay trained them for marketplace tasks including product titles, descriptions, attribute extraction, and pricing.</p></li><li><p><strong>Intuit.</strong> <a href="https://investors.intuit.com/news-events/press-releases/detail/1210/intuit-accelerates-development-velocity-with-major-enhancements-to-proprietary-generative-ai-operating-system-genos">GenOS</a> gives Intuit teams access to commercial and open models alongside its own custom-trained models for tax, personal finance, accounting, and marketing work.</p></li><li><p><strong>Bloomberg.</strong> <a href="https://arxiv.org/abs/2303.17564">BloombergGPT</a> is a 50 billion parameter model trained on general text and Bloomberg&#8217;s financial corpus. The project showed how a company with a valuable information archive could turn that archive into model weights.</p></li><li><p><strong>Mastercard.</strong> <a href="https://www.mastercard.com/sea/en/news-and-trends/stories/2026/mastercard-new-generative-ai-model.html">Mastercard is training a tabular foundation model</a> on billions of anonymized transactions, with plans to expand into hundreds of billions. It expects the model to support fraud detection, loyalty, personalization, portfolio tools, and other payment workloads.</p></li><li><p><strong>Thomson Reuters.</strong> The company starts with open Qwen models and adds professional content, tools, and feedback from hundreds of subject-matter experts. Its <a href="https://www.thomsonreuters.com/en/press-releases/2026/august/thomson-reuters-leverages-its-world-class-data-assets-to-launch-its-own-frontier-model">Thomson model</a> will operate inside CoCounsel alongside models from outside providers.</p></li><li><p><strong>Harvey.</strong> Harvey began as a defining application of GPT-4, then added models from Anthropic and Google. It has now <a href="https://www.harvey.ai/blog/post-training-update-harvey-tenet">post-trained an open-weight Kimi model</a> inside long-running legal environments using public law, synthetic examples, expert work, and legal rubrics.</p></li><li><p><strong>Cursor.</strong> Cursor operates a router across outside and internal models, while training its own <a href="https://cursor.com/blog/composer">Composer coding models</a> inside the harness where they will be used. The product can keep difficult tasks on frontier systems and move repeatable coding work to models it controls.</p></li><li><p><strong>C3 AI.</strong> <a href="https://c3.ai/blog/why-c3-ai-fine-tunes-its-own-models-even-with-frontier-ai-this-strong">Narwhal</a> is a 27 billion parameter model trained further on C3 AI&#8217;s proprietary programming system. The company turns failed developer questions into evaluation rubrics and training examples.</p></li><li><p><strong>ServiceNow.</strong> ServiceNow&#8217;s <a href="https://downloads.docs.servicenow.com/resource/enus/infocard/sn-slm-v2.pdf">Apriel 13B</a> was continued from an open Mistral model and trained for enterprise workflows, tool use, and Now Assist applications. Its smaller size also makes restricted and private deployments more practical.</p></li><li><p><strong>Salesforce.</strong> <a href="https://www.salesforce.com/news/stories/agentforce-ai-models-announcement/">xGen-Sales and xLAM</a> are proprietary model families built for CRM work and actions. Salesforce uses synthetic-data pipelines to train models that can select tools and carry out work inside Agentforce.</p></li><li><p><strong>IBM.</strong> <a href="https://www.ibm.com/granite">Granite</a> is IBM&#8217;s open family of enterprise models for language, code, documents, speech, time series, and safety. IBM sells the models through watsonx while allowing customers to run and customize them elsewhere.</p></li><li><p><strong>Snowflake.</strong> <a href="https://www.snowflake.com/en/news/press-releases/snowflake-launches-arctic-the-most-open-enterprise-grade-large-language-model-2/">Arctic</a> was trained for enterprise work including SQL generation and instruction following. Snowflake released its weights and training information while serving it inside Cortex alongside outside models.</p></li><li><p><strong>Databricks.</strong> <a href="https://www.databricks.com/blog/introducing-dbrx-new-state-art-open-llm">DBRX</a> is Databricks&#8217; open mixture-of-experts model. It also serves as proof for the company&#8217;s larger pitch that enterprises can use their own data to train and govern custom models through Mosaic AI.</p></li><li><p><strong>SAP.</strong> <a href="https://help.sap.com/docs/sap-ai-core/generative-ai/sap-rpt-1?locale=en-US">RPT-1</a> is a relational foundation model built for tables and connected business data. It performs classification and regression from examples without requiring a separate model to be trained for each prediction task.</p></li><li><p><strong>Adobe.</strong> <a href="https://news.adobe.com/news/2025/04/adobe-revolutionizes-ai-assisted-creativity-firefly">Firefly</a> is a family of Adobe-trained models for images, video, audio, vectors, and design. Adobe also lets companies customize Firefly models on approved brand assets for production work.</p></li><li><p><strong>Autodesk.</strong> <a href="https://www.research.autodesk.com/projects/project-bernini/">Project Bernini</a> is an experimental model trained on ten million 3D shapes. It generates functional geometry from text, images, sketches, voxels, and point clouds.</p></li><li><p><strong>Cisco.</strong> <a href="https://blogs.cisco.com/security/foundation-sec-cisco-foundation-ai-first-open-source-security-model">Foundation-sec-8B</a> starts with Llama and adds continued training on a cybersecurity corpus built inside Cisco. It is designed for security workflows that can&#8217;t send sensitive material to a hosted general-purpose API.</p></li><li><p><strong>Siemens.</strong> Its <a href="https://blogs.sw.siemens.com/nx-manufacturing/teaching-ai-to-speak-the-language-of-engineering-and-manufacturing-through-industrial-foundation-model/">Industrial Foundation Model</a> is being developed for engineering drawings, 3D models, technical specifications, and manufacturing data that general LLMs rarely understand.</p></li><li><p><strong>Recursion.</strong> Recursion trained <a href="https://ir.recursion.com/static-files/aefba7ce-e070-404d-8cfe-605609c9a586">Phenom-1</a>, a vision foundation model built on billions of cellular images from its proprietary phenomics library. It is extending the same approach into transcriptomics, chemistry, and patient biology.</p></li><li><p><strong>Insilico Medicine.</strong> <a href="https://insilico.com/news/n4y7nk0f31-insilico-medicine-launches-nach01-founda">Nach01</a> is a foundation model for chemical structures and molecular prediction. Insilico sells access to the model while using related systems throughout its own drug-discovery platform.</p></li><li><p><strong>Apple.</strong> <a href="https://machinelearning.apple.com/research/introducing-apple-foundation-models">Apple trains its own on-device and server foundation models</a> for Apple Intelligence. Small models run locally, while larger workloads can move to Apple&#8217;s Private Cloud Compute infrastructure.</p></li><li><p><strong>Samsung.</strong> <a href="https://news.samsung.com/global/samsung-electronics-hosts-samsung-developer-conference-korea-2024-unveils-its-improved-gen-ai-model">Gauss2</a> is Samsung&#8217;s multimodal language, code, and image model. It already powers internal coding and productivity tools and is being adapted for Samsung devices.</p></li><li><p><strong>LG.</strong> LG runs a smaller model derived from <a href="https://www.lg.com/global/newsroom/news/media-entertainment-solution/lgs-hybrid-ai-gram-laptops-offer-the-best-of-both-worlds-with-on-device-and-cloud-ai-services/">EXAONE</a> locally on its laptops, giving users private search, summarization, and device assistance without requiring a cloud connection.</p></li><li><p><strong>Naver.</strong> <a href="https://www.navercorp.com/en/tech/hyperclovax">HyperCLOVA X</a> is a from-scratch model family built around Korean language and business needs. Naver uses it in its own consumer services and sells access and dedicated infrastructure to other enterprises.</p></li><li><p><strong>Rakuten.</strong> <a href="https://global.rakuten.com/corp/news/press/2024/0321_01.html?category=corp&amp;month=3&amp;year=2024">Rakuten AI 7B</a> was continually trained from Mistral on Japanese and English data using Rakuten&#8217;s own GPU cluster. Rakuten uses its model family across its businesses and has created versions that can run locally on PCs.</p></li></ul><h2>Sovereignty without isolation</h2><p>None of this means every enterprise needs its own LLM. The economics only work when a task happens often, the company has data or expertise that improves it, and success can be measured reliably. Plenty of work will remain on frontier APIs because the best general model is difficult to reproduce and cheaper to rent than build. Training a model requires a team and dedicated resources to put together.</p><p>Sovereignty also doesn&#8217;t require a company to cut itself off from OpenAI or Anthropic. Shopify, Thomson Reuters, Harvey, Cursor, Walmart, and Intuit all use outside models while building models of their own. The frontier model can handle unfamiliar work, teach a smaller model, or judge its output. A specialized model can absorb high-volume work where cost and company context matter more.</p><p>Tokenmaxxing was about proving that a company could consume AI. The next phase is deciding what methods yield the highest ROI per token burned.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.trovereport.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">If you&#8217;re interested in model creation and the future of data and AI, subscribe!</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Harvey, from App to Frontier AI Lab]]></title><description><![CDATA[It started as an application on top of frontier models. Its next advantage may come from models trained for the profession, then for a single firm.]]></description><link>https://www.trovereport.com/p/harvey-is-becoming-a-model-company</link><guid isPermaLink="false">https://www.trovereport.com/p/harvey-is-becoming-a-model-company</guid><dc:creator><![CDATA[Brian D'Erario]]></dc:creator><pubDate>Mon, 24 Aug 2026 11:03:43 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/89a7a1d7-cdce-4ccf-9b68-2c0b4b512f5f_800x450.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Last week, legal AI company Harvey unveiled Tenet, a model it created with Fireworks by post-training Moonshot AI&#8217;s open-weight Kimi K3.</p><p>Harvey began as an application built on top of frontier models. In 2022, integrating those models into its customers&#8217; legal workflows was a meaningful edge. Harvey wrapped GPT-3 and later models in tools for research, drafting, document review, and firm controls, making lawyers more productive by automating work they and their teams previously handled manually.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.trovereport.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Follow the Data</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>That advantage gets thinner as the underlying models improve. Any legal AI startup can call the same OpenAI or Anthropic API. Every release, models are getting better at long documents, tool use, and complex reasoning, absorbing more of the work that once belonged to the application. The industry is quickly realizing model + workflows is no longer a moat.</p><p>Netflix went through similar struggles in a different industry. Its streaming service originally depended on content licensed from studios and when those suppliers began launching competing services, Netflix moved into original programming. It gave them a moat its partners couldn&#8217;t take away or sell to every competitor, original content found only on Netflix that may not have been studio quality, but it didn&#8217;t need to be. As they scaled, they were able to invest more into higher quality original content that widened the gap from the studios.</p><p>Harvey&#8217;s version of original programming is a specialist model. It can take a near-frontier open-weight model and apply reinforcement learning using the legal tasks, feedback, and evaluations it has developed inside customer workflows. Kimi K3 supplies the general intelligence, but Harvey teaches it what good legal work looks like. While the model may be specialist and Harvey is likely still using frontier models in some work, as Harvey scales and gathers more resources for research, there&#8217;s a world where all the work they&#8217;re doing is cycling through exclusively Harvey trained models.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!q8Y7!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55aa77b9-fac7-4e52-8dd6-0e0ff3ee2580_1123x472.svg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!q8Y7!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55aa77b9-fac7-4e52-8dd6-0e0ff3ee2580_1123x472.svg 424w, https://substackcdn.com/image/fetch/$s_!q8Y7!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55aa77b9-fac7-4e52-8dd6-0e0ff3ee2580_1123x472.svg 848w, https://substackcdn.com/image/fetch/$s_!q8Y7!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55aa77b9-fac7-4e52-8dd6-0e0ff3ee2580_1123x472.svg 1272w, https://substackcdn.com/image/fetch/$s_!q8Y7!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55aa77b9-fac7-4e52-8dd6-0e0ff3ee2580_1123x472.svg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!q8Y7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55aa77b9-fac7-4e52-8dd6-0e0ff3ee2580_1123x472.svg" width="1456" height="612" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/55aa77b9-fac7-4e52-8dd6-0e0ff3ee2580_1123x472.svg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:612,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:405992,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/svg+xml&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.trovereport.com/i/212342766?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55aa77b9-fac7-4e52-8dd6-0e0ff3ee2580_1123x472.svg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!q8Y7!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55aa77b9-fac7-4e52-8dd6-0e0ff3ee2580_1123x472.svg 424w, https://substackcdn.com/image/fetch/$s_!q8Y7!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55aa77b9-fac7-4e52-8dd6-0e0ff3ee2580_1123x472.svg 848w, https://substackcdn.com/image/fetch/$s_!q8Y7!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55aa77b9-fac7-4e52-8dd6-0e0ff3ee2580_1123x472.svg 1272w, https://substackcdn.com/image/fetch/$s_!q8Y7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55aa77b9-fac7-4e52-8dd6-0e0ff3ee2580_1123x472.svg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Harvey went from app to frontier (legal) model lab</figcaption></figure></div><p>Tenet wasn&#8217;t trained on customer data, but Harvey&#8217;s broader plan includes specialized models that firms can build and own. Customers can opt into a bespoke model trained on their data and kept exclusive to them. Harvey has already built one for Cuatrecasas that searches the firm&#8217;s internal knowledge and produces work in its style. That turns a firm&#8217;s documents, standards, and expertise into model performance that a generalized LLM, or another firm, can&#8217;t access.</p><p>Harvey won&#8217;t be the last app-layer company to make this move. As frontier models absorb more of the capabilities that once had to be built into workflows, wrapping those models in an interface will become less defensible. Companies will need to own models refined around the specific tasks, data, and standards of their customers. Thanks to a healthy open-source ecosystem, they&#8217;ll start with near-frontier models and make them better at completing real, end-to-end, work.</p><p></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Wrnv!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61d9071a-d587-40bb-bbbe-9fe187876892_1200x445.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Wrnv!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61d9071a-d587-40bb-bbbe-9fe187876892_1200x445.png 424w, https://substackcdn.com/image/fetch/$s_!Wrnv!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61d9071a-d587-40bb-bbbe-9fe187876892_1200x445.png 848w, https://substackcdn.com/image/fetch/$s_!Wrnv!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61d9071a-d587-40bb-bbbe-9fe187876892_1200x445.png 1272w, https://substackcdn.com/image/fetch/$s_!Wrnv!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61d9071a-d587-40bb-bbbe-9fe187876892_1200x445.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Wrnv!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61d9071a-d587-40bb-bbbe-9fe187876892_1200x445.png" width="1200" height="445" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/61d9071a-d587-40bb-bbbe-9fe187876892_1200x445.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:445,&quot;width&quot;:1200,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:59515,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.trovereport.com/i/212342766?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61d9071a-d587-40bb-bbbe-9fe187876892_1200x445.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Wrnv!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61d9071a-d587-40bb-bbbe-9fe187876892_1200x445.png 424w, https://substackcdn.com/image/fetch/$s_!Wrnv!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61d9071a-d587-40bb-bbbe-9fe187876892_1200x445.png 848w, https://substackcdn.com/image/fetch/$s_!Wrnv!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61d9071a-d587-40bb-bbbe-9fe187876892_1200x445.png 1272w, https://substackcdn.com/image/fetch/$s_!Wrnv!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61d9071a-d587-40bb-bbbe-9fe187876892_1200x445.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Partnerships and deals</h2><ul><li><p><strong>Nvidia</strong> agreed to pay $6B to license <strong>Poolside</strong>&#8217;s Laguna models and hire more than 100 employees. Separately, it opened preliminary talks with Korean AI-chip startup <strong>Rebellions</strong> about a partnership, investment, or acquisition.</p></li><li><p><strong>Ode</strong> with <strong>Anthropic</strong> acquired Claude implementation firm <strong>Casper Studios</strong>. Terms weren&#8217;t disclosed.</p></li><li><p><strong>DoiT</strong> acquired AI cost-management startup <strong>Attribute</strong> for an estimated $65M. The companies didn&#8217;t confirm the price.</p></li><li><p><strong>Serve Robotics</strong> partnered with <strong>Wonder</strong> and <strong>Grubhub</strong> to launch robot delivery in Chicago, Los Angeles, and Alexandria, Virginia. Its Moxi 2.0 rollout adds a new robotic foundation model and 15x faster perception.</p></li><li><p><strong>Higgsfield</strong> raised a $400M Series B at a $5.4B valuation led by <strong>DST Global</strong>.</p></li><li><p><strong>Wispr</strong> raised a $280M Series B at a $2B valuation led by <strong>Menlo Ventures</strong> and launched a new speech model.</p></li><li><p><strong>Gravis Robotics</strong> raised a $200M Series A from <strong>SoftBank</strong> to scale autonomous earthmoving across mixed fleets and real construction sites.</p></li><li><p><strong>Veeda AI</strong> raised more than $90M in seed funding from <strong>Khosla Ventures</strong> and <strong>Radical Ventures</strong> to build simulated worlds where embodied agents can generate training experience.</p></li><li><p><strong>Rillet</strong> raised a $100M Series C at a $1B valuation led by <strong>ICONIQ</strong> to expand its AI-native accounting platform.</p></li><li><p><strong>Palona AI</strong> closed a Series A that brought total funding to $20M and unveiled a multimodal AI operating layer for restaurants and other physical businesses.</p></li><li><p><strong>Etched</strong> raised $700M at a $21B valuation led by <strong>Jane Street</strong> and shipped its first customer rack.</p></li><li><p><strong>Orbbec</strong> and <strong>Linkerbot</strong> partnered to combine 3D vision with dexterous robotic hands for embodied AI.</p></li><li><p><strong>LG Electronics</strong> and <strong>Nvidia</strong> are targeting 100,000 hours of robot-training data by year-end across home, manufacturing, logistics, and robotic-hand environments.</p></li><li><p><strong>Canary Data</strong> and <strong>Perplexity</strong> partnered to bring proprietary investment-research datasets into Perplexity workflows through MCP.</p></li><li><p><strong>Orbbec</strong> introduced four physical-AI data-collection systems and reported more than 37 hours of continuous collection with zero dropped RGB or IMU frames.</p></li><li><p><strong>Cursor</strong> launched Origin, an early-beta code-hosting service with repositories, pull requests, code browsing, and GitHub sync.</p></li><li><p><strong>Anything AI</strong> introduced a roster of more than 150 AI agents for engineering, operations, design, and administrative work, extending beyond its app-building platform.</p></li><li><p><strong>OpenAI</strong> previewed Private Safety Processing to preserve Zero Data Retention across related interactions. <strong>Anthropic</strong>, which still requires 30-day retention for its most capable models, reportedly plans to let customers keep that data in their own cloud.</p></li><li><p><strong>Generalist AI</strong> released GEN-1.5, which learns a new task from a 3-to-12-second demonstration without weight updates. It averaged 59% success across 10 tasks and reached 83% after 10 gradient steps using five minutes of data.</p></li><li><p><strong>Foxglove</strong> launched an agentic data platform for physical AI that lets teams query robot data in natural language, compare recordings, and search images and video semantically.</p></li><li><p><strong>Robo Robotics</strong> launched a standardized teachable-automation platform connecting demonstrations, data collection, model training, and fleet deployment. Policies taught on one unit can be deployed across identical hardware.</p></li></ul><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.trovereport.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Follow the Data</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Data: The Next Asset Class]]></title><description><![CDATA[Why Google&#8217;s $10 million bid for Spirit Airlines&#8217; data matters to every company]]></description><link>https://www.trovereport.com/p/data-the-next-asset-class</link><guid isPermaLink="false">https://www.trovereport.com/p/data-the-next-asset-class</guid><dc:creator><![CDATA[Brian D'Erario]]></dc:creator><pubDate>Thu, 20 Aug 2026 22:25:02 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/1d5889aa-1916-4d50-ac7e-c11f85d40310_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Google bid $10 million for Spirit Airlines&#8217; internal business data, including roughly 100 million emails, 500 million Teams messages, 30 million lines of code and operational records, according to the <a href="https://news.bloomberglaw.com/new-york-brief/google-aims-to-boost-ai-with-purchase-of-spirit-airlines-data">bankruptcy-court notice</a>.</p><p>That number stopped me. If Spirit&#8217;s leftover emails and Teams messages are worth eight figures, every company sitting on years of its own work history should be paying attention.</p><p>Subscribe to get Trove&#8217;s reporting on the AI data market every Monday.</p><h2>From Books to the Internet</h2><p>In eight years, model training has moved from books, to the public internet, to human feedback and proprietary data. GPT-1 trained largely on books. GPT-2 shifted toward WebText, a collection of webpages discovered through links shared on Reddit. The bet was simple: instead of training AI for a single task like chess or Go, train one model on a broad sample of human language. GPT-3 scaled that approach with more webpages, books and Wikipedia. OpenAI details the WebText construction in its <a href="https://cdn.openai.com/better-language-models/language_models_are_unsupervised_multitask_learners.pdf">GPT-2 report</a>.</p><p>GPT-3 became very good at continuing text, but it struggled to answer questions, follow complex instructions or admit uncertainty. Some of those problems remain. What came next was the human-feedback race. InstructGPT, the predecessor to ChatGPT&#8217;s GPT-3.5, added human demonstrations, output rankings and user prompts through reinforcement learning from human feedback. Later generations of GPT, Claude, Llama and Gemini incorporated first-party feedback and licensed third-party data into their training and post-training. The method is detailed in the <a href="https://arxiv.org/abs/2203.02155">InstructGPT paper</a>.</p><h2>Digital Data as a Commodity</h2><p>Human feedback showed that better-targeted data could improve a model without simply making it larger. At the same time, the public internet was becoming less differentiated. Every major model company could crawl similar websites, books and code repositories.</p><p>Model companies began racing to assemble datasets their competitors did not have. Shutterstock, Reddit and major publishers began licensing their archives for model training. The prices are public enough now to sketch the market. News Corp reportedly sold OpenAI access to its archives for more than $250 million over five years. Reddit&#8217;s licensing deal with Google was reported at around $60 million a year. The AP, Axel Springer, the Financial Times, Time, all of them have cut deals in the tens of millions. Add up the announced contracts and model builders have already spent something north of a billion dollars on content they used to scrape for free.</p><p>But most of these licenses were nonexclusive, which means everyone&#8217;s model got the same books. You can see that sameness in the AI-isms developed by every model (looking at you, em dash). Shared data can&#8217;t be an edge. Which raises the question: what do you buy when everyone has already bought the internet?</p><h2>The Race for Proprietary Data</h2><p>As access to capable base models becomes more common, competition is shifting toward the data used to specialize them.</p><p>This is where Spirit comes in. Google bid $10 million for Spirit&#8217;s Teams messages, code and operational records, beating a $7.5 million backup bid from Mercor. As agents enter specific business functions, data showing how work actually gets done could become more valuable than isolated code, manuals or messages. The figures come from the <a href="https://www.axios.com/2026/08/17/google-spirit-airlines-bankruptcy">auction result</a>.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!5hIY!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65758103-64fe-4b14-ad2c-ff5b96082745_1360x936.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!5hIY!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65758103-64fe-4b14-ad2c-ff5b96082745_1360x936.png 424w, https://substackcdn.com/image/fetch/$s_!5hIY!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65758103-64fe-4b14-ad2c-ff5b96082745_1360x936.png 848w, https://substackcdn.com/image/fetch/$s_!5hIY!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65758103-64fe-4b14-ad2c-ff5b96082745_1360x936.png 1272w, https://substackcdn.com/image/fetch/$s_!5hIY!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65758103-64fe-4b14-ad2c-ff5b96082745_1360x936.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!5hIY!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65758103-64fe-4b14-ad2c-ff5b96082745_1360x936.png" width="1360" height="936" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/65758103-64fe-4b14-ad2c-ff5b96082745_1360x936.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:936,&quot;width&quot;:1360,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:94637,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:&quot;&quot;,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://trovemedia.substack.com/i/211995147?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65758103-64fe-4b14-ad2c-ff5b96082745_1360x936.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!5hIY!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65758103-64fe-4b14-ad2c-ff5b96082745_1360x936.png 424w, https://substackcdn.com/image/fetch/$s_!5hIY!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65758103-64fe-4b14-ad2c-ff5b96082745_1360x936.png 848w, https://substackcdn.com/image/fetch/$s_!5hIY!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65758103-64fe-4b14-ad2c-ff5b96082745_1360x936.png 1272w, https://substackcdn.com/image/fetch/$s_!5hIY!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65758103-64fe-4b14-ad2c-ff5b96082745_1360x936.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The value lies in connecting instructions, decisions and outcomes into a usable training set.</p><p>The race extends beyond the frontier labs. Harvey says it post-trained Tenet from Kimi K3 with Fireworks for long-horizon legal work. The company is using that effort to develop proprietary model intelligence for legal work, as described in its <a href="https://www.harvey.ai/blog/post-training-update-harvey-tenet">Tenet update</a>.</p><p>If Tenet outperforms general models on legal work, Harvey owns an advantage that competitors cannot copy simply by calling the same model API.</p><h2>From Digital Data to Physical Data</h2><p>Digital agents can learn from emails, Slack messages and operational documents. Robots need an even scarcer asset: physical-world data. If we believe in a future with abundant robots, those robots will need enormous amounts of data showing how physical work is performed.</p><p>Atoms, Travis Kalanick&#8217;s new company, is assembling operating businesses around physical automation. Its portfolio includes Pronto in mining and Lab37 in food robotics. Pronto says its software learns a haul route after a human operator drives it once. As Atoms expands into more industries, access to those environments and the data they produce could become one of its biggest advantages, according to <a href="https://pronto.ai/solution/">Pronto&#8217;s route-learning description</a>.</p><p>A supplier market is already forming around that demand. MicroAGI&#8217;s Shift records physical work and turns it into training-ready data. Its New York offer exchanges free home cleaning for recorded work, as reported by <a href="https://arstechnica.com/ai/2026/05/robot-training-startup-will-send-humans-wearing-cameras-to-clean-your-home/">Ars Technica</a>.</p><p>If you run a business, your Slack archive might be your most underpriced asset. Worth thinking about before a bankruptcy auction decides its price for you.</p><h2>Why I started Trove</h2><p>Eight years ago, data was the thing companies kept in backups nobody priced. Now it&#8217;s turning up in bankruptcy auctions with eight-figure bids.</p><p>I started Trove because I kept noticing the gap between what companies think their data is worth and what someone will actually pay for it. Spirit didn&#8217;t know it was sitting on a $10 million asset until the auction notice went up. Most companies still don&#8217;t.</p><p>So that&#8217;s the beat: every Monday, one piece of this market &#8212; a deal broken down, a company profiled, or someone who sells or buys data explaining how it works. If you own data, need data, or just like watching a new asset class get built in public, stick around.</p><p>See you Monday.</p><p>&#8212; Brian</p><p>Subscribe to get Trove&#8217;s reporting on the AI data market every Monday.</p>]]></content:encoded></item></channel></rss>