Perspective

Who Owns the Data That Trains the Models

By Kyle Harrison

Updated

May 11, 2024

Reading Time

3 min

The cycle of breakneck AI news is just spinning faster and faster. Everyone just close your eyes, hold your breath, and dive in. Because this week is no exception. As usual, OpenAI is at the center of a lot of the attention. From a ChatGPT feature that can search the web and cite sources, or maybe a Google search competitor, or maybe… none of those things? It’s unclear.

But one thing that was clear is that OpenAI’s relationship with data and copyright really spilled out into the open this week. First up, Stack Overflow and OpenAI announced a partnership for OpenAI to get access to the company’s data. We’ve seen similar partnerships between the likes of Reddit and Google after companies with data pushed to get compensated for their platforms being used as training data for LLMs.

Stack Overflows users didn’t react positively to the news, with many of them deleting their contributions to the site en masse. And Stack Overflow responded by banning those users. As Gergely Orosz points out:

“The reality is that now StackOverflow and Reddit are both sites where they are free to use to provide answers because you are providing training data for AI models with every one of your answers (and past ones). The only way to opt out is to stop contributing to them.”

Stack Overflow wasn’t the only data set in OpenAI’s crosshairs. In a class action lawsuit brought against OpenAI by the Author’s Guild, it was alleged that OpenAI had used 100K+ copyrighted books in training its models, and then deleted the datasets. That’s just one of the many lawsuits from artists, writers, and publishers that are pushing back on their work being used in training data.

OpenAI is attempting to respond by announcing a tool called “Media Manager,” which would “allow ‘creators and content owners to tell [OpenAI] what they own’ and specify ‘how they want their works to be included or excluded from machine learning research and training.’” But details are sparse, and the execution seems complicated. Even companies like Reddit, who have made access deals, are attempting to introduce public content policies to better regulate how the data on its platform is being used.

Ironically, OpenAI is the pot calling the kettle black on the other side of copyright law. The company recently made a copyright complaint against the ChatGPT subreddit for using the OpenAI logo.

The unfortunate reality is that the same copyright battle will continue to play out across writing, art, music, and more, because what we thought was the unstoppable force of the open internet is now running into an immovable object of the exorbitant demand for data that AI companies have. And that demand is only going up as we see ever more competition in the space. Even Microsoft, the sugar daddy of OpenAI, is working on its own in-house LLM called MAI-1, which is explicitly meant to compete with OpenAI.

Important Disclosures

This material has been distributed solely for informational and educational purposes only and is not a solicitation or an offer to buy any security or to participate in any trading strategy. All material presented is compiled from sources believed to be reliable, but accuracy, adequacy, or completeness cannot be guaranteed, and Contrary LLC (Contrary LLC, together with its affiliates, “Contrary”) makes no representation as to its accuracy, adequacy, or completeness.

The information herein is based on Contrary beliefs, as well as certain assumptions regarding future events based on information available to Contrary on a formal and informal basis as of the date of this publication. The material may include projections or other forward-looking statements regarding future events, targets or expectations. Past performance of a company is no guarantee of future results. There is no guarantee that any opinions, forecasts, projections, risk assumptions, or commentary discussed herein will be realized. Actual experience may not reflect all of these opinions, forecasts, projections, risk assumptions, or commentary.

Contrary shall have no responsibility for: (i) determining that any opinions, forecasts, projections, risk assumptions, or commentary discussed herein is suitable for any particular reader; (ii) monitoring whether any opinions, forecasts, projections, risk assumptions, or commentary discussed herein continues to be suitable for any reader; or (iii) tailoring any opinions, forecasts, projections, risk assumptions, or commentary discussed herein to any particular reader’s objectives, guidelines, or restrictions. Receipt of this material does not, by itself, imply that Contrary has an advisory agreement, oral or otherwise, with any reader.

Contrary is registered with the Securities and Exchange Commission as an investment adviser under the Investment Advisers Act of 1940. The registration of Contrary in no way implies a certain level of skill or expertise or that the SEC has endorsed Contrary. Investment decisions for Contrary clients are made by Contrary. Please note that, although Contrary manages assets on behalf of Contrary clients, Contrary clients may take any position (whether positive or negative) with respect to the company described in this material. The information provided in this material does not represent any investment strategy that Contrary manages on behalf of, or recommends to, its clients.

Different types of investments involve varying degrees of risk, and there can be no assurance that the future performance of any specific investment, investment strategy, company or product made reference to directly or indirectly in this material, will be profitable, equal any corresponding indicated performance level(s), or be suitable for your portfolio. Due to rapidly changing market conditions and the complexity of investment decisions, supplemental information and other sources may be required to make informed investment decisions based on your individual investment objectives and suitability specifications. All expressions of opinions are subject to change without notice. Investors should seek financial advice regarding the appropriateness of investing in any security of the company discussed in this presentation.

Please see www.contrary.com/legal for additional important information.

Authors

Kyle Harrison

General Partner @ Contrary

Kyle leads Contrary’s investing efforts for companies from seed to scale. He’s previously worked at firms like Index and Coatue investing in companies like Databricks, Snowflake, Snyk, Plaid, Toast, and Persona.

See articles

© 2026 Contrary Research · All rights reserved

Privacy Policy

By navigating this website you agree to our privacy policy.