← 목록으로

Cloudflare’s legal chief on AI, copyright and the future of publisher monetisation

Cloudflare’s legal chief on AI, copyright and the future of publisher monetisation

요약

AI is changing how content is discovered, and who profits from it. Publishers are already feeling the impact, but the bigger shift may still be ahead. The post Cloudflare’s legal chief on AI, copyright and the future of publisher monetisation appeared first o…

본문

Douglas Kramer is the Chief Legal Officer at Cloudflare, where he works at the intersection of internet infrastructure, cybersecurity, and global regulation. In this segment of the conversation, MediaNama Editor Nikhil Pahwa speaks with Kramer about how artificial intelligence is reshaping copyright, the economics of the web, and the emerging tensions between AI companies and publishers.

The discussion focuses on the rise of AI crawlers, the collapse of traditional traffic-based monetisation models, the challenges posed by retrieval-based AI systems, and Cloudflare’s attempt to build a marketplace for content usage. It also explores whether publishers can meaningfully distinguish between different types of bots, and the regulatory debates emerging around this question.

This last section of the interview has been lightly edited for clarity and readability.

Read Part 1 of the interview: [link]
Read Part 2 of the interview: [link]
Read Part 3 of the interview: [link]
Read Part 4 of the interview: [link]
Read Part 5 of the interview: [link]

AI Crawlers and the Collapse of the Traffic–Revenue Model

Nikhil Pahwa: I’m going to talk about copyright. AI is pushing copyright into a new phase. It’s no longer just 500 copies, but systematic extraction of value from the open web with training crawlers.

Cloudflare took a lead in addressing bot traffic, amongst the first to recognise what this means to publishers. And first, thank you for taking on that role as an organisation. I think as a publisher it’s extremely important for us to protect our copyright.

And you were the first to block bot traffic to publishers, and now you have a marketplace for monetisation of copyrighted content. From Cloudflare’s network-level vantage point, what has changed in the last year? And what is driving the biggest stress for publishers in terms of copyright?

Is it user-triggered retrieval bots like we discussed? Is it hybrid search and training bots? What are you seeing?

Douglas Kramer: So this has been one of the more fascinating stories I’ve seen in my tenure at Cloudflare for almost 10 years, because, as we were discussing before, we’ve had to work very closely and in a cooperative way, but maybe not fully satisfying rights holders over the years, with their concerns about copyright.

Because, as I said before, by offering free services, cybersecurity services, which we think is a public good, we can’t necessarily be in a position where we now have to adjudicate all these copyright infringement challenges.

But in those conversations, and this is always what we do, even if we may have some disagreement or not be able to satisfy all of their requests, we sort of understood where they were coming from, understanding what their concerns were.

And what we heard starting about two years ago is that advertising revenue for publishers was just falling off a cliff. The number of visitors they were getting to their websites that they could monetise through showing ads, things like that, was just going away.

And as we started to dig into it, we got a little bit of a sense of what had been happening with AI. And then there was a real watershed moment when Google went from merely providing links to websites to giving you that little AI summary, that little description, so people, in a lot of cases, didn’t have to go through and click that link.

And so we started to run the numbers, because this is what we can do, this is, again, our bread and butter. Looking online, we can see the bots that come through.

And for those that don’t realise, I mean, the way that Google funds or organises its search, or the way that AI companies train their LLMs, their large language models, is they have these crawlers that go around and check everything online and then sort of synthesise all of that.

And it used to be the case that Google would scrape a website maybe somewhere close to one-to-one for the number of eyeballs that it sent. So it would scrape your website, and then in return for that, the exchange was Google got to scrape your website, learn from it, but then they would send you traffic.

When Google started to introduce their AI search summaries, that went down to something like six to one, that they would crawl six times for every one set of eyeballs that they sent you. DuckDuckGo is still somewhat in that category as a search engine.

The AI companies were never built to do that. So when we started to dig in on the large AI companies and the scrape-to-eyeball ratio, we saw that they were crawling websites thousands of times for every user that they would send to that website organically to look at.

And I think that’s been all of our experience, we don’t really know where AI gets the answer. The large language models, we just sort of trust it and don’t click through anymore.

The problem is that it has been the monetisation model for the Internet. The publishers create new content, good content, because they then get eyeballs that they can sell ads to. But if they’re not getting those click-throughs anymore, then that dries up immediately.

So the question has been: what is the new monetisation model? Because there’s clearly still a transfer of value, that the training bots, the search bots, are clearly getting value off those websites, and there needs to be some way to return that.

So we’ve worked with publishers and worked with other website operators to set up systems to identify bots, potentially block bots, and potentially then set up systems by which there can be at least enough engagement, a marketplace, something, to lead to an exchange of value.

It’s by no means mature. We’re just still in the very early stages of rolling this out, and it’s not clear if that model is going to look more like Google Ads, is it going to look like Spotify, is it going to look like a licensing agreement. So that I turn on access to those bots for people that I have my own discrete contract with. It’ll probably be some combination of all of that. But in the absence of doing really sophisticated things in understanding bot traffic and being able to categorize it and block it selectively, that’s just the starting point for this. And that’s what we’ve been working with publishers to establish.

RAG Models and Control Over Content Access

Nikhil Pahwa: So that’s on the training side. But on the other side, you also have RAG models that go out and retrieve information live. So Perplexity AI’s argument is: we are user-driven, we’re like a browser fetching content for the users, so robots.txt shouldn’t stop us. Someone needs an answer, we’re going and getting it right here, right now, as if the human being is doing this. So we’re not training anything on it, we’re just taking the information because someone wants an answer.

Just how do you see, and you had a run-in with them in terms of bypassing your restrictions, tell us a little bit about your take.

Douglas Kramer: So there are two things on that. One is, I don’t think these are necessarily our restrictions. We always want these to be in the hands of website operators. Now, we’ve changed the default, light switches on or off when you walk into the room, but website operators have the ability to say, in most cases, I want the default to be this way and then change it.

And so that’s one of the two essential things here. So even for a RAG model, even if you’re not training, even if you’re just going out and getting a discrete answer, I still think the ability to decide whether or not they want to answer that question should be with the website operator.

Do I want to give that answer under these terms? And just because they’ve been giving answers for so long, usually in exchange for the eyeballs of those getting the answer, do they want to continue to do that even for discrete requests? That’s something that the information owner should be able to decide for themselves. And so we give them the power to do that

The second piece of this is we need to make sure that the people who are operating on the other side are being genuine about their purpose for crawling and their use for crawling. And I think we have certainly seen some AI companies that aren’t always as clear about that as they should be.

So we think having standardised ways to communicate your purpose in crawling a website for any purpose, I think those need to be mutually agreed to. Both sides have to follow them.

But ultimately, at the end of the day, it’s going to be the control of the website operator who decides for what purposes they’re comfortable being crawled, for what purposes they’re comfortable providing access.

Building a Marketplace for AI–Publisher Value Exchange

Nikhil Pahwa: So tell me how the marketplace model works. What is it? How does a publisher monetise this content? Who sets the rate for the content? How do you prevent, I mean, the challenge with training is that it’s a one-time activity. Once it’s done, I was talking to someone about this and he said that it’s like, with wheat flour, asking where the grains came from, it’s all mixed together and you get an output, right?

So how are publishers thinking about pricing? Who sets the pricing in the marketplace?

Douglas Kramer: Yeah, so this is something that we are still working on. I think what we’ve done so far is realising there’s no basis for a marketplace until you establish a couple of things around just a common set of terms, a common way of operating.

Unless you have, like, the currency in which it gets paid, and I’m not talking about real currency, but what exactly is being sold, unless there are agreed terms on that to set a market. And I think that’s what we’ve done.

We continue to examine all the different options for what might work here, along with, we’re talking to competition authorities to make sure that we’re doing this in a way that keeps an open marketplace. We’re talking to AI companies. We’re talking to everybody about how they want to transact in this space.

And we’re hearing all different things from those parties. Some of them have already engaged in discrete contracts. If you’re a larger content creator, you might have the ability to go negotiate on your own. And then they have the ability to come back to a provider like Cloudflare and turn the dials on their website to allow in those that they’ve operated with.

That is probably the most effective way of monetisation, but it’s open to very few. The people who can do discrete deals as opposed to other models are going to run out pretty quickly.

So then we look at what we see, some people looking at the model that has worked for the Internet up till now, which was sort of a Google ad auction. When people were searching for something or going to a website, people would pay some percentage of a cent or rupee or whatever to get that ad placed where they wanted it, that there’d be that sort of ad auction or revenue base for those sorts of things, where it’s more of a transactional thing.

Now, because it won’t be eyeballs and ads, it’s maybe a negotiation that happens at the scraping, the authorisation to scrape a site or not.

And then another one is just whether or not you license, opt into being part of a broad licence, something like a Spotify model, that I’ve got a bunch of songs over here, I’m not going to sell them individually, but I’m part of a larger model, and you find some way to figure out the value of what people are paying for.

And the hope there really is that although the LLMs have a very big head start because they’ve accumulated so much information already, human beings are going to continue to evolve. We’re going to have new questions that haven’t been considered before. We’re going to have new developments, hopefully, because we incentivise all of this, that we will continue to learn and grow from.

And so that content will continue to be valuable. And you certainly do hear the AI companies talking about once we have gotten through the adoption phase and into the competition-for-keeping-customers phase, they will want to be known as the people that are providing not just answers, but good answers, quality answers.

And so there will be that interest in making sure that they’re not just coming up with any data, but really quality data, really quality analysis, quality journalism, those sorts of things.

And so that’s kind of the path we’re on to figure out for, among all the players, whether you’re the network providers like Cloudflare, the content creators, or the AI companies, a sort of groundwork to have that conversation and then make sure that there’s some exchange of value there.

The Bot Dilemma: Search vs Training

Nikhil Pahwa: So one last question, as a publisher my interest is in getting the search traffic but not getting the crawling traffic that’s going to use it to take my content to train an LLM.

Is there any way of distinguishing between these two types of bots, or is it like if you block the bot traffic, you block the search traffic as well?

Douglas Kramer: So this is one of the really contentious issues right now being considered by some competition authorities around the world: if they are different bots and they are authentically identifying what they are doing, are they different bots in here?

They are often not different bots for the large search engines. And so if they do that, yes, you can block that. But there is a real catch-22 here when it comes to the search engine bots, and obviously Google is the largest name in search.

And generally, the same bot that is crawling websites for search, which publishers very much still want to be a part of, because they might get clicks and still have ad revenue, is the same bot that’s going to be crawling for training as well.

We have seen that since we have given the ability for websites to block AI bots, there’s quite a bit of adoption of that, that the default really has changed and people block a lot of AI crawlers. The very clear exception is Google, because to shut off the training bot would also mean to shut off the search bot, and they don’t want to do that yet.

So the CMA, which is the competition regulator in the United Kingdom, is probably furthest down this road considering it. They have come out with draft language trying to figure out whether or not Google should be required to split those two bots.

Google says that that becomes more expensive to them, it’s inefficient, operationally it’s not great. So far, the CMA is pushing Google at least to come up with some way to say that if someone describes a preference that, okay, I’m going to let you in but only for search purposes, not for training purposes, that Google on the back end would be careful not to use that in training.

I think content creators have some concern that that is harder to audit and prove once they’ve been forced to give access. But that’s what is happening right now a lot in the regulatory space, to figure out how to empower content creators to be able to make those choices and really have that autonomy.

Read More:

For You

← 목록으로