Content teams inherited two copyright questions nobody assigned them. The first: AI engines are reading and citing your content, so what rights do you actually have over that? The second: your team now ships AI-assisted content, so who owns it, and what could it infringe? In 2026 both questions have real answers, mostly of the form "here's what's settled, here's what isn't, here's what to do meanwhile." That's this guide. One thing to say plainly before any of it: this is a practitioner's map of the AI search copyright terrain, and none of it is legal advice; the calls specific to your business belong with counsel.
Three Questions the Courts Keep Separate
The most useful thing to understand about AI copyright law right now is its structure. Courts and regulators are treating three activities as legally distinct: training (a model ingesting your content to learn), retrieval (an engine fetching and quoting your pages to build answers), and output (what the model generates, and who owns it). Different cases attack different layers, different defenses apply, and a headline about one layer usually says nothing about the others. Most confusion in marketing conversations about "AI stealing content" comes from mixing the three.
A heavily upvoted r/webdev thread asked the question underneath all of it, where did AI companies get legal permission to train on copyrighted data, and the honest answer as of September 2026 is that most never asked; they made a fair-use bet, and the courts haven't finished grading it:
Where did AI companies get legal permission to train on copyrighted data?
This is a question I have thought about a lot lately. If it wasn’t for the repositories hosted in Github and other platforms, none of the models would exist now. Specifically, I don’t remember agreeing on anything that said something about...
Where the Big Cases Actually Stand
Three matters define the September 2026 map, and each one moved this summer.
New York Times v. OpenAI remains the main event on the training question, still actively litigated with no final merits ruling. The development that reset expectations: on September 2, the US government filed a brief backing OpenAI's position that training on copyrighted material is generally fair use, warning that restricting it could hold back creative and scientific progress. The reaction in creator communities was grim, with one widely shared thread reading it as the end of the argument. The scale is still moving, though, and a brief is a weight on one side rather than the verdict.
The Anthropic books settlement went final in July, when a judge approved the $1.5 billion class settlement with authors over pirated books used in training. Hold on, one distinction matters here first, because this case gets miscited constantly: the settlement resolved past claims about pirated source material, and it set no precedent on whether training itself is lawful. What it did do is put a visible market price on getting provenance wrong, which is why every serious lab now cares where its data came from.
Reddit v. Perplexity covers the layer closest to AI search. In late July a Manhattan judge rejected most of Perplexity's bid to dismiss Reddit's suit over data scraping, letting claims proceed that scrapers circumvented protective measures. Procedural, not final. But it keeps alive the question content owners care most about: whether a site's technical controls carry legal force when a crawler ignores them.
The pattern across all three: settlements and briefs instead of clean precedents, which means the working rules for the next few years are being set by deals and risk tolerance rather than by decided law.
AI search traffic grew 835% this year. Is your content ready?
RankControl generates 26 content formats optimized for ChatGPT, Claude, and Perplexity. Published on your domain, matched to your brand.

Your Content Inside the Engines
So what does the unsettled law mean for the content you publish? Practically, your controls are technical and your compensation is visibility, and both deserve a deliberate decision.
The technical controls are the crawler directives: GPTBot, ClaudeBot, PerplexityBot, Google-Extended and the rest of the allow-block roster. Two properties of those controls matter for strategy. They're voluntary signals whose legal force is precisely what the scraping cases are testing. And blocking doesn't remove you from AI answers anyway, because engines keep learning about you from third-party content; one community analysis found roughly a third of sites blocking GPTBot still got cited by ChatGPT through exactly those leaks.
Licensing, the other path, is real money for large publishers and mostly unavailable to everyone else. There's no standard rate card, terms stay confidential, and nobody is writing checks to a B2B SaaS blog. For most content teams the honest frame is the one the Figma blocking episode made concrete: the real choice sits between visibility in the answer layer and a principled absence from it, while protection with payment attached stays a menu reserved for the largest publishers. Make that call deliberately, per crawler, and write it down. A default you never chose is the only wrong answer.
The Copyright Status of What You Ship
Now the second inherited question, ownership of your own output, where the ground is firmer than most teams assume. The US Copyright Office's position has held steady: copyright requires human authorship. Purely machine-generated text is not protectable, while AI-assisted work is protectable to the extent a human exercised genuine creative control through direction, substantial editing, selection, and arrangement. Registration requires disclosing the AI-generated portions and claiming only the human-authored ones.
Sit with the business implication of that first clause for a second. An article your pipeline generated and nobody meaningfully touched may be property nobody owns. A competitor lifting it wholesale would face, at most, a thin claim. Which means full automation with zero human authorship has a hidden cost beyond quality: it produces pages you may not be able to defend. The fix lives in workflow rather than in law. Human editing that changes structure and adds judgment, your own data and customer evidence woven in, and a record of who did what. Those steps strengthen the copyright position, and they make the content better at the same time.
What about infringement risk in the other direction, your AI-assisted article reproducing someone's protected text? The headline cases target the AI companies, and no wave of claims against their business customers has materialized. The narrow real exposure is verbatim reproduction sneaking into output, which an originality check in the publishing pipeline catches. Boring, cheap insurance.

Built by the team that got cited in 48 hours.
Content generation, backlink building, AI visibility tracking, and Google rankings. One platform, zero guesswork.
Vendor Diligence: The Questions Your Tools Should Answer
One more exposure surface hides in the stack itself. Content teams don't train models; they buy tools built on them, and the copyright posture of those tools varies more than their marketing suggests. The major model providers moved early here, with OpenAI's Copyright Shield and Microsoft's Copilot copyright commitment both promising to defend enterprise customers against infringement claims arising from output, and similar indemnities have spread through the serious end of the market since.
That history gives you a clean diligence script for any AI content tool. Does the vendor indemnify you against copyright claims on generated output, and does your tier actually qualify? Can they say what their underlying models were trained on, or at least which providers they build on? And do they run originality screening before output reaches you, or is that left as your job? A vendor with good answers took the Anthropic settlement's provenance lesson seriously. A vendor with none is quietly handing the risk down the chain to you, priced as if it weren't there.
The Five-Point Checklist
Everything above compresses into five workflow decisions a content lead can implement this month.
- Set your crawler policy on purpose. Audit which AI crawlers you allow and block today, decide per engine with the visibility tradeoff in view, and record why.
- Build documented human authorship into the pipeline. Named editors, substantive revision passes, original data in every important piece, and a record of all of it. This is your copyright position.
- Run originality checks before publishing. One automated pass per article closes the realistic infringement exposure.
- Keep provenance records. Which model, which sources, which human edits. The Anthropic settlement's lesson generalizes: provenance is what gets expensive to reconstruct later.
- Recheck the legal map quarterly, never daily. Cases move in months. A quarterly note from counsel beats a panic per headline.
The Strategic Read
Step back from the case names and the 2026 picture organizes itself around one idea a sharp X commentator put well: the fight looks like copyright, and underneath it's about who participates in the value AI builds from human work. Publishers with negotiating power are converting that participation into licenses and settlements. Everyone else participates through the answer layer itself, as citations and mentions and the trust they carry.
Which is why the copyright question and the visibility question end up on the same desk. The team that decides its crawler policy deliberately, ships defensible content, and then actually watches what the engines do with it is playing the whole board. That last part is the instrument we build: per-engine citation tracking that shows what ChatGPT, Perplexity, Claude, Gemini, and Google's AI surfaces are doing with your content week over week, so your policy calls run on your own data instead of on litigation headlines. The law will take years to settle. Your visibility is being decided now, and it's the part you can measure.
See your first AI citation report in under 5 minutes.
No setup calls. No onboarding meetings. Connect your domain and see where AI mentions your brand right now.




