Protecting Publisher Content in the Age of AI | HighWire Best Practice Webinar

Protecting Publisher Content in the Age of AI | HighWire Best Practice Webinar

Best Practices Webinar Series states 29 October 2025

About This Webinar

The rapid rise of large language models (LLMs) and generative AI is reshaping the foundation of scholarly publishing. From zero-click searches and aggressive web scraping to the disruption of traditional discovery channels, AI’s impact on content integrity and publisher business models is profound and complex.

In this session, experts from across the scholarly communications ecosystem discuss how publishers can actively protect their content while still embracing AI’s transformative potential. The panel explores the realities of Generative Engine Optimization (GEO), the nuances of licensing agreements for AI training, and practical strategies for making research FAIR (Findable, Accessible, Interoperable, and Reusable) for both humans and machines. They also unpack actionable steps publishers can take to adapt their metadata, safeguard their intellectual property, and transition from simply blocking AI to building viable new licensing models.

Featured Speakers

  • Joshua Routh — Director of Hosting Products, HighWire Press
  • Pascal Hetzscholdt — Senior Director of AI Strategy & Content Integrity, Wiley
  • Mark Hahnel — VP of Open Research, Digital Science
  • Josh Nicholson — Chief Strategy Officer, Research Solutions; Co-founder, scite

Webinar Q&A Highlights

1. What business models are used for licensing journal content to AI platforms, and how is revenue shared with societies?

 
Pascal: Every technology company uses data differently. Some models focus on conversational ability, others on domain-specific applications, and others on data repositories built for developer experimentation. Each use case brings its own legal and financial framework, so every partnership becomes an individually tailored agreement based on the sensitivity and confidentiality of the content.

Josh: We are exploring models that open up fragments, what we call smart citations, for retrieval-augmented generation. This lets publishers track usage when an AI interrogates an article, generating recurring revenue while solving attribution. It allows content to surface relevantly within these tools without exposing the full version of record.

2. Are there technical tools to prevent or limit AI platforms from scraping content without permission or licensing agreements?

 
Pascal: Yes. Vendors provide granular tools to identify scrapers, even ones that constantly change IPs or names. The goal isn’t blunt-force blocking, since that risks blocking benign bots that boost brand visibility, but proving that scraping is happening so publishers can negotiate licensing fees.

Mark: Bots are clever. If content is on the web, a bot will likely find it. That said, the industry is shifting from a “wild west” era of asking forgiveness to a more regulated space.

Joshua: Publishers should prepare for future regulation now. Adding TDM (Text and Data Mining) rights reservation metadata to content clearly signals to LLMs that scraping isn’t permitted. When firmer legislation arrives, that signaling becomes evidence publishers can point to.

3. How reliable are enterprise data protections, and are tech companies finding ways around them?

 
Pascal: AI now permeates everything, from chips to networks to operating systems, which amplifies potential threat vectors. Development is moving so fast that there’s little time for thorough testing, so vulnerabilities, prompt injections, and exploits are inevitable. Organizations should experiment with these tools in sandbox environments and build strong internal guardrails before layering AI onto sensitive corporate systems.

Mark: I’m somewhat reassured by models that explicitly state they won’t train on enterprise data. Since this isn’t a monopoly, any major company that breaks that promise risks losing ground to competitors who better protect user data.

4. Why do LLMs still surface retracted studies even when publishers have properly marked them as retracted?

 
Pascal: It’s neither a lack of willingness nor pure ignorance. It largely comes down to the speed of development and deployment to hundreds of millions of users without adequate testing. AI developers are generally receptive to this feedback, particularly since highly regulated sectors like legal and healthcare demand absolute accuracy. Fixing every flagged issue simply takes time.

Josh: LLMs have consumed the entire web, so the roughly 41,000 retracted documents in existence aren’t a large share of their core training data. The bigger challenge isn’t the initial training, it’s making sure that when LLMs dynamically summarize the web today, they’re drawing on high-value, newly published, and properly attributed content.

 

Latest news and blog articles