Split the problem in two. Published books are a legal and collective question — you
reduce exposure there by registering copyright, joining an authors' organization and
negotiating AI clauses in contracts. Unpublished work is a storage question, and it is
entirely in your hands: fewer services holding the file, terms you have actually read, and
sync you turned on deliberately.
The longer version
What you cannot do
Being honest about this first, because a lot of advice implies otherwise:
- You cannot remove your work from a model already trained. Models are not databases
- You cannot retroactively opt out of datasets already assembled and used
- A "no AI training" notice on your copyright page has no established legal effect,
though it costs nothing and states your position
For published books
- Register your copyright. In the US this is what makes infringement enforceable and
determines eligibility for statutory damages and class settlements. It is the most
concrete single step available
- Join an authors' organization. Individual authors have almost no leverage; collective
bodies negotiate and litigate
- Read the AI clauses in publishing contracts. New contracts increasingly grant training
rights. This is negotiable and frequently negotiated
- Set robots.txt on your own website. This governs your site, not your books — but if
you post sample chapters, serials or substantial excerpts, it is the difference between
offering them to crawlers and not
On that last point, there is a useful distinction most sites miss: training crawlers and
retrieval crawlers are different things. You can block the crawlers that collect training
corpora while allowing the ones that fetch a page because a user asked a question and then
cite the source. "Don't train on us, do cite us" is a coherent, implementable position.
For unpublished work
This is where the leverage actually is.
- Inventory which tools receive the file. Word processor, cloud storage, notes app,
grammar checker, beta-reader platform, formatting tool, anything with a browser
extension that reads page contents. The list is usually longer than expected
- Read one section of each tool's terms — Your Content, User Content, or License.
Look for sublicensable, derivative works, and language about improving,
developing or training. Then check whether any opt-out is on by default and whether it
survives a policy change
- Turn off the sync you did not choose. OS-level backup sweeping your Documents folder,
mobile autosave-to-cloud, "connected experiences" toggles, browser extensions with
read-page permission
- Split deliberately. Published blurbs, buy links and public bios can live anywhere.
Unreleased manuscripts, unannounced series and pen-name material are worth keeping local
with an encrypted off-site backup
- Fix backups first. An author who goes local and neglects backups has traded an
exposure risk for a far likelier data-loss risk. That is a bad trade
What this buys you
Realistically: it stops new copies of unpublished work accumulating on servers you do not
control, and it reduces the number of terms-of-service documents governing your draft from
eleven to two.
It does not solve the industry problem. A policy fix would be better, and it is not
arriving this quarter.
Common exceptions
- Accessibility tools — dictation, screen readers, text-to-speech — are not what this
debate is about.
- On-device processing sends nothing regardless of branding.
- Business tiers of consumer products often carry stricter terms.
- The EU's approach, including text-and-data-mining exceptions and rights reservation,
differs from US fair-use analysis.
Sources
- US Copyright Office, Circular 1; 17 U.S.C. §§411–412.
- Authors Guild survey, December 2023; University of Cambridge Minderoo Centre survey, 2025.
General information, not legal advice. The landscape changes frequently. Verified
9 August 2026.