Split the problem in two. Published books are a legal and collective question — you reduce exposure there by registering copyright, joining an authors' organization and negotiating AI clauses in contracts. Unpublished work is a storage question, and it is entirely in your hands: fewer services holding the file, terms you have actually read, and sync you turned on deliberately.
Being honest about this first, because a lot of advice implies otherwise:
On that last point, there is a useful distinction most sites miss: training crawlers and retrieval crawlers are different things. You can block the crawlers that collect training corpora while allowing the ones that fetch a page because a user asked a question and then cite the source. "Don't train on us, do cite us" is a coherent, implementable position.
This is where the leverage actually is.
Realistically: it stops new copies of unpublished work accumulating on servers you do not control, and it reduces the number of terms-of-service documents governing your draft from eleven to two.
It does not solve the industry problem. A policy fix would be better, and it is not arriving this quarter.
General information, not legal advice. The landscape changes frequently. Verified 9 August 2026.