AuthorAZ
Legal

Author Data Privacy: What Writing Apps Know

Beyond your manuscript, writing apps collect telemetry, identifiers and behavioral data. What each category is, and how to check what yours collects.

Alexandru Filip
Aug 10, 2026
8 min read

Most privacy conversations among authors stop at the manuscript. That's the biggest thing, but it isn't the only thing, and the rest is collected far more routinely because nobody thinks to ask about it.

Here's the full picture of what a writing, formatting or publishing app can know, how to find out which parts apply to yours, and which of it is worth caring about.


Six categories of data

1. Content

Your actual words. The manuscript, notes, blurbs, character names. Covered in detail in What Happens to Your Writing When You Store It in the Cloud.

Any app that syncs, spellchecks against a server, offers cloud backup, or shows your document in a browser has your content by architectural necessity.

2. Metadata about content

Frequently overlooked and surprisingly revealing even when the content itself is protected:

  • Document titles and file names — "Untitled Romance Book 4" tells a story
  • Word counts and how they change over time
  • Creation and modification timestamps — which reveal your working hours, and whether you wrote during contracted hours somewhere else
  • Folder structure
  • Number of projects

A service can hold all of this while genuinely not reading your prose.

3. Account and identity data

Email address, name, payment details, IP address (which gives approximate location), and whatever you entered at signup. Standard, necessary for paid services, and the thing that links every other category to you specifically.

For authors running a firewalled pen name, this is the category that matters most — the account is the connection point between identities.

4. Device and telemetry data

Collected automatically, usually without any specific prompt:

  • Device model, OS version, screen size, language, timezone
  • Advertising identifier (IDFA on iOS, GAID on Android)
  • Crash reports, which can include fragments of app state
  • Performance metrics
  • Install source

Mostly benign in isolation. The advertising identifier is the exception — it's designed to be a stable cross-app identifier, which is exactly what makes it useful for linking your behavior across unrelated services.

5. Behavioral data

What you do inside the app: features used, session length, time of day, buttons clicked, screens viewed, where you abandoned a flow. Standard product analytics.

Reasonable for a company improving its product. Worth knowing exists.

6. Third-party SDK data

The category most people never consider. Apps embed other companies' code — analytics, crash reporting, advertising attribution, customer support chat, A/B testing. Each SDK may transmit to its own vendor, under its own terms.

So an app's privacy policy can be honest and still not be the whole story: the developer isn't collecting your data, but the four SDKs bundled in the app are collecting it for four other companies.


How to actually find out

Four methods, in ascending order of effort and descending order of how much you'll enjoy them.

The app store privacy label — five minutes

Both Apple and Google require developers to declare data collection. On iOS it's the "App Privacy" section of the listing; on Android it's "Data safety."

Read for three things:

  1. "Data Used to Track You" — the strongest signal. This means data is linked to you across other companies' apps and websites
  2. "Data Linked to You" versus "Data Not Linked to You" — the difference between identified and aggregate
  3. Whether Content appears at all in the list

These labels are self-declared and occasionally inaccurate, but they're structured, comparable, and take five minutes. Comparing two tools' labels side by side is the single most efficient privacy research an author can do.

The privacy policy — twenty minutes, and you can skip most of it

Don't read it linearly. Search for these terms:

Search forYou're looking for
"third part"Who else receives data
"sell" / "share"Whether data goes to other companies commercially
"retain" / "retention"How long they keep it after you delete
"train" / "machine learning" / "improve our"Model training language
"affiliate"Data sharing within a corporate group, often unrestricted
"aggregate" / "de-identified"Data they consider outside your control — re-identification is often easier than implied
"legal process"Response to subpoenas

Your device's own tools — ten minutes

  • iOS: Settings → Privacy & Security → App Privacy Report. Shows which domains each app actually contacted. This is observed behavior, not a claim
  • Android: Settings → Privacy → Privacy Dashboard
  • Desktop: a firewall like Little Snitch (Mac) or GlassWire (Windows) shows outbound connections in real time

This is the only method that tells you what an app did rather than what it says it does.

The airplane-mode test — ninety seconds

Turn off the network. Use the app normally. Quit and reopen it, still offline. Then reconnect and watch what transmits.

An app that works fully offline and transmits nothing on reconnection isn't sending anything, whatever its policy says.


What's worth caring about

Not everything on this list deserves equal concern. A rough ranking for a working author:

High: unpublished manuscript content. Account data linking a firewalled pen name to your legal identity. Anything shared with data brokers or advertising networks.

Medium: document metadata, if it reveals unannounced projects or working patterns you'd rather not disclose. Retention periods, if you've deleted something for a reason.

Low: crash reports, device model, aggregate feature analytics. This is how software gets fixed, and objecting to all of it costs you more than it protects.

The goal isn't zero collection. It's knowing which of the six categories a given tool is in and deciding whether that's an acceptable trade for what the tool does. A formatting tool that sees your finished, about-to-be-public book is a very different proposition from a notes app holding an unannounced series bible.


Questions worth asking before you adopt a tool

  1. Does it require an account? (Strong signal of a server relationship)
  2. Does the store listing declare Data Used to Track You?
  3. Does the privacy policy mention using content to train or improve models?
  4. What's the retention period after deletion?
  5. Can I export everything, in a format something else can read?
  6. Does it work in airplane mode after a restart?
  7. Which third-party SDKs are in it? (Often findable in the policy's sub-processor list)

Seven questions, maybe twenty minutes per tool. Worth doing once for the two or three tools your unpublished work actually passes through, and not worth doing for the rest.


Where AuthorAZ fits

AuthorAZ is made by the author of this site, so apply extra skepticism here.

AuthorAZ's position on all six categories is that there's no server to send anything to: no account, no cloud sync, no analytics SDK, no advertising identifier, no AI processing. Data stays on the device, and backups are encrypted exports you create and place yourself.

Don't take that on trust. Run the airplane-mode test, then check the iOS App Privacy Report to see what domains it contacts. That applies to us and to every other tool in your stack — a claim you verified yourself is worth more than one you read in an article written by the developer.


Verified 9 August 2026. Privacy labels and policies change; re-check when a tool updates.

Key Takeaways

Six categories of collection; only one of them is your actual prose.

Compare app store privacy labels — five minutes, structured, comparable.

The iOS App Privacy Report shows what an app did, not what it claims.

Free · No account · 79 lessons

Read the free Publishing Course

The whole Publishing Manual, free and online: checklists, templates and decision trees for every stage, with platform figures verified against the Publishing Database.