A PDF carries a second document nobody reads: who made it, with what, when, and occasionally the folder it was sitting in.
What is in there
| Field | Typically contains | Risk |
|---|---|---|
| Author | Your account or full name | Identity |
| Title | Often the original filename | Context you did not intend |
| Producer | The software and exact version | Low, but fingerprints your stack |
| Creator | The application that authored it | Low |
| CreationDate / ModDate | Timestamps to the second | Reveals timing and revision history |
| Keywords / Subject | Whatever a template left there | Occasionally embarrassing |
| XMP packet | A fuller record, sometimes edit history | Highest |
Check any file in two clicks: File, then Properties, in Acrobat, Preview or most readers. It is worth doing once on a document you have already sent, purely to see what has been going out.
The file path problem
The field that catches people is not a metadata field at all. Some producers embed the source document's full local path, and design tools embed the paths of linked assets.
A path such as C:\Users\firstname.lastname\Documents\Clients\Acme\Q3-restructure\draft.pdf discloses a full name, a client, and a project — none of which appeared anywhere on the page.
This has produced real disclosures. It is not hypothetical, and it is invisible unless you go looking.
Why metadata is also worth keeping
Stripping everything is not automatically right. Metadata does useful work.
A correct Title field is what a browser shows in the tab and what many document systems display in a list — without it, users see the filename, which is often final_v3_REALLY-final.pdf. Search engines read the Title of an indexed PDF. Document management systems sort and search on Author and Keywords. And for archival formats like PDF/A, metadata is part of the specification rather than an optional extra.
The sensible position is not "strip everything" but "know what is there, set what helps, remove what does not". A published PDF should have a real Title. It should probably not have your Windows username.
Removing it
There are two levels, and the difference matters.
Editing the fields — clearing Author, fixing Title — is what a PDF editor's properties dialog does. It handles the standard document information dictionary and is enough for ordinary sharing.
Removing everything, including XMP and any retained revision data, needs a full rewrite of the file. An incremental save appends changes rather than replacing content, so an edited-out field can survive in an earlier revision inside the same file.
- For ordinary documents: edit the fields in a reader, then use Save As rather than Save, which rewrites rather than appends.
- For anything sensitive: render the pages to images with PDF to JPG and rebuild with JPG to PDF. That discards every field, every XMP packet and every retained revision, because nothing survives except pixels.
- The cost of that route is the text layer. The result is a picture of a document — not searchable, not accessible. Use it when the disclosure risk outweighs both.
Metadata after merging and compressing
Both operations change what is carried, usually without telling you.
Merging typically keeps the metadata of the first document and discards the rest — so a combined file can be attributed to whoever authored page one, which may be a colleague or a template you downloaded.
Compressing rewrites the file, which often drops metadata as a side effect. Convenient if you wanted it gone, unhelpful if you had set a Title deliberately.
Either way, check the properties after assembling and before sending. It is the last point at which it costs nothing to fix.
A short habit
Three checks, thirty seconds, before anything leaves.
- Open document properties and read the Author, Title and Keywords fields.
- Set Title to something meaningful — it is what a reader and a search engine will display.
- Clear anything that names a person, a client or a path unless you intend it.
For documents going to a counterparty, a regulator or the public, treat that as part of the send, in the same way you would check the attachment is the right one. Both mistakes are equally easy and only one of them is obvious afterwards.
Frequently asked questions
What metadata does a PDF contain?
Title, Author, Subject, Keywords, the producing software and its version, and creation and modification timestamps. Many also carry an XMP packet with a fuller record, and some embed the full local file path of the source document.
Can PDF metadata reveal personal information?
Yes. The Author field is usually your account or full name, and embedded file paths can disclose a name, a client and a project that appeared nowhere on the page. Both are visible in any reader's document properties in two clicks.
How do I remove PDF metadata completely?
Editing the fields handles ordinary cases, but use Save As rather than Save — an incremental save appends changes, so a cleared field can survive in an earlier revision inside the same file. For anything sensitive, render to images and rebuild, which discards everything.
Should I remove all metadata?
No. A correct Title is what browsers, document systems and search engines display, and archival formats require metadata. The right approach is to set what helps and remove what identifies you, rather than stripping indiscriminately.
Does merging PDFs affect metadata?
Usually yes — most tools keep the first document's metadata and discard the rest. A merged file can therefore be attributed to whoever authored page one, which might be a colleague or a downloaded template. Check the properties after assembling.
Is metadata visible to the recipient?
Yes, in two clicks — File then Properties in essentially any reader. It is not hidden, merely unnoticed. It is worth opening a document you have already sent, just to see what has been travelling with it.