What's Hiding in the Metadata of the Files You Share and Publish?

Every file you share, post, and publish carries a second, invisible layer of information that you may overlook. When you hit send on a file or post a file to your organizations Sharepoint or CMS, metadata travels with it, including who made it, when, with what software and sometimes a geographic location.

Circulus in Definiendo
You know the recursive definition: metadata is data about data. Metadata decides whether a file can be found, trusted, and reused. There are two sides to the metadata coin
The useful side: search, rights, provenance
The risky side: unintended disclosure
Documents and images move through our organizations every day and it’s important to think about the metadata that travels along with them. Let’s look inside the metadata layers of some common formats.
PDF: Two Metadata Layers
You can think of a PDF document having two layers of metadata:
The Document Information Dictionary (legacy)
XMP Fields (modern)
The document information dictionary is where typical document properties are stored. Document properties are the original metadata fields created by Adobe. This metadata is restricted and supports a handful of hard coded fields.
Document Information Fields – Some of the key document property metadata fields include
File name, title, author, subject, keywords
Creator: date created, date modified, Application (e.g., Adobe InDesign 17.3 Macintosh)
Producer: the software that wrote the PDF (e.g., Adobe PDF Library 16.0.7)
PDF Version
Location
File Size
Page Size

Extensible Metadata Platform (XMP) was introduced in April 2001 and represents a chunk of raw, uncompressed XML text embedded entirely inside a PDF object stream. Because this is XML-based metadata, it’s extensible. You can define custom namespaces to suit your business needs.
XMP fields – Some of the key XMP metadata fields include
dc:title, dc:creator, dc:description and dc:rights (copyright statement)
pdf:Keywords, pdf:Producer and xmp:CreatorTool
xmp:CreateDate, xmp:ModifyDate and xmp:MetadataDate
xmpMM:DocumentID and xmpMM:InstanceID: unique IDs that track a file and its versions
xmpMM:History: a log of save events, present in some files

What Can Leak
The utility of all this metadata is clear and can provide information about a document before it is even read. However, it’s important to consider the exact metadata traveling with your document when it enters the world. Do you want the original file name stored in the metadata if the document title has changed? Does a document created months before it’s published detail information you would rather not share?
Word: Properties Plus Hidden Content
Word files carry the .docx extension, which designates a file that is a zip archive of XML files. Word metadata is captured in three basic parts:
Core properties
Extended (app) properties
Optional custom properties

Core properties
Title, Subject, Author (creator), Keywords and Comments (description)
Last modified by and revision number
Created, Modified, and Last printed dates
Category
Extended properties
Application and application version
Company and Manager
Template name
Total editing time
Pages, Words, Characters, Lines and Paragraphs counts
Document security setting and hyperlink base
Custom properties
User-defined name and value pairs, often added by document management systems or sensitivity-label tools
There is some “hidden content” that behaves like metadata in Word files such as
Tracked changes and comments, each stamped with an author name and time
Hidden text and the names of previous editors
Headers, footers, and file paths in links or templates
What Can Leak
A common practice when reusing a Word file is “Save As.” In which you save a current file as a new file with a new name. However, the metadata in the new file will still reflect what was in the original Word file. Additional metadata such as who last edited the file, how long it took to write, and a full review history that you might not want the recipient to read also travels with the file.
Excel: The Workbook's Inner Structure
An .xlsx workbook uses the same general metadata parts as Word, so the document-level fields look familiar. The surprise is how much structural information a workbook adds.
Document properties
Title, Subject, Author (creator), Manager, Company, Category, Keywords, Comments
Created and Modified dates, Last saved by
Category and Content status
Application, application version, Company and Manager
Document security setting

Workbook-specific information
Worksheet names and defined (named) ranges, listed in the file's extended properties
Hidden sheets, plus "very hidden" sheets that don't appear in the normal Unhide menu
Hidden rows and columns
Cell comments and notes, with author names
Links to other workbooks, which can expose file paths and server names
Data connections and Power Query sources
Pivot caches, which can keep a copy of the underlying source data
What Can Leak
There’s a lot that can leak in an Excel file. Things like internal file paths, database or server names, information on a hidden tab, and the names of anyone who left a comment. If your organization uses Sharepoint or OneDrive, custom properties and hidden data folders are captured in the metadata and might not be something you want to expose for public consumption.
Review Your Own Metadata
A short look inside your own files might be very revealing. Did someone purchase a PowerPoint template and forget to update the metadata? Does that “final” PDF you’re sending still list the draft’s file name and the person who started it? Is that hidden worksheet with your internal costs and margins being sent along with that client quote? This kind of information can often be viewed in just a few clicks.
From Hidden to Helpful
Metadata can be a blessing and a curse. Left unmanaged, it can reveal more than was ever intended. Managed well and updated, it makes content findable, trustworthy, and ready to use for people, systems, and AI.
DCL can help ensure you do metadata right. We help organizations extract, standardize, and enrich the metadata with validation and QA to keep it consistent at scale.




Comments