Metadata Scrubbing With mat2 — Cleaning Files Before Upload
When you’re preparing a document or image for upload to a secure channel, you likely focus on the visible content: the wording, the resolution, the redactions. But the file itself is wrapped in layers of invisible data-metadata-that can betray far more than what’s on the surface. For anyone operating in high-stakes environments where anonymity is the entire point, failing to scrub this data is a catastrophic and entirely avoidable OPSEC failure.
This guide focuses on mat2, a command-line tool designed specifically for stripping metadata from a wide range of file formats. We’ll cover what it actually does, how it fits into a robust privacy workflow, and why its approach differs from the kludgy “Save As” or “Export” functions built into office suites.
Why Metadata Is a Liability, Not Just a Convenience
Metadata is often described as “data about data.” In the context of a digital file, it’s the hidden attributes attached to the visible content-author names, software versions, creation timestamps, and even GPS coordinates in the case of photos. For a privacy-conscious researcher or vendor, this is a trail of breadcrumbs leading back to your identity.
The threat is concrete. A resume submitted as a Word document might contain your previous employers in the revision history, hidden comments, or a local file path revealing your internal directory structure. Worse, if you embed an image, the EXIF data can reveal the camera model and precise GPS coordinates of where the photo was taken, effectively drawing a map to your physical location. As one analysis of metadata risks notes, sharing photos on social media without cleaning EXIF data can expose your residential address and travel trajectory through GPS positioning.
It’s not just about personal photos either. Standard document properties-the “Author” or “Company” fields automatically injected by word processors-can reveal organizational structures and internal naming conventions. Even the time you saved a file can be correlated with your work habits to identify you. In the digital marketplace ecosystem, there is no room for this kind of sloppiness. If an adversary on the other side of the transaction receives a file that leaks your OS username or the hostname of your computer, you’ve just handed them a spearphishing vector or a direct link to a centralized database they can query.
What mat2 Does Different
While many operating systems and desktop applications offer built-in “Remove Properties” functions, these are often incomplete. They might strip the author name but leave the core XML metadata intact, or they might remove document properties but fail to touch embedded thumbnails in JPEGs. This is where command-line tools like mat2 shine-they perform deep, recursive scrubbing across the entire file structure.
mat2 (Metadata Anonymisation Toolkit 2) is a free, open-source utility that takes a “parse and remove” approach. Unlink simple strippers that attempt to remove a single field, mat2 parses the file into its component streams, cleans the metadata from each segment, and then reassembles the file without the statistical fingerprints of the original authoring software. It doesn’t just blank the “created by” line; it removes the ability to trace the file back to a specific software build or machine.
This is a critical distinction for those in high-stakes environments. A PDF might contain hidden annotations or JavaScript that script-kiddie tools miss. Office documents (DOCX, XLSX) are zip archives containing XML where metadata lives in multiple places simultaneously. mat2 understands these container formats and knows where to look, providing a level of assurance that a simple “Print to PDF” or “Clear All Properties” checkbox cannot offer.
The Practical Workflow: Input, Scrub, Verify
Before integrating mat2 into your pipeline, you need to accept a hard truth: metadata scrubbing is not a “set and forget” operation. It is a per-file, pre-flight check that must be verified after the fact.
Here is a baseline workflow for rigorous file hygiene before upload:
- Isolate the environment: Perform all scrubbing operations in a disposable VM or a live Tails session. This ensures that the tool itself isn’t logging your activity to the host OS.
- Strip before you package: Always clean files before placing them into an archive. If you copy a JPEG into a DOCX and then scrub the DOCX, mat2 might not clean the JPEG’s EXIF data nested inside the document. You must clean the images first, then assemble the document, then clean the document again.
- Run mat2 on the target file: Use the command to generate a cleaned copy of the file (e.g., `
mat2 -i document.pdf` for in-place cleaning, or let it output a `.clean.pdf` file). - Verify the output: This is the step most users skip. After cleaning, run mat2 again in “list” mode (
mat2 -l file.clean.pdf) to see if any metadata remains. If you see output, the scrubbing failed, or the file format is unsupported. Never trust a single pass; if a tool fails to parse the file structure, it will often say “Unsupported file format” instead of failing to clean it. If you see that, you need a different approach.
A common pitfall involves file types that are simply containers for other streams. For instance, a PDF file often carries thumbnails of embedded images. While mat2 strips the metadata from the PDF catalog, it might leave the thumbnail’s EXIF data intact if the thumbnail is a raw JPEG stream within the PDF. Understanding your file structure is essential-if you are dealing with high-risk documents, consider converting them to formats that carry less metadata or rasterizing the output to a pure image, though this sacrifices text selectability.
Limitations and the False Sense of Security
It would be negligent not to point out where mat2 and similar tools fall short. Metadata stripping is about file attributes, not content characteristics. Here are the key blind spots:
- Statistical fingerprints: Even after scrubbing, the byte size of the file and the specific way the text is encoded can reveal that it was created in a specific version of a specific word processor. mat2 removes the labeled attributes but cannot alter the core syntax of the language used to write the document content itself.
- Hidden content: mat2 does not perform OCR (Optical Character Recognition) or scan for hidden text layers in images. If you have a PNG with text baked into the pixels that you forgot to white-out, mat2 won’t help you.
- The “comments” problem: If you have a PDF with editorial comments or sticky notes, mat2 will remove the metadata describing those comments, but it often asks whether you want to remove the annotations themselves. If you don’t, you end up with a clean file full of leaked intellectual property.
- Flash and complex objects: mat2 has notoriously poor support for certain file types involving Flash or embedded multimedia players that hold their own metadata stores.
For those dealing with crypto wallet exports or transaction logs, understand that wallet metadata is a separate beast entirely. Off-chain signals-RPC requests, IPs, browser headers-are observable before a transaction even hits the chain, and no file scrubber can fix those leaks because they happen at the network layer, not in the file. mat2 cleans the document you save; it cannot clean the requests your machine makes to sync that document.
Alternative Methods and Fallback Strategies
While mat2 is an excellent tool, it is not the only one. The context of your threat model dictates which method you should choose. If you are in a rush and only need to remove the obvious “Author” field from a DOCX, using the built-in “Inspect Document” function in Microsoft Word (found under File > Info > Check for Issues) is a passable temporary measure. But as one source rightly points out, these in-built functions often fail to catch nested attachment metadata, such as the EXIF in a pasted photo or the revision history in an embedded spreadsheet.
| Torzon Market |
torzon7aphar3x4l5b77nsylgyw26kntbi4m2wemrjh72aczeh27f6qd.onion
|
| Omega Market |
omega7yhz7n4vg4yhf2na2qaaaeatdlqvjbj2juc245mr5muxtnuvgyd.onion
|
| BlackOps |
blackoogcnxogvymmebfwfjhx4k7efpgeoeytxtsev2lc4pqlbz54qad.onion
|
| Nexus |
nexusbem4wmo67jt723niftkejivtgxbsbxkb6aesj5gyzj7b3v3mxid.onion
|
| DrugHub |
drughuj7l72ig56pza77eriu7yh6qsao4xb4yasq2qfjusxzuq6rlwqd.onion
|
A more comprehensive fallback is the creation of “flattened” files. Instead of sending a native DOCX or XLSX, convert the final version to a PDF, and then rasterize the PDF to a flat TIFF or JPG image. This removes all searchable metadata and all editable text layers. You lose utility, but you gain a massive reduction in attack surface. According to security guidance on preventing metadata exposure, encouraging the use of flattened PDFs or screenshots instead of original document formats is a best practice to prevent the accidental sharing of revision histories and comments.
For those wedded to web-based tools, services like Metadata Online offer recursive scanning to detect nested issues and cleanup functions. However, uploading a sensitive document to a third-party web service to clean it is a ludicrous OPSEC failure-you are literally introducing your data into a system outside your control. Stick to local tools; if you need a GUI, use a fork of mat2 like MAT (the original Metadata Anonymisation Toolkit) or a hardened Linux distribution, but remember that command-line usage inherently offers more control over the output path and logs.
Integrating mat2 Into Your Base OPSEC
File scrubbing is just one layer in a multi-faceted security stack. In the ecosystem of privacy tools, it sits alongside Tor Browser for traffic routing and PGP for message encryption. You should treat every file you touch as hostile until proven clean.
Adopt a policy of “clean at the point of creation.” Before you even type the first character of a sensitive document, configure your word processor’s default settings to strip the author name and disable the saving of “last printed” timestamps. If you are using LibreOffice on Tails, this is often already configured, but it’s worth double-checking. If you are building a resume or a product listing, write the text in a basic text editor like Gedit or Notepad++, which saves files without heavy metadata, and then paste it into the final template only after you have scrubbed the template itself.
To those who argue that metadata leaks are a “low-risk” issue because they only reveal information to sophisticated LE analysts, remember that the darknet is a place where a single slip-a username found in a PDF author field-connects your market account to a real-world social media profile via a search engine query. The tools used for attribution are becoming more robust, and the techniques involve profiling based on off-chain exhaust and file artifacts (Wang et al., arXiv). The cost of running mat2 is negligible; the cost of skip it is the end of your operational security.
Final Checks Before Upload
Manual verification is the only way to ensure your data is clean. Do not rely on the assumption that the tool worked. Run the list command frequently and, if you are dealing with images, cross-check with an EXIF viewer to ensure the GPS tag is absent. Create a checkbook for your regular file types-a small shell script that runs mat2 and then greps the output for “No metadata found.”
Even when you think you’ve scrubbed everything, consider the “zero trust” model: assume the file you upload will be parsed by an adversary with forensic tools better than yours. If the file does not absolutely need to be in a vector format, convert it to a plain text file. If it must be a PDF, minimize the amount of time between creation and upload-longer retention times increase the risk of the file being copied to an unsecured cache.
In a landscape where your digital life-documents, photos, and crypto wallets-generates a constant stream of identifying metadata, taking the extra seconds to strip it manually with tools like mat2 is a low-cost investment in your long-term safety. The interface isn’t pretty, but it is precise. And in this game, precision is worth more than convenience.