Skip to content

Metagoofil

Extract metadata from public documents found on a target domain.

Self-hostedNot yet verified
Report issue
MIT★ 1800Source project only — not browser-runnable

External Tool

This open-source tool is maintained externally. View the source on GitHub to learn more or run it yourself.

Browse text & tools tools →

What's next with Metagoofil?

Choose how you want to get started.

Use it free

Open the official tool or demo — no account needed.

Free

Self-host it

Run the open-source version on your own infrastructure.

Open

What is Metagoofil?

Metagoofil is an open-source metadata extraction tool designed to harvest information from public documents hosted on target websites. It specializes in analyzing files such as PDFs, Word documents, Excel spreadsheets, and PowerPoint presentations to retrieve embedded metadata. This metadata often includes details like author names, creation dates, revision history, and internal references, which can reveal sensitive information about organizations. The tool is primarily used by security researchers, penetration testers, and digital forensicators to conduct reconnaissance during ethical hacking operations. By automating the extraction of metadata, Metagoofil addresses the challenge of manually sifting through vast amounts of publicly available data to identify potential vulnerabilities or intelligence gaps. Its ability to aggregate metadata from multiple sources makes it a critical asset in information-gathering phases of cybersecurity assessments.

How it works

Metagoofil is a metadata harvester that leverages existing web infrastructure to extract information from documents hosted online. It targets files uploaded to public-facing websites, where metadata may contain unsecured data about the organization or individuals associated with the documents. The tool’s primary purpose is to assist in reconnaissance by uncovering details that could be exploited in cyberattacks. For example, metadata might reveal internal network configurations, employee names, or project timelines, which are often overlooked in standard security audits. Metagoofil supports a wide range of file formats, including PDF, DOC, XLS, PPT, and more, using libraries like hachoir_metadata and pdfminer for parsing. It can recursively scan websites to identify and download documents, then extract metadata such as author, title, and revision history. For instance, a PDF file might contain a hidden comment field indicating the creator’s email address.

How to use it

  1. 1Install dependencies like Python 2.7, hachoir libraries, and pdfminer. 2. Use the downloader.py script to crawl target websites and download documents. 3. Run metagoofil.py with parameters specifying the target URL and output directory. 4. Process the extracted metadata using processor.py to filter and export results in HTML or CSV formats. Practical tips include using proxies to bypass website blocks, limiting the depth of website crawling to avoid excessive data retrieval, and verifying file formats before extraction to prevent errors.

What it can do

  • metadata harvesting

Use cases

Assumptions and limitations

Assumptions

  • source: https://github.com/laramies/metagoofil
  • license: GPL-2.0 — free to use
  • privacy: Self-hosted — you control your data

Limitations

  • Relies on publicly accessible documents, ignoring private or password-protected files
  • Limited support for newer file formats or encrypted metadata fields
  • Requires manual configuration for complex website structures or dynamic content
  • May violate terms of service if used without explicit authorization
  • Processing large datasets can be resource-intensive and time-consuming

Understanding the result

Extract metadata from public documents found on a target domain.

Tool details

  • Clearly flagged when a network request is needed.
  • No account, no sign-up, and no tracking of your content.
  • Powered by (MIT).
Built with
(laramies/metagoofil)
License
MIT
Runs locally
No — requires a network request
Verification
Not yet verified
Input
Query
Output
Text
Open-source source & license

Built with laramies/metagoofil. OpenToolVault provides the discovery and browser interface while crediting the original project maintainers.

Built with
License
MIT
View source on GitHub

Open-source project

OpenToolVault is an independent directory. We are not affiliated with or endorsed by this project.

References

Frequently asked

What types of files does Metagoofil support?

Metagoofil supports common document formats including PDF, DOC, XLS, PPT, and others through libraries like hachoir_metadata and pdfminer. It focuses on extracting embedded metadata rather than content text, making it effective for discovering hidden organizational details.

How does Metagoofil parse metadata from documents?

The tool uses libraries such as hachoir_core and pdfminer to analyze file structures and extract metadata fields. For example, PDF metadata is parsed by reading the file’s trailer section, while Word documents leverage the DOCX format’s XML structure to retrieve author and revision information.

How do I run Metagoofil on a target website?

First, install Python 2.7 and required dependencies. Use the downloader.py script to crawl the target domain and download documents. Then execute metagoofil.py with parameters specifying the target URL and output directory. Finally, process results with processor.py to filter and export metadata in HTML or CSV format.

How does Metagoofil compare to ExifTool?

Metagoofil is specifically designed for web-based document metadata harvesting, while ExifTool focuses on image and video metadata. Metagoofil integrates web crawling capabilities, whereas ExifTool requires manual file input. Both tools use open-source libraries but serve different primary use cases.

What should I do if Metagoofil fails to extract metadata?

Check if the file format is supported, verify the document isn’t encrypted or corrupted, and ensure the target website allows public access. If the issue persists, update dependencies like hachoir_metadata or try alternative parsing libraries. Use the --debug flag to identify specific parsing errors.

Spotted something wrong with Metagoofil, or want to maintain it? See how to help.