FOCA
Tool to analyze metadata in documents to discover hidden information.
External Tool
This open-source tool is maintained externally. View the source on GitHub to learn more or run it yourself.
Browse text & tools tools →What's next with FOCA?
Choose how you want to get started.
Use it free
Open the official tool or demo — no account needed.
Self-host it
Run the open-source version on your own infrastructure.
What is FOCA?
FOCA (Fingerprinting Organizations with Collected Archives) is an open-source tool designed to extract metadata and hidden information from documents. It primarily scans web-based documents using search engines like Google, Bing, and DuckDuckGo to identify files containing sensitive data. The tool is widely used by cybersecurity professionals, digital forensics experts, and researchers to uncover organizational fingerprints, such as email addresses, server details, or internal network information. FOCA addresses the challenge of discovering latent data in publicly available documents, which can reveal vulnerabilities or intellectual property theft. Its ability to analyze diverse file formats makes it a critical tool for proactive security assessments and competitive intelligence gathering.
How it works
FOCA is a metadata extraction tool that systematically scans web pages and documents to uncover hidden information. It leverages search engines to locate files and then parses their contents for embedded data, such as EXIF tags, document properties, or embedded URLs. The primary purpose of FOCA is to assist in digital forensics and security audits by revealing organizational footprints. It helps users identify patterns in document metadata that may indicate compromised systems, insider leaks, or unauthorized data sharing. FOCA supports analysis of Microsoft Office, Open Office, PDF, Adobe InDesign, and SVG files. It can extract EXIF data from graphic files and parse hidden metadata like author names, creation dates, and geolocation coordinates. The tool integrates with three search engines to aggregate results, enabling large-scale document collection.
How to use it
- 1Configure search engines: Set up Google, Bing, or DuckDuckGo queries to target specific domains or keywords. 2. Add local files: Upload documents or folders for manual analysis. 3. Run scans: Initiate metadata extraction using the configured search terms or uploaded files. 4. Review results: Analyze extracted data in the output interface, filtering for relevant information like email addresses or server paths. Practical tips include using plugins to extend functionality, prioritizing high-traffic domains for broader results, and verifying file integrity before analysis to avoid false positives.
What it can do
- metadata document analysis
Use cases
Assumptions and limitations
Assumptions
- source: https://github.com/ElevenPaths/FOCA
- license: GPL-2.0 — free to use
- privacy: Self-hosted — you control your data
Limitations
- Dependence on search engine indexing may miss recently uploaded files
- Limited support for non-English document formats or encoding standards
- Processing large datasets requires significant computational resources
- No built-in GUI for non-technical users; requires command-line interface
- Potential legal risks from scraping copyrighted or sensitive documents
Understanding the result
Tool to analyze metadata in documents to discover hidden information.
Tool details
- Clearly flagged when a network request is needed.
- No account, no sign-up, and no tracking of your content.
- Powered by (MIT).
- Built with
- (ElevenPaths/FOCA)
- License
- MIT
- Runs locally
- No — requires a network request
- Verification
- Not yet verified
- Input
- Query
- Output
- Text
Built with ElevenPaths/FOCA. OpenToolVault provides the discovery and browser interface while crediting the original project maintainers.
- Built with
- License
- MIT
Open-source project
OpenToolVault is an independent directory. We are not affiliated with or endorsed by this project.
References
- / — GitHub Repository
Upstream project · GitHub
- GPL-2.0 License
Upstream project
Frequently asked
How does FOCA collect and analyze documents?
FOCA uses search engines to locate documents containing specific keywords or domains. It then downloads and parses these files to extract metadata, such as authorship information, embedded URLs, or geolocation data. The tool supports multiple file formats, including Microsoft Office, PDF, and SVG, and can process both web-hosted and locally uploaded files.
What technical methods does FOCA use for metadata extraction?
FOCA employs a combination of regex pattern matching, file format-specific parsers, and metadata readers to extract information. For example, it uses Python libraries to decode PDF headers, Office XML structures, and EXIF tags. The tool also leverages search engine APIs to crawl web pages and identify downloadable documents containing hidden data.
How do I extract email addresses from a PDF using FOCA?
Upload the PDF file to FOCA's local processing module. The tool will automatically scan the document for text patterns matching email formats (e.g., name@domain.com). Results are displayed in the output interface, where you can filter and export the extracted addresses for further analysis.
How does FOCA compare to tools like TheHarvester or Goolag?
FOCA focuses on metadata extraction from documents, while TheHarvester specializes in email harvesting from search engine results. Goolag is designed for Google-specific searches. FOCA's strength lies in its ability to analyze file contents for hidden data, whereas other tools prioritize surface-level information retrieval.
What should I do if FOCA fails to process a specific file format?
Check if the file format is supported by FOCA's built-in parsers. If not, try converting the file to a compatible format (e.g., PDF to TXT) or use a plugin. If the issue persists, verify the file's integrity and ensure it's not corrupted. For advanced cases, modify the source code to add custom parsing logic.