gau
Get All URLs - fetches known URLs for a domain from public web archives like the Wayback Machine.
External Tool
This open-source tool is maintained externally. View the source on GitHub to learn more or run it yourself.
Browse security tools →What's next with gau?
Choose how you want to get started.
Use it free
Open the official tool or demo — no account needed.
Self-host it
Run the open-source version on your own infrastructure.
What is gau?
gau is an open-source command-line tool designed to collect known URLs associated with a domain by querying multiple public intelligence sources. It aggregates data from AlienVault's Open Threat Exchange, the Wayback Machine, Common Crawl, and URLScan, providing cybersecurity professionals with a comprehensive list of historical and active URLs. This tool is primarily used by ethical hackers, red teams, and threat intelligence analysts to identify potential attack vectors, monitor domain activity, and uncover hidden assets. By consolidating data from diverse sources, gau simplifies the process of domain reconnaissance, which is critical in penetration testing and incident response scenarios. Its ability to bypass traditional search engine limitations makes it a valuable asset in uncovering outdated or forgotten web resources. The tool addresses the challenge of manually gathering URL data from fragmented sources, offering a streamlined solution for security researchers. Its lightweight design and modular architecture allow users to customize data collection parameters, such as excluding certain file types or specifying output formats. gau's integration with threat intelligence platforms like AlienVault enables real-time analysis of malicious activity, while its reliance on the Wayback Machine helps identify defunct websites that may still host sensitive information. This makes it an essential tool for proactive security assessments and digital forensics.
How it works
gau (getallurls) is a command-line utility that aggregates URLs from multiple public repositories, including AlienVault's Open Threat Exchange, the Wayback Machine, Common Crawl, and URLScan. It is designed to help security professionals discover historical and current web assets linked to a domain, enabling targeted reconnaissance and threat analysis. The tool's primary purpose is to automate the collection of URLs that might otherwise require manual effort across disparate sources. By consolidating this data, gau accelerates the identification of potential vulnerabilities, such as exposed admin panels, outdated files, or malicious redirects. gau leverages AlienVault's Open Threat Exchange to identify domains associated with known threats, the Wayback Machine to retrieve historical URLs, and Common Crawl for archived web content. It also supports URLScan for real-time analysis of suspicious domains. Users can specify domains via stdin, files, or direct input, with options to filter results (e.g., excluding image files) and control output formats.
How to use it
- 1Install gau via Go (requires Go 1.18+), or use precompiled binaries from the release page. 2. Run the tool by piping domains into stdin: `printf example.com | gau` or provide a list via a file: `cat domains.txt | gau --threads 5`. 3. Specify output files with `--o output.txt` to save results. 4. Filter results using `--blacklist` to exclude unwanted file types (e.g., `--blacklist jpg,png`). Practical tips: Use `--threads` to parallelize requests, combine with tools like `sort` or `uniq` to deduplicate results, and verify API access credentials if using premium services like URLScan. Always validate outputs against the sources to ensure accuracy.
What it can do
- URL fetching tool
Use cases
Assumptions and limitations
Assumptions
- source: https://github.com/lc/gau
- license: MIT — free to use
- privacy: Self-hosted — you control your data
Limitations
- Dual-use tool — use only with explicit authorization on systems you own or have permission to test.
- Reliance on external APIs may introduce latency or availability issues
- Data from Common Crawl and Wayback Machine may be incomplete or outdated
- Lacks real-time monitoring capabilities for dynamic URL generation
- Requires internet access to query external sources
Understanding the result
Get All URLs - fetches known URLs for a domain from public web archives like the Wayback Machine.
Tool details
- Clearly flagged when a network request is needed.
- No account, no sign-up, and no tracking of your content.
- Powered by (MIT).
- Built with
- (lc/gau)
- License
- MIT
- Runs locally
- No — requires a network request
- Verification
- Not yet verified
- Input
- Query
- Output
- Text
Built with lc/gau. OpenToolVault provides the discovery and browser interface while crediting the original project maintainers.
- Built with
- License
- MIT
Open-source project
OpenToolVault is an independent directory. We are not affiliated with or endorsed by this project.
References
- / — GitHub Repository
Upstream project · GitHub
- MIT License
Upstream project
Frequently asked
What is gau and what problem does it solve?
gau is a command-line tool that solves the problem of manually gathering URLs associated with a domain by aggregating data from multiple public sources. It automates the collection of historical and current URLs from AlienVault, the Wayback Machine, Common Crawl, and URLScan, enabling security professionals to identify potential attack vectors, monitor domain activity, and uncover hidden assets without manually querying each source individually.
How does gau collect data from these sources?
gau queries AlienVault's Open Threat Exchange for threat-related domains, the Wayback Machine for archived URLs, Common Crawl for historical web content, and URLScan for real-time analysis. It uses API integrations to fetch data, then consolidates and deduplicates results. The tool respects rate limits and authentication requirements for premium services like URLScan, ensuring efficient and compliant data retrieval.
How do I use gau with a list of domains?
To process multiple domains, create a file (e.g., domains.txt) with each domain on a separate line. Run `cat domains.txt | gau --threads 5` to process all domains in parallel. Use `--o output.txt` to save results to a file. For example: `cat domains.txt | gau --threads 5 --o results.txt` will output all collected URLs to results.txt, using 5 parallel threads for faster processing.
How does gau compare to alternatives like waybackurls or crt.sh?
gau extends Tomnomnom's waybackurls by adding data from AlienVault and URLScan, providing broader threat intelligence. Unlike crt.sh, which focuses on certificate transparency logs, gau emphasizes URL discovery across multiple sources. It is more comprehensive than tools like subfinder for reconnaissance but lacks the subdomain enumeration capabilities of dedicated subdomain finders. gau is ideal for combining historical and threat data, while crt.sh is better for SSL certificate analysis.
What should I do if gau returns an error about API access?
If gau encounters API access issues, verify that your internet connection is stable and that you have valid credentials for premium services like URLScan. Check the tool's documentation for required API keys or authentication steps. If using a local installation, ensure that the tool's configuration files (e.g., .gau.toml) are correctly set up. For rate limit errors, try reducing the number of parallel threads with `--threads` or pause between requests.