Skip to content

dupe Guru

Find duplicate files, images, and music on your system.

Self-hostedNot yet verified
Report issueDemo online
GPL-3.0★ 1500

Open the official app on dupeguru.voltaicideas.net

This tool is hosted by its maintainers. Click below to open dupeguru.voltaicideas.net in a new tab — it's their official demo.

Browse converters tools →

What's next with dupe Guru?

Choose how you want to get started.

Use it free

Open the official tool or demo — no account needed.

Free

Self-host it

Run the open-source version on your own infrastructure.

Open

What is dupe Guru?

dupeGuru is an open-source cross-platform tool designed to identify duplicate files on Linux, macOS, and Windows systems. Its primary purpose is to help users efficiently locate and manage redundant files, whether they are exact duplicates or similar in content. The tool addresses the common problem of disk space wastage caused by duplicated media, documents, or other file types. It is particularly useful for users with large media libraries, system administrators managing file servers, and individuals organizing personal data. By leveraging fuzzy matching algorithms and specialized modes for music and images, dupeGuru streamlines the process of cleaning up duplicate files, saving time and storage space. The tool’s flexibility allows users to customize matching criteria, such as filename similarity thresholds or content comparison methods. Its ability to scan both filenames and file contents makes it adaptable to various use cases, from eliminating redundant photos to managing duplicate audio files. With a user-friendly GUI and support for multiple operating systems, dupeGuru balances simplicity with advanced functionality, making it a reliable choice for both casual users and power users seeking precise duplicate detection.

How it works

dupeGuru is a cross-platform GUI tool written primarily in Python 3, with platform-specific UI layers for macOS (Objective-C/Cocoa) and Linux/Windows (Qt5). It scans systems to identify duplicate files based on filenames or content, solving the problem of wasted storage and clutter caused by redundant files. The tool’s core functionality includes fuzzy filename matching, which detects near-duplicates even when filenames differ slightly. It also supports content-based scanning, making it effective for finding identical or similar files regardless of their names. dupeGuru excels at handling media files, with specialized modes for music and pictures. Its Music mode scans metadata like artist, album, and track names, while Picture mode uses fuzzy algorithms to detect visually similar images. These features make it ideal for organizing music libraries or photo collections.

How to use it

  1. 1Launch dupeGuru and select the directories to scan via the 'Add Folder' button. 2. Choose a scan mode (e.g., Music, Pictures, or General) to tailor the detection criteria. 3. Initiate the scan and review the results in the main window, which highlights potential duplicates. 4. Use the 'Mark Duplicates' or 'Delete' options to manage identified files. Practical tips: For music files, enable the Music mode to leverage metadata matching. For images, use the Picture mode to find visually similar files. Always verify scan results before deleting to avoid accidental data loss.

What it can do

  • duplicate finder

Use cases

Assumptions and limitations

Assumptions

  • source: https://github.com/arsenetar/dupeguru
  • license: GPL-3.0 — free to use
  • privacy: Self-hosted — you control your data

Limitations

  • Lacks integration with cloud storage services for scanning remote files.
  • Does not support real-time file monitoring for new duplicates.
  • Content-based scanning may be resource-intensive for very large files.
  • Limited support for non-ASCII filenames in certain configurations.
  • No built-in export options for duplicate reports beyond the GUI.

Understanding the result

Find duplicate files, images, and music on your system.

Tool details

  • Clearly flagged when a network request is needed.
  • No account, no sign-up, and no tracking of your content.
  • Powered by (GPL-3.0).
Built with
(arsenetar/dupeguru)
License
GPL-3.0
Runs locally
No — requires a network request
Verification
Not yet verified
Input
Query
Output
Text
Open-source source & license

Built with arsenetar/dupeguru. OpenToolVault provides the discovery and browser interface while crediting the original project maintainers.

Built with
License
GPL-3.0
View source on GitHub

Open-source project

OpenToolVault is an independent directory. We are not affiliated with or endorsed by this project.

References

Frequently asked

How does dupeGuru handle different file types?

dupeGuru uses a combination of filename fuzzy matching and content analysis. For text files, it compares byte-by-byte content, while media files (like MP3s) use metadata and audio fingerprints. The Music and Picture modes further refine this by prioritizing tags or visual similarity, ensuring accurate results for specialized file types.

How does the fuzzy matching algorithm work?

dupeGuru’s fuzzy matching algorithm calculates similarity scores based on filename patterns, ignoring case and special characters. For example, 'photo1.jpg' and 'Photo_1.jpg' would be flagged as duplicates. The algorithm allows users to adjust thresholds, balancing between false positives and missed duplicates. Content-based matching uses hash comparisons or byte-level checks, depending on the file type.

How do I find duplicate music files by artist and album?

Select the 'Music' mode in dupeGuru, then add your music folders. The tool will scan ID3 tags for artist, album, title, and track numbers. Duplicates will be grouped by these metadata fields, allowing you to easily identify and remove tracks with identical metadata.

How does dupeGuru compare to alternatives like Duplicate Cleaner or fdupes?

dupeGuru offers a GUI with specialized modes for media files, while tools like Duplicate Cleaner focus on simpler filename matching. fdupes, a command-line utility, provides faster content-based scanning but lacks dupeGuru’s customizable fuzzy matching and media-specific features. dupeGuru’s cross-platform support and Python-based architecture also differentiate it from purely CLI tools.

What should I do if dupeGuru fails to detect duplicates?

First, verify that the scan directories include all relevant files. Check for permission issues that might block access to certain folders. If content-based scanning is unreliable, try switching to filename-only mode. For media files, ensure the correct mode (Music/Picture) is enabled. If issues persist, review the log files for error messages or consult the GitHub issue tracker.

Spotted something wrong with dupe Guru, or want to maintain it? See how to help.