To find duplicate files on Linux from the command line without installing anything, hash every file and let uniq flag the repeats: find d -type f -exec md5sum {} + | sort | uniq -w32 -dD. Nearly every search result is a review of a tool you have to install first — fdupes, rdfind, rmlint. On someone else's server or in a throwaway container, you often can't. Below is a pipeline built entirely from coreutils, plus the honest line where it's worth installing a real tool.
1. Duplicates by content, no packages required
$ clai find files with identical content in folder d→ find d -type f -exec md5sum {} + | sort | uniq -w32 -dD4f784201526479699db75d69c7abad6d d/a/one.txt4f784201526479699db75d69c7abad6d d/b/two-copy.txt
uniq -w32 compares only the first 32 characters of each line — exactly the length of an md5 hash, so the comparison is by content, not by name. -d keeps only repeats, -D prints every line in each group instead of just one.
2. Count before you decide
$ clai how many groups of duplicates are there→ find d -type f -exec md5sum {} + | sort | uniq -w32 -d | wc -l
One line per group. If there are three groups, eyeball them. If there are three thousand, you need a tool with an interactive mode, not a pipeline.
3. Don't hash terabytes for nothing
$ clai find duplicates only among files bigger than 10 megabytes→ find . -type f -size +10M -exec md5sum {} + | sort | uniq -w32 -dD
md5 reads the whole file. On a large tree, filtering by size first speeds everything up by an order of magnitude — and it's the big files that eat the disk anyway.
4. Faster: size first, hash second
$ clai find duplicates fast by filtering on size first→ find . -type f -printf '%s\n' | sort -n | uniq -d | while read s; do find . -type f -size "${s}c" -exec md5sum {} + ; done | sort | uniq -w32 -dD
Files of different sizes can't be duplicates. Only the ones whose size collides get hashed. On a tree with a million files, that's the difference between minutes and hours.
5. Keep one copy from each group
$ clai show which duplicates I can delete, keeping one copy of each→ find . -type f -exec md5sum {} + | sort | uniq -w32 -d --all-repeated=separate | awk 'seen[$1]++ {print $2}'
awk 'seen[$1]++' skips the first occurrence of a hash and prints every one after it — the candidates for deletion. It's a list to read, not to pipe into xargs rm. Read it.
6. When it's finally time to install a tool
$ clai find duplicates and let me delete them interactively→ fdupes -r -d .
If there are hundreds of groups and you want a prompt for each one, fdupes -d asks what to keep. rdfind can replace duplicates with hard links without deleting anything. The pipeline above is for the places where you can't install packages at all.
Gotchas
- md5 is fine for finding duplicates, not for security. An md5 collision can be built on purpose; it won't happen by accident. Use sha256 for integrity checks — md5 is faster and good enough for spotting copies.
- Hard links look like duplicates. The same file under two names hashes the same but takes up space only once. Filter by inode number:
find . -type f -printf '%i %p\n' | sort -u -k1,1. - Never pipe the list straight into rm. One wrong filter and you delete both copies. Save the list to a file, read it, then delete.
Related questions
How do I find duplicates by name instead of content? find . -type f -printf '%f\n' | sort | uniq -d shows filenames that repeat across different directories.
Does this work on macOS? Replace md5sum with md5 -r — the output format matches. You'll also need -exec stat -f instead of find's -printf.
What's fastest for hundreds of thousands of files? The approach from section 4, or rdfind, which sorts by inode before touching the disk.
See also
- Which folder is eating the disk? du, sorted by biggest
- df says the disk is full, du says it isn't
- Stop memorizing find flags
CliAI won't hash your files for you, but it saves you from re-deriving the uniq -w32 flag every time. Install it in one line.