CLI AI

Find duplicate files on the Linux command line

2026-06-13

To find duplicate files on Linux from the command line without installing anything, hash every file and let uniq flag the repeats: find d -type f -exec md5sum {} + | sort | uniq -w32 -dD. Nearly every search result is a review of a tool you have to install first — fdupes, rdfind, rmlint. On someone else's server or in a throwaway container, you often can't. Below is a pipeline built entirely from coreutils, plus the honest line where it's worth installing a real tool.

1. Duplicates by content, no packages required

clai
$ clai find files with identical content in folder d→ find d -type f -exec md5sum {} + | sort | uniq -w32 -dD4f784201526479699db75d69c7abad6d  d/a/one.txt4f784201526479699db75d69c7abad6d  d/b/two-copy.txt

uniq -w32 compares only the first 32 characters of each line — exactly the length of an md5 hash, so the comparison is by content, not by name. -d keeps only repeats, -D prints every line in each group instead of just one.

2. Count before you decide

clai
$ clai how many groups of duplicates are there→ find d -type f -exec md5sum {} + | sort | uniq -w32 -d | wc -l

One line per group. If there are three groups, eyeball them. If there are three thousand, you need a tool with an interactive mode, not a pipeline.

3. Don't hash terabytes for nothing

clai
$ clai find duplicates only among files bigger than 10 megabytes→ find . -type f -size +10M -exec md5sum {} + | sort | uniq -w32 -dD

md5 reads the whole file. On a large tree, filtering by size first speeds everything up by an order of magnitude — and it's the big files that eat the disk anyway.

4. Faster: size first, hash second

clai
$ clai find duplicates fast by filtering on size first→ find . -type f -printf '%s\n' | sort -n | uniq -d | while read s; do find . -type f -size "${s}c" -exec md5sum {} + ; done | sort | uniq -w32 -dD

Files of different sizes can't be duplicates. Only the ones whose size collides get hashed. On a tree with a million files, that's the difference between minutes and hours.

5. Keep one copy from each group

clai
$ clai show which duplicates I can delete, keeping one copy of each→ find . -type f -exec md5sum {} + | sort | uniq -w32 -d --all-repeated=separate | awk 'seen[$1]++ {print $2}'

awk 'seen[$1]++' skips the first occurrence of a hash and prints every one after it — the candidates for deletion. It's a list to read, not to pipe into xargs rm. Read it.

6. When it's finally time to install a tool

clai
$ clai find duplicates and let me delete them interactively→ fdupes -r -d .

If there are hundreds of groups and you want a prompt for each one, fdupes -d asks what to keep. rdfind can replace duplicates with hard links without deleting anything. The pipeline above is for the places where you can't install packages at all.

Gotchas

  • md5 is fine for finding duplicates, not for security. An md5 collision can be built on purpose; it won't happen by accident. Use sha256 for integrity checks — md5 is faster and good enough for spotting copies.
  • Hard links look like duplicates. The same file under two names hashes the same but takes up space only once. Filter by inode number: find . -type f -printf '%i %p\n' | sort -u -k1,1.
  • Never pipe the list straight into rm. One wrong filter and you delete both copies. Save the list to a file, read it, then delete.

Related questions

How do I find duplicates by name instead of content? find . -type f -printf '%f\n' | sort | uniq -d shows filenames that repeat across different directories.

Does this work on macOS? Replace md5sum with md5 -r — the output format matches. You'll also need -exec stat -f instead of find's -printf.

What's fastest for hundreds of thousands of files? The approach from section 4, or rdfind, which sorts by inode before touching the disk.

See also

CliAI won't hash your files for you, but it saves you from re-deriving the uniq -w32 flag every time. Install it in one line.