Find Duplicate Files (based on size first, then MD5 hash)

find . -type f -not -empty -printf "%-25s%p\n"|sort -n|uniq -D -w25|cut -b26-|xargs -d"\n" -n1 md5sum|sed "s/ /\x0/"|uniq -D -w32|awk -F"\0" 'BEGIN{l="";}{if(l!=$1||l==""){printf "\n%s\0",$1}printf "\0%s",$2;l=$1}END{printf "\n"}'|sed "/^$/d"

* Find all file sizes and file names from the current directory down (replace "." with a target directory as needed). * sort the file sizes in numeric order * List only the duplicated file sizes * drop the file sizes so there are simply a list of files (retain order) * calculate md5sums on all of the files * replace the first instance of two spaces (md5sum output) with a \0 * drop the unique md5sums so only duplicate files remain listed * Use AWK to aggregate identical files on one line. * Remove the blank line from the beginning (This was done more efficiently by putting another "IF" into the AWK command, but then the whole line exceeded the 255 char limit). >>>> Each output line contains the md5sum and then all of the files that have that identical md5sum. All fields are \0 delimited. All records are \n delimited.

By: alafrosty

2013-10-22 13:34:19

awk cut find sed sort uniq xargs size find md5sum fdupes duplicate files

10 Alternatives + Submit Alt

Find Duplicate Files (based on size first, then MD5 hash)

This dup finder saves time by comparing size first, then md5sum, it doesn't delete anything, just lists them.
This is sample output - yours may be different.
82

find -not -empty -type f -printf "%s\n" | sort -rn | uniq -d | xargs -I{} -n1 find -type f -size {}c -print0 | xargs -0 md5sum | sort | uniq -w32 --all-repeated=separate

grokskookum · 2009-09-21 00:24:14 58
Find Duplicate Files (based on size first, then MD5 hash)

If you have the fdupes command, you'll save a lot of typing. It can do recursive searches (-r,-R) and it allows you to interactively select which of the duplicate files found you wish to keep or delete.
This is sample output - yours may be different.
23

fdupes -r .

Vilemirth · 2011-02-19 17:02:30 9
Find Duplicate Files (based on MD5 hash)

Calculates md5 sum of files. sort (required for uniq to work). uniq based on only the hash. use cut ro remove the hash from the result.
This is sample output - yours may be different.
18

find -type f -exec md5sum '{}' ';' | sort | uniq --all-repeated=separate -w 33 | cut -c 35-

infinull · 2009-08-04 07:05:12 6
Find Duplicate Files, excluding .svn-directories (based on size first, then MD5 hash)

Improvement of the command "Find Duplicate Files (based on size first, then MD5 hash)" when searching for duplicate files in a directory containing a subversion working copy. This way the (multiple dupicates) in the meta-information directories are ignored. Can easily be adopted for other VCS as well. For CVS i.e. change ".svn" into ".csv": find -type d -name ".csv" -prune -o -not -empty -type f -printf "%s\n" | sort -rn | uniq -d | xargs -I{} -n1 find -type d -name ".csv" -prune -o -type f -size {}c -print0 | xargs -0 md5sum | sort | uniq -w32 --all-repeated=separate Show Sample Output
This is sample output - yours may be different.
```
[...]
f2e6bb247f110dcab63b4d38ff7b2dee  ./themes/darkblue_orange/img/b_relations.png
f2e6bb247f110dcab63b4d38ff7b2dee  ./themes/original/img/b_relations.png

f5309bd2a2fc5e512a0cc38ac6f10c09  ./themes/darkblue_orange/img/b_deltbl.png
f5309bd2a2fc5e512a0cc38ac6f10c09  ./themes/original/img/b_deltbl.png

f60bfbb7ce218a55650c1abbbbee06ae  ./themes/darkblue_orange/img/s_lang.png
f60bfbb7ce218a55650c1abbbbee06ae  ./themes/original/img/s_lang.png

f63a5ad833147eeb94adb4496ddbec41  ./themes/darkblue_orange/img/s_theme.png
f63a5ad833147eeb94adb4496ddbec41  ./themes/original/img/s_theme.png

f6ae61146ce3de8fa11b9e84e086bd04  ./themes/darkblue_orange/img/bd_drop.png
f6ae61146ce3de8fa11b9e84e086bd04  ./themes/original/img/bd_drop.png

f95d66c11bfed9198d13a278269c32b2  ./themes/darkblue_orange/img/s_loggoff.png
f95d66c11bfed9198d13a278269c32b2  ./themes/original/img/s_loggoff.png
[...]
```
2

find -type d -name ".svn" -prune -o -not -empty -type f -printf "%s\n" | sort -rn | uniq -d | xargs -I{} -n1 find -type d -name ".svn" -prune -o -type f -size {}c -print0 | xargs -0 md5sum | sort | uniq -w32 --all-repeated=separate

2chg · 2010-01-28 09:45:29 5
Find Duplicate Files (based on MD5 hash) -- For Mac OS X

This works on Mac OS X using the `md5` command instead of `md5sum`, which works similarly, but has a different output format. Note that this only prints the name of the duplicates, not the original file. This is handy because you can add `| xargs rm` to the end of the command to delete all the duplicates while leaving the original.
This is sample output - yours may be different.
2

find . -type f -exec md5 '{}' ';' | sort | uniq -f 3 -d | sed -e "s/.*($.*$).*/\1/"

noahspurrier · 2012-01-14 08:54:12 10

What Others Think

You can learn so many such Unix commands through this platform. This is the best platform for students flutter training who are interested in learning such commands. The above command is used to find duplicate files that are based on size first, then MD5 hash. Keep sharing more updates here.

Alyssalauren · 91 weeks and 6 days ago

Amazing! This blog looks just like my old one! It's on a completely different subject but it has pretty much the same layout and design. Wonderful choice of colors! Pug Puppies for Sale Near Me PUG PUPPY FOR SALE NEAR ME PUG PUPPIES FOR SALE pug puppies for sale in kentucky Pug Puppies for Sale Under $500 Near Me pug puppies for sale in texas

rahimhh21 · 87 weeks and 3 days ago

Perfecthomepugs · 78 weeks and 1 day ago

SENSA838 selaku situs agen slot online terbaik no 1 di Indonesia akan selalu menyediakan layanan yang terbaik dan bisa memuaskan para member setia kami. Diantara nya seperti memberikan informasi mengenai daftar provider slot gacor dan terpopuler saat ini. SENSA838 Bukan hanya sekedar memberikan statement sebagai situs tergacor 2022 saat ini, kami membuktikan dengan memberikan jaminan garansi gacor 100% atau biasa di sebut garansi kekalahan 100%. Yang di mana apabila member baru deposit dan mengalami kekalahan pertamanya maka member berhak mengklaim garansi kekalahan 100% sesuai dengan syarat dan ketentuan yang berlaku. Menarik kan ? Main slot kalah dan bisa klaim kembali saldo nya. Promo ini hanya berlaku untuk member yang baru melakukan deposit untuk pertama kali nya. Cara klaim nya juga sangat mudah sekali. Member hanya perlu melapor ke CS Livechat SENSA838 yang bertugas 24 jam dimana fitur klaim garansi nya terdapat di bagian pojok kanan bawah di halaman situs kami.

loanabilly · 63 weeks and 4 days ago

pugpuppies95 · 40 weeks and 2 days ago

liltommy · 40 weeks and 1 day ago

donperi1 · 32 weeks and 3 days ago

What do you think?

Any thoughts on this command? Does it work on your machine? Can you do the same thing with only 14 characters?

You must be signed in to comment.

What's this?

commandlinefu.com is the place to record those command-line gems that you return to again and again. That way others can gain from your CLI wisdom and you from theirs too. All commands can be commented on, discussed and voted up or down.

Share Your Commands

Similar Commands

DELETE all those duplicate files but one based on md5 hash comparision in the current directory tree

Find Duplicate Files (based on size, name, and md5sum)

Find all files over a set size and displays accordingly

Find Malware in the current and sub directories by MD5 hashes

Stay in the loop…

Follow the Tweets.

Every new command is wrapped in a tweet and posted to Twitter. Following the stream is a great way of staying abreast of the latest commands. For the more discerning, there are Twitter accounts for commands that get a minimum of 3 and 10 votes - that way only the great commands get tweeted.

» http://twitter.com/commandlinefu
» http://twitter.com/commandlinefu3
» http://twitter.com/commandlinefu10

Subscribe to the feeds.

Use your favourite RSS aggregator to stay in touch with the latest commands. There are feeds mirroring the 3 Twitter streams as well as for virtually every other subset (users, tags, functions,…):

Subscribe to the feed for:

» all commands
» commands with 3 up-votes (commandlinefu3)
» commands with 10 up-votes (commandlinefu10)