CLI AI

Remove duplicate lines, with or without sorting

2026-08-01

sort -u file.txt removes duplicate lines and sorts what is left; awk '!seen[$0]++' file.txt removes them and keeps the original order. Which one you want depends on whether the order carries meaning — in a log it always does. uniq on its own does neither.

1. Why uniq alone doesn't help

clai
$ clai remove duplicates from access.log→ uniq access.log10.0.0.710.0.0.310.0.0.710.0.0.910.0.0.310.0.0.710.0.0.1

The file came back unchanged. uniq only collapses adjacent identical lines, and in a real file the repeats are scattered. Hence the classic sort | uniq pairing.

2. Deduplicate and sort

clai
$ clai remove duplicates and sort the file→ sort -u access.log10.0.0.110.0.0.310.0.0.710.0.0.9

sort -u does both jobs in one pass — shorter and faster than piping into uniq.

3. Deduplicate and keep the order

clai
$ clai remove duplicates but keep the original line order→ awk '!seen[$0]++' access.log10.0.0.710.0.0.310.0.0.910.0.0.1

This is the real difference from sort -u: the first occurrence of each line survives and the order is untouched, which matters for logs and anything else written in time order. The cost is memory — awk holds every distinct line, which shows on multi-gigabyte files.

4. How often each line appears

clai
$ clai count how many times each line appears→ sort access.log | uniq -c | sort -rn      3 10.0.0.7      2 10.0.0.3      1 10.0.0.9      1 10.0.0.1

The "top N" recipe for logs. uniq -c counts, the second sort -rn puts the frequent ones on top.

5. Only the lines that repeated

clai
$ clai show only lines that appear more than once→ sort access.log | uniq -d10.0.0.310.0.0.7

-d keeps only duplicates; -u does the opposite and keeps only the one-offs — 10.0.0.1 and 10.0.0.9 in this file. Handy when you're diffing two lists.

6. Deduplicate on one field

clai
$ clai keep one line per URL→ awk '!seen[$2]++' req.logGET /a 200GET /b 404

$2 is the second field. That is how you collapse a log by URL, by IP or by session id while keeping the whole first matching line.

Gotchas

  • sort -u file > file empties the file. Verified: the shell truncates the file for writing before sort ever reads it, and you end up with zero lines. Use sort -u file -o file instead — sort has a flag for writing back to its own input, and it works.
  • uniq needs sorted input. Without sort in front it quietly returns almost the same file. Not an error, just a wrong answer with no warning.
  • Order and sorting are different requirements. sort -u gives you alphabetical order, awk '!seen[$0]++' gives you the original. For logs you almost always want the second, and most articles show only the first.

Related questions

How do I find lines two files have in common? comm -12 <(sort a.txt) <(sort b.txt). Both inputs must be sorted, which is what the process substitution is for.

How do I ignore case? sort -uf compares case-insensitively; awk '!seen[tolower($0)]++' does the same while keeping the order.

What about a file too big for memory? sort -u spills to temporary files on disk and handles inputs larger than RAM; the awk version can run out of memory on the same file.

See also

CliAI picks between the sorting and the order-preserving version based on how you phrase the task, and shows the command first. Install it in one line.