sort -u file.txt removes duplicate lines and sorts what is left; awk '!seen[$0]++' file.txt removes them and keeps the original order. Which one you want depends on whether the order carries meaning — in a log it always does. uniq on its own does neither.
1. Why uniq alone doesn't help
$ clai remove duplicates from access.log→ uniq access.log10.0.0.710.0.0.310.0.0.710.0.0.910.0.0.310.0.0.710.0.0.1
The file came back unchanged. uniq only collapses adjacent identical lines, and in a real file the repeats are scattered. Hence the classic sort | uniq pairing.
2. Deduplicate and sort
$ clai remove duplicates and sort the file→ sort -u access.log10.0.0.110.0.0.310.0.0.710.0.0.9
sort -u does both jobs in one pass — shorter and faster than piping into uniq.
3. Deduplicate and keep the order
$ clai remove duplicates but keep the original line order→ awk '!seen[$0]++' access.log10.0.0.710.0.0.310.0.0.910.0.0.1
This is the real difference from sort -u: the first occurrence of each line survives and the order is untouched, which matters for logs and anything else written in time order. The cost is memory — awk holds every distinct line, which shows on multi-gigabyte files.
4. How often each line appears
$ clai count how many times each line appears→ sort access.log | uniq -c | sort -rn 3 10.0.0.7 2 10.0.0.3 1 10.0.0.9 1 10.0.0.1
The "top N" recipe for logs. uniq -c counts, the second sort -rn puts the frequent ones on top.
5. Only the lines that repeated
$ clai show only lines that appear more than once→ sort access.log | uniq -d10.0.0.310.0.0.7
-d keeps only duplicates; -u does the opposite and keeps only the one-offs — 10.0.0.1 and 10.0.0.9 in this file. Handy when you're diffing two lists.
6. Deduplicate on one field
$ clai keep one line per URL→ awk '!seen[$2]++' req.logGET /a 200GET /b 404
$2 is the second field. That is how you collapse a log by URL, by IP or by session id while keeping the whole first matching line.
Gotchas
sort -u file > fileempties the file. Verified: the shell truncates the file for writing beforesortever reads it, and you end up with zero lines. Usesort -u file -o fileinstead — sort has a flag for writing back to its own input, and it works.uniqneeds sorted input. Withoutsortin front it quietly returns almost the same file. Not an error, just a wrong answer with no warning.- Order and sorting are different requirements.
sort -ugives you alphabetical order,awk '!seen[$0]++'gives you the original. For logs you almost always want the second, and most articles show only the first.
Related questions
How do I find lines two files have in common? comm -12 <(sort a.txt) <(sort b.txt). Both inputs must be sorted, which is what the process substitution is for.
How do I ignore case? sort -uf compares case-insensitively; awk '!seen[tolower($0)]++' does the same while keeping the order.
What about a file too big for memory? sort -u spills to temporary files on disk and handles inputs larger than RAM; the awk version can run out of memory on the same file.
See also
- Extract CSV columns with awk
- Find and replace across files
- Count files in a directory, recursively and by type
CliAI picks between the sorting and the order-preserving version based on how you phrase the task, and shows the command first. Install it in one line.