split -b 10M -d archive.tar.gz archive.tar.gz.part- cuts a file into 10 MB pieces, and cat archive.tar.gz.part-* > restored.tar.gz puts them back. Matching sizes prove nothing — the only honest check is comparing checksums, and there is one flag that drops data without failing.
1. Cut it by size
$ clai split archive.tar.gz into 10 megabyte pieces→ split -b 10M -d archive.tar.gz archive.tar.gz.part-10485760 archive.tar.gz.part-0010485760 archive.tar.gz.part-014028480 archive.tar.gz.part-02
-b 10M is the piece size, -d asks for numeric suffixes instead of letters. The last piece is smaller, which is expected. The command itself prints nothing — the sizes above come from ls -l.
2. Put it back together
$ clai rejoin the file from its parts→ cat archive.tar.gz.part-* > restored.tar.gz
There is no special "unsplit" command: cat in glob order is the whole trick. That is exactly why the suffixes matter — the numeric and lettered ones split generates always sort into the right order.
3. Prove the result is identical
$ clai check that the rejoined file matches the original→ sha256sum archive.tar.gz restored.tar.gz93f9298124fc680c archive.tar.gz93f9298124fc680c restored.tar.gz
The hashes are shortened to 16 characters here so they fit on one line. Equal file sizes are not evidence: a flipped byte inside a piece doesn't change the size at all.
4. By count instead of by size
$ clai split the file into exactly three parts→ split -n 3 -d archive.tar.gz piece-8333334 piece-008333333 piece-018333333 piece-02
-n splits into equal parts and gives the remainder to the first one. Useful when you're spreading a file across three drives rather than fitting under an upload limit.
5. Leave a checksum for whoever rejoins it
$ clai save the checksum next to the file→ sha256sum archive.tar.gz > archive.tar.gz.sha256
After rejoining, the other side runs sha256sum -c archive.tar.gz.sha256 and gets archive.tar.gz: OK. On a corrupted file it prints archive.tar.gz: FAILED plus sha256sum: WARNING: 1 computed checksum did NOT match and exits 1 — verified by flipping a single byte inside the second piece.
6. The trap: a fixed suffix length
$ clai split the file into 200 kilobyte pieces with two digit suffixes→ split -b 200k -d -a 2 archive.tar.gz fixed-split: output file suffixes exhausted⚠ CAUTION — part of the file is missing and the exit code is still 0
Verified: -a 2 allows 100 names, the file needs 123 pieces, so split writes 100 files, prints the error and exits 0 anyway. Without -a, split extends the suffix itself — names run small-00…small-89, then small-9000 — and the rejoin order stays correct, which the checksums confirm.
Gotchas
- Equal sizes prove nothing. The only honest criterion is the checksum of the original against the checksum of the rejoined file.
-awith a fixed length loses data quietly. The error is printed but the exit code is 0, so a script happily carries on with an incomplete set of pieces. Either leave-aout or count the pieces in advance.- A piece is not an archive.
archive.tar.gz.part-00is the first third of an archive, not a small archive. Rejoin first, extract second —taron a single piece will just complain about an unexpected end of file.
Related questions
How do I split by lines instead of bytes? split -l 100000 big.csv part- cuts on line boundaries, which is what you want for text and logs.
Can split compress at the same time? No — compress first with tar czf, then split the archive. The other order produces pieces that no tool can read.
What if the other side is on Windows? They can rejoin with copy /b part-00+part-01+part-02 archive.tar.gz. The order in that list is what matters.
See also
- Extract any archive
- Resume an interrupted download with curl
- Copy files preserving the directory structure
CliAI writes the split, the rejoin and the checksum check from one sentence, and shows each command before it runs. Install it in one line.