Is it safe to let AI run terminal commands? Yes, on your own machine, in a repo you can revert, and only after you've read what's about to run — no, the moment a tool pipes model output straight into a shell with no confirmation step. Everything below is the threat model behind that answer.
1. What exactly you're handing over
An agent with shell access can do anything your user account can do: read environment variables — and the tokens sitting inside them; read ~/.ssh and ~/.aws/credentials; rewrite configs; install services that survive a reboot. This isn't a hypothetical about a rogue model. One mistaken command is enough.
2. The difference between "suggest" and "execute"
$ clai remove all stopped containers and unused images→ docker system prune -a⚠ CAUTION — removes every image not used by a running containerRun it? type yes to confirm:
Generating and running are two separate steps. A tool that pipes model output straight into the shell has removed the one barrier standing between a model's typo and your disk.
3. Three risk levels worth telling apart
SAFE — read-only commands: ls, ps, df, git log. A mistake costs time.
CAUTION — commands that change state: docker prune, git reset, systemctl restart, chmod -R. A mistake costs work.
DANGER — irreversible ones: rm -rf, dd, mkfs, DROP TABLE, force push. A mistake costs data.
The label itself isn't what helps — it's that it forces a pause exactly where the cost of a mistake jumps by an order of magnitude.
4. Ask for a dry run
$ clai sync this folder to prod, but show me what would happen first→ rsync -av --dry-run ./build/ deploy@prod:/var/www/
Almost every destructive operation has a "show only" mode: rsync --dry-run, git clean -n, find without -delete. Ask for it first whenever production is involved.
5. Limit the blast radius, not the trust
A separate user for agent tasks. Tokens with minimal scope and a short lifetime. Production touched only through a separate, reviewed channel. These are the same rules you'd give a new hire on their first day, and they work for the same reason: they don't depend on how good the model is.
6. Check what actually ran
$ clai show me the last 20 commands I ran through clai→ cliai log --tail 20
A local request log gives you a postmortem. If something goes wrong, you can see exactly which phrasing produced which command.
Gotchas
- The dangerous part isn't only the destructive command — it's the leak. An agent that dumps environment variables while debugging just put your tokens in a log, and that log may leave the machine.
- "I'll check it later" doesn't work. You read the command before Enter. After Enter, you're reading the consequences.
- Autonomous mode on your own machine and on production are different decisions. In a disposable sandbox or a git repo, auto-run is reasonable. On a server holding real data, it isn't.
Related questions
Is it okay to give AI terminal access? For reversible work — a git repo, a sandbox, your own machine — yes, as long as it confirms before running. For production, only the "show me, I'll decide" mode.
How is this different from copy-pasting from Stack Overflow? In principle, it isn't — and that's the point. You read a command from the internet before running it too. The only difference is this one's generated for your system.
What if the wrong command still runs? Stop the process, check the local request log, assess the damage. That's exactly why irreversible operations require typed confirmation, not a single keypress.
See also
- SAFE, CAUTION, DANGER: how CliAI classifies what's about to run
- Why a CLI, not a chat: the design philosophy of CliAI
- Natural language to shell command: how it actually works
CliAI shows you the command and waits for confirmation before anything destructive runs — get started here.