A command is a list of words
A program does not receive the line you typed. It receives an array of strings, and between your keyboard and that array sits a program whose entire job is rewriting.
That program is the shell — on almost every Linux system, bash.
It is not a menu and not a text box wired to the kernel. It is an ordinary program that reads a
line, performs a fixed sequence of transformations on it, and then asks the kernel to run
whatever the first surviving word names, handing it the rest.
The sequence matters more than any individual command, so here it is up front. The shell splits your line into words on whitespace; expands anything that looks like a glob, a variable or a substitution; removes the quotes it used to decide where words ended; and only then executes. Four steps, always in that order, for every line you have ever typed.
wc was handed two filenames and has no way to tell that a glob was ever involved.
Everything amber on this page vanishes before the green line exists — that is the whole
distinction the page is built on.You can watch this happen with no theory at all. printf '[%s]\n' prints each
argument it is given on its own line, wrapped in brackets. Whatever comes out is exactly what
the shell handed over.
$printf '[%s]\n' *.txt [my report.txt] [notes.txt]$printf '[%s]\n' '*.txt' [*.txt]$printf '[%s]\n' my report.txt [my] [report.txt]← one argument became two$printf '[%s]\n' "my report.txt" [my report.txt]
printf. They are instructions to the shell about
where words end. That is the single most useful sentence on this page, and §09 is nothing but
its consequences.Nothing in this tool is simulated. Each case was run as
printf '[%s]\n' … in the demo directory described in the footer, and the words
shown are the lines that command printed. Where a case depends on which files exist, that is
said in the explanation.
The line you cannot read yet
Here is a command of the kind that makes people decide the terminal is not for them. Do not try to decode it. Just notice how much of it is amber — how much of it is not a program at all.
$find . -name '*.log' -print0 | xargs -0 grep -h ERROR | awk '{print $4}' | sort | uniq -c | sort -rn > top-offenders.txt
sort runs twice — five pipes,
one redirect, one quoted glob and two flags that exist purely because filenames can contain
spaces. There is no syntax here you will not have met by §10. The last section runs it one
stage at a time.Where you are, and what is there
The shell always has a working directory, and every relative path you type is resolved against it. Most early confusion is not knowing where you are.
pwd prints the working directory. ls lists what is in it.
cd changes it. Those three carry more traffic than everything else on this page.
$pwd /home/skydude/demo$ls archive data.csv my report.txt notes.txt scripts server.log todo.md
ls tells you the names and nothing else. It does not say
which of these are directories, how big they are or when they changed — and it hides everything
beginning with a dot. Each of those omissions has a flag.$ls -l total 28 drwxr-xr-x 2 skydude skydude 4096 Jul 22 16:40 archive -rw-r--r-- 1 skydude skydude 64 Jul 24 08:15 data.csv -rw-r--r-- 1 skydude skydude 29 Jul 25 13:07 my report.txt -rw-r--r-- 1 skydude skydude 130 Jul 28 09:30 notes.txt drwxr-xr-x 2 skydude skydude 4096 Jul 27 10:48 scripts -rw-r--r-- 1 skydude skydude 429 Jul 28 09:24 server.log -rw-r--r-- 1 skydude skydude 69 Jul 26 19:33 todo.md$ls -F archive/ data.csv my report.txt notes.txt scripts/ server.log todo.md$ls -a . .. archive data.csv my report.txt notes.txt scripts server.log todo.md
-l line is the type — d
for directory, - for an ordinary file. The next nine characters are the permission
bits, which §08 takes apart. -F is the cheap version of the same question, and
-a reveals the two entries every directory contains: itself and its parent.
(The total 28 is a count of disk blocks, not of files or bytes, so it varies
with the filesystem — this was taken on ext4. The same tree reports a smaller number on a
tmpfs.)Add -t to sort by modification time, newest first. On a directory you do not
know, that is usually the most informative single command available:
$ls -lt total 28 -rw-r--r-- 1 skydude skydude 130 Jul 28 09:30 notes.txt -rw-r--r-- 1 skydude skydude 429 Jul 28 09:24 server.log drwxr-xr-x 2 skydude skydude 4096 Jul 27 10:48 scripts -rw-r--r-- 1 skydude skydude 69 Jul 26 19:33 todo.md -rw-r--r-- 1 skydude skydude 29 Jul 25 13:07 my report.txt -rw-r--r-- 1 skydude skydude 64 Jul 24 08:15 data.csv drwxr-xr-x 2 skydude skydude 4096 Jul 22 16:40 archive
Moving around
cd takes a path. Four of them are worth memorising outright, and only one of
them is a directory name.
$cd archive; pwd /home/skydude/demo/archive$cd ..; pwd /home/skydude/demo$cd scripts; pwd; cd -; pwd /home/skydude/demo/scripts /home/skydude/demo /home/skydude/demo$cd; pwd /home/skydude
cd - printed the directory it moved to, which is why
/home/skydude/demo appears twice. .. is the parent,
- is wherever you were last, and cd with no argument is home.
The semicolons are shell syntax — they just end one command and start the next.The tilde is not a directory. ~ is expanded by the
shell into your home directory during step 2, exactly like a glob. Running
echo ~ in the demo directory printed /home/skydude — the program
received an absolute path and never saw a tilde. That is also why ~ inside
single quotes stays literal.
Two keys do more for your speed than any command. Tab completes a path — press it after a few characters and the shell fills in the rest, or beeps if it is ambiguous; press it twice to see the candidates. ↑ walks back through your history. Between them they eliminate most typing and most typos, and neither appears in any tutorial's command list because neither is a command.
Reading a file without opening it
Most of the time you do not want to read a file. You want one fact out of it, and there is a command that produces exactly that fact.
cat prints a whole file. It is the right tool only for short ones — on a
100 000-line log it will scroll past uselessly.
$cat notes.txt Shopping list for the week. TODO: ring the plumber back Bread, milk, coffee. TODO: renew the car insurance Nothing else pressing.$head -3 server.log 2026-07-28 09:14:02 INFO 10.0.0.4 GET /index.html 200 2026-07-28 09:14:05 INFO 10.0.0.9 GET /style.css 200 2026-07-28 09:15:41 WARN 10.0.0.4 GET /missing.png 404$tail -2 server.log 2026-07-28 09:21:57 INFO 10.0.0.9 GET /about.html 200 2026-07-28 09:24:30 ERROR 10.0.0.4 GET /admin 403$tail -n +7 server.log 2026-07-28 09:21:57 INFO 10.0.0.9 GET /about.html 200 2026-07-28 09:24:30 ERROR 10.0.0.4 GET /admin 403
tail -2 and tail -n +7 printed the same two lines here
by coincidence — the file is eight lines long. They ask different questions:
-2 means "the last two", -n +7 means "from line seven onward". On a
file of unknown length only the first is predictable.For anything long, use less. It pages a file without
loading all of it: Space and b move a screen at a time, /
searches forward, n repeats the search, G jumps to the end and
q quits. It is interactive, so there is no output to show you here — that is
exactly why it is not in any of the listings on this page.
Facts about a file rather than its contents
$wc -l server.log 8 server.log$wc server.log 8 56 429 server.log← lines, words, bytes$file notes.txt scripts/backup.sh archive notes.txt: ASCII text scripts/backup.sh: Bourne-Again shell script, ASCII text executable archive: directory$du -sh archive 12K archive$diff notes.txt todo.md 1,5c1,3 < Shopping list for the week. < TODO: ring the plumber back < Bread, milk, coffee. < TODO: renew the car insurance < Nothing else pressing. --- > # Todo > - [ ] TODO: write the quarterly report > - [x] book the flights
file reads the contents, not the name. It identified
backup.sh as a shell script from its first line, and would have said the same about
a file called backup with no extension at all. Linux has no concept of a file
extension — the .sh is a convention for humans.du reports disk usage, not file size. The
archive directory holds two files of 12 and 14 bytes, and du says
12K. Space is allocated in blocks — usually 4096 bytes — so a 12-byte file
occupies 4096, and the directory itself occupies another. The exact numbers here depend on the
filesystem; these were taken on ext4. On a tmpfs the same tree reports different figures.
stat is the exhaustive version, and shows something worth knowing: a file has
three timestamps, not one.
$stat notes.txt File: notes.txt Size: 130 Blocks: 8 IO Block: 4096 regular file Device: 259,9 Inode: 19018539 Links: 1 Access: (0644/-rw-r--r--) Uid: ( 1000/ skydude) Gid: ( 1000/ skydude) Access: 2026-08-01 01:31:02.184978867 -0400 Modify: 2026-07-28 09:30:00.000000000 -0400 Change: 2026-08-01 01:30:51.470929024 -0400 Birth: 2026-08-01 01:30:09.059730095 -0400
ls -l shows. The device
and inode numbers, and every timestamp except Modify, are specific to this machine and this run —
yours will differ, and that is not the page being wrong.Making things, and unmaking them
Creating files and directories is unremarkable. Deleting them is the one place on this page where a typo is permanent, so it gets its own warning rather than a footnote.
$mkdir reports$mkdir -p 2026/q3/drafts$ls -R .: 2026 reports ./2026: q3 ./2026/q3: drafts ./2026/q3/drafts: ./reports:
-p created three nested directories in one call and would have
succeeded silently had they already existed. Without it, mkdir fails if the parent
is missing — which is why -p is nearly always what a script wants.$touch reports/blank.txt$cp ~/demo/notes.txt reports/$ls reports blank.txt notes.txt$cp -r reports backup-of-reports$ls backup-of-reports blank.txt notes.txt$mv reports/blank.txt reports/scratch.txt$ls reports notes.txt scratch.txt$mv reports 2026/q3/$ls 2026 backup-of-reports
mv did two different jobs with the same syntax. Given a name that
does not exist it renames; given a directory that does, it moves into it. The trailing slash on
2026/q3/ is a habit worth keeping — it makes the intent explicit and turns a typo
into an error instead of a surprise rename.$rm 2026/q3/reports/scratch.txt$ls 2026/q3/reports notes.txt$rmdir 2026/q3/drafts$ls 2026/q3 reports$rm -r backup-of-reports$ls 2026$rmdir 2026 rmdir: failed to remove '2026': Directory not empty
rmdir removes only empty
directories, so it refuses rather than guessing. rm -r has no such scruple — it is
the command that removed a whole tree two lines earlier without a word.There is no recycle bin. rm does not move a file
anywhere; it unlinks it, and on a normal desktop filesystem the space becomes reusable
immediately. Three habits are worth building now rather than after the first accident:
Run the ls before you run the rm — if
ls *.log lists what you expect, rm *.log will delete
exactly that. Quote every expansion — rm "$file", never rm $file
— for the reason §09 demonstrates with real output, and be especially careful with
$DIR/, which becomes a bare / the moment $DIR is unset.
And treat rm -rf as a command you type deliberately, never one you reach for by
reflex.
Three streams, and the pipe between them
Every process starts life with three open channels it did not have to ask for. Redirection and pipes are nothing but the shell attaching those channels to somewhere other than your terminal.
The channels are numbered. 0 is standard input, where a program reads from. 1 is standard output, where its results go. 2 is standard error, where its complaints go. By default all three are your terminal, which is why output and errors appear interleaved and why you never think about it.
sort. Two channels exist so that diagnostics and
results can be separated by a machine, not by a human reading the screen.Redirection: pointing a channel at a file
$echo "hello" > greeting.txt$cat greeting.txt hello$echo "again" > greeting.txt$cat greeting.txt again← the first line is gone$echo "and again" >> greeting.txt$cat greeting.txt again and again
> truncates the file before the command even starts; >>
appends. The word "hello" was not overwritten by "again" — the file was emptied first, by
the shell, before echo ran. Getting these two backwards is how people lose a log
they meant to add to.Redirecting channel 2 needs its number, because > on its own means
1>. This is the mechanism behind every 2>/dev/null you have
copied off the internet.
$ls greeting.txt nosuchfile > out.txt 2> err.txt$cat out.txt greeting.txt$cat err.txt ls: cannot access 'nosuchfile': No such file or directory$ls greeting.txt nosuchfile > both.txt 2>&1$cat both.txt ls: cannot access 'nosuchfile': No such file or directory greeting.txt$ls nosuchfile 2>/dev/null; echo "exit was $?" exit was 2
ls's own.
It checks its operands before it lists anything, so the complaint about nosuchfile
genuinely comes first — on a terminal too, not just here. 2>&1 reads as "make
channel 2 go wherever channel 1 is currently going", and it must come after the
>, because it copies the destination as it stands at that moment. Written the
other way round, 2>&1 > both.txt, channel 2 is aimed at the terminal —
where channel 1 still pointed — and only then does channel 1 move to the file.< is the mirror image and is rarer than you would think.
wc -l < greeting.txt printed 2 — with no filename, because
wc was handed a stream and never learnt a name. wc -l greeting.txt
prints 2 greeting.txt. Most commands take filenames directly, so
< earns its keep mainly in scripts and in while read loops (§10).
The pipe
A pipe is redirection without a file in the middle: the kernel gives you a buffer, the left command's channel 1 writes into it, the right command's channel 0 reads out of it, and both run at the same time.
$ls | wc -l 7$grep ERROR server.log | wc -l 3
ls printed one name per line even though a bare ls in a
terminal prints columns. It checks whether channel 1 is a terminal and changes its output when
it is not — which is why the count is right. Programs that get this wrong are the reason
ls output is famously unsafe to parse.This is the whole design. Small programs that read a stream and
write a stream compose into things nobody wrote. Nothing in sort knows about log
files; nothing in grep knows about counting. The composition happens in the shell,
which knows about neither.
Finding things, and reshaping what you find
Two commands search: grep looks inside files, find looks
for files. Everything after them is turning what you found into an answer.
$grep ERROR server.log 2026-07-28 09:16:03 ERROR 10.0.0.7 POST /upload 500 2026-07-28 09:19:10 ERROR 10.0.0.7 POST /upload 500 2026-07-28 09:24:30 ERROR 10.0.0.4 GET /admin 403$grep -c ERROR server.log 3$grep -n TODO notes.txt 2:TODO: ring the plumber back 4:TODO: renew the car insurance$grep -v INFO server.log 2026-07-28 09:15:41 WARN 10.0.0.4 GET /missing.png 404 2026-07-28 09:16:03 ERROR 10.0.0.7 POST /upload 500 2026-07-28 09:19:10 ERROR 10.0.0.7 POST /upload 500 2026-07-28 09:24:30 ERROR 10.0.0.4 GET /admin 403$grep -ri todo . ./todo.md:# Todo ./todo.md:- [ ] TODO: write the quarterly report ./notes.txt:TODO: ring the plumber back ./notes.txt:TODO: renew the car insurance
-c counts instead of printing,
-n gives line numbers, -v inverts the match, -i ignores
case and -r walks a directory tree. The last command found Todo,
TODO and prefixed each hit with the file it came from, because more than one file
was searched. (-r visits files in whatever order the directory hands them back,
which is not alphabetical and not stable across filesystems — expect the same four lines in a
different order. Pipe through sort if the order matters.)The pattern is a regular expression, not a plain string — which matters the first time you search for something containing a dot or a bracket.
$grep -E "40[34]" server.log 2026-07-28 09:15:41 WARN 10.0.0.4 GET /missing.png 404 2026-07-28 09:24:30 ERROR 10.0.0.4 GET /admin 403$grep -o "10\.0\.0\.[0-9]" server.log | sort -u 10.0.0.4 10.0.0.7 10.0.0.9
-o prints only the matched text rather than the whole line, which
turns grep from a filter into an extractor. Note the backslashes: an unescaped
. matches any character, so 10.0.0.4 as a pattern would also match
10000004. It matches exactly one character, though, not any number of
them — an eight-character pattern cannot match a six-character
string.find looks for files, not inside them
$find . -name '*.log' ./archive/old.log ./archive/2024.log ./server.log$find . -type d . ./scripts ./archive$find . -name '*.log' -exec wc -l {} + 1 ./archive/old.log 1 ./archive/2024.log 8 ./server.log 10 total
'*.log' are load-bearing and this is the classic
mistake. Unquoted, the shell would expand the glob first — against the current directory
only — and hand find the single word server.log, which would then
search for files with that exact name and never look in archive/. Quoting passes
the asterisk through so that find does the matching, at every depth.
(find emits results in directory order, not sorted — yours will list the same
paths in a different sequence.)Four small tools that turn lines into answers
cut takes fields by position. sort orders lines. uniq -c
collapses adjacent duplicates and counts them. awk does the same job as
cut but smarter about whitespace. Together they answer most "how many of each"
questions.
$cut -d" " -f3 server.log | sort | uniq -c | sort -rn 4 INFO 3 ERROR 1 WARN
uniq only collapses adjacent duplicates, which is why
sort must come before it and is the single most common reason this idiom fails. The
second sort -rn then orders by the count numerically and in reverse, putting the
worst offender at the top.Now the reason to prefer awk. The log lines are not evenly spaced —
INFO and WARN are padded with an extra space to line up with
ERROR. Ask cut for the fourth space-separated field and watch what
happens:
$cut -d" " -f4 server.log 10.0.0.7 10.0.0.7 10.0.0.4$awk '{print $4}' server.log | sort | uniq -c | sort -rn 4 10.0.0.4 2 10.0.0.9 2 10.0.0.7
INFO and
WARN line is that empty field. cut is being exactly right and
completely useless. awk treats any run of whitespace as one separator, which is
what you meant. Reach for cut on a real delimiter like a comma, and
awk on anything aligned by eye.$awk -F, 'NR>1 {sum+=$3} END {print "total hours:", sum}' data.csv total hours: 153$awk -F, 'NR>1 && $2=="ops" {print $1}' data.csv alice carol$sort -t, -k3 -n data.csv name,dept,hours dan,dev,29 alice,ops,38 bob,dev,41 carol,ops,45$sed 's/ERROR/FAILURE/' server.log | head -4 2026-07-28 09:14:02 INFO 10.0.0.4 GET /index.html 200 2026-07-28 09:14:05 INFO 10.0.0.9 GET /style.css 200 2026-07-28 09:15:41 WARN 10.0.0.4 GET /missing.png 404 2026-07-28 09:16:03 FAILURE 10.0.0.7 POST /upload 500
NR>1 is how you skip a header row — NR is the
current line number, and an awk program is a list of condition/action pairs. Look at
the sort output: the header sorted to the top by luck, not by design.
sed edited a stream and left the file untouched — nothing here has modified
server.log.xargs exists because some programs do not read stdin.
rm takes filenames as arguments, not as input, so … | rm does
nothing. xargs reads a stream and turns it into arguments. The
-print0/-0 pair used with it separates names by a zero byte instead
of a newline, which is the only separator that cannot occur in a filename — the reason it
appears in the line from §01.
Processes, jobs and the status nobody prints
Running a command creates a process; the shell waits for it and then reads the number it exited with. That number is invisible, and half of shell scripting is built on it.
Put & at the end of a line and the shell does not wait. The command runs in
the background as a job, and you get your prompt back.
$sleep 300 &$sleep 400 &$jobs [1]- Running sleep 300 & [2]+ Running sleep 400 &$ps -o pid,stat,etime,cmd --ppid $$ PID STAT ELAPSED CMD 1497480 S 00:00 sleep 300 1497481 S 00:00 sleep 400 1497482 R 00:00 ps -o pid,stat,etime,cmd --ppid 1497479$kill %1$jobs [1]- Terminated sleep 300 [2]+ Running sleep 400 &
ps listed itself. It is a process too, started by the same shell,
and it appears in its own output because the snapshot is taken after it starts. The PIDs are
from one particular run on one machine and will never reproduce; $$ is the shell's
own PID, which is why the fourth column of the last row shows a number the listing does not
otherwise contain.%1 is a job number, not a process ID. The shell
translates it. kill 1497480 would have done the same thing for that one run;
kill %1 works every time without looking anything up.
A bare ps is narrower than people expect: it lists only processes attached to
your own terminal, which on a quiet prompt is a handful of lines. The listing above asks for
--ppid $$ because backgrounded jobs are children of the shell, and
ps aux is the one that shows every process on the machine — close to seven
hundred on the desktop this was written on, and a different number every time it is asked.
(Counting them with ps aux | wc -l gives one more than that: the header is
a line too.)
Three keystrokes are part of this system. Ctrl+C
asks the foreground process to stop. Ctrl+Z suspends it — it stays alive
and stopped, and fg resumes it in front, bg behind.
Ctrl+D is not a signal at all: it delivers end-of-file to whatever is
reading standard input — the terminal driver turns the next read into a zero-length one,
which is what every program treats as "no more input". Nothing is closed; your shell still has
a standard input afterwards. That is why it ends cat with no arguments and why it
logs you out of a shell.
The number nobody prints
Every process exits with a status: 0 means success, anything else is a
failure, and the shell keeps the last one in $?. Because zero is success, the
convention is backwards from every other truth value you have used.
$grep again greeting.txt; echo "exit status: $?" again and again exit status: 0$grep zebra greeting.txt; echo "exit status: $?" exit status: 1$true; echo $? 0$false; echo $? 1
grep printing nothing and grep failing are the same
event. "No match" is exit 1, deliberately, so that a script can branch on it. That is also
why a pipeline ending in grep can fail a strict script for reasons that are not
errors — §10 shows the fix.$grep -q again greeting.txt && echo "it is in there" it is in there$grep -q zebra greeting.txt || echo "no zebra here" no zebra here$grep -q zebra greeting.txt && echo "found" || echo "not found" not found
&& runs the right side only if the left succeeded;
|| only if it failed. They are not logical operators on values — they branch on
exit status. -q suppresses the matched line, so the command becomes a pure question.
This is the shell's if statement, written on one line.Permissions: nine bits and a number
Every file carries nine permission bits: read, write and execute, for the owner, the group and everyone else. The octal number you have copied off Stack Overflow is those nine bits written in binary, three at a time.
Executable is a bit, not a file type. A script that lacks it will not run no matter what is inside it — which is the single most common "but the file is right there" failure:
$ls -l hello.sh -rw-r--r-- 1 skydude skydude 42 Aug 1 01:32 hello.sh$./hello.sh /usr/bin/bash: line 1: ./hello.sh: Permission denied$chmod +x hello.sh$ls -l hello.sh -rwxr-xr-x 1 skydude skydude 42 Aug 1 01:32 hello.sh$./hello.sh the script ran
ls
lines: same size, same timestamp, three new x characters. The error's prefix names
whichever shell reported it; in an ordinary login session it reads bash: rather than
the full path shown here.Now the number. Read, write and execute are worth 4, 2 and 1. Add up the ones you want and you have one digit per audience — owner, group, other.
Owner
Group
Everyone else
$stat -c "%a %A %n" hello.sh greeting.txt 755 -rwxr-xr-x hello.sh 644 -rw-r--r-- greeting.txt$chmod u=rw,g=r,o= greeting.txt$stat -c "%a %A" greeting.txt 640 -rw-r-----$umask 0022
chmod also takes the symbolic form, which is easier to read and does not require
knowing the other six bits. umask 0022 is why new files arrive as 644 and not 666 —
it is a mask of bits to remove.chmod 777 is almost never the answer. It means
"anyone on this machine may rewrite this file", and it is reached for because it makes a
permissions error go away without diagnosing it. The question worth asking first is which user
the failing process actually runs as — whoami answers it. On this run,
whoami printed skydude and id -u printed
1000.
sudo runs one command as another user, normally root.
It is not a mode you enter. The trap worth knowing now: in
sudo echo hi > /root/f the redirection is performed by your shell,
before sudo starts, so it fails on permissions even though the command was
elevated. The > is amber on this page for exactly that reason — it belongs to
the shell, not the command.
Quoting, variables and everything that expands
Section 01 said the shell expands before it executes. This is the section where that costs you something, because expansion happens whether or not the result still makes sense.
A variable is assigned with no spaces around the =, and read back with a
$. The $ is shell syntax: it never reaches the program.
$NAME=world$echo "hello $NAME" hello world$echo 'hello $NAME' hello $NAME$echo "today the count is $(wc -l < greeting.txt)" today the count is 2$echo $HOME /home/skydude
$( ) runs a command and substitutes its output — the
shell ran wc, collected 2, and only then built the string
echo received.The failure that eats a weekend
The demo directory contains a file called my report.txt. Watch what an unquoted
variable does to it.
$FILE="my report.txt"$wc -l $FILE wc: my: No such file or directory wc: report.txt: No such file or directory 0 total$wc -l "$FILE" 1 my report.txt
$echo rm -rf "$DIR/" rm -rf /$set -u$echo rm -rf "$DIR/" /usr/bin/bash: line 1: DIR: unbound variable
echo is in front of both commands here because the point can be made without making
it twice. set -u turns the empty expansion into an error instead of a catastrophe,
and is the reason §10 puts it in the first line of every script.Everything the shell expands
| Syntax | Example | Becomes |
|---|---|---|
| Glob | *.txt | Every matching name in the directory — or, if nothing matches, the literal text *.txt |
| Any single character | file?.log | Matches exactly one character where the ? is |
| Character set | log[0-9].txt | One character from the set |
| Brace expansion | a{1,2,3}b | a1b a2b a3b — no files involved, pure text |
| Brace range | file{01..04}.txt | file01.txt … file04.txt |
| Tilde | ~/demo | Your home directory, absolute |
| Variable | $HOME | Its value, or nothing at all if unset |
| Default value | ${NAME:-none} | Its value, or none if unset or empty |
| Command substitution | $(date) | Whatever that command printed, trailing newlines removed |
| Arithmetic | $((2 + 2)) | 4 |
Every row was checked with the splitter in §01 — scroll back and step through the presets; the brace, glob and substitution cases are all there with the words they actually produced.
A glob that matches nothing is left alone. Running
printf '[%s]\n' *.zzz in the demo directory printed [*.zzz] — bash's
default is to hand the pattern through unexpanded. This is why a loop over
*.log in an empty directory runs once, with the literal string
*.log, instead of not running at all.
Writing it down
There is no separate scripting language. A shell script is the same commands in a file, and everything in the previous nine sections works unchanged inside one.
Three things make a file a script: a first line naming the interpreter, the executable bit from §08, and — for anything you intend to keep — a line of safety flags.
#!/usr/bin/env bash
set -euo pipefail
usage() {
echo "usage: $(basename "$0") FILE..." >&2
exit 2
}
[ "$#" -ge 1 ] || usage
total=0
for file in "$@"; do
if [ ! -r "$file" ]; then
echo "skipping $file — cannot read it" >&2
continue
fi
n=$(grep -c ERROR "$file" || true)
printf '%-16s %3d\n' "$file" "$n"
total=$(( total + n ))
done
echo "----------------- ---"
printf '%-16s %3d\n' TOTAL "$total"
$( ),
||, grep -c and $(( )) all appeared between §06 and §09,
and >&2 is §05's 2>&1 pointed the other way. What §10
adds is the scripting layer around them, none of which the page has shown before: a
usage function, [ ] tests, if, for,
continue, and the three ways a script sees its own invocation —
$0, $# and $@. $@ is the one worth staring
at: quoted, so that a filename with a space stays one argument, exactly as in
§09.$~/lab/errcount.sh server.log archive/old.log archive/2024.log server.log 3 archive/old.log 0 archive/2024.log 0 ----------------- --- TOTAL 3$~/lab/errcount.sh *.log server.log 3 ----------------- --- TOTAL 3$~/lab/errcount.sh usage: errcount.sh FILE...$~/lab/errcount.sh server.log /etc/shadow server.log 3 skipping /etc/shadow — cannot read it ----------------- --- TOTAL 3
-r test doing real work: /etc/shadow exists and
is unreadable to an ordinary user.set -euo pipefail is four decisions in one line, and it
is what separates a script you run twice from a script you trust. -e exits on
the first failing command instead of ploughing on. -u makes an unset variable an
error — the §09 catastrophe. -o pipefail makes a pipeline fail if any
stage failed, not just the last one. Verified: without it, false | true reports
exit 0; with it, exit 1.
And immediately, the cost. Under -e, the
grep -c in that script would kill it the first time a file contains no errors,
because "no match" is exit 1 (§07). That is what || true is doing on the end of
the line — it is not noise, it is the price of -e, and forgetting it is the most
common way a strict script dies on correct input.
The building blocks, each actually run
$for f in *.txt; do echo "[$f]"; done [my report.txt] [notes.txt]$for f in $(ls *.txt); do echo "[$f]"; done [my] [report.txt] [notes.txt]$cut -d, -f1 data.csv | while read -r name; do echo "hello $name"; done hello name hello alice hello bob hello carol hello dan
$(ls) re-splits that output on whitespace, so
my report.txt arrives as two iterations. Never loop over the output of
ls.$if [ -f server.log ]; then echo "it is a file"; fi it is a file$if [ -d archive ]; then echo "it is a directory"; fi it is a directory$n=5; if (( n > 3 )); then echo "$n is greater than 3"; fi 5 is greater than 3$greet(){ echo "hello, $1"; }; greet world; greet "everyone here" hello, world hello, everyone here$files=(*.log); echo "${#files[@]} log file(s): ${files[*]}" 1 log file(s): server.log
[ is a command, not punctuation. That is why it needs spaces around
it and why its arguments follow the same quoting rules as any other command — there is a real
program at /usr/bin/[, though bash uses its own builtin. A function takes arguments
as $1, $2 exactly as a script does, which is the whole reason functions
in bash feel like little scripts.The classic [ failure, and the classic fix.
Running x=""; if [ $x = "" ] produced
/usr/bin/bash: line 1: [: =: unary operator expected — the empty variable expanded
to nothing at all, so [ received two arguments instead of three and could not
parse them. Both [ "$x" = "" ] and bash's [[ $x = "" ]] work; the
double-bracket form does not word-split its contents, which is why it is the better default in
bash scripts and why it is not available in sh.
The cheat sheet
A reference rather than a lesson. Grouped by what you are trying to do, because that is how you will be looking. Anything marked with a bullet appears in a worked example further up the page.
Getting around §02
| Command | What it does |
|---|---|
| pwd | Print the working directory · §02 |
| ls | List names only · §02 |
| ls -l | Long form: type, permissions, owner, size, modified time · §02 |
| ls -la | Long form including dotfiles · §02 |
| ls -lt | Newest first — the best first command in an unfamiliar directory · §02 |
| ls -lh | Sizes as K/M/G rather than bytes |
| ls -F | Mark directories with /, executables with * · §02 |
| cd DIR | Change directory · §02 |
| cd .. | Up one level · §02 |
| cd - | Back to wherever you just were · §02 |
| cd | Home, with no argument at all · §02 |
| tree -L 2 | Directory tree, two levels deep (often not installed by default) |
Making and unmaking §04
| Command | What it does |
|---|---|
| touch FILE | Create it empty, or update its timestamp if it exists · §04 |
| mkdir DIR | Make a directory · §04 |
| mkdir -p a/b/c | Make parents as needed; no error if it already exists · §04 |
| cp SRC DST | Copy a file · §04 |
| cp -r SRC DST | Copy a directory and its contents · §04 |
| cp -a SRC DST | Copy preserving permissions, times and links |
| mv OLD NEW | Rename, or move into a directory if NEW is one · §04 |
| rm FILE | Delete. No recycle bin · §04 |
| rm -r DIR | Delete a directory and everything under it · §04 |
| rm -i FILE | Ask before each deletion |
| rmdir DIR | Delete only if empty — refuses otherwise · §04 |
| ln -s TARGET NAME | Make a symbolic link |
| readlink -f PATH | Resolve a path to its real absolute location |
Looking inside files §03
| Command | What it does |
|---|---|
| cat FILE | Print the whole thing — short files only · §03 |
| less FILE | Page through it. / search, n next, G end, q quit · §03 |
| head -20 FILE | First 20 lines · §03 |
| tail -20 FILE | Last 20 lines · §03 |
| tail -n +7 FILE | From line 7 to the end · §03 |
| tail -f FILE | Follow a growing file — the way to watch a live log |
| wc -l FILE | Count lines · §03 |
| wc FILE | Lines, words, bytes · §03 |
| file FILE | What kind of file it is, judged by contents · §03 |
| stat FILE | Size, permissions, owner, all three timestamps · §03 |
| du -sh DIR | Total disk usage of a directory · §03 |
| df -h | Free space per mounted filesystem |
| diff A B | Line-by-line difference between two files · §03 |
| sha256sum FILE | Checksum, for verifying a download |
Streams and redirection §05
| Syntax | What it does |
|---|---|
| cmd > FILE | stdout to a file, truncating it first · §05 |
| cmd >> FILE | stdout appended to a file · §05 |
| cmd < FILE | Read stdin from a file · §05 |
| cmd 2> FILE | stderr to a file · §05 |
| cmd > FILE 2>&1 | Both channels to one file. Order matters · §05 |
| cmd &> FILE | Bash shorthand for the same thing |
| cmd 2>/dev/null | Discard errors · §05 |
| cmd1 | cmd2 | cmd1's stdout becomes cmd2's stdin · §05 |
| cmd | tee FILE | Write to a file and pass it on |
| cmd | tee -a FILE | The same, appending |
| cmd <<< "text" | Feed one string in as stdin |
Text, pipes and reshaping §06
| Command | What it does |
|---|---|
| grep PAT FILE | Print matching lines · §06 |
| grep -i | Ignore case · §06 |
| grep -v | Invert: print what does not match · §06 |
| grep -n | Prefix line numbers · §06 |
| grep -c | Count matching lines instead of printing them · §06 |
| grep -r PAT DIR | Search a whole tree · §06 |
| grep -o PAT | Print only the matched text · §06 |
| grep -q PAT | Print nothing; use the exit status · §07 |
| grep -E PAT | Extended regular expressions · §06 |
| sort | Sort lines · §06 |
| sort -n / -r / -u | Numeric · reversed · drop duplicates · §06 |
| sort -t, -k3 -n | Comma-separated, sort on field 3, numerically · §06 |
| uniq -c | Collapse adjacent duplicates and count. Sort first · §06 |
| cut -d, -f2 | Field 2, comma-delimited · §06 |
| awk '{print $4}' | Field 4, any run of whitespace as separator · §06 |
| awk -F, 'NR>1 {…}' | Comma-separated, skipping a header row · §06 |
| sed 's/A/B/' | Replace the first A on each line · §06 |
| sed 's/A/B/g' | Replace every A on each line |
| sed -i 's/A/B/g' F | The same, editing the file in place. No undo |
| tr a-z A-Z | Translate characters · §06 |
| nl | Number the lines |
| column -t | Align whitespace-separated columns |
Finding files §06
| Command | What it does |
|---|---|
| find . -name '*.log' | By name, recursively. Quote the pattern · §06 |
| find . -iname '*.LOG' | The same, ignoring case |
| find . -type f | Files only. -type d for directories · §06 |
| find . -size +10M | Larger than 10 megabytes |
| find . -mtime -7 | Modified in the last seven days |
| find . -name X -delete | Delete every match. Run it without -delete first |
| find . -name X -exec CMD {} + | Run a command over the matches · §06 |
| find . -print0 | xargs -0 CMD | The same, safe with spaces in names · §06 |
| which CMD | Which file would run |
| type CMD | Whether it is a builtin, function, alias or file |
| command -v CMD | The portable form, for scripts |
Processes and jobs §07
| Command | What it does |
|---|---|
| ps | Processes in this session · §07 |
| ps aux | Every process on the machine |
| ps --ppid $$ | Only the children of this shell · §07 |
| pgrep -f PATTERN | PIDs whose command line matches |
| top | Live process table. htop if installed |
| cmd & | Run in the background · §07 |
| jobs | Background jobs of this shell · §07 |
| fg / bg | Resume a stopped job in front / behind · §07 |
| kill %1 | Terminate job 1 · §07 |
| kill PID | Ask a process to stop (SIGTERM) |
| kill -9 PID | Force it. Last resort — no cleanup happens |
| pkill PATTERN | Kill by name. Check with pgrep first |
| nohup cmd & | Keep running after the terminal closes |
| time cmd | How long it took |
| watch -n 5 cmd | Re-run it every five seconds |
| Ctrl+C | Interrupt the foreground process · §07 |
| Ctrl+Z | Suspend it; fg brings it back · §07 |
| Ctrl+D | Close stdin — ends cat, logs you out · §07 |
Permissions and identity §08
| Command | What it does |
|---|---|
| ls -l | The nine bits, as -rwxr-xr-x · §08 |
| stat -c "%a %A" F | The same, octal and symbolic · §08 |
| chmod +x FILE | Make it executable · §08 |
| chmod 644 FILE | Owner read/write, everyone else read · §08 |
| chmod 755 FILE | The above plus execute for all — scripts and directories · §08 |
| chmod 600 FILE | Owner only. What an SSH private key needs · §08 |
| chmod u=rw,g=r,o= F | Symbolic form; sets rather than adds · §08 |
| chmod -R 755 DIR | Recursively. Think before using it on a tree of files |
| chown USER FILE | Change owner. Needs root |
| chgrp GROUP FILE | Change group |
| umask | Which bits are removed from new files · §08 |
| whoami / id -u | Who you are, by name and by number · §08 |
| groups | Which groups you are in |
| sudo cmd | Run one command as root · §08 |
| sudo -u USER cmd | Run it as somebody else |
Archives and moving data
| Command | What it does |
|---|---|
| tar -czf out.tar.gz DIR | Create a gzipped archive of a directory |
| tar -xzf in.tar.gz | Extract one |
| tar -tzf in.tar.gz | List the contents without extracting — always do this first |
| gzip FILE / gunzip FILE.gz | Compress or decompress in place |
| zip -r out.zip DIR | Zip archive, for sending to other systems |
| unzip in.zip | Unpack one |
| scp FILE host:path | Copy over SSH |
| rsync -av SRC/ DST/ | Sync directories, copying only what changed |
| rsync -avn SRC/ DST/ | The same as a dry run. Use it every time first |
| curl -O URL | Download to a file of the same name |
| curl -s URL | Fetch quietly, to stdout — pipe it onward |
Networking
| Command | What it does |
|---|---|
| ping -c 4 HOST | Four probes and stop. Without -c it runs forever |
| curl -I URL | Response headers only |
| curl -sS -o /dev/null -w '%{http_code}\n' URL | Just the status code |
| ip addr | Interfaces and addresses. ip -brief addr for one line each |
| ip route | The routing table, including the default gateway |
| ss -tlnp | Which processes are listening on which TCP ports |
| dig +short NAME | Resolve a name. host NAME is the terser one |
| ssh user@host | Shell on another machine |
| ssh user@host 'cmd' | Run one command there and come straight back |
Shell syntax §01 · §05 · §09
| Syntax | What it means |
|---|---|
| 'single quotes' | Literal. Nothing inside expands · §09 |
| "double quotes" | Keeps it one word; $ still expands · §09 |
| \x | Escape one character |
| $VAR · ${VAR} | A variable's value · §09 |
| ${VAR:-default} | Its value, or a fallback if unset or empty · §09 |
| $(cmd) | Substitute what that command printed · §09 |
| $((a + b)) | Arithmetic · §09 |
| $? | Exit status of the last command · §07 |
| $$ | PID of the shell itself · §07 |
| $0 · $1 · $# · $@ | Script name · first argument · count · all of them · §10 |
| * ? [ ] | Glob: any text · any one character · one from a set · §09 |
| {a,b} {1..9} | Brace expansion — text, not filenames · §09 |
| ~ | Your home directory · §02 |
| a ; b | Run a, then b, regardless · §02 |
| a && b | Run b only if a succeeded · §07 |
| a || b | Run b only if a failed · §07 |
| # comment | To end of line |
| \ at end of line | Continue a long command on the next line |
Scripting §10
| Construct | What it does |
|---|---|
| #!/usr/bin/env bash | First line: which interpreter runs this file · §10 |
| set -euo pipefail | Exit on error · unset variable is an error · a failing pipe stage fails the pipeline · §10 |
| cmd || true | Deliberately ignore a failure under -e · §10 |
| for x in *.log; do … done | Loop over a glob — safe with spaces · §10 |
| while read -r line; do … done | Loop over input lines · §10 |
| if [ -f "$f" ]; then … fi | Test a file exists · §10 |
| [ -d "$d" ] · [ -r "$f" ] | Is a directory · is readable · §10 |
| [ -z "$s" ] · [ -n "$s" ] | String is empty · is not empty · §10 |
| [[ $a == b* ]] | Bash test: no word splitting, and patterns work · §10 |
| (( n > 3 )) | Arithmetic test · §10 |
| case $x in a) … ;; *) … ;; esac | Branch on a pattern |
| name() { … ; } | Define a function; arguments arrive as $1… · §10 |
| arr=(a b c) · "${arr[@]}" | An array, and every element as separate words · §10 |
| echo "msg" >&2 | Send a message to stderr, not stdout · §10 |
| exit 2 | Stop with a non-zero status · §10 |
| trap 'rm -f "$tmp"' EXIT | Run a cleanup command however the script ends |
| mktemp | Create a temporary file safely and print its name |
| bash -n script.sh | Check the syntax without running anything |
| bash -x script.sh | Trace every line as it executes — the debugger |
At the prompt
| Key | What it does |
|---|---|
| Tab | Complete a command or path; twice to list the candidates · §02 |
| ↑ / ↓ | Walk through history · §02 |
| Ctrl+R | Search history backwards as you type |
| Ctrl+A / Ctrl+E | Jump to start / end of the line |
| Ctrl+U / Ctrl+K | Delete to start / to end |
| Ctrl+W | Delete the word before the cursor |
| Ctrl+L | Clear the screen |
| !! | The previous command — sudo !! is the common use |
| !$ | The last argument of the previous command |
| history | Numbered list of what you have run |
| man CMD | The manual. / searches, q quits |
| cmd --help | Usually shorter and usually enough |
The line you could not read
This is the command from section 01, unchanged. Step through it one stage at a time and watch a pile of log lines become a ranked answer.
The pipeline
What arrives at this point
Read the finished line once more with the palette in mind. There are only four kinds of thing in it — a prompt, program words, shell syntax and one filename. The largest is the green one, and every amber character in it stops existing before any program starts:
$find . -name '*.log' -print0 | xargs -0 grep -h ERROR | awk '{print $4}' | sort | uniq -c | sort -rn > top-offenders.txt$cat top-offenders.txt 2 10.0.0.7 1 10.0.0.4
sort
runs twice. Everything amber is gone by the time the first one starts. The quotes around
'*.log' exist so that find receives the asterisk rather than the
shell's guess at it; -print0 and -0 exist because the demo directory
contains a file called my report.txt and a newline is not a safe separator; the
five pipes exist because none of these programs can do more than one
thing.You have not learnt six commands. You have learnt one idea six times. Each program reads a stream and writes a stream, knows nothing about its neighbours, and is joined to them by a shell that knows nothing about any of them. That is why the cheat sheet in §11 is a list of small things rather than a list of features: the combinations are not in the manual, and there is no ceiling on them.
The line at the top of section 01 said the shell reads first. Nothing in this final command contradicts it — the pipes, the redirect, the quotes and the glob were all consumed by bash before a single one of those six programs existed. The only thing you ever hand to Linux is a list of words.