| |

LFCS 18 ๐Ÿง Text Processing โ€” awk Basics

awk is a complete programming language designed for text processing. Where grep selects lines and sed transforms them, awk understands fields. It splits each input line into columns, gives each column a name, and lets you write expressions and actions that operate on those columns. A single awk command can find the lines that match a condition, extract specific fields from them, compute totals, and format the output โ€” all without a pipeline of separate tools. The LFCS exam tests awk as part of the text-processing objective, and the command’s field-oriented model is what distinguishes it from the other filters .

The name is an acronym of its authors’ surnames: Aho, Weinberger, and Kernighan. The program reads input line by line, and each line is automatically split into fields based on a field separator. The fields are referenced as $1, $2, $3, and so on. $0 refers to the entire line. NF is the number of fields in the current line, and NR is the current line number. These variables are available in every awk program, and they are the foundation of the field-oriented model .

An awk program is a sequence of pattern-action rules. A pattern is a condition; an action is what to do when the condition matches. The pattern can be a regular expression, a comparison, or a combination. The action is a block of statements enclosed in braces. If the pattern is omitted, the action applies to every line. If the action is omitted, the default action is to print the line. This pattern-action structure is what makes awk expressive: $3 > 100 { print $1 } prints the first field of every line where the third field is greater than 100 .

This chapter covers three areas. First, why awk exists and how the field model works โ€” the problem of column-oriented data and the automatic field splitting. Second, how the basic program structure works โ€” patterns, actions, the BEGIN and END blocks, and the built-in variables. Third, how to use the core features โ€” field references, field separators, comparison and regular-expression patterns, arithmetic, and formatted output with printf. The chapter ends with a complete example session, a quick reference, best practices, common pitfalls, real-world examples, and diagrams showing the field model.

Key point: awk reads input line by line and splits each line into fields. $1, $2, … are the fields; $0 is the whole line; NF is the field count; NR is the line number. A program is a sequence of pattern-action rules. BEGIN runs before input; END runs after. The default action is to print the line .


Why awk exists

The column-oriented problem. Many text files are structured as columns: /etc/passwd uses colons to separate seven fields, ps output uses whitespace to separate columns, and CSV files use commas. grep can find lines, and sed can transform them, but neither understands the concept of a field. Extracting the third column of a file with grep and sed requires a chain of commands with fragile regular expressions. awk splits the line into fields automatically and lets you reference them by number. This is the core advantage: the data is structured, and the tool understands the structure .

The computation problem. Some text-processing tasks require arithmetic. Summing a column of numbers, computing an average, or counting records that match a condition are not line-selection or line-transformation tasks; they are computations over the data. awk has variables, arithmetic operators, and control flow. It can accumulate a total, compute a percentage, and print a summary. This makes it a programming language, not just a filter .

The formatting problem. awk‘s printf function provides formatted output: field widths, decimal places, alignment, and padding. A report that needs columns aligned to a fixed width is a natural awk task. The print statement is the simple form, and printf is the precise form .

The pattern-action problem. An awk program is a set of rules, each with a pattern and an action. The pattern determines when the action runs, and the action determines what happens. This is a declarative style: the programmer describes the conditions and the results, and awk handles the iteration. A rule like /error/ { count++ } counts the lines matching error, and the END block prints the count. The pattern-action structure is what makes awk concise and expressive .

The one-liner problem. awk is often used as a one-liner in a pipeline. ps aux | awk '$3 > 50 { print $2, $11 }' prints the PID and command of every process using more than 50% CPU. The command is short, readable, and does exactly what it says. This is the use case that made awk famous: complex field-based filtering in a single command.

The trade-off. awk is a programming language, and with that power comes complexity. A short one-liner is easy to read, but a long awk program can become difficult to maintain. The syntax is different from shell scripting, and the quoting rules interact with the shell. The language has limitations: it is not designed for binary data, it does not handle nested structures, and its regular expressions are less powerful than those in modern languages. The trade-off is between the expressiveness of a field-aware language and the simplicity of a line-oriented filter. For column-oriented data, awk is the right tool.


a. Program structure and built-in variables

An awk program consists of a sequence of pattern-action rules. The basic invocation is awk 'program' file or awk -f script.awk file .

awk '{ print $1 }' file.txt

The program is quoted to prevent the shell from interpreting the dollar signs and braces. The braces enclose the action, which runs for every line. The pattern is omitted, so the action applies to all lines.

The BEGIN and END blocks are special patterns that run before and after the input:

awk 'BEGIN { print "Start" } { print $1 } END { print "End" }' file.txt

The BEGIN block runs once, before any input is read. The main rules run for each line. The END block runs once, after all input is processed. The BEGIN block is used for initialization, such as setting the field separator or printing a header. The END block is used for summaries and totals .

The built-in variables are available throughout the program:

  • $0 โ€” the entire current line.
  • $1, $2, …, $n โ€” the fields of the current line.
  • NF โ€” the number of fields in the current line.
  • NR โ€” the total number of records (lines) read so far.
  • FNR โ€” the number of records read from the current file.
  • FS โ€” the field separator. The default is whitespace (spaces and tabs).
  • OFS โ€” the output field separator. The default is a space.
  • ORS โ€” the output record separator. The default is a newline .
  • RS โ€” the record separator. The default is a newline.

The FS variable can be set in the BEGIN block or with the -F option. The -F option is the most common way:

awk -F: '{ print $1 }' /etc/passwd

This sets the field separator to a colon and prints the first field (the username) of each line.


b. Patterns and actions

A pattern determines which lines the action applies to. An action is a block of statements that runs when the pattern matches. If the pattern is omitted, the action runs for every line. If the action is omitted, the default action is { print } .

Regular-expression patterns:

awk '/error/ { print $0 }' file.txt

The pattern /error/ matches lines containing error. The action prints the line. This is equivalent to grep error file.txt.

Comparison patterns:

awk '$3 > 100 { print $1, $2 }' file.txt

The pattern $3 > 100 matches lines where the third field is numerically greater than 100. The action prints the first two fields. This is the kind of field-aware condition that grep cannot express.

Compound patterns:

awk '$3 > 100 && $4 == "active" { print $1 }' file.txt

The pattern combines two conditions with &&. awk supports && (and), || (or), and ! (not).

Range patterns:

awk '/start/, /end/ { print }' file.txt

The pattern /start/, /end/ matches from the line containing start to the line containing end. This is the same range syntax as sed.

BEGIN and END:

awk 'BEGIN { FS=":"; print "Users:" } { print $1 } END { print NR " total" }' /etc/passwd

The BEGIN block sets the field separator and prints a header. The main rule prints the first field of each line. The END block prints the total number of lines.

Actions:

An action can contain multiple statements, separated by semicolons or newlines. It can declare variables, perform arithmetic, use control flow, and call functions. The print statement writes to standard output, and the printf statement writes formatted output .

awk '{ total += $1 } END { print "Sum:", total }' numbers.txt

This accumulates the sum of the first field in the total variable and prints it in the END block. The variable total is automatically initialized to 0 on first use.


c. Fields, separators, and formatted output

The field references are the core of awk. $1 is the first field, $2 is the second, and $NF is the last field. The expression $(NF-1) is the second-to-last field .

awk '{ print $1, $NF }' file.txt

This prints the first and last fields of each line.

Field separators:

The default field separator is whitespace, which means one or more spaces or tabs. The -F option sets a different separator. For a single character, the option is -F:. For a regular expression, the option is -F'[[:space:]]+' .

awk -F: '{ print $1, $3 }' /etc/passwd
awk -F',' '{ print $2 }' data.csv

Output field separator:

The OFS variable determines what separates the fields in the output of print. The default is a space. Setting OFS changes the output format without changing the input parsing .

awk 'BEGIN { FS=":"; OFS=" - " } { print $1, $3 }' /etc/passwd

This reads colon-separated fields and prints them separated by -.

Formatted output with printf:

The printf statement provides formatted output. The format string contains placeholders: %s for strings, %d for integers, %f for floating-point numbers, and %5.2f for a float with a minimum width of 5 and 2 decimal places .

awk '{ printf "%-10s %5d\n", $1, $2 }' file.txt

The %-10s left-aligns a string in a field of 10 characters. The %5d right-aligns an integer in a field of 5 characters. The \n is the newline. This is the standard way to produce aligned columnar output.

Arithmetic:

awk supports the standard arithmetic operators: +, -, *, /, % (modulo), and ^ (exponent). It also has assignment operators like +=, -=, and ++ .

awk '{ total += $3 } END { print "Average:", total / NR }' data.txt

This computes the average of the third field across all lines. The NR variable is the number of lines, so total / NR is the average.

Conditional statements:

awk supports if, else, and while:

awk '{ if ($3 > 100) print "High:", $1; else print "Low:", $1 }' file.txt

This is the pattern-action structure extended with conditional logic. The action contains an if statement that branches based on the value of the third field.


Complete Example Session

# ============================================
# PART 1: PRINT THE FIRST FIELD
# ============================================

awk '{ print $1 }' file.txt

# Each line is split into fields.
# $1 is the first field.


# ============================================
# PART 2: COLON-SEPARATED FIELDS
# ============================================

awk -F: '{ print $1, $3 }' /etc/passwd

# -F: sets the field separator to a colon.
# $1 is the username, $3 is the UID.


# ============================================
# PART 3: PRINT THE LAST FIELD
# ============================================

awk '{ print $NF }' file.txt

# $NF is the last field, regardless of how many fields there are.


# ============================================
# PART 4: FILTER BY CONDITION
# ============================================

awk '$3 > 100 { print $1 }' data.txt

# The pattern $3 > 100 matches lines where the third field is > 100.


# ============================================
# PART 5: REGULAR EXPRESSION PATTERN
# ============================================

awk '/error/ { print $0 }' log.txt

# /error/ matches lines containing "error".
# This is equivalent to grep error log.txt.


# ============================================
# PART 6: BEGIN AND END BLOCKS
# ============================================

awk 'BEGIN { print "Users:" } { print $1 } END { print NR " total" }' /etc/passwd

# BEGIN runs before input, END runs after.
# NR is the total number of lines.


# ============================================
# PART 7: ACCUMULATE A TOTAL
# ============================================

awk '{ total += $3 } END { print "Total:", total }' data.txt

# total accumulates the third field.
# The variable is initialized to 0 automatically.


# ============================================
# PART 8: COMPUTE AN AVERAGE
# ============================================

awk '{ total += $3 } END { print "Average:", total / NR }' data.txt


# ============================================
# PART 9: FORMATTED OUTPUT
# ============================================

awk '{ printf "%-10s %5d\n", $1, $2 }' file.txt

# %-10s: left-aligned string, width 10
# %5d: right-aligned integer, width 5


# ============================================
# PART 10: THE LFCS PIPELINE
# ============================================

# Find processes using more than 50% CPU
ps aux | awk '$3 > 50 { print $2, $11 }'

# Sum a column of numbers
awk '{ sum += $1 } END { print sum }' numbers.txt

# Extract usernames with UID >= 1000
awk -F: '$3 >= 1000 { print $1 }' /etc/passwd

# Count lines matching a pattern
awk '/error/ { count++ } END { print count }' log.txt

# Print the last field of every line
awk '{ print $NF }' file.txt

# Print lines where the second field is "active"
awk '$2 == "active" { print $1 }' file.txt

# Print a report with headers
awk 'BEGIN { print "USER\tUID" } { print $1 "\t" $3 }' /etc/passwd

The ten parts show printing the first field, colon-separated fields, the last field, filtering by condition, regular expression patterns, BEGIN and END blocks, accumulating a total, computing an average, formatted output, and the LFCS pipeline scenarios.


Quick Reference

Built-in Variables

VariableMeaning
$0Entire current line
$1, $2, …Fields 1, 2, …
$NFLast field
NFNumber of fields in current line
NRTotal records (lines) read
FNRRecords read from current file
FSField separator (default whitespace)
OFSOutput field separator (default space)
ORSOutput record separator (default newline)
RSRecord separator (default newline)

Special Patterns

PatternMeaning
BEGINRuns before any input
ENDRuns after all input
/regex/Lines matching regex
$3 > 100Lines where field 3 > 100
condition1 && condition2Both conditions
condition1 || condition2Either condition
/start/, /end/Range from start to end

Field Separators

CommandEffect
awk -F: '{...}'Colon separator
awk -F',' '{...}'Comma separator
awk -F'\t' '{...}'Tab separator
BEGIN { FS=":" }Set FS in BEGIN
BEGIN { OFS=" - " }Set output separator

Common Actions

ActionEffect
print $1Print field 1
print $1, $2Print fields 1 and 2, space-separated
print $1, $2 with OFSFields separated by OFS
printf "%-10s %5d\n", $1, $2Formatted output
total += $3Accumulate
count++Increment
if ($3 > 100) ... else ...Conditional

Comparison Operators

OperatorMeaning
==Equal
!=Not equal
<Less than
>Greater than
<=Less than or equal
>=Greater than or equal
~Matches regex
!~Does not match regex

Best Practices

โœ… Do This:

# Use -F for field separators
awk -F: '{ print $1 }' /etc/passwd                                          # โœ…
# Use BEGIN for headers and initialization
awk 'BEGIN { FS=":"; OFS="\t" } { print $1, $3 }' /etc/passwd              # โœ…
# Use END for summaries
awk '{ sum += $1 } END { print sum }' numbers.txt                          # โœ…
# Use printf for aligned output
awk '{ printf "%-20s %5d\n", $1, $2 }' file.txt                            # โœ…
# Quote the program
awk '{ print $1 }' file.txt                                                 # โœ…
# Use comparison for field filtering
awk '$3 > 100 { print $1 }' file.txt                                       # โœ…

โŒ Don’t Do This:

# Don't forget to quote the program
awk { print $1 } file.txt  # shell interprets braces and $1                  # โŒ
# Don't use $0 when a specific field is needed
awk '{ print $0 }' file.txt  # prints the whole line                       # โŒ
# Don't hardcode the last field number
awk '{ print $7 }' file.txt  # use $NF instead                             # โŒ
# Don't forget OFS when output needs a specific separator
awk 'BEGIN { FS=":" } { print $1, $3 }' # output is space-separated        # โŒ
# Don't use awk for simple line selection
grep "error" file.txt  # simpler than awk '/error/' file.txt               # โŒ

Common Pitfalls

PitfallWhy It HappensFix
Fields not splitting correctlyWrong FSSet -F or FS in BEGIN
Numbers compared as stringsMissing numeric contextUse $3 + 0 > 100 or ensure numeric
Output not alignedUsing print instead of printfUse printf with format specifiers
$1 empty in shellSingle quotes neededUse single quotes around the program
NR vs FNR confusionMultiple filesNR is total, FNR is per-file
OFS not appliedSet after first printSet in BEGIN
Regex not matchingSpecial charactersEscape or use character classes
Last field number variesHardcoded field indexUse $NF

Real-World Examples

1. Print First Field

awk '{ print $1 }' file.txt

2. Colon-Separated

awk -F: '{ print $1, $3 }' /etc/passwd

3. Filter by Numeric Condition

awk '$3 > 100 { print $1 }' data.txt

4. Regular Expression Match

awk '/error/ { print $0 }' log.txt

5. Sum a Column

awk '{ sum += $1 } END { print sum }' numbers.txt

6. Average

awk '{ sum += $1 } END { print sum / NR }' numbers.txt

7. Count Matches

awk '/error/ { count++ } END { print count }' log.txt

8. Formatted Report

awk '{ printf "%-10s %5d\n", $1, $2 }' file.txt

9. Processes Over 50% CPU

ps aux | awk '$3 > 50 { print $2, $11 }'

10. Usernames with UID >= 1000

awk -F: '$3 >= 1000 { print $1 }' /etc/passwd

Visual

The Field Model

โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚  THE FIELD MODEL                                             โ”‚
โ”‚                                                              โ”‚
โ”‚  Input line:                                                 โ”‚
โ”‚  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”    โ”‚
โ”‚  โ”‚  alice:1001:1001:Alice Smith:/home/alice:/bin/bash   โ”‚    โ”‚
โ”‚  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜    โ”‚
โ”‚         โ”‚      โ”‚      โ”‚        โ”‚           โ”‚         โ”‚       โ”‚
โ”‚         โ–ผ      โ–ผ      โ–ผ        โ–ผ           โ–ผ         โ–ผ       โ”‚
โ”‚  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”    โ”‚
โ”‚  โ”‚  $1     $2     $3     $4          $5        $6       โ”‚    โ”‚
โ”‚  โ”‚  alice  1001   1001   Alice Smith /home/... /bin/bashโ”‚    โ”‚
โ”‚  โ”‚                                                      โ”‚    โ”‚
โ”‚  โ”‚  FS = ":"                                            โ”‚    โ”‚
โ”‚  โ”‚  NF = 6 (number of fields)                           โ”‚    โ”‚
โ”‚  โ”‚  $0 = the entire line                                โ”‚    โ”‚
โ”‚  โ”‚  $NF = $6 = /bin/bash                                โ”‚    โ”‚
โ”‚  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜    โ”‚
โ”‚                                                              โ”‚
โ”‚  awk splits the line into fields automatically.              โ”‚
โ”‚  The fields are referenced as $1, $2, etc.                   โ”‚
โ”‚                                                              โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜

Pattern-Action Structure

โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚  PATTERN-ACTION STRUCTURE                                    โ”‚
โ”‚                                                              โ”‚
โ”‚  awk 'PATTERN { ACTION }' file                               โ”‚
โ”‚                                                              โ”‚
โ”‚  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”    โ”‚
โ”‚  โ”‚  /error/ { print $1 }                                โ”‚    โ”‚
โ”‚  โ”‚    โ”‚           โ”‚                                     โ”‚    โ”‚
โ”‚  โ”‚    โ”‚           โ””โ”€โ”€ ACTION: runs when pattern matches โ”‚    โ”‚
โ”‚  โ”‚    โ”‚                                                 โ”‚    โ”‚
โ”‚  โ”‚    โ””โ”€โ”€ PATTERN: condition for the action             โ”‚    โ”‚
โ”‚  โ”‚                                                       โ”‚    โ”‚
โ”‚  โ”‚  If PATTERN is omitted: action runs for every line   โ”‚    โ”‚
โ”‚  โ”‚  If ACTION is omitted: default is { print }          โ”‚    โ”‚
โ”‚  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜    โ”‚
โ”‚                                                              โ”‚
โ”‚  Examples:                                                   โ”‚
โ”‚  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”    โ”‚
โ”‚  โ”‚  awk '{ print $1 }'          โ†’ every line            โ”‚    โ”‚
โ”‚  โ”‚  awk '/error/'               โ†’ print matching lines  โ”‚    โ”‚
โ”‚  โ”‚  awk '$3 > 100 { print $1 }' โ†’ conditional action    โ”‚    โ”‚
โ”‚  โ”‚  awk 'END { print NR }'      โ†’ after input           โ”‚    โ”‚
โ”‚  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜    โ”‚
โ”‚                                                              โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜

BEGIN and END Flow

โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚  BEGIN AND END FLOW                                          โ”‚
โ”‚                                                              โ”‚
โ”‚  BEGIN { ... }                                               โ”‚
โ”‚    โ”‚                                                         โ”‚
โ”‚    โ”‚  Runs once, before any input                            โ”‚
โ”‚    โ”‚  Used for: FS, OFS, headers, initialization             โ”‚
โ”‚    โ–ผ                                                         โ”‚
โ”‚  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”    โ”‚
โ”‚  โ”‚  Main rules { ... }                                  โ”‚    โ”‚
โ”‚  โ”‚                                                       โ”‚    โ”‚
โ”‚  โ”‚  Run once for each input line                        โ”‚    โ”‚
โ”‚  โ”‚  Used for: field processing, filtering, accumulation โ”‚    โ”‚
โ”‚  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜    โ”‚
โ”‚    โ”‚                                                         โ”‚
โ”‚    โ–ผ                                                         โ”‚
โ”‚  END { ... }                                                 โ”‚
โ”‚    โ”‚                                                         โ”‚
โ”‚    โ”‚  Runs once, after all input                             โ”‚
โ”‚    โ”‚  Used for: summaries, totals, averages                  โ”‚
โ”‚    โ–ผ                                                         โ”‚
โ”‚  Output                                                      โ”‚
โ”‚                                                              โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜

Field Separators

โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚  FIELD SEPARATORS                                            โ”‚
โ”‚                                                              โ”‚
โ”‚  Default (whitespace):                                       โ”‚
โ”‚  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”    โ”‚
โ”‚  โ”‚  alice  1001  1001  Alice Smith                      โ”‚    โ”‚
โ”‚  โ”‚  $1     $2    $3    $4                               โ”‚    โ”‚
โ”‚  โ”‚                                                       โ”‚    โ”‚
โ”‚  โ”‚  Spaces and tabs separate fields.                    โ”‚    โ”‚
โ”‚  โ”‚  Multiple spaces count as one separator.             โ”‚    โ”‚
โ”‚  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜    โ”‚
โ”‚                                                              โ”‚
โ”‚  Colon separator (-F:):                                      โ”‚
โ”‚  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”    โ”‚
โ”‚  โ”‚  alice:1001:1001:Alice Smith:/home/alice:/bin/bash   โ”‚    โ”‚
โ”‚  โ”‚  $1    $2   $3   $4          $5          $6          โ”‚    โ”‚
โ”‚  โ”‚                                                       โ”‚    โ”‚
โ”‚  โ”‚  FS = ":"                                            โ”‚    โ”‚
โ”‚  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜    โ”‚
โ”‚                                                              โ”‚
โ”‚  Comma separator (-F,):                                      โ”‚
โ”‚  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”    โ”‚
โ”‚  โ”‚  alice,1001,Alice Smith                              โ”‚    โ”‚
โ”‚  โ”‚  $1    $2   $3                                       โ”‚    โ”‚
โ”‚  โ”‚                                                       โ”‚    โ”‚
โ”‚  โ”‚  FS = ","                                            โ”‚    โ”‚
โ”‚  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜    โ”‚
โ”‚                                                              โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜

Summary

ItemValue
awkField-oriented text-processing language
Basic syntaxawk 'program' file
Fields$1, $2, …, $NF
Whole line$0
Field countNF
Line numberNR
Field separatorFS, set with -F
Output separatorOFS
Special patternsBEGIN, END
Default action{ print }
Formatted outputprintf

Key takeaways:

  • awk is field-oriented. It splits each input line into fields automatically, and the fields are referenced as $1, $2, and so on. $0 is the whole line, $NF is the last field, and NF is the number of fields. This is what distinguishes awk from grep and sed .
  • The program is a sequence of pattern-action rules. A pattern determines when the action runs, and the action determines what happens. If the pattern is omitted, the action runs for every line. If the action is omitted, the default is to print the line .
  • The BEGIN and END blocks run before and after the input. BEGIN is used for initialization, such as setting FS or printing a header. END is used for summaries, totals, and averages .
  • The field separator is set with -F or FS. The default is whitespace, which means one or more spaces or tabs. For colon-separated files like /etc/passwd, use -F:. For CSV, use -F,. The OFS variable controls the output separator .
  • Comparison and regular-expression patterns filter lines. $3 > 100 matches lines where the third field is greater than 100. /error/ matches lines containing error. The conditions can be combined with &&, ||, and ! .
  • awk performs arithmetic. Variables are automatically initialized to 0, and the standard operators (+, -, *, /, %, ^) are available. The END block is the place to print totals and averages .
  • printf provides formatted output. The format string uses %s for strings, %d for integers, and %f for floats. Width and precision specifiers like %-10s and %5.2f align the output into columns .
  • awk is often used as a one-liner in a pipeline. ps aux | awk '$3 > 50 { print $2, $11 }' is a complete field-based filter in a single command. The command is composable, and it reads from standard input when no file is given .

Remember: awk is the standard tool for column-oriented text processing. It reads input line by line, splits each line into fields, and applies pattern-action rules. The fields are the core abstraction: $1, $2, and so on. The BEGIN and END blocks provide initialization and summarization. The -F option sets the field separator, and printf provides formatted output. For the LFCS exam, the key patterns are extracting fields, filtering by field conditions, summing columns, and printing formatted reports. The command is a complete programming language, but the common cases are one-liners, and the one-liners are what make awk indispensable.



Stop using slow, ad-bloated tool sites! ๐Ÿคฎ

๐Ÿ”Ž Search “KandZ Tools” on Google to use many professional utilities for free.

KandZ.me is the ultimate minimalist hub for:
โœ… Finance (Mortgage, Interest, Inflation)
โœ… Tech (Base64, JSON, Dev Suite, IP)
โœ… Health (BMI, BMR, TDEE)
โœ… Productivity (Timer, Workspace, QR)

โšก๏ธ Fast & Private
๐Ÿ”’ No data leaves your device
๐Ÿ’Ž 100% Free

๐Ÿ”— Use it now: https://tools.kandz.me
๐Ÿ”– Bookmark itโ€”youโ€™ll need it later!