System
Administration
CO NT E NT S
01
Operating
Systems
&
Kernel
02
Scripting
&
Automation
03
Networking
&
DNS
04
Identity
&
Directory
Services
05
Virtualization
&
Containers
06
Cloud
Infrastructure
07
Storage
,
Filesystems
&
Backup
08
Configuration
Management
&
IaC
09
Monitoring
,
Logging
&
Observability
10
Security
,
Hardening
&
Compliance
11
Application
,
Web
&
Database
Servers
12
Performance
,
Availability
&
DR
13
IT
Service
Management
&
Operations
14
Adjacent
&
Emerging
Practice
01
Operating
Systems
&
Kernel
Day
-
to
-
day
fluency
across
the
server
platforms
that
carry
production
workloads
—
from
boot
loader
to
kernel
-
level
tuning
.
Linux
distributions
RHEL
9,
Ubuntu
24.04
LTS
,
Debian
12,
SUSE
SLES
15
Windows
Server
Server
2019 / 2022,
Server
Core
,
roles
and
features
,
licensing
Boot
and
recovery
GRUB
2,
initramfs
,
rescue
mode
,
single
-
user
,
PXE
netboot
Kernel
tuning
sysctl
, /
proc
and
/
sys
,
ulimits
,
huge
pages
,
I
/
O
schedulers
systemd
Units
,
targets
,
timers
,
journald
,
cgroup
v
2
resource
limits
Package
management
dnf
,
apt
,
zypper
,
RPM
spec
files
,
patch
baselines
Process
and
resource
control
ps
,
top
,
nice
/
ionice
,
cgroups
,
OOM
handling
Filesystem
hierarchy
and
permissions
FHS
,
POSIX
permissions
,
ACLs
,
setuid
/
setgid
,
umask
Unix
shell
environments
bash
,
zsh
,
profile
.
d
,
PATH
hygiene
,
SSH
agent
forwarding
macOS
and
BSD
administration
Homebrew
,
launchd
,
plist
management
,
MDM
enrollment
02
Scripting
&
Automation
Writing
the
small
programs
that
remove
repetitive
work
,
and
the
pipelines
that
run
them
unattended
.
Bash
POSIX
-
safe
scripts
,
set
-
euo
pipefail
,
error
traps
,
process
substitution
PowerShell
7.
x
,
cmdlets
,
remoting
,
DSC
,
module
packaging
Python
3.12,
argparse
,
subprocess
,
boto
3,
paramiko
,
requests
Go
Static
cross
-
compiled
operational
CLIs
,
single
-
binary
deployment
1 / 7
Perl
,
Ruby
,
Node
.
js
Legacy
automation
,
Ansible
modules
,
tooling
glue
Version
control
Git
branching
and
rebase
,
hooks
,
GitHub
/
GitLab
review
flow
CI
/
CD
pipelines
GitLab
CI
,
GitHub
Actions
,
Jenkins
,
Argo
CD
API
integration
REST
,
webhooks
,
OAuth
2.0
tokens
,
JSON
and
YAML
parsing
Scheduled
execution
cron
,
systemd
timers
,
Task
Scheduler
,
Airflow
DAGs
Secrets
in
automation
Vault
,
SOPS
,
AWS
Secrets
Manager
,
no
credentials
in
Git
03
Networking
&
DNS
Configuring
and
diagnosing
the
layer
that
everything
else
depends
on
,
from
host
firewalls
to
name
resolution
.
TCP
/
IP
fundamentals
Subnetting
,
CIDR
,
routing
tables
,
MTU
/
MSS
,
ARP
Packet
analysis
tcpdump
,
Wireshark
,
ss
/
netstat
,
mtr
,
traceroute
Host
firewalls
nftables
/
iptables
,
firewalld
,
UFW
,
Defender
Firewall
Load
balancing
HAProxy
,
NGINX
,
F
5
LTM
,
AWS
ALB
/
NLB
,
health
checks
DNS
services
BIND
9,
Unbound
,
dnsmasq
,
Route
53,
zone
files
,
DNSSEC
Name
resolution
design
Split
-
horizon
,
conditional
forwarders
,
record
hygiene
,
TTL
policy
DHCP
and
IP
address
management
ISC
Kea
,
Infoblox
,
scopes
,
reservations
,
subnet
planning
Remote
access
WireGuard
,
OpenVPN
,
IPsec
,
ZTNA
,
bastion
hosts
Switching
and
routing
VLANs
, 802.1
Q
trunking
,
LACP
,
OSPF
and
BGP
fundamentals
Supporting
network
services
chrony
/
NTP
,
syslog
relay
,
SNMP
polling
and
traps
,
TFTP
04
Identity
&
Directory
Services
Owning
the
source
of
truth
for
who
can
reach
what
—
and
being
able
to
prove
it
at
audit
time
.
Active
Directory
Forests
,
OUs
,
GPOs
,
FSMO
roles
,
trusts
,
replication
health
Microsoft
Entra
ID
Hybrid
join
,
conditional
access
,
entitlement
management
LDAP
directories
OpenLDAP
,
schemas
,
ACLs
,
referrals
,
replication
Kerberos
Realms
,
keytabs
,
ticket
lifetimes
,
cross
-
realm
trust
Federation
and
single
sign
-
on
SAML
2.0,
OIDC
,
OAuth
2.0,
Okta
,
Keycloak
MFA
and
passwordless
TOTP
,
FIDO
2 /
WebAuthn
,
smart
cards
,
account
recovery
flows
Authorization
models
RBAC
,
ABAC
,
sudoers
policy
design
,
least
privilege
Privileged
access
management
Credential
vaulting
,
session
recording
,
break
-
glass
accounts
Certificate
services
AD
CS
,
internal
PKI
,
ACME
,
certificate
lifecycle
management
Identity
lifecycle
Joiner
/
mover
/
leaver
automation
,
SCIM
provisioning
2 / 7
05
Virtualization
&
Containers
Building
,
resizing
,
and
retiring
the
compute
layer
—
from
hypervisor
clusters
to
orchestrated
containers
.
VMware
vSphere
ESXi
,
vCenter
,
vMotion
,
DRS
,
HA
clusters
,
snapshots
Microsoft
Hyper
-
V
Failover
clustering
,
live
migration
,
checkpoints
KVM
and
Proxmox
libvirt
,
virt
-
manager
,
bridged
and
SR
-
IOV
networking
Container
runtimes
containerd
,
Docker
,
Podman
,
OCI
image
specifications
Kubernetes
operations
kubeadm
,
workloads
,
ingress
,
RBAC
,
Helm
releases
Cluster
lifecycle
Node
pools
,
taints
and
tolerations
,
drain
and
cordon
,
upgrades
Autoscaling
HPA
,
VPA
,
cluster
autoscaler
,
node
pool
sizing
Image
management
Multi
-
stage
builds
,
vulnerability
scanning
,
registry
retention
Registry
and
artifact
operations
Harbor
,
Amazon
ECR
,
Artifactory
,
image
signing
Virtual
desktop
infrastructure
Horizon
,
Citrix
,
golden
images
,
profile
containers
06
Cloud
Infrastructure
Running
the
same
operational
discipline
across
hyperscaler
environments
,
with
cost
and
governance
under
control
.
Amazon
Web
Services
EC
2,
S
3,
VPC
,
IAM
,
RDS
,
CloudWatch
,
Organizations
,
IAM
Identity
Center
Microsoft
Azure
Virtual
Machines
,
Blob
Storage
,
VNet
,
Entra
ID
,
Monitor
,
Bicep
Google
Cloud
Compute
Engine
,
Cloud
Storage
,
VPC
,
IAM
,
Cloud
Logging
Cloud
networking
VPC
peering
,
Transit
Gateway
,
private
endpoints
,
NAT
gateways
Cloud
identity
and
access
IAM
policies
,
service
accounts
,
roles
,
workload
federation
Cost
management
Tagging
standards
,
budgets
,
savings
plans
,
rightsizing
Migration
execution
Landing
zones
,
wave
planning
,
cutover
and
rollback
plans
Hybrid
connectivity
AWS
Direct
Connect
,
Azure
ExpressRoute
,
VPN
gateways
Managed
service
operations
Patching
windows
,
backup
policies
,
failover
drills
Cloud
governance
Service
control
policies
,
guardrails
,
audit
trails
07
Storage
,
Filesystems
&
Backup
Keeping
data
available
,
performant
,
and
recoverable
—
including
the
part
everyone
skips
,
which
is
testing
the
restore
.
Block
storage
SAN
,
iSCSI
,
Fibre
Channel
,
LUN
masking
,
multipathing
NAS
and
file
shares
NFSv
4,
SMB
/
CIFS
,
DFS
namespaces
,
ACL
mapping
3 / 7
Object
storage
S
3-
compatible
APIs
,
MinIO
,
lifecycle
and
tiering
rules
Filesystems
ext
4,
XFS
,
ZFS
,
Btrfs
,
NTFS
,
ReFS
,
mount
options
,
quotas
RAID
and
disk
management
RAID
levels
,
mdadm
,
hardware
controllers
,
SMART
health
LVM
and
volume
management
PV
/
VG
/
LV
,
thin
provisioning
,
online
resize
Backup
platforms
Veeam
,
Commvault
,
Bacula
,
restic
,
Duplicati
Backup
strategy
3-2-1-1-0,
RPO
/
RTO
targets
,
GFS
retention
,
immutability
Restore
validation
Bare
-
metal
recovery
,
sample
restores
,
integrity
verification
Storage
performance
IOPS
vs
throughput
,
queue
depth
,
tiering
,
deduplication
08
Configuration
Management
&
Infrastructure
as
Code
Defining
infrastructure
in
version
control
so
it
can
be
reviewed
,
reproduced
,
and
rolled
back
like
any
other
change
.
Ansible
Playbooks
,
roles
,
dynamic
inventory
,
Galaxy
,
AWX
/
AAP
Terraform
HCL
,
modules
,
remote
state
,
workspaces
,
provider
versioning
Puppet
Manifests
,
Hiera
data
,
module
design
,
compliance
reporting
Chef
Recipes
,
cookbooks
,
InSpec
profiles
,
policyfiles
SaltStack
States
,
grains
,
pillars
,
event
-
driven
orchestration
Golden
image
pipelines
Packer
,
AMI
builds
,
CIS
-
hardened
base
templates
GitOps
Argo
CD
,
Flux
,
declarative
cluster
state
,
drift
detection
Policy
as
code
Open
Policy
Agent
,
Sentinel
,
Checkov
,
tflint
Idempotency
and
drift
Convergence
testing
,
scheduled
compliance
scans
Environment
promotion
Dev
/
stage
/
prod
parity
,
change
gates
,
release
tagging
09
Monitoring
,
Logging
&
Observability
Knowing
what
is
broken
before
a
user
reports
it
,
and
being
able
to
explain
why
afterwards
.
Metrics
collection
Prometheus
,
node
_
exporter
,
recording
rules
,
Grafana
Long
-
term
metrics
storage
Thanos
,
VictoriaMetrics
,
retention
,
downsampling
Log
pipelines
rsyslog
,
Fluent
Bit
,
Logstash
,
Loki
,
Elasticsearch
Tracing
and
APM
OpenTelemetry
,
Jaeger
,
Datadog
,
New
Relic
Alerting
Alertmanager
,
PagerDuty
,
escalation
policy
,
alert
deduplication
Synthetic
monitoring
Blackbox
exporter
,
Pingdom
,
Uptime
Kuma
SNMP
and
infrastructure
polling
Zabbix
,
LibreNMS
,
Nagios
/
Icinga
,
threshold
tuning
SLOs
and
error
budgets
SLI
design
,
burn
-
rate
alerts
,
reliability
reviews
4 / 7
Dashboards
Service
health
views
,
capacity
trends
,
on
-
call
quick
panels
Health
checks
Liveness
and
readiness
probes
,
dependency
mapping
,
deep
checks
10
Security
,
Hardening
&
Compliance
Reducing
the
attack
surface
on
every
system
you
own
,
and
producing
the
evidence
that
says
you
did
.
Operating
system
hardening
CIS
Benchmarks
,
DISA
STIGs
,
SELinux
,
AppArmor
,
seccomp
Vulnerability
management
Nessus
,
Qualys
,
OpenVAS
,
patch
SLAs
,
risk
-
based
triage
Endpoint
protection
EDR
/
XDR
,
CrowdStrike
,
Microsoft
Defender
,
quarantine
workflow
Patch
management
Red
Hat
Satellite
,
WSUS
,
unattended
-
upgrades
,
maintenance
windows
Encryption
LUKS
,
BitLocker
,
TLS
1.3,
KMS
,
at
-
rest
and
in
-
transit
controls
Secrets
management
HashiCorp
Vault
,
key
rotation
,
HSM
basics
,
secret
-
sprawl
cleanup
Network
security
IDS
/
IPS
,
WAF
,
segmentation
,
microsegmentation
,
zero
trust
Incident
response
Triage
,
containment
,
evidence
preservation
,
blameless
postmortems
Compliance
frameworks
SOC
2,
ISO
27001,
PCI
DSS
,
HIPAA
,
GDPR
Audit
and
log
integrity
SIEM
(
Splunk
,
Microsoft
Sentinel
),
immutable
logs
,
retention
policy
11
Application
,
Web
&
Database
Servers
The
service
layer
that
sits
above
the
operating
system
and
below
the
application
team
'
s
code
.
Web
servers
NGINX
,
Apache
httpd
,
IIS
;
virtual
hosts
,
TLS
,
reverse
proxy
rules
Application
platforms
Apache
Tomcat
,
WildFly
, .
NET
hosting
,
Node
service
runtimes
Relational
databases
PostgreSQL
,
MySQL
/
MariaDB
,
SQL
Server
operations
NoSQL
services
MongoDB
,
Cassandra
,
Redis
persistence
modes
Database
operations
Backup
and
restore
,
replication
,
failover
,
vacuum
and
indexes
Message
queues
and
brokers
RabbitMQ
,
Kafka
,
ActiveMQ
;
topics
,
consumers
,
lag
monitoring
Caching
layers
Redis
,
Memcached
,
eviction
policy
,
invalidation
strategy
Email
systems
Postfix
,
Exchange
Online
,
SPF
/
DKIM
/
DMARC
,
relay
hygiene
File
transfer
SFTP
,
FTPS
,
rsync
,
managed
file
transfer
,
key
rotation
Deployment
support
Blue
/
green
,
canary
,
rollback
plans
,
release
verification
12
Performance
,
Availability
&
Disaster
Recovery
Making
systems
fast
enough
at
peak
,
and
survivable
when
a
site
or
region
goes
dark
.
5 / 7
Performance
analysis
sar
,
iostat
,
vmstat
,
perf
,
eBPF
,
flame
graphs
Bottleneck
isolation
CPU
,
memory
,
disk
,
network
,
and
lock
contention
triage
System
tuning
Kernel
parameters
,
JVM
heap
,
database
config
,
web
worker
counts
Load
testing
k
6,
JMeter
,
Locust
,
baseline
versus
peak
profiling
High
availability
design
Clustering
,
quorum
,
fencing
,
active
/
passive
versus
active
/
active
Failover
automation
keepalived
,
Pacemaker
,
health
-
based
traffic
routing
Disaster
recovery
DR
runbooks
,
RTO
/
RPO
targets
,
site
failover
exercises
Business
continuity
Dependency
mapping
,
tabletop
scenarios
,
communication
plans
Capacity
planning
Growth
modeling
,
headroom
thresholds
,
procurement
lead
times
Post
-
incident
review
Root
cause
analysis
,
action
tracking
,
MTTR
and
repeat
-
incident
trends
13
IT
Service
Management
&
Operations
The
process
discipline
that
makes
technical
work
predictable
,
auditable
,
and
repeatable
by
more
than
one
person
.
Ticketing
systems
ServiceNow
,
Jira
Service
Management
,
Zendesk
,
queue
hygiene
Incident
management
Severity
matrix
,
major
incident
bridge
,
communication
cadence
Change
management
CAB
review
,
risk
assessment
,
rollback
and
back
-
out
plans
Problem
management
Recurring
issue
analysis
,
permanent
corrective
actions
Request
fulfillment
Service
catalog
,
SLAs
,
automated
approval
workflows
Documentation
Runbooks
,
network
diagrams
,
build
standards
,
CMDB
accuracy
Asset
and
license
management
Hardware
inventory
,
software
entitlements
,
compliance
audits
Knowledge
management
KB
articles
,
self
-
service
deflection
,
scheduled
article
review
Vendor
management
Support
escalations
,
contract
renewals
,
RMA
process
On
-
call
operations
Rotations
,
escalation
trees
,
handover
notes
,
fatigue
limits
14
Adjacent
&
Emerging
Practice
Skills
that
increasingly
separate
a
senior
systems
administrator
from
a
mid
-
level
one
.
Site
reliability
engineering
Error
budgets
,
toil
reduction
,
blameless
culture
Platform
engineering
Internal
developer
platforms
,
Backstage
,
golden
paths
AI
-
assisted
operations
Runbook
copilots
,
log
summarization
,
anomaly
triage
Security
automation
SOAR
playbooks
,
auto
-
remediation
,
drift
containment
6 / 7
Edge
and
IoT
operations
Fleet
management
,
over
-
the
-
air
updates
,
constrained
hardware
7 / 7