Statistical
Analysis
—
Hard
Skills
01
Statistical
Theory
&
Inference
18
skills
02
Experimental
Design
&
Causal
Inference
11
skills
Probability
theory
—
discrete
,
continuous
,
joint
&
conditional
distributions
–
Statistical
inference
—
point
estimation
,
confidence
intervals
,
hypothesis
testing
–
Bayesian
inference
—
priors
,
posteriors
,
credible
intervals
,
hierarchical
models
–
Estimation
theory
—
maximum
likelihood
,
method
of
moments
,
M
-
estimators
–
Sampling
distributions
—
central
limit
theorem
and
asymptotic
theory
–
Hypothesis
testing
frameworks
—
t
-
tests
,
chi
-
square
,
F
-
tests
,
likelihood
ratio
tests
–
Power
analysis
&
sample
size
determination
–
Effect
size
estimation
—
Cohen
'
s
d
,
odds
ratios
,
hazard
ratios
,
eta
-
squared
–
Analysis
of
variance
—
ANOVA
,
ANCOVA
,
MANOVA
,
repeated
measures
–
Regression
modeling
—
OLS
,
logistic
,
Poisson
,
generalized
linear
models
–
Nonparametric
methods
—
Mann
–
Whitney
,
Kruskal
–
Wallis
,
permutation
tests
–
Resampling
&
simulation
—
bootstrap
,
jackknife
,
Monte
Carlo
methods
–
Multivariate
analysis
—
PCA
,
factor
analysis
,
discriminant
and
cluster
analysis
–
Survival
&
reliability
analysis
—
Kaplan
–
Meier
,
Cox
proportional
hazards
,
Weibull
–
Time
series
analysis
—
ARIMA
,
SARIMA
,
GARCH
,
VAR
,
spectral
analysis
–
Missing
data
handling
—
multiple
imputation
,
EM
algorithm
,
MCAR
/
MAR
/
MNAR
–
Meta
-
analysis
&
evidence
synthesis
–
Linear
algebra
&
calculus
for
statistics
—
matrix
operations
,
optimization
,
derivatives
–
Randomized
controlled
trial
design
—
randomization
,
allocation
,
blinding
–
Factorial
&
fractional
factorial
designs
–
Response
surface
methodology
—
central
composite
,
Box
–
Behnken
–
Blocking
&
stratification
—
variance
reduction
techniques
–
A
/
B
&
multivariate
testing
at
scale
–
Sequential
testing
&
alpha
spending
—
group
sequential
,
O
'
Brien
–
Fleming
boundaries
–
Quasi
-
experimental
methods
—
difference
-
in
-
differences
,
IV
,
regression
discontinuity
–
Propensity
score
methods
—
matching
,
inverse
probability
weighting
–
Causal
graph
analysis
—
directed
acyclic
graphs
,
confounding
identification
–
Minimum
detectable
effect
calculation
–
Adaptive
allocation
&
multi
-
armed
bandits
–
1 / 3
03
Programming
&
Statistical
Computing
11
skills
04
Data
Management
&
Engineering
10
skills
05
Predictive
Modeling
&
Machine
Learning
12
skills
06
Data
Visualization
&
Reporting
8
skills
R
—
tidyverse
,
data
.
table
,
ggplot
2,
lme
4,
caret
–
Python
—
pandas
,
NumPy
,
SciPy
,
statsmodels
,
scikit
-
learn
–
SAS
—
Base
,
STAT
,
PROC
SQL
,
SAS
Macro
language
–
Stata
—
panel
data
,
survey
commands
,
post
-
estimation
–
SPSS
&
JMP
&
Minitab
—
point
-
and
-
click
statistical
workflows
–
MATLAB
&
Julia
—
numerical
computation
,
optimization
routines
–
Bayesian
programming
—
Stan
,
PyMC
,
brms
,
JAGS
,
MCMC
/
HMC
sampling
–
Reproducible
research
authoring
—
R
Markdown
,
Quarto
,
Jupyter
,
knitr
–
Version
control
—
Git
,
GitHub
,
branching
,
code
review
workflows
–
Package
&
function
development
—
reusable
modules
,
testing
,
documentation
–
Performance
optimization
—
vectorization
,
Rcpp
,
Cython
,
Numba
,
parallel
computing
–
SQL
for
analytics
—
joins
,
window
functions
,
CTEs
,
query
tuning
–
Data
wrangling
&
reshaping
—
tidying
,
merging
,
pivoting
,
cleaning
–
ETL
/
ELT
pipeline
development
–
Data
quality
validation
—
schema
checks
,
outlier
and
anomaly
detection
–
Database
design
—
relational
,
columnar
,
NoSQL
data
stores
–
Cloud
data
platforms
—
Snowflake
,
BigQuery
,
Redshift
,
Databricks
–
Workflow
orchestration
—
Airflow
,
dbt
,
Dagster
,
scheduled
pipelines
–
Large
-
scale
data
processing
—
Spark
,
Dask
,
Hadoop
MapReduce
–
API
integration
&
web
scraping
–
Data
governance
&
lineage
—
cataloging
,
metadata
,
access
control
–
Supervised
learning
—
regression
,
classification
,
decision
trees
–
Unsupervised
learning
—
k
-
means
,
hierarchical
clustering
,
DBSCAN
–
Dimensionality
reduction
—
PCA
,
t
-
SNE
,
UMAP
,
factor
rotation
–
Regularization
—
ridge
,
LASSO
,
elastic
net
,
shrinkage
–
Cross
-
validation
&
resampling
—
k
-
fold
,
leave
-
one
-
out
,
nested
CV
–
Model
selection
&
diagnostics
—
AIC
,
BIC
,
residual
analysis
,
calibration
–
Ensemble
methods
—
random
forests
,
bagging
,
stacking
–
Gradient
boosting
—
XGBoost
,
LightGBM
,
CatBoost
–
Mixed
-
effects
&
hierarchical
models
–
Feature
engineering
&
selection
–
Model
interpretability
—
SHAP
,
partial
dependence
,
permutation
importance
–
Model
deployment
&
monitoring
—
MLOps
,
drift
detection
,
model
retraining
–
Statistical
graphics
—
ggplot
2,
matplotlib
,
seaborn
,
Plotly
,
Altair
–
BI
&
dashboard
platforms
—
Tableau
,
Power
BI
,
Looker
,
Qlik
–
Interactive
visualization
—
Shiny
,
Streamlit
,
D
3.
js
,
R
htmlwidgets
–
Chart
selection
&
perceptual
design
–
KPI
dashboard
design
—
metric
hierarchy
,
drill
-
downs
,
filters
–
Data
storytelling
—
executive
presentation
,
narrative
framing
–
Technical
writing
—
methodology
sections
,
statistical
documentation
–
Automated
reporting
—
parametrized
reports
,
scheduled
delivery
–
2 / 3
07
Domain
-
Specific
Statistical
Applications
10
skills
08
Quality
,
Compliance
&
Reproducibility
9
skills
Biostatistics
&
clinical
trial
analysis
—
ICH
E
9,
CDISC
SDTM
/
ADaM
–
Econometrics
—
panel
data
,
fixed
&
random
effects
,
instrumental
variables
–
Epidemiology
&
public
health
statistics
–
Finance
&
risk
modeling
—
VaR
,
volatility
models
,
credit
scoring
–
Marketing
analytics
—
attribution
,
lift
measurement
,
customer
lifetime
value
–
Survey
methodology
—
sampling
frames
,
weighting
,
nonresponse
adjustment
–
Statistical
process
control
—
control
charts
,
capability
analysis
,
Six
Sigma
–
Actuarial
&
insurance
modeling
—
life
tables
,
loss
distributions
,
credibility
theory
–
Genomics
&
bioinformatics
statistics
—
high
-
dimensional
testing
,
FDR
control
–
Social
science
&
psychometrics
—
scale
reliability
,
item
response
theory
–
Statistical
Analysis
Plan
authoring
–
Regulatory
compliance
— 21
CFR
Part
11,
GxP
,
ICH
guidelines
–
Submission
support
—
FDA
,
EMA
,
eCTD
documentation
packages
–
Programming
validation
—
independent
double
programming
,
QC
checks
–
Peer
code
review
—
analytical
review
,
reproducibility
audits
–
Unit
testing
for
analytical
code
—
testthat
,
pytest
,
assertions
on
outputs
–
Study
pre
-
registration
&
protocols
–
Data
privacy
&
ethics
—
HIPAA
,
GDPR
,
de
-
identification
,
IRB
requirements
–
Audit
trails
&
documentation
standards
–
3 / 3