Skill v1.0.1
currentAutomated scan100/100+4 new
version: "1.0.1" name: kubernetes-patterns description: Kubernetes workload patterns, resource management, RBAC, probes, autoscaling, ConfigMap/Secret handling, and kubectl debugging for production-grade deployments. Use when writing or reviewing Kubernetes manifests, or debugging probes, RBAC, autoscaling, or resource limits. metadata: origin: ECC
Kubernetes Patterns
Production-grade Kubernetes patterns for deploying, managing, and debugging workloads reliably.
When to Activate
- Writing Kubernetes manifests (Deployments, Services, Ingress, Jobs)
- Configuring resource requests/limits, liveness/readiness probes
- Setting up RBAC, namespaces, or ServiceAccounts
- Managing configuration and secrets in K8s
- Debugging CrashLoopBackOff, OOMKilled, pending pods, or image pull errors
- Configuring HPA (Horizontal Pod Autoscaler) or PodDisruptionBudgets
- Reviewing K8s YAML for security or correctness
When to Use
Same as When to Activate above. This alias satisfies repo skill-format conventions. Use this skill any time you are writing, reviewing, or debugging Kubernetes YAML and workloads.
How It Works
This skill provides copy-pasteable, production-grade YAML patterns and kubectl debugging commands organized by task:
- Deployment template — A fully configured production
Deploymentwith security context, rolling update strategy, all three probe types, resource limits, and environment injection from ConfigMap/Secret. - Probes — Decision table for startup vs liveness vs readiness, with correct
failureThreshold × periodSecondsmath. - Services & Ingress — ClusterIP, LoadBalancer, and TLS Ingress patterns with cert-manager annotations.
- ConfigMaps & Secrets —
envFrom, file-mount, and external secrets guidance. - Resource management — Requests vs limits rules of thumb by workload type (web API, JVM, worker, sidecar).
- RBAC — Least-privilege ServiceAccount → Role → RoleBinding chain.
- HPA & PDB — Autoscaling and node-drain safety configurations.
- Jobs & CronJobs — One-off and scheduled workload patterns with correct
restartPolicy. - kubectl cheatsheet — Logs, exec, rollback, port-forward, dry-run, and common error diagnosis commands.
- Anti-patterns & checklist — What NOT to do, and a security/reliability/observability checklist.
Examples
See the sections below for complete, runnable examples. Quick references:
| Task | Jump to | |
|---|---|---|
| Full production Deployment YAML | Core Workload Patterns | |
| Probe configuration | Probes | |
| RBAC least-privilege setup | RBAC | |
| Debug a CrashLoopBackOff | kubectl Debugging Cheatsheet | |
| Autoscaling | HPA |
Core Workload Patterns
Deployment — Production Template
apiVersion: apps/v1kind: Deploymentmetadata:name: my-appnamespace: my-namespacelabels:app: my-appversion: "1.0.0"spec:replicas: 3selector:matchLabels:app: my-appstrategy:type: RollingUpdaterollingUpdate:maxSurge: 1 # Allow 1 extra pod during updatemaxUnavailable: 0 # Never reduce below desired counttemplate:metadata:labels:app: my-appversion: "1.0.0"spec:# Security context at pod levelsecurityContext:runAsNonRoot: truerunAsUser: 1001fsGroup: 1001# Graceful shutdownterminationGracePeriodSeconds: 30containers:- name: my-appimage: ghcr.io/org/my-app:1.0.0 # Never use :latestimagePullPolicy: IfNotPresentports:- containerPort: 8080protocol: TCP# Resource requests AND limits are both requiredresources:requests:cpu: "100m"memory: "128Mi"limits:cpu: "500m"memory: "256Mi"# Container security contextsecurityContext:allowPrivilegeEscalation: falsereadOnlyRootFilesystem: truecapabilities:drop:- ALL# Probes (see Probes section below)startupProbe:httpGet:path: /healthport: 8080failureThreshold: 30periodSeconds: 5livenessProbe:httpGet:path: /healthport: 8080initialDelaySeconds: 0periodSeconds: 30failureThreshold: 3readinessProbe:httpGet:path: /readyport: 8080initialDelaySeconds: 5periodSeconds: 10failureThreshold: 2# Environment from ConfigMap and SecretenvFrom:- configMapRef:name: my-app-configenv:- name: DB_PASSWORDvalueFrom:secretKeyRef:name: my-app-secretskey: db-password# Writable tmp directory when readOnlyRootFilesystem: truevolumeMounts:- name: tmpmountPath: /tmpvolumes:- name: tmpemptyDir: {}
Probes — Liveness, Readiness, Startup
Understanding when to use each probe is critical:
| Probe | Failure Action | Use For | |
|---|---|---|---|
startupProbe | Kills container if slow to start | Slow-starting apps (JVM, Python) | |
livenessProbe | Restarts container | Deadlock / hung process detection | |
readinessProbe | Removes from Service endpoints | Temporary unavailability (DB reconnect) |
# Correct pattern: startupProbe covers slow startup,# then liveness/readiness take overstartupProbe:httpGet:path: /healthport: 8080failureThreshold: 30 # 30 * 5s = 150s max startup timeperiodSeconds: 5livenessProbe:httpGet:path: /healthport: 8080periodSeconds: 30failureThreshold: 3 # 3 * 30s = 90s before restartreadinessProbe:httpGet:path: /ready # Separate endpoint: checks DB, cache, etc.port: 8080periodSeconds: 10failureThreshold: 2
# WRONG: initialDelaySeconds without startupProbe# If the app takes 60s to start, set a startupProbe insteadlivenessProbe:httpGet:path: /healthport: 8080initialDelaySeconds: 60 # BAD: Arbitrary wait, race condition
Services and Ingress
Service Types
# ClusterIP (default) — internal-onlyapiVersion: v1kind: Servicemetadata:name: my-appnamespace: my-namespacespec:selector:app: my-appports:- port: 80targetPort: 8080protocol: TCPtype: ClusterIP
# LoadBalancer — external traffic (cloud providers)spec:type: LoadBalancerports:- port: 443targetPort: 8080
Ingress with TLS
apiVersion: networking.k8s.io/v1kind: Ingressmetadata:name: my-appnamespace: my-namespaceannotations:nginx.ingress.kubernetes.io/ssl-redirect: "true"cert-manager.io/cluster-issuer: "letsencrypt-prod"spec:ingressClassName: nginxtls:- hosts:- myapp.example.comsecretName: my-app-tlsrules:- host: myapp.example.comhttp:paths:- path: /pathType: Prefixbackend:service:name: my-appport:number: 80
ConfigMaps and Secrets
ConfigMap — Non-sensitive configuration
apiVersion: v1kind: ConfigMapmetadata:name: my-app-confignamespace: my-namespacedata:LOG_LEVEL: "info"APP_ENV: "production"MAX_CONNECTIONS: "100"# Mount as a file for complex configapp.yaml: |server:port: 8080timeout: 30s
# Mount ConfigMap as a filevolumes:- name: configconfigMap:name: my-app-configitems:- key: app.yamlpath: app.yamlvolumeMounts:- name: configmountPath: /etc/appreadOnly: true
Secrets — Sensitive data
# Create secret from literal (CLI, then store in Vault/SOPS)kubectl create secret generic my-app-secrets \--from-literal=db-password='s3cr3t' \--namespace=my-namespace \--dry-run=client -o yaml | kubectl apply -f -
apiVersion: v1kind: Secretmetadata:name: my-app-secretsnamespace: my-namespacetype: Opaque# Values are base64-encoded (NOT encrypted — use Sealed Secrets or ESO for real encryption)data:db-password: czNjcjN0 # base64 of 's3cr3t'
Important: Raw Kubernetes Secrets are only base64-encoded, not encrypted at rest unless your cluster has encryption configured. Use Sealed Secrets or External Secrets Operator for production.
Resource Requests and Limits
resources:requests: # Scheduler uses this to place the podcpu: "100m" # 100 millicores = 0.1 CPUmemory: "128Mi"limits: # Container is killed/throttled above thiscpu: "500m"memory: "256Mi"
Rules of thumb:
| Workload Type | CPU Request | Memory Request | Notes | |
|---|---|---|---|---|
| Web API | 100–250m | 128–256Mi | Set limits 2-4x requests | |
| Worker/consumer | 250–500m | 256–512Mi | Memory limit = request for predictability | |
| JVM app | 500m–1 | 512Mi–2Gi | Allow headroom above -Xmx for JVM overhead | |
| Sidecar | 10–50m | 32–64Mi | Keep minimal |
# WRONG: No requests or limits — unpredictable scheduling, OOM evictionscontainers:- name: appimage: myapp:latest# Missing resources: {} — this is dangerous in production# WRONG: Limits without requests — requests default to limits, over-reserves capacityresources:limits:cpu: "2"memory: "1Gi"# requests missing — will default to limits values
RBAC — Roles and ServiceAccounts
Principle of Least Privilege
Two patterns depending on whether the app calls the Kubernetes API:
Pattern A — App does NOT need the Kubernetes API (most apps)
Disable token automounting on the ServiceAccount. The Role/RoleBinding are not needed.
# ServiceAccount with token disabled — safest defaultapiVersion: v1kind: ServiceAccountmetadata:name: my-app-sanamespace: my-namespaceautomountServiceAccountToken: false # No K8s API token injected into pods
# Reference in Deployment — no token, no API accessspec:template:spec:serviceAccountName: my-app-saautomountServiceAccountToken: false # Belt-and-suspenders: also set at pod level
Pattern B — App DOES need the Kubernetes API (operators, controllers, config watchers)
Enable the token and grant only the permissions actually required.
# 1. ServiceAccount — enable token for this SAapiVersion: v1kind: ServiceAccountmetadata:name: my-app-sanamespace: my-namespaceautomountServiceAccountToken: true # Token required: app calls K8s API
# 2. Role — grant only what the app needs (namespace-scoped)apiVersion: rbac.authorization.k8s.io/v1kind: Rolemetadata:name: my-app-rolenamespace: my-namespacerules:- apiGroups: [""]resources: ["configmaps"]verbs: ["get", "list", "watch"] # Read-only, specific resource- apiGroups: [""]resources: ["secrets"]resourceNames: ["my-app-secrets"] # Restrict to specific secret by nameverbs: ["get"]
# 3. Bind Role to ServiceAccountapiVersion: rbac.authorization.k8s.io/v1kind: RoleBindingmetadata:name: my-app-rolebindingnamespace: my-namespacesubjects:- kind: ServiceAccountname: my-app-sanamespace: my-namespaceroleRef:kind: RoleapiGroup: rbac.authorization.k8s.ioname: my-app-role
# 4. Reference SA in Deploymentspec:template:spec:serviceAccountName: my-app-sa# automountServiceAccountToken defaults to true from SA — token is injected
Horizontal Pod Autoscaler (HPA)
apiVersion: autoscaling/v2kind: HorizontalPodAutoscalermetadata:name: my-app-hpanamespace: my-namespacespec:scaleTargetRef:apiVersion: apps/v1kind: Deploymentname: my-appminReplicas: 2 # Always at least 2 for HAmaxReplicas: 10metrics:- type: Resourceresource:name: cputarget:type: UtilizationaverageUtilization: 70 # Scale up when avg CPU > 70%- type: Resourceresource:name: memorytarget:type: UtilizationaverageUtilization: 80
HPA requiresresources.requeststo be set on all containers — it calculates utilization ascurrent / request.
PodDisruptionBudget (PDB)
Prevent too many pods going down during node drains or rolling updates:
apiVersion: policy/v1kind: PodDisruptionBudgetmetadata:name: my-app-pdbnamespace: my-namespacespec:minAvailable: 2 # OR use maxUnavailable: 1selector:matchLabels:app: my-app
Namespaces and Multi-Tenancy
# Create namespace with resource quotaskubectl create namespace my-namespace# Apply ResourceQuota to limit namespace consumptionkubectl apply -f - <<EOFapiVersion: v1kind: ResourceQuotametadata:name: my-namespace-quotanamespace: my-namespacespec:hard:requests.cpu: "4"requests.memory: 4Gilimits.cpu: "8"limits.memory: 8Gipods: "20"EOF
Jobs and CronJobs
# One-off Job (DB migration, data processing)apiVersion: batch/v1kind: Jobmetadata:name: db-migratenamespace: my-namespacespec:backoffLimit: 3 # Retry up to 3 times on failurettlSecondsAfterFinished: 3600 # Auto-delete after 1htemplate:spec:restartPolicy: OnFailure # Never for Jobs (not Always)containers:- name: migrateimage: ghcr.io/org/my-app:1.0.0command: ["python", "manage.py", "migrate"]resources:requests:cpu: "100m"memory: "256Mi"
# CronJobapiVersion: batch/v1kind: CronJobmetadata:name: cleanup-jobnamespace: my-namespacespec:schedule: "0 2 * * *" # 2am dailyconcurrencyPolicy: Forbid # Don't run if previous still runningsuccessfulJobsHistoryLimit: 3failedJobsHistoryLimit: 1jobTemplate:spec:template:spec:restartPolicy: OnFailurecontainers:- name: cleanupimage: ghcr.io/org/cleanup:1.0.0resources:requests:cpu: "50m"memory: "64Mi"
kubectl Debugging Cheatsheet
# --- Pod status and logs ---kubectl get pods -n my-namespacekubectl get pods -n my-namespace -o wide # Show node assignmentkubectl describe pod <pod-name> -n my-namespace # Events and state detailskubectl logs <pod-name> -n my-namespace # Current logskubectl logs <pod-name> -n my-namespace --previous # Logs from crashed containerkubectl logs <pod-name> -n my-namespace -c <container> # Multi-container pod# --- Execute into a running container ---kubectl exec -it <pod-name> -n my-namespace -- shkubectl exec -it <pod-name> -n my-namespace -- bash# --- Check resource usage ---kubectl top pods -n my-namespacekubectl top nodes# --- Deployment operations ---kubectl rollout status deployment/my-app -n my-namespacekubectl rollout history deployment/my-app -n my-namespacekubectl rollout undo deployment/my-app -n my-namespace # Rollbackkubectl rollout undo deployment/my-app --to-revision=2 -n my-namespace# --- Scale manually ---kubectl scale deployment my-app --replicas=5 -n my-namespace# --- Inspect events (cluster-wide issues) ---kubectl get events -n my-namespace --sort-by='.lastTimestamp'# --- Port-forward for local debugging ---kubectl port-forward pod/<pod-name> 8080:8080 -n my-namespacekubectl port-forward svc/my-app 8080:80 -n my-namespace# --- Dry-run to validate YAML ---kubectl apply -f deployment.yaml --dry-run=clientkubectl apply -f deployment.yaml --dry-run=server # Validates against live cluster
Diagnosing Common Errors
# CrashLoopBackOff: container keeps crashingkubectl logs <pod-name> --previous -n my-namespace # Check crash logskubectl describe pod <pod-name> -n my-namespace # Check exit code & OOMKilled# ImagePullBackOff: can't pull imagekubectl describe pod <pod-name> -n my-namespace # Check Events section# Causes: wrong image tag, missing imagePullSecret, private registry# Pending pod: not scheduledkubectl describe pod <pod-name> -n my-namespace# Causes: insufficient resources, no matching node selector, taint/toleration mismatch# OOMKilled: out of memory# Increase memory limits, check for memory leakskubectl describe pod <pod-name> -n my-namespace | grep -A5 "Last State"
Anti-Patterns
# BAD: Using :latest tag — non-deterministic deploymentsimage: myapp:latest# GOOD: Pin to a specific immutable tag (SHA or semver)image: ghcr.io/org/myapp:1.4.2# orimage: ghcr.io/org/myapp@sha256:abc123...# ---# BAD: Running as rootsecurityContext: {} # Defaults to root# GOOD: Non-root with explicit UIDsecurityContext:runAsNonRoot: truerunAsUser: 1001# ---# BAD: No resource limits — one pod can starve the entire nodecontainers:- name: appimage: myapp:1.0.0# No resources defined# GOOD: Always set requests and limitsresources:requests:cpu: "100m"memory: "128Mi"limits:cpu: "500m"memory: "256Mi"# ---# BAD: Storing plaintext secrets in ConfigMapsapiVersion: v1kind: ConfigMapdata:DB_PASSWORD: "mysecretpassword" # NEVER — use Secret or external secrets manager# ---# BAD: ClusterAdmin for application service accountsapiVersion: rbac.authorization.k8s.io/v1kind: ClusterRoleBindingroleRef:kind: ClusterRolename: cluster-admin # Grants god-mode to your app# ---# BAD: minAvailable: 0 in PDB — defeats the purposespec:minAvailable: 0# ---# BAD: restartPolicy: Always in a Job (causes infinite restart loop)spec:restartPolicy: Always # Use OnFailure or Never for Jobs
Best Practices Checklist
Security
- [ ] Container runs as non-root (
runAsNonRoot: true,runAsUserset) - [ ]
readOnlyRootFilesystem: truewithemptyDirfor writable paths - [ ]
allowPrivilegeEscalation: false - [ ] All capabilities dropped (
capabilities.drop: [ALL]) - [ ] Dedicated ServiceAccount per app, not
default - [ ]
automountServiceAccountToken: falseunless needed - [ ] RBAC follows least privilege (use
Role, notClusterRoleunless needed) - [ ] Secrets managed via Sealed Secrets or External Secrets Operator
Reliability
- [ ] All 3 probe types configured (startup + liveness + readiness)
- [ ] Resource requests AND limits set on every container
- [ ]
minReplicas: 2+for any production workload - [ ] PodDisruptionBudget defined for stateful or critical services
- [ ]
RollingUpdatestrategy withmaxUnavailable: 0 - [ ] HPA configured for variable-load services
Observability
- [ ] App exposes
/health(liveness) and/ready(readiness) endpoints - [ ] Structured JSON logging (no PII in logs)
- [ ] Resource labels:
app,version,environment
Related Skills
docker-patterns— Multi-stage Dockerfiles and image securitydeployment-patterns— CI/CD pipelines, rollback strategy, health check endpointssecurity-review— Broader security hardening contextgit-workflow— GitOps integration with K8s (ArgoCD / Flux patterns)