Kubernetes Runner
Kubernetes Executor
Auto-Scaling
Helm Chart
Pod Security
A complete guide to the GitLab Kubernetes Runner covering installation, configuration, running jobs in Kubernetes pods, auto-scaling, and production best practices.
Kubernetes
Auto-Scaling
Helm
GitLab
What is the Kubernetes Runner?
The Kubernetes Runner is GitLab Runner configured with the Kubernetes executor. Instead of running jobs on a single host (shell executor) or in containers on the runner's Docker daemon (docker executor), it runs each job as a pod in a Kubernetes cluster. This makes the runner cloud-native: it integrates with Kubernetes' scheduling, resource management, and — most importantly — auto-scaling.
When a job arrives, the runner talks to the Kubernetes API server and requests a new pod for that job. Kubernetes schedules the pod onto a node, pulls the image, starts the pod, and runs the job's script. When the job finishes, the pod is deleted. Because pods are ephemeral, the environment is fresh every time — the same guarantee the Docker executor provides, but now delivered by Kubernetes itself.
The real power of the Kubernetes Runner is that it inherits Kubernetes' auto-scaling. If you run a Cluster Autoscaler (or Karpenter), then when there are more jobs than the cluster has capacity for, the cluster will add nodes automatically. When the jobs finish, the cluster shrinks back down. This means your CI/CD infrastructure scales with demand — you pay for compute only when jobs are running, and you never have to manually provision build servers.
For teams already operating Kubernetes, the Kubernetes Runner is usually the right choice: it reuses the same cluster, the same observability stack, and the same operational practices. For teams without Kubernetes, the setup is more involved but the auto-scaling benefits can justify the investment.
GitLab
Pipeline
Job
Kubernetes Cluster
GitLab Runner
Runner Pod
Job Execution
Build Pod
Service Pod
Helper Pod
Auto-Scaling
Cluster Autoscaler
Nodes added/removed
Key Concept: The Kubernetes Runner doesn't manage nodes itself — it delegates that to Kubernetes. The runner asks for pods; the Cluster Autoscaler ensures there's capacity to run them. This separation is what makes the Kubernetes executor so scalable.
Kubernetes Runner Architecture
Understanding the Kubernetes Runner's architecture is essential for configuring and troubleshooting it. The runner is itself deployed as a pod in the cluster, and it creates additional pods for each job. Three types of pods work together to run a single job: the helper pod, the build pod, and any service pods.
The helper pod is a small container that GitLab Runner uses for internal coordination — it handles things like cloning the repository into a shared volume, executing commands, and streaming logs back to GitLab. The build pod is where your job's script actually runs; its image is the one specified in .gitlab-ci.yml. The service pods are created for the services keyword in your job — databases, caches, and other dependencies that the build pod needs.
All three pods share volumes and network namespaces so they can communicate with each other. The helper pod can execute commands in the build pod via an exec-like mechanism, and services are reachable from the build pod by their alias hostname. Understanding this three-pod architecture explains why jobs sometimes need a little extra memory — the helper and service pods consume some of the requested resources too.
Helper Pod
Manages the job's lifecycle: clones repo, executes commands, streams logs. Runs a small image (usually the runner's helper image).
Build Pod
Where the job's script runs. Uses the image specified in .gitlab-ci.yml. This is what most people think of as "the job".
Service Pods
Started for services declared in the job (Postgres, Redis, etc.). Accessible from the build pod by their alias hostname.
Shared Volumes
Pods share volumes for the repo, cache, and artifacts. This is how the helper pod passes data to the build pod.
Shared Network
Pods share the same network namespace, so they can communicate over localhost or by alias hostname.
Poll Timeout
The runner waits for a pod to become ready. If the timeout is exceeded, the job fails. Tune for slow-starting images.
# Pod structure for a single job
#
# ┌─────────────────────────────────────────────┐
# │ Pod (job pod) │
# ├─────────────────────────────────────────────┤
# │ │
# │ ┌──────────────┐ ┌──────────────────┐ │
# │ │ Helper Pod │ │ Build Pod │ │
# │ │ (gitlab- │ │ (user image, │ │
# │ │ runner- │ │ e.g. node:18) │ │
# │ │ helper) │ │ │ │
# │ └──────────────┘ └──────────────────┘ │
# │ │
# │ ┌──────────────┐ ┌──────────────────┐ │
# │ │ Service Pod │ │ Service Pod │ │
# │ │ (postgres) │ │ (redis) │ │
# │ └──────────────┘ └──────────────────┘ │
# │ │
# │ Shared volumes: repo, cache, artifacts │
# │ Shared network: localhost + aliases │
# └─────────────────────────────────────────────┘
#
# All pods are in the same pod (or same network namespace)
# Viewing the pods during a job:
kubectl get pods -n gitlab-runner
# NAME READY STATUS
# runner-abc123-project-456-concurrent-0 3/3 Running
# - helper
# - build
# - svc-0 (postgres)
# Pod logs:
kubectl logs -n gitlab-runner runner-abc123-project-456-concurrent-0 -c helper
kubectl logs -n gitlab-runner runner-abc123-project-456-concurrent-0 -c build
kubectl logs -n gitlab-runner runner-abc123-project-456-concurrent-0 -c svc-0
# Describe pod for events:
kubectl describe pod -n gitlab-runner runner-abc123-project-456-concurrent-0
Installing the Kubernetes Runner
There are two ways to install the Kubernetes Runner: using the official Helm chart (recommended) or using a manual Deployment manifest. The Helm chart is recommended because it handles the RBAC, ConfigMap, Deployment, and Secret resources for you, and makes upgrades easy.
Before installing, you need to decide how the runner will authenticate to GitLab. There are two approaches: using a runner authentication token (obtained by registering a runner manually and reusing its token) or using the newer runner authentication tokens created via the GitLab UI or API. The Helm chart supports both approaches. The newer approach is recommended because it doesn't require a registration token to be present in the cluster.
You also need to decide on the namespace, service account, and RBAC permissions. The runner needs permission to create, get, list, watch, and delete pods in its namespace, plus permission to create secrets and configmaps for cache and artifacts. The Helm chart creates an appropriately scoped service account and role for you.
# Method 1: Install with Helm (recommended)
helm repo add gitlab https://charts.gitlab.io
helm repo update
# Create a namespace
kubectl create namespace gitlab-runner
# Install with a runner authentication token
helm install gitlab-runner gitlab/gitlab-runner \
--namespace gitlab-runner \
--set gitlabUrl=https://gitlab.com/ \
--set runnerToken="glrt-xxxxxxxxxxxxxxxxxxxx" \
--set rbac.create=true \
--set runners.privileged=false
# Or install with values file
helm install gitlab-runner gitlab/gitlab-runner \
--namespace gitlab-runner \
--values runner-values.yaml
# Method 2: Manual Deployment (for advanced users)
apiVersion: apps/v1
kind: Deployment
metadata:
name: gitlab-runner
namespace: gitlab-runner
spec:
replicas: 1
selector:
matchLabels:
app: gitlab-runner
template:
metadata:
labels:
app: gitlab-runner
spec:
serviceAccountName: gitlab-runner
containers:
- name: gitlab-runner
image: gitlab/gitlab-runner:latest
args:
- run
- --config=/etc/gitlab-runner/config.toml
volumeMounts:
- name: config
mountPath: /etc/gitlab-runner
volumes:
- name: config
configMap:
name: gitlab-runner-config
# Get a runner authentication token
# GitLab UI: Project → Settings → CI/CD → Runners → New project runner
# The token looks like: glrt-xxxxxxxxxxxxxxxxxxxx
# Verify installation
kubectl get pods -n gitlab-runner
kubectl get svc -n gitlab-runner
kubectl logs -n gitlab-runner deployment/gitlab-runner
# Check the runner's configuration
kubectl exec -n gitlab-runner deployment/gitlab-runner -- \
cat /etc/gitlab-runner/config.toml
# Check the runner's registration
# GitLab UI: Project → Settings → CI/CD → Runners
# The runner should appear with a green dot (online)
Runner Token Security:
- Store the runner token as a Kubernetes Secret, not in a ConfigMap
- Use a dedicated service account for the runner
- Limit RBAC permissions to the runner's namespace
- Rotate runner tokens if they're exposed
- Use protected runners for production workloads
Configuring the Kubernetes Runner
The Kubernetes Runner is configured via config.toml, which is mounted into the runner pod. The most important section is [runners.kubernetes], which controls the pod template used for each job: the image, CPU and memory limits, node selectors, service account, and security context. Getting this section right is the key to a performant and secure runner.
One of the most important settings is cpu_limit and memory_limit. These determine the resources requested for the build pod, which in turn determine which nodes the job can be scheduled on. If you set these too low, jobs may be OOM-killed or starved for CPU. If you set them too high, jobs may not fit on smaller nodes, or you may waste resources. The right values depend on your workload — start conservative and tune based on observed usage.
The namespace setting determines where job pods are created. Typically this is the same namespace as the runner itself. The service_account setting determines the identity used by the pods — by default, the default service account of the namespace, which is fine unless your jobs need special permissions. The node_selector and tolerations settings let you control which nodes jobs run on — useful for dedicating specific nodes to CI/CD or for using spot instances.
# ConfigMap: runner-values.yaml (Helm) or config.toml
[[runners]]
name = "Kubernetes Runner"
url = "https://gitlab.com/"
token = "glrt-xxxxxxxxxxxxxxxxxxxx"
executor = "kubernetes"
[runners.kubernetes]
# Namespace for job pods
namespace = "gitlab-runner"
# Default image for jobs
image = "alpine:latest"
# Allow privileged containers
privileged = false
allow_privilege_escalation = false
# Resource limits for the build pod
cpu_limit = "1"
memory_limit = "2Gi"
cpu_request = "500m"
memory_request = "512Mi"
# Helper pod resources
helper_cpu_limit = "500m"
helper_memory_limit = "512Mi"
# Service pod resources
service_cpu_limit = "1"
service_memory_limit = "1Gi"
# Timeout for pod startup (seconds)
poll_timeout = 180
# Pull policy for images
image_pull_secrets = []
# Node selector (which nodes to run on)
[runners.kubernetes.node_selector]
"kubernetes.io/os" = "linux"
"ci" = "true"
# Node tolerations (for tainted nodes)
[runners.kubernetes.node_tolerations]
"ci=true" = "NoSchedule"
# Pod labels
[runners.kubernetes.pod_labels]
"app" = "gitlab-runner"
# Pod annotations
[runners.kubernetes.pod_annotations]
"prometheus.io/scrape" = "true"
# Security context (pod level)
[runners.kubernetes.pod_security_context]
run_as_non_root = true
run_as_user = 1000
fs_group = 1000
# Security context (container level)
[runners.kubernetes.container_security_context]
allow_privilege_escalation = false
read_only_root_filesystem = false
# Helm values file
gitlabUrl: https://gitlab.com/
runnerToken: glrt-xxxxxxxxxxxxxxxxxxxx
rbac:
create: true
rules:
- apiGroups: [""]
resources: ["pods", "pods/exec", "pods/attach"]
verbs: ["get", "list", "watch", "create", "delete"]
- apiGroups: [""]
resources: ["secrets", "configmaps"]
verbs: ["get", "list", "watch", "create", "update", "delete"]
runners:
config: |
[[runners]]
[runners.kubernetes]
namespace = "gitlab-runner"
cpu_limit = "1"
memory_limit = "2Gi"
poll_timeout = 180
Key Configuration Settings:
- namespace: Where job pods are created
- image: Default image if the job doesn't specify one
- cpu_limit / memory_limit: Resource limits for build pod
- poll_timeout: How long to wait for pod startup
- node_selector: Which nodes to run on
- pod_security_context: Security settings for pods
Auto-Scaling with Kubernetes
The most compelling feature of the Kubernetes Runner is auto-scaling — and it comes essentially for free if you already use the Cluster Autoscaler. The runner doesn't need to know anything about auto-scaling: it just requests pods. When there aren't enough nodes to schedule those pods, the Cluster Autoscaler notices the pending pods and adds nodes. When the jobs finish and nodes are idle, the Cluster Autoscaler removes them.
There are two distinct scaling mechanisms to understand. First, Kubernetes itself schedules as many job pods as it can fit on available nodes. Second, the Cluster Autoscaler adds nodes when pods can't be scheduled. Together, these give you the effect of a runner that scales from zero to hundreds of concurrent jobs and back, all without manual intervention. This is particularly powerful for spiky workloads — a big release might create 50 jobs at once, but the cluster can grow to handle them and shrink back within minutes.
There are some caveats. Auto-scaling takes time — new nodes take 30 seconds to a few minutes to join the cluster, so jobs may be pending for a while during a spike. And there are limits: node pools have max sizes, and large clusters cost money. But the flexibility is enormous, and for most teams the auto-scaling savings (both time and money) far outweigh the complexity.
# Auto-scaling architecture
#
# 1. Jobs arrive at the runner
# 2. Runner requests pods from K8s API
# 3. If no node can fit the pod, it stays "Pending"
# 4. Cluster Autoscaler sees Pending pods
# 5. Cluster Autoscaler adds nodes (or scales up node group)
# 6. Pods get scheduled on new nodes
# 7. Jobs run
# 8. Jobs finish, pods deleted
# 9. Cluster Autoscaler removes idle nodes (after cooldown)
# Cluster Autoscaler configuration (AWS EKS example)
# Install the Cluster Autoscaler
helm repo add autoscaler https://kubernetes.github.io/autoscaler
helm install cluster-autoscaler autoscaler/cluster-autoscaler \
--namespace kube-system \
--set autoDiscovery.clusterName=my-cluster \
--set awsRegion=us-east-1 \
--set rbac.create=true
# Configure the node group to scale
# EKS managed node group with:
# minSize: 1
# maxSize: 20
# desiredSize: 1
# Karpenter (modern alternative to Cluster Autoscaler)
helm repo add karpenter https://charts.karpenter.sh
helm install karpenter karpenter/karpenter \
--namespace karpenter \
--create-namespace \
--set clusterName=my-cluster \
--set clusterEndpoint=$(aws eks describe-cluster \
--name my-cluster --query "cluster.endpoint" --output text)
# Karpenter NodePool (K8s CRD)
apiVersion: karpenter.sh/v1beta1
kind: NodePool
metadata:
name: ci-nodes
spec:
template:
spec:
requirements:
- key: "karpenter.sh/capacity-type"
operator: In
values: ["spot", "on-demand"]
- key: "kubernetes.io/arch"
operator: In
values: ["amd64"]
nodeClassRef:
name: default
limits:
cpu: "1000"
memory: 1000Gi
disruption:
consolidationPolicy: WhenUnderutilized
expireAfter: 720h
# Runner config for auto-scaling
# The runner needs no special config — just ensure
# the pods it creates have appropriate resource requests
# and node selectors.
# Pending pod example (during a spike)
kubectl get pods -n gitlab-runner --field-selector status.phase=Pending
# NAME READY STATUS
# runner-abc-project-123-concurrent-0 0/3 Pending
# Cluster Autoscaler triggers
kubectl logs -n kube-system deployment/cluster-autoscaler | grep "Triggered scale up"
# Auto-scaling best practices:
# 1. Set resource requests (not just limits) so the scheduler can make decisions
# 2. Use node pools with appropriate instance types
# 3. Consider spot instances for cost savings (with tolerations)
# 4. Set reasonable poll_timeout to handle node startup
# 5. Monitor pending pod counts and cluster size
Auto-Scaling Best Practices:
- Set resource requests — the scheduler uses them for placement
- Use multiple node pools (on-demand for baseline, spot for spikes)
- Set
poll_timeout to accommodate node startup time
- Use Karpenter for modern, fast auto-scaling (alternative to Cluster Autoscaler)
- Monitor cluster size and pending pods to tune node pool sizes
- Consider pod disruption budgets for critical infrastructure
Configuring Job Pods
Individual jobs can override many of the runner's Kubernetes settings using CI/CD variables. This is powerful because different jobs have different resource needs — a small lint job doesn't need 4GB of memory, while a large build might need 8GB. By setting these variables in .gitlab-ci.yml, you can right-size each job's pod and help the scheduler place it optimally.
The most commonly used variables are KUBERNETES_CPU_REQUEST, KUBERNETES_MEMORY_REQUEST, KUBERNETES_CPU_LIMIT, and KUBERNETES_MEMORY_LIMIT. These directly map to the pod's resource requests and limits. Setting requests is especially important for auto-scaling: if a job's request is too small, it may be scheduled on a node that can't actually run it efficiently; if it's too large, the job may not fit anywhere.
Beyond resources, jobs can override the service account, node selector, annotations, and even the entire pod spec. The KUBERNETES_OVERWRITE_CONTAINER_SPEC variable lets a job specify its own build container spec, giving it full control. These features let sophisticated jobs customize their environment while keeping the runner's defaults sensible for the common case.
# Override Kubernetes settings in .gitlab-ci.yml
# Resource overrides
build_large:
stage: build
image: node:18
variables:
KUBERNETES_CPU_REQUEST: "2"
KUBERNETES_MEMORY_REQUEST: "4Gi"
KUBERNETES_CPU_LIMIT: "4"
KUBERNETES_MEMORY_LIMIT: "8Gi"
script:
- npm run build:large
# Node selector overrides
gpu_job:
stage: test
image: nvidia/cuda:12.0
variables:
KUBERNETES_NODE_SELECTOR_OPERATOR: "In"
KUBERNETES_NODE_SELECTOR: "accelerator=nvidia-tesla-t4"
script:
- nvidia-smi
# Service pod overrides
integration_test:
stage: test
image: node:18
services:
- name: postgres:15
alias: db
variables:
KUBERNETES_SERVICE_CPU_REQUEST: "500m"
KUBERNETES_SERVICE_MEMORY_REQUEST: "1Gi"
KUBERNETES_SERVICE_CPU_LIMIT: "1"
KUBERNETES_SERVICE_MEMORY_LIMIT: "2Gi"
script:
- npm run test:integration
# Helper pod overrides
log_heavy_job:
script:
- generate-lots-of-logs.sh
variables:
KUBERNETES_HELPER_CPU_REQUEST: "500m"
KUBERNETES_HELPER_MEMORY_REQUEST: "512Mi"
KUBERNETES_HELPER_CPU_LIMIT: "1"
KUBERNETES_HELPER_MEMORY_LIMIT: "1Gi"
# Pod annotations
monitored_job:
script:
- run-app.sh
variables:
KUBERNETES_POD_ANNOTATIONS_1: "prometheus.io/scrape=true"
KUBERNETES_POD_ANNOTATIONS_2: "prometheus.io/port=9090"
# Pod labels
labeled_job:
script:
- echo "labeled"
variables:
KUBERNETES_POD_LABELS_1: "app=my-app"
KUBERNETES_POD_LABELS_2: "tier=backend"
# Tolerations (run on tainted nodes)
spot_job:
script:
- echo "Running on spot"
variables:
KUBERNETES_NODE_TOLERATIONS: "spot=true:NoSchedule"
# Service account override
privileged_job:
script:
- kubectl get pods
variables:
KUBERNETES_SERVICE_ACCOUNT_OVERWRITE: "ci-service-account"
# Runtime class (for gVisor, Kata, etc.)
sandboxed_job:
script:
- echo "Running in sandbox"
variables:
KUBERNETES_POD_SECURITY_CONTEXT_RUN_AS_NON_ROOT: "true"
KUBERNETES_POD_SECURITY_CONTEXT_RUN_AS_USER: "1000"
# Complete list of KUBERNETES_* variables:
# KUBERNETES_CPU_REQUEST
# KUBERNETES_CPU_LIMIT
# KUBERNETES_MEMORY_REQUEST
# KUBERNETES_MEMORY_LIMIT
# KUBERNETES_HELPER_CPU_REQUEST
# KUBERNETES_HELPER_CPU_LIMIT
# KUBERNETES_HELPER_MEMORY_REQUEST
# KUBERNETES_HELPER_MEMORY_LIMIT
# KUBERNETES_SERVICE_CPU_REQUEST
# KUBERNETES_SERVICE_CPU_LIMIT
# KUBERNETES_SERVICE_MEMORY_REQUEST
# KUBERNETES_SERVICE_MEMORY_LIMIT
# KUBERNETES_NODE_SELECTOR
# KUBERNETES_NODE_SELECTOR_OPERATOR
# KUBERNETES_NODE_TOLERATIONS
# KUBERNETES_SERVICE_ACCOUNT
# KUBERNETES_POD_ANNOTATIONS_n
# KUBERNETES_POD_LABELS_n
# KUBERNETES_POD_SECURITY_CONTEXT_*
# KUBERNETES_OVERWRITE_CONTAINER_SPEC
Resource Requests vs Limits:
- Requests: What the pod is guaranteed. Used for scheduling.
- Limits: What the pod is capped at. Exceeding leads to throttling or OOM.
- Always set requests. Set limits generously to avoid false OOMs.
- For auto-scaling, requests determine node sizing — get them right.
Security for the Kubernetes Runner
Security is especially important for the Kubernetes Runner because it runs untrusted code inside your cluster. If a job pod is compromised, an attacker could try to attack other workloads or the cluster itself. The good news is that Kubernetes provides excellent security primitives — you just need to enable them.
The most important controls are pod security contexts (run as non-root, drop capabilities, read-only filesystem), network policies (isolate job pods from production), RBAC (minimal permissions for the runner service account), and namespace isolation (run the runner in its own namespace). Combining these creates strong defense in depth: even if a job is compromised, its ability to affect anything else is limited.
For untrusted code (public repositories, contributed MRs), consider additional controls like gVisor or Kata Containers for stronger isolation, and ephemeral runners that are destroyed after each job. The most security-sensitive teams run one runner per project, or use the "runners autoscaling" feature to run each job in a fresh pod with no persistence.
# Pod security context (restrictive)
[runners.kubernetes.pod_security_context]
run_as_non_root = true
run_as_user = 1000
run_as_group = 1000
fs_group = 1000
seccomp_profile_type = "RuntimeDefault"
# Container security context
[runners.kubernetes.container_security_context]
allow_privilege_escalation = false
read_only_root_filesystem = true
capabilities:
drop = ["ALL"]
# Network policy (isolate job pods)
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: gitlab-runner-isolation
namespace: gitlab-runner
spec:
podSelector:
matchLabels:
app: gitlab-runner
policyTypes:
- Ingress
- Egress
ingress: [] # No ingress
egress:
# Allow DNS
- to:
- namespaceSelector: {}
podSelector:
matchLabels:
k8s-app: kube-dns
ports:
- port: 53
protocol: UDP
# Allow GitLab
- to:
- ipBlock:
cidr: 0.0.0.0/0
ports:
- port: 443
protocol: TCP
# RBAC for the runner (minimal)
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
name: gitlab-runner
namespace: gitlab-runner
rules:
- apiGroups: [""]
resources: ["pods", "pods/exec", "pods/attach", "pods/log"]
verbs: ["get", "list", "watch", "create", "delete"]
- apiGroups: [""]
resources: ["secrets", "configmaps"]
verbs: ["get", "list", "watch", "create", "update", "delete"]
- apiGroups: [""]
resources: ["serviceaccounts"]
verbs: ["get"]
# gVisor runtime class (stronger isolation)
apiVersion: node.k8s.io/v1
kind: RuntimeClass
metadata:
name: gvisor
handler: runsc
# Job using gVisor
sandboxed_job:
script:
- echo "Running in gVisor"
variables:
KUBERNETES_RUNTIME_CLASS: "gvisor"
# Ephemeral runner with GitLab Runner Operator
# Each job runs in a fresh runner pod that is deleted after use
# Security best practices:
# 1. Use a dedicated namespace for the runner
# 2. Use a dedicated service account with minimal RBAC
# 3. Enforce pod security contexts (non-root, drop caps)
# 4. Apply network policies to isolate job pods
# 5. Use image pull secrets for private registries
# 6. Scan job images for vulnerabilities
# 7. Audit runner and pod activity
# 8. Consider gVisor/Kata for untrusted code
# 9. Rotate runner tokens regularly
# 10. Keep the runner image updated
Security Checklist:
- Dedicated namespace for the runner
- Minimal RBAC permissions
- Pod security contexts (non-root, drop capabilities)
- Network policies to isolate job pods
- Image pull secrets for private registries
- gVisor or Kata for untrusted code
- Regular runner and image updates
- Audit logs and monitoring
Troubleshooting the Kubernetes Runner
# Runner pod not starting
kubectl get pods -n gitlab-runner
kubectl describe pod -n gitlab-runner
kubectl logs -n gitlab-runner
# Job stuck in "pending"
# 1. Check job pod status
kubectl get pods -n gitlab-runner
kubectl describe pod -n gitlab-runner
# 2. Check for scheduling issues:
kubectl get events -n gitlab-runner --sort-by='.lastTimestamp'
# Common "pending" causes:
# - Insufficient CPU/memory on nodes
# - Node selector matches no nodes
# - Taints without tolerations
# - Image pull secret missing
# - Resource quota exceeded
# Job pod running but job failing
kubectl logs -n gitlab-runner -c helper
kubectl logs -n gitlab-runner -c build
# Image pull errors
kubectl describe pod -n gitlab-runner | grep -A 5 Events
# Fix: ensure image exists, or add imagePullSecret
# "This job is stuck because you don't have any active runners"
# - Runner pod is not running or not registered
# - Check runner pod logs
kubectl logs -n gitlab-runner deployment/gitlab-runner
# RBAC issues
kubectl auth can-i create pods \
--as=system:serviceaccount:gitlab-runner:gitlab-runner \
-n gitlab-runner
# Should be "yes"
# Check runner registration
kubectl exec -n gitlab-runner deployment/gitlab-runner -- \
gitlab-runner list
# Check runner connectivity
kubectl exec -n gitlab-runner deployment/gitlab-runner -- \
gitlab-runner verify
# Increase debug logging
kubectl set env -n gitlab-runner deployment/gitlab-runner \
LOG_LEVEL=debug
# Restart runner
kubectl rollout restart -n gitlab-runner deployment/gitlab-runner
# Check the runner's configuration
kubectl exec -n gitlab-runner deployment/gitlab-runner -- \
cat /etc/gitlab-runner/config.toml
# View job pod events
kubectl get events -n gitlab-runner \
--field-selector involvedObject.name=
# Check resource usage
kubectl top pods -n gitlab-runner
kubectl top nodes
# Check cluster autoscaler
kubectl logs -n kube-system deployment/cluster-autoscaler | tail -50
# Debug with an interactive shell
kubectl run debug --rm -it --image=alpine -- sh
# Then test network access to GitLab, registry, etc.
Common Issues:
- Pending pods: Insufficient resources, node selectors, taints
- Image pull failures: Missing imagePullSecret, wrong registry
- RBAC errors: Missing permissions on pods, secrets, configmaps
- Network issues: Pods can't reach GitLab or registry
- Timeout: Increase poll_timeout for slow-starting images
- Helper pod OOM: Increase helper_memory_limit
Frequently Asked Questions
Do I need the Cluster Autoscaler for the Kubernetes Runner?
No, the runner works without auto-scaling. Auto-scaling is an optional optimization. But if you want the runner to scale with demand, install the Cluster Autoscaler (or Karpenter) and configure node pools that it can scale.
How do I isolate job pods from production workloads?
Use a dedicated namespace for the runner, apply network policies to isolate the namespace, use node selectors or taints to run job pods on dedicated nodes, and enforce pod security contexts.
Why are my job pods stuck in "Pending"?
Common causes: insufficient cluster resources (CPU/memory), node selectors that match no nodes, taints without matching tolerations, missing image pull secrets, or resource quotas. Check kubectl describe pod for scheduling events.
What's the difference between the helper pod and the build pod?
The helper pod is a small internal container that coordinates the job (clones the repo, streams logs, executes commands). The build pod is where the user's script runs, using the image specified in .gitlab-ci.yml.
How do I use Docker-in-Docker with the Kubernetes Runner?
Use the docker:dind service and set privileged = true in the runner configuration (or override per-job). Note that privileged mode has security implications — prefer Kaniko or Buildah for building images without DinD.
Can I use multiple Kubernetes runners?
Yes. You can install multiple runners with different configurations, or scale a single runner deployment to multiple replicas. The runner's concurrent setting controls how many jobs each replica handles.
How do I use spot instances for jobs?
Configure a node pool using spot instances, add a taint to those nodes (e.g., spot=true:NoSchedule), and add a toleration to the runner's job pods so they can be scheduled there. Jobs that must run on-demand nodes should avoid the toleration.
What's the maximum number of concurrent jobs?
Limited by the runner's concurrent setting, the cluster's capacity, and the node pool's max size. The runner's default is 10. Adjust concurrent in config.toml and ensure the cluster can scale to support it.
The Kubernetes Runner combines the isolation of containers with the auto-scaling of Kubernetes. For teams already operating Kubernetes, it's usually the best choice for CI/CD.