Pipeline Stages & Jobs

A complete guide to GitLab pipeline stages and jobs covering stage ordering, job parallelization, dependencies, needs (DAG), and best practices for pipeline design.

Stages Parallelization Needs (DAG) Dependencies
Understanding Pipeline Stages and Jobs

A GitLab CI/CD pipeline is built from two fundamental concepts: stages and jobs. Understanding how these two concepts interact is essential for designing pipelines that are fast, reliable, and easy to maintain. Stages provide the structure — they define the order in which work happens. Jobs provide the content — they define what work actually gets done. Together, they form the execution graph that GitLab follows every time your code is pushed.

When you push code, GitLab reads your .gitlab-ci.yml, extracts all the jobs, groups them by their assigned stage, and then executes stage by stage. Within each stage, jobs run in parallel by default. This two-level hierarchy gives you a powerful combination: sequential ordering where correctness demands it, and parallel execution where speed matters. Mastering this hierarchy is the difference between a pipeline that takes 30 minutes and one that takes 3.

The default execution model — stages in order, jobs in parallel — is called "stage-based execution." It's simple and works well for many cases. But GitLab also supports a more advanced model called "Directed Acyclic Graph" (DAG) execution, which allows jobs to declare their exact dependencies on other jobs, bypassing the stage barrier entirely. We'll cover both models in this guide.

Stage 1: build
build_app build_docs
Stage 2: test
unit_tests integration_tests lint
Stage 3: deploy
deploy_staging deploy_prod
Key Concept: Stages run sequentially (build → test → deploy). Jobs within a stage run in parallel. This model gives you ordering where you need it and speed where you can get it.
Stages: The Pipeline Backbone

Stages are the sequential backbone of a GitLab pipeline. When you define a stages list at the top of your .gitlab-ci.yml, you're telling GitLab the order in which groups of jobs should run. Every job must be assigned to exactly one stage (via the stage keyword), and if a job isn't assigned explicitly, it defaults to the test stage. This default is a common source of confusion for beginners, so it's usually best to assign every job to a stage explicitly.

The execution rule is simple but powerful: all jobs in stage N must complete successfully before any job in stage N+1 can begin. This means your test stage won't start until your build stage finishes, and your deploy stage won't start until your test stage finishes. If any job fails, subsequent stages are skipped (unless the failing job has allow_failure: true). This is what makes stages a natural fit for the build → test → deploy workflow that most software projects follow.

GitLab provides five default stages if you don't define your own: .pre, build, test, deploy, and .post. The dotted stages (.pre and .post) are special — they always run first and last, regardless of where they appear in your stages list. You can use .pre for environment setup and .post for cleanup, and they'll be respected no matter how you order the others.

# Defining custom stages stages: - .pre # Always runs first - validate # Custom stage - build # Custom stage - test # Custom stage - security # Custom stage - deploy # Custom stage - .post # Always runs last # Jobs assigned to stages validate_lint: stage: validate script: - npm run lint build_app: stage: build script: - npm run build test_unit: stage: test script: - npm run test:unit scan_security: stage: security script: - trivy scan . deploy_prod: stage: deploy script: - deploy.sh production # Default stages (if not defined) # .pre → build → test → deploy → .post # The default stage for jobs without a stage no_stage_job: script: - echo "This runs in the test stage" # Equivalent to: stage: test # Stage execution rules: # 1. Stages run in order of definition # 2. All jobs in a stage must finish before next stage # 3. A failure stops subsequent stages (unless allow_failure) # 4. .pre and .post are special and always run first/last

Sequential Execution

Stages run one after another. This guarantees ordering: build finishes before test begins.

Fail-Fast

If a job fails, subsequent stages are skipped. This stops broken code from reaching production.

.pre Stage

Always runs first. Use for environment setup, dependency download, or configuration.

.post Stage

Always runs last. Use for cleanup, notifications, or reporting, even after failures.
Stage Pitfalls:
  • Too many stages slow down pipelines unnecessarily
  • A single slow job in a stage blocks the entire next stage
  • Jobs without an explicit stage default to test
  • Failures in one job block all later stages (use allow_failure for non-critical jobs)
  • Stage names are case-sensitive and must match exactly
Job Parallelization

Within a single stage, all jobs run in parallel by default. This is one of the most powerful features of GitLab CI/CD — it means a test stage with five test jobs completes in roughly the time of the slowest single test, not the sum of all five. Parallelization is what makes GitLab pipelines fast, and understanding how to exploit it is essential for building efficient pipelines.

Beyond the automatic parallelization of independent jobs, GitLab also supports explicit parallelization within a single job definition. Using the parallel keyword, you can split one job into N identical copies that run simultaneously, each with access to $CI_NODE_INDEX and $CI_NODE_TOTAL environment variables. This is perfect for test suites that can be split across multiple runners — instead of running 1000 tests sequentially in one job, you run 10 jobs with 100 tests each, dramatically reducing wall-clock time.

The most advanced form of parallelization is the parallel: matrix keyword, which generates one job per combination of variables. This is ideal for testing across multiple language versions, operating systems, or configurations. A single matrix definition can generate dozens of jobs, each testing a different combination, all running in parallel.

# Default parallelization: independent jobs in a stage test: stage: test script: - npm test lint: stage: test script: - npm run lint # These two jobs run in parallel automatically # Explicit parallel: split one job into N copies parallel_test: stage: test parallel: 5 script: - echo "Node $CI_NODE_INDEX of $CI_NODE_TOTAL" - npm run test -- --shard=$CI_NODE_INDEX/$CI_NODE_TOTAL # Environment variables for parallel jobs: # $CI_NODE_INDEX - Index of this job (1-based) # $CI_NODE_TOTAL - Total number of parallel jobs # Matrix builds: one job per combination matrix_test: stage: test parallel: matrix: - NODE_VERSION: [16, 18, 20] OS: [ubuntu, alpine] image: node:$NODE_VERSION script: - echo "Testing on $OS with Node $NODE_VERSION" - npm test # The above generates 6 jobs: # - Node 16 on ubuntu # - Node 16 on alpine # - Node 18 on ubuntu # - Node 18 on alpine # - Node 20 on ubuntu # - Node 20 on alpine # Matrix with more variables comprehensive_matrix: stage: test parallel: matrix: - OS: [ubuntu, alpine, debian] NODE: [16, 18, 20] DATABASE: [postgres, mysql] script: - test.sh --os $OS --node $NODE --db $DATABASE # The above generates 18 jobs (3 × 3 × 2) # Matrix with include for extra variables matrix_with_include: stage: test parallel: matrix: - NODE: 16 DB: postgres - NODE: 18 DB: postgres EXTRA: "yes" - NODE: 20 DB: mysql script: - echo "Node: $NODE, DB: $DB, Extra: $EXTRA" # Testing parallel jobs locally # Each parallel job gets a unique $CI_NODE_INDEX # and the same $CI_NODE_TOTAL

Automatic Parallel

All jobs in a stage run in parallel by default. No configuration needed.

parallel: N

Creates N copies of a job. Use $CI_NODE_INDEX to differentiate.

parallel: matrix

One job per combination of variables. Generates jobs from a variable matrix.

Speed Gains

Parallel execution dramatically reduces wall-clock time for large test suites.
Parallelization Best Practices:
  • Split test suites using parallel: N for large suites
  • Use parallel: matrix for testing across multiple versions
  • Ensure parallel jobs are truly independent
  • Consider runner capacity before setting high parallel counts
  • Use interruptible: true to allow auto-cancellation on new commits
Dependencies: Passing Artifacts Between Jobs

By default, every job in a later stage downloads the artifacts produced by every job in every earlier stage. This is convenient — you don't have to think about it — but it becomes a problem in large pipelines. If a build stage produces 500MB of artifacts across five jobs, and your deploy stage has ten jobs that don't need any of them, you're wasting bandwidth and time downloading unnecessary files. The dependencies keyword solves this by letting you specify exactly which jobs' artifacts a job should download.

The dependencies keyword is important to understand in terms of what it does and doesn't do. It does control artifact download — a job only receives artifacts from the jobs listed in its dependencies. It does not control execution order — the stage ordering still applies, and the job won't run until its stage is reached regardless of dependencies. For controlling execution order, you need needs, which we'll cover next.

A common pattern is to use dependencies: [] (an empty list) when a job doesn't need any artifacts at all. This tells GitLab to download nothing, which speeds up the job. On the other hand, if you want a job to receive artifacts from a specific earlier job, list that job by name. The dependencies keyword gives you fine-grained control over data flow between stages, which is essential for optimizing large pipelines.

# Default behavior: download all artifacts from previous stages build_app: stage: build script: - npm run build artifacts: paths: - dist/ build_docs: stage: build script: - npm run build:docs artifacts: paths: - docs/ # This job downloads BOTH dist/ and docs/ (default) test_default: stage: test script: - ls dist/ # Available - ls docs/ # Available # This job downloads only dist/ (explicit dependencies) test_limited: stage: test dependencies: - build_app script: - ls dist/ # Available # - ls docs/ # NOT available # This job downloads nothing test_no_artifacts: stage: test dependencies: [] script: - echo "No artifacts downloaded" # dependencies does NOT affect execution order: # The stage ordering still applies # Multiple dependencies deploy: stage: deploy dependencies: - build_app - build_docs script: - ls dist/ - ls docs/ # dependencies vs artifacts: # - artifacts: what a job PRODUCES # - dependencies: what a job CONSUMES # # dependencies is a consumer-side control # artifacts is a producer-side declaration # Note: dependencies only works within the same pipeline # Artifacts from parallel jobs can also be dependencies # Cross-pipeline artifact sharing requires artifacts # from parent or child pipelines
Dependencies vs Needs:
  • dependencies: Controls which artifacts a job downloads
  • needs: Controls which jobs must complete before this job runs
  • dependencies affects data flow; needs affects execution order
  • You can use both together for maximum control
Needs (DAG): Breaking the Stage Barrier

The needs keyword introduces a fundamentally different execution model called "Directed Acyclic Graph" (DAG). In the standard stage-based model, all jobs in a stage wait for all jobs in the previous stage, even if they don't actually depend on each other. In a DAG pipeline, jobs declare their exact dependencies using needs, and GitLab runs them as soon as those dependencies complete — even if their own stage hasn't "officially" started yet.

To understand why this matters, imagine a pipeline with a "build" stage containing a fast job (30 seconds) and a slow job (5 minutes), followed by a "test" stage with a job that only needs the fast build job. In the standard model, the test job waits 5 minutes for the slow build job to finish, even though it doesn't need it. With needs, the test job runs as soon as the fast build job is done — saving 4.5 minutes of idle time. Multiply this across a large pipeline, and the savings become enormous.

When you use needs, you can optionally skip stage semantics entirely by omitting stages from your config or by mixing needs with stages. If needs references a job in the same stage, that's fine too — the DAG model is more flexible. The only rule is that the graph must be acyclic — you can't have circular dependencies. GitLab validates this and rejects pipelines with cycles.

# DAG example: needs breaks the stage barrier stages: - build - test - deploy # Fast job in build stage build_fast: stage: build script: - echo "Fast build (30s)" artifacts: paths: - fast-output/ # Slow job in build stage build_slow: stage: build script: - sleep 300 - echo "Slow build (5m)" # This test job only needs the fast build # With needs, it runs as soon as build_fast completes # WITHOUT waiting for build_slow test_early: stage: test needs: - build_fast script: - echo "Testing early (runs before build_slow finishes)" # This deploy job needs everything deploy: stage: deploy needs: - build_fast - build_slow - test_early script: - echo "Deploying" # DAG with multiple dependencies test_comprehensive: stage: test needs: - build_fast - build_slow script: - echo "Testing everything" # needs with parallel jobs parallel_job: stage: build parallel: 3 script: - echo "Parallel build $CI_NODE_INDEX" dependent_job: stage: test needs: - parallel_job # Waits for all 3 parallel jobs script: - echo "All parallel builds complete" # needs with optional optional_dependency: stage: test needs: - job: build_fast optional: true # Don't fail if build_fast is missing script: - echo "Optional dependency" # needs with artifacts # needs also downloads artifacts from the needed jobs # by default (equivalent to dependencies) # DAG best practices: # 1. Use needs to reduce pipeline duration # 2. Declare only real dependencies # 3. Combine needs with dependencies for artifact control # 4. Avoid circular dependencies # 5. Test the DAG with different failure scenarios # needs vs stages: # - Stages: all jobs in stage N must finish before stage N+1 # - Needs: jobs run as soon as their dependencies complete # - You can use both together # - Omitting stages + using needs = pure DAG # Pure DAG (no stages) build_a: script: echo "Build A" build_b: script: echo "Build B" test_a: needs: [build_a] script: echo "Test A" test_b: needs: [build_b] script: echo "Test B" deploy: needs: [test_a, test_b] script: echo "Deploy" # This pipeline has no stages — pure DAG
Standard
build_fast (30s) build_slow (5m)
Standard
test (waits 5m)
DAG
build_fast (30s) test (runs at 30s)
DAG
build_slow (5m)
Needs (DAG) Best Practices:
  • Use needs to parallelize jobs across stage boundaries
  • Only declare real dependencies to maximize parallelism
  • Combine needs with dependencies for artifact control
  • Use needs: [] to start a job immediately (no dependencies)
  • Consider needs with parallel: matrix for fast matrix testing
  • Validate that the DAG has no cycles
How Stages and Needs Work Together

Stages and needs can be combined in the same pipeline, and understanding how they interact is key to designing sophisticated pipelines. When you use needs, you're not throwing away stages — you're layering a finer-grained execution model on top of them. The stage ordering still applies as a default, but needs can override it for specific jobs.

The interaction is straightforward: if a job has needs, it starts as soon as those needs complete, regardless of stage boundaries. If a job doesn't have needs, it follows the standard stage-based rule. This means you can have a pipeline where most jobs follow stage order, but a few critical-path jobs use needs to run earlier. This hybrid model is very powerful and is used in many production pipelines.

One subtlety to be aware of: if you use needs with jobs in the same stage, the stage semantics are relaxed — GitLab allows jobs within the same stage to depend on each other via needs, which is not possible in the pure stage model. This lets you express intra-stage dependencies, which is useful for things like "test B only after test A passes."

# Combining stages and needs stages: - build - test - deploy # build stage (standard, runs in parallel) build_a: stage: build script: echo "Build A" artifacts: paths: [a.txt] build_b: stage: build script: echo "Build B" artifacts: paths: [b.txt] # test stage (standard, waits for build stage to complete) test_standard: stage: test script: echo "Standard test (waits for both builds)" # test stage with needs (runs early) test_a_only: stage: test needs: [build_a] # Runs as soon as build_a finishes script: echo "Test A only (early)" # test with same-stage needs test_b_after_a: stage: test needs: [build_a] # Runs after build_a, regardless of stage script: echo "Test B after A" # deploy stage (waits for test stage) deploy: stage: deploy needs: - test_standard - test_a_only - test_b_after_a script: echo "Deploy" # Summary of execution: # 1. build_a and build_b run in parallel (stage: build) # 2. test_a_only starts as soon as build_a finishes (needs) # 3. test_b_after_a starts as soon as build_a finishes (needs) # 4. test_standard starts after both builds complete (stage ordering) # 5. deploy starts after all tests complete (needs) # Pure stage model (no needs): # All jobs in each stage wait for all jobs in the previous stage # # Hybrid model (stages + needs): # Jobs with needs run as soon as their dependencies complete # Jobs without needs follow stage ordering # When to use stages vs needs: # - Use stages for coarse-grained ordering # - Use needs for fine-grained control # - Combine for optimal pipeline performance # - Pure DAG (no stages) for maximum parallelism # Visualizing with needs: # build_a ──┬─→ test_a_only ──┬─→ deploy # └─→ test_b_after_a ┤ # build_b ──→ test_standard ───┘ # Note that test_a_only and test_b_after_a both # depend on build_a, so they can run in parallel # even though they're in the "test" stage.
Choosing Between Models:
  • Simple pipelines: Use stages only (default model)
  • Complex pipelines: Use stages + needs for critical-path optimization
  • Maximum speed: Use pure DAG (needs only, no stages)
  • Mixed: Use stages for coarse ordering, needs for fine control
Complete Pipeline Example

Here's a complete, production-ready pipeline that demonstrates stages, parallelization, dependencies, and needs in a realistic workflow. This example is optimized to minimize pipeline duration using DAG execution.

# Complete .gitlab-ci.yml with optimized stages and needs stages: - .pre - build - test - security - deploy - .post variables: NODE_VERSION: "18" default: image: node:$NODE_VERSION cache: key: files: - package-lock.json paths: - node_modules/ # Stage: .pre (setup) setup: stage: .pre script: - echo "Setting up environment" # Stage: build (parallel jobs) build_app: stage: build script: - npm ci - npm run build artifacts: paths: - dist/ expire_in: 1 day build_docs: stage: build script: - npm ci - npm run build:docs artifacts: paths: - docs/ expire_in: 1 day # Stage: test (using needs to run early) test_unit: stage: test needs: [build_app] # Runs as soon as build_app finishes script: - npm ci - npm run test:unit artifacts: reports: junit: test-results/unit.xml parallel: 4 # 4 parallel test jobs test_integration: stage: test needs: [build_app] # Runs as soon as build_app finishes script: - npm ci - npm run test:integration artifacts: reports: junit: test-results/integration.xml test_e2e: stage: test needs: [build_app] # Runs as soon as build_app finishes script: - npm ci - npm run test:e2e artifacts: reports: junit: test-results/e2e.xml # Stage: security (parallel, needs nothing) sast: stage: security image: registry.gitlab.com/security-products/sast:latest script: - /analyzer run needs: [] # Starts immediately dependency_scanning: stage: security image: registry.gitlab.com/security-products/dependency-scanning:latest script: - /analyzer run needs: [] # Starts immediately # Stage: deploy (using needs for fine control) build_image: stage: deploy image: docker:latest services: - docker:dind needs: - test_unit - test_integration - test_e2e script: - docker build -t $CI_REGISTRY_IMAGE:$CI_COMMIT_SHORT_SHA . - docker push $CI_REGISTRY_IMAGE:$CI_COMMIT_SHORT_SHA only: - main deploy_staging: stage: deploy image: alpine/helm:latest needs: [build_image] script: - helm upgrade --install my-app ./charts/my-app --namespace staging --set image.tag=$CI_COMMIT_SHORT_SHA --wait environment: name: staging url: https://staging.example.com only: - main deploy_production: stage: deploy image: alpine/helm:latest needs: - build_image - deploy_staging script: - helm upgrade --install my-app ./charts/my-app --namespace production --set image.tag=$CI_COMMIT_SHORT_SHA --atomic --wait environment: name: production url: https://example.com when: manual only: - main # Stage: .post (cleanup) cleanup: stage: .post script: - echo "Cleaning up" when: always needs: [] # Runs after everything else # Execution flow summary: # 1. setup (.pre) runs first # 2. build_app and build_docs run in parallel # 3. test_unit, test_integration, test_e2e run as soon as build_app finishes # 4. sast and dependency_scanning run in parallel from the start # 5. build_image runs after all tests pass # 6. deploy_staging runs after build_image # 7. deploy_production waits for deploy_staging (manual) # 8. cleanup (.post) runs last, always
What This Pipeline Demonstrates:
  • Stages: .pre, build, test, security, deploy, .post
  • Parallel jobs: build_app and build_docs run together
  • needs: test jobs run as soon as build_app finishes
  • parallel: test_unit runs as 4 parallel jobs
  • artifacts: build outputs are passed to tests and deploy
  • cache: node_modules cached between jobs
  • needs: []: security scans start immediately
  • manual deploy: production requires manual approval
Frequently Asked Questions
What happens if a job in a stage fails?
By default, subsequent stages are skipped and the pipeline is marked as failed. Use allow_failure: true on a job to let the pipeline continue even if it fails.
What's the difference between needs and dependencies?
needs controls execution order (which jobs must complete first). dependencies controls artifact download (which jobs' artifacts this job receives). You can use both together.
Can I use needs without stages?
Yes! If you omit the stages keyword and only use needs, you get a pure DAG pipeline where jobs run based solely on their declared dependencies.
How does parallel work with needs?
A job with parallel: N creates N jobs. A job with needs pointing to a parallel job waits for all N copies to complete.
What's the maximum number of parallel jobs?
The limit depends on your runner capacity. On GitLab.com, it's typically limited by the number of concurrent jobs your plan allows. Self-hosted runners can be configured with concurrent in config.toml.
Can I have circular dependencies in needs?
No. GitLab validates the DAG and rejects pipelines with cycles. You must ensure your needs graph is acyclic.
How do I make a job wait for another job in the same stage?
Use needs to reference a job in the same stage. This creates an intra-stage dependency, which the pure stage model doesn't support.
What's the performance difference between stages and DAG?
DAG can dramatically reduce pipeline duration by starting jobs as soon as their dependencies complete, rather than waiting for the entire previous stage. Savings depend on how well your jobs parallelize.
Previous: .gitlab-ci.yml Explained Next: Variables & Secrets

Stages and jobs are the foundation of GitLab CI/CD pipelines. Master them to build fast, reliable, and maintainable pipelines that scale with your project.