---
title: "Actions Runner Controller on EKS: a reproducible setup"
description: "Deploy GitHub Actions runner scale sets on EKS with scoped credentials, separate worker capacity, a real job, and a complete teardown procedure."
url: https://9apes.com/blog/actions-runner-controller-eks/
---

# Actions Runner Controller on EKS: a reproducible setup

A runner pod reaching `Running` is only one part of setting up Actions Runner
Controller. The controller must authenticate to GitHub, a scale-set listener must
receive work, a worker node must have capacity, and the runner must execute a job.
You also need a way to remove everything when the experiment ends.

This walkthrough uses a disposable EKS cluster with one system node and separate
runner capacity. It uses GitHub's runner scale-set charts, a repository-scoped
credential, and a shell-only workflow. The
[lab configuration and evidence](/downloads/ci-lab/eks-arc-evidence.zip) accompany
the article. We use a dedicated cluster name and tags to identify the resources
that need to be removed during cleanup.

## What the lab provisions

The system node runs CoreDNS, the ARC controller and listener, and Karpenter.
Runner pods select a separate NodePool. Karpenter creates worker nodes when those
pods cannot be scheduled and removes eligible idle nodes afterward.

| Component | Lab configuration |
| --- | --- |
| Region | AWS `us-east-1` |
| EKS Kubernetes minor | `1.36` |
| EKS node AMI release | AL2023 `1.36.3-20260911` |
| System capacity | One `m7i.large`, fixed at one node |
| Worker capacity | On-demand `c7i.xlarge`, 12-vCPU and 24-GiB NodePool limits |
| ARC charts | Controller and scale set `0.14.2` |
| Runner image | `ghcr.io/actions/actions-runner:2.337.0` |
| Karpenter | `1.14.1` |
| Runner request and limit | One CPU, 2 GiB per runner |
| Scale-set bounds | Zero minimum idle runners, eight maximum runners |

The cluster has public subnets in two availability zones and no NAT gateway.
Worker public addresses provide outbound access to GitHub and image registries.
There is no SSH access or internet-facing Kubernetes Service. The public
Kubernetes API endpoint accepts only the administrator's current IPv4 `/32`;
the nodes also use private endpoint access.

That network choice keeps this short experiment inexpensive. It is a lab
configuration, not a recommendation to copy the same network design into a
production environment with different connectivity requirements.

## Use current runner scale sets

There are two charts to install: `gha-runner-scale-set-controller` and
`gha-runner-scale-set`. The first manages the Kubernetes resources. The second
configures a particular GitHub scope and runner pool.

Start with the current [GitHub ARC quickstart](https://docs.github.com/en/actions/tutorials/use-actions-runner-controller/get-started)
and pin the chart versions you actually test. Instructions based on older
`RunnerDeployment` and `HorizontalRunnerAutoscaler` resources describe a
different configuration surface.

The controller, listener, runner pod and EC2 worker each have a separate job:

1. The controller reconciles the desired runner resources.
2. The listener communicates with GitHub and reacts to assigned work.
3. A runner pod starts the Actions runner process and executes a job.
4. Karpenter supplies EC2 capacity when Kubernetes cannot place the pod.

This distinction matters when diagnosing a job that has not started. Reinstalling
ARC will not fix an EC2 capacity problem, and increasing a NodePool limit will
not fix a rejected GitHub credential.

## Prepare two separate credentials

The local operator needs AWS permissions to create and delete the dedicated
lab resources. ARC needs permission to register runners for the private lab
repository. These credentials serve different purposes.

The supplied example uses a dedicated fine-grained GitHub token selected only
for the lab repository, with repository Administration read/write. GitHub also
documents [GitHub App authentication](https://docs.github.com/en/actions/how-tos/manage-runners/use-actions-runner-controller/authenticate-to-the-api).
The credential becomes a Kubernetes Secret referenced by name. It is not placed
in Helm values, a workflow file or the runner container's environment.

```yaml
githubConfigUrl: https://github.com/9apes/ci-lab
githubConfigSecret: ci-lab-github
runnerScaleSetName: ci-lab-eks
minRunners: 0
maxRunners: 8
```

The lab provisioning scripts require a non-root AWS caller. An optional identity
bootstrap accepts a separate authorized administrator or SSO profile and creates
a temporary user that can assume only one dedicated provisioning role. These
IAM operations do not require account-root credentials. The role's inline policy
scopes CloudFormation to lab stacks, EKS to the lab cluster, IAM changes to lab
names, and network changes to tagged lab resources where the API supports those
conditions. The temporary user, key and role are deleted after cloud cleanup.

A reader with an existing suitable SSO role can use that role directly. Do not
create another credential merely to reproduce the optional bootstrap step.

## Run preflight before starting the meter

Extract the lab files under `examples/ci-lab` in a local working directory. Follow
the archive README to select the AWS profile and set the expected account ID,
GitHub token-file path and administrator CIDR. Set `LAB_GITHUB_REPO` to your own
`owner/repository` when reproducing the lab. `preflight.sh` performs read-only
checks before provisioning:

```bash
examples/ci-lab/eks/preflight.sh
```

It checks the caller, local file permissions, GitHub runner API access, EKS
version support, the pinned AMI and account quota inventory. A successful GET
against the runner API verifies read access; it does not replace checking the
token's Administration write permission.

Check existing usage as well as the quota. An account with a 64-vCPU quota does
not necessarily have 64 vCPUs available for a new experiment. Use an EKS version
still in [standard support](https://docs.aws.amazon.com/eks/latest/userguide/kubernetes-versions.html)
so that an old example does not unexpectedly incur extended-support pricing.

The supplied configuration was validated with eksctl `0.230.0` and Helm
`v3.19.0`. The archive records the tool and chart versions, while the live evidence
records the managed EKS platform version and resolved container images.

## Create the cluster and controllers

After reviewing the configuration and cost estimate:

```bash
examples/ci-lab/eks/bootstrap.sh
export KUBECONFIG="$PWD/examples/ci-lab/eks/.local/kubeconfig"
export PATH="$PWD/examples/ci-lab/eks/.local/bin:$PATH"
```

The script creates the dedicated Karpenter IAM stack, EKS cluster and fixed
system node group, then installs the Karpenter and ARC charts. The worker
NodePool remains separate from the system node group. Both the ARC controller
and its listener select `ci-lab-role: system`; runner pods select
`ci-lab-role: runner` and tolerate only the lab worker taint.

The important runner-template choices are visible in the configuration:

```yaml
template:
  spec:
    nodeSelector:
      ci-lab-role: runner
    automountServiceAccountToken: false
    containers:
      - name: runner
        image: ghcr.io/actions/actions-runner:2.337.0
        command: [/home/runner/run.sh]
        resources:
          requests:
            cpu: '1'
            memory: 2Gi
          limits:
            cpu: '1'
            memory: 2Gi
```

The complete file includes the worker taint toleration and security context.
This runner has no Docker daemon, host Docker socket or privileged sidecar. It
can execute shell steps. Adding container jobs or an image builder is a separate
configuration change with separate permissions and tests.

## Run a test workflow

Inspect each stage before sending a large burst:

```bash
kubectl get pods -n arc-system
kubectl get pods -n arc-runners
kubectl get ec2nodeclass ci-lab
kubectl get nodepool ci-lab
kubectl get nodes -o wide
```

With zero idle runners, `arc-runners` can be empty. In this configuration the
listener remains in `arc-system` waiting for work.

Then dispatch one job using `runs-on: ci-lab-eks`. The smoke workflow in the
archive records its first shell-command timestamp, CPU visibility and
architecture, holds the runner briefly, and exits. It needs no repository
checkout, external action or AWS credential.

`maxRunners` is a runner limit. A NodePool's CPU and memory limits govern a
different resource pool and are eventually consistent during provisioning.
Neither is an account-level spending cap. `minRunners: 0` also does not mean a
zero-cost cluster: the EKS control plane and fixed system node remain active.

For startup analysis, retain the GitHub workflow creation timestamp, job
metadata, first-command timestamp and Kubernetes node/pod observations.
Workflow creation to first command includes scheduling, capacity provisioning,
image pull and runner initialization. Calling the whole interval pure queue
time obscures the part you need to improve.

## The smoke result

The initial verification ran on 15 September 2026 UTC.
GitHub recorded workflow creation at `22:52:50Z`; the first shell
command ran at `22:54:05.256Z`. The job completed successfully after the
20-second synthetic hold.

| Observed property | Result |
| --- | --- |
| Workflow creation to first command | 75.3 seconds in this single smoke run |
| Runner architecture | `x86_64` |
| Host CPUs visible to the container | 4 |
| Cgroup CPU quota / period | `100000 / 100000`, equivalent to one CPU |
| Cgroup memory limit | `2147483648` bytes, equivalent to 2 GiB |
| Worker instance type | `c7i.xlarge` |
| Workflow source revision | `5fd84bb81bf1fc0966a2878deaf6d5ab3b2c922d` |

Treat this as an end-to-end verification sample. The 75.3-second interval
includes worker startup and runner initialization; repeated trials are needed
to compare capacity policies. The distinction between four visible host CPUs
and a one-CPU quota is also why recording only `nproc` or `getconf` is
insufficient to establish what CPU allocation a CI job received.

## Reconstruct from the frozen recipe

We removed the first cluster before trying the recipe again. The retained
inventory at `01:42:11Z` on 16 September showed no lab EKS cluster, active
instances, volumes, VPCs, network interfaces, launch templates, active stacks or
instance profiles. We retained the dedicated provisioning identity for the
second run.

The reconstruction used a separate copy of the 26 source files, excluding
generated configuration, kubeconfig, results and previous local state. Its
per-file hashes match both the frozen manifest and the supplied recipe. The
manifest fingerprint is
`5f5d9ef28c5f456912ee5a5d444c00eb7dfd9c21c650782a856167c323965e4d`.
Existing verified tools, the AWS provisioning profile and the repository-scoped
GitHub credential remained external prerequisites; this was not a new-account
bootstrap.

The second EKS control plane was created on 16 September 2026 UTC at
`01:45:39.371Z`. The bootstrap log
records the cluster ready at `02:04:48Z`, followed by deployed Karpenter, ARC
controller and runner scale-set releases. One new smoke job then exercised the
reconstructed installation:

| Replay observation | Result |
| --- | --- |
| GitHub workflow run | `35046687226` |
| Workflow created | `2026-09-16T02:06:20Z` |
| First shell command | `2026-09-16T02:07:05.710401992Z` |
| Workflow creation to first command | 45.7 seconds |
| Initial worker nodes, NodeClaims and ready runners | Zero of each |
| Job completed | `2026-09-16T02:07:26Z`, success |
| Requested synthetic hold | 20 seconds |

The job reported the same `x86_64` architecture, four visible host CPUs,
one-CPU cgroup quota and 2-GiB memory limit as the initial smoke. The evidence
includes the resolved runner image digest and checks the command markers
against the retained job log. The workflow source revision was unchanged.

These two individual smoke runs verify that jobs executed on both installations.
Their different startup times are not a performance comparison. The complete
scaling experiment was not repeated during reconstruction. See
`eks/results/replay.json` and `eks/results/replay-source-manifest.json` in the
evidence bundle for the reconstruction record and source hashes.

## Diagnose the stage that failed

| Symptom | Check first |
| --- | --- |
| Controller pod cannot start | System-node capacity, image pull and its Kubernetes service account |
| Listener reports authentication errors | Repository scope, token permissions, expiry and Secret reference |
| Runner pods remain Pending | Node selector, taint toleration, requests and NodePool status |
| NodeClaims exist but nodes do not join | EC2NodeClass conditions, node IAM role, EKS access entry and network access |
| Runner starts but the workflow fails immediately | Requested execution mode and the tools actually in the runner image |
| Fresh provisioning fails with AccessDenied | The exact AWS action and resource, including existing-resource checks during create operations |

Several provisioning failures are worth recording because a broader admin policy
would hide them. CloudFormation can truncate generated physical names, so IAM
patterns must match the actual dedicated resource prefix. EC2 create operations
can also authorize both the new resource and an existing VPC: request-tag checks
on a new subnet do not authorize use of its parent VPC. The supplied policy
keeps those permissions separate.

The default security group created by EKS has its own
[cluster ownership tags](https://docs.aws.amazon.com/eks/latest/userguide/sec-group-reqs.html).
It does not automatically inherit the example's `Project` tag. The policy therefore
authorizes the required ingress changes using the exact EKS cluster-name tag,
while continuing to reject a security group belonging to another cluster.

Finally, eksctl enables CloudFormation termination protection. A failed stack
can finish rollback and still reject deletion until that protection is disabled.
Keep the disable permission scoped to the lab stacks, inspect rollback status,
and verify that the failed resources are removed before retrying.

A managed node group with a launch template can also fail before launching a
node because EKS checks the caller's `ec2:RunInstances` permission. The supplied
policy restricts that check to the lab's selected image, system-node instance
type and AWS service-mediated requests. The complete policy lives in the
download rather than being scattered through a sequence of broad console grants.

Inspect the failure before changing permissions. The fact that a policy blocks a
request does not establish that every action for the entire service is needed.

## Remove workers before their controller

After canceling unfinished lab jobs, run:

```bash
examples/ci-lab/eks/teardown.sh
```

The sequence is deliberate: remove the runner scale set, delete the worker
NodePool and wait for NodeClaims to disappear while Karpenter and its IAM
permissions still exist. Delete the EC2NodeClass so that its generated instance
profile can be released, then remove the controllers, cluster and dedicated
Karpenter stack.

Deleting the controller first can leave you without the component responsible
for completing worker termination. Removing finalizers to make Kubernetes
objects disappear is not proof that the EC2 instances are gone.

Verify the EKS cluster is absent, the dedicated CloudFormation stacks are deleted,
and no lab instances, volumes, interfaces, public addresses, launch templates or
instance profiles remain. Only then delete the temporary provisioning identity
and its local credential file. The repository token can be revoked when all lab
experiments are finished.

The planned eight-hour estimate uses the selected instance rates, EKS
standard-support pricing, public addresses, storage and a traffic allowance.
Keep the experiment supervised and complete teardown within the documented
session deadline. A useful tutorial should leave behind reproducible evidence,
not an unexplained recurring bill.

Our reconstruction passed its job check, but the session was interrupted before
the second teardown. The original cleanup deadline, `2026-09-16T05:44:21Z`, was
missed. A successful smoke test therefore does not certify cleanup or prove the
planned runtime cost. The replay evidence records this limit explicitly. Tracked infrastructure
removal was separately verified on September 17 at 19:18 UTC through resource
inventories and completed CloudFormation deletions. The cleanup record lists
the scoped role's inventory limits and the temporary provisioning identity,
whose removal still requires renewed administrator authentication. A collector
deadline stops new trials; it does not itself delete a running cluster.

## ARC on EKS or 9apes?

ARC on EKS gives you control over the runner image, network and worker capacity
in your AWS account. Your team also owns cluster upgrades, runner availability,
scaling failures and the bill for idle infrastructure. Running it in production
needs engineers with dedicated responsibility for that infrastructure. As usage
grows, that can mean a dedicated platform team handling maintenance and incidents.
That investment can make sense when your workloads need the control.

With 9apes, managing runners is simpler because we operate the infrastructure.
You connect the GitHub App, choose a runner and update your workflow's `runs-on`
label. Your team manages the workflow, its dependencies and secrets, while we
handle the runner infrastructure. You do not need to build a team to operate
an EKS cluster for those jobs. To try it with one workflow, follow the
[9apes quickstart](/docs/quick-start/).
