Buildkite agents: tokens, queues, one-job agents and autoscaling

Buildkite runs the control plane and you run the agents. You create a cluster and a queue, copy the cluster's agent token into buildkite-agent.cfg on your machine, start the agent with a queue= tag, and target that queue from your pipeline with agents: queue: .... For autoscaling, Buildkite maintains the Elastic CI Stack for AWS and the Agent Stack for Kubernetes.

How Buildkite agents work

The Buildkite agent is a single Go binary, buildkite-agent. It opens an outbound HTTPS connection to Buildkite, polls for work and runs jobs on the host where you installed it. Buildkite opens no inbound connections to your machines, so agents work behind NAT and firewalls without port forwarding.

Agents belong to a cluster, and inside a cluster each agent listens on one queue. A pipeline step names the queue it wants, and Buildkite hands the job to any connected agent on that queue. Buildkite also sells hosted agents, but this guide covers the self-hosted path, where you own the machines, the images and the network.

When self-hosting makes sense

Self-hosted agents fit when you need hardware Buildkite's hosted agents do not offer (GPUs, large-memory hosts, specific CPU architectures), when builds must reach private networks or databases, or when you want warm caches on local disk between builds. You also control the toolchain: the agent runs whatever you install on the host or bake into the container image you pick.

The price is operations work. You patch the hosts, size the fleet, clean up disk and keep secrets off long-lived machines.

Prerequisites: cluster, queue and agent token

Buildkite organizations created after the release of clusters on February 26, 2024 can only use clustered agents. Older organizations may still have unclustered agent tokens, which Buildkite has deprecated. Set things up in this order:

  1. In the Buildkite UI, open Agents and create a cluster (or use the default one).
  2. Inside the cluster, create a self-hosted queue, for example linux-medium-x86. Agents that start without a queue tag join the cluster's default queue. If the cluster has no default queue, the agent fails to connect.
  3. Open the cluster's Agent Tokens page and click New Token. Buildkite shows the token value once. Copy it into your secret store right away; if you lose it, you create a new one.

Tokens you create in the UI do not expire. Tokens you create through the REST or GraphQL API can carry an expiry date, which is useful for short-lived automation. A token belongs to one cluster and cannot connect agents to a different cluster.

Install the agent on Linux

On Ubuntu or Debian, add Buildkite's signed apt repository and install the package. These commands come from the official Ubuntu instructions:

curl -fsSL https://keys.openpgp.org/vks/v1/by-fingerprint/32A37959C2FA5C3C99EFBC32A79206696452D198 \
  | sudo gpg --dearmor -o /usr/share/keyrings/buildkite-agent-archive-keyring.gpg

echo "deb [signed-by=/usr/share/keyrings/buildkite-agent-archive-keyring.gpg] https://apt.buildkite.com/buildkite-agent stable main" \
  | sudo tee /etc/apt/sources.list.d/buildkite-agent.list

sudo apt-get update && sudo apt-get install -y buildkite-agent

Write the token into the config file and start the systemd service:

sudo sed -i "s/xxx/INSERT-YOUR-AGENT-TOKEN-HERE/g" /etc/buildkite-agent/buildkite-agent.cfg
sudo systemctl enable buildkite-agent && sudo systemctl start buildkite-agent

The package keeps its config in /etc/buildkite-agent/buildkite-agent.cfg and checks out code under /var/lib/buildkite-agent/builds/. On other distributions, or when you lack root, the install script puts everything under your home directory:

TOKEN="INSERT-YOUR-AGENT-TOKEN-HERE" bash -c "`curl -sL https://raw.githubusercontent.com/buildkite/agent/main/install.sh`"
~/.buildkite-agent/bin/buildkite-agent start

That variant reads ~/.buildkite-agent/buildkite-agent.cfg. Once the agent starts, it appears under the cluster's queue in the Buildkite UI within a few seconds.

Configure buildkite-agent.cfg

The config file uses key="value" lines. Each key also has an environment variable and a CLI flag (tags maps to BUILDKITE_AGENT_TAGS and --tags). A typical file for a Linux build host:

token="file:///etc/buildkite-agent/token"
name="%hostname-%spawn"
tags="queue=linux-medium-x86,os=linux,arch=amd64"
spawn=2
build-path="/var/lib/buildkite-agent/builds"
hooks-path="/etc/buildkite-agent/hooks"

Notes on the keys you will touch most:

  • token: the cluster agent token. A file:// prefix reads it from a file, and fd:// reads it from an inherited file descriptor, so the secret stays out of the config file itself.
  • tags: a comma-separated list of key=value pairs. The queue tag decides which cluster queue the agent joins. You can also set the queue with the queue key or --queue; that value overrides the tag.
  • spawn: the number of agent processes to run in parallel on this host (default 1). Each process runs one job at a time.
  • name: supports %hostname, %spawn, %random and %pid placeholders.

Restart the service after edits with sudo systemctl restart buildkite-agent.

Target a queue from your pipeline

Steps choose agents through the agents attribute. Set it at the top level of pipeline.yml to apply a default, and override it per step:

agents:
  queue: "linux-medium-x86"

steps:
  - label: "Unit tests"
    command: "make test"

  - label: "GPU benchmarks"
    command: "make bench"
    agents:
      queue: "gpu-a10"

A job with no matching agent sits in the scheduled state until one connects to that queue. Extra tags (os=linux, arch=arm64) narrow the match inside a queue; keep the number of tag combinations small. Each combination fragments your capacity and makes queue-based autoscaling harder to reason about.

One-job agents with --disconnect-after-job

Long-lived agents reuse the same filesystem, Docker cache and process table between jobs. That speeds up builds and leaks state between them. For untrusted code or reproducible builds, run each agent for one job and throw the machine away afterwards:

buildkite-agent start \
  --disconnect-after-job \
  --tags "queue=linux-ephemeral"

In the config file, the equivalent is disconnect-after-job=true (environment variable BUILDKITE_AGENT_DISCONNECT_AFTER_JOB). Pair it with disconnect-after-idle-timeout, a value in seconds, so an agent that never receives a job also exits instead of sitting idle. Your provisioning layer watches for the agent process to exit and then destroys the VM or pod.

If you already know which job a machine should run, buildkite-agent start --acquire-job <job-id> starts an agent that runs that one job. Schedulers that react to Buildkite webhooks use this to avoid races between agents.

Elastic CI Stack for AWS

The Elastic CI Stack for AWS is Buildkite's maintained autoscaling setup. You deploy it into your AWS account with CloudFormation or Terraform, and it creates an Auto Scaling group with a launch template whose instance count follows your queue's build activity. It supports Linux and Windows instances, and Mac through the CloudFormation setup.

The parameters you set first:

ParameterPurposeDefault
BuildkiteAgentToken or BuildkiteAgentTokenParameterStorePathCluster agent token, inline or from SSM Parameter Storenone
BuildkiteQueueQueue the agents joindefault
InstanceTypesUp to 25 comma-separated EC2 typest3.large
MinSize / MaxSizeInstance count floor and ceiling0 / 10
OnDemandPercentageShare of on-demand versus Spot instances100
BuildkiteTerminateInstanceAfterJobTerminate the instance after one jobfalse
ScaleInIdlePeriodSeconds all agents on an instance must idle before scale-in600

Store the token in Parameter Store instead of pasting it into the template. Set BuildkiteTerminateInstanceAfterJob to true for queues that run pull requests from forks. Lowering OnDemandPercentage moves capacity to Spot, which cuts cost but lets AWS reclaim instances mid-build, so add retries to steps that run on Spot.

Agent Stack for Kubernetes

The Agent Stack for Kubernetes (agent-stack-k8s) is a controller that watches your queue through Buildkite's Agent API. For each job it creates a Kubernetes Job with one pod: init containers copy the agent binary and check images, a checkout container clones the repo, and your step's containers run the commands. Each pod handles one job and goes away afterwards, which gives you one-job isolation by default.

Install it with Helm 3.8.0 or newer:

helm upgrade --install agent-stack-k8s oci://ghcr.io/buildkite/helm/agent-stack-k8s \
  --namespace buildkite \
  --create-namespace \
  --set agentToken=<buildkite-cluster-agent-token> \
  -f values.yaml

With a values.yaml that sets the queue:

config:
  tags:
    - queue=kubernetes

The controller polls the kubernetes queue unless you set another one. Use agentStackSecret instead of agentToken to reference an existing Kubernetes Secret, and tune the controller's max-in-flight setting (default 25) to cap concurrent jobs. On controller 0.30.0 and later, a step picks its container with the image attribute:

steps:
  - label: "Tests in a pod"
    agents:
      queue: kubernetes
    image: "golang:1.23"
    command: "go test ./..."

Older controllers use the kubernetes plugin with a podSpec block, which also covers sidecars, resource requests and node selectors.

Security hardening

  • Separate queues by trust. Run pull requests from forks on a queue with one-job agents and no deploy credentials. Keep release and deploy steps on a different queue whose agents hold the secrets.
  • Lock down what agents run. --no-command-eval stops the agent from running arbitrary step commands (only scripts already on disk), --no-plugins blocks plugins, and --allowed-repositories takes regular expressions for the repositories an agent may clone.
  • Inject secrets with hooks. Put an environment hook in hooks-path that fetches secrets from your secret manager per job, so nothing sensitive sits in pipeline YAML or the agent config.
  • Keep the token out of shell history. Use file://, SSM Parameter Store or a Kubernetes Secret, and rotate the token by creating a new one and revoking the old one in the cluster's Agent Tokens page.

Troubleshooting

Agent fails to connect at startup

Check that the token belongs to the cluster you expect and that nobody has revoked it. If you started the agent without a queue tag and the cluster has no default queue, it refuses to connect. Add queue=... to tags or create a default queue.

Jobs stay in "scheduled"

No connected agent matches the step's agents block. Compare the step's queue and tags with the tags shown on the agent's page in the UI. A typo in the queue name is the usual cause.

Disk fills up on long-lived agents

Checkouts under build-path and Docker images accumulate. Add a cron job for docker system prune, or move the queue to one-job agents on fresh VMs.

Builds cancelled mid-step on Spot

AWS reclaimed the instance. Raise OnDemandPercentage for that queue or add an automatic retry rule (retry: with automatic: true) to affected steps.

Cost trade-offs

Buildkite bills for its platform; the compute is yours. A fleet of always-on VMs gives you warm caches and zero boot time but bills you around the clock. Autoscaling with MinSize=0 on the Elastic CI Stack or pod-per-job on Kubernetes cuts idle cost, and you pay for it with boot latency for each fresh instance and with cold caches. Budget engineering time too: someone has to update AMIs, agent versions and the Helm chart, and debug the autoscaler when jobs stop picking up.

FAQ

Can I run several Buildkite agents on one machine?

Yes. Set spawn in buildkite-agent.cfg (or --spawn) to start that many agent processes on the host. Each process runs one job at a time and shares the host's disk and Docker daemon with the others.

Do Buildkite agent tokens expire?

Tokens created in the UI do not expire. Tokens created through the REST or GraphQL API can have an expiry set at least 10 minutes in the future.

Can one agent listen to several queues?

No. In a cluster, each agent joins one queue. Run separate agent processes or hosts for separate queues.

Is there a Buildkite equivalent of GitHub's ephemeral runners?

Start the agent with --disconnect-after-job and destroy the host when it exits, or use the Agent Stack for Kubernetes, which runs each job in its own pod. The setup mirrors ephemeral GitHub Actions runners and GitLab runners with one job per VM.

Running Buildkite agents with cirunner.dev

cirunner.dev provisions a fresh VM per job, registers it with your CI using a token you provide, scales with the queue and destroys the machine after the job. You choose CPU, RAM, x86_64 or arm64, region, base image and an optional GPU, and pay per compute minute.

The service is in early access with GitHub Actions and GitLab CI as launch platforms. Buildkite is on the roadmap; vote for it on the early-access form if you want one-job Buildkite agents without running the Elastic CI Stack yourself.

Skip the runner fleet

cirunner.dev boots a fresh VM for every Buildkite job, registers it, and destroys it when the job ends. Buildkite is on our roadmap. Vote for it on the early-access form.

Join early access