BK
← Writing
AIEngineering LeadershipInfrastructure

The control boundary for agents in software delivery

The important decision is not whether an agent can deploy. It is where the company chooses to put authority, review, and production access.

If I were deciding how a company should introduce agents into software delivery, I would start with a governance question before a tooling question: where do we actually want the agent to have authority? An agent can technically be given a kubeconfig, an Argo CD token, or access to an internal deployment API. That does not mean it should be. The architecture should assume that agents will sometimes misunderstand intent, produce an incomplete change, or make a locally reasonable decision that is wrong for the larger system.

For that reason, I would put the agent on the proposal side of the deployment boundary rather than the production side. In a GitOps environment, the agent can interpret a request, inspect the repository and the operational context it is allowed to see, propose a narrow change, and open a pull request. CI and policy systems validate the artifact, humans or automated approval rules decide whether it should merge, and Argo CD remains responsible for reconciling the approved desired state into the cluster. This is less about distrusting AI than it is about preserving a control model the organization already understands.

The architecture I would want looks roughly like this:

issue / alert / deployment request
              |
              v
         agent runner
              |
              |  proposes a change
              v
        branch + commit
              |
              v
         pull request
              |
       +------+------+
       |             |
       v             v
   CI checks     review / policy
       |             |
       +------+------+
              |
              v
            merge
              |
              v
        config repository
              |
              v
           Argo CD
              |
              v
       Kubernetes cluster

The strategic point is what is deliberately missing from the left half of that diagram. The agent does not have a kubeconfig, it does not call kubectl apply, and it does not need permission to call argocd app sync. Argo CD already knows how to watch Git and reconcile the application after the desired state changes, so I would keep execution inside that existing control plane. The Argo CD CI documentation already treats updating Git as the normal deployment path, and an explicit sync call is unnecessary when automated synchronization is enabled.

At the implementation level, a simple Application for this pattern could look like this:

apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
  name: payments
  namespace: argocd
spec:
  project: production

  source:
    repoURL: https://github.com/example/platform-config.git
    targetRevision: main
    path: environments/prod/payments

  destination:
    server: https://kubernetes.default.svc
    namespace: payments

  syncPolicy:
    automated:
      enabled: true
      prune: false
      selfHeal: true
    syncOptions:
      - ApplyOutOfSyncOnly=true

I would probably start with prune: false for an agent-authored path unless there were a strong reason not to. Updating an image or changing a resource limit is one class of action; deciding that a resource should no longer exist is a different blast radius. That distinction matters at the policy level because one of the easiest mistakes in agent adoption is granting an entire capability just because the platform supports it.

The config repository itself does not need to become more complicated because an agent is involved. I would prefer to keep the structure boring and make the agent operate inside an explicit surface area:

platform-config/
  apps/
    payments/
      base/
        deployment.yaml
        service.yaml

  environments/
    staging/
      payments/
        kustomization.yaml

    prod/
      payments/
        kustomization.yaml

  ci/
    validate-agent-diff.py

For a routine deployment, the agent may only need to change an image reference. That is a much easier action to reason about organizationally than granting a generic "deploy" permission because the proposed state is reviewable before anything reaches production:

apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization

resources:
  - ../../../apps/payments/base

images:
  - name: ghcr.io/example/payments
    newName: ghcr.io/example/payments
    newTag: 9f13c42

I would also treat agent-created changes as a distinct trust class in CI. The point is not that an agent is inherently malicious. The point is that generation is cheap, so it is easy to produce a syntactically valid change whose blast radius is much larger than the request that triggered it. A company-wide agent strategy should define the kinds of changes agents are allowed to propose, not just the systems they can technically reach.

For example, even a very small policy check can demonstrate the boundary:

# ci/validate-agent-diff.py

from pathlib import Path
import subprocess
import sys
import yaml

ALLOWED_PREFIX = "environments/prod/"
BLOCKED_KINDS = {
    "ClusterRole",
    "ClusterRoleBinding",
    "CustomResourceDefinition",
    "Namespace",
}

changed = subprocess.check_output(
    ["git", "diff", "--name-only", "origin/main...HEAD"],
    text=True,
).splitlines()

for filename in changed:
    if not filename.startswith(ALLOWED_PREFIX):
        sys.exit(f"agent may not modify {filename}")

    if not filename.endswith((".yaml", ".yml")):
        continue

    with Path(filename).open() as f:
        for doc in yaml.safe_load_all(f):
            if not doc:
                continue

            kind = doc.get("kind")
            if kind in BLOCKED_KINDS:
                sys.exit(f"agent may not create or modify {kind}")

            pod_spec = (
                doc.get("spec", {})
                   .get("template", {})
                   .get("spec", {})
            )

            for container in pod_spec.get("containers", []):
                security = container.get("securityContext", {})
                if security.get("privileged") is True:
                    sys.exit(
                        f"agent may not enable privileged mode in {filename}"
                    )

print("agent diff is inside the allowed deployment surface")

That is not meant to be a complete Kubernetes security policy. In a mature environment I would rather express durable rules in an existing policy system, but the principle is the important part: the agent gets a defined change surface. It should not be able to turn "update the payments image" into "also modify a cluster role and create a privileged container" just because both actions happen to be possible through the same repository.

The corresponding CI job can remain conventional, which is another reason I like this design. The organization does not have to invent a second deployment system for agent-authored changes; it can reuse the same validation machinery and add narrower controls where needed:

name: validate deployment change

on:
  pull_request:
    paths:
      - "environments/**"

permissions:
  contents: read
  pull-requests: read

jobs:
  validate:
    runs-on: ubuntu-latest

    steps:
      - uses: actions/checkout@v4
        with:
          fetch-depth: 0

      - name: Install Python dependencies
        run: pip install pyyaml

      - name: Apply agent-specific policy checks
        if: contains(github.event.pull_request.labels.*.name, 'agent-authored')
        run: python ci/validate-agent-diff.py

      - name: Render production manifests
        run: |
          kubectl kustomize environments/prod/payments > /tmp/rendered.yaml

      - name: Client-side Kubernetes validation
        run: |
          kubectl apply \
            --dry-run=client \
            -f /tmp/rendered.yaml

The agent workflow itself should have Git permissions, not deployment permissions. In the simplest version it can create a branch and a pull request, but it cannot bypass branch protection or directly mutate cluster state. That keeps the model replaceable and prevents the agent framework from becoming part of the production control plane:

name: deployment agent

on:
  workflow_dispatch:
    inputs:
      request:
        description: "What deployment change should be proposed?"
        required: true

permissions:
  contents: write
  pull-requests: write

jobs:
  propose:
    runs-on: ubuntu-latest

    steps:
      - uses: actions/checkout@v4

      - name: Run deployment agent
        run: |
          python tools/deployment_agent.py \
            --request "${{ inputs.request }}"

      - name: Create branch and commit
        run: |
          BRANCH="agent/deploy-${GITHUB_RUN_ID}"

          git config user.name "deployment-agent"
          git config user.email "deployment-agent@example.com"

          git switch -c "$BRANCH"
          git add environments/
          git commit -m "Propose deployment change"
          git push origin "$BRANCH"

          gh pr create \
            --base main \
            --head "$BRANCH" \
            --title "Agent-proposed deployment change" \
            --body "Generated from: ${{ inputs.request }}" \
            --label "agent-authored"
        env:
          GH_TOKEN: ${{ github.token }}

I left the model-specific implementation inside deployment_agent.py on purpose. Whether the company uses one model provider, an internal model, or a more elaborate tool-calling framework should not change the deployment trust model. If the AI layer can be swapped without redesigning production access, the architecture is less coupled to today's agent stack.

The same separation applies when the agent needs operational context. GitHub Actions can use OIDC to obtain short-lived cloud credentials instead of storing long-lived keys, and the role can be limited to the metrics, logs, or metadata the agent needs to inspect. The GitHub OIDC documentation is useful here because the workflow identity can be constrained by repository and workflow claims rather than sharing a permanent AWS credential.

I would also make the approval model risk-based rather than treating every agent-authored change the same way. Image updates and bounded configuration changes may eventually be good candidates for automatic approval if the organization has enough confidence in its tests and policy layer. Cluster-scoped RBAC, namespace deletion, CRDs, network policy changes, production database changes, and anything that materially expands privileges should have a different path. The company should decide those boundaries deliberately rather than letting the agent framework define them by accident.

For multi-step deployments, I would keep sequencing declarative where possible instead of teaching the agent an imperative production runbook. Argo CD already supports hooks and sync phases, so a migration that must happen before an application rollout can remain part of the declared deployment:

apiVersion: batch/v1
kind: Job
metadata:
  name: payments-migrate
  annotations:
    argocd.argoproj.io/hook: PreSync
    argocd.argoproj.io/hook-delete-policy: HookSucceeded
spec:
  template:
    spec:
      restartPolicy: Never
      containers:
        - name: migrate
          image: ghcr.io/example/payments:9f13c42
          command: ["./bin/migrate"]

Argo CD runs PreSync hooks before applying the normal sync resources and stops the sync if the hook fails. From a leadership perspective, that is the kind of separation I want: the agent can decide that the desired state should include a migration, but the existing deployment system owns sequencing, failure behavior, and reconciliation.

The broader company decision is not whether agents are capable of deploying software. They are. The more important question is whether introducing agents requires us to discard the controls, auditability, and separation of responsibility we already built around production. I do not think it should.

If the agent is wrong in this model, the failure mode is usually a failed check, a rejected pull request, or a proposed change that never merges. That is a much better organizational default than turning a reasoning error directly into a production incident. The technical design follows from that leadership choice: let agents interpret intent and accelerate proposals, but keep production authority inside systems whose behavior and blast radius the company already knows how to govern.