BK
← Writing
EngineeringInfrastructure

A quick mental model for AWS ECS

ECS gets much easier to reason about once you separate clusters, services, tasks, task definitions, and the EC2 instances underneath them.

I've been spending more time with AWS ECS lately, and the terminology is much more intimidating than the actual model. There are clusters, container instances, services, tasks, task definitions, target groups, and a bunch of different places in the AWS console where all of these things seem to live. Once I stopped treating them as one giant ECS blob, it got a lot easier to understand.

The simplest mental model I have is this: a task definition describes what to run, a task is an actual running copy of that definition, and a service makes sure the number of tasks you asked for keeps running. All of that happens inside a cluster. If you're using the EC2 launch type, the cluster has EC2 machines registered to it as container instances, and ECS decides which machine should run each task. AWS also has Fargate now, which lets you skip managing those EC2 instances yourself.

Here is roughly how I picture the EC2 version:

                         AWS ECS Cluster
                              |
              +---------------+---------------+
              |                               |
      EC2 Container Instance          EC2 Container Instance
        (ECS agent)                      (ECS agent)
              |                               |
         +----+----+                     +----+----+
         |         |                     |         |
       Task      Task                  Task      Task
         |         |                     |         |
    Container  Container            Container  Container


Task Definition
      |
      +---- image
      +---- CPU / memory
      +---- ports
      +---- environment
      +---- logging
      |
      v
   ECS Service
      |
      +---- desired count: 4
      +---- keeps 4 Tasks alive
      +---- places them in the Cluster
      +---- can attach them to a load balancer

If this is a web service, there is usually another chain in front of it. An Application Load Balancer receives traffic, forwards it through a target group, and ECS registers the tasks created by the service as targets. So the request path looks more like this:

Internet
   |
   v
Application Load Balancer
   |
   v
Target Group
   |
   v
ECS Service
   |
   +----------+----------+
   |          |          |
 Task       Task       Task
   |          |          |
Container  Container  Container

The task definition is probably the most important object to understand first. It is basically the blueprint for a task. It tells ECS which Docker image to run, how much CPU and memory the containers need, which ports they expose, which environment variables they get, where logs go, and what IAM role the task should use. Registering a new task definition creates a new revision rather than changing the old one in place, which is useful because a service can be moved from one revision to another during a deployment.

A task is the running thing. One task can technically contain multiple containers that need to live together, although for a basic service I usually find it easier to think of one application container per task. The important distinction is that you normally do not care about a specific task for very long. Tasks die, deployments replace them, and services create new ones. The service is the durable object that says, "I want three copies of this task definition running." ECS tries to keep reality matching that desired state.

Clusters are simpler than the name makes them sound. With the EC2 launch type, a cluster is basically the logical pool of EC2 capacity ECS can place tasks onto. Each EC2 instance runs the ECS agent, which talks to the ECS control plane and lets the scheduler know what resources are available. With Fargate, you still use a cluster, but AWS handles the machines underneath it, so the cluster becomes much more of a logical boundary.

For infrastructure I still prefer declaring this stuff in CloudFormation rather than creating it manually in the console. A very stripped-down Fargate service might look something like this. I'm using Fargate here mostly because it keeps the example focused on ECS itself rather than also needing an Auto Scaling group and ECS-optimized EC2 AMI.

AWSTemplateFormatVersion: '2010-09-09'

Parameters:
  ContainerImage:
    Type: String
  SubnetA:
    Type: AWS::EC2::Subnet::Id
  SubnetB:
    Type: AWS::EC2::Subnet::Id
  SecurityGroup:
    Type: AWS::EC2::SecurityGroup::Id
  TargetGroupArn:
    Type: String

Resources:
  Cluster:
    Type: AWS::ECS::Cluster

  LogGroup:
    Type: AWS::Logs::LogGroup
    Properties:
      LogGroupName: /ecs/example-service
      RetentionInDays: 14

  TaskExecutionRole:
    Type: AWS::IAM::Role
    Properties:
      AssumeRolePolicyDocument:
        Version: '2012-10-17'
        Statement:
          - Effect: Allow
            Principal:
              Service:
                - ecs-tasks.amazonaws.com
            Action:
              - sts:AssumeRole
      ManagedPolicyArns:
        - arn:aws:iam::aws:policy/service-role/AmazonECSTaskExecutionRolePolicy

  TaskDefinition:
    Type: AWS::ECS::TaskDefinition
    Properties:
      Family: example-service
      RequiresCompatibilities:
        - FARGATE
      NetworkMode: awsvpc
      Cpu: '256'
      Memory: '512'
      ExecutionRoleArn: !GetAtt TaskExecutionRole.Arn
      ContainerDefinitions:
        - Name: web
          Image: !Ref ContainerImage
          Essential: true
          PortMappings:
            - ContainerPort: 8080
          LogConfiguration:
            LogDriver: awslogs
            Options:
              awslogs-group: !Ref LogGroup
              awslogs-region: !Ref AWS::Region
              awslogs-stream-prefix: web

  Service:
    Type: AWS::ECS::Service
    Properties:
      Cluster: !Ref Cluster
      TaskDefinition: !Ref TaskDefinition
      LaunchType: FARGATE
      DesiredCount: 2
      NetworkConfiguration:
        AwsvpcConfiguration:
          AssignPublicIp: DISABLED
          SecurityGroups:
            - !Ref SecurityGroup
          Subnets:
            - !Ref SubnetA
            - !Ref SubnetB
      LoadBalancers:
        - ContainerName: web
          ContainerPort: 8080
          TargetGroupArn: !Ref TargetGroupArn

There are a couple of Fargate-specific things hiding in that template. Fargate tasks use awsvpc networking, which means each task gets networking in your VPC rather than sharing the host's network namespace. The task definition also has task-level CPU and memory values, and the service uses LaunchType: FARGATE instead of relying on EC2 container instances. Those were some of the main differences AWS called out when Fargate was introduced.

In a real stack, the VPC, subnets, security groups, Application Load Balancer, listener, target group, ECR repository, alarms, and autoscaling rules would probably be CloudFormation too. I left most of that out because it obscures the ECS relationship I'm actually trying to show. The container image can live in ECR, the task definition points at that image, the service points at the task definition, and the service runs tasks in the cluster.

One thing I initially found confusing is that ECS itself does not build your image. Your build pipeline still creates a Docker image and pushes it somewhere like ECR. ECS starts from the already-built image. That separation is useful because it means application packaging and application scheduling are different concerns, even though the AWS console makes it easy to mentally mash them together.

Deployments also make more sense once the objects are separated. A typical deployment can be: build a new image, push it to ECR, register a new revision of the task definition with the new image tag, and update the ECS service to use that revision. ECS then replaces the old tasks with new ones while trying to maintain the service's desired count and deployment health. There are more sophisticated ways to do this, but that basic sequence explains a lot of what is happening underneath the UI.

What I like about ECS is that you get container orchestration without having to operate quite as much machinery yourself. It is definitely very AWS-shaped, and there are still a lot of nouns, but it gets a lot less confusing once you know what each piece is responsible for. For me, the useful shortcut is still: definition describes it, task runs it, service keeps it running, cluster gives it somewhere to run.