Deploying a service on ECS (Amazon Elastic Container Service) has always been a two-tool job. You start the rollout with aws ecs update-service, then you go somewhere else to find out if it worked: describe-services for the deployment state, the service event log, CloudWatch Logs for the task, and CloudTrail if none of those explain it. You keep the state machine in your head and poll it.

This month ECS started drawing that state machine for you, in the console. That part is good news, and I will cover it first.

Then the part the console cannot fix. In my own test environment, a small app of two Fargate services, a rollout reported success three different ways while nothing was running. And when I went back to check the command most pipelines use to wait for a deploy, aws ecs wait services-stable, it passed on a service with zero tasks. I reproduced that on purpose, and the result is below.

What the new Deployments tab gives you

Open any ECS service in the console and go to the Deployments tab. For services on the native deployment types (rolling, linear, canary, and blue/green), it shows the deployment as it happens:

  • Phase, service events, and task progress on one timeline, live. You can see whether it is starting new tasks, waiting on a lifecycle hook, or sitting in bake time, without reading JSON.
  • How much traffic has moved from the old revision to the new one. For canary and blue/green, that is the number you care about during the rollout.
  • Health signals in one place: the circuit breaker (the ECS feature that stops a deployment after too many task failures and can roll it back), with live failure counts against its threshold; alarm state; container and load balancer health checks; and lifecycle hooks.
  • Failed tasks with the reason and a link into CloudTrail, so “why did it die” and “who changed what” are one click away.

It costs nothing extra and is in all commercial Regions and GovCloud.

QuestionBeforeNow
Which phase is it in?Read deployments[].rolloutStateThe timeline shows the phase
How much traffic has moved?Work it out from load balancer statsShown in the timeline
Did the circuit breaker trip?See rolloutState = FAILED and guessCircuit breaker status with thresholds
Did the alarm fire?Check CloudWatch separatelyAlarm state on the timeline
Why did the task die?CloudWatch Logs, then CloudTrailFailed tasks with the reason and CloudTrail links

The old loop looked like this, three query shapes where status and rolloutState are easy to mix up under pressure:

1
2
3
4
5
aws ecs update-service --cluster my-cluster --service my-service --force-new-deployment
aws ecs describe-services --cluster my-cluster --services my-service \
  --query 'services[0].deployments[].[status,rolloutState,desiredCount,runningCount,failedTasks]'
aws ecs describe-services --cluster my-cluster --services my-service \
  --query 'services[0].events[:10].[createdAt,message]'

Keep the scope in mind. This is deployment observability: what happened to one rollout. It does not replace your metrics, traces, or logs. It is not an API you can query later, and it is not a deployment history.

  • Pro: if you deploy by hand, it replaces three polling commands and a lot of guessing.
  • Con: it lives only in the console. Nothing in a pipeline can read it, and it only covers the native deployment types. Services still on the CodeDeploy controller show nothing.

My setup, and why it is the worst case

My test app is two Fargate services, each running one small task (0.5 vCPU, 1 GB). They are Slack agents that connect out over Socket Mode, so there is no load balancer and no HTTP port. The health check is a no-op, and each service allows the old task to stop before the new one starts:

1
2
3
4
circuitBreaker: { enable: true, rollback: true },
minHealthyPercent: 0,   // one task; there is no room for two
maxHealthyPercent: 100,
healthCheck: { command: ["CMD-SHELL", "true"] }, // no port to check

With one task and minHealthyPercent: 0, the rollout is the outage window. And with a health check that always passes, the circuit breaker is the only thing that can notice a crash loop. So this setup leans entirely on the deployment machinery. That is exactly why the next part hurt.

Three ways a rollout reported success with nothing running

These all happened on one evening in my test environment. The times are local.

1. The infrastructure code reset the task count on every deploy.

My CDK code had desiredCount: 0. I had put it there so the services would ship “dark”, and I scaled them up by hand afterwards. CloudFormation writes the value from the code back on every deploy, so every release quietly scaled both services back to zero:

18:45:00  task definition registered
18:45:06  has stopped 1 running tasks
18:46:25  deployment completed
18:46:25  has reached a steady state      <- desiredCount 0, runningCount 0

Every line is true. The deployment did complete. The service is in a steady state. It is also down, and it stays down until someone notices. The console timeline shows this as a clean finish, because by its own definition it is one.

The fix is in the code: put the real count in the source, and take a service down by changing it there. A count you set by hand undoes itself on the next release.

It got worse before it got better. My fix for this was later undone by accident, when a stale copy of the file was copied over the fixed one and committed with a message that only described the other change in it. Nothing in the pipeline noticed, because nothing in the pipeline checked the count.

2. The error message named the wrong cause.

After I scaled the services back up, ECS could not place a task:

18:59:51  was unable to place a task. The reason for failure is
          Capacity is unavailable at this time. Please try again later
          or in a different availability zone.

“Capacity is unavailable” sends you to check your task size and subnets. So before touching anything, I built a control: a test that removes my changes from the picture. A plain busybox task, no volumes, no secrets, same size, same subnets and security group, in a different cluster. I ran it once in each AZ (availability zone, one of the separate data centers in a Region):

TestResult
busybox task, other cluster, AZ aRunTask -> ServerException: Internal Error
busybox task, other cluster, AZ bsame
A request that is invalid on purposeInvalidParameterException, a normal validation error
Create a network interface directly in each AZsucceeded
Fargate vCPU quota4,000, with 1.5 in use

The invalid request mattered most. ECS answered it with a correct validation error, so the API was up and reading requests. Only the step that places a valid task was failing, with a server error, in both AZs, for a task that had nothing to do with my change. The message the service showed me was a guess about the cause, not a diagnosis. I could not check the AWS Health dashboard for my account, because that needs a Business support plan.

3. Silence was not recovery, and not failure either.

18:59:51  unable to place a task
          ... 46 minutes of nothing ...
19:45:52  deployment failed: tasks failed to start
19:48:50  has started 1 tasks
19:52:00  deployment completed / steady state

For 46 minutes the service posted no new events. ECS backs off after repeated placement failures, and it also stops repeating the same event, so a service stuck in backoff looks exactly like a quiet, healthy one. The 19:45 line is the circuit breaker doing its job, and it was the first honest signal in three quarters of an hour.

One more thing I did not expect: the service scheduler and the RunTask API are not the same path. My other service placed its task at 19:33, while my own run-task probes at 19:32 and 19:34 were still getting ServerException. I had set up a watcher that waited for run-task to succeed before nudging the services. It would have waited forever while the scheduler was already placing tasks. A synthetic probe built on run-task can be red while your services are fine, and, less comfortably, green while they are not.

The pipeline check that passes on zero tasks

Most pipelines wait for an ECS deploy like this:

1
aws ecs wait services-stable --cluster my-cluster --services my-service

I went and read what that waiter actually checks. In botocore, the library behind the AWS CLI, its success condition is:

length(services[?!(length(deployments) == `1` && runningCount == desiredCount)]) == `0`

In words: one deployment, and the running count equals the desired count. 0 equals 0. So I built exactly that case in a test account and ran it: one ECS service with desiredCount: 0, deployed with CDK, then a forced new deployment, then the waiter. No task ever ran.

10:37:31.9  force-new-deployment   -> new PRIMARY deployment, IN_PROGRESS, desired 0, running 0
10:38:34    services-stable        -> exit 0 after 63 s
            describe-services      -> desired 0, running 0, PRIMARY still IN_PROGRESS
10:38:59    ECS event              -> deployment completed, steady state

The waiter said “stable” with zero tasks running, and it said so 25 seconds before ECS itself called the deployment complete. A pipeline gated on this would have gone green on the exact failure from Finding 1.

The same condition also passes after a circuit breaker rollback. Once ECS rolls back, there is one deployment again (the old revision) and its counts match. I read that from the waiter’s definition and did not trigger a rollback, so treat that one as from the source, not executed.

Here is the check I use now. It fails unless the new revision really finished with the number of tasks I meant to run:

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
#!/usr/bin/env bash
# usage: check-rollout.sh <cluster> <service> <expected-task-def-arn> <expected-count>
set -euo pipefail
cluster=$1 service=$2 want_td=$3 want_n=$4
read -r state td running desired < <(aws ecs describe-services --cluster "$cluster" --services "$service" \
  --query 'services[0].deployments[?status==`PRIMARY`] | [0].[rolloutState, taskDefinition, runningCount, desiredCount]' \
  --output text)
fail=0
[ "$state" = "COMPLETED" ]   || { echo "FAIL: rolloutState is $state"; fail=1; }
[ "$td" = "$want_td" ]       || { echo "FAIL: running ${td##*/}, expected ${want_td##*/} (rolled back?)"; fail=1; }
[ "$desired" -ge "$want_n" ] || { echo "FAIL: desiredCount is $desired (did the IaC reset it?)"; fail=1; }
[ "$running" -ge "$want_n" ] || { echo "FAIL: runningCount is $running"; fail=1; }
exit "$fail"

On the same test service, right after the waiter passed:

primary: rolloutState=IN_PROGRESS running=0 desired=0 taskDef=...:1
FAIL: rolloutState is IN_PROGRESS, not COMPLETED
FAIL: desiredCount is 0, expected at least 1 (IaC reset it?)
FAIL: runningCount is 0, expected at least 1
exit 1

Run it after the waiter, or in a loop with a timeout of your own. The waiter tells you the service stopped moving. It does not tell you the service is doing what you meant.

The test setup, the check, and the raw output are in yopa-experiments.

Pitfalls

  • “Steady state” is not “serving”. A deployment can complete, reach steady state, and leave runningCount at zero. Check the count against the number you intended, not rolloutState alone.
  • wait services-stable passes on 0/0, and after a rollback. Pair it with a check like the one above.
  • Don’t scale a service by hand if your IaC also sets the count. The next deploy puts the code’s number back.
  • Treat the reason in a service event as a guess. Before you change anything, run a control that removes your change: a plain task, somewhere unrelated.
  • Silence is not good news. No new events can mean ECS is backing off, not that things are fine.
  • Don’t build a health probe on run-task and assume it tells you about your services. They took different paths in my test.
  • minHealthyPercent: 0 gives you a gap whose length you don’t control. On a normal day mine is about a minute. That evening it was fifty minutes, and the deployment had reported success in under two.

What to do

  • If you deploy by hand: open the Deployments tab before you start guessing at describe-services output. The phase is right there.
  • If you deploy from a pipeline: this launch changes nothing for you, because nothing in a pipeline can read the console. Check what your pipeline actually waits on. If it is wait services-stable alone, add a check on the primary deployment’s state, its task definition, and the running count.
  • If your IaC sets desiredCount: make it the real number, and change it in code.

For running a real service on ECS end to end, including version pinning, secrets, and rollback, see Run n8n Queue Mode on AWS ECS with Recoverable Operations. For a single-task agent on Fargate, see how I first deployed an AI agent to ECS Fargate.

References