Vision is our AI product photography app, and it runs Laravel on AWS ECS. One container image runs as four services: web, queue worker, image generation worker and scheduler. Traffic reaches it only through Cloudflare, and the instances need no NAT gateway. Every deploy is zero-downtime. In the last 60 days a healthy target served traffic 99.98% of the time.

The app

Vision turns product photos into lifestyle images for online stores. It is our own product, so we pick the stack and carry the pager.

The app is Laravel 13 on Octane with FrankenPHP. It stores data in Postgres and keeps its cache and API rate limits in a Redis-compatible store. Image generation runs in background jobs that call Google’s AI APIs.

What we needed

  • Deploys without downtime. Customers start generation jobs at any hour, and a deploy must not drop them.
  • A small attack surface. The servers should not answer the internet directly.
  • Infrastructure we can rebuild. Dev and prod should come from the same code.
  • No static cloud keys in the app’s environment.

The architecture

Architecture: visitors reach Cloudflare, which forwards only to the ALB. The ALB forwards only to ECS on EC2, where web, queue, worker and scheduler run in a public subnet with no NAT and no SSH. The team signs in through AWS SSO and SSM Session Manager. RDS Postgres and Valkey sit in private subnets.

One image, four ECS services

CI builds one image per release and stores it in Amazon ECR. ECS runs it as four services with different commands:

  • web serves HTTP through Octane
  • queue runs Laravel’s queued jobs
  • generation worker runs the long image jobs on their own
  • scheduler runs Laravel’s scheduled commands, started by EventBridge

The generation worker runs apart from the other queued jobs, so long image jobs scale on their own.

Cloudflare in front, no NAT behind

Requests go through Cloudflare to an Application Load Balancer and on to the ECS tasks. The load balancer’s security group accepts ports 80 and 443 only from Cloudflare’s published IP ranges. Anyone who finds the load balancer’s address still cannot reach it directly.

The ECS instances sit in public subnets and make outbound calls through the VPC’s internet gateway. That removes the NAT gateway and its hourly and per-gigabyte charges. A public subnet does not mean an open server. The instances’ security group accepts traffic from the load balancer only, and no SSH port is open. We sign in through AWS SSO and reach an instance or the database through SSM Session Manager.

Managed data stores

Postgres runs encrypted on RDS in private subnets with 7-day backups. The cache is Valkey on ElastiCache, also in a private subnet. Secrets Manager holds the credentials, and ECS injects them into each task at start.

Infrastructure as code

The whole stack is OpenTofu: VPC, load balancer, ECS cluster, IAM, database, cache, registry and Cloudflare DNS. Dev and prod are two environments built from the same modules. CI itself runs on AWS CodeBuild, connected to GitLab as a runner.

No static keys for Google

Vision calls Google’s APIs from inside AWS. The ECS task role signs a Workload Identity Federation token, and Google trusts that token instead of a stored service-account key. Nothing in the environment can leak a long-lived Google key.

Zero-downtime deploys

A release runs in a fixed order:

Deploy order: CodeBuild builds one image into ECR, all four ECS services roll onto it while old tasks keep serving, then a one-off task runs the migrations. Across releases a schema change goes expand, migrate, contract.

  1. CI rolls all four ECS services onto the new image. ECS starts new tasks before it stops the old ones.
  2. Only then does a one-off task run php artisan migrate --force.

The app containers never run migrations at start. For a few minutes the new code runs against the old schema, so every migration has to work in that window. That makes every schema change a multi-release change, known as the expand/contract pattern:

  1. Expand. Add the new column or table. Old code ignores it, and new code can use it once the migration has run.
  2. Migrate. Move the data and switch the code to the new shape.
  3. Contract. Drop the old column or table in a later release, once no deployed code reads it.

A release never renames or drops something the running code still reads.

One incident added a second rule. A scheduled command shipped in the same release as the migration that created its table. It fired before the migration ran and failed with relation does not exist. Now new code never reads a table created in the same release.

The result

Measured over the last 60 days:

  • 99.98% of the time a healthy target served traffic. There were four five-minute windows with no healthy target, so at most 20 minutes of hard downtime. Two of them came from one crash-loop incident on 2 September 2026.
  • Zero-downtime deploys. Releases roll onto new tasks while the old ones keep serving.

Run your Laravel or Symfony app on AWS

We set up the same stack for client apps, in Laravel or Symfony. See Laravel and Symfony on AWS ECS for what the setup includes, or tell us about your app.

Frequently asked questions

Laravel on AWS ECS

How do you run Laravel on AWS ECS?

Build one container image and run it as several ECS services. Vision runs a web service, a queue worker, a second worker for image generation and a scheduler from the same image. An Application Load Balancer sends web traffic to the tasks. RDS holds the database, and a Valkey cache sits next to it.

How do Laravel migrations run on ECS without downtime?

Run them as a one-off ECS task after the new image is live, never at container start. Every migration is backward-compatible, so new code works against the old schema for the few minutes in between. Columns get added in one release and dropped in a later one.

Do ECS tasks need a NAT gateway?

Not when the instances sit in public subnets. Vision's instances make outbound calls through the VPC's internet gateway, so there is no NAT gateway to pay for. Their security group accepts traffic only from the load balancer, and the load balancer accepts traffic only from Cloudflare. No SSH port is open: we reach the instances through AWS SSO and SSM Session Manager.

Does the same setup work for Symfony?

Yes. A Symfony app with Messenger workers fits the same pattern: one image, a web service, worker services and scheduled tasks. Both frameworks deploy with far fewer moving parts than a Magento store.

Share