Category Archives: Cloud Services

Where Your Secrets Really Live: From IaC State to a Running Container

The credential leaks that actually make it into a postmortem tend to follow the same shape. A secret that was supposed to exist in exactly one place ends up existing in four: a .tfvars file someone forgot to delete, a CI log that echoed an environment variable during a failed-deploy debugging session, a state file with the value sitting in plaintext because Terraform needed it to create a resource, and the running container, which is the only place it was actually meant to be.

Secrets Management!

None of the tooling in this space is really about generating or storing a string securely. That part’s not hard. What’s hard is controlling how many places that string touches on its way from “created” to “used by a workload,” and closing off every one of those places except the last.

The map

A secret’s actual path through a modern deployment crosses four systems, and each one leaks it in its own particular way: the infrastructure-as-code tool that provisions the store, the CI/CD or GitOps engine that ships the workload, the store itself, and the runtime platform the workload runs on.

Secret Management Flow

Almost every incident I’ve dug into traces back to one of those four boundaries being noticeably weaker than the other three. Hardening one while ignoring the rest doesn’t close the leak, it just relocates it.

Infrastructure as code: the tool that provisions the vault, imperfectly

CloudFormation and Terraform dominate this layer for AWS-centric shops, and they fail in genuinely different ways.

CloudFormation’s answer is dynamic references: a string like {{resolve:secretsmanager:secret-id:SecretString:json-key}} that CloudFormation resolves at deploy time instead of you typing a password into a template. Pair that with the NoEcho attribute on the parameter and the value won’t print in the console or in describe-stacks. It’s a solid pattern with one sharp edge: NoEcho hides the value, it doesn’t encrypt it, and it offers no protection against someone who already has stack-read permissions. CloudFormation also won’t resolve a dynamic reference inside resource metadata like AWS::CloudFormation::Init, because that would print the secret straight to the console anyway. It’s a small but telling detail about where the real exposure risk actually sits. One advantage CloudFormation has by default: there’s no separate state file for you to secure, because the state lives inside the service.

Terraform’s version of this problem is structural rather than incidental. Mark a variable sensitive = true and Terraform redacts it from your terminal, your plan output, and your CI logs. It does nothing to the state file. Any sensitive value that becomes a resource attribute — a database password, an environment variable on a Lambda function — lands in terraform.tfstate in plain text regardless of how you flagged it, because Terraform needs the real value there to detect drift. Explaining that distinction to a team that assumed the sensitive flag was doing more than it actually does is a conversation most Terraform users eventually have.

For years the honest answer was: encrypt your state backend, restrict who can read it, and treat the state file itself as sensitive data. That’s still correct, but it’s no longer the whole story. Terraform 1.10 introduced ephemeral values, and 1.11 extended that to write-only arguments on managed resources, and together they’re an actual fix rather than a mitigation. An ephemeral resource fetches a value (a secret from Vault, a freshly generated password) and Terraform uses it during the apply without ever writing it to plan or state. A write-only argument, password_wo instead of password, accepts that ephemeral value directly on a resource like aws_db_instance, and the value is gone the moment the apply finishes. If you’re standing up new Terraform configurations this year, there’s not much reason to keep pulling secrets through a data source and hoping your state encryption is strong enough. The one thing worth planning for: migrating an existing attribute to its _wo form can trigger a resource replacement depending on the provider, so treat that migration like a credential rotation and test it somewhere that isn’t production.

CI/CD and GitOps are solving two different secret problems

People use “CI/CD” and “GitOps” almost interchangeably, but they hand secrets to infrastructure in opposite directions, and that difference matters more than which specific tool you’ve standardized on.

A CI/CD pipeline (GitHub Actions, Jenkins, Azure DevOps, or Spacelift running your Terraform) is push-based. Something happens, a runner spins up somewhere outside your infrastructure, that runner needs credentials to reach into your cloud account and your secret store, and then it pushes a change. The secret problem here is entirely about the runner: how does this ephemeral, often shared compute prove who it is without you handing it a long-lived key it might leak in a stray log line?

GitOps (Argo CD, Flux) is pull-based. An agent already lives inside the cluster and continuously reconciles actual state against what’s declared in Git. There’s no discrete pipeline-run moment to inject a secret from outside; the reconciler has to fetch it itself, from inside the cluster, and Git can never hold the plaintext value because Git is exactly where everyone with repo access can see it. That’s a different problem than “authenticate a runner,” which is why GitOps secret tooling looks nothing like CI/CD secret tooling.

On the CI/CD side, most of the industry has converged on the same answer: stop handing runners long-lived cloud credentials at all. GitHub Actions federates over OIDC: the workflow requests a short-lived token from GitHub’s own identity provider, and an IAM role’s trust policy accepts that token only from specific repositories, branches, or environments. No access key sits in GitHub Secrets waiting to be exfiltrated by a compromised dependency. Azure DevOps gets most of the way there with federated service connections, though the pattern I still see most often is a variable group linked to an Azure Key Vault: the group stores secret names rather than values and pulls the live value from the vault at queue time, so a vault outage fails the run early instead of mid-deploy. One thing that catches people off guard the first time: a variable group can’t link to a Key Vault that uses Azure RBAC for its permission model — the vault has to be running the older access-policy model instead. If secrets change often enough that a queue-time pull isn’t fresh, the AzureKeyVault@2 task reads the vault mid-run instead, at the cost of an extra step in every pipeline.

Jenkins is the odd one out because it predates all of this, and the right way to hand it secrets depends entirely on where the agents actually run. Agents on Kubernetes should use Vault’s Kubernetes auth method and let the pod’s own service account token do the authenticating. Agents on long-lived VMs typically use AppRole instead. The Vault plugin for Jenkins handles the retrieval and masks secrets in the build log, but “the plugin masks it” is a claim worth being a little skeptical of. An advisory against the plugin found that masking silently failed for secrets printed from shell steps on an agent when a specific durable-task logging mode was enabled; a workaround shipped in a related credentials plugin, but the underlying issue was still open at the time the advisory was published. The lesson generalizes past this one CVE: log masking is a convenience feature, not a security boundary, and you should assume that anything a build step can print, eventually it will.

Spacelift, which is closer to purpose-built CI/CD for Terraform than a general-purpose runner, handles this with contexts: reusable bundles of environment variables and mounted files attached to one or many stacks. Mark a variable or a mounted file as secret and it’s hidden from the UI and the API, visible only to the run itself. It composes nicely with Terraform’s own write-only arguments — Spacelift keeps the value out of its dashboard, Terraform keeps it out of state, and neither tool has to do the other’s job.

On the GitOps side, three patterns cover almost every cluster I’ve come across, and they trade off differently enough that picking one is a real decision rather than a coin flip.

ApproachWhere the secret livesRotationBest fit
Sealed SecretsEncrypted in Git, decrypted only by an in-cluster controllerManual, no external store to pollSmall teams with no existing secret store, comfortable with cluster-key lock-in
SOPSEncrypted in Git via KMS or ageManual re-encryption on changeConfig that spans Kubernetes and non-Kubernetes systems
External Secrets OperatorReferenced in Git; the real value stays in Vault, Secrets Manager, or Key VaultAutomatic, on a sync intervalTeams that already run a real secret store and want rotation to just work

I’ve moved teams off Sealed Secrets more than once, and it’s rarely about the encryption being weak. It’s that the cluster’s private key becomes a single point of failure people forget about until they’re migrating clusters and discover nothing sealed two years ago can be unsealed anywhere else. External Secrets Operator avoids that by keeping Git as a set of pointers instead of a vault substitute, which is honestly the more defensible GitOps posture: Git should describe which secret a workload needs, not contain it in any form.

The stores themselves: what you’re actually paying for

AWS Secrets Manager, Azure Key Vault, and HashiCorp Vault all do the core job (encrypted storage, access control, an audit trail) well enough that the choice usually comes down to which cloud you’re already committed to, plus one real conceptual difference.

Secrets Manager and Key Vault are, underneath the branding, secured key-value stores with a rotation feature bolted on. That’s not a criticism. For a team running in one cloud, storing mostly static credentials and rotating them on a schedule, that’s exactly the right amount of tool, and you get IAM or Entra ID integration for free instead of standing up a separate identity story.

Vault’s dynamic secrets are the actual differentiator, and they change the security model rather than just the storage location. Instead of storing a long-lived database password and rotating it on a schedule, Vault’s database engine creates a brand-new database user with a short lease every time a workload asks for one, and revokes it the moment the lease expires or the workload shuts down. Most of the time, there is no standing credential to steal, because most of the time there isn’t a credential at all. That’s a genuinely stronger posture, but it comes with real operational weight: an engine to configure per database type, lease-renewal logic your applications need to handle, and a Vault cluster that now needs its own high availability, because your workloads can’t start without it. I’d only take that on if you’re actually going to use dynamic secrets for something: databases, cloud IAM, PKI certificates. If you’re going to end up storing mostly static values and checking a rotation box, you’ve built a more complicated Secrets Manager, not a better one.

Getting the secret into a running container

This is where the container platform’s own identity model decides how much of the pain above you actually feel.

ECS keeps it simple by splitting responsibility across two IAM roles that get conflated constantly the first few times someone sets this up: the execution role, which the ECS agent assumes to pull the image from ECR and to resolve any secrets block in the task definition (calling Secrets Manager or SSM Parameter Store before the container even starts), and the task role, which the application itself assumes at runtime for everything else. Put the secret permission on the task role instead of the execution role and the container just fails to start, with an error message that doesn’t always point at which role is missing what.

EKS and AKS have more moving parts because Kubernetes wasn’t originally built with cloud IAM in mind, so there’s a federation step before you even reach the secret. On EKS, that step used to mean IRSA exclusively (mapping a Kubernetes service account to an IAM role through the cluster’s OIDC provider), and IRSA still works and remains the practical choice for Fargate and Windows node groups. But EKS Pod Identity, generally available since late 2023, has become the default recommendation for new EC2-based clusters: no per-cluster OIDC provider to wire up, a universal trust policy instead of one keyed to each cluster’s issuer URL, and IAM session tags for namespace and service account applied automatically. Once that identity exists, the actual secret retrieval is usually External Secrets Operator syncing a value from Secrets Manager or Vault into a native Kubernetes Secret on an interval, or the Secrets Store CSI driver mounting it as a volume without ever materializing a Kubernetes Secret object. That’s a real difference if you’re trying to keep secrets out of etcd entirely. AKS follows the same shape with its own managed identity and CSI driver combination, including autorotation that polls Key Vault every two minutes by default and updates the mounted file and any synced Secret without a pod restart.

If you’re running EKS on mostly EC2 node groups and still wiring up IRSA for new workloads out of habit, that’s worth a second look. Pod Identity isn’t dramatically more secure by itself, but it removes an entire category of trust-policy maintenance that scales badly past one or two clusters.

Rotation: the step most projects only half-finish

Storing a secret securely and rotating it are different problems, and plenty of “secrets management” initiatives quietly solve only the first one.

AWS Secrets Manager’s rotation runs through a Lambda function that Secrets Manager calls in four steps on a schedule, and the mechanics explain why it can be close to zero-downtime instead of a coordinated cutover.

AWS Secret Rotation

Secrets Manager tracks versions with staging labels instead of overwriting anything in place. AWSCURRENT is whatever your application is using right now. A new version is created and labeled AWSPENDING while createSecret and setSecret do the actual work of generating a new credential and pushing it to the database or service, and only once testSecret confirms the new credential logs in does finishSecret flip the labels: AWSPENDING becomes the new AWSCURRENT, and the old current version becomes AWSPREVIOUS. Nothing reading the secret mid-rotation ever sees a half-updated value, because the label that matters doesn’t move until the new credential is proven to work.

Vault handles rotation by mostly avoiding the need for it. Dynamic secrets carry a short lease instead of a rotation schedule, so “rotation” for a database credential is really Vault refusing to hand out a lease longer than its TTL and issuing a fresh one on the next request. That’s arguably the cleaner model, but it only covers secrets Vault generates itself; anything stored as a static key-value pair rotates exactly as manually as it would anywhere else.

The IaC layer has its own rotation trap, and both major tools share it. CloudFormation’s dynamic references and Terraform’s data-source lookups both resolve a secret’s value once, at deploy time, and neither one watches the store for changes afterward. Pin a CloudFormation dynamic reference to a specific version-id and a background rotation in Secrets Manager will never touch your stack, not until you push a template change that touches the resource. That’s exactly why AWS’s own guidance is to use versionless references, so a stack update always picks up whatever AWSCURRENT happens to be. It’s a small detail that’s generated more than one confused ticket asking why a rotation “didn’t take.”

Hardening: shrinking the blast radius

Everything above assumes the access paths are already reasonably tight. Hardening is about making sure a compromised credential, a leaked log, or an overprivileged role can’t turn into anything worse than it already is.

Network path matters more than people give it credit for. Secrets Manager supports a VPC interface endpoint, and once that endpoint exists, a secret’s resource policy can add an aws:SourceVpce condition that denies any request not arriving through it, collapsing the attack surface from “anyone with the right IAM permissions from anywhere” down to “traffic that physically transited this one endpoint inside the VPC.” AWS also recommends BlockPublicPolicy: true on any identity allowed to attach resource policies, which uses AWS’s own automated reasoning to reject a resource policy granting broad or public access before it ever takes effect, rather than relying on someone catching it in review.

Identity is the other half of this. Every federation pattern described above, and every IRSA or Pod Identity association, exists for the same reason: a short-lived, narrowly scoped credential that expires on its own beats a long-lived one that has to be remembered, stored, and eventually rotated by a person. A workload or pipeline still authenticating with a static access key at this point is usually not a deliberate choice, it’s just the oldest unresolved thing in the account.

And then there’s the secret that never should have left a laptop. Gitleaks and TruffleHog cover this from two angles. Gitleaks runs fast, offline pattern matching that’s cheap enough as a pre-commit hook to block a leak before it’s even recorded in history. TruffleHog goes further and verifies whether a detected credential is still live with a real, read-only call against the provider it belongs to, which matters when triaging which of the dozens of things a full history scan just turned up are actual emergencies. Running both, one at commit time and one on a schedule against full history, catches more than either alone would. It’s worth taking seriously specifically because of how coding assistants have changed the failure mode: a GitGuardian report earlier this year found AI-assisted commits leaking secrets at roughly double the rate of human-typed ones, which tracks. Pasting a working credential into a prompt or a generated config file is a much easier way to leak one than typing it from memory ever was.

None of these controls substitute for each other. Scoping a secret to a VPC endpoint does nothing if the IAM role allowed to read it is wide open. Rotating on a schedule doesn’t help if the previous version is still sitting in three CI logs somewhere. The actual work is closing every path on the map from the start of this post, not just the one that happens to look most broken today.

If there’s one rule that generalizes

Count the number of places a secret’s plaintext value could theoretically be read by something other than the workload that needs it: a state file, a build log, a Kubernetes Secret sitting in etcd, a debug terraform output, a .env pasted into a chat message. Every tool in this post is, underneath its specific syntax, either reducing that count or increasing it. Write-only arguments reduce it. A Sealed Secret sitting in a public repo doesn’t touch it at all, which is the point. A sensitive = true flag that a team believes is doing more than it actually is increases it, quietly, until someone finds out the hard way. Pick tools by asking which direction they move that number, and most of the decisions in this post get easier to make on their own.

VPC Lattice: What It Actually Replaces, What It Costs, and When I’d Reach for It

AWS App Mesh shuts down on September 30, 2026. Not “enters maintenance mode” — the console and the resources stop working. If you’re one of the teams that bet on it in 2019 for EKS service-to-service traffic, you’ve got a couple of months left, and AWS has been pointing everyone at the same replacement for ECS workloads and EKS workloads alike: Amazon VPC Lattice.

That deadline is a decent excuse to actually understand what Lattice is, because it’s not just an App Mesh clone. It quietly replaces a chunk of what Transit Gateway and PrivateLink do too, and it changes some assumptions network engineers have had baked in since VPC peering existed. Worth twenty minutes even if you have no App Mesh to migrate.

What VPC Lattice actually is

Strip away the marketing and VPC Lattice is a managed, regional application-layer proxy that sits between your consumers and your services, handles service discovery, applies IAM-based authorization, and routes requests — all without you deploying a load balancer, a sidecar, or a peering connection. AWS’s own description is that it manages network connectivity and application layer routing between services across different VPCs and AWS accounts, and that’s a fair summary, if a little dry.

The mental model that clicked for me: think of it as a load balancer that doesn’t live in any one VPC. You publish a service into a service network (a logical boundary, not a piece of infrastructure you provision), and any VPC or account associated with that service network can reach it by DNS name — no route tables, no CIDR planning, no peering mesh growing into an unmanageable tangle.

Lattice Service Network.

A service network can span accounts via AWS Resource Access Manager, so a platform team can own the network and share it out to application teams without those teams ever touching a route table. That’s the part that made this interesting to me as an architect rather than as a networking hobbyist — it moves network topology out of the app team’s problem space entirely.

What it’s quietly replacing

Nobody built this in a vacuum. Every one of these had a gap Lattice was built to close.

VPC peering doesn’t scale past a certain account count — it’s point-to-point, so N accounts means something close to N² connections, and overlapping CIDRs break it outright. Lattice’s data plane assigns consumers a link-local address and doesn’t care what your VPC CIDR looks like, overlapping or not.

Transit Gateway solved the peering-mesh problem at layer 3/4 — it’s still the right tool for routing IP traffic between VPCs and on-premises networks at scale, and I wouldn’t rip one out to replace it with Lattice. But TGW has no concept of an HTTP request, a path, or an identity. It moves packets, not requests. If your actual problem is “service A needs to call /v2/orders on service B and I want to enforce that with IAM,” TGW can’t help you get there on its own.

AWS PrivateLink is the closest sibling — it’s also a proxy-based, non-peering way to expose a service across accounts. The difference is PrivateLink is fundamentally point-to-point: each consumer VPC needs its own interface endpoint, and there’s no native request-level routing or IAM authorization baked into the data path. Lattice centralizes that into one service network that many consumers attach to, and layers on HTTP-aware routing rules on top.

AWS App Mesh is the one with the deadline. It gave you real service-mesh features — client-side retries, circuit breaking, fine control — but it did that with an Envoy sidecar next to every task, which is real operational weight: certificate rotation, sidecar upgrades, resource overhead per pod. Lattice deliberately drops the sidecar. You lose some of App Mesh’s client-side sophistication in the trade — Lattice’s routing and auth decisions happen server-side, not in a proxy next to your code — but for most teams that’s a fair trade for not operating a mesh control plane.

VPC PeeringTransit GatewayPrivateLinkVPC Lattice
OSI layerL3L3/L4L3/L4L7 (app layer)
Cross-account model1:1 meshHub-and-spoke1:1 endpoint per consumerMany-to-many via service network
CIDR overlap tolerantNoNoYesYes
Identity-aware authNoNoNo (SG/NACL only)Yes (IAM/SigV4)
Request-level routingNoNoNoYes (path, method, header, weighted)
Sidecars requiredNoNoNoNo

How it’s actually built

Four concepts do the work. A service is the logical front door — listeners, rules, target groups — conceptually a load balancer, but one that isn’t tied to a single VPC. A listener is a protocol and port (HTTP, HTTPS, or TLS_PASSTHROUGH). A rule matches on path, HTTP method, or header and forwards to a target group, which can hold EC2 instances, IP addresses, Lambda functions, or an Application Load Balancer, and — as of the more recent releases — ECS and Fargate tasks natively alongside EC2 and EKS.

Worth calling out explicitly: Lattice terminates HTTPS itself using an ACM-managed certificate, which is convenient, but it only does server-side TLS. If you need mutual TLS, you have to use a TLS_PASSTHROUGH listener and let the target negotiate the client certificate — Lattice’s own data plane won’t do mTLS termination for you.

The newer half of the story is VPC Resources and Resource Gateways, which let you expose non-HTTP TCP endpoints — a database, an internal domain name, an on-premises system reachable over Direct Connect or VPN — through the same service network model, complete with the same IAM controls. This is the part that pushes Lattice past “App Mesh replacement” and into “how do I expose an RDS instance to another account without a peering connection or a bastion.” It’s genuinely useful, and it’s the feature I’d bring up first if a platform team asked me why this deserves attention beyond the App Mesh deadline.

For Kubernetes specifically, the AWS Gateway API Controller implements the open-source Kubernetes Gateway API and translates Gateway and HTTPRoute objects into Lattice service network objects behind the scenes. If your EKS team already writes Gateway API manifests, adopting Lattice barely changes their workflow — it’s the controller that’s doing the AWS-specific work.

Authorization runs on IAM. Auth policies attached at the service network or service level use standard IAM policy JSON, and clients authenticate with SigV4-signed requests, the same signing scheme every other AWS API call uses. This is a genuine strength if your org already centers identity around IAM roles, and a genuine adoption cost if you have legacy clients that have never had to sign a request in their life. Budget time for that conversation — it comes up in almost every migration writeup I’ve read.

When I’d actually reach for it

When Lattice makes sense.

Lattice earns its place when the actual requirement is cross-account or cross-VPC service exposure with identity-aware access control and some request-level routing, and you don’t want to own a control plane to get it. New microservice builds, platform teams centralizing how application teams expose services to each other, EKS shops trying to get off App Mesh before the clock runs out, and anyone who’s been maintaining a PrivateLink endpoint-per-consumer sprawl are the clearest fits.

I’d think twice before using it as a wholesale Transit Gateway replacement. It’s not built for raw IP-layer routing at TGW’s scale, and — as I’ll get to in pricing — the per-service hourly charge adds up fast if you’re running hundreds of services with modest traffic each. I’d also pause if you need genuine client-side mesh behavior: retries with backoff, circuit breakers, fault injection for chaos testing. That logic sits server-side in Lattice, which is simpler to operate but less flexible than a sidecar that runs next to your code.

Pros and cons, plainly

The upside is real: no sidecars to patch, no peering mesh to keep untangled, IAM auth you already understand, weighted routing for blue/green and canary out of the box, and a service directory that gives you an actual inventory of what’s exposed to what. Observability is baked in too — access logs and CloudWatch metrics per service without instrumenting anything yourself.

The downside is mostly about maturity and rigidity. It’s newer than TGW or PrivateLink, so you’ll hit rough edges — G2 reviewers flag protocol support currently limited to HTTP, HTTPS, and gRPC, and region coverage, while it’s grown a lot since GA, still isn’t everywhere. Security group design has a real gotcha: targets need to allow inbound traffic from the Lattice association security group, not from the consumer’s VPC CIDR, and that trips people up in early deployments because it looks wrong at first glance. And moving to IAM/SigV4 auth is a genuine client-side change — nothing you can skip past.

Pricing — the part that actually decides adoption

Three dimensions drive the bill: an hourly charge per service, a per-GB data processing charge, and a per-request (or per-connection, for TLS_PASSTHROUGH) charge above a free tier. In US East (N. Virginia), that’s $0.025 per service-hour, $0.025 per GB processed, and $0.10 per million requests beyond the first 300,000 free per hour. Compare that to PrivateLink at roughly $0.01 per hour per AZ plus $0.01/GB with volume discounts, and it’s clear Lattice costs more per unit — you’re paying for the extra L7 intelligence and the simpler operating model, not for cheaper bytes.

Run the numbers on a single, modestly busy service and it’s not scary: one service processing 100 GB and 200,000 requests an hour for a month lands around $18–20 in hourly charges alone before data and requests, and a heavier example AWS publishes — one service with HTTPS and TLS listeners together processing 2,100 GB and millions of requests a month — comes out to roughly $268 a month. The bill gets serious at fleet scale. One cloud architect’s published comparison modeled 200 services across 100 accounts pushing 30TB a month: Transit Gateway came out around $4,250, Lattice around $4,840 — roughly a 14% premium, driven almost entirely by the per-service hourly charge rather than data processing, which was actually cheaper for TGW in that scenario. If you’re running that many low-traffic services, the fixed hourly cost per service matters more than the per-GB rate.

VPC Resources (the database/on-prem exposure feature) bill separately and use a tiered per-GB rate: $0.01/GB for the first petabyte in a Region each month, dropping to $0.006 and then $0.004/GB at higher volumes, plus a small hourly charge per resource. Worth modeling separately from your service traffic if you’re planning to route a lot of data through it.

Quotas worth knowing before you design around them

These are the ones that actually shape an architecture, not just trivia:

QuotaDefaultAdjustable
Services per Region2,000Yes
Service networks per Region50Yes
Service associations per service network500Yes
VPC associations per service network500Yes
Target groups per service10Yes
Targets per target group1,000Yes
Listeners per service2Yes
Rules per listener10Yes
Requests/sec per service per AZ10,000Contact your SA/TAM
Bandwidth per service per AZ10 GbpsContact your SA/TAM
Connection idle timeout (HTTP/gRPC)1 minuteContact your SA/TAM
Max connection lifetime10 minutesFixed
Service networks a VPC can associate with directly1— (use service-network VPC endpoints for more)

Full, current numbers are on the official quotas page — check it before you design, not after you hit a wall. The one that catches people off guard most is the one-service-network-per-VPC-via-direct-association limit; if you need a VPC in more than one service network, you’re routing through service-network VPC endpoints instead, which is an extra layer to plan for.

A few concrete use cases

A platform team centralizing how forty application teams expose internal APIs to each other, replacing forty sets of ad-hoc security-group rules and a slowly decaying peering mesh with one service network and a consistent IAM auth policy. A company running canary deployments where 5% of production traffic gets weighted onto a new service version before a full cutover — Lattice’s weighted target groups do this natively, no custom load balancer logic required. An EKS shop migrating off App Mesh before the September deadline, using the Gateway API Controller so app teams keep writing the same HTTPRoute manifests they already know. And the resource-gateway pattern: exposing a shared RDS instance in a data-platform account to a dozen consuming accounts without provisioning a PrivateLink endpoint in every one of them.

Why this belongs on your radar even without the App Mesh deadline

The App Mesh shutdown is the forcing function, but the more interesting shift is architectural: AWS is pulling network topology out of individual VPCs and into a service-oriented abstraction that’s owned centrally and consumed by reference. That’s a meaningful change to how you’d write a network architecture standard or review an ADR — the question stops being “how do these two VPCs route to each other” and starts being “which service network does this belong to, and what’s the IAM policy governing who can call it.” If you own architecture review or network standards for your org, that’s worth getting ahead of before it shows up in someone else’s design doc and you’re reviewing it cold.

It’s not a wholesale replacement for Transit Gateway, and it’s not free — budget the per-service hourly charge honestly before you commit a few hundred services to it. But for the specific problem it targets, cross-account and cross-VPC service exposure with identity-aware access control, it’s a cleaner answer than anything that came before it, and worth prototyping now rather than in month eleven of an App Mesh migration deadline.

AWS Landing Zone, Part 3: Transit Gateway, Centralized Egress, and Where IPAM Earns Its Keep

Five VPCs connected by peering need ten peering connections. Ten VPCs need forty-five. Every new VPC means going back into the route tables of every existing VPC to add the new relationship, and every peering connection is a distinct thing that can be misconfigured, forgotten, or left open longer than it should be. That math is the whole argument for hub-and-spoke, and it’s why almost nobody designs a multi-account AWS network as a full peering mesh past the first few accounts.

AWS Landing Zone – Part 3

Part 1 and Part 2 covered the account structure and the guardrails that govern it. This post is about the network underneath: Transit Gateway as the hub almost everyone reaches for, centralized egress and what it costs against the alternative, IPAM (IP Address Management) as the CIDR governance layer, and where VPC Lattice fits without pretending it replaces Transit Gateway wholesale.

Transit Gateway as the hub

AWS Transit Gateway is a managed, regional routing hub. Spokes, VPCs, VPN connections, Direct Connect, attach to it once, and it handles routing between them using its own route tables instead of point-to-point relationships. One Transit Gateway per Region is typically enough since it’s highly available by design, though there’s a legitimate case for more than one where you want to limit the blast radius of a routing misconfiguration or separate control-plane operations between teams. Cross-region connectivity happens through Transit Gateway peering; hybrid connectivity, Direct Connect or VPN, attaches the same way a VPC would.

The account this lives in matters. Transit Gateway, IPAM, and centralized egress infrastructure all belong in the Infrastructure OU’s networking account, sometimes literally called Network or Network Hub, not scattered across workload accounts and not in the Control Tower management account. That account becomes the place a network engineer actually works day to day, and it’s worth treating its own access and change process with the same care as the Security OU from Part 2.

Laid out, the topology looks like this:

Transit Gateway topology!

Every spoke VPC gets one attachment to the hub instead of a direct relationship with every other spoke. Whether Prod and Test can actually reach each other is a routing table decision inside the Transit Gateway, not something the topology forces one way or the other, which is exactly the control you lose with a flat peering mesh.

Centralized egress and what it actually costs

The other big reason to centralize on a hub isn’t just routing hygiene, it’s cost. Every VPC that needs internet access wants its own NAT gateway per Availability Zone for resilience, and NAT Gateway data processing charges add up fast once you’re running dozens of VPCs, each with its own pair or trio of gateways running independently. Centralizing egress means routing outbound traffic from every spoke through the Transit Gateway into a single egress VPC in the network account, where it exits through one set of NAT gateways instead of dozens. One practitioner audit found organizations spending around $15,000 a month on NAT Gateway data processing alone, spread thin across VPC after VPC, and cut that by 40 to 70 percent by consolidating.

Adding traffic inspection to that same choke point is a natural next step. AWS Network Firewall sits in the egress VPC and inspects traffic before it reaches the NAT gateway, at a real but bounded cost, roughly $0.40 an hour per endpoint plus a per-gigabyte processing charge, which is a fair trade in a regulated environment and a harder sell if you’re mostly optimizing for spend. One genuine gotcha here: DNS doesn’t follow this path by default. Route 53 Resolver and DNS Firewall are a separate egress route entirely, so centralizing your data-plane egress through Transit Gateway and Network Firewall doesn’t automatically mean your DNS queries are inspected the same way. If DNS-based filtering matters to your threat model, it needs its own explicit design, not an assumption that it rides along with everything else.

IPAM: get the CIDR plan right before account one

IP address planning is one of those things that’s cheap to fix before it exists and expensive to fix after. Amazon VPC IPAM gives you a hierarchical pool structure, a top-level pool subdivided into regional pools, then further into business-unit or environment pools, so that a new VPC in the Workloads_Test OU pulls its CIDR from a pool already guaranteed not to overlap with Prod, Sandbox, or anything else in the organization. IPAM should be delegated to the network account rather than run from the Control Tower management account, the same separation-of-duties instinct as everything else in this series.

The failure mode IPAM prevents is duller than a security incident but just as disruptive: two teams independently pick 10.0.0.0/16 for their VPCs, everything works fine in isolation, and then someone needs to connect the two networks and discovers the ranges collide. Fixing that after the fact means readdressing a live VPC, which is exactly the kind of maintenance window nobody wants to schedule. Getting the CIDR plan right before the first account exists costs an afternoon. Getting it wrong costs a migration.

Route 53 Profiles solve the equivalent problem for DNS. Instead of manually associating private hosted zones, resolver rules, and DNS firewall rule groups to every VPC individually, you bundle them into a single profile in the network account and share it across accounts through AWS RAM. New VPCs associate with the profile once and inherit the whole DNS configuration, rather than someone remembering to wire up each piece by hand every time an account gets vended.

Where VPC Lattice actually fits

VPC Lattice gets pitched sometimes as a Transit Gateway replacement, and that framing oversells it. Transit Gateway operates at Layer 3, it moves packets between IP addresses and doesn’t know or care which identity sent them; security has to come from route tables, security groups, and NACLs. VPC Lattice operates at Layer 7, HTTP, HTTPS, and gRPC specifically, and it’s service-centric rather than network-centric: services register into a service network, consumers discover them by DNS name, and access is governed by IAM policy rather than which subnet you happen to be in. It’s also currently single-region, so it isn’t a drop-in for anything that needs to span regions the way Transit Gateway peering does.

In practice these coexist rather than compete. Transit Gateway keeps doing the job of moving bulk traffic, hybrid connectivity, and anything that isn’t a clean HTTP service call. VPC Lattice picks up new service-to-service communication where IAM-based authorization is a better fit than managing security group rules across account boundaries. I’d reach for Lattice for a new internal API a team wants to expose across accounts without punching new holes in the network layer, not as a project to migrate an existing Transit Gateway backbone onto.

As a decision, it collapses to one real question about the shape of the traffic:

Shaping traffic!

Bulk data, non-HTTP protocols, or anything crossing regions stays on Transit Gateway. A specific, well-defined service boundary within one region is where Lattice earns its keep, and it’s rarely an either-or choice at the level of the whole network.

What’s next

The network is the part of a landing zone that’s genuinely painful to change once workloads depend on it, which is exactly why it deserves the same up-front discipline as the OU structure in Part 1. The last post in this series moves from design to operations: how account vending actually scales once you’re provisioning dozens of accounts a month, why Control Tower can tell you about drift automatically but won’t fix it for you, and the CI/CD pipeline that should sit in front of every change to the guardrails from Part 2 before it reaches a real account.

AWS Landing Zone, Part 2: SCPs, RCPs, and the Logging That Holds Up Under Audit

There’s a specific error message in AWS Control Tower that ruins a Tuesday. Someone edits or detaches a managed SCP (Service Control Policy) on the Security OU (Organization Unit), and the console locks you out with a warning that the shared accounts may no longer be working, and that you shouldn’t provision new accounts until it’s fixed. AWS documents this exact failure mode in its own knowledge base, which tells you it happens often enough to need a canonical writeup. The lesson isn’t “don’t touch SCPs.” It’s that the guardrails Control Tower manages for you and the guardrails you write yourself live in the same policy type, and mixing them up on the wrong OU is how you find out which one was load-bearing.

AWS Landing Zone – Part 2

This is Part 2 of the landing zone series. Part 1 covered what a landing zone is and the build-vs-buy decision behind Control Tower, LZA, and AFT. This one is about the layer that actually does the enforcing: service control policies and their newer sibling, resource control policies, the logging architecture that has to hold up when an auditor asks for it, and IAM Identity Center as the one door every human uses to get into any of this.

SCPs and RCPs aren’t the same guardrail

It’s worth being precise here because the names invite confusion. A service control policy caps what your own identities, users and roles inside your organization, are allowed to do, regardless of what their IAM policy says. A resource control policy caps who can touch a given resource, an S3 bucket, a KMS key, an SQS queue, and a growing list of others, regardless of what that resource’s own policy allows. AWS puts it simply: use an SCP to limit your own principals, use an RCP to restrict access to your resources from principals outside your organization. Neither one grants anything. Both are ceilings, not floors, and the permission a principal ends up with is whatever’s left after intersecting the SCP, the RCP, and the identity or resource policy underneath.

RCPs are the newer of the two, and they’re moving fast. They launched in November 2024 covering five services: S3, STS, KMS, SQS, and Secrets Manager. Since then AWS has kept extending the list. Cognito and CloudWatch Logs picked up support in January 2026, DynamoDB followed a few weeks later, and the per-organization quota doubled to 2,000 RCPs this past July. That pace tells you something: RCPs are still filling in gaps, not yet the mature, complete tool that SCPs are. Check the current supported-service list before designing a control around a service it doesn’t cover yet.

SCPs got their own significant upgrade in September 2025, when AWS gave them full IAM policy language support: conditions inside Allow statements, individual resource ARNs, NotAction with Allow, wildcards in the middle of an Action string. Before that update, SCPs were noticeably blunter than a regular IAM policy, which pushed a lot of teams toward broad deny-everything-except statements because anything more surgical wasn’t expressible. Existing SCPs kept working unchanged after the update, but if yours predate September 2025, there’s a real case for revisiting them now that more precise allow-with-conditions patterns are possible.

Where these guardrails get scoped

Tie this back to the OU structure from Part 1. In practice, almost every SCP and RCP you write gets attached at the Security OU, the Infrastructure OU, the Workloads OU, or somewhere in between, and the Security OU is the one to treat with real caution. That’s where Control Tower’s own managed SCPs live, the ones protecting the Log Archive and Audit accounts, and it’s exactly the OU where the lockout scenario above happens. If you need custom guardrails for security tooling, write them, but write them as additions at a level below where Control Tower’s own policies sit, not as edits to the managed ones.

Put visually, the two guardrail types gate different paths to the same resource:

SCPs,RCPs

The distinction matters operationally, not just semantically. An SCP written to block a risky action only stops your own users and roles from doing it. If the same S3 bucket is reachable by a principal from another AWS account entirely, only an RCP, or the bucket policy underneath it, actually stops that. Teams that treat SCPs as a complete security boundary and skip RCPs are leaving exactly this gap open, usually without realizing it until an access review turns it up.

Centralized logging that holds up under audit

Control Tower sets up an organization-level CloudTrail trail automatically, which, since landing zone version 3.0, replaced the older model of a separate trail per account. One trail logs everything, management account and every member account, and delivers into an S3 bucket that lives in the Log Archive account, the one nobody logs into day to day. AWS Config runs alongside it, aggregating configuration history into the Audit account rather than Log Archive, a distinction worth remembering when someone asks where a specific piece of evidence actually lives.

The pattern that makes this scale past a handful of accounts is delegated administration. Instead of every security service being manageable only from the management account, you designate a member account, almost always the Audit account, as the delegated admin for GuardDuty, Security Hub, Config, and similar services, and manage all of them centrally from there without ever touching the management account for day-to-day work. This has expanded well beyond security services specifically; one recent count put the number of AWS services supporting delegated administration at 37, up from roughly a dozen when the pattern first appeared. GuardDuty is regional, so this has to be repeated per Region you actually monitor, which is easy to miss if you’ve only ever tested in one.

Here’s what that centralization looks like end to end for a single finding:

Sequence of tracking

None of this is exotic engineering. It’s mostly turning on the right delegation and pointing things at the Audit account instead of leaving them scattered. The payoff shows up specifically during an audit, when “show me every API call across every account for the last year” is a single query against one bucket instead of a scavenger hunt across forty.

IAM Identity Center as the only door in

Every human touching any of these accounts should be going through IAM Identity Center, not IAM users with long-lived access keys. The mechanics are straightforward: permission sets define what a person can do, assignments connect a permission set to a group and an account or set of accounts, and Identity Center issues short-lived credentials rather than anything standing. Assign to groups, not individuals. When someone changes teams you move the group membership and every downstream permission follows automatically.

Identity Center federates cleanly with an external identity provider over SAML, Okta, Entra ID, whatever your organization already runs, with SCIM (System for Cross-domain Identity Management) handling user and group provisioning so you’re not managing a second identity store by hand. The one thing every landing zone still needs underneath all of this is break-glass access: a small number of IAM roles or users, outside Identity Center entirely, that work even if federation itself is broken or misconfigured. It’s not a contradiction to have SSO for everything and also keep a locked-down emergency path around it. It’s the same reason a building keeps a physical key next to the electronic badge reader.

One detail worth knowing if you’re serious about treating this whole layer as code: permission sets and their assignments can be managed through a CI/CD pipeline just like the SCPs and RCPs above, JSON templates in a repository, a pipeline reacting to changes and pushing updates to Identity Center. It’s a small thing, but it means access changes go through the same review process as everything else instead of being a console click nobody remembers making six months later.

What’s next

Guardrails and identity get most of the attention in landing zone design because they’re where the compliance conversations happen, but the networking layer underneath all of this has its own decisions and its own ways to quietly overspend. Part 3 covers Transit Gateway and the hub-and-spoke pattern almost everyone converges on, centralized egress and what it actually costs against a pile of per-VPC NAT gateways, IPAM as the CIDR governance layer you want in place before account number one, and where VPC Lattice fits next to Transit Gateway instead of replacing it.

Amazon CloudWatch Managed Prometheus Collectors: Retiring the Self-Managed OTel Collector

Ask anyone who has wired Prometheus metrics into CloudWatch what the real bottleneck was, and it was never Prometheus. It was the collector. You’d stand up an OpenTelemetry Collector as a Deployment or a DaemonSet, hand-tune its scrape config, guess at memory limits for cardinality you hadn’t measured yet, and then spend the next year treating it like one more piece of cluster infrastructure that needed patching, right-sizing, and on-call attention of its own.

AWS Managed Prometheus Collectors

On July 31, AWS announced managed Prometheus collectors for Amazon CloudWatch, a fully managed, agentless scraper for Amazon EKS, Amazon EC2, Amazon ECS, Amazon MSK, and Amazon OpenSearch Service. You supply a scrape configuration and a way to reach your resources, and CloudWatch takes it from there: provisioning the collector, scaling it, and keeping it running. What it scrapes arrives in OpenTelemetry format, queryable with PromQL right next to your AWS vended metrics.

That’s the announcement. The more useful question, if you’ve spent any time in AWS’s observability stack already, is what’s actually new here versus what’s just wearing a new badge.

The new part is the destination, not the collector

This isn’t new collector technology. AWS shipped this same agentless scraper (automatic target discovery, no in-cluster agent, multi-AZ, metrics that never leave your VPC) as the Amazon Managed Service for Prometheus collector back in November 2023, initially scoped to EKS and writing into an AMP workspace. That lineage is visible directly in the API today: you still create one of these with the CreateScraper operation, under the aws amp CLI namespace, not some new CloudWatch-specific command. What’s new in this announcement is the destination and the source list. Alongside an AMP workspace, the same managed collector can now write into a CloudWatch dataset, and the supported sources have grown from EKS-only to include EC2, ECS, MSK, and OpenSearch.

Here’s an EKS scraper creation, adapted from AWS’s setup docs:

aws amp create-scraper \
  --alias "eks-metrics-scraper" \
  --source eksConfiguration="{clusterArn='arn:aws:eks:us-west-2:111122223333:cluster/prod-cluster', \
    securityGroupIds=['sg-0123456789abcdef0'], \
    subnetIds=['subnet-0aaa111','subnet-0bbb222']}" \
  --scrape-configuration configurationBlob=$(cat eks-scrape-config.yaml | base64 -w 0) \
  --destination cloudWatchConfiguration="{datasetArn='arn:aws:cloudwatch:us-west-2:111122223333:dataset/default'}"

Swap that final --destination flag for an ampConfiguration with a workspace ARN and you’re back to the exact collector AWS shipped in 2023. If you’re already running AMP collectors for Grafana dashboards, that’s the useful takeaway: this is an incremental destination option on infrastructure you may already trust, not a new product to evaluate from a cold start.

It also builds on something CloudWatch shipped only six weeks earlier: native OTLP ingestion with PromQL querying, launched June 9. That release is what made CloudWatch a legitimate destination for Prometheus-shaped metrics at all, with per-GB pricing and curated Container Insights dashboards for EKS. Managed Prometheus collectors are the agentless front door to that same pipeline, extended past what Container Insights auto-instruments.

How a scrape actually reaches CloudWatch

The mechanics will feel familiar if you’ve used any managed AWS scraper before. You hand the collector a set of subnets, and it creates an Elastic Network Interface in each one. It scrapes your targets over those ENIs using OTLP, then delivers the result to your CloudWatch dataset through a VPC endpoint. Nothing crosses the public internet.

What isn’t obvious from the diagram is that target discovery differs meaningfully by source, and that’s what determines how much of an existing scrape config actually survives the move:

  • EKS uses Kubernetes service discovery against the cluster API, plus control-plane metrics (kube-scheduler and kube-controller-manager, exposed starting at Kubernetes 1.28).
  • ECS uses DNS-based discovery through AWS Cloud Map, so tasks register once and get picked up without a static IP list.
  • EC2 is static_configs against private IPs, no different from pointing a self-managed setup at a fixed target list.
  • MSK uses DNS-based discovery against the cluster’s own bootstrap DNS name, which resolves to every broker and survives broker replacement without reconfiguration.

MSK is worth a closer look because it needs a prerequisite step: enabling Open Monitoring on the cluster, which exposes a JMX Exporter on port 11001 and a Node Exporter on port 11002. A scrape config covering both looks like this:

global:
  scrape_interval: 60s
  external_labels:
    cluster_name: my-msk-cluster

scrape_configs:
  - job_name: 'msk-jmx'
    dns_sd_configs:
      - names: ['my-cluster.abc123.c4.kafka.us-west-2.amazonaws.com']
        type: A
        port: 11001
    relabel_configs:
      - source_labels: [__address__]
        target_label: instance

  - job_name: 'msk-node'
    dns_sd_configs:
      - names: ['my-cluster.abc123.c4.kafka.us-west-2.amazonaws.com']
        type: A
        port: 11002
    relabel_configs:
      - source_labels: [__address__]
        target_label: instance

That gets you broker-level Kafka metrics (topic throughput, consumer lag, under-replicated partitions) from the JMX exporter, and host-level CPU, memory, and disk metrics from Node Exporter, both queryable with PromQL once they land. Two things to check before you plan around this: MSK Serverless and MSK Express aren’t supported, and public access combined with KRaft metadata mode is excluded too.

The scrape config is Prometheus-compatible, not Prometheus

This is worth reading closely before you migrate anything, because “Prometheus-compatible YAML” undersells how much is actually missing. The supported configuration surface covers global settings, scrape_configs with static_configs and dns_sd_configs, and relabeling. What isn’t there matters more than what is:

  • Minimum scrape interval is 30 seconds. An SLO built on 10 or 15-second scrapes doesn’t fit here.
  • No file_sd_configs, and no Consul, Eureka, or other external service discovery. If your target list comes from anywhere other than Kubernetes, Cloud Map, a static list, or MSK’s broker DNS, this collector can’t reach it.
  • remote_write and remote_read don’t apply, because delivering to CloudWatch is the whole job. A fan-out to a second backend doesn’t carry over.
  • Configuration blobs cap out at 256 KB, base64-encoded.

None of this is really a flaw. It’s the standard trade-off of a managed service. AWS narrowed the surface to what it can operate reliably at scale, and the narrowing happens to remove exactly the parts of Prometheus configuration that are hardest to run safely inside someone else’s control plane. But “copy your existing prometheus.yml over” is rarely literally true. Diff your current scrape configs against this list before you commit to a cutover date, not after.

What ships automatically, and what doesn’t

EKS and MSK both get a curated dashboard the moment metrics start flowing — EKS OTel and MSK OTel in the CloudWatch console — plus attribute enrichment on every metric: account, region, an inferred unit, and for EKS specifically the cluster name and ARN. That enrichment is what makes PromQL filtering pleasant instead of an exercise in label archaeology. The EC2/ECS path doesn’t come with an automatic dashboard, which is a reasonable trade given how varied EC2/ECS metric shapes are in practice, but it’s worth knowing before cutover rather than after.

Operationally, a managed collector vends its own logs to CloudWatch Logs covering target discovery, scrape successes and failures, and configuration errors, and you manage the scraper itself with aws amp list-scrapers, describe-scraper, and delete-scraper. Deleting one tears down its ENIs, and that cleanup takes a few minutes, so don’t script an immediate re-create into the same subnet.

Cross-account guidance has shifted too. Instead of the role-chaining pattern AMP has historically used for cross-account scrapers, AWS’s current recommendation for both EKS and MSK is CloudWatch metric centralization: replicate the metrics to a monitoring account rather than granting the scraper cross-account reach in the first place.

What this actually costs

Pricing has two components that scale independently, and it’s worth keeping them separate. The collector itself is billed hourly, similar in shape to an ENI or a NAT gateway, though the specific per-hour rate isn’t broken out yet on AWS’s public pricing page as of this writing. What is confirmed is that everything the collector delivers rides on standard CloudWatch OpenTelemetry ingestion pricing: currently $0.50 per GB for general OTel metric publishing, a flat rate that bundles in fifteen months of storage with no separate per-metric or per-API charge. A typical data point with ten to fifteen attributes runs 300 to 600 bytes, for a sense of scale. Console PromQL queries (Query Studio, dashboards) are free; programmatic PromQL queries cost $0.01 per million samples scanned.

That per-GB model is a real shift in how you’d think about cost here. Classic CloudWatch metrics billed per unique metric-dimension combination, so cardinality was the thing to fear. Under OTel ingestion, cardinality by itself is free; what costs money is the serialized size of each data point, which grows with the number and length of the labels attached to it. Fifteen short labels on a metric barely move the needle regardless of series count. A handful of long, high-cardinality values move it fast: unbounded request IDs, verbose ARNs repeated on every point. If you’re porting scrape configs over from a self-managed setup, that’s the moment to prune labels you’ve been carrying out of habit rather than actual query need.

One more line item: scraping itself can rack up VPC data transfer charges between the collector and your targets, separate from CloudWatch ingestion. AWS’s own suggestion is to enable gzip on your /metrics endpoints, which cuts transfer cost without changing what actually gets ingested.

Where I’d actually reach for this

If your team is already CloudWatch-centric and has been putting off a real Prometheus setup mainly because nobody wanted to own collector infrastructure, this closes that gap cleanly. For EKS and MSK metrics specifically, it’s a solid default: alarms and dashboards in the same place as everything else, without standing up Grafana and an AMP workspace just to get there.

I’d still keep a self-managed collector, or the AMP-workspace version of this same scraper, in a handful of specific situations: sub-30-second scrape requirements, target discovery through Consul, Eureka, or file-based service discovery, or a remote_write fan-out to more than one backend because you’re mid-migration or running a deliberately multi-vendor setup on purpose. None of those are rare for a mature platform team, so this doesn’t replace the self-managed path outright. It’s a genuinely good default for the common case, which happens to be most of it.

Availability is broad but not universal — it’s live everywhere CloudWatch’s OTLP endpoint already exists, minus Asia Pacific (New Zealand) for now. Worth checking before you write the migration runbook, not after you’ve scheduled it.

Self-Managed S3 Buckets for Lambda Code: You Finally Own the Artifact


Picture a platform team running a few hundred Lambda functions across a handful of accounts. Every function has several published versions kept around for safe rollbacks. One morning a deploy fails with a quota error, and the culprit isn’t the new function at all. It’s the 75 GB account-level limit on function and layer code storage, quietly filled up by copies of packages you thought lived in your own S3 bucket. If you’ve ever had to explain to a security reviewer why your deployment artifacts also sit in an internal bucket you can’t see, encrypt, or tag, this one is for you.

S3 buckets for Lambda code!

AWS has added a way to point Lambda at your own S3 bucket and have it read your code from there directly, with no hidden second copy.

What actually changed

There’s a new function configuration setting called S3ObjectStorageMode. It has two values.

  • The default is COPY, which behaves exactly like Lambda always has. If you don’t set the field at all, you get COPY, so nothing about your existing functions changes on its own.
  • The new value is REFERENCE. Set it when you create or update a function, and Lambda stops copying your .zip package into its internal storage. Instead it keeps a reference to your S3 object and reads the code from your bucket when it needs it. Your object becomes the one canonical artifact.

The feature works in all AWS standard regions where Lambda is available, and there’s no extra charge for it beyond the normal S3 storage and request costs you’d pay anyway.

How it works, old way and new

Under the old model (now COPY), the flow looks like this: you upload a .zip to your bucket, call CreateFunction or UpdateFunctionCode with the bucket name and object key, and Lambda pulls that artifact into a service-managed bucket. It builds the optimized, runnable version of your function from that internal copy. The catch is that this copy counts against your 75 GB account quota, and you have no say over how it’s stored.

With REFERENCE, that copy step disappears. Lambda records where your object lives and reads it directly. Two things fall out of that. First, your package no longer counts toward the 75 GB limit, because there’s no Lambda-side copy to count. Second, creating and updating functions gets faster, since Lambda skips the copy-into-internal-bucket step. AWS describes this as a faster time to first invoke for new functions and after updates.

Architecture

The diagram below contrasts the two modes and sketches the multi-account pattern that, in my experience, is the real reason to adopt this.

Two modes for S3 buckets of Lambda code.

In a COPY deployment, the artifact exists twice: once in your bucket, once inside Lambda. In REFERENCE mode there’s a single object, and your function holds a pointer to it. Extend that to an organization and the shape gets interesting: put every artifact in one bucket in a central “artifact” or shared-services account, then grant s3:GetObject to each workload account’s Lambda execution role through the bucket policy. Now one bucket is the source of truth for what’s deployed everywhere, with one place to enforce encryption, versioning, and retention.

When to reach for it

  • CI/CD and artifact management. Your pipeline uploads a package once, and the function references that same object. One set of lifecycle rules and access controls covers everything, and a rollback becomes “point the function at the previous S3 object version.”
  • Multi-account, multi-team setups. Centralize artifacts in one account, hand out cross-account s3:GetObject via bucket policies, and keep a single inventory of what code runs where.
  • Disaster recovery. Because you own the bucket, you can turn on Cross-Region Replication (CRR) or Same-Region Replication (SRR), and pair it with S3 Versioning and Object Lock to keep a tamper-resistant archive that survives an accidental delete or a corrupted deploy. See the S3 replication docs and S3 lifecycle examples.
  • Quota and compliance pressure. If you’re brushing up against 75 GB, or you need your own encryption, access logging, Object Lock, or compliance tags on the artifact, this gives you that control.

When to leave it alone

For plenty of workloads, COPY is fine and simpler. If you aren’t near the quota, don’t need custom encryption or tagging on the artifact, and have no DR requirement on your deployment packages, there’s little reason to change anything. The default exists for a reason.

Security notes

  • The trade in REFERENCE mode is control for responsibility. You now own the bucket’s posture: its encryption, access policies, lifecycle transitions, and audit trail are yours to configure and to get right.
  • For cross-account access, you grant s3:GetObject to the calling function’s execution role in the bucket policy.
  • Versioning plus Object Lock is the combination worth setting up early if you care about a durable, tamper-proof code archive.
  • The full IAM and bucket-policy details live in the Lambda developer guide.

Pricing

No additional charge for the feature. You pay the standard S3 storage and request costs for the object, which you were largely paying already if your artifacts lived in S3.

Limitations and open questions

It doesn’t state whether REFERENCE needs KMS decrypt permissions beyond s3:GetObject when you use customer-managed encryption (It should), exactly how a referenced object version is pinned for rollback, or what happens to a live function if the referenced object is later deleted or modified. Any latency effect from reading code directly at runtime isn’t discussed either. Confirm these against the Lambda console and the developer guide before you roll it out widely.

Bottom line

This is a small setting with an outsized effect for teams operating at scale. If the 75 GB quota or “we can’t audit that copy” has ever slowed you down, S3ObjectStorageMode: REFERENCE is worth a look. Start on a non-critical function, get your bucket policy and versioning right, and expand from there. Original announcement on the AWS Compute Blog.

Lambda MicroVMs: When Functions Aren’t Enough and EC2 Is Too Much

Picture this: you’re building a browser-based notebook where data analysts paste in Python, load a 3 GB dataframe, generate a few charts, then wander off to a meeting. Ninety minutes later they come back and expect their kernel, their variables, and their half-finished plot to still be sitting there.

AWS Lambda MicroVMs

Now try to build that on what AWS gave you before June 2026. Lambda? Fifteen-minute execution ceiling, no persistent process between invocations, no way to hold that dataframe in memory across the analyst’s coffee break. ECS or EC2? Sure, but now you’re running per-user containers or VMs, paying for idle capacity, and writing your own scheduler to reap dead sessions. Fargate gets you closer, but you still own the isolation story when the code being run was generated by an LLM you can’t fully trust.

This is the gap Lambda MicroVMs is aimed at.

What actually launched

AWS Lambda MicroVMs is a new compute primitive that exposes the Firecracker virtualization layer (the same one that has always run underneath Lambda) as something you can address directly. You get per-instance hardware isolation, snapshot-based startup, and — this is the part that matters — the ability to keep a single execution environment alive for an entire working session rather than a single request.

Alongside it, AWS shipped a companion resource called the Lambda Network Connector (LNC), which is how you attach a MicroVM to a private VPC when you need it to reach a database or an internal API.

How it actually works

Two resource types, and once you understand these the rest falls into place:

  • MicroVM image: a versioned artifact you build from a Dockerfile. When you create one, the service runs your Dockerfile, boots your application inside a MicroVM, and takes a Firecracker snapshot of the memory and disk state. Think of it as a “warm” image — dependencies already imported, JIT already warmed, whatever init your app does already done.
  • MicroVM: an instance launched from that image. Because it’s restored from a snapshot rather than cold-booted, it comes up close to instantly.

Each MicroVM gets its own HTTPS endpoint, and that endpoint speaks to individual ports on the guest — plain HTTPS, WebSockets, and gRPC all work, so you connect to it the same way you’d connect to any container you were running yourself.

On sizing: the default baseline is 2 GB of memory and 1 vCPU. You can configure that up to 8 GB and 4 vCPUs at launch, with vCPU pinned to memory at a 2:1 ratio (memory in GB is double the vCPU count). From whatever baseline you pick, a MicroVM will auto-scale vertically up to 4x during peak demand. Horizontally, the service claims you can launch several hundred MicroVMs inside a minute.

Lambda MicroVMs Instance Sizes

Sessions can run from a few minutes up to eight hours. Egress to the public internet works out of the box; VPC access requires the LNC.

Architecture

See the diagram below. The pattern is the per-session model: user requests come into your control plane, the control plane looks up whether that user already has a live MicroVM, and either routes traffic to the existing HTTPS endpoint or launches a new instance from a MicroVM image. When the user goes idle, you decide the lifecycle — keep it warm, snapshot it, or let it go.

AWS Lambda MicroVMs Flow

The interesting design question isn’t really “how do I launch a MicroVM” — it’s “who owns the session-to-endpoint mapping, and what’s my policy when the user disconnects?” You will end up building a small state machine. Plan for it.

When this is the right tool

  • Browser-based IDEs, notebooks, and the current wave of vibe-coding platforms — anywhere users bring their own code and expect their environment to feel persistent.
  • Analytics platforms running user- or LLM-generated queries where the working set is large and the session is long.
  • AI coding assistants that iterate on generated code and want to hold context between iterations, including RL-style loops that spin environments up and down to compare execution paths.
  • Security and vulnerability scanning where you need real isolation between scans and sometimes elevated OS privileges inside the guest.
  • CI/CD build and test runners where each job wants a clean, isolated box that starts fast.

When it probably isn’t

The launch material doesn’t call out anti-patterns directly, so treat this as my read rather than AWS’s guidance:

  • If your workload is stateless per-invocation and short, regular Lambda is still simpler and cheaper reasoning-wise.
  • If you need more than 8 GB of memory or 4 vCPUs per instance, you’re outside the MicroVM baseline envelope and you should look at Fargate or EC2.
  • If your sessions genuinely need to outlive eight hours without a snapshot/resume, this isn’t your primitive either.

Security notes worth reading twice

The isolation story is genuinely strong — hardware-level, one guest per user or job, which is the whole point of using Firecracker as a primitive rather than just a container runtime. That’s what makes it defensible as a sandbox for untrusted or model-generated code.

Two things to design around, though.

  1. First, outbound internet access is on by default. If you’re running untrusted code, you almost certainly want egress controls, and the launch material doesn’t spell out how granular those are — treat this as a question to answer before you go to production.
  2. Second, private VPC connectivity is not a checkbox; it requires configuring an LNC. Factor that into your networking design early, not late.

Pricing

Lambda MicroVMs are priced per instance-second. Check the Lambda pricing page (MicroVMs tab) before you commit to a design, especially for long-lived sessions — an eight-hour idle MicroVM has very different economics from a 200 ms function invocation.

Limitations to keep on the whiteboard

  • 2 GB / 1 vCPU baseline, 8 GB / 4 vCPU max, 2:1 memory-to-vCPU ratio.
  • Vertical auto-scale capped at 4x baseline.
  • Horizontal scale advertised as “several hundred per minute” during spikes — quantify this against your own burst profile before betting on it.
  • Sessions in the “few minutes to 8 hours” range; the exact hard upper bound isn’t explicitly stated.
  • VPC access is opt-in via LNC.

Wrapping up

The most useful way to think about Lambda MicroVMs is that AWS unbundled Firecracker from Lambda’s request/response model and let you address it directly, while keeping the parts of Lambda that were actually pleasant — no capacity planning, no patching, no scheduler to write. If you’ve been building per-user sandboxes on top of ECS or bare EC2 and feeling like you were re-implementing Firecracker badly, this is the primitive to evaluate.

Start with the product page and the developer guide, and if snapshot-based startup is new to you, the older SnapStart post is worth a read for the mental model. Networking details live in the LNC docs, and quotas are on the service limits page. And refer to this AWS Blog which walk you through AWS MicroVms.

Know about Amazon GuardDuty Investigation: AI-powered security analysis

Security teams routinely burn hours triaging a single GuardDuty finding. The workflow is familiar: pull the finding, pivot into CloudTrail, cross-reference VPC flow logs, chase the principal across accounts, map the behavior to a known technique, and finally decide whether the alert is a real incident or noise. Multiply that across an organization producing dozens or hundreds of findings a day, and investigation backlog becomes the bottleneck — not detection.

Amazon GuardDuty is now addressing that bottleneck directly. The GuardDuty investigation is available in public preview and performs on-demand, AI-powered investigations of GuardDuty findings, returning a structured assessment that a human analyst would otherwise assemble by hand.

What changed

The investigation agent adds a new capability to Amazon GuardDuty: rather than only surfacing findings, GuardDuty can now investigate them for you. Each investigation produces a structured output that includes a risk level, a confidence score, a natural-language summary, investigation details, MITRE ATT&CK technique mappings, resource mappings, and prioritized recommended actions — including specific AWS CLI commands the analyst can execute.

The agent is accessible through the AWS Management Console, the AWS CLI, AWS APIs, AWS SDKs, and the AWS MCP server — meaning it slots into both console-driven workflows and agentic security tooling built on top of the Agent Toolkit for AWS.

How it works

An investigation is initiated by an caller with appropriate permissions through the GuardDuty console, by calling the CreateInvestigation API, or via the AWS CLI. The target can be:

  • a specific GuardDuty finding
  • a member account
  • an entire AWS organization

CLI and API callers can supply a free-form natural-language trigger prompt of up to 2,048 characters to steer the investigation — for example, focusing the agent on a particular hypothesis or scope.

Under the hood, the agent correlates findings and evidence and returns a structured assessment. Results are retrieved with GetInvestigation, and prior investigations can be enumerated with ListInvestigations.

Processing uses cross-Region inference via the Cross-Region Inference Service (CRIS). Investigation data remains stored in the Region where the investigation was created, but inference and summary generation may occur in another Region within the same geography, transmitted over Amazon’s encrypted network.

Please go through this very detailed AWS blog that walks you through the whole investigation process – Amazon GuardDuty investigation agent: on-demand AI-powered threat assessment.

Architecture

The accompanying diagram shows the request flow: an administrator triggers an investigation from the console, CLI, SDK, API, or an MCP-connected agent; GuardDuty accepts the request in the origin Region; CRIS performs inference in-geography; and the structured assessment is returned to the caller while the underlying investigation data stays resident in the origin Region.

When to use it

The investigation agent fits naturally when you want to:

  • Triage a single GuardDuty finding without a manual pivot chase
  • Assess security posture across an entire AWS organization
  • Reduce manual correlation of evidence spread across CloudTrail, VPC flow logs, and other GuardDuty data sources
  • Wire GuardDuty investigations into AI-assisted security workflows through the AWS MCP server
  • Produce on-demand, structured threat assessments consumable by downstream automation

When not to use it

The agent is not the right tool if:

  • GuardDuty is not enabled in the account or organization
  • You are operating in a Region that does not support the preview
  • You need a member account to view another member’s or the administrator’s investigations, which is not permitted

Security considerations

Architecturally, three points matter for a review.

First, cross-Region inference. Investigation data is stored only in the origin Region, but processing and summary results may traverse another Region in the same geography. If your data-residency posture is scoped to a single Region rather than a geography, this is worth an explicit review before enabling the feature in regulated workloads. Please refer geography and their inference regions here.

Second, transport. Data moves across Amazon’s internal, encrypted network — but the residency point above still stands independently of encryption in transit.

Third, authorization. Investigations honor the existing GuardDuty authorization model. Callers can investigate only accounts they are already authorized to access. The three IAM actions to allow on principals that will drive investigations are:

  • guardduty:CreateInvestigation
  • guardduty:GetInvestigation
  • guardduty:ListInvestigations

Scope these on the administrator account role that runs the SOC’s investigation workflow, not broadly.

Pricing

In Preview mode, GuardDuty Investigations are made available at free of cost. Confirm current pricing in the GuardDuty product page before enabling at scale once its GA.

Limitations

  • The feature is in public preview.
  • GuardDuty must already be enabled, the account must be in a supported Region, and only administrator accounts can create investigations in admin as well as member accounts.
  • Member accounts cannot access investigations belonging to peer members or to the administrator.
  • Cross-Region inference behavior described above also applies.
  • During preview, 10 investigations/account/day with total limit of 100 investigations/account.Failed investigations do not count toward these quotas.
  • This feature is available only in the following 10 commercial AWS Regions: US East (N. Virginia), US East (Ohio), US West (Oregon), Canada (Central), Europe (Frankfurt), Europe (Ireland), Europe (London), Europe (Paris), Europe (Stockholm), and Asia Pacific (Tokyo).
  • The trigger prompt is capped at 2,048 characters when invoked through API/CLI.

Verify latest limitations against the GuardDuty investigation documentation before committing to a design.

Conclusion

The investigation agent shifts GuardDuty from a detection surface to a triage surface. For teams whose incident response cost is dominated by evidence correlation rather than detection, that is the interesting move. Preview status means the mechanics — Regions, quotas, latency, retention — need to be validated against your environment before it lands in a runbook, but the shape of the workflow is clear enough to start piloting against a defined scope, most sensibly a single administrator account or a bounded set of high-signal findings. Details and getting-started guidance are on the AWS Security Blog and in the GuardDuty documentation.