perf(infra): size the EU ingest fleet to its traffic - #1215
Conversation
|
Navigate logical layers of code changes, visualize relationships, and explore their blast radius. Note Currently processing new changes in this PR. This may take a few minutes, please wait... ⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Advanced Run ID: 📒 Files selected for processing (4)
✨ Finishing Touches📝 Generate docstrings
🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Maple review🟢 Confidence 4/5 · likely safe to merge Right-sizes the EU prd ingest fleet behind one
What was checked
|
Over the last 30 days the EU ingest ALB served 34k requests (~2 GB through the hosts) against the US's 136M (~2.3 TB in, 8 TB out), yet EU prd ran the US footprint: 2x c7gd.large, collector at 512/1024, ~$300/mo. EU prd now runs one c7gd.medium (task 1024/1536, autoscaling 1-3 on the same CPU target and cooldowns) and the collector at the non-prd size. Sizing functions take the region; US values are unchanged. Electric keeps its size. A single-host deploy still rolls through managed scaling with ECS's default 100/200 deployment config.
3d03a98 to
1392dab
Compare
Maple review🟢 Confidence 4/5 · likely safe to merge Sizes only the EU prd ingest fleet to its traffic: one c7gd.medium with a 1024/1536 task, autoscaling 1-3, and the non-prd collector size, via a private
What was checked
|
Stacked on #1214.
Problem
The EU prd instance runs the same footprint as US: 2x c7gd.large, autoscaling 2-6, collector at 512/1024, and two load balancers. Over the last 30 days its ingest ALB served 34k requests (US: 136M), its hosts moved ~2 GB, and memory peaked at 0.6% of the task. That's roughly $300/mo for almost no traffic.
Change
The EU decision lives in one private
isEuPrd(stage, region)inpackages/infra/src/aws/stage.ts. US prd and previews behave exactly as before.INGEST_EC2_INSTANCE_TYPEandINGEST_EC2_TASK_SIZEbecomeresolveIngestEc2InstanceTypeandresolveIngestEc2TaskSize, which take(stage, region).resolveIngestDesiredCount,resolveIngestScalingandresolveCollectorTaskSizenow also takeregion.docs/infra.mdCost decisions gains a bullet explaining how to raise EU back.Rolling deploys with one host
No deployment config change. ECS defaults to 100% minimum healthy / 200% maximum. The new task can't bind the port on the old host, so it waits; managed scaling adds a host (the ASG maximum is
maxTasks * 2= 6), and the old task drains once the new one is healthy. This is the same path US already uses. Minimum healthy 0 would drop EU ingest on every deploy.Risks
Tests
bunx vitest run src/aws/stage.test.tsinpackages/infra: 25 passed, covering EU vs US for every resolver.Need help on this PR? Tag
@codesmith-botwith what you need. Autofix is disabled.