Skip to content

perf(infra): size the EU ingest fleet to its traffic - #1215

Merged
Makisuo merged 1 commit into
perf/ingest-self-telemetry-samplingfrom
perf/infra-eu-ingest-right-size
Oct 2, 2026
Merged

Makisuo merged 1 commit into
perf/ingest-self-telemetry-samplingfrom
perf/infra-eu-ingest-right-size

Conversation

@Makisuo

@Makisuo Makisuo commented Oct 2, 2026 •

Copy link
Copy Markdown
Collaborator

Stacked on #1214.

Problem

The EU prd instance runs the same footprint as US: 2x c7gd.large, autoscaling 2-6, collector at 512/1024, and two load balancers. Over the last 30 days its ingest ALB served 34k requests (US: 136M), its hosts moved ~2 GB, and memory peaked at 0.6% of the task. That's roughly $300/mo for almost no traffic.

Change

The EU decision lives in one private isEuPrd(stage, region) in packages/infra/src/aws/stage.ts. US prd and previews behave exactly as before.

US prd / previews EU prd
Ingest tasks / autoscaling 2, 2-6 1, 1-3 (same CPU target and cooldowns)
Instance c7gd.large c7gd.medium (keeps local NVMe for WAL fsync)
EC2 task size 2048 / 3072 1024 / 1536
Collector 512 / 1024 256 / 1024
  • INGEST_EC2_INSTANCE_TYPE and INGEST_EC2_TASK_SIZE become resolveIngestEc2InstanceType and resolveIngestEc2TaskSize, which take (stage, region). resolveIngestDesiredCount, resolveIngestScaling and resolveCollectorTaskSize now also take region.
  • Electric is unchanged. Its 512 MiB size has never been measured under sync load.
  • docs/infra.md Cost decisions gains a bullet explaining how to raise EU back.

Rolling deploys with one host

No deployment config change. ECS defaults to 100% minimum healthy / 200% maximum. The new task can't bind the port on the old host, so it waits; managed scaling adds a host (the ASG maximum is maxTasks * 2 = 6), and the old task drains once the new one is healthy. This is the same path US already uses. Minimum healthy 0 would drop EU ingest on every deploy.

Risks

  • EU loses AZ redundancy. A dead host means a few minutes of EU ingest downtime while it's replaced, with WAL segments recovered from S3 by the replacement.
  • The instance type change replaces both EU hosts on the next deploy.
  • The 48 GiB WAL cap on a 59 GB NVMe leaves about 7 GiB free. That's fine at EU volume; a per-region cap is the follow-up if it ever matters.

Tests

bunx vitest run src/aws/stage.test.ts in packages/infra: 25 passed, covering EU vs US for every resolver.


View with [code]smith Autofix with [code]smith
Need help on this PR? Tag @codesmith-bot with what you need. Autofix is disabled.

@coderabbitai

coderabbitai Bot commented Oct 2, 2026 •

Copy link
Copy Markdown

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

Note

Currently processing new changes in this PR. This may take a few minutes, please wait...

⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Advanced

Run ID: c0f06189-c90d-483f-a251-e3585bda3bda

📥 Commits

Reviewing files that changed from the base of the PR and between de19115 and 1392dab.

📒 Files selected for processing (4)
  • apps/ingest/alchemy.run.ts
  • docs/infra.md
  • packages/infra/src/aws/stage.test.ts
  • packages/infra/src/aws/stage.ts
 ______________________________________________
< Looking for trouble in all the right places. >
 ----------------------------------------------
  \
   \   \
        \ /\
        ( )
      .( o ).
✨ Finishing Touches
📝 Generate docstrings
  • Commit to this branch
  • Create a new PR
🧪 Generate unit tests (beta)
  • Commit to this branch
  • Create a new PR
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Autopilot is currently an internal CodeRabbit preview.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@maple-review-bot

maple-review-bot Bot commented Oct 2, 2026 •

Copy link
Copy Markdown

Maple review

🟢 Confidence 4/5 · likely safe to merge
Sizing-only change confined to the isEuPrd branches, with a test per resolver; the one deliberate cost is that a lost EU host now means EU ingest downtime.
quality 100/100 · no findings · tests covered · risk medium

Right-sizes the EU prd ingest fleet behind one isEuPrd(stage, region) gate: 1 task (1-3), a c7gd.medium with a 1024/1536 task, collector 256 CPU, Electric unchanged. US prd and previews keep their sizing. Safe to merge.

  • isEuPrd gates EU prd sizing in packages/infra/src/aws/stage.ts
  • resolveIngestDesiredCount and resolveIngestScaling now take region: EU 1, 1-3
  • EC2 fleet resolves to c7gd.medium with a 1024/1536 task
  • EU prd collector drops to 256 CPU; Electric keeps the prd size
What was checked
  • Every call site of the five region-taking resolvers was updated, and no consumer of the removed INGEST_EC2_* constants is left at the head commit
  • The EU collector keeps 1024 MiB, so the 768 MiB hard memory_limiter in packages/infra/otel-collector still fires
  • EU ASG maxSize is scaling.max * 2 = 6, which covers the two hosts a 100/200 rolling deploy needs

3d03a98 · Updated on every push. Reply "won't fix" to dismiss a finding, or mention @maple-review-bot to ask about one.

Over the last 30 days the EU ingest ALB served 34k requests (~2 GB through
the hosts) against the US's 136M (~2.3 TB in, 8 TB out), yet EU prd ran the
US footprint: 2x c7gd.large, collector at 512/1024, ~$300/mo.

EU prd now runs one c7gd.medium (task 1024/1536, autoscaling 1-3 on the same
CPU target and cooldowns) and the collector at the non-prd size. Sizing
functions take the region; US values are unchanged. Electric keeps its size.
A single-host deploy still rolls through managed scaling with ECS's default
100/200 deployment config.
@Makisuo
Makisuo force-pushed the perf/infra-eu-ingest-right-size branch from 3d03a98 to 1392dab Compare October 2, 2026 18:42
@maple-review-bot

maple-review-bot Bot commented Oct 2, 2026 •

Copy link
Copy Markdown

Maple review

🟢 Confidence 4/5 · likely safe to merge
EU prd's ingest drops to one task on one host: I verified the ASG ceiling (6, from scaling.max * 2) still allows a two-host roll, but that is the choice to accept here.
quality 100/100 · no findings · tests covered · risk medium

Sizes only the EU prd ingest fleet to its traffic: one c7gd.medium with a 1024/1536 task, autoscaling 1-3, and the non-prd collector size, via a private isEuPrd(stage, region). US prd, previews and Electric are unchanged.

  • isEuPrd(stage, region) gates every EU prd branch
  • resolveIngestEc2InstanceType/resolveIngestEc2TaskSize replace the two constants and take (stage, region)
  • resolveIngestDesiredCount, resolveIngestScaling, resolveCollectorTaskSize take region; EU prd is 1 task, 1-3 scaling, 256 CPU collector
What was checked
  • US prd and previews are byte-identical: every resolver returns the old value for ("prd", "us"), pr-12, dev-alice (stage.ts:94-207, 310)
  • ASG ceiling still rolls one host: maxTasks = 3 for EU, so maxSize = 6 (alchemy.run.ts:383-389)
  • Disk math holds: 59 GB NVMe is ~55 GiB usable against a 48 GiB cap, and orphan-recovered frames also increment live_bytes (telemetry.rs:1506)

1392dab · Updated on every push. Reply "won't fix" to dismiss a finding, or mention @maple-review-bot to ask about one.

@Makisuo
Makisuo added this pull request to stack #1216 October 2, 2026 20:05
@Makisuo
Makisuo merged commit b2062e2 into main Oct 2, 2026
39 of 50 checks passed
@Makisuo
Makisuo deleted the perf/infra-eu-ingest-right-size branch October 2, 2026 20:07
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant