FinOpsAugust 29, 20266 min
New post

The end-of-month Pantitlán station: why scaling well is more than scaling cheap

Sizing infrastructure so it never fails, without looking at real usage data, almost always means overpaying the whole month to cover a peak that happens a few days. How to detect silent overprovisioning and why autoscaling needs to know the business calendar, not just the time of day.

As clients grew larger —more collections portfolio, more placement volume—, the infrastructure had to scale along with them to hold the SLA commitments on asynchronous processing. There, there was no room to over-save. There was room to tune better. This second part is the mirror image of the first: production, where the challenge is the exact opposite of the non-production environments. If you haven't read the first installment yet — about cost per user, benchmarks, and contract gaps — start there.

Read Part 1: users, non-production environments, and the contract gap

How do you detect silent compute overprovisioning?

That tuning started with something as simple as observing. At some point eight instances of a certain compute role had been provisioned, under the logic of "just set this up so it never fails." Monitoring, over time, revealed something different: 70% of the time, those eight instances were more than enough. One, sometimes two, were left over.

Sizing for zero failure without looking at real usage data is almost always sizing over the whole 100% of the time to cover a peak that happens 10% of the time. The cost of overprovisioning doesn't show up as a dramatic incident. It shows up as a monthly bill slightly higher than necessary, month after month, silently accumulating.

What does Mexico City's metro have to do with autoscaling?

But that 10%, in the banking and microfinance business, had a very clear pattern. As the end of the month approached, when all the field advisors wanted to close their collections or hit their placement goal, traffic spiked in a way that can only be described with a very Mexican image: the Pantitlán metro station between 7 and 8 in the morning. Everyone, at the same time, trying to pass through the same place.

Why isn't office-hours autoscaling enough?

Scaling by office hours solves the daily pattern. But in industries with closing cycles —banking, microfinance, collections, payroll, seasonal retail— there is also a periodic pattern tied to specific dates of the month or year. Designing autoscaling rules that only consider the time of day leaves your infrastructure blind to the most predictable —and paradoxically most avoidable— business peaks: the ones you already know, with anticipation, when they'll arrive.

Containing the organic growth of cost was never about turning off spending. It was about learning to read, with enough notice, when Pantitlán was going to fill up again.

Does your infrastructure know what day of the month it is, or only what time it is?

#FinOps#Autoscaling#CloudOps#Fintech#Microfinance#CapacityManagement#Azure#CloudScars
Share:LinkedIn
Quick answerDetail

Why does sizing infrastructure for zero failure end up costing more?

Because sizing for zero failure without looking at real usage data almost always means sizing over the whole 100% of the time to cover a peak that happens 10% of the time. Overprovisioning doesn't show up as an incident, but as a slightly higher bill month after month. The real fix has two steps: use monitoring to detect the surplus compute, and make autoscaling aware of the business calendar (end of month, payday, seasons) — not just office hours — so it anticipates the peaks you already know are coming.

Written and reviewed by Rogelio Barajas González — certified Lead Auditor ISO 27001:2022 and ISO 9001:2015, with direct experience in SOC 1 Type 2 and SOC 2 Type 2. Founder of Barajas Advisory.

Verify his credentials on LinkedIn:linkedin.com/in/rogelio-barajas-gonzalez

Last updated: August 2026

This is one of nine real cases

Cicatrices de Nube — do you want the rest of the stories?

All nine documented cases —FinOps, Release Management, Service Delivery, Compliance, and AI governance— with a self-assessment checklist per chapter and an overall scorecard.

Download the free playbook

Does this resonate?

If you lead operations, technology, or teams at a SaaS company and recognize these situations, let's talk. No strings attached.

Schedule your diagnosis
Usually available

I respond within 2 hours max
Monday to Friday