[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"blog-faqs-peak-load-capacity-planning-before-go-live-en":3,"blog-post-peak-load-capacity-planning-before-go-live-en":28},[4,8,12,16,20,24],{"answer":5,"id":6,"question":7},"A load test checks that the system meets its criteria at the target load, such as the concurrent users estimated for opening day. A stress test keeps increasing the load until the system starts to fail, to show how much headroom there is and which component breaks first. For a system with a clear peak day, we recommend running both before go-live.","ddhb7z32rtcz1xs","What is the difference between a load test and a stress test?",{"answer":9,"id":10,"question":11},"There is no universal figure. The margin should reflect how confident you are in the assumptions behind the peak estimate. If the number comes from real data from a previous opening day, a smaller margin is reasonable. For a new system with no history, allow more, and keep testing until something breaks so you know the real ceiling.","jybe7oh9u5z4xzm","How far above the estimated peak should a load test target be?",{"answer":13,"id":14,"question":15},"In some cases, but it risks affecting real users and real data. We usually recommend a separate environment with instance sizes and data volumes close to production. If testing on production is unavoidable, do it before real users are let in, with an agreed plan for removing the test data afterwards.","ibwva4ufv0oiv8h","Can we load test directly on production?",{"answer":17,"id":18,"question":19},"Replace that service with a stub during the test and get the provider's rate limit in writing. If their ceiling is below your estimated peak, the options are to negotiate a higher limit, redesign the flow so it calls that service less often, or use a waiting room to control the arrival rate.","2pumrzdp48fprdh","What if a third-party provider, such as the SMS OTP service, will not allow load testing?",{"answer":21,"id":22,"question":23},"Not necessarily. A waiting room is a decision made in advance that once demand passes the proven level, users queue in order instead of everyone facing slowdowns and random errors. Even a system that scales well has ceilings from quota, the database or third-party services, and a waiting room keeps those ceilings from turning into an outage.","f8nqup9v7papt69","Does using a waiting room mean the system cannot handle the load?",{"answer":25,"id":26,"question":27},"Before the next known peak, and after any change that affects the main user path, such as adding steps to a form, switching a third-party provider or moving regions. If re-runnable test scripts were handed over, the in-house team can do this without waiting for the vendor.","yu21qg752lbi9zn","After go-live, when should we load test again?",{"id":29,"parentId":30,"title":31,"excerpt":32,"content":33,"image":34,"date":35,"isoDate":36,"views":37,"tag":38,"categorySlug":39,"slug":40,"to":41,"readTime":42,"isHighlight":43,"sortOrder":44,"video":45,"video_id":45,"keywords":46},"b68ks35n6j5qby4","avgdrdxnmsg448l","How Many Users Can It Take at Once? Capacity Planning and Load Testing Before Go-Live","Estimate peak from real business events, load-test the database and third parties too, know your cloud quota ceilings, and put capacity targets into acceptance criteria.","\u003Cp>A registration system that opens at 9:00 a.m. on a date announced in advance will usually carry the heaviest load of its life in the first ten minutes, not in month twelve. That is when a question that belonged in the design phase comes back: how many people can this system handle at the same time, and has anyone actually proven that number? The people who have to answer it to management that morning are on your team. The cloud provider is not in the room.\u003C\u002Fp>\u003Cp>This piece covers the decisions that come before go-live: where the peak estimate comes from, what a load test has to include before it proves anything, where the cloud ceilings are, how the system should behave when demand exceeds what you tested, and how to write all of it into acceptance criteria. Alerting after launch is a separate problem, which we cover in \u003Ca target=\"_blank\" rel=\"noopener noreferrer\" href=\"https:\u002F\u002Fwww.superdev.tech\u002Fen\u002Fblogs\u002Fmonitoring-alerting-what-to-page-on\">our article on monitoring and alerts\u003C\u002Fa>.\u003C\u002Fp>\u003Ch2>The peak comes from the business calendar, not the average\u003C\u002Fh2>\u003Cp>Design documents usually carry a monthly or daily user count. That number is useful for sizing storage and nearly useless for sizing the peak. A system with 200,000 users a month can run comfortably for weeks and still fall over on the one morning when everyone arrives at once.\u003C\u002Fp>\u003Cp>Most peaks in enterprise and government systems are predictable, because they are tied to business events: the day registration opens, the day results are announced, the last day to submit documents, the day a story runs in the news. Each event has its own shape. Opening day tends to spike within minutes. A deadline tends to build steadily until the final hour.\u003C\u002Fp>\u003Cp>Our approach starts with the number of people who are actually eligible to use the system, then asks the business two questions: what share of them will arrive in the opening window, and how long does each person stay in the system? Little's Law turns those answers into concurrency: the number of people in the system at once equals the arrival rate multiplied by the time each person spends there.\u003C\u002Fp>\u003Cp>Here is an example with hypothetical numbers. With 200,000 eligible users and an expectation that 20% arrive in the first 30 minutes, that is 40,000 people in 30 minutes, or roughly 1,300 per minute. If each person takes about 10 minutes to fill in the form and submit documents, the system will be holding around 13,000 concurrent users.\u003C\u002Fp>\u003Cp>That figure is still a half-hour average, and the first minute after the announcement is usually heavier. What matters most at this stage is that every assumption is written down and signed off by someone on the business side. If more people show up than expected, everyone can go back and see which assumption missed.\u003C\u002Fp>\u003Cp>A number nobody owns tends to become IT's fault.\u003C\u002Fp>\u003Ch2>A load test proves something only if it walks the user's path\u003C\u002Fh2>\u003Cp>A common load test fires a large volume of requests at the home page and reports tens of thousands of requests per second. That result usually measures the CDN or the cache more than the system, because the home page barely touches the database. What breaks on the day usually sits behind the submit button.\u003C\u002Fp>\u003Cp>A load test worth its report covers at least four things:\u003C\u002Fp>\u003Cul>\u003Cli>\u003Cp>Realistic user journeys: log in, request an OTP, fill in the form, attach files, submit, and come back to check status, in proportions close to the real day. People refreshing the status page again and again are load too.\u003C\u002Fp>\u003C\u002Fli>\u003Cli>\u003Cp>A database with production-scale data. A query that is fast on a thousand rows can be slow on ten million, and any row that many users update at once, such as a seat or quota counter, produces lock waits you never see when one tester clicks through.\u003C\u002Fp>\u003C\u002Fli>\u003Cli>\u003Cp>The third-party services the system depends on: the SMS provider behind the OTP, the payment gateway, another agency's verification API. Each has its own rate limit, and none of them scales with your instance count.\u003C\u002Fp>\u003C\u002Fli>\u003Cli>\u003Cp>A spike profile: from zero to target within a few minutes. We prefer this to a slow ramp, because opening day does not give the system time to grow into the load.\u003C\u002Fp>\u003C\u002Fli>\u003C\u002Ful>\u003Cp>Third-party services are the hardest call. Many providers do not allow load testing against their production endpoints. What we do is replace the service with a stub during the test, and separately get the provider's rate limit in writing. If their ceiling is below your estimated peak, adding capacity on your side will not fix it.\u003C\u002Fp>\u003Cp>The pass criteria have to be agreed before the test starts: at 13,000 concurrent users, 95% of requests must complete within how many seconds, and what error rate is acceptable. Criteria set after seeing the results tend to get adjusted until the system passes. Once the target is met, we recommend pushing the load further until something breaks, so you know how much headroom there is and which component gives way first.\u003C\u002Fp>\u003Ch2>Autoscaling has ceilings, and some have nothing to do with budget\u003C\u002Fh2>\u003Cp>The appeal of the cloud is that you can add machines in minutes. Those minutes, and the ceilings that never show up on the bill, are the two things that most often undo a peak plan.\u003C\u002Fp>\u003Cp>The first ceiling is quota. Google Cloud sets CPU quota per region, counting the vCPUs of every VM in that region, and a managed instance group can only scale if there is quota available for every resource it uses. Google's documentation also states plainly that quotas do not guarantee resources will always be available (source: \u003Ca target=\"_blank\" rel=\"noopener noreferrer\" href=\"https:\u002F\u002Fdocs.cloud.google.com\u002Fcompute\u002Fresource-usage\">Google Cloud, Allocation quotas\u003C\u002Fa>, read 29 September 2026).\u003C\u002Fp>\u003Cp>Microsoft Azure enforces vCPU quota in two tiers per region: a regional total and a per-VM-family limit. Exceed either one and the deployment is not allowed. Azure also checks quota and capacity separately, so a deployment can fail even with enough quota if the region or zone has no capacity left for the size you asked for. Azure's documented answer for guaranteed capacity is on-demand capacity reservation (source: \u003Ca target=\"_blank\" rel=\"noopener noreferrer\" href=\"https:\u002F\u002Flearn.microsoft.com\u002Fen-us\u002Fazure\u002Fvirtual-machines\u002Fquotas\">Microsoft Learn, vCPU quotas\u003C\u002Fa>, read 29 September 2026).\u003C\u002Fp>\u003Cp>In practice, an approved budget does not add machines by itself. Check quota in every region you use, request increases weeks before the event, and for a system that cannot afford to fail, weigh whether reserving capacity is worth it.\u003C\u002Fp>\u003Cp>The second ceiling is time. The autoscaler has to see the load before it adds instances, and each new instance has to boot and warm up before it can serve traffic. Google Cloud calls that warm-up the initialization period and sets it to 60 seconds by default. Predictive autoscaling, which scales out ahead of load, works best when a workload varies predictably on daily or weekly cycles (source: \u003Ca target=\"_blank\" rel=\"noopener noreferrer\" href=\"https:\u002F\u002Fdocs.cloud.google.com\u002Fcompute\u002Fdocs\u002Fautoscaler\">Google Cloud, Autoscaling groups of instances\u003C\u002Fa>, read 29 September 2026). A registration day that happens once a year does not fit that pattern. The more direct approach is to raise the minimum instance count to peak level before opening time, and lower it once the peak has passed.\u003C\u002Fp>\u003Cp>The third ceiling is the database. The web tier scales out easily. The primary database usually cannot be resized mid-event the same way, and every new web instance opens more connections to it. A common pattern is autoscaling that works exactly as configured, followed by an outage because the database ran out of connections first. Database size and connection pool settings have to be decided from load test results, before the event.\u003C\u002Fp>\u003Ch2>A waiting room is choosing to be slow in order, instead of failing at random\u003C\u002Fh2>\u003Cp>However good the estimate, more people may turn up than planned. The question is what the system should do once demand passes the level you proved. If nobody decides in advance, the system decides for you: everyone slows down together, then errors start at random. People halfway through a form get dropped, try again, and add more load.\u003C\u002Fp>\u003Cp>A waiting room is a decision made in advance to hold some visitors back so the people already inside can finish. Cloudflare Waiting Room uses two thresholds to decide when to start queuing: total active users and new users per minute. The queueing method available to everyone is first in, first out; other methods such as random order are limited to certain plans. The waiting page refreshes to show an estimated wait time (source: \u003Ca target=\"_blank\" rel=\"noopener noreferrer\" href=\"https:\u002F\u002Fdevelopers.cloudflare.com\u002Fwaiting-room\u002Fabout\u002F\">Cloudflare Waiting Room\u003C\u002Fa>, read 29 September 2026). Before you count on it in the plan, confirm that your organization's Cloudflare plan actually includes it.\u003C\u002Fp>\u003Cp>The thresholds must come from the number your load test proved. If the system passed at 13,000 concurrent users, setting the threshold at 20,000 so that fewer people wait is closing the door after the system has already fallen over.\u003C\u002Fp>\u003Cp>A waiting room has a cost you have to accept. Users will see that they are waiting, and some will complain. That makes it a decision for the business owner of the system, together with a message that tells users plainly why they are waiting. For public services, a fair and explainable order often matters more than speed, because the question after the event will be who got in first, and why.\u003C\u002Fp>\u003Ch2>Writing the capacity target into acceptance criteria\u003C\u002Fh2>\u003Cp>All of these numbers carry weight only once they are in the acceptance documents. If the TOR only says the system must support a large number of users, nobody can prove whether it passed. Verifiable criteria should state at least the following:\u003C\u002Fp>\u003Cul>\u003Cli>\u003Cp>The target number of concurrent users, and the mix of user journeys used in the test.\u003C\u002Fp>\u003C\u002Fli>\u003Cli>\u003Cp>The 95th-percentile response time and the maximum acceptable error rate at that load.\u003C\u002Fp>\u003C\u002Fli>\u003Cli>\u003Cp>How long the load must be sustained, for example 30 minutes continuously.\u003C\u002Fp>\u003C\u002Fli>\u003Cli>\u003Cp>The test environment, the data volume, and which third-party services were real and which were stubbed.\u003C\u002Fp>\u003C\u002Fli>\u003Cli>\u003Cp>Who from the organization witnesses the test, the report to be handed over, and test scripts that can be re-run.\u003C\u002Fp>\u003C\u002Fli>\u003C\u002Ful>\u003Cp>The last point matters more than it looks. Test scripts handed over to the organization let the in-house team re-run the test before the next peak without waiting for the vendor.\u003C\u002Fp>\u003Ch2>When this is more than you need\u003C\u002Fh2>\u003Cp>A formal load-testing programme has real costs: time to write scripts, prepare test data, build an environment, and coordinate with third-party providers. For an internal system with a few dozen users, usage spread through the day, and no date when everyone has to log in at once, that effort is not worth it.\u003C\u002Fp>\u003Cp>In that case, we recommend checking only two things: that the heaviest report or query is still fast enough on real data volumes, and that the infrastructure has room as data grows over the next year. If the system is later opened to outside users, or gains a deadline that everyone hits at once, that is the time to revisit this.\u003C\u002Fp>\u003Ch2>A common objection: \"We're on the cloud, it scales on its own\"\u003C\u002Fh2>\u003Cp>There is truth in this. Autoscaling on Google Cloud and Microsoft Azure works, and it spares you from buying peak capacity for the whole year the way on-premises forced you to. But autoscaling only scales what it is configured to scale, within the quota you have, and only after it sees the load.\u003C\u002Fp>\u003Cp>The database, row locks and third-party rate limits do not scale with it. A load test works alongside autoscaling: it tells you how to configure it, what minimum instance count to set, and which parts of the system need fixing before the day. The team that runs the system every day often knows some of these answers already. The test turns that knowledge into numbers that management can make a decision on.\u003C\u002Fp>\u003Ch2>Where to start with only a few weeks left\u003C\u002Fh2>\u003Col>\u003Cli>\u003Cp>Write the peak estimate on a single page with every assumption stated, and have the business owner sign off on it.\u003C\u002Fp>\u003C\u002Fli>\u003Cli>\u003Cp>List every third-party service on the main user path and ask each provider for its rate limit in writing.\u003C\u002Fp>\u003C\u002Fli>\u003Cli>\u003Cp>Check quota in the regions you use today, and file increase requests immediately if it falls short.\u003C\u002Fp>\u003C\u002Fli>\u003Cli>\u003Cp>Test one main user journey at target load on production-scale data before expanding to other journeys.\u003C\u002Fp>\u003C\u002Fli>\u003Cli>\u003Cp>Set waiting room thresholds from the test results, and prepare the message users will see.\u003C\u002Fp>\u003C\u002Fli>\u003Cli>\u003Cp>Schedule when the minimum instance count goes up before opening, and when it comes back down after the peak.\u003C\u002Fp>\u003C\u002Fli>\u003C\u002Fol>\u003Cp>The order deliberately puts anything that depends on other people first. Answers from providers and quota approvals take time your team cannot speed up, while writing test scripts can run in parallel.\u003C\u002Fp>\u003Cdiv data-type=\"horizontalRule\">\u003Chr>\u003C\u002Fdiv>\u003Ch2>Summary\u003C\u002Fh2>\u003Cp>\"How many users can it take at once?\" should have a tested number as its answer before go-live. If you take one thing from this piece, write the capacity target and its pass criteria into the acceptance documents. Once the number is on paper, the peak estimate, the test and the quota requests each get an owner and a deadline.\u003C\u002Fp>\u003Cp>If your system has an opening day or an announcement day coming up and you would like someone to review the peak estimate or the test plan, we are glad to talk. There is no deadline and no need to decide quickly.\u003C\u002Fp>\u003Cp>To discuss the details of a project, you can reach us at 088-983-9386 or contact@superdev.tech. Our office is in Bang Kapi, Bangkok, and more about us is at \u003Ca target=\"_blank\" rel=\"noopener noreferrer\" href=\"https:\u002F\u002Fwww.superdev.tech\">superdev.tech\u003C\u002Fa>. For organizations preparing to launch a system, we are happy to meet first just to understand the problem, and if it turns out your system does not need a full load-testing programme, we will tell you so.\u003C\u002Fp>","https:\u002F\u002Ftwsme-r2.tumwebsme.com\u002Fpbc_2128795511\u002Fb68ks35n6j5qby4\u002Fcover_cover_j57f9c1i0l.webp","September 29, 2026","2026-09-29 04:16:29.596Z",106,"Enterprise","enterprise","peak-load-capacity-planning-before-go-live","\u002Fblogs\u002Fpeak-load-capacity-planning-before-go-live",12,false,0,"",[47,48,49,50,51,52,53],"capacity planning","load testing","autoscaling","Cloudflare Waiting Room","Google Cloud","Microsoft Azure","go-live readiness"]