Measure the workflow before choosing a scaling change. A slow screen can come from a database query, a large response, an external API, model latency, or work repeated unnecessarily.
Those causes need different fixes. Adding application instances without locating the constraint can leave the user waiting for the same dependency.
Define the workload you need to support
Describe the tasks, data sizes, concurrency, and acceptable response behavior. Include background jobs and integration traffic. A daily import can affect a user-facing workflow even when the number of signed-in users is small.
Record a baseline with representative data. Track successful completions, latency, errors, and resource use where those measurements are available. Keep the test conditions with the results so a later comparison means something.
Do not turn a single demonstration into a capacity claim. A result observed with synthetic data and limited concurrency applies to those conditions.
Find the expensive part of a completed task
Inspect the queries and calls the task makes. Look for unbounded reads, repeated requests, large payloads, and dependencies that dominate the response time. Apply pagination, caching, indexing, or asynchronous processing only where they fit the observed behavior and correctness requirements.
Caching needs an explicit freshness policy. Background processing needs a visible job state and a way to investigate failures. A retry needs limits and a defined duplicate policy. Each optimization introduces behavior the operator must understand.
For model-backed features, distinguish time spent retrieving data from time spent generating or reviewing the output. Check whether repeated calls add useful information before making them faster.
Match the proposed fix to the observed delay
Consider an order dashboard that reads complete order histories just to show the latest status. If the large response and repeated processing dominate the task, request only the needed fields and use an appropriate bounded query. If the query itself is slow, inspect its filters and access pattern before selecting an index. These are diagnostic possibilities, not measured results from a particular application.
If the dashboard instead waits on a delivery provider, adding database capacity may do little for that wait. Decide whether the task can show a recently retrieved event with its timestamp, or whether it requires a fresh response. Caching changes the freshness promise, so the interface and business rules must agree with it.
A long-running import may belong in background processing. That requires a visible job identifier, progress or status, and a defined response to retries and partial completion. Moving the work out of the request does not remove the need to know whether it finished.
For each change, write the expected effect and the correctness condition it must preserve. Then compare the same workload before and after. This prevents a faster response from being mistaken for an improvement when it returns stale or incomplete information the workflow cannot accept.
Plan operational ownership alongside capacity
Define alerts that point to user-visible problems, such as tasks that cannot complete. Give the operator an investigation path and a procedure for reducing load or disabling a failing feature.
Establish retention, archival, and recovery requirements as records accumulate. Recheck access boundaries as the number of teams or customers grows. More capacity does not establish stronger isolation.
OBTO provides the platform destination for applications built through its tools, but your app still needs workload-specific evaluation. Review its current commercial terms and verify any capacity or service commitment relevant to your deployment.
After an indexing change, rerun the query and the user workflow with the same data. If query time improves but the screen still waits on an external API, the remaining delay has a different owner.