On this page4 sections
A tuning change that completes more requests can still make a service worse for some of its users. Increasing worker concurrency, for example, may improve total throughput while extending the wait for a smaller request class. Whether to keep it depends on the workload and objective, not on whether the largest number in the benchmark rose.
Linux performance tuning is most useful as an investigation: identify the user-visible problem, find a plausible limiting condition and compare a reversible change under the workload that exposed it. The settings become the means of testing an explanation rather than a checklist of adjustable values.
The slow workload and its baseline
Record the latency distribution, throughput, errors and workload shape in both a healthy and an affected interval. Include kernel version, instance type, container limits and recent changes. These details prevent a before-and-after comparison from quietly becoming a comparison between different systems or different traffic.
They also explain when the conclusion needs another look. A concurrency value chosen for one container limit may no longer fit after that limit or the machine type changes. Retaining the original conditions gives the next operator something to compare with the new environment.
These read-only tools can narrow the first investigation when installed; their availability and fields depend on the distribution and packages:
uptime
vmstat 1 5
iostat -xz 1 5
free -m
ss -s
Interpret the fields using local manuals. The first vmstat or iostat report may summarize time since boot instead of the following interval, so it should not be compared blindly with an interval measurement. Read load average separately from CPU utilization, and consider available memory as well as free memory when caches can be reclaimed.
Brendan Gregg's USE-method Linux checklist organizes resource investigation around utilization, saturation and errors. It helps turn the symptom into a question, such as whether requests are waiting on a constrained resource, without requiring a collection of every available counter.
The limits that apply to this process
Spare capacity on the host does not settle the application's situation. A container may be throttled by its CPU quota, and a process may reach its open-file limit while the system-wide limit remains generous. Inspect the relevant cgroup and process configuration before increasing a broader host setting.
The kernel's cgroup v2 documentation documents the controllers and interfaces. Paths from cgroup v1 should not be pasted into a v2 system. On managed platforms, make changes through supported workload configuration rather than editing control files behind the orchestrator.
Storage needs the same attention to the actual limiting condition. Latency may reflect queueing, a cloud-volume limit or the application's access pattern. An I/O scheduler used on an older kernel may not exist or suit the current device, so selecting a familiar name does not establish that the suspected bottleneck changed.
An experiment that could disprove the hypothesis
Write the expected effect before changing a setting. If excessive worker concurrency is overloading a dependency, reduced concurrency should be tested against throughput and tail latency together. Lower peak throughput may be worthwhile if it prevents queue collapse under sustained load; that is a service tradeoff to evaluate, not an automatic failure of the experiment.
Use an isolated test or a canary exposing a limited part of the workload where appropriate. Record the old value, restoration method, test interval and stopping condition. Repeat the load pattern that produced the original problem, keeping unrelated deployments out of the comparison so they do not obscure attribution.
In a hypothetical concurrency test, the higher setting improves average throughput but worsens tail latency for a smaller request class. The broad benchmark passes while that class misses its user-facing objective. The result supports a narrower setting or reversal unless the affected population can meet its objective under the new configuration.
Acceptance and rejection matter too. A lower latency value may result from refusing more work, as the diagram below illustrates. Compare accepted requests, rejected requests, completions and latency under the same offered load before concluding that useful work improved.
Read diagram description
A lower latency value can result from rejecting more requests. Compare acceptance, rejection, completions and latency together before deciding whether the change meets the user-facing objective. Diagram labels: Same offered workload: Keep the request mix and test conditions comparable; Accepted requests: Measure completed work + latency; Rejected requests: Measure the work refused or deferred; Service decision: Does the tradeoff meet the user-facing objective?.
If the observations point to a memory constraint rather than excessive concurrency, choose an experiment that tests that explanation. Virtual-memory controls, including swappiness, dirty-page thresholds, huge pages and swap compression, carry workload-specific tradeoffs. The kernel’s virtual-memory documentation is a starting point for understanding the controls; it does not supply a universal low-latency preset. State why the proposed setting should affect the observed constraint before adding it to the experiment.
Emergency containment may require several coordinated changes. Preserve their sequence and uncertainty, then isolate effects afterward; experimental neatness should not delay a necessary mitigation.
When a tested setting becomes persistent configuration
Before making a setting persistent through normal configuration management, compare user-facing outcomes, resource consumption and failure behavior with the baseline. A throughput gain does not erase the adverse effect on the slower request class, and an isolated successful run does not describe every future workload.
Keep the hypothesis, representative comparison, adverse effect and reversal condition with the setting. Those details let the next owner understand why this value was chosen and reassess it when the request mix or execution limits change.
The filesystem guide helps investigate storage state, while the metrics guide explains distributions. For the concurrency change, the final decision should record what became faster, who waited longer and whether each affected request class still met its service objective under the tested conditions.
Sources & context
Sources linked in this article. Read alongside the author’s analysis; a citation does not independently verify a publisher’s claims.
- USE-method Linux checklistwww.brendangregg.com
- cgroup v2 documentationdocs.kernel.org
- virtual-memory documentationdocs.kernel.org
