Amirhossein Hosseinpouramirhp
CV
Service · Servers, caching, migrations

Performance and infrastructure

When a store falls over, it is almost never one setting.

When a site gets slow, expensive or unstable, I trace it across code, database, cache, PHP and server instead of guessing at plugin settings. Around that sits the rest of the layer underneath: Linux server configuration, hosting migrations, backups, and monitoring that reaches a person.

Tell me what is falling over Read the outage case

Where a slow request loses its time

A request passes through six layers before it becomes a page. I read all of them, in order, because the cause is rarely in the layer where the symptom shows. The findings come from one store, written up in full further down.

  1. Browser

    JavaScript, admin-ajax, heartbeat

    What I look at: What the page asks for after it has loaded: AJAX calls, heartbeats, counters that go around the cache.

    On one store: A page-view counter fired an admin-ajax POST on every front-end view, bypassing the page cache. The editor heartbeat ran every 10 seconds.

  2. Edge

    Cloudflare, CDN, WAF

    What I look at: Cache rules, bot and rate limits, and what reaches the origin at all.

    On one store: Search was 21% of all traffic, from a widely distributed set of addresses. This is where I would stop it first next time. Why.

  3. Web server

    LiteSpeed, Nginx, .htaccess

    What I look at: Cache headers, rewrite rules, and whatever an old plugin left in the config.

    On one store: An orphaned cache block in .htaccess forced a four-minute max-age over the LiteSpeed TTL.

  4. PHP

    PHP 8.3, OPcache, workers

    What I look at: Whether the opcode cache is loaded, worker and memory limits, and the version the host actually runs.

    On one store: The main cause. OPcache was not loaded at all on the web SAPI. Every request compiled WordPress and every plugin from source.

  5. Object cache

    Redis

    What I look at: Whether repeated queries come from memory or are rebuilt on every request.

    On one store: Redis object caching was switched on as one of the smaller fixes.

  6. Database

    MySQL, options, autoload

    What I look at: Slow queries, table sizes, autoloaded options, and rows nobody needs any more.

    On one store: Each search ran a LIKE scan across 757,000 rows. 615,750 of them were legacy log entries. One security plugin option had grown to 33MB.

What I do

The layer under the site: where I get called when something is slow or falling over, and where most of the cost of running a store quietly sits.

Performance rescue

A store that is slow, expensive or down, traced to its causes and fixed in order of cost. Sometimes the first finding explains most of it, as it did in the case below.

Caching, layer by layer

OPcache, Redis, LiteSpeed page cache and Cloudflare rules, set up to work with each other. On one store that brought cached pages to 54ms time to first byte.

Server configuration

Linux, cPanel and WHM, HestiaCP, LiteSpeed and Nginx. PHP versions, workers and limits, and the defaults a host ships that do not suit WordPress.

Hosting migrations

What moves, in what order, and what has to stay up while it happens, written down as a checklist the next move can repeat.

Monitoring that reaches a person

Alerts in the chat your team already reads, sent when something changes rather than on every check.

Backups that really exist

Off-provider copies, checked on disk rather than trusted from a plugin's own records. Especially for a database the host's backups do not cover.

Case study · Client not named, at their preference

From constant 503s to 15% CPU

A software download store with 30 million unique visitors a month was returning HTTP 503 for most of every business day. I traced it to three compounding causes and brought the account from a pinned 100/100 CPU to roughly 15/100, without changing hosts.

Account CPU
Before

100/100

After

about 15/100

Within minutes of enabling OPcache through the host's PHP selector.

CPU faults per hour
Before

100 to 183

After

near zero

Outside the search bursts. The sustained 503s stopped the same day.

21% of all traffic was search, each a LIKE scan across 757,000 rows, about 20 seconds long
615,750 of the 757,000 rows were legacy log entries, with 3.6 million rows of metadata behind them
33MB in a single security plugin option, cleared
120s between editor heartbeats, up from 10 seconds, set in a must-use plugin

How it was traced

I worked alone, with the store owner and the hosting provider's support team.

cPanelLiteSpeedLSCachePHP 8.3OPcacheMySQLRedisEasy Digital Downloadsmu-pluginsBash

Context

A large Persian software download store, about 30 million unique visitors a month, running Easy Digital Downloads rather than WooCommerce, on shared cPanel hosting with LiteSpeed and PHP 8.3. The database sits on a separate server, which means the host's own backups do not cover it.

The problem

The account hit its CloudLinux CPU ceiling every morning from 08:00 onward. Most of the 503s landed on admin-ajax.php. A resource upgrade raised the disk and MySQL quotas but left the CPU limit untouched, so the outages continued. At one point during the incident the host suspended terminal access.

Constraints

Shared hosting with a hard CPU ceiling and no root. A 4.4GB database outside the host's backup coverage. Live revenue traffic that could not be paused. And a theme and plugin set the client did not want replaced.

What I did

The first ten minutes

OPcache was not loaded at all on the web SAPI. Every single request was compiling WordPress and every plugin from source. Enabling it through the host's PHP selector dropped account CPU from 100/100 to about 15/100 within minutes.

The second finding took the access logs. Search requests were roughly 21% of all traffic, arriving from a widely distributed set of IP addresses, and each one ran a LIKE scan across 757,000 rows in the posts table, occupying a worker for about twenty seconds. Third, a page-view counter plugin was firing an admin-ajax POST on every front-end view, which bypassed the page cache entirely.

Around those, a set of smaller fixes. An orphaned cache block left in .htaccess was forcing a four-minute max-age over the LiteSpeed TTL, so I rewrote the file. The editor heartbeat went from 10 seconds to 120 through a must-use plugin. Redis object caching was enabled. A security plugin option had grown to 33MB and was cleared.

Finally I traced the row count itself: 615,750 of those posts were legacy log entries left behind by the Easy Digital Downloads 3.0 migration, along with 3.6 million rows of their metadata. Legacy logging had stopped in 2022 while the replacement table had been running since 2019, which made the old rows safe to purge.

Result

The sustained 503s stopped the same day OPcache was enabled. CPU faults fell from between 100 and 183 per hour to near zero outside the search bursts. The client received a full written report in Persian documenting every change and every file that should not be edited without care.

Moving sites, and watching them

Two pieces of infrastructure work at Pigment Agency, where I run development: one move, and the system that has watched the sites since.

Twelve production sites, cPanel to HestiaCP

From

cPanel hosting

Twelve production sites, plus a Node and pm2 application and an SSO service.

To

One VPS on HestiaCP

  • Redis, with per-site cache prefixes
  • Backups held off the provider
  • Monitoring from a Cloudflare Worker

The move followed a documented checklist, written so the next migration can repeat it rather than rediscover it.

A monitor that tells a person

I built a Cloudflare Worker to watch those sites: KV for state, Telegram for alerts, with downtime duration tracking and stale-backup detection.

Uptime, the monitoring system Pigment now runs its WordPress fleet on, started the same way: a free monitor on someone else's cloud, posting to Telegram when a site went down. Its rules are the ones I apply to any monitoring I set up.

Change, not state
One message when something breaks, one when it recovers.
Warnings never mean down
Only reachability decides whether a site is up.
Read the disk, not the record
A backup is real when the archive is on disk.
Fail soft
An outage somewhere else must not paint the fleet red.

How Uptime works

How an engagement goes

When a site is down, the order still matters: evidence first, then the change that explains the most.

  1. Access and evidence

    Hosting panel, logs and database, with whatever access you can give. I read the access logs and resource graphs before I change anything, and I say what I could not see.

  2. The cause that explains the most

    Fixed first, and measured. On the store above, the first finding took ten minutes and explained most of it.

  3. The smaller fixes around it

    Configuration, cleanup and caching, each one recorded with what it changed and why.

  4. A written report

    Every change, and every file that should not be edited without care, in English or Persian, so your team can maintain it without me.

  5. Monitoring or maintenance, if you want it

    So the next problem reaches a person before it reaches your customers. Optional, and separate.

A good fit when

A store has become slow, expensive or unstable and nobody knows why, or a migration cannot afford downtime.

Not a fit when

You want a speed score raised with one more plugin setting. I work on the cause, across the whole stack, or not at all.

branch main 6 active projects ↑ 113 releases services/performance-infrastructure.md Sari --:-- UTC+3:30 its@amirhp.com