Amirhossein Hosseinpouramirhp
CV
Active Internal system · WordPress fleet monitoring

Uptime

Every site we look after, in one channel.

Uptime watches our WordPress fleet around the clock and speaks in Telegram only when something changes: whether a site is reachable, what its plugins did overnight, who was made an administrator and from where, what the store sold and what failed, when the hosting is due, and whether the backups actually exist.

How it works What it watches

An internal system, built at Pigment Dev for ourselves and used every day. It is not for sale and not something you install. It is shown here because it explains how we work.

Pigment Dev system2025 to today
Fleet alerts · Telegram channel · illustrative

Morning

Domain digest

  • Approaching expiry
  • Crossed a threshold overnight
  • Fine

Once, grouped. Birthday reminders arrive in the same slot.

During the day

Plugin changes on a client store

  • updated: the store plugin, a minor version
  • downgrade: the page builder, one version back
  • deactivated: the spam filter
  • installed: a classic editor plugin

Grouped into one message. Sent once. The site stays up.

Open dashboardOpen Mini App

Hosting due this week

A client store, on the monthly plan

The amount is shown in the currency it is billed in, and links to the payment page.

If something breaks

A booking site is down

With a screenshot of what a visitor sees, from our own headless browser.

A booking site is back

Down for

One message each way. Nothing in between.

End of day

The summary

SiteUpChecks
A client store
An agency site
A booking site
A publisher

A day in the channel

  • Morning: one digest for domains and dates. It reads like a calendar, not like an alarm.
  • During the day: a message only when something moved, each one sent once.
  • If something breaks: one alert with a screenshot, then one when it recovers, with how long it was gone.
  • End of day: one table for the whole fleet. On most days it is a column of full marks, which is the point of reading it.

Illustrative. Sites are anonymised and values are hidden; the real channel is private.

// why

Why we built it instead of buying one

There was no single incident. There was a list of things we kept finding out too late, and a market that answered a different question from the one we were asking.

Before

Most mornings started the same way. Open a site, wait for it to load, open the next one. Dozens of tabs, to confirm that nothing had fallen over in the night. It felt like diligence. It caught almost nothing that mattered.

The things that hurt were never full outages. A plugin would update itself and quietly break a page nobody visited that morning, and the first we heard of it was the client asking why their customers could not check out. There is a particular kind of embarrassment in being told about your own site by the person paying for it.

Then there were the bills. A hosting renewal came due on a date that lived in someone's inbox and nowhere else, and we found out when the site went dark. A domain drifted toward expiry because the registrar's reminder went to an address nobody read any more.

The tools we looked at

They answered one question, is the site up, and charged per site to answer it.

None of them looked inside WordPress. None knew what a Jalali date was, or that a hosting invoice might be billed on a day of the month that drifts against the Gregorian calendar, or that our support team needed something they could paste into a message to a client. And every one of them wanted to email us, when the place our team actually lives is a Telegram channel.

So we built it

One autumn we started small: a free monitor on someone else's cloud that visited each site in turn and posted to Telegram when one went down. For a few weeks that was enough.

Then we added the next question we kept asking, and the one after that. It moved to our own server when it outgrew the free tier, grew a database when the files got unwieldy, and grew a small plugin that lives on every site and reports back. Somewhere along the way it stopped being an uptime checker and became the place we go to find out what is happening. It has never been for sale. It is the tool we would have paid for, if it had existed.

// how-it-works

How it works

Three parts, and the third is the one that matters. Most monitoring is loud. This is built to stay quiet until it has something to say.

First

A reporter on every site

One file in each WordPress install. It answers a cheap question about plugins, themes, core, errors and disk constantly, and a heavier one about the database, users and store only when someone opens a screen. The split keeps watching a site from becoming a load on it.

Then

The monitor asks, every few minutes

It loads each site the way a visitor would, checks the page is real and not a white screen, and asks the reporter what changed. A failed connection is retried before anything is decided: a false alarm at night teaches a team to ignore the channel.

Finally

Telegram hears about changes, not checks

An alert goes out once when a state changes, and once when it recovers. Never on every run. A plugin or disk warning never marks a site down. Each signal is its own message, grouped per site, so a real outage is never buried under routine.

// what-it-watches

What it watches

Fifteen signals per site, each checked on its own schedule and reported on its own. Only reachability decides whether a site is up.

  • Reachability
  • Plugins
  • Errors
  • Disk
  • SSL
  • Domains
  • Hosting renewals
  • Backups
  • Core version
  • Security plugin
  • New admins
  • Database
  • Store
  • Forms
  • Findability

Reachability, with judgement

A site is up when it returns a real page of a sane size that carries WordPress's own REST marker. A blank white screen with a success code is not up. Connection failures are retried first.

Required plugins

Each site can name the plugins that must stay installed and active. If someone deactivates the store, we hear about it before the store's customers do.

Plugin drift

Every install, removal, activation, deactivation and version change, in one message per site. Downgrades are flagged on their own: a rollback is either a fix or a sign of tampering.

WordPress errors

Recovery mode, paused plugins and paused themes, reported from inside WordPress even when the front page still renders. The homepage is scanned for the critical error message too.

Disk and capacity

Host usage on a VPS, quota through the cPanel API where there is no shell, or the WordPress folder's own size. An alert when a threshold is crossed, and one when it is back under.

SSL and domain expiry

Certificates checked directly. Domain expiry read from the registry, refreshed daily, delivered as one morning digest. An expiry date is a calendar entry, not an emergency.

WordPress core

Whether each site runs the current release, checked against WordPress itself. It fails soft: if WordPress.org cannot be reached, nothing turns red.

Whether the site can be found

The robots file, fetched over HTTP the way a crawler gets it, since a cache or CDN can serve something different from the disk. The question is whether the site has quietly told search engines to leave.

Backups that really exist

The newest archive is read from disk, not from the backup plugin's records. The plugin knows what it believes it built. The disk knows what survived a cleanup job or a full volume.

The security plugin, present and awake

Installed, active, and which version. A security plugin someone switched off to "fix something" and forgot is a common way a site ends up unguarded.

New administrators

The monitor keeps the list of administrators and compares it on every run, rather than trusting an event. A new one, however it was created, raises an alert with when it appeared and the address it came from.

A picture of the failure

When a site goes down, a screenshot of what a visitor sees is taken through our own headless browser and attached to the alert. No third-party screenshot service.

// inside-a-site

Inside each site

Open a site in the dashboard or the Mini App and the reporter answers the expensive questions: scanning tables, summing orders. Asked only when a screen opens, never polled.

Overview

Core version against the current release, with a red asterisk beside the site's name wherever it appears when it is behind. Security plugin, backup plugin and the newest archive on disk. PHP, memory limit and the address the site answers from.

Database

Size, table count and the largest tables. How much of the options table loads on every request, and its heaviest entries. Transients and trash. A posts table that doubled since last week is a story the front page will not tell you.

Users and admins

Every administrator, when the account was registered and when it last signed in. The view you open when the new-admin alert fires, to see who else is in there.

WooCommerce

Orders by status, sales for the month day by day, low stock kept apart from out of stock, and which order emails are on and who receives them. A spike in failed orders is how you find a gateway that quietly stopped.

Elementor forms

Every form, the page it lives on, where it sends, whether it saves submissions, how many it has taken and when the last one arrived. If someone points a form at a different address, it shows here.

Per-site switches

Plugin drift, errors, disk, SSL, domain expiry and new-admin alerts can each be muted per site, and a site can be marked as having no reporter at all. A brochure site and a busy store do not need the same attention.

// beyond-watching

Beyond watching

The parts of running a fleet that are about people and money rather than servers.

Hosting renewals

Every account with its due date, cycle, price and provider. Reminders arrive with three buttons: paid, snooze, let it expire. A missed date goes overdue instead of silently rolling forward. Dollars, euros or Toman, as billed.

Two calendars, one date

Gregorian and Jalali shown together, written the same way and joined as one unit, so nobody converts in their head.

Cycles as hosting is sold

Monthly, two-monthly, quarterly, six-monthly, yearly or one fixed date, on a Gregorian or a Jalali day of the month. Each advances from the calendar it was set in, so a month-end plan does not creep earlier.

A report for people who are not us

Plain text the support team can hand to a client: what is due, when, for how much, by domain rather than internal name. Sites nobody is tracking are listed too, because that is the point.

The months ahead

Hosting renewals, domain expiries and certificate expiries in one list, by date, with anything overdue pinned to the top. Per site, or for everything.

And birthdays

The one thing here that is not about a server. Stored in whichever calendar they were given in, with a reminder a week before and the day before, in the channel we already read every morning.

// where-it-lives

Where it lives

Telegram first, because that is where the team already is. A Mini App and a dashboard for when a message is not enough.

The channel

Every message carries the site's own favicon as an inline emoji, from a custom set the bot builds and owns, so a row of alerts reads as a row of recognisable sites. The same two buttons, dashboard and Mini App, sit under every message.

The Mini App

Runs inside Telegram. Heatmaps by year and month, and an up-next list that merges renewals, domains and SSL, because "what do I deal with soon" does not care which system owns the date.

The dashboard

Private, with passkey login and a rate-limited password as fallback. Per-site settings, live probing of a site's reporter and disk before you save, and mutes per alert type.

A signed public API

For the cases where a client or a status page needs to read uptime without a login.

// watchdog

It watches itself

The one failure a monitor cannot see is its own.

One day in August the server that runs the monitor wedged. Cloudflare showed an error page. SSH accepted the connection and hung up before saying hello. And the Telegram channel stayed completely silent, because the thing that sends the alerts was the thing that had died.

So there is now a watchdog somewhere else entirely, on Cloudflare's edge. Every few minutes it asks the monitor one question, and if the answer is too old, it speaks into the same channel the monitor normally uses.

The question

Not "are you running", which a wedged process answers cheerfully, but "when did you last actually check a site".

A pause is a decision, not a fault. When monitoring is paused on purpose, the watchdog knows, and stays quiet.

// lessons

Things it got wrong first

A tool used every day for a year has a list of its own mistakes. These are the ones worth telling, because each one changed how it works.

It said a plugin was missing when it was there

A site with broken permalinks answered the REST path with a page instead of data. The monitor read that as "reporter not installed" and sent us to upload a file that was already in place.

Now It tries the second way WordPress exposes the same route before concluding anything, and tells a routing problem from a missing plugin.

It blamed the token when the door was locked

A security plugin had closed the REST API to anonymous callers. The monitor reported an authentication failure, and we rotated a token that was fine.

Now It recognises that specific refusal for what it is, and says so.

Pressing "paid" twice walked a bill a year ahead

Each tap advanced the billing cycle. On a snoozed account it did nothing visible at all, because an early return left the account asleep.

Now "Paid" is two things: recording a payment, which always happens and clears any snooze, and advancing the cycle, which happens only for a bill that is actually due.

The confirmation appeared behind the tab bar

In the Mini App, the toast confirming an action rendered inside the glass of the bottom bar. Paid and Snooze looked as if they did nothing.

Now Feedback sits above the bar, where a thumb can see it.

It painted the fleet red when WordPress.org was slow

The core check compared each site with the latest release. When that source was unreachable, everything looked out of date.

Now An unknown latest version turns nothing red. A stale answer beats a false one.

It could not report its own death

The August outage. The alerter was the casualty, so nothing alerted.

Now A watchdog off the server asks when it last did its job, and speaks when the answer is wrong.

// rules

Rules it follows

Most of these were learned the hard way, and are written into the code so they cannot be unlearned.

Change, not state
One message when something goes wrong, one when it recovers. A channel that repeats itself gets muted, and a muted channel is worse than none.
Warnings never mean down
Plugin drift, disk, SSL, domains and WordPress errors are separate signals. Only reachability decides.
Fail soft
If the thing we compare against is unreachable, nothing turns red.
Read the disk, not the record
A backup plugin's table says what it believes it built. The filesystem says what is there.
Read as a visitor would
The robots file is fetched like a crawler fetches it, because that is what matters.
Cheap checks constantly, heavy ones on request
So watching a site never becomes a load on it.
Never blind without saying so
A reporter that stops answering is its own alert. Silence from a checker is not good news.
Pause is open, not closed
A corrupt or ambiguous pause flag reads as running. Failing closed would silence everything with nobody aware.

Built for us, and shown because it is the clearest example we have of what we mean by building things properly.

Separate and much smaller, I have published a Cloudflare Worker uptime monitor with KV state, Telegram alerts, downtime duration tracking and stale-backup detection, under the MIT licence.

cloudflare-simple-uptime-monitor

WordPressREST APITelegram Bot APITelegram Mini AppsCloudflare WorkersHeadless browsercPanel APIPasskeys
branch main 6 active projects ↑ 113 releases products/uptime.md Sari --:-- UTC+3:30 its@amirhp.com