Description
WordPress has no honest health endpoint. The front page, admin-ajax.php and the usual
health.php stubs all return 200 while the database is down, the object cache is gone, or
WordPress is serving its “database update required” interstitial. Every probe in common
use calls that pod ready.
Klarsmith Ops Kit gives a containerised WordPress site three things operators actually need:
- Honest readiness β
GET /wp-json/ops/v1/readyzreturns 200 only when the database
answers, the schema matches the running core version, the object cache round-trips, and
the uploads directory is writable. Anything else is a 503 with the failing check names.
Liveness is deliberately left alone: a database outage must drain pods, never restart
them. - Prometheus metrics β
GET /wp-json/ops/v1/metricsserves the text exposition
format from a snapshot thatwp ops collectwrites to the object cache on a schedule.
Nothing expensive runs on scrape.wp_ops_snapshot_age_secondstells you when the
collector has stopped. Pod-level series (wp_ops_pod_*) and site-level series
(wp_ops_site_*) are split so replicas do not duplicate site numbers. - Structured JSON logs β with
WP_OPS_LOG_JSON=true, PHP errors and fatals go to
stderr as one JSON object per line, ready for a log pipeline.
The plugin fails closed. /metrics is disabled until a token is configured, and the
anonymous readiness response names failing checks without leaking detail.
Configuration is by environment variable, which is how containers are configured:
WP_OPS_TOKENβ bearer token for/metrics(Authorization: Beareror
X-Ops-Token). Unset disables the endpoint.WP_OPS_SITE_NAMEβsitelabel on JSON log records (metrics carry no site
label; the scraper’s own labels identify the site).WP_OPS_EXPECT_OBJECT_CACHEβ fail readiness when no external object cache is active.WP_OPS_REQUIRED_PLUGINSβ comma-separated plugin files (dir/plugin.php) that
must be active for readiness.WP_OPS_REST_BYPASS_AUTHβ set tofalseto stop the plugin allowing anonymous
access to its own REST namespace.WP_OPS_LOG_JSONβ set totrueto emit JSON log lines on stderr.
WP-CLI commands: wp ops check (exit 1 on any failing check), wp ops collect
(refresh the snapshot), wp ops metrics (print the exposition locally).
Source, issue tracker and Kubernetes manifests for probes, the collector CronJob and
scraping live at https://github.com/klarsmith/wp-ops-kit.
Installation
- Install from the WordPress plugin directory, or with Composer
(composer require klarsmith/wp-ops-kit; the Composer package installs as
wp-content/plugins/wp-ops-kit). - Activate the plugin.
- Point your readiness probe at
/wp-json/ops/v1/readyz. - Run
wp ops collecton a schedule (a five-minute CronJob is the reference setup). - Set
WP_OPS_TOKENand scrape/wp-json/ops/v1/metricswith that bearer token.
If a security plugin blocks anonymous REST requests, the plugin allows its own ops/v1
namespace through at priority 1 unless WP_OPS_REST_BYPASS_AUTH=false. Add ops to the
security plugin’s allowlist as well where it has one.
FAQ
-
Why not use the front page as the readiness probe?
-
Because it returns 200 while the site is broken. The database-upgrade interstitial, a
missing object cache and a read-only uploads volume all serve a 200 front page. -
Why does readiness not cover liveness too?
-
If liveness consulted the database, one database blip would fail liveness on every pod
of every site at once and restart-storm the fleet. Readiness drains; liveness restarts.
Only readiness should depend on WordPress. -
Why is /metrics computed from a snapshot?
-
Walking the cron array and summing autoloaded option sizes on every scrape would load a
shared database for no benefit. The collector does the work once; scrapes read the
result. If the collector dies the snapshot age climbs, which is alertable. -
Does it work behind WPML or a language prefix?
-
Yes. Route detection uses the REST route query variable with a request-URI fallback, so
/en/wp-json/ops/v1/readyz works the same as/wp-json/ops/v1/readyz.
Reviews
There are no reviews for this plugin.
Contributors & Developers
“Klarsmith Ops Kit” is open source software. The following people have contributed to this plugin.
ContributorsTranslate “Klarsmith Ops Kit” into your language.
Interested in development?
Browse the code, check out the SVN repository, or subscribe to the development log by RSS.
Changelog
0.1.7
- Directory listing name is “Klarsmith Ops Kit” (slug
klarsmith-ops-kit);
Composer/GitHub name unchanged.
0.1.6
- Earlier directory listing name “Ops Kit” (superseded by 0.1.7).
Tested up to WordPress 7.1.
0.1.3
- Code conforms to the WordPress Coding Standards; PHPStan level 6 clean. No
behaviour change beyond sanitising two request inputs used for route detection. - Copy-paste manifests in
examples/for Kubernetes, stock Prometheus,
VictoriaMetrics and a Grafana dashboard.
0.1.2
/metricsno longer appends its exposition to a response another handler has already
served.- Zero-valued post statuses are dropped at collection time;
publishis always exported
so a drop to zero stays alertable.
0.1.1
- Allow the plugin’s own
ops/v1REST namespace through site-level anonymous-REST
lockdowns (hooked at priority 1, never overriding an earlier decision).
Escape hatch:WP_OPS_REST_BYPASS_AUTH=false.
0.1.0
- Initial release: readiness endpoint, snapshot-backed Prometheus metrics, JSON logging,
wp ops check|collect|metrics commands.
