Skip to content

Internal Maintenance Jobs ​

Hegemony runs internal maintenance from the scheduler service. The scheduler periodically calls the API endpoint POST /internal/maintenance/tick with X-Internal-Token; the API owns job due checks, leases, execution, and last-run state in the maintenance_jobs table.

Jobs ​

  • run_status_sync reconciles stale non-terminal run rows with Temporal.
  • approval_timeouts expires pending approval requests whose DB timeout has elapsed. Temporal normally handles active approval steps; this job repairs drift. Each expiry is recorded as an approval.expired audit entry carrying the request's state before and after, so an approval that lapsed is told apart from one a person decided.
  • stale_monitors marks non-terminal monitor rows as STOPPED when their parent run is terminal. Monitor workers observe this DB stop signal and stop their in-memory probe loops.
  • run_event_pruning deletes old run_events rows for terminal runs. It is disabled by default with interval 0.
  • audit_log_retention deletes audit entries older than HEGEMONY_AUDIT_RETENTION_DAYS in bounded batches, and records each pass as an audit_log.pruned entry. It is a no-op until that retention is set, so nothing is deleted by default.
  • inventory_provider_sync scans configured inventory providers whose per-provider sync interval has elapsed and materializes provider-backed devices/sites. Rows a provider stops reporting are marked stale rather than deleted, and restored if it reports them again; each tick logs those counts, with sites counted apart from devices.
  • file_repository_backfill provisions the managed internal file repository for organizations that predate that wiring or whose creation-time provisioning failed, and re-runs the bundled object store's bucket bootstrap (HEGEMONY_S3_BUCKET, HEGEMONY_S3_EXTRA_BUCKETS, the seed directory) so a store that was down at API startup gets its buckets without a restart. It never waits on the store: an unreachable endpoint is recorded in the run's notes rather than failing the run, since uploads already surface that.
  • sandbox_image_pruning reclaims dangling images and build cache on the Docker-in-Docker sandbox daemon, whose dind-data volume otherwise grows without bound. It touches only regenerable data — tagged images, containers, and volumes are left to their owners — and deletes nothing on a daemon that does not carry the sandbox marker volume with this stack's HEGEMONY_SANDBOX_TOKEN. Every stack runs the sandbox, so each way of not reaching it is a broken deployment rather than one without the feature: a missing or non-tcp:// HEGEMONY_CONTAINER_DOCKER_HOST, a missing token, a daemon that cannot be reached, and a daemon that answers without the marker or with a different token (sandbox marker token mismatch, typically a token rotated without restamping the marker; see Rotating the sandbox token). In each case the job fails, with the reason in last_error, and platform health turns degraded, where a skip would report a clean sweep of a sandbox that is never pruned.

Anything a job changes is recorded in the audit log as system:maintenance, so a change no person made is still attributable to the job that made it. The job runner binds that actor around every handler, so a job needs no audit wiring of its own.

The three platform sync jobs are the exception. They drive the same sync engine a person drives from the UI, and they bind system:scheduler around their own work so a scheduled export, drift plan, or auto-apply is attributed to the scheduler rather than to maintenance in general. Each operation is recorded whoever started it, with trigger_source in the entry's details telling a scheduled run apart from a manual one.

Configuration ​

All settings use the HEGEMONY_ prefix and are read from environment variables:

SettingDefaultNotes
HEGEMONY_MAINTENANCE_ENABLEDtrueEnables the scheduler loop.
HEGEMONY_MAINTENANCE_TICK_INTERVAL_SECONDS30Scheduler tick cadence.
HEGEMONY_MAINTENANCE_TICK_TIMEOUT_SECONDS120HTTP timeout for one tick.
HEGEMONY_MAINTENANCE_JOB_LEASE_SECONDS600Reclaim stale running jobs after this lease.
HEGEMONY_MAINTENANCE_RUN_STATUS_SYNC_INTERVAL_SECONDS3000 disables the job.
HEGEMONY_MAINTENANCE_RUN_STATUS_SYNC_BATCH_SIZE200Maximum runs reconciled per run status sync execution.
HEGEMONY_MAINTENANCE_APPROVAL_TIMEOUTS_INTERVAL_SECONDS3000 disables the job.
HEGEMONY_MAINTENANCE_APPROVAL_TIMEOUTS_BATCH_SIZE200Maximum pending approvals scanned per approval timeout execution.
HEGEMONY_MAINTENANCE_STALE_MONITORS_INTERVAL_SECONDS3000 disables the job.
HEGEMONY_MAINTENANCE_STALE_MONITORS_BATCH_SIZE500Maximum stale monitors stopped per job execution.
HEGEMONY_MAINTENANCE_RUN_EVENT_PRUNING_INTERVAL_SECONDS0Disabled by default.
HEGEMONY_MAINTENANCE_RUN_EVENT_RETENTION_DAYS90Applies only to terminal runs.
HEGEMONY_MAINTENANCE_RUN_EVENT_PRUNING_BATCH_SIZE5000Maximum run_events rows deleted per pruning batch.
HEGEMONY_MAINTENANCE_RUN_EVENT_PRUNING_MAX_BATCHES5Maximum pruning batches per job execution.
HEGEMONY_MAINTENANCE_PLATFORM_SYNC_EXPORT_INTERVAL_SECONDS3000 disables platform sync export maintenance.
HEGEMONY_MAINTENANCE_PLATFORM_SYNC_DRIFT_INTERVAL_SECONDS9000 disables platform sync drift-plan maintenance.
HEGEMONY_MAINTENANCE_PLATFORM_SYNC_AUTO_APPLY_INTERVAL_SECONDS0Disabled by default; controls platform sync auto-apply maintenance.
HEGEMONY_MAINTENANCE_PLATFORM_SYNC_BATCH_SIZE10Profiles processed per platform sync maintenance tick.
HEGEMONY_MAINTENANCE_INVENTORY_PROVIDER_SYNC_INTERVAL_SECONDS600 disables inventory provider auto-sync maintenance.
HEGEMONY_MAINTENANCE_INVENTORY_PROVIDER_SYNC_BATCH_SIZE20Provider configs processed per inventory sync maintenance tick.
HEGEMONY_MAINTENANCE_SANDBOX_IMAGE_PRUNING_INTERVAL_SECONDS216000 disables the job. Fails, pruning nothing, when the sandbox endpoint or token is missing, the daemon is unreachable, or it is not the attested sandbox.
HEGEMONY_MAINTENANCE_SANDBOX_PRUNE_BUILD_CACHEtrueAlso prune the sandbox daemon's build cache.

The scheduler and API must share the same HEGEMONY_INTERNAL_API_TOKEN; otherwise the maintenance tick endpoint fails closed with internal-token auth errors.

Admin UI ​

Admins can inspect maintenance health from Settings → Maintenance Jobs. The page is read-only: it lists the registered jobs, derived health state, last run, next scheduled run, counters, duration, and persisted diagnostics from the last execution. It intentionally does not call POST /internal/maintenance/tick or trigger jobs manually; scheduling remains owned by the scheduler service.

Health states are derived by the API from registry configuration plus the maintenance_jobs row:

HealthMeaning
healthyLast completed run succeeded and the next run is not due yet.
failedLast completed run failed; inspect last_error.
overdueThe job is enabled and past its next expected run time.
runningA current API worker lease is active.
stale_runningA previous running lease expired and can be reclaimed.
disabledGlobal maintenance is disabled or the job interval is 0.
never_runThe job is registered but has no completed run yet.

Released as open source under the AGPL-3.0-or-later license. Development is sponsored by Rexonix s.r.o.. Contact — [email protected].