Internal Maintenance Jobs
Hegemony runs internal maintenance from the scheduler service. The scheduler periodically calls the API endpoint POST /internal/maintenance/tick with X-Internal-Token; the API owns job due checks, leases, execution, and last-run state in the maintenance_jobs table.
Jobs
run_status_syncreconciles stale non-terminal run rows with Temporal.approval_timeoutsexpires pending approval requests whose DB timeout has elapsed. Temporal normally handles active approval steps; this job repairs drift. Each expiry is recorded as anapproval.expiredaudit entry carrying the request's state before and after, so an approval that lapsed is told apart from one a person decided.stale_monitorsmarks non-terminal monitor rows asSTOPPEDwhen their parent run is terminal. Monitor workers observe this DB stop signal and stop their in-memory probe loops.run_event_pruningdeletes oldrun_eventsrows for terminal runs. It is disabled by default with interval0.audit_log_retentiondeletes audit entries older thanHEGEMONY_AUDIT_RETENTION_DAYSin bounded batches, and records each pass as anaudit_log.prunedentry. It is a no-op until that retention is set, so nothing is deleted by default.inventory_provider_syncscans configured inventory providers whose per-provider sync interval has elapsed and materializes provider-backed devices/sites. Rows a provider stops reporting are marked stale rather than deleted, and restored if it reports them again; each tick logs those counts, with sites counted apart from devices.file_repository_backfillprovisions the managed internal file repository for organizations that predate that wiring or whose creation-time provisioning failed, and re-runs the bundled object store's bucket bootstrap (HEGEMONY_S3_BUCKET,HEGEMONY_S3_EXTRA_BUCKETS, the seed directory) so a store that was down at API startup gets its buckets without a restart. It never waits on the store: an unreachable endpoint is recorded in the run's notes rather than failing the run, since uploads already surface that.sandbox_image_pruningreclaims dangling images and build cache on the Docker-in-Docker sandbox daemon, whosedind-datavolume otherwise grows without bound. It touches only regenerable data — tagged images, containers, and volumes are left to their owners — and deletes nothing on a daemon that does not carry the sandbox marker volume with this stack'sHEGEMONY_SANDBOX_TOKEN. Every stack runs the sandbox, so each way of not reaching it is a broken deployment rather than one without the feature: a missing or non-tcp://HEGEMONY_CONTAINER_DOCKER_HOST, a missing token, a daemon that cannot be reached, and a daemon that answers without the marker or with a different token (sandbox marker token mismatch, typically a token rotated without restamping the marker; see Rotating the sandbox token). In each case the job fails, with the reason inlast_error, and platform health turnsdegraded, where a skip would report a clean sweep of a sandbox that is never pruned.
Anything a job changes is recorded in the audit log as system:maintenance, so a change no person made is still attributable to the job that made it. The job runner binds that actor around every handler, so a job needs no audit wiring of its own.
The three platform sync jobs are the exception. They drive the same sync engine a person drives from the UI, and they bind system:scheduler around their own work so a scheduled export, drift plan, or auto-apply is attributed to the scheduler rather than to maintenance in general. Each operation is recorded whoever started it, with trigger_source in the entry's details telling a scheduled run apart from a manual one.
Configuration
All settings use the HEGEMONY_ prefix and are read from environment variables:
| Setting | Default | Notes |
|---|---|---|
HEGEMONY_MAINTENANCE_ENABLED | true | Enables the scheduler loop. |
HEGEMONY_MAINTENANCE_TICK_INTERVAL_SECONDS | 30 | Scheduler tick cadence. |
HEGEMONY_MAINTENANCE_TICK_TIMEOUT_SECONDS | 120 | HTTP timeout for one tick. |
HEGEMONY_MAINTENANCE_JOB_LEASE_SECONDS | 600 | Reclaim stale running jobs after this lease. |
HEGEMONY_MAINTENANCE_RUN_STATUS_SYNC_INTERVAL_SECONDS | 300 | 0 disables the job. |
HEGEMONY_MAINTENANCE_RUN_STATUS_SYNC_BATCH_SIZE | 200 | Maximum runs reconciled per run status sync execution. |
HEGEMONY_MAINTENANCE_APPROVAL_TIMEOUTS_INTERVAL_SECONDS | 300 | 0 disables the job. |
HEGEMONY_MAINTENANCE_APPROVAL_TIMEOUTS_BATCH_SIZE | 200 | Maximum pending approvals scanned per approval timeout execution. |
HEGEMONY_MAINTENANCE_STALE_MONITORS_INTERVAL_SECONDS | 300 | 0 disables the job. |
HEGEMONY_MAINTENANCE_STALE_MONITORS_BATCH_SIZE | 500 | Maximum stale monitors stopped per job execution. |
HEGEMONY_MAINTENANCE_RUN_EVENT_PRUNING_INTERVAL_SECONDS | 0 | Disabled by default. |
HEGEMONY_MAINTENANCE_RUN_EVENT_RETENTION_DAYS | 90 | Applies only to terminal runs. |
HEGEMONY_MAINTENANCE_RUN_EVENT_PRUNING_BATCH_SIZE | 5000 | Maximum run_events rows deleted per pruning batch. |
HEGEMONY_MAINTENANCE_RUN_EVENT_PRUNING_MAX_BATCHES | 5 | Maximum pruning batches per job execution. |
HEGEMONY_MAINTENANCE_PLATFORM_SYNC_EXPORT_INTERVAL_SECONDS | 300 | 0 disables platform sync export maintenance. |
HEGEMONY_MAINTENANCE_PLATFORM_SYNC_DRIFT_INTERVAL_SECONDS | 900 | 0 disables platform sync drift-plan maintenance. |
HEGEMONY_MAINTENANCE_PLATFORM_SYNC_AUTO_APPLY_INTERVAL_SECONDS | 0 | Disabled by default; controls platform sync auto-apply maintenance. |
HEGEMONY_MAINTENANCE_PLATFORM_SYNC_BATCH_SIZE | 10 | Profiles processed per platform sync maintenance tick. |
HEGEMONY_MAINTENANCE_INVENTORY_PROVIDER_SYNC_INTERVAL_SECONDS | 60 | 0 disables inventory provider auto-sync maintenance. |
HEGEMONY_MAINTENANCE_INVENTORY_PROVIDER_SYNC_BATCH_SIZE | 20 | Provider configs processed per inventory sync maintenance tick. |
HEGEMONY_MAINTENANCE_SANDBOX_IMAGE_PRUNING_INTERVAL_SECONDS | 21600 | 0 disables the job. Fails, pruning nothing, when the sandbox endpoint or token is missing, the daemon is unreachable, or it is not the attested sandbox. |
HEGEMONY_MAINTENANCE_SANDBOX_PRUNE_BUILD_CACHE | true | Also prune the sandbox daemon's build cache. |
The scheduler and API must share the same HEGEMONY_INTERNAL_API_TOKEN; otherwise the maintenance tick endpoint fails closed with internal-token auth errors.
Admin UI
Admins can inspect maintenance health from Settings → Maintenance Jobs. The page is read-only: it lists the registered jobs, derived health state, last run, next scheduled run, counters, duration, and persisted diagnostics from the last execution. It intentionally does not call POST /internal/maintenance/tick or trigger jobs manually; scheduling remains owned by the scheduler service.
Health states are derived by the API from registry configuration plus the maintenance_jobs row:
| Health | Meaning |
|---|---|
healthy | Last completed run succeeded and the next run is not due yet. |
failed | Last completed run failed; inspect last_error. |
overdue | The job is enabled and past its next expected run time. |
running | A current API worker lease is active. |
stale_running | A previous running lease expired and can be reclaimed. |
disabled | Global maintenance is disabled or the job interval is 0. |
never_run | The job is registered but has no completed run yet. |