Skip to content

Connectivity Monitors ​

Connectivity monitors run health probes concurrently with a flow's steps - optionally in the background, so the flow keeps working while the probes watch. A monitor targets:

  • one or more device roles from the flow, and/or
  • one or more explicit IPv4/IPv6 addresses

Every measurement is stored as a time series, surfaced in the run's Monitors tab as a table plus graphs, and streamed live to the UI while the run is in progress.


0) Terms ​

  • Monitor: a configured probe running on an interval.
  • TargetSelector: how user selects targets (role/ip).
  • TargetInstance: resolved concrete endpoint (device+address or ip).
  • Sample: one measurement row (per tick, per target instance).

1) Flow step type ​

1.1 monitor.connectivity ​

Continuously probes targets in the background during a flow run. The step starts the monitor and returns immediately, so the flow keeps executing while the probes tick. Targets come from the step's target selectors (device roles and/or explicit IP addresses - see section 2); everything else is step config:

fieldtypedefaultnotes
check_typeenumtcp_connecticmp_ping, tcp_connect, dns_resolve, or http_health
portinteger22port for the TCP-based checks; ignored by icmp_ping and dns_resolve
interval_msinteger5000tick interval in milliseconds
schedule_modeenumcountcount, duration, or until_join
countinteger10ticks to run when schedule_mode: count
duration_secinteger60run time when schedule_mode: duration
timeout_secinteger10per-probe timeout
url_pathstring""request path for http_health (defaults to /health when omitted)
socks_proxystring""optional socks5[h]://[user:pass@]host:port proxy for the TCP-based checks (tcp_connect, http_health)

The field names above are the step handler's own configuration keys, which the handler maps onto the platform's monitor API (for example duration_sec becomes the API's schedule.duration_s). The check types listed are those the bundled probe wheel registers; the platform's check vocabulary also defines tls_handshake and ssh_banner for probe plugins that implement them.

The monitor.start and monitor.stop handlers are internal lifecycle controls and are not offered in the step palette.


2) Targeting model ​

2.1 TargetSelector (input) ​

A monitor step can define multiple selectors. Each selector can expand to multiple concrete TargetInstances.

fieldtyperequirednotes
typeenumyesrole or ip
rolestringif type=roleflow device role name
addresseslist of stringif type=ipIPv4/IPv6 literals

2.2 TargetInstance (resolved at runtime) ​

fieldtypenotes
target_instance_idstringunique stable key
selector_indexintindex in target_selectors[]
selector_typestringrole or ip
rolestring?if role-based
device_idstring?if role-based
addressstringconcrete IPv4/IPv6

2.3 Resolution rules ​

  • type=role: resolves to the devices bound to that role in the run, using each device's management address; devices without a management address are skipped
  • type=ip: creates one TargetInstance per address in addresses
  • If no selector resolves to a usable address, the step fails.

3) Scheduling / lifetime (no fail/recover semantics) ​

Scheduling is part of the step config (section 1): interval_ms sets the tick interval, and schedule_mode (with count / duration_sec) decides when sampling stops.

Rules:

  • count: Runs exactly N ticks, then completes and routes token (standard step behavior).
  • duration: Runs for the specified duration, then completes and routes token (standard step behavior).
  • until_join: Fire-and-forget - stopped when termination edge target (join/terminal) fires. Does NOT route token; the monitor is cancelled when the linked join completes.

4) Check execution contract ​

4.1 Common configuration ​

Every check type shares the step config described in section 1: interval_ms, schedule_mode (with count / duration_sec), and a per-probe timeout_sec. port applies to tcp_connect and http_health, url_path to http_health, and socks_proxy to the TCP-based checks (tcp_connect, http_health).

4.2 Sample (stored + streamed) — one per tick per TargetInstance ​

fieldtypenotes
tsRFC3339 stringsample timestamp
run_idstringcorrelation
step_idstring?step that started monitor (optional)
monitor_idstring
check_idstring
target_instance_idstring
selector_typestringrole/ip
rolestring?
device_idstring?
addressstringresolved IPv4/v6
okboolprobe succeeded
metrics_jsonmapkeep numeric keys stable for graphs
error_kindstring?timeout, conn_refused, dns_nxdomain, ...
error_detailstring?short message
raw_textstring?optional raw probe output

5) Storage model (table + graph) ​

Persist every Sample row as timeseries.

5.1 The monitor_samples table ​

columntypenotes
iduuidpk
run_iduuid
step_idstring?index (optional)
monitor_idstringindex
check_idstring
tsdatetime
target_instance_idstringindex
selector_indexintposition in target_selectors[]
selector_typestring
rolestring?
device_iduuid?index
addressstringindex
okbool
metrics_jsonjsonflat-ish map
error_kindstring?
error_detailstring?
raw_texttext?

The indexed columns are seq, step_id, monitor_id, target_instance_id, device_id, and address.

5.2 Database Migration ​

The monitor_samples table is created by the standard database migrations; apply with task db:migrate.


6) API + streaming ​

6.1 REST endpoints ​

  • GET /runs/{run_id}/monitors
    • list a run's monitors (status, config, resolved targets, tick/sample counts)
  • GET /runs/{run_id}/monitors/{monitor_id}
    • fetch a single monitor
  • GET /runs/{run_id}/monitors/{monitor_id}/samples?limit=...&offset=...&since=...&target_instance_id=...
    • fetch samples for table/graphs, newest first; since filters by timestamp, target_instance_id narrows to one target

6.2 Live streaming ​

Server-sent events (SSE):

  • GET /runs/{run_id}/monitors/stream

Event: sample Payload: the stored Sample row (plus seq for client-side ordering)

When a monitor reaches a terminal state, the stream also emits a monitor_status event with the monitor's final status and total sample count.

Event: stream_error Payload: message, plus detail naming the exception class (for example OperationalError). The stream recovers on its own and keeps polling, so this is advisory rather than terminal. detail is deliberately the class name and not the driver's message, which would carry the database host and user; the full text goes to the server log.

UI behavior:

  • Monitor list: every monitor in the run with its status; selecting one shows its detail
  • Graphs: per-target charts of the numeric metrics over time, with toggleable series and summary statistics
  • Table: samples grouped per target, with each sample's ok state, metrics, and any error
  • Auto-refresh: while monitors are running, the Monitors tab falls back to polling whenever the live stream is not connected

7) Supported checks ​

check_typeNameCheck-specific configMetrics
icmp_pingICMP Echo—rtt_ms, jitter_ms, packet_loss_pct
tcp_connectTCP Connectport, socks_proxyconnect_ms
dns_resolveDNS Resolve—resolve_ms, addresses
http_healthHTTP(S) Healthport, url_path, socks_proxyhttp_status, connect_ms

8) Flow YAML examples ​

8.1 Background ping monitor until join (role + explicit IP) ​

yaml
- id: mon_ping_core
  type: step
  name: "Ping core devices"
  tags: {kind: CHECK}
  handler: monitor.connectivity
  target_selectors:
    - type: role
      role: upgraded_device
    - type: ip
      addresses: ["2001:4860:4860::8888"]
  params:
    check_type: icmp_ping
    interval_ms: 2000
    schedule_mode: until_join
    timeout_sec: 2

8.2 TCP connect for duration (all upgraded devices) ​

yaml
- id: mon_tcp_ssh
  type: step
  name: "Watch SSH reachability"
  tags: {kind: CHECK}
  handler: monitor.connectivity
  target_selectors:
    - type: role
      role: upgraded_device
  params:
    check_type: tcp_connect
    port: 22
    interval_ms: 1000
    schedule_mode: duration
    duration_sec: 30

8.3 Stopping a monitor ​

An until_join monitor is stopped by the flow engine: wire a termination edge from the monitor step to a join node, and the monitor is cancelled when that join completes. count and duration monitors complete on their own after their configured ticks or run time.


9) Runtime behavior ​

Targets are resolved once when the monitor starts. On every tick the monitor probes each target and stores one sample per target, publishing it to the live stream as soon as it is stored. The UI derives up/down status from the most recent samples.

The monitor's own run events (its egress report and its completion line) carry the phase tag of the step that started it, and none when that step has no such tag - phase is an ordinary, optional tag on a step.

Released as open source under the AGPL-3.0-or-later license. Development is sponsored by Rexonix s.r.o.. Contact — [email protected].