Connectivity Monitors
Connectivity monitors run health probes concurrently with a flow's steps - optionally in the background, so the flow keeps working while the probes watch. A monitor targets:
- one or more device roles from the flow, and/or
- one or more explicit IPv4/IPv6 addresses
Every measurement is stored as a time series, surfaced in the run's Monitors tab as a table plus graphs, and streamed live to the UI while the run is in progress.
0) Terms
- Monitor: a configured probe running on an interval.
- TargetSelector: how user selects targets (role/ip).
- TargetInstance: resolved concrete endpoint (device+address or ip).
- Sample: one measurement row (per tick, per target instance).
1) Flow step type
1.1 monitor.connectivity
Continuously probes targets in the background during a flow run. The step starts the monitor and returns immediately, so the flow keeps executing while the probes tick. Targets come from the step's target selectors (device roles and/or explicit IP addresses - see section 2); everything else is step config:
| field | type | default | notes |
|---|---|---|---|
check_type | enum | tcp_connect | icmp_ping, tcp_connect, dns_resolve, or http_health |
port | integer | 22 | port for the TCP-based checks; ignored by icmp_ping and dns_resolve |
interval_ms | integer | 5000 | tick interval in milliseconds |
schedule_mode | enum | count | count, duration, or until_join |
count | integer | 10 | ticks to run when schedule_mode: count |
duration_sec | integer | 60 | run time when schedule_mode: duration |
timeout_sec | integer | 10 | per-probe timeout |
url_path | string | "" | request path for http_health (defaults to /health when omitted) |
socks_proxy | string | "" | optional socks5[h]://[user:pass@]host:port proxy for the TCP-based checks (tcp_connect, http_health) |
The field names above are the step handler's own configuration keys, which the handler maps onto the platform's monitor API (for example duration_sec becomes the API's schedule.duration_s). The check types listed are those the bundled probe wheel registers; the platform's check vocabulary also defines tls_handshake and ssh_banner for probe plugins that implement them.
The monitor.start and monitor.stop handlers are internal lifecycle controls and are not offered in the step palette.
2) Targeting model
2.1 TargetSelector (input)
A monitor step can define multiple selectors. Each selector can expand to multiple concrete TargetInstances.
| field | type | required | notes |
|---|---|---|---|
| type | enum | yes | role or ip |
| role | string | if type=role | flow device role name |
| addresses | list of string | if type=ip | IPv4/IPv6 literals |
2.2 TargetInstance (resolved at runtime)
| field | type | notes |
|---|---|---|
| target_instance_id | string | unique stable key |
| selector_index | int | index in target_selectors[] |
| selector_type | string | role or ip |
| role | string? | if role-based |
| device_id | string? | if role-based |
| address | string | concrete IPv4/IPv6 |
2.3 Resolution rules
type=role: resolves to the devices bound to that role in the run, using each device's management address; devices without a management address are skippedtype=ip: creates one TargetInstance per address inaddresses- If no selector resolves to a usable address, the step fails.
3) Scheduling / lifetime (no fail/recover semantics)
Scheduling is part of the step config (section 1): interval_ms sets the tick interval, and schedule_mode (with count / duration_sec) decides when sampling stops.
Rules:
count: Runs exactly N ticks, then completes and routes token (standard step behavior).duration: Runs for the specified duration, then completes and routes token (standard step behavior).until_join: Fire-and-forget - stopped when termination edge target (join/terminal) fires. Does NOT route token; the monitor is cancelled when the linked join completes.
4) Check execution contract
4.1 Common configuration
Every check type shares the step config described in section 1: interval_ms, schedule_mode (with count / duration_sec), and a per-probe timeout_sec. port applies to tcp_connect and http_health, url_path to http_health, and socks_proxy to the TCP-based checks (tcp_connect, http_health).
4.2 Sample (stored + streamed) — one per tick per TargetInstance
| field | type | notes |
|---|---|---|
| ts | RFC3339 string | sample timestamp |
| run_id | string | correlation |
| step_id | string? | step that started monitor (optional) |
| monitor_id | string | |
| check_id | string | |
| target_instance_id | string | |
| selector_type | string | role/ip |
| role | string? | |
| device_id | string? | |
| address | string | resolved IPv4/v6 |
| ok | bool | probe succeeded |
| metrics_json | map | keep numeric keys stable for graphs |
| error_kind | string? | timeout, conn_refused, dns_nxdomain, ... |
| error_detail | string? | short message |
| raw_text | string? | optional raw probe output |
5) Storage model (table + graph)
Persist every Sample row as timeseries.
5.1 The monitor_samples table
| column | type | notes |
|---|---|---|
| id | uuid | pk |
| run_id | uuid | |
| step_id | string? | index (optional) |
| monitor_id | string | index |
| check_id | string | |
| ts | datetime | |
| target_instance_id | string | index |
| selector_index | int | position in target_selectors[] |
| selector_type | string | |
| role | string? | |
| device_id | uuid? | index |
| address | string | index |
| ok | bool | |
| metrics_json | json | flat-ish map |
| error_kind | string? | |
| error_detail | string? | |
| raw_text | text? |
The indexed columns are seq, step_id, monitor_id, target_instance_id, device_id, and address.
5.2 Database Migration
The monitor_samples table is created by the standard database migrations; apply with task db:migrate.
6) API + streaming
6.1 REST endpoints
GET /runs/{run_id}/monitors- list a run's monitors (status, config, resolved targets, tick/sample counts)
GET /runs/{run_id}/monitors/{monitor_id}- fetch a single monitor
GET /runs/{run_id}/monitors/{monitor_id}/samples?limit=...&offset=...&since=...&target_instance_id=...- fetch samples for table/graphs, newest first;
sincefilters by timestamp,target_instance_idnarrows to one target
- fetch samples for table/graphs, newest first;
6.2 Live streaming
Server-sent events (SSE):
GET /runs/{run_id}/monitors/stream
Event: sample Payload: the stored Sample row (plus seq for client-side ordering)
When a monitor reaches a terminal state, the stream also emits a monitor_status event with the monitor's final status and total sample count.
Event: stream_error Payload: message, plus detail naming the exception class (for example OperationalError). The stream recovers on its own and keeps polling, so this is advisory rather than terminal. detail is deliberately the class name and not the driver's message, which would carry the database host and user; the full text goes to the server log.
UI behavior:
- Monitor list: every monitor in the run with its status; selecting one shows its detail
- Graphs: per-target charts of the numeric metrics over time, with toggleable series and summary statistics
- Table: samples grouped per target, with each sample's ok state, metrics, and any error
- Auto-refresh: while monitors are running, the Monitors tab falls back to polling whenever the live stream is not connected
7) Supported checks
| check_type | Name | Check-specific config | Metrics |
|---|---|---|---|
icmp_ping | ICMP Echo | — | rtt_ms, jitter_ms, packet_loss_pct |
tcp_connect | TCP Connect | port, socks_proxy | connect_ms |
dns_resolve | DNS Resolve | — | resolve_ms, addresses |
http_health | HTTP(S) Health | port, url_path, socks_proxy | http_status, connect_ms |
8) Flow YAML examples
8.1 Background ping monitor until join (role + explicit IP)
- id: mon_ping_core
type: step
name: "Ping core devices"
tags: {kind: CHECK}
handler: monitor.connectivity
target_selectors:
- type: role
role: upgraded_device
- type: ip
addresses: ["2001:4860:4860::8888"]
params:
check_type: icmp_ping
interval_ms: 2000
schedule_mode: until_join
timeout_sec: 28.2 TCP connect for duration (all upgraded devices)
- id: mon_tcp_ssh
type: step
name: "Watch SSH reachability"
tags: {kind: CHECK}
handler: monitor.connectivity
target_selectors:
- type: role
role: upgraded_device
params:
check_type: tcp_connect
port: 22
interval_ms: 1000
schedule_mode: duration
duration_sec: 308.3 Stopping a monitor
An until_join monitor is stopped by the flow engine: wire a termination edge from the monitor step to a join node, and the monitor is cancelled when that join completes. count and duration monitors complete on their own after their configured ticks or run time.
9) Runtime behavior
Targets are resolved once when the monitor starts. On every tick the monitor probes each target and stores one sample per target, publishing it to the live stream as soon as it is stored. The UI derives up/down status from the most recent samples.
The monitor's own run events (its egress report and its completion line) carry the phase tag of the step that started it, and none when that step has no such tag - phase is an ordinary, optional tag on a step.