Skip to content

Repository files navigation

Agent Smith

Test Coverage Lint CodeQL Release

Rewst's lean, open-source command executor that fits right into your Rewst workflows. See community corner for more details.

Installation

Agent Smith runs as a system service on Windows, Linux, and macOS. Installation involves configuring the agent with your organization credentials and starting the service.

Prerequisites

  • A Rewst organization ID
  • Configuration URL and secret from your Rewst platform
  • Administrative/root privileges for service installation

Basic Installation

  1. Download the appropriate binary for your platform from the releases page
  2. Configure the agent with your organization credentials:

Windows:

rewst_agent_config.win.exe --org-id YOUR_ORG_ID --config-url CONFIG_URL --config-secret CONFIG_SECRET

Linux/macOS:

./rewst_agent_config.linux.bin --org-id YOUR_ORG_ID --config-url CONFIG_URL --config-secret CONFIG_SECRET
# or
./rewst_agent_config.mac-os.bin --org-id YOUR_ORG_ID --config-url CONFIG_URL --config-secret CONFIG_SECRET

Configuration Options

  • --logging-level: Set logging verbosity (info, warn, error, debug)
  • --syslog: Write logs to system log instead of file (Linux/macOS)
  • --disable-agent-postback: Disable agent postback
  • --no-auto-updates: Disable auto updates
  • --mqtt-qos: MQTT subscription QoS level (0 = at-most-once, 1 = at-least-once). Defaults to 1 when omitted. Azure IoT Hub does not support QoS 2.

Example with optional parameters:

./rewst_agent_config --org-id YOUR_ORG_ID --config-url CONFIG_URL --config-secret CONFIG_SECRET --logging-level info --syslog --disable-agent-postback --no-auto-updates --mqtt-qos 1

Update

Once installed, the agent can be updated and configured using the config executable. The optional parameters are also available.

./rewst_agent_config --org-id YOUR_ORG_ID --update --logging-level info --syslog --disable-agent-postback --no-auto-updates --mqtt-qos 1

Service Mode

Once configured, the agent can run in service mode using the generated configuration:

./rewst_agent_config --org-id YOUR_ORG_ID --config-file /path/to/config.json --log-file /path/to/agent.log

Diagnostic Mode

The diagnostic mode provides an interactive menu to validate an installed agent without needing to inspect log files or know platform-specific service commands. It is useful for troubleshooting connectivity issues, verifying permissions, and confirming the agent is healthy.

Usage

Run without an org ID to scan all installed agents:

Windows:

rewst_agent_config.win.exe --diagnostic

Linux:

sudo ./rewst_agent_config.linux.bin --diagnostic

macOS:

sudo ./rewst_agent_config.mac-os.bin --diagnostic

To target a specific organization directly:

./rewst_agent_config --org-id YOUR_ORG_ID --diagnostic

Interactive menu

Once launched, the menu guides you through the following checks:

[1] Scan installed agents and check status
[2] Test command execution
[3] Test MQTT/WebSocket connectivity
[4] Test temp directory write access
[5] View live log data
[6] Run all checks
[0] Exit
Option What it checks
1 Lists all installed agents with running/stopped status and config details (device ID, IoT Hub, engine host, log level)
2 Runs a test command using the platform shell (PowerShell on Windows, Bash on Linux/macOS) and confirms execution succeeds
3 Attempts TLS connections to the agent's IoT Hub on port 8883 (MQTT) and port 443 (WebSocket). Prints troubleshooting tips if both fail
4 Creates a test file in the scripts temp directory and reads it back to confirm write access
5 Opens the agent log file and tails it in real time. Press Ctrl+C to stop
6 Runs checks 1–4 in sequence

Example output

  ╔══════════════════════════════════════════════════╗
  ║         Agent Smith Diagnostic Mode              ║
  ║         Version: v1.1.0                          ║
  ║         Platform: windows/amd64                  ║
  ╚══════════════════════════════════════════════════╝

  ── Installed Agents ──

    [PASS] a1b2c3d4-... - RUNNING (RewstRemoteAgent_a1b2c3d4-...)
      Device ID:    device-xyz
      IoT Hub:      abc123.azure-devices.net
      Engine Host:  engine.rewst.io
      Log Level:    info
      Syslog:       false
      Auto-Updates: true
      MQTT QoS:     1

  ── MQTT/WebSocket Connectivity ──

    Host: abc123.azure-devices.net
    Testing MQTT (TLS port 8883)... OK
    [PASS] MQTT TLS connection to abc123.azure-devices.net:8883
    Testing WebSocket (port 443)... OK
    [PASS] WebSocket connection to abc123.azure-devices.net:443

Uninstallation

To remove Agent Smith from your system:

# Replace with your organization ID
./rewst_agent_config --org-id YOUR_ORG_ID --uninstall

This will stop the service, remove configuration files, and clean up system service registrations. If a directory cannot be removed — a file held open by an AV scanner, for example — the remaining directories are still removed and the paths that survived are named in the log (see "Completing Uninstall When a Directory Cannot Be Removed" below).

Features

  • Cross-platform: Runs on Windows, Linux, and macOS
  • Secure: Uses Azure IoT Hub MQTT for encrypted communication
  • Extensible: Plugin system for custom notifications and integrations
  • Reliable: Automatic reconnection and error handling
  • Lightweight: Minimal resource footprint

How It Works

  1. Agent connects to your Rewst organization via Azure IoT Hub MQTT
  2. Receives command execution requests from Rewst workflows
  3. Executes commands using PowerShell (Windows) or Bash (Unix/Linux/macOS)
  4. Returns results back to the Rewst platform
  5. Supports system information collection and custom plugins

Message Delivery Guarantee

Incoming commands are handed to a buffered queue drained by a pool of command-execution workers. Two mechanisms together make sure a command the broker has handed over is never silently lost.

Back-pressure. When the queue fills (a burst of commands, or execution slow enough to keep every worker busy), the subscribe callback blocks instead of dropping: it waits until a worker frees a slot. The agent acknowledges a message only after the callback returns, so a saturated agent stops acknowledging and the broker holds later commands rather than the agent discarding them.

A durable command journal. Before a command is acknowledged to the broker it is written to <data directory>/command_journal, one file per command. The acknowledgement therefore means "durably accepted", not "buffered in memory". If the agent process dies — a crash, an OOM kill, a force-stop, host power loss — with commands queued or executing, their journal entries survive and the next start replays them:

  • a command that had not started executes exactly as it would have;
  • a command that had started is reported back to the engine as interrupted ("interrupted": true in the result) rather than run again — a partial first run of a non-idempotent script followed by a second full run is the failure mode at-least-once delivery is most often criticised for, so the receiving workflow gets the facts and decides whether to re-issue;
  • a command older than one hour (the broker's own default message TTL) is reported as expired and discarded, so a device that was off for a day does not wake up and run a day-old script.

A completed command leaves a tombstone for an hour, so a redelivery of the same message — the broker resends an unacknowledged QoS 1 message on reconnect, and the acknowledgement for a message the agent has since finished can be lost in transit — is acknowledged and not executed a second time. The de-duplication key is the message's post_id (what the engine correlates results by), or a digest of the payload for a message without one.

Why the acknowledgement cannot simply wait for execution

The obvious design — acknowledge only once the command has run and posted back — does not work against Azure IoT Hub. From Microsoft's documentation: a cloud-to-device message the device has not acknowledged returns to the queue "after a visibility timeout (or lock timeout). The length of this timeout is one minute and can't be changed", and "If the lock expires, the message returns to Enqueued but the delivery count does not increment" — so it is redelivered, and "messages continue cycling … indefinitely until they expire." A command that runs longer than a minute (the default timeout is thirty) would be redelivered every minute for as long as it ran. Microsoft's recommendation for long-running work is exactly what the journal does: "Complete the cloud-to-device message after the device persists the task description in local storage."

Where a command can still be lost

The journal is best effort in the same way the postback spool, the syslog forwarder and the log rotator are: a disk problem must not stop commands from running. If the journal write fails (disk full, permissions), the command is accepted in memory and acknowledged anyway — the behaviour before the journal existed — and the failure is logged once at Error level with an AgentCommandJournalDegraded plugin notification, then once more when it recovers. A crash in that degraded window loses the queued commands, as it always did. The journal is bounded to 1000 pending entries; past that it behaves as if the write failed.

The one case in which the agent leaves a message unacknowledged on purpose is a command that arrives while a connection cycle is tearing down and could not be journaled: there is nowhere to put it, so it is left for the broker to redeliver after its lock expires — which is what the corresponding Error log now says. It is counted (AgentMessageDropped notification) so it is visible in monitoring. A journaled command arriving during teardown is acknowledged and replayed on the next connection. The set that is replayed is captured before the cycle connects and excludes every entry this process has accepted itself, so a command the cycle accepts - or one the previous cycle's workers are still finishing across a reconnect - is never mistaken for a leftover and reported as interrupted or run a second time. Replay is for what a previous process left behind.

Tuning queue capacity and concurrency

Two optional fields in the device configuration file let high-volume deployments tune the queue without code changes:

Config key Default Description
worker_count 10 Number of concurrent command-execution workers draining the queue.
message_queue_size 100 Capacity of the buffered inbound message queue before back-pressure begins.

Both fall back to their defaults when omitted or set to a non-positive value. Raising message_queue_size absorbs larger bursts before back-pressure begins; raising worker_count widens execution parallelism. Example snippet:

{
  "worker_count": 25,
  "message_queue_size": 500
}

Bounding per-command execution time

Every received command runs under an execution deadline, on by default, so a script that hangs (infinite loop, blocked on a prompt/stdin, stuck network call) cannot occupy its worker indefinitely. Without a bound, once as many hung commands accumulate as there are workers the whole pool is exhausted and no further commands run until the agent reconnects.

Set command_timeout_seconds to override how long any single command may run:

Config key Default Description
command_timeout_seconds 1800 (30 minutes) Maximum seconds a single command may run before it is killed and its worker released.

Each command runs under a derived context with that deadline; if it is exceeded the command's process group is killed, the worker is freed, and the result posted back is flagged with "timed_out": true (distinct from a normal non-zero exit) while the event is logged at Error level with the post_id. It falls back to the default when omitted or set to a non-positive value, so the bound can never be disabled by configuration; raise it for workflows with legitimately long-running commands. Example snippet:

{
  "command_timeout_seconds": 300
}

Killing the Full Process Tree on Windows

Killing a command's process on timeout or cancellation must also kill whatever that command spawned — a Start-Process call, an installer, a stuck helper — or the child is reparented and keeps running on the endpoint after the worker is released, leaking a process per hang. On Unix this is handled by placing the shell in its own process group and killing the group (internal/interpreter/proc_unix.go); Windows has no process-group equivalent, so the same guarantee is provided with a job object (internal/interpreter/proc_windows.go): the shell process is assigned to a job created with JOB_OBJECT_LIMIT_KILL_ON_JOB_CLOSE, and closing the job's handle — which cancellation now does instead of a plain Process.Kill — terminates every process still assigned to it, direct or descendant. The job handle is released once the command finishes either way, so a command that completes normally leaks no handle.

Bounding per-command output size

A command's stdout and stderr are each captured through a bounded writer, so a script that writes a very large volume of output cannot exhaust memory on the endpoint and get the agent OOM-killed. Once a stream reaches its ceiling, further bytes from that stream are discarded instead of accumulated, which keeps the memory one command's output costs the agent a small constant multiple of the ceiling no matter how much the script writes.

Config key Default Description
max_output_bytes 10485760 (10 MiB) Maximum bytes of command output kept, applied independently to stdout and stderr.

The default is far above any legitimate observed command result, so existing workflows are unaffected. It falls back to the default when omitted or set to a non-positive value, so the bound can never be disabled by configuration. The value can also be set at install time with --max-output-bytes <N>, or changed later with --update --max-output-bytes <N>.

Truncation is deliberately non-fatal:

  • The command is not killed for being verbose. It runs to completion (or to its command_timeout_seconds deadline) with the excess output dropped, so a script whose real work succeeds still reports success.
  • The output produced before the ceiling was reached is still delivered — a truncated result is never turned into an empty or error-only one.
  • The result says so explicitly, so a workflow can tell a truncated result from a complete one instead of silently trusting a partial one:
{
  "output": "...the first max_output_bytes of stdout...",
  "error": "...the first max_output_bytes of stderr...",
  "truncated": true,
  "output_bytes_produced": 2097152000,
  "output_bytes_kept": 20971520
}

output_bytes_produced and output_bytes_kept are totals across both streams. All three keys are omitted when the output was captured in full, so a complete result is byte-identical to what previous releases posted back. A command that was both verbose and hung carries "truncated": true alongside "timed_out": true.

Each truncation is logged once per command at Warn level with the message_id, the ceiling in effect, and both byte counts — never once per write.

Bounding the agent log file on disk

The agent's own log file (rewst_agent.log in the data directory) is written through a size-bounded rotating writer, so a long-running installation can no longer fill the endpoint's system volume — the volume that also holds the data directory, the postback spool and the operating system. Every other file the agent writes there was already bounded (the spool by entries and age, stale scripts and installers by the startup sweeps, command output by max_output_bytes); the log was the last one that was not, and the only one whose growth was proportional to how long the agent had been doing its job correctly.

When a write would carry the active file past log_max_bytes, the file is rotated first: rewst_agent.log becomes rewst_agent.log.1, the previous .1 becomes .2, and so on up to .log_max_files, which is discarded. Rotation happens in-process on the write that crosses the threshold, not only at startup, so an agent that stays up for months rotates without a restart. The worst-case footprint is (log_max_files + 1) × log_max_bytes plus one line's overshoot — 60 MiB at the defaults.

Config key Default Description
log_max_bytes 10485760 (10 MiB) Size at which the active log file is rotated.
log_max_files 5 Rotated copies kept (rewst_agent.log.1 … .5); the oldest is discarded on each rotation.

Both fall back to their defaults when omitted or set to a non-positive value, so existing deployments need no config change. They can be set at install time with --log-max-bytes <N> / --log-max-files <N>, or changed later with --update. Lowering log_max_files also removes any copies numbered above the new value on the next rotation. A pre-existing oversized log left by an agent that never rotated is rotated on the first write that crosses the threshold after upgrade, rather than left in place.

Rotated copies are created by renaming the active file, so they keep its permissions and stay in the (owner-only) data directory. Rotation never loses a line: it happens before the write, under the same lock, and the line then lands in whichever file is open afterwards. It also prefers a line boundary: a writer that delivers one line in several chunks (a notification plugin's stderr is copied through in whatever pieces the pipe returns) never has that line split across two files. A crossing that arrives mid-line is deferred until the line completes; a writer that never terminates a line is rotated regardless once the file reaches twice the ceiling, so the bound still holds.

Rotation is best effort and cannot take the service down. The writer owns the file handle and closes it before renaming, which is what lets the rename succeed on Windows — where a file cannot be renamed while any handle lacking FILE_SHARE_DELETE is open on it. The agent's own handle is the common one. The one remaining handle that can still block a rename there is the detached --update helper's inherited stdout, for the seconds it takes to start. In that case (or on a full disk) the current file is reopened and appending continues, one [WARN] line records the failure, rotation is not retried for a minute so a lasting block cannot flood the file, and one line records the eventual recovery. This mirrors the syslog forwarder's "counted, not fatal" posture, and the two writers share one formatter for those lines so they parse identically. Syslog forwarding itself is unchanged: the on-disk write is always performed and its result is what Write returns.

A rename that succeeds but whose reopen then fails (out of descriptors, a full disk refusing a new inode) is a different case and is reported as one: the log is momentarily closed, every write until it reopens is counted, and when it does reopen a single [WARN] line says how many writes were lost. That is never reported as a rotation failure and never produces a "recovered" line, because rotation did not fail.

Diagnostic mode's live log viewer (menu option 5) follows the active file across a rotation on every platform. At EOF it compares the identity of the file it holds open with the file at the log path; when they differ it first drains whatever the agent wrote to the old file after the viewer's last read, then reopens the new one and prints a marker line, so no line is skipped. On Windows the viewer opens the log with FILE_SHARE_DELETE (which Go's os.Open does not request) precisely so that its own handle can never be the thing blocking the agent's rotation for the length of a support session.

Hardening the Command Scripts Directory

Each received command is written to a temporary script file before it is executed. On Linux and macOS, that file used to be written to a subdirectory of the shared, world-writable system temp directory (os.TempDir()/rewst_remote_agent/scripts/<orgId>). Because directory creation is a no-op against a directory that already exists, any local unprivileged user could pre-create that directory — or simply wait for a tmpfs /tmp to reset on reboot — with permissive ownership and mode before the agent ever ran, and the agent would silently reuse it exactly as found. Combined with the write, close, and path-based re-open the executor used to hand the script to the shell, a local user able to write into that directory could in principle swap the script's contents in the gap between the two opens and have it run with the agent's privilege (root/SYSTEM).

Three changes address this, matching the precedent already set for auto-update installers (see "Reclaiming Downloaded Installer Binaries") where they apply:

  • On Linux and macOS, scripts land in a directory the agent owns. The scripts directory is now <data directory>/scripts (/etc/rewst_remote_agent/<orgId>/scripts, /Library/Application Support/rewst_remote_agent/<orgId>/scripts) instead of the shared system temp directory — a location an unprivileged local user cannot pre-create in the first place.

    Windows deliberately keeps its historical location, C:\RewstRemoteAgent\scripts\<orgId> (the system drive root, not under ProgramData), and does not follow this move. Some customers have their endpoint security software configured to whitelist exactly this path so the dynamically-written PowerShell scripts the agent executes here are not flagged or blocked as they run; relocating it would silently break command execution on those endpoints. It also was not the vulnerability described above to begin with — that depended on a shared, world-writable temp directory, and an unprivileged user cannot create a new top-level directory at the Windows system drive root under the OS's default ACLs.

  • Its permissions are re-asserted on every command, not only when first created. EnsureSecureDir (internal/utils/filesystem_unix.go, filesystem_windows.go) runs before every command, not only the first, on all three platforms. On Linux and macOS it refuses to follow a symlink or non-directory planted at that path, reclaims ownership via chown if the directory belongs to another uid, and re-applies mode 0700 if it does not already have it — failing loud rather than proceeding if either correction itself fails. A directory that already has the right owner and mode is left untouched. On Windows, where there are no POSIX ownership/mode bits to re-assert, it only refuses a symlink or non-directory planted at that path.

  • The write-then-exec window is verified, not just trusted. After the script file is written and closed, its contents are read back and compared byte-for-byte against what was just written; a mismatch aborts the command instead of executing whatever is now on disk. This is defense in depth on top of the directory hardening above — on Linux/macOS, with the directory locked to 0700 and owned by the agent's own account, no other local user can write into it at all, so the window this check exists for should never legitimately fire there.

Command Result Delivery

After a command runs, the agent posts its result back to the Rewst engine with retry and exponential backoff. If every in-line attempt fails (network error or 5xx across the whole retry budget), the result is not dropped:

  • The failure is surfaced beyond the log with a best-effort AgentPostbackFailed:<post_id> plugin notification so monitoring can observe it.
  • The result is written to a bounded on-disk spool (under the agent's data directory) and re-attempted on the next successful connection cycle. A transient engine outage therefore recovers automatically once connectivity returns, instead of losing the result.

The spool is bounded by count and age (the oldest/expired entries are evicted) so it cannot grow without limit, and the flush is bound to the connection cycle so it never blocks shutdown.

One undeliverable result never blocks the rest

A flush distinguishes two failures that look alike from a distance:

  • The engine is unreachable — a transport error, or a connection that broke while the response was being read. Nothing behind this entry could be delivered either, so the flush stops and every remaining entry is retried on a later cycle. This costs the entry nothing: an outage is not the entry's fault and consumes none of its attempt budget.
  • The engine rejected this entry — it answered with a 5xx, or with a body that could not be parsed. The engine is plainly up, so the flush passes over this entry and keeps going; every entry behind it still gets its attempt.

The second case used to be read as the first. One result the engine consistently rejected would pin the queue: the flush restarted from it every cycle, retried it forever with no attempt bound, and the healthy results behind it were never attempted until the age check discarded them — logged as expired, which reads like stale-data cleanup rather than the delivery failure it was.

An entry that is rejected carries a persisted attempt counter and last error, so its budget survives an agent restart. After 5 counted rejections the entry is abandoned: removed with the distinct reason attempts_exhausted, counted separately, and surfaced with a best-effort AgentPostbackAbandoned:<post_id> plugin notification — never silently reported as stale.

Rejections are counted at most once every 10 minutes. An engine that is failing wholesale answers 5xx for every entry, which at the HTTP layer is indistinguishable from it rejecting each one specifically; without that spacing, a flapping connection could spend an entry's whole budget in minutes and abandon a result the engine would have accepted on recovery. Spacing bounds the budget in time rather than in reconnects, so a result survives at least 40 minutes of a wholesale outage however often the agent reconnects. It never delays the pass-over itself — a rejected entry is skipped on every flush regardless; only the counting is spaced.

Drop reasons are counted separately (expired, capacity, attempts_exhausted, corrupt) so a spool shedding entries under pressure is distinguishable in diagnostics from one abandoning a result the engine refuses. Spool entry files written by an older agent have no attempt counter; they are read as never-attempted and delivered normally, not discarded.

The in-line retry budget is tunable per deployment:

Config key Default Description
postback_max_attempts 3 Total postback attempts (including the first try) before the result is spooled.
postback_base_retry_backoff_seconds 1 Base delay for exponential backoff between attempts (base * 2^(n-2)).

The per-attempt backoff is capped at 64s (with up to ±25% jitter) regardless of how high postback_max_attempts is raised, mirroring the reconnect backoff. This keeps a wide retry window from overflowing into a busy-loop or blocking a worker for days, so raising postback_max_attempts only ever adds more bounded retry slots.

Both fall back to their defaults when omitted or set to a non-positive value, so existing configurations are unaffected. Example snippet:

{
  "postback_max_attempts": 6,
  "postback_base_retry_backoff_seconds": 2
}

Staying Connected (SAS token renewal)

The agent authenticates to Azure IoT Hub with a short-lived SAS token, and the hub forcibly disconnects a client the instant its token expires. To avoid a forced disconnect on a fixed cadence, the agent mints a long-lived token and proactively reconnects with a fresh one a safety margin before expiry. The old connection is therefore never torn down by an expired token: the reconnect is a routine, Info-level Renewing SAS token before expiry log line rather than an Error-level Connection lost. An Error Connection lost now reflects a genuine fault (network drop, broker-side disconnect), and reconnect behavior for those real losses is unchanged.

Neither a renewal nor a lost connection interrupts a command that is running. The command workers belong to the service, not to the connection cycle: when a cycle ends they stop taking new work, finish (or time out on their own per-command deadline) whatever they are executing, post the result back over HTTP - which needs no broker connection - and exit once the queue they were fed from is drained. Until sc-118039 every cycle end cancelled the workers, so a routine renewal killed any command spanning it and reported it as failed (or, since the command journal, as interrupted); with the default 24-hour token that was once a day on every device. Only a service stop cancels a running command, and a command cancelled that way is reported as interrupted by the next start. While an outgoing pool finishes its last commands the next cycle's pool is already receiving, so the live worker count can briefly reach twice worker_count, never more.

Config key Default Description
sas_token_lifetime_hours 24 Lifetime (in hours) of the Azure IoT Hub SAS token minted for each connection.

The renewal margin is 10% of the lifetime, floored at 1 minute and capped at 15 minutes, so the token is used for almost its entire lifetime yet always refreshed ahead of expiry. Falls back to the default when omitted or set to a non-positive value. Example snippet:

{
  "sas_token_lifetime_hours": 12
}

Bounded MQTT Operations

Every MQTT operation in the connection cycle waits with a deadline, and the subscribe wait is additionally interruptible by a service stop. Without that, a broker which holds the connection open and answers keepalive pings but stops responding to control packets — exactly what Azure IoT Hub does when it throttles a device, and what a stateful firewall, VPN idle-timeout, captive portal, or SD-WAN appliance that half-opens a connection produces — left the agent parked mid-operation forever. It looked healthy (connected, keepalive fine, no errors logged) while silently never subscribing, so every command sent to it was queued at the broker and never ran, and a stop, upgrade, or uninstall hung until the platform's stop deadline expired and force-killed the process.

With bounded waits, a broker in that state produces a clean, logged failure and a normal reconnect instead:

  • Subscribe — a SUBACK that does not arrive within the timeout is logged at Error as Failed to subscribe: timed out waiting for broker acknowledgement (with the timeout and elapsed wait), ends the cycle, and is retried on the reconnect backoff schedule. The backoff is deliberately not cleared, so a throttling broker gets progressively longer waits rather than being hammered. A stop arriving during the wait is honored immediately rather than after the timeout.
  • Connect — bounded as a backstop above paho's own connect timeout, and likewise stop-interruptible.
  • Unsubscribe (teardown) — bounded by a short fixed timeout so total teardown stays inside the Windows SCM stop window and systemd's TimeoutStopSec. Exceeding it logs a Warn and proceeds to disconnect; against a black-holed connection an unbounded wait would otherwise block until keepalive plus ping timeout elapsed, which alone exceeds the Windows default.
  • Device-twin reported properties — bounded so the informational version report cannot wedge the cycle before the agent has subscribed. Its failure stays non-fatal (Warn).

Because the service now always exits through its own teardown rather than being force-killed, deferred cleanup still runs: temp script files are removed, and the spooled-postback and plugin subprocesses are shut down rather than orphaned.

Config key Default Description
mqtt_connect_timeout_seconds 30 Bounds a single MQTT connect attempt.
mqtt_subscribe_timeout_seconds 30 How long to wait for the broker's SUBACK before treating the connection attempt as failed.

Both fall back to their defaults when omitted or set to a non-positive value, and can also be set at install or update time with --mqtt-connect-timeout-seconds and --mqtt-subscribe-timeout-seconds. The teardown-side timeouts (unsubscribe, device-twin publish) are fixed constants rather than per-device knobs, because a tunable value there could push total teardown past a platform stop deadline the operator cannot see. Example snippet:

{
  "mqtt_subscribe_timeout_seconds": 45
}

Bounded Windows Service Stop (Windows only)

Update, install and uninstall all stop the running service first, and on Windows that means asking the Service Control Manager to stop it and then waiting for the service to report Stopped. That wait used to be unbounded: it polled the service state every 250 ms forever. A service that never reaches Stopped — a wedged agent process, or Windows itself holding the service in StopPending behind a stuck operation — therefore hung the caller indefinitely. Because the auto-updater runs unattended, the visible symptom was a device that went offline during a routine update with no error logged anywhere, needing hands on the endpoint to recover; an interactive uninstall simply hung with no output.

The wait is now bounded:

  • The stop waits at most 5 minutes for the service to reach Stopped. The bound is deliberately generous — it exists to catch a wedge, not to race a normal shutdown, so a healthy agent draining long-running commands is never cut short.
  • A service observed in StopPending keeps being polled until the deadline, while a state the service cannot reach Stopped from fails immediately rather than burning the full deadline. Running and StartPending are treated as still-stopping, because the SCM can report the pre-stop state before the service thread picks the control up.
  • Overrunning the deadline logs at Error and names the service, the deadline and the last observed state — for example service rewst_agent_smith_<org> did not stop within 5m0s: last observed state StopPending — so the reason an update aborted is visible in the log.
  • Callers abort instead of proceeding as if the stop succeeded. Update does not overwrite the agent executable or the config file, since the old process may still hold the executable open, leaving the existing installation intact and the service recoverable. Uninstall does not delete the registration or any files out from under a live process, and logs plainly that nothing was removed. Install (--config) likewise aborts before deleting a pre-existing wedged service.

Recovery is the ordinary one: end the wedged agent process, start the service, and re-run the update or uninstall. Linux and macOS are unaffected — their service implementations do not use this polling loop.

Reporting Stop Progress to the SCM (Windows only)

The bound above is the caller's side of a stop: how long the updater waits for the service to report Stopped. The service's own side used to say nothing during that wait. windowsRunner.Execute published StartPending at entry and Running once the agent was up, but on receiving a stop control it only signalled its internal stop channel and then published Stopped after the agent had finished shutting down — so the service went from Running straight to Stopped with a silent gap in between.

Windows reads that silence as a hang. The Service Control Manager expects a stopping service to acknowledge the control with StopPending and then keep showing progress — an incremented CheckPoint within the WaitHint the service published — and a service that goes quiet is logged as not responding (events 7009 and 7043), with services.msc or net stop reporting failure. The gap is not short: the agent's shutdown drains in-flight command workers, each of which an operator can let run for as long as command_timeout_seconds (30 minutes by default), and then waits on plugin shutdown. A stop, restart, auto-update or uninstall issued while a long command was in flight therefore looked like a wedged service for its whole duration, even though the agent was following its own bounded shutdown sequence correctly — enough to trigger alerting or an external forced kill mid-update.

The service now reports its progress:

  • The stop control is acknowledged immediately, before the agent starts draining: StopPending with CheckPoint 1 and a WaitHint of 30 seconds.
  • While the shutdown runs, StopPending is republished every 10 seconds with an incremented CheckPoint, for as long as the shutdown takes. The WaitHint is three times the checkpoint interval, so a merely late checkpoint on a busy endpoint does not read as a wedge either.
  • Stopped is still published exactly once, after the agent's shutdown returns, and is always the last status sent. All of the service's status updates now go through one reporter that enforces the order the SCM expects: nothing is published after Stopped — so a checkpoint that wakes up late can neither contradict the terminal status nor block on a channel the SCM has stopped reading — and Running is dropped once a stop has been acknowledged, which matters for a stop issued while the agent is still starting up.

The caller-side wait is unchanged — this fixes what the service reports during that wait, not the wait itself — and a shutdown that genuinely wedges still fails the same way, on the same 5-minute bound, with the same error.

Bounded Service and Host-Info Shell-Outs (macOS, Linux, Windows)

The Windows service Stop() wait above bounds one path to a wedged OS-level tool. Several adjacent shell-outs used to have no bound at all: launchctl on macOS, systemctl on Linux, and sc query / dsregcmd / the WMI queries PowerShell issues on Windows host-info gathering. A D-Bus stall, launchd in a bad state, or a broken WMI repository — a known real-world failure mode, especially on domain controllers — blocked install, update, uninstall and config generation indefinitely, the same failure class already fixed for Stop(), just left open here.

Every one of these commands is now run with exec.CommandContext under a bounded, documented timeout:

  • macOS (internal/service/service_darwin.go): every launchctl invocation (load, start, stop, unload, print) is bounded by launchctlTimeout, 5 minutes — the same bound as the Windows service stop wait, generous enough to never cut off a legitimate stop.
  • Linux (internal/service/service_linux.go): every systemctl invocation (start, stop, enable, disable, is-active, is-enabled, daemon-reload) is bounded by systemctlTimeout, also 5 minutes.
  • Windows (internal/agent/host_windows.go): the WMI queries behind ADDomain/IsADDomainController, the sc query calls behind IsEntraConnectServer, and the dsregcmd /status call behind EntraDomain are each bounded by hostCommandTimeout, 30 seconds — these normally complete in well under a second, and the caller-supplied context passed in from config.go/update.go carries no deadline of its own, so the bound has to be enforced internally rather than relied on from the caller.

A command that exceeds its bound is killed and returns an error naming the command and its arguments (for example systemctl stop rewst_agent_smith_<org> timed out after 5m0s: ...), which callers treat the same as any other failure from that command — an install/update/uninstall aborts without writing or deleting anything, and host-info gathering logs the field as unavailable (NewHostInfo already warns per-field and continues) rather than blocking the rest of the flow.

IsEntraConnectServer is the one host-info field that shells out more than once: it probes four candidate service names in sequence. Each probe was originally given its own fresh hostCommandTimeout, which made the bound on the field four times the documented bound — two minutes against a wedged Service Control Manager — and, worse, a probe killed by its deadline was treated exactly like a probe that came back "service does not exist", so the check moved on to the next name and ultimately returned false with no error. A timeout was silently reported as "this host is not an Entra Connect server".

Both are fixed:

  • One shared budget for the whole check. IsEntraConnectServer derives a single hostCommandTimeout context up front and runs all four sc query calls under it, so the field is bounded by 30 seconds no matter how many names are probed.
  • A timeout aborts the check with an error. A failed probe with the shared budget expired (or the caller canceled) says nothing about whether the service exists, and every remaining name would fail the same way, so the check returns an error naming the service and the bound (sc query "ADSync" timed out after 30s: ...) instead of a made-up false. NewHostInfo logs it as an unavailable field and keeps the rest of the host info, the same as any other per-field failure.

This bound matters on the get_installation path in particular. That request is handled by a worker out of the small MQTT command-processing pool, and it is not covered by the per-command execution timeout described in "Bounding per-command execution time" — that bound applies to the interpreter's own command execution, not to host-info gathering. Before this fix a wedged SCM occupied the worker for as long as it stayed wedged; with enough concurrent get_installation requests that degrades or stalls command execution on the device, the same failure the per-command timeout was introduced to prevent.

All three timeouts are overridable via -ldflags for integration testing — service.launchctlTimeoutOverrideStr, service.systemctlTimeoutOverrideStr, and agent.hostCommandTimeoutOverrideStr — the same mechanism stopTimeoutOverrideStr uses for the Windows service stop wait, so a wedged command can be observed aborting in seconds rather than the production bound.

Bounded, Non-Fatal Syslog Forwarding (Linux, macOS)

With use_syslog enabled, every log line the agent writes is also forwarded to the system logger by shelling out to logger (internal/syslog/syslog_unix.go). That forward happens synchronously inside the writer hclog holds, so it runs on whatever goroutine is logging — the MQTT client, a command worker, the plugin supervisor. It used to be an unbounded exec.Command(...).Run(), and it used to return early on failure:

  • A hung logger froze the logging goroutine forever. A syslog daemon in a bad state, or logger blocked writing to /dev/log, had no timeout to recover from — the same failure class as the unbounded systemctl/launchctl shell-outs above, but reachable from every subsystem, since every subsystem logs.
  • A failed logger deleted the line from the on-disk log too. The write returned (0, err) before ever reaching the log file, so a transient hiccup (daemon restart, missing logger binary) punched a hole in the file operators read afterwards — precisely at the moments something is already going wrong and the on-disk log most needs to be complete.

Both are fixed, and the fix has three parts:

  • Each logger call is bounded. It runs under exec.CommandContext with a 5 second timeout (syslogCommandTimeout), enormously generous for a one-shot that normally finishes in milliseconds. As with the interpreter's per-command timeout, expiry kills the whole process group, not just logger itself: a surviving descendant inherits the output pipe and would keep cmd.Wait blocked, leaving the call hung despite the deadline. The group kill reaches that descendant only while it stays in the group, so the command also carries a 1 second WaitDelay: a descendant that calls setsid, or is re-homed by an init system, escapes the group still holding the pipe, and without that backstop Wait would block forever — the same permanent freeze, reached through the pipe rather than through the process.
  • The on-disk write always happens. The syslog forward is best effort and its outcome never propagates: Write returns the log file write's (n, err), so a genuine file error still surfaces to hclog while a syslog failure never costs the log file a line. This matches how the Windows event-log writer has always behaved (event-log errors are ignored, the file write proceeds) and the "counted, not fatal" treatment plugin notify failures get — hclog cannot react to a writer error beyond discarding the line, so surfacing one buys nothing and loses the record.
  • A degraded daemon is not paid for on every line. Bounding one call is not enough when logger runs once per log line: at 5 seconds each, with hclog serializing writers, all agent logging would crawl. After a failed or timed-out call the agent stops shelling out for syslogSuppressWindow (1 minute) and then lets one line through as a probe, so a wedged daemon costs at most one timeout per minute while the log file keeps every line at full speed.

Forwarding failures are not silent. The transition into failing and the transition back to healthy are each recorded once, in hclog's own line format, in the log file itself:

2026-01-01T00:00:00.000+0000 [WARN]  rewst_agent_smith_<org>: syslog forwarding failed, suppressing it for 1m0s (log lines still reach this file): logger -p daemon.info -t rewst_agent_smith_<org> timed out after 5s:
2026-01-01T00:01:00.000+0000 [WARN]  rewst_agent_smith_<org>: syslog forwarding recovered after 1 failure(s); 128 line(s) reached this log only

Nothing is written in between, so a syslog outage lasting hours cannot flood the log file it is, at that point, the only copy of.

Separately, the severity mapping was wrong on all three platforms. Each writer matched the level bracket in the formatted line itself, and all three looked for [WARNING] — a spelling hclog never emits; it writes [WARN] . Every agent warning was therefore forwarded as daemon.info on Linux/macOS and as an informational event-log entry on Windows. The classification now lives once in internal/syslog/syslog.go (levelForLine, which accepts both spellings) and each platform only maps it onto its own sink, so the three writers cannot drift apart again.

Because Linux and macOS differed only in whether the source is repeated in the message body (macOS' unified logging does not surface logger -t the way syslogd does), the two byte-identical implementations were collapsed into one syslog_unix.go; syslog_linux.go and syslog_darwin.go now hold just that per-platform syslogMessage formatting.

Waiting for the Old Agent Process to Exit

Stopping the service is not the same as the old agent process being gone. Install, update and uninstall used to bridge that gap with a fixed time.Sleep of five seconds and then act regardless — overwriting the agent executable, or deleting the installation directory. On a loaded endpoint the old process is frequently still alive at the five second mark, because its shutdown legitimately drains the commands in flight, tears down MQTT and kills its plugin subprocesses one at a time. Windows then refuses to replace a running image, so the update failed with a sharing violation after the service was already stopped, leaving the device offline until someone intervened. Uninstall had the mirror-image problem: files deleted out from under a live process, leaving an installation that neither ran nor reinstalled.

Nothing is written or deleted now until the process is observed to be gone:

  • The wait polls three real signals, all of which must clear: the service manager no longer reports the service active, no running process is executing the agent binary, and the executable is no longer held open as a running image (a sharing violation on Windows, ETXTBSY on Linux). The three overlap on purpose — the file signal is what actually blocks the write on Windows, but macOS permits writing to a running image, and a service manager can report a service stopped while its process is still winding down. An elapsed poll interval is never by itself treated as evidence of anything.
  • It returns as soon as the process is gone, so a healthy update is faster than the unconditional five second sleep it replaces, not slower.
  • It is bounded by a documented 2 minute deadline, sized for a slow but legitimate shutdown (many workers, long-running commands, several plugin subprocesses). Overrunning it logs at Error what was still outstanding and for how long, and the caller aborts rather than proceeding.
  • Update aborts before writing anything and leaves the installation fully intact, then restarts the service it stopped so a failed update never leaves the endpoint silently offline. Uninstall aborts before deleting the registration or any files, and logs that nothing was removed. Install (--config) aborts before deleting the existing registration and restarts the service it stopped.
  • A probe that cannot run at all (a restrictive ACL, a process table that cannot be enumerated) is logged once at Warn and the remaining signals are used. A probe failure is not evidence a process is alive, and must not wedge every update on an endpoint where it can never succeed.

The agent executable and the config file are also written atomically — to a temporary file in the destination directory, then renamed into place, the same pattern the postback spool uses. An interrupted or failed write therefore leaves the previous file byte-identical rather than truncated: the endpoint keeps running the old agent instead of a binary that cannot start.

Completing Uninstall When a Directory Cannot Be Removed

Uninstall removes three directories: the org's data directory, its program directory, and its scripts directory. It used to remove them sequentially and return on the first RemoveAll failure, after the service registration had already been deleted. A single locked file — an AV scanner mid-scan, a stale open handle on Windows, both routine during uninstall — therefore orphaned every directory behind the one that failed, potentially tens of megabytes, with no service registration left to retry the uninstall cleanly through.

Every directory is now attempted independently:

  • A failure on one directory is logged with its path and the removal moves on to the next, so the directories that can be removed always are.
  • The failures are reported together at the end, at Error, naming every path that survived (failed_directories, plus failed_count of attempted_count), so whoever picks up the cleanup does not have to guess which directories were reached. A clean uninstall logs Uninstall completed.
  • The order stays data, program, scripts. On Linux and macOS the scripts directory lives inside the data directory, so removing the data directory normally takes it along and the later attempt is a no-op (RemoveAll on a missing path succeeds); when the data directory removal fails, that separate attempt is a second, narrower chance to reclaim the scripts underneath it.

This applies only to removing the installed files. Everything before that — a service that will not stop, an agent process that will not exit, a registration that cannot be deleted — still aborts the uninstall with nothing removed, since those failures mean the agent may still be running (see "Waiting for the Old Agent Process to Exit" above).

The integration suite exercises this end to end on Windows: a fixture holds an exclusive handle (no FILE_SHARE_DELETE) on a file it plants inside the program directory, then runs a real uninstall against a real installed service and asserts the data and scripts directories are gone, the registration is gone, the program directory is still there, and the summary line names it. Windows only, because that is the only platform where the failure exists — an open handle does not block unlink on Linux or macOS, so there is no equivalent fixture to build; the platform-independent half (every directory attempted, every failed path reported) is covered by unit tests on all three. The handle is taken on a planted file rather than on the agent executable on purpose: holding the executable open would make the process-exit wait conclude the old agent is still running and abort before removing anything, which is the different scenario above.

Surviving Its Own systemd Stop (Linux)

The --update helper that an auto-update spawns has to stop the running service before it can replace the binary and start it again. On Linux that service runs as a systemd unit, and systemd's default KillMode=control-group tears down every process in the unit's cgroup — not just its main process — when the unit is stopped. The helper is launched as a child of the running service, so without intervention it inherits that cgroup: the moment it calls systemctl stop on its own unit, systemd kills the helper along with the service it just asked to stop, mid-update. The service is left stopped, the binary and config were never touched, and — because the kill is a signal, not a normal return — the helper's own deferred recovery never runs either. Restart=always does not help: the unit was stopped by an explicit systemctl stop, which systemd treats as a clean, intentional exit, not the unexpected one Restart= reacts to.

The helper now runs inside its own transient systemd scope (systemd-run --scope --collect) rather than as a plain child process, so it is never a member of the unit's cgroup in the first place. Stopping the unit it was launched from tears down only that unit's cgroup; the helper's scope is untouched, so it survives to replace the binary, update the config, and start the service again — the same flow already used on Windows and macOS. macOS needed no equivalent change: launchd tears down a stopped job by BSD process group (killpg), and the Setsid the helper already sets moves it into a new process group, which is enough to escape that teardown. Linux's cgroup-based KillMode is inherited across fork() and untouched by setsid(), so the same call that protects the helper on macOS does not protect it on Linux.

Reporting the Agent Version

internal/version.Version is the release tag stamped by the build script (v1.5.7), and its default is v0.0.0 — the same shape — so a plain go build or go test binary differs from a release only in the number. The bare number that the x-rewst-agent-smith-version header and the AGENT_SMITH_VERSION variable in every command script have always carried comes from version.Number(), which strips a leading v only when one is present and returns 0.0.0 for an empty stamp instead of panicking. Before this, both call sites sliced Version[1:] against a bare 0.0.0 default, so every developer and CI build reported .0.0, an empty -X injection would have panicked on the first HTTP request (the config fetch during install), and the only test of the header derived its expectation from the same Version[1:] expression and therefore could not fail. The device twin's agent_version, the startup log line and --diagnostic intentionally keep the tag form; that difference is documented at each call site.

Verified, Version-Gated Auto-Updates

Every auto-update downloads a full agent binary and executes it as the installer, so two questions have to be answered before that binary is trusted: is it actually the release the agent asked for, and is it actually newer than what is already running? Neither was checked before.

  • Checksum verification. Download hashes the installer as it streams to disk and aborts — removing the temp file it already cleans up on any other download failure — unless the hash matches the SHA-256 digest GitHub's Releases API returns natively for that asset (Asset.Digest, format sha256:<hex>): GitHub computes this itself, server-side, from the bytes it received when the release was published, so there is nothing our own release job has to compute or upload alongside the binary. A missing digest, one using an algorithm other than sha256, or one that isn't a well-formed 64-character hex string fails the same way: verification is required, not best-effort, so a corrupted-but-200-OK download, a tampered release asset, or a broken release job can never be executed. (An earlier version of this mechanism instead published a hand-computed <binary-name>.sha256 sidecar asset per binary — .github/workflows/sign.yml still emits it for now as a safety net for agents built before this change, which keep expecting it on every future release until they themselves update, but nothing in the current agent reads it.)
  • Newer-than, not not-equal. Run used to update whenever the latest tag differed from the running version at all. That also updates on an older tag — a release process mistake that republishes or points the check endpoint at a stale release would silently downgrade the whole fleet. The comparison is now a proper semantic version check (isNewerVersion): an update proceeds only when the latest release's MAJOR.MINOR.PATCH is greater than the running version's, and a tag that fails to parse aborts the check with an error instead of being guessed at in either direction.
  • A size ceiling on the download. Download bounds the installer to 200 MiB regardless of the Content-Length header (which can be absent or wrong) — generous headroom over the compiled binary's actual size, so a legitimate release is never at risk, while a misbehaving or compromised release endpoint cannot fill the updates directory, and the volume under it, by serving an oversized or endless body. downloadTimeout already bounds how long the request runs; this bounds how many bytes it can deliver in that time.

Retrying the Config Fetch

The install-time POST to --config-url is answered by a Rewst workflow behind the engine's front door, and a fraction of requests come back with the engine's own ceiling (408), a gateway error (5xx), a rate limit (429) or a transient routing 404 whose body says Workflow was not found; a fresh attempt a few seconds later succeeds. Config mode now classifies the answer: those, and a request that never completed (connection refused, DNS, the request timeout), are transient and retried on a jittered exponential backoff (utils.JitteredBackoff) starting at 5 seconds and capped at 30, three attempts in total; any other non-2xx is a refusal (wrong trigger URL, wrong secret, bad payload) and fails immediately with the status and a body excerpt in the error. Each retry is an Info line naming the attempt and the reason, and an exhausted transient says so: failed to fetch configuration: status 408 (transient; gave up after 3 attempts). Before sc-118306 the first transient answer failed the install outright, which in the integration suite was the largest single source of red runs and for a technician was an install that fails once and works on the second try. The budget is adjustable with --config-max-attempts and --config-base-retry-backoff-seconds.

Bounded Config-Fetch Response

The auto-update download is not the only response body the agent buffers in memory. Config mode's install-time POST to --config-url reads the Rewst config endpoint's reply the same way, and it gets the same treatment: the body is read through an io.LimitReader bounded to 10 MiB (maxConfigResponseSize, cmd/agent_smith/config_context.go), and a response that runs past the ceiling aborts the install with configuration response exceeds maximum allowed size of <n> bytes rather than parsing the prefix that arrived.

The ceiling is deliberately enormous relative to what a legitimate reply looks like — a single small JSON object holding a device id, two hostnames, a key and a handful of tuning fields, on the order of a kilobyte — so no real configuration is ever at risk, while still sitting far below a size that could put memory pressure on the machine being installed. configHTTPTimeout (5 minutes) already bounds how long the request may run, but a slow-but-steady sender can still deliver an effectively unbounded body inside that window, so a compromised, misconfigured or DNS-hijacked config endpoint — or a proxy or middlebox on the path to it — could otherwise drive unbounded allocation on the client during install. Unlike the auto-update path this runs once, at install time, rather than repeating on a schedule; the bound costs nothing either way.

Rejection happens before the body is unmarshalled, so a truncated payload can never be partially applied: the install fails with the size error and the existing installation, if any, is left untouched.

Capped and Jittered Auto-Update Retries

When an update check or download fails, the agent retries on an exponential schedule (base 5 minutes, doubling per retry, 5 retries) before waiting out the next 48-hour check interval. Two bounds keep that schedule safe at fleet scale:

  • A cap. No single retry sleep exceeds 1 hour, or a quarter of the check interval when that is shorter (integration builds shorten the interval via ldflags). Without a ceiling the doubling reaches a 42-hour sleep by retry 10 — longer than the interval the retries are nested inside — and a large retry count overflows the doubling into a negative duration, which makes the wait return immediately and spin the retry loop against the release endpoint. The schedule is computed by iterated doubling with an early exit at the cap, so no intermediate value can overflow and every slot is strictly positive and bounded.
  • Jitter. Each slot is spread by up to ±25%, mirroring the reconnect and postback backoffs. An unjittered schedule makes every agent that failed at the same moment retry at the same instants, so a GitHub releases outage or rate limit turns the whole fleet into a synchronized retry storm that sustains the condition it is recovering from and keeps endpoints on older versions long after the outage ends. Jitter is applied after the cap, and a slot that would exceed the cap is reflected back under it rather than clamped to it — clamping would land half of every capped slot on exactly the ceiling, so a fleet held at the cap by a long outage would re-synchronize there.

The backoff wait stays interruptible by the service stop signal, so a stop is never delayed by a pending retry, and a retry that succeeds resets the schedule for the next cycle.

Reclaiming Downloaded Installer Binaries

Every auto-update downloads a full agent binary and executes it as the installer. That file has to survive the download — the installer is spawned detached and the process that could delete it afterwards is the one the installer replaces — so the agent cannot clean up after itself on the update path. Nothing else did either, so one orphaned binary (tens of megabytes) accumulated per update for the lifetime of the installation. On the space-constrained systems where that matters most — thin VDI images, small VM system disks, appliances — a full temp volume is not just an agent problem: it breaks Windows Installer, application logging, and anything else that needs scratch space, and the agent's own next update fails because it can no longer allocate a temp file.

Two changes reclaim the space:

  • Downloads land in a directory the agent owns. Installers are written to <data directory>/updates (C:\ProgramData\RewstRemoteAgent\<orgId>\updates, /etc/rewst_remote_agent/<orgId>/updates, /Library/Application Support/rewst_remote_agent/<orgId>/updates) instead of the shared system temp directory, with the directory created 0700. A full agent binary is no longer left executable and world-readable, the sweep below only ever runs against a directory this agent created, and endpoints that mount /tmp noexec — a common hardening baseline — can execute the installer at all. Uninstall already removes the data directory wholesale, so nothing is left behind.
  • A startup sweep removes what previous updates left. On every service start, after the service has reported itself running, installer binaries older than 24 hours are removed. Because a successful update restarts the agent, each start reclaims the previous update's installer and leaves the current one alone, so steady-state usage is a single file rather than one per update. The legacy shared temp directory is swept as well, so an upgraded endpoint reclaims everything it has accumulated since it was installed rather than only stopping the growth from here on.

The sweep is deliberately conservative, matching the existing stale-script sweep: only regular files (never symlinks or device nodes) whose name is exactly the installer-<digits>.bin pattern os.CreateTemp produces, and only those past the age threshold — which is why it is safe to point at a directory shared with the rest of the system. The pattern is a shared constant used by both the download and the sweep, so the two cannot drift. It is best effort throughout: an unreadable directory or an unremovable file (a Windows installer still running holds its own image open) is logged and skipped, never failing or delaying agent startup. A non-zero number of removals is logged at Info with the count and directory; individual removals are logged at Debug.

Notification Plugin Supervision

Notification plugins run as separate subprocesses reached over RPC, and every notification the agent sends is best effort — a delivery failure never interferes with command execution. On its own that combination hides a plugin that dies: once the subprocess exits, its RPC client stays broken forever and every later notification (AgentStatus:Online/Offline/Reconnecting, AgentPostbackFailed, AgentMessageDropped) silently goes nowhere.

Loaded plugins are therefore supervised for the lifetime of the service:

  • A subprocess that exits or crashes is detected — by a health check that polls every 30s, and immediately on the notification path if one arrives first — and relaunched, so notifications resume without restarting the agent.
  • Relaunches use an exponential backoff (5s up to 5 minutes) so a plugin that crashes on startup cannot turn into a process-spawn loop. A plugin that ran for at least 2 minutes before dying is treated as a one-off and relaunched immediately.
  • Failures are observable instead of silent: a failed delivery is logged at Error level once per failure transition (with a matching recovery line), so a persistently broken plugin cannot flood the log, and cumulative counters for missed notifications, restarts and failed restarts are reported in a Plugin notification health summary line on shutdown.
  • An error the plugin itself returns is counted and logged but leaves the healthy subprocess running; only a broken RPC channel triggers a relaunch.

Deployments with no plugins configured are unaffected — no supervision runs.

Bounded Notification Plugin RPC Calls

The health check above only detects a subprocess that has actually exited — it reads an in-process exit flag and performs no RPC. A plugin whose process is still alive but has deadlocked internally (blocked on a channel that's never signaled, stuck on a downstream call) is invisible to it, and the Notify RPC call itself used to have no deadline: net/rpc's Call blocks until a response arrives, however long that takes. Since every message and status transition calls Notify on every loaded plugin, one plugin hanging without crashing could silently exhaust the whole worker pool over time — command execution stalling fleet-wide with MQTT connectivity still looking perfectly healthy, and no error anywhere pointing at the cause.

Every Notify call is now bounded by a 10 second timeout, matching the deadline pattern already used for MQTT operations:

  • A call that does not return within the timeout is abandoned and treated as a plugin failure — counted and logged once per failure transition, exactly like a crash — rather than blocking the calling worker any longer.
  • The failure is tracked in its own counter, separate from other notify failures, so a hang is observable as distinct from a crash or a plugin- returned error in both the logs and the Plugin notification health summary.
  • Because the subprocess is still alive at this point, dropping the handle also kills it (the same teardown Kill uses), so the next Notify or health check relaunches a fresh subprocess on the existing backoff schedule instead of repeatedly timing out against a permanently wedged process.

The timeout is a fixed constant rather than a device config knob: unlike the MQTT timeouts, it bounds RPC to a subprocess the host itself launched on the same machine, not a network endpoint an operator might need to tune.

Build

Required tools and packages:

  • commitizen: To use a standardized description of commits.

    pipx ensurepath
    pipx install commitizen
    pipx upgrade commitizen
    
  • go-winres: To embed icons and file versions to windows executables.

    go install github.com/tc-hib/go-winres@latest
    

Run the following command using powershell or pwsh to build the binary:

./scripts/build.ps1

Testing and Coverage

Agent Smith maintains high code quality through comprehensive testing with an 80% coverage threshold.

Running Tests

Run all tests:

go test ./...

Run tests with verbose output:

go test ./... -v

Run tests for a specific package:

go test ./cmd/agent_smith -v
go test ./internal/service -v
go test ./plugins -v

Run a specific test:

go test ./cmd/agent_smith -v -run TestLoadConfig

Coverage Reports

Generate coverage report:

./scripts/coverage.ps1

This script:

  • Runs tests across all packages
  • Generates coverage profiles
  • Enforces 80% minimum coverage threshold

Note: When running tests locally on Linux, some tests write to /tmp/rewst_remote_agent/scripts. If that directory was created by root (e.g., via sudo), your user won't have write access. Fix it by running:

sudo chmod -R o+w /tmp/rewst_remote_agent

Test Categories

Unit Tests: Test individual functions and components in isolation

  • Message parsing and validation
  • Configuration loading
  • SAS token generation
  • Path resolution

Integration Tests: Test component interactions

  • MQTT message flow (with test broker)
  • Service lifecycle (start/stop/restart)
  • Plugin loading and notifications
  • Command execution and postback

The suite's send-command action (.github/actions/send-command/send-command.sh) classifies every answer from the Rewst engine so a red step names its cause: request-failed (curl itself), engine-timeout (the engine's own ceiling), engine-transient (a 5xx, or the transient 404 "Workflow was not found"), engine-refusal (any other 4xx), wrong-result (a 2xx that is not this command's postback, usually another agent on the same device id) and success. Timeouts and transients retry on a bounded schedule with a warning per attempt; refusals, wrong results and curl failures fail on the first attempt. Steps whose output is fixed pass expected_output so a foreign result fails at the send rather than at the log assertion after it. test/sendcommand runs the script against stub engines for every class under a plain go test ./....

Race detector in CI

The Test workflow runs the unit suite twice on ubuntu and macOS: once plainly and once under go test -race, as a separate step so a data-race report is distinguishable from a test failure in the job summary. The agent's command path is concurrent in ways a plain run cannot check - a worker pool that outlives its connection cycle and overlaps the next, the journal ownership map, the postback spool flushed on one goroutine while workers write to it, the log rotator's drain-before-swap - and every one of those was verified under -race locally before it merged with nothing enforcing it (sc-119837). Windows is excluded because -race needs cgo and a C toolchain the hosted runner does not reliably provide; the Windows-only files are covered by the GOOS=windows vet and lint jobs. Run it locally with go test -race ./....

test/workflowlint guards the workflow file itself under the same go test ./...: no matrix variable may be spliced into the agent's arguments (a Windows-only --disable-agent-postback used to ride along in one and was inherited by seven unrelated update steps, so five scenarios silently ran with the agent postback disabled on Windows alone - sc-117885), and --disable-agent-postback may appear in exactly one, Windows-only, step: the one that verifies it.

Bounding the integration jobs

Every job in integration-test.yml carries a timeout-minutes set to roughly twice the p95 of its duration over the last 20 runs (measured 2026-09-28; test p95 27.4 min → 55, test-service-user 3.7 → 10, test-config-response-cap 1.2 → 5, build-integration-test 1.1 → 10 with cold-cache headroom, set-matrix 0.3 → 5; the reusable build.yml job carries 15 because Staging and Release reuse it). Without a bound a hung step ran to GitHub's six-hour default and, under the integration-test concurrency group, held every dispatch behind it (sc-117886). Every step that waits on something outside the runner - the agent binary, the engine, a stub server, a log line - also carries its own bound above its internal one (install-agent 20, send-command 20, fixture actions that set up Go and build a stub 15 - setup-go's cache restore alone has taken over five minutes on a Windows runner - run-agent 10, waits and assertions 5-6), so a hang is attributed to a named step in the job summary rather than to the job. Re-measure with gh run list --workflow integration-test.yml --limit 20 and the per-job startedAt/completedAt when the suite grows; a job that trips its bound on a healthy run is the signal to raise it, and the guard below keeps the bounds present.

test/workflowlint guards the workflow file itself under the same go test ./...: no matrix variable may be spliced into the agent's arguments (a Windows-only --disable-agent-postback used to ride along in one and was inherited by seven unrelated update steps, so five scenarios silently ran with the agent postback disabled on Windows alone - sc-117885), and --disable-agent-postback may appear in exactly one, Windows-only, step: the one that verifies it. It also requires every job (including build.yml's) to carry timeout-minutes, every step that waits on an external system to carry its own, and every curl under .github/ to pass --max-time (sc-117886).

Platform-Specific Tests: Test OS-specific functionality

  • Windows service management
  • Linux systemd integration
  • macOS launchd integration
  • System information collection

Writing Tests

When contributing new code, ensure:

  1. Test coverage: Aim for >80% coverage for new code
  2. Table-driven tests: Use for multiple test cases
    tests := []struct {
        name     string
        input    string
        expected string
    }{
        {"case1", "input1", "expected1"},
        {"case2", "input2", "expected2"},
    }
  3. Clean up resources: Use t.TempDir() and defer statements
  4. Avoid flaky tests: Use proper synchronization and timeouts
  5. Mock external dependencies: Don't rely on network or filesystem in unit tests

CI/CD

Tests run automatically on:

  • Every pull request
  • Every push to main branch
  • Pre-release validation

GitHub Actions Workflows:

  • .github/workflows/test.yml - Runs test suite
  • .github/workflows/coverage.yml - Validates coverage threshold

Pull requests must:

  • ✅ Pass all tests
  • ✅ Maintain ≥80% coverage
  • ✅ Pass all linters
  • ✅ Pass CodeQL security scanning

Code Quality and Linting

Agent Smith uses golangci-lint for strict security and code formatting enforcement.

Running Locally

Install golangci-lint:

See this guide to learn how to install golangci-lint on your local machine.

Run linter:

golangci-lint run

Auto-fix formatting:

golangci-lint run --fix

CI/CD

Linting runs automatically on:

  • Every pull request
  • Every push to main branch

Contributing

Contributions are always welcome. Please submit a PR!

Please use commitizen to format the commit messages. After staging your changes, you can commit the changes with this command.

cz commit

License

Agent Smith is licensed under GNU GENERAL PUBLIC LICENSE. See license file for details.

About

No description, website, or topics provided.

Resources

Stars

8 stars

Watchers

2 watching

Forks

Releases

Packages

Used by

Contributors

Languages