Skip to content

Make the Nexus Dashboard HTTP request timeout configurable - #22

Open
noppanut15 wants to merge 1 commit into
netascode:mainfrom
noppanut15:feat/configurable-request-timeout
Open

noppanut15 wants to merge 1 commit into
netascode:mainfrom
noppanut15:feat/configurable-request-timeout

Conversation

@noppanut15

Copy link
Copy Markdown
Contributor

Summary

The ND client's httpx timeout was hard-coded to 60 seconds, with no flag, environment
variable, or YAML key able to reach it. This exposes it as request_timeout_seconds
under the nexus_dashboard: section — reachable the same three ways as poll_interval
— keeping 60 seconds as the default.

Why

GET /api/v1/manage/fabrics normally answers in well under a second. On a cluster where
at least one site is down, it still returns HTTP 200 but takes roughly 64 seconds
to do so.

Every ND command reaches that endpoint before doing anything useful
(validate_fabric → aci_fabric_names → list_fabrics), so on such a cluster the run
died with httpx.ReadTimeout seconds before the response landed. nd prechange was
where this bit us in practice — an upload that should have kicked off an analysis
instead failed at the inventory lookup.

What changed

Surface Name
CLI flag --request-timeout
Environment variable ND_REQUEST_TIMEOUT_SECONDS
YAML key (under nexus_dashboard:) request_timeout_seconds
Default 60 seconds
  • config.py — DEFAULT_REQUEST_TIMEOUT_SECONDS = 60 as the single source of truth
    for both the dataclass field and the CLI option default. __post_init__ rejects a
    non-positive value alongside the existing poll_interval_seconds and
    job_timeout_minutes guards.
  • commands/_helpers.py — new shared RequestTimeoutOpt, and _build_config takes
    request_timeout instead of hard-coding 60.0.
  • All six ND commands wired up. doctor and snapshots carry no other timing flags
    today, but both hit the fabric inventory, so both get this one.
  • settings.py — the key joins KNOWN_KEYS / ENV_MAP and the existing integer
    coercion set, so a YAML value reaches the environment the same way poll_interval
    does.
  • Docs and templates — ND_HELP, docs/nexus-dashboard.md, docs/configuration.md,
    both config.example.yaml copies, both env.example copies, and the CHANGELOG. The
    aligned columns were re-padded because ND_REQUEST_TIMEOUT_SECONDS is wider than any
    previous variable name, which is most of the diff noise in those files.

Usage

# Flag
nac-analytics nd prechange --request-timeout 180 plan.json

# Environment
ND_REQUEST_TIMEOUT_SECONDS=180 nac-analytics nd prechange plan.json
# nac-analytics.yaml
nexus_dashboard:
  host: nd.example.com
  fabric: FABRIC-A
  request_timeout_seconds: 180

Precedence is unchanged: CLI flag → environment variable → YAML → default.

Validation

--request-timeout 0 exits 4 (InputError) with
--request-timeout must be greater than 0 seconds. rather than reaching httpx, because:

  • timeout=0 makes httpx fail the connection instantly — it means "give the request zero
    time," not "no timeout," so every command would fail instantly.
  • timeout=-1 raises ValueError inside the httpx constructor, which would surface as a
    raw traceback and exit 1.

Testing

241 passed — ruff, ruff format, and mypy clean. New coverage:

  • tests/unit/test_config.py — non-positive values rejected; the default is 60.
  • tests/unit/test_settings.py — a YAML value reaches ND_REQUEST_TIMEOUT_SECONDS, and
    the key is recognised (no "Ignoring unknown config key" warning).
  • tests/unit/test_cli_commands.py — parametrised over flag / env var / default /
    flag-beats-env, asserting the value that lands on Config; a separate test asserts the
    configured value reaches httpx.Client.timeout, and one asserts the exit-4 path.

The use_lab fixture injects its own httpx.Client, so it cannot observe the timeout —
those assertions are made on Config and on a directly constructed NDClient instead.

Verified manually with ND 4.3.

The ND client's httpx timeout was hard-coded to 60 seconds in
_build_config, so a cluster that answered slowly failed every command
with httpx.ReadTimeout and the operator had no way to raise the ceiling.

Expose it the same three ways as poll_interval: --request-timeout,
ND_REQUEST_TIMEOUT_SECONDS, and request_timeout_seconds under the
nexus_dashboard YAML section, defaulting to 60 seconds. It is typed as
an int like the other timing settings, so it shares their env coercion.
Every ND command gets the flag, including doctor and snapshots, since
both reach the fabric inventory before doing anything else. A
non-positive value is rejected as InputError (exit 4) rather than handed
to httpx, which reads 0 as "fail immediately".
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant