Caasi v0.3.0 Observability: monitor · logs · audit · provenance

Observability

Once a run is going you need to watch it, read its logs, account for what happened and prove you can reproduce it. Caasi gives each of those jobs its own verb — monitor, run logs, audit and the provenance bundle — and none of them import the heavy stack. They read /proc, nvidia-smi, the run store under ~/.caasi/runs/ and your project's .caasi/ metadata.

The five verbs are distinct

doctor, check, test, validate and audit answer five different questions. doctor does not absorb check or audit; audit is the fifth verb, added in 0.3.0. See the five-verb rule below.

The five-verb rule

The organising principle of the whole CLI. Every diagnostic-ish command is one of five verbs, and each verb has exactly one job:

VerbQuestion it answersScope
doctorWhat's wrong with my environment? (is the machine capable?)machine / toolchain
checkCan this work here? (preflight / compatibility)project ↔ machine, per-domain
testDoes it actually work when executed? (live smoke)per-domain
validateIs this config / data / resource structurally valid?per-domain
auditCan I account for and reproduce what happened?project + runs + environment

inspect is the per-domain deep-dive (env inspect, project inspect, run inspect); setup provisions the machine/toolchain; project manages the robotics project. All check entry points share one Check Contract.

caasi monitor

A live resource panel for a run and the machine — the gpu monitor pattern widened to CPU, RAM, GPU, VRAM, throughput and runtime. It defaults to the latest run; with no runs recorded it shows the machine alone.

shellcaasi monitor
# live panel, refreshed every 2s — Ctrl+C stops watching, the run keeps going
caasi monitor --once
Run         wave (20260905-142301-wave)
Status      running
CPU         12.5%
RAM         8.5 / 64.0 GiB
GPU         30.0%
VRAM        6.2 / 24.0 GiB
Simulation  980.0 steps/s
Runtime     00:02:15
caasi monitor --once --json | jq '{cpu,ram_used,gpu,runtime_s}'
{
  "cpu": 12.5,
  "ram_used": 8.5,
  "gpu": 30.0,
  "runtime_s": 135.2
}
Argument / optionEffect
[query]Run id, unique prefix, name or latest (the default). With no runs, the machine is shown alone.
--interval <float>Seconds between refreshes (default 2.0, minimum 0.1).
--oncePrint a single snapshot and exit — the honest replacement for the dropped root status.
--jsonEmit the snapshot as JSON (implies a single sample).

The JSON payload carries the run identity (run, name, status, backend) plus the sample: cpu, ram_used, ram_total, gpu, vram_used, vram_total, processes, runtime_s, and steps_per_s only when a throughput metric was parsed (it is omitted, never faked).

Stays in its lane

gpu monitor remains GPU-only; system status|memory|processes remain machine-only. monitor is the whole-run view that ties them together.

Aggregated run logs

run logs reads a run's captured stdout/stderr and can filter it. The root caasi logs is a thin alias of caasi run logs — same command, one underlying operation.

shellcaasi run logs latest -f
# follow until the run ends — Ctrl+C stops following, the run keeps going
caasi run logs latest --component nav2 --errors
[14:23:31] [nav2] [ERROR] planner: no valid path to goal
[14:23:44] [nav2] [ERROR] controller: aborted — recovery triggered
caasi run logs latest --since 10m -n 200
caasi logs latest --stream stderr
# root alias
OptionShortEffect
--lines-nTrailing lines to show (default 50).
--follow-fFollow the log until the run ends.
--stream-sstdout (default) or stderr.
--component—Only lines tagged with this component (e.g. nav2).
--errors—Only lines that look like errors.
--since—Time window: 30s, 10m, 2h, 1d. Untimestamped lines are excluded.
--json—Structured log records.

caasi audit

The fifth verb: account for a project and its reproducibility. audit walks six scopes and renders a ✓ / ⚠ / ✗ report. It never fails on warnings — exit is 1 only when there are errors.

shellcaasi audit
CAASI Audit

PROJECT
  ✓ Valid caasi.yaml
  ⚠ Working tree has uncommitted changes

DEPENDENCIES
  ✓ isaac.lab 4.5.0
  ⚠ nav2: version not verified

ENVIRONMENT
  ✓ Environment fingerprint recorded
  ✓ GPU available
  ✓ CUDA 13.0

FILES
  ✓ All CAASI directories present
  ✓ Component definitions look valid

RUNS
  ✓ 2 runs recorded
  ✓ Latest run has logs

REPRODUCIBILITY
  ✓ git HEAD 1a2b3c4
  ✓ caasi.lock present

Result: 6 warnings, 0 errors
echo $?
0
ScopeWhat it accounts for
projectcaasi.yaml present, schema-valid and named; git tree clean or dirty.
dependenciesDeclared requires vs. what is detected — reuses the Check Contract detector, so "is it there?" has one owner.
environmentFingerprint recorded, GPU available, CUDA detected, Python drift vs. the recorded fingerprint.
filesCAASI directories present; component definitions well-formed.
runsRecorded runs, orphaned/stale run directories, latest-run logs, parsed throughput metrics.
reproducibilitygit commit, environment fingerprint, caasi.lock, and provenance-bundle completeness for the latest run.

Omit the scope to audit all six, or name one: caasi audit reproducibility. An unknown scope is a usage error (exit 1).

audit --fix is deliberately conservative

--fix applies only CAASI-owned safe fixes. Today that means exactly one thing: creating missing CAASI directories (runs/, artifacts/, logs/, .caasi/ …). It never rewrites or deletes a user file, and never touches user source, packages, drivers, upstream installs, ROS packages or Isaac installs — those categories stay report-only. An audit fix must never destroy evidence.

shellrm -rf artifacts runs && caasi audit files --fix
Applied 2 safe fix(es).
FILES
  ✓ Created missing CAASI directory: artifacts
  ✓ Created missing CAASI directory: runs

Result: 0 warnings, 0 errors

Fixes are off unless asked for. Set audit.fix: true in caasi.yaml (or pass --fix) to enable them; the key is absent by default. See configuration.

Machine-readable audit

shellcaasi audit reproducibility --json
{
  "scope": "reproducibility",
  "result": "ok",
  "warnings": 0,
  "errors": 0,
  "fixed": [],
  "groups": [
    {
      "scope": "reproducibility",
      "findings": [
        { "severity": "ok", "message": "git HEAD 1a2b3c4" },
        { "severity": "ok", "message": "Environment fingerprint recorded" },
        { "severity": "ok", "message": "caasi.lock present" }
      ]
    }
  ]
}

result is ok, warning or error; each finding carries a severity, a message and — when a safe fix exists — a fix hint. The process exits 1 only when errors > 0.

The provenance bundle & run identity

Every run start writes a provenance bundle into the run directory, so a finished run can be accounted for and reproduced. run inspect lists every file in that directory — the bundle plus the captured logs — as {path, size} objects; jq -r '.files[].path' shows just the names:

shellcaasi run inspect latest --json | jq -r '.files[].path'
configuration.yaml
dependencies.yaml
environment.yaml
git.json
hardware.yaml
manifest.yaml
processes.json
stderr.log
stdout.log
FileRecords
manifest.yamlRun identity: id, name, backend, the resolved command, start time, status.
environment.yamlThe environment fingerprint at launch (OS, Python, ROS, Isaac, compilers …).
hardware.yamlCPU, RAM and GPU/VRAM captured at launch.
dependencies.yamlResolved versions of the project's declared requirements.
configuration.yamlThe effective Caasi configuration (secrets redacted).
git.jsonRepo HEAD, branch and dirty state — or {"available": false} when there is no git or no repo.
processes.jsonThe launched process tree (pids, commands).

Provenance degrades honestly: no git → {"available": false}; no parsed throughput → the metric is omitted. Run identity is recorded locally now so it survives the SSH hop when remote execution lands in v0.4.0. audit reproducibility checks this bundle for completeness; env fingerprint and env lock (see Environment) feed it.