ClouisleClouisle

System Observability & Health Monitoring

Look up admin observability metrics, health thresholds, aggregation limits, and diagnostic endpoints

Clouisle provides centralized system observability and health monitoring for platform administrators (requires the admin:dashboard:access global permission). All monitoring routes are mounted under /api/v1/admin/observability, and responses are cached in Redis for 30 seconds (namespace admin:observability:v1).

Time Range and Bucketing Rules

Most metric endpoints accept the time_range query parameter:

  • Supported values: 7d, 30d (default), 90d, all (all historical records). Invalid values fall back to 30d.
  • Trend granularity (granularity): hour or day. When omitted, 7d defaults to hour buckets and other ranges default to day buckets.

Metric Categories

1. Overview

Endpoint: GET /api/v1/admin/observability/overview

Metric FieldTypeDefinition
totals.agent_requestsintegerCanonical assistant-final messages in the period
totals.workflow_runsintegerTotal workflow run instances in the period
totals.total_requestsintegerCombined agent requests and workflow runs
totals.total_tokensintegerTotal tokens consumed
rates.agent_success_ratenumber (0-100)Agent success rate (round_status == 'completed')
rates.workflow_success_ratenumber (0-100)Workflow run success rate (status == 'success')
rates.overall_success_ratenumber (0-100)Combined execution success rate
rates.timeout_ratenumber (0-100)Combined error/timeout rate (workflow status == 'timeout' plus agent round_status == 'error')
latency.p50_ms / p90_ms / p95_ms / p99_msnumberContinuous percentiles of total duration (agent duration_ms and workflow total_duration_ms)
ttft.p50_ms / p90_ms / p95_ms / p99_msnumberTime-to-first-token percentiles, computed from agent first_token_ms only
throughput.current_qpsnumberEvents created in the preceding 60 seconds divided by 60
throughput.peak_hourly_requestsintegerHighest number of requests in a single hourly bucket

2. Entity Performance (Agents & Workflows)

Endpoints:

  • Agent list: GET /api/v1/admin/observability/agents (sortable by requests, p50, p90, p95, p99, timeout_rate, success_rate, tokens)
  • Agent detail and trend: GET /api/v1/admin/observability/agent/{agent_id}
  • Workflow list: GET /api/v1/admin/observability/workflows (additionally sortable by runs, failed_nodes)
  • Workflow detail and node bottlenecks: GET /api/v1/admin/observability/workflow/{workflow_id}

The workflow detail endpoint also returns node-type statistics for the top 20 most-executed node types (execution_count, failed_count, avg_duration_ms), which makes it fast to identify which node (for example http_request or llm) is blocking a pipeline.

3. Timeout and Failure Diagnosis

Endpoint: GET /api/v1/admin/observability/timeouts (supports source=all|agent|workflow)

  • Workflow timeout events: runs with status = 'timeout', reported as timeout_type = 'workflow'.
  • Agent error events: assistant interactions with round_status = 'error'. Because historical messages do not separate the error subtype, the type is reported as unknown, and the response explicitly sets agent_timeout_type_available: false.

4. Throughput and Token Consumption

  • GET /api/v1/admin/observability/throughput: returns real-time qps, running_workflows, and time-bucketed request series.
  • GET /api/v1/admin/observability/tokens: aggregates token usage by source (agent, workflow, other) and by model (by_model). For 7d and 30d it also compares against team model quota counters and takes the larger value, preventing undercounting when events are missing.

System Health and Thresholds

Endpoint: GET /api/v1/admin/observability/system/health

The system collects hardware metrics and retains roughly the most recent 120 health snapshots in a Redis list (LTRIM 0..120; the whole list expires after 24 hours — retrieve the trend via GET .../system/trend). Thresholds and fields:

MonitorWarning / Danger ThresholdsKey Fields
CPUusage_percent >= 70% warning; >= 90% dangerstatus, usage_percent, cores, architecture
Memoryusage_percent >= 80% warning; >= 90% dangerstatus, usage_percent, used_bytes, total_bytes
Diskusage_percent >= 80% warning; >= 90% dangerstatus, usage_percent, used_bytes, total_bytes
DatabaseConnection usage derived from PostgreSQL pg_stat_activity and max_connectionsstatus, active_connections, max_connections
RedisConnectivity and cache efficiencystatus, used_memory, connected_clients, ops_per_sec, hit_rate
WorkersCelery nodes and backlog across 5 core queuesSee the worker section below

Celery Workers and Core Queues

Endpoint: GET /api/v1/admin/observability/system/workers

Clouisle defines 5 fixed queues:

  • default: general system tasks and email delivery
  • knowledge: document parsing, chunk extraction, and vector indexing
  • workflow: asynchronous workflow execution and long-running scheduling
  • agent: background agent reasoning and async inference
  • sandbox: code node sandbox execution

Queue inspection scans Redis lists in batches of 500 messages with a per-queue cap of 5000. Backlog beyond that cap is reported as unscanned:{queue}.

Slow Query Diagnostics

Endpoint: GET /api/v1/admin/observability/system/slow-queries

  • Default threshold is threshold_ms=1000 (queries slower than 1 second), returning duration, call count, and returned rows with pagination.
  • Prerequisite: PostgreSQL must have the pg_stat_statements extension enabled. If unavailable, the endpoint returns available: false with the reason.

Database Aggregate Concurrency Limit

Statistics and observability analysis fan out independent aggregates with asyncio.gather. Without a bound, one request can hold every slot of the shared connection pool and stall regular APIs.

Clouisle bounds this with a global semaphore:

  • Environment variable: DB_AGGREGATE_CONCURRENCY (default 4, must be greater than 0).
  • Mechanism: a single process-wide asyncio.Semaphore, applied per query via run_bounded() rather than once around an entire gather().
  • Rationale: the Tortoise shared pool defaults to maxsize=5 and PostgreSQL defaults to max_connections=100. The default deployment already runs roughly 17 processes, each holding a pool, so limiting a single process's aggregate concurrency to 4 keeps analytics from exhausting database connections at any moment.

How is this guide?

On this page