> ## Documentation Index
> Fetch the complete documentation index at: https://docs.ocient.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Prometheus Metrics Endpoint

> Access Ocient System metrics in Prometheus exposition format using the /metrics endpoint for integration with Prometheus, Grafana, Datadog, and other monitoring tools.

export const Telegraf = "Telegraf™";

export const PromQL = "PromQL®";

export const Prometheus = "Prometheus®";

export const Datadog = "Datadog®";

export const Ocient = "Ocient®";

export const Grafana = "Grafana®";

export const OcientDataIntelligencePlatform = "OcientAIQ™ Unified Data Platform";

The {OcientDataIntelligencePlatform} exposes a {Prometheus}-compatible `/metrics` endpoint on each node that serves system performance metrics in the [OpenMetrics](https://openmetrics.io/) text format. Use this endpoint to integrate with Prometheus, {Grafana}, {Datadog}, and other monitoring tools that support Prometheus scraping.

The `/metrics` endpoint is available on the same [REST interface](/system-information-rest-endpoints) (port 9090) as the legacy [`/v1/stats`](/statistics-monitoring) JSON endpoint.

## Access the Endpoint

Query the endpoint using `curl` or any HTTP client.

```curl CURL theme={null}
curl http://oc1-lts0:9090/metrics
```

The response uses the Prometheus text exposition format.

Output

```text Prometheus theme={null}
# HELP jemalloc_stats_bytes jemalloc: memory statistics
# TYPE jemalloc_stats_bytes gauge
# UNIT jemalloc_stats_bytes bytes
jemalloc_stats_bytes{type="active"} 1073741824
jemalloc_stats_bytes{type="allocated"} 1048576000
jemalloc_stats_bytes{type="resident"} 1610612736
# HELP local_storage_service_free_space_bytes Free space on storage device in bytes
# TYPE local_storage_service_free_space_bytes gauge
# UNIT local_storage_service_free_space_bytes bytes
local_storage_service_free_space_bytes{device="SDM00001A616"} 356599726080
# HELP cmdcomp_queries Number of queries by state
# TYPE cmdcomp_queries gauge
cmdcomp_queries{status="execution"} 5
cmdcomp_queries{status="parsing"} 2
cmdcomp_queries{status="queue"} 12
```

## Query Parameters

The `/metrics` endpoint supports these query parameters to filter the response.

| Parameter | Type | Default | Description |
| - | - | - | - |
| `level` | string | `info` | Filter metrics by visibility level with these supported values: `info`, `debug`, or `default`. |
| `filter` | string | None | Regular expression pattern to match metric names. The endpoint returns only metrics whose names match the pattern. |

### Filter by Level

Return only production-level (`info`) metrics. The {Ocient} system exports these metrics to Datadog for production monitoring.

```curl CURL theme={null}
curl "http://oc1-lts0:9090/metrics?level=info"
```

Include debug-level metrics for troubleshooting.

```curl CURL theme={null}
curl "http://oc1-lts0:9090/metrics?level=debug"
```

### Filter by Name

Use the `filter` parameter with a regular expression to return only metrics whose names match the pattern. Narrow the response to a specific functional area instead of retrieving all available metrics.

Return only query-processing metrics by matching the `cmdcomp` prefix.

```curl CURL theme={null}
curl -G --data-urlencode "filter=^cmdcomp" http://oc1-lts0:9090/metrics
```

Return all storage-related metrics.

```curl CURL theme={null}
curl -G --data-urlencode "filter=^local_storage_service" http://oc1-lts0:9090/metrics
```

Combine both parameters to return only `info`-level ingestion metrics.

```curl CURL theme={null}
curl -G --data-urlencode "filter=^stream_loader" "http://oc1-lts0:9090/metrics?level=info"
```

## Metric Types

Each metric has a type that determines how you interpret and aggregate it.

| Type | Description | Example |
| - | - | - |
| `counter` | A monotonically increasing value that resets only on process restart. Use `rate()` or `increase()` in queries. | `stream_loader_api_push_rows_requests_total` |
| `gauge` | A value that can increase or decrease. Represents a current measurement. | `cmdcomp_queries{status="execution"}` |

## Metric Units

When a metric has a defined unit, the system appends the unit name as a suffix to the metric name using Prometheus conventions.

| Unit | Suffix | Description |
| - | - | - |
| Bytes | `_bytes` | Memory, disk space, and data transfer sizes. |
| Seconds | `_seconds` | Durations and time measurements. |
| Percent | `_percent` | Ratios expressed as percentages (0–100). |
| None | No suffix | Dimensionless counts or ratios. |

## Metric Labels

Labels provide dimensional metadata that distinguishes different instances of the same metric. For example, `cmdcomp_queries` uses a `status` label to differentiate query states.

```text Prometheus theme={null}
cmdcomp_queries{status="execution"} 5
cmdcomp_queries{status="parsing"} 2
cmdcomp_queries{status="queue"} 12
```

Common labels include these categories.

| Label | Description | Example Values |
| - | - | - |
| `device` | Storage device identifier | Device Universally Unique IDentifier (UUID) |
| `protocol` | Internal protocol name | `storage_cluster`, `vm`, or `health` |
| `status` | Operation or query state | `success`, `error`, or `execution` |
| `type` | Metric subtype or category | `active`, `allocated`, or `page` |
| `stage` | Pipeline processing stage | `batch_preparation` or `segment_transfer` |
| `action` | Protocol action name | `on_get_segment_data` or `activate_segment` |

## Metric Categories

This table summarizes the categories of metrics available with the `/metrics` endpoint.

| Category | Metric Name Prefix | Description |
| - | - | - |
| Query processing | `cmdcomp_*` | Query states, metadata cache operations, SQL Node load, and server connections. |
| Storage devices | `local_storage_service_*` | Disk free space, total space, NVMe SMART health attributes, I/O timeouts, and device status. |
| Ingestion | `stream_loader_*` | Push rows API request counts, data durability, pipeline stage operations, and durations. |
| Memory | `jemalloc_stats_*`, `memory_heap_*`, `memory_huge_*` | Memory allocator statistics and per-subsystem memory usage. |
| Query execution | `vm_*` | Active queries, datablock router throughput, scheduler state, and dispatch queue events. |
| Virtual segments | `virtual_segment_*`, `virtual_read_cache_*` | Segment rebuild activity, read cache hit rates, and block counts. |
| Storage cluster | `storage_cluster_*` | Segment counts, OSN lifecycle, and cluster data statistics. |
| Protocol timing | `protocol_time_*`, `protocol_actions_*` | Internal protocol dispatch delays and action durations. |
| Raft consensus | `protocol_raft_*`, `raft_engine_*` | Raft log sizes, snapshot sizes, participant state, and leader method timing. |
| Disk spill | `disk_based_operator_instance_*` | Bytes spilled to disk during query execution. |
| System configuration | `system_configuration_store_*` | Configuration refresh durations and operation counts. |
| Health | `operator_summary_oldest_age_*`, `healthProtocolInstance_*` | Oldest running operator age and distributed task counts. |
| TKT pipeline | `tkt_*`, `partition_provider_*` | Pipeline operator I/O buffer stalls and emits rows. |

## Key Metrics Reference

These metric families capture monitoring system health, query performance, ingestion throughput, and storage capacity.

### Query States and Connections

These metrics provide visibility into the query lifecycle and client connectivity on each SQL Node.

#### `cmdcomp_queries`

The number of queries currently in each processing state on this node. Use this metric to monitor query queue depth, identify bottlenecks in the compilation pipeline, and detect query pile-ups.

| Attribute | Value |
| - | - |
| Type | `gauge` |
| Unit | None (count) |
| Labels | `status` |

Label values for `status`.

| Value | Description |
| - | - |
| `queue` | Queries waiting to begin processing. |
| `parsing` | Queries that the database parses from SQL text. |
| `compilation` | Queries that the database compiles into execution plans. |
| `optimization` | Queries the query planner optimizes. |
| `validation` | Queries undergoing validation before execution. |
| `vm_initialization` | Queries initializing the virtual machine execution state. |
| `execution` | Queries actively executing. |

This example alerts when the database queues more than 50 queries.

```promql PromQL theme={null}
cmdcomp_queries{status="queue"} > 50
```

#### `cmdcomp_server_connections`

The number of active client connections to the command compiler server on this node. A sudden drop can indicate a network issue, whereas a sustained increase can indicate connection leaks.

| Attribute | Value |
| - | - |
| Type | `gauge` |
| Unit | None (count) |
| Labels | None |

#### `cmdcomp_sql_node_load`

A composite metric representing the total SQL Node load across all nodes. Use this metric to assess whether the query workload is balanced across the cluster.

| Attribute | Value |
| - | - |
| Type | `gauge` |
| Unit | None |
| Labels | None |

#### `cmdcomp_metadata_cache_operations`

Counts of metadata cache operations by result type. A low hit rate (high misses relative to requests) might indicate cache pressure or frequent schema changes.

| Attribute | Value |
| - | - |
| Type | `counter` |
| Unit | None (count) |
| Labels | `type` |

This table describes label values for `type`.

| Value | Description |
| - | - |
| `hit` | Metadata found without recomputation (cache hit). |
| `misses` | Metadata fetched or recomputed (cache miss). |
| `requests` | Total requests to the metadata cache. |

This example calculates the metadata cache hit rate over a 5-minute period.

```promql PromQL theme={null}
rate(cmdcomp_metadata_cache_operations{type="hit"}[5m])
  /
rate(cmdcomp_metadata_cache_operations{type="requests"}[5m])
```

### Ingestion Metrics

These metrics track the data ingestion pipeline from API request receipt through segment generation and transfer.

#### `stream_loader_api_push_rows_requests`

Total count of push-row API requests that the Loader Node receives by outcome. Use this metric to monitor ingestion request volume and error rates.

| Attribute | Value |
| - | - |
| Type | `counter` |
| Unit | None (count) |
| Labels | `status` |

This table describes label values for `status`.

| Value | Description |
| - | - |
| None | Total requests (all outcomes). |
| `success` | Requests that completed successfully. |
| `error` | Requests that failed. |
| `rate_limit_exceeded` | Requests rejected due to rate limiting. |

This example calculates the error rate as a percentage over a 5-minute period.

```promql PromQL theme={null}
rate(stream_loader_api_push_rows_requests{status="error"}[5m])
  /
rate(stream_loader_api_push_rows_requests[5m]) * 100
```

#### `stream_loader_api_row_count`

Total number of rows pushed using the API. Use this metric along with the `stream_loader_api_push_rows_requests` metric to calculate the average batch size.

| Attribute | Value |
| - | - |
| Type | `counter` |
| Unit | None (count) |
| Labels | None |

#### `stream_loader_data_durable_bytes`

Bytes that the Loader Node has made durable (persisted) by the storage tier.

| Attribute | Value |
| - | - |
| Type | `gauge` |
| Unit | Bytes |
| Labels | `type` |

This table describes the label values for `type`.

| Value | Description |
| - | - |
| `api` | Durable bytes received from the API. |
| `page` | Durable bytes written to pages. |

#### `stream_loader_data_row_count`

Current number of rows in the Loader Node by durability status.

| Attribute | Value |
| - | - |
| Type | `gauge` |
| Unit | None (count) |
| Labels | `durability` |

This table describes label values for `durability`.

| Value | Description |
| - | - |
| `durable` | Rows persisted to storage. |
| `non_durable` | Rows still in-memory awaiting persistence. |

#### `stream_loader_operations_count`

Counts of pipeline stage operations by stage and outcome. Use this metric to identify the stage in the ingestion pipeline that experiences the most errors or throughput.

| Attribute | Value |
| - | - |
| Type | `counter` |
| Unit | None (count) |
| Labels | `stage` or `status` |

This table describes label values for `stage`.

| Value | Description |
| - | - |
| `batch_preparation` | Prepare raw data into batches. |
| `segment_generation` | Generate storage segments from batches. |
| `segment_grouping` | Group segments for transfer. |
| `segment_partitioning` | Partition segments across the cluster. |
| `segment_transfer` | Transfer segments to storage nodes. |

This table describes label values for `status`.

| Value | Description |
| - | - |
| None | Total operations. |
| `success` | Successful completions. |
| `error` | Failed operations. |

#### `stream_loader_operations_duration`

Most recent duration, in seconds, of each pipeline stage. Use this metric to identify slow stages in the ingestion pipeline.

| Attribute | Value |
| - | - |
| Type | `gauge` |
| Unit | Seconds |
| Labels | `stage` |

### Disk Health and Space

These metrics monitor storage device capacity, health, and performance on each Foundation Node.

#### `local_storage_service_free_space_bytes`

Free space, in bytes, on each storage device. Monitor this value to detect nodes approaching full capacity.

| Attribute | Value |
| - | - |
| Type | `gauge` |
| Unit | Bytes |
| Labels | `device` |

This example finds devices with less than 10% free space.

```promql PromQL theme={null}
local_storage_service_space_free_pct < 10
```

#### `local_storage_service_total_space_bytes`

Total capacity, in bytes, of each storage device.

| Attribute | Value |
| - | - |
| Type | `gauge` |
| Unit | Bytes |
| Labels | `device` |

#### `local_storage_service_space_free_pct`

Percentage of free space on each storage device (0–100).

| Attribute | Value |
| - | - |
| Type | `gauge` |
| Unit | Percent |
| Labels | `device` |

#### `local_storage_service_io_timeouts`

Count of NVMe I/O timeouts by the `opcode` category. Any non-zero value warrants investigation as it might indicate drive degradation.

| Attribute | Value |
| - | - |
| Type | `counter` |
| Unit | None (count) |
| Labels | `opcode` |

This table describes label values for `opcode`.

| Value | Description |
| - | - |
| `admin` | Administrator command timeouts. |
| `read` | Read command timeouts. |
| `write` | Write command timeouts. |
| `dsm` | Data set management (TRIM or deallocate) timeouts. |
| `other` | Other `opcode` timeouts. |

#### `local_storage_service_device_endurance`

Estimated percentage of device life used (0–100). As this value approaches 100, schedule the drive for replacement.

| Attribute | Value |
| - | - |
| Type | `gauge` |
| Unit | Percent |
| Labels | `device` |

#### `local_storage_service_available_spare_pct`

Available spare capacity on the NVMe device, normalized to a percentage. When this value drops below the threshold (see the `local_storage_service_available_spare_threshold_pct` metric), the device controller asserts a warning.

| Attribute | Value |
| - | - |
| Type | `gauge` |
| Unit | Percent |
| Labels | `device` |

#### `local_storage_service_crc_errors`

Cumulative count of CRC errors detected by the device. A rising count indicates potential data integrity issues on the link between the host and drive.

| Attribute | Value |
| - | - |
| Type | `counter` |
| Unit | None (count) |
| Labels | `device` |

### Active Queries and Operator Age

These metrics provide high-level visibility into query execution activity and help detect stuck or long-running operations.

#### `vm_active_queries_count`

Total number of queries currently executing on the virtual machine of this node. Use this metric to gauge overall query concurrency and correlate with resource consumption.

| Attribute | Value |
| - | - |
| Type | `gauge` |
| Unit | None (count) |
| Labels | None |

This example alerts when active queries exceed the expected concurrency limit.

```promql PromQL theme={null}
vm_active_queries_count > 200
```

#### `operator_summary_oldest_age_seconds`

Age, in seconds, of the oldest currently running operator instance on this node. A large value might indicate a stuck query or an operation waiting on an unavailable resource.

| Attribute | Value |
| - | - |
| Type | `gauge` |
| Unit | Seconds |
| Labels | None |

This example alerts when any operator has been running for more than one hour.

```promql PromQL theme={null}
operator_summary_oldest_age_seconds > 3600
```

#### `disk_based_operator_instance_bytes_spilled`

Total bytes that query operators have spilled to disk when the database exceeds in-memory budgets. A high spill volume indicates that queries are exceeding available memory and resorting to disk-based processing, which is significantly slower.

| Attribute | Value |
| - | - |
| Type | `counter` |
| Unit | Bytes |
| Labels | `compression` |

This table describes label values for `compression`.

| Value | Description |
| - | - |
| `compressed` | Bytes written to disk after compression. |
| `uncompressed` | Logical bytes spilled before compression. |

This example monitors the spill rate in bytes per second.

```promql PromQL theme={null}
rate(disk_based_operator_instance_bytes_spilled{compression="uncompressed"}[5m])
```

## Configure Prometheus Scraping

If you run a Prometheus server, you can configure it to periodically scrape the `/metrics` endpoint on each Ocient node. Prometheus stores the scraped metrics in its time series database, making them available for {PromQL} queries and Grafana dashboards.

Add a scrape job to your `prometheus.yml` configuration file. List every Ocient node as a target and set `metrics_path` to `/metrics`. The `params` section passes query parameters, such as `level`, with each scrape request.

```yaml prometheus.yml theme={null}
scrape_configs:
  - job_name: 'ocient'
    scrape_interval: 30s
    params:
      level: ['info']
    static_configs:
      - targets:
          - 'oc1-lts0:9090'
          - 'oc1-lts1:9090'
          - 'oc1-sql0:9090'
    metrics_path: '/metrics'
```

After you save and reload the Prometheus configuration, the targets appear on the Prometheus **Status > Targets** page. Prometheus begins collecting metrics at the configured `scrape_interval` interval.

## Configure the Datadog Agent

If you use Datadog instead of Prometheus, Ocient provides an official Datadog integration that scrapes the same `/metrics` endpoint and forwards metrics to the Datadog cloud platform. For details, see [Datadog Integration](/datadog-integration).

## Example Queries

These example PromQL queries demonstrate common monitoring use cases. You can execute them in the Prometheus expression browser, in a Grafana dashboard panel, or with any tool that supports PromQL.

**Disk Space Utilization**

Calculate the percentage of free disk space on each storage device. A declining value indicates that the node is running low on storage and might need capacity planning or data retention adjustments.

```promql PromQL theme={null}
local_storage_service_free_space_bytes / local_storage_service_total_space_bytes * 100
```

**Query Queue Depth**

Return the current number of queries waiting in the queue. A consistently high value might indicate that the system is under heavy query load or that additional SQL Nodes are needed to handle the workload.

```promql PromQL theme={null}
cmdcomp_queries{status="queue"}
```

**Loader Node Error Rate**

Calculate the ratio of failed push-row API requests to total requests over the last 5 minutes. The `rate()` function converts the raw counter into a per-second rate. A rising error rate might indicate issues with the data source, schema mismatches, or pipeline configuration errors.

```promql PromQL theme={null}
rate(stream_loader_api_push_rows_requests_total{status="error"}[5m])
  /
rate(stream_loader_api_push_rows_requests_total[5m])
```

**Memory Allocator Pressure**

Compare the amount of memory actively in use by `jemalloc` to the total memory mapped from the operating system. A ratio approaching 1.0 indicates that most mapped memory is in active use, which might signal memory pressure on the node.

```promql PromQL theme={null}
jemalloc_stats_bytes{type="active"} / jemalloc_stats_bytes{type="mapped"}
```

## Comparison with the Legacy `/v1/stats` Endpoint

This table compares the legacy `/v1/stats` endpoint features to the Prometheus metrics.

| Feature | `/v1/stats` (Legacy) | `/metrics` (Prometheus) |
| - | - | - |
| Format | JSON | Prometheus text exposition |
| Naming | Dot-delimited (e.g., `localStorageService.device.spaceFree`) | Underscore-delimited with unit suffix (e.g., `local_storage_service_free_space_bytes`) |
| Labels | Encoded in metric name (e.g., `device[UUID]`) | Prometheus labels (e.g., `{device="UUID"}`) |
| Filtering | `filter` query parameter (regex) | `filter` and `level` query parameters |
| Tooling | Custom JSON parsers or {Telegraf} | Prometheus, Grafana, Datadog, or any OpenMetrics client |
| Level filtering | Not supported | `level=info\|debug\|default` |

Both endpoints remain available on port 9090 and serve the same underlying metric data.

## Related Links

* [Datadog Integration](/datadog-integration)
* [Statistics Monitoring](/statistics-monitoring)
* [System Information REST Endpoints](/system-information-rest-endpoints)
* [Set Up System Monitoring with the TIG Stack and Kapacitor](/set-up-system-monitoring-with-the-tig-stack-and-kapacitor)
