Skip to main content
The OcientAIQ™ Unified Data Platform exposes a Prometheus®-compatible /metrics endpoint on each node that serves system performance metrics in the OpenMetrics text format. Use this endpoint to integrate with Prometheus, Grafana®, Datadog®, and other monitoring tools that support Prometheus scraping. The /metrics endpoint is available on the same REST interface (port 9090) as the legacy /v1/stats JSON endpoint.

Access the Endpoint

Query the endpoint using curl or any HTTP client.
CURL
The response uses the Prometheus text exposition format. Output
Prometheus

Query Parameters

The /metrics endpoint supports these query parameters to filter the response.

Filter by Level

Return only production-level (info) metrics. The Ocient® system exports these metrics to Datadog for production monitoring.
CURL
Include debug-level metrics for troubleshooting.
CURL

Filter by Name

Use the filter parameter with a regular expression to return only metrics whose names match the pattern. Narrow the response to a specific functional area instead of retrieving all available metrics. Return only query-processing metrics by matching the cmdcomp prefix.
CURL
Return all storage-related metrics.
CURL
Combine both parameters to return only info-level ingestion metrics.
CURL

Metric Types

Each metric has a type that determines how you interpret and aggregate it.

Metric Units

When a metric has a defined unit, the system appends the unit name as a suffix to the metric name using Prometheus conventions.

Metric Labels

Labels provide dimensional metadata that distinguishes different instances of the same metric. For example, cmdcomp_queries uses a status label to differentiate query states.
Prometheus
Common labels include these categories.

Metric Categories

This table summarizes the categories of metrics available with the /metrics endpoint.

Key Metrics Reference

These metric families capture monitoring system health, query performance, ingestion throughput, and storage capacity.

Query States and Connections

These metrics provide visibility into the query lifecycle and client connectivity on each SQL Node.

cmdcomp_queries

The number of queries currently in each processing state on this node. Use this metric to monitor query queue depth, identify bottlenecks in the compilation pipeline, and detect query pile-ups. Label values for status. This example alerts when the database queues more than 50 queries.
PromQL

cmdcomp_server_connections

The number of active client connections to the command compiler server on this node. A sudden drop can indicate a network issue, whereas a sustained increase can indicate connection leaks.

cmdcomp_sql_node_load

A composite metric representing the total SQL Node load across all nodes. Use this metric to assess whether the query workload is balanced across the cluster.

cmdcomp_metadata_cache_operations

Counts of metadata cache operations by result type. A low hit rate (high misses relative to requests) might indicate cache pressure or frequent schema changes. This table describes label values for type. This example calculates the metadata cache hit rate over a 5-minute period.
PromQL

Ingestion Metrics

These metrics track the data ingestion pipeline from API request receipt through segment generation and transfer.

stream_loader_api_push_rows_requests

Total count of push-row API requests that the Loader Node receives by outcome. Use this metric to monitor ingestion request volume and error rates. This table describes label values for status. This example calculates the error rate as a percentage over a 5-minute period.
PromQL

stream_loader_api_row_count

Total number of rows pushed using the API. Use this metric along with the stream_loader_api_push_rows_requests metric to calculate the average batch size.

stream_loader_data_durable_bytes

Bytes that the Loader Node has made durable (persisted) by the storage tier. This table describes the label values for type.

stream_loader_data_row_count

Current number of rows in the Loader Node by durability status. This table describes label values for durability.

stream_loader_operations_count

Counts of pipeline stage operations by stage and outcome. Use this metric to identify the stage in the ingestion pipeline that experiences the most errors or throughput. This table describes label values for stage. This table describes label values for status.

stream_loader_operations_duration

Most recent duration, in seconds, of each pipeline stage. Use this metric to identify slow stages in the ingestion pipeline.

Disk Health and Space

These metrics monitor storage device capacity, health, and performance on each Foundation Node.

local_storage_service_free_space_bytes

Free space, in bytes, on each storage device. Monitor this value to detect nodes approaching full capacity. This example finds devices with less than 10% free space.
PromQL

local_storage_service_total_space_bytes

Total capacity, in bytes, of each storage device.

local_storage_service_space_free_pct

Percentage of free space on each storage device (0–100).

local_storage_service_io_timeouts

Count of NVMe I/O timeouts by the opcode category. Any non-zero value warrants investigation as it might indicate drive degradation. This table describes label values for opcode.

local_storage_service_device_endurance

Estimated percentage of device life used (0–100). As this value approaches 100, schedule the drive for replacement.

local_storage_service_available_spare_pct

Available spare capacity on the NVMe device, normalized to a percentage. When this value drops below the threshold (see the local_storage_service_available_spare_threshold_pct metric), the device controller asserts a warning.

local_storage_service_crc_errors

Cumulative count of CRC errors detected by the device. A rising count indicates potential data integrity issues on the link between the host and drive.

Active Queries and Operator Age

These metrics provide high-level visibility into query execution activity and help detect stuck or long-running operations.

vm_active_queries_count

Total number of queries currently executing on the virtual machine of this node. Use this metric to gauge overall query concurrency and correlate with resource consumption. This example alerts when active queries exceed the expected concurrency limit.
PromQL

operator_summary_oldest_age_seconds

Age, in seconds, of the oldest currently running operator instance on this node. A large value might indicate a stuck query or an operation waiting on an unavailable resource. This example alerts when any operator has been running for more than one hour.
PromQL

disk_based_operator_instance_bytes_spilled

Total bytes that query operators have spilled to disk when the database exceeds in-memory budgets. A high spill volume indicates that queries are exceeding available memory and resorting to disk-based processing, which is significantly slower. This table describes label values for compression. This example monitors the spill rate in bytes per second.
PromQL

Configure Prometheus Scraping

If you run a Prometheus server, you can configure it to periodically scrape the /metrics endpoint on each Ocient node. Prometheus stores the scraped metrics in its time series database, making them available for PromQL® queries and Grafana dashboards. Add a scrape job to your prometheus.yml configuration file. List every Ocient node as a target and set metrics_path to /metrics. The params section passes query parameters, such as level, with each scrape request.
prometheus.yml
After you save and reload the Prometheus configuration, the targets appear on the Prometheus Status > Targets page. Prometheus begins collecting metrics at the configured scrape_interval interval.

Configure the Datadog Agent

If you use Datadog instead of Prometheus, Ocient provides an official Datadog integration that scrapes the same /metrics endpoint and forwards metrics to the Datadog cloud platform. For details, see Datadog Integration.

Example Queries

These example PromQL queries demonstrate common monitoring use cases. You can execute them in the Prometheus expression browser, in a Grafana dashboard panel, or with any tool that supports PromQL. Disk Space Utilization Calculate the percentage of free disk space on each storage device. A declining value indicates that the node is running low on storage and might need capacity planning or data retention adjustments.
PromQL
Query Queue Depth Return the current number of queries waiting in the queue. A consistently high value might indicate that the system is under heavy query load or that additional SQL Nodes are needed to handle the workload.
PromQL
Loader Node Error Rate Calculate the ratio of failed push-row API requests to total requests over the last 5 minutes. The rate() function converts the raw counter into a per-second rate. A rising error rate might indicate issues with the data source, schema mismatches, or pipeline configuration errors.
PromQL
Memory Allocator Pressure Compare the amount of memory actively in use by jemalloc to the total memory mapped from the operating system. A ratio approaching 1.0 indicates that most mapped memory is in active use, which might signal memory pressure on the node.
PromQL

Comparison with the Legacy /v1/stats Endpoint

This table compares the legacy /v1/stats endpoint features to the Prometheus metrics. Both endpoints remain available on port 9090 and serve the same underlying metric data.
Last modified on September 24, 2026