OCI Instance Pool monitoring
Oracle Cloud Infrastructure Instance Pools let you manage multiple compute instances as a group. Site24x7 integrates with OCI Instance Pools to provide visibility into instance pool size, running and provisioning instances, compute resource utilization, network and disk activity, GPU and RDMA metrics, autoscaling configuration, and instance pool status.
Overview
Site24x7's OCI Instance Pool monitoring provides visibility into the performance, capacity, and status of instance pools in Oracle Cloud Infrastructure. It helps you monitor key resource metrics such as CPU and memory utilization, disk and network activity, GPU performance, and RDMA activity. You can also track instance pool capacity, including the number of running, provisioning, and terminated instances, and monitor autoscaling-related activity. The monitor supports alerts for stopped, failed, and scaling states, along with IT automation actions to start, stop, reset, or soft reset an instance pool when specific monitoring conditions are met.
Use case
Instance Pools are commonly used to manage groups of compute instances that can be scaled according to workload requirements. Changes in instance count, resource utilization, provisioning activity, or autoscaling configuration can affect application capacity and performance. OCI Instance Pool monitoring in Site24x7 helps you track these changes and configure alerts for conditions that require attention.
For example, when an instance pool is scaling, you can monitor the number of provisioning and running instances to verify that the expected capacity is being established. Similarly, CPU, memory, disk, network, GPU, and RDMA metrics can help identify resource utilization or infrastructure issues affecting instances in the pool.
Benefits of Site24x7's OCI Instance Pool integration
Site24x7's OCI Instance Pool monitoring provides the following benefits:
- Instance pool visibility: Monitor instance pool size and the number of running, provisioning, and terminated instances.
- Resource monitoring: Track CPU, memory, disk, and network activity across instances in the pool.
- GPU monitoring: Monitor GPU utilization, memory utilization, temperature, power draw, and ECC errors for supported workloads.
- RDMA monitoring: Track RDMA traffic and identify PCIe, cable, link speed, and other RDMA-related faults.
- Autoscaling visibility: View autoscaling configurations and policies associated with an instance pool.
- Proactive alerting: Configure thresholds for supported metrics and receive alerts when monitored conditions breach the configured limits.
- Capacity monitoring: Use capacity monitoring to track instance pool capacity and resource utilization.
- IT automation: Automate instance pool operations when monitoring conditions or alerts are triggered.
Setup and configuration
To get started with OCI Instance Pool monitoring, complete the following setup steps:
- Site24x7 uses cross tenancy access to monitor OCI resources using the Site24x7 tenancy user. Create the required OCI policy to allow Site24x7 to view your resources.
- Log in to your Site24x7 account and navigate to Cloud > OCI > Integrate OCI Monitor.
- On the Integrate OCI Monitor page, select Instance Pool from the Services to be discovered list.
Permissions
Ensure that Site24x7 receives the following permissions to monitor OCI Instance Pool:
- GetInstancePool
- ListInstancePools
- ListAutoScalingPolicies
- ListAutoScalingConfigurations
- GetAutoScalingPolicy
Polling frequency
Site24x7 queries OCI service level APIs according to the configured polling frequency, which can range from once a minute to once a day, to collect performance data and metadata from the OCI Instance Pool monitor.
Supported metrics
The supported metrics for the OCI Instance Pool monitor are given below.
| Metric name | Description | Statistics | Unit |
|---|---|---|---|
| CpuUtilization | CPU utilization of the instances in the pool. | Mean | Percentage |
| DiskBytesRead | Amount of data read from disk. | Mean | MB |
| DiskBytesWritten | Amount of data written to disk. | Mean | MB |
| DiskIopsRead | Number of disk read operations. | Mean | Count |
| DiskIopsWritten | Number of disk write operations. | Mean | Count |
| LoadAverage | Load average of the instances in the pool. | Mean | Count |
| MemoryAllocationStalls | Number of memory allocation stalls. | Mean | Count |
| MemoryUtilization | Memory utilization of the instances in the pool. | Mean | Percentage |
| NetworksBytesIn | Number of bytes received by the instances in the pool. | Mean | Bytes |
| NetworksBytesOut | Number of bytes sent by the instances in the pool. | Mean | Bytes |
| GpuUtilization | GPU utilization. | Mean | Percentage |
| GpuMemoryUtilization | GPU memory utilization. | Mean | Percentage |
| GpuPowerDraw | GPU power draw. | Sum | Count |
| GpuTemperature | GPU temperature. | Maximum | Count |
| GpuEccSingleBitErrors | Number of GPU ECC single bit errors. | Sum | Count |
| GpuEccDoubleBitErrors | Number of GPU ECC double bit errors. | Sum | Count |
| Fault | Number of GPU faults. | Maximum | Count |
| RdmaLinkSpeedFault | RDMA link speed faults. | Maximum | Count |
| RdmaPcieAddressFault | RDMA PCIe address faults. | Maximum | Count |
| RdmaPcieBerCheckFault | RDMA PCIe BER check faults. | Maximum | Count |
| RdmaPcieCableFlapFault | RDMA PCIe cable flap faults. | Maximum | Count |
| RdmaPcieCablePlugFault | RDMA PCIe cable plug faults. | Maximum | Count |
| RdmaPcieCableStateFault | RDMA PCIe cable state faults. | Maximum | Count |
| RdmaTxBytes | Number of bytes transmitted through RDMA. | Sum | Bytes |
| RdmaRxBytes | Number of bytes received through RDMA. | Sum | Bytes |
| RdmaTxPackets | Number of packets transmitted through RDMA. | Sum | Count |
| RdmaRxPackets | Number of packets received through RDMA. | Sum | Count |
| InstancePoolSize | Total number of instances in the instance pool. | Sum | Count |
| ProvisioningInstances | Number of instances currently being provisioned. | Sum | Count |
| RunningInstances | Number of running instances in the pool. | Sum | Count |
| TerminatedInstances | Number of terminated instances in the pool. | Sum | Count |
| Total Instances | Monitors the total number of instances associated with the pool. | Average | Count |
Threshold configuration
To configure thresholds for OCI Instance Pool monitor:
- Log in to your Site24x7 account and navigate to Admin > Configuration Profiles > Threshold and Availability.
- Click Add Threshold Profile.
- Select OCI Instance Pool from the Monitor Type drop-down menu.
- Provide an appropriate name in the Display Name field.
- The supported metrics are displayed in the Threshold Configuration section. You can set threshold values for the supported metrics.
- Click Save.
The following custom status notifications are available:
- Notify When Instance Pool Is in Stopped Status: Enable this option to receive an alert when the instance pool enters the Stopped status.
- Notify When Instance Pool Is in Failed Status: Enable this option to receive an alert when the instance pool enters the Failed status.
- Notify When Instance Pool Is in Scaling Status: Enable this option to receive an alert when the instance pool enters the Scaling status.
IT automation
IT automation support is available for OCI Instance Pool monitors. You can configure automation actions to perform predefined operations on the instance pool when specific monitoring conditions or alerts are triggered. The supported action is Start/Stop/Reset/Soft Reset Instance Pool, which helps automate instance pool power operations and reduce manual intervention.
Capacity planning
Capacity monitoring is supported for OCI Instance Pool monitors, helping you track the available and utilized capacity of an instance pool. You can monitor the instance pool size along with the number of provisioning, running, and terminated instances to understand changes in pool capacity. Resource metrics such as CPU utilization, memory utilization, disk activity, and network activity can also be used to monitor resource usage and identify changes in workload or capacity requirements.
Monitor Groups
Monitor Groups can be used to organize OCI Instance Pool monitors and view their overall status from a single location.
You can group Instance Pool monitors based on your application, environment, region, business unit, or other operational requirements. This can help you monitor multiple instance pools without reviewing each monitor individually.
Licensing
Each OCI Instance Pool monitor utilizes one basic monitor license.
Viewing OCI Instance Pool data
To monitor your OCI Instance Pools, log in to the Site24x7 console and navigate to Cloud > OCI > Instance Pool.
Select an Instance Pool to view its monitoring data.
Monitor data
Summary
The Summary tab provides an overview of the OCI Instance Pool monitor and displays its key performance metrics through charts. You can use this view to track resource utilization and identify changes in the performance of the instance pool.
Attached Instances
The Attached Instances tab displays the compute instances associated with the selected Instance Pool. This helps you view the instances currently attached to the pool and review their associated details.
Instance Configuration
The Instance Configuration tab displays the configuration details associated with the Instance Pool. This includes information such as the image ID, boot volume size, subnet ID, and whether a public IP is assigned.
The tab also displays the Autoscaling Configuration, including whether autoscaling is enabled, the autoscaling configuration name, cooldown period, minimum, or maximum.
Autoscaling Policies
The Autoscaling Policies tab is displayed only when schedule-based autoscaling policies are configured for the Instance Pool. You can view details such as the policy name, enabled status, schedule, action, and time zone.
Click an autoscaling policy to view additional details, including the policy ID, policy name, enabled status, schedule, time zone, action, target instance count, minimum instance count, maximum instance count, and creation time. The monitor reference also shows scheduled power and scaling actions configured for the Instance Pool.
Configuration
The Configuration tab displays the configuration and metadata associated with the Instance Pool. Use this tab to review the configuration details retrieved from OCI for the monitored pool.
Outages
The Outages tab displays the outage details associated with the Instance Pool, including the outage start time, end time, duration, and comments, when available.
Notes
The Notes tab displays the notes associated with the Instance Pool monitor.
Log Report
The Log Report tab provides the log details associated with the Instance Pool monitor, allowing you to review the monitor's logged events and status information.
Alert Logs
The Alert Logs tab displays the alerts generated for the Instance Pool monitor. You can use this tab to review alert details and track monitoring conditions that have triggered notifications.
Audit Logs
The Audit Logs tab provides audit information related to actions and changes associated with the Instance Pool monitor.
