Amazon MemoryDB monitoring
Amazon MemoryDB is a durable, in‑memory database service compatible with Valkey and Redis OSS. It provides microsecond reads and single digit millisecond writes with Multi AZ durability.
Integration with Amazon MemoryDB provides visibility into the health, availability, configuration, and performance of your MemoryDB resources. Monitor cluster, shard, and node performance, track memory and CPU utilization, and identify replication and network issues. Monitor cache efficiency and backup status and receive alerts when resource health or performance deviates from expected levels.
Integrating with Amazon MemoryDB provides monitoring for the following monitor types:
- MemoryDB Cluster: Monitors the overall health, availability, configuration, and performance of a MemoryDB cluster.
- MemoryDB Cluster Shard: Monitors individual shards within a MemoryDB cluster, including shard status, node count, and engine CPU utilization.
- MemoryDB Cluster Node: Monitors individual nodes within a shard and provides detailed visibility into resource utilization, database activity, replication, network performance, security, data tiering, and vector search.
- MemoryDB Multi-Region Cluster: Monitors a MemoryDB Multi-Region (global) cluster, including its configuration, status, and the regional clusters that belong to it.
MemoryDB Cluster monitors use one basic monitor license, while MemoryDB Cluster Shard and MemoryDB Multi-Region Cluster monitors are free. MemoryDB Cluster Node monitors use 0.2 basic monitor license per node, with five nodes consuming one basic monitor.
Use cases
Monitor MemoryDB health and availability
Monitor the status of MemoryDB clusters, shards, and nodes to quickly identify availability issues and resource failures. Node and shard status can propagate to the cluster monitor, helping you identify issues at the cluster level.
Identify resource and performance issues
Monitor CPU and engine utilization, memory usage, connections, network traffic, and database activity to identify performance degradation and resource constraints.
Monitor replication and cache performance
Track replication lag, replication throughput, keyspace hits and misses, evictions, and Cache Hit Rate to understand replication health and cache efficiency.
Monitor backup health
Track Time Since Last Backup to identify when the latest manual or automated snapshot is older than the expected backup interval.
Identify security and network issues
Monitor authentication failures, authorization failures, IAM authentication activity, and network allowance metrics to identify security and connectivity related issues.
Monitor Multi-Region cluster health
Track the regional cluster count, available regional clusters, and unavailable regional clusters of a MemoryDB Multi-Region cluster to identify when a regional member degrades.
Benefits of Amazon MemoryDB integration
Integrating with Amazon MemoryDB provides the following benefits:
- Monitor clusters, shards, and nodes from a single console.
- Monitor Multi-Region clusters along with their regional cluster members.
- Track MemoryDB resource and database performance.
- Detect memory, CPU, network, and replication issues.
- Monitor cache efficiency and backup health.
- Track authentication and authorization failures.
- Configure thresholds and receive proactive alerts.
- Use anomaly detection and forecasting for supported metrics.
- Organize MemoryDB resources using Monitor Groups.
- Analyze historical performance trends for capacity planning.
Setup and configuration
Follow the steps below to configure the AWS integration:
- Log in to your Site24x7 account.
- Go to Cloud > AWS > Integrate AWS Account and create a cross-account IAM role to enable Site24x7 to access your AWS resources.
- On the Integrate AWS Account page, select MemoryDB Cluster from the Services to be discovered list based on your requirement.
Permissions
Ensure that Site24x7 receives the following permissions to monitor Amazon MemoryDB:
- cloudwatch:GetMetricData
- memorydb:DescribeClusters
- memorydb:DescribeMultiRegionClusters
- memorydb:DescribeSnapshots
- memorydb:DescribeEvents
- memorydb:DescribeServiceUpdates
- memorydb:DescribeSubnetGroups
- memorydb:ListTags
Polling frequency
Site24x7 queries AWS service-level APIs according to the set polling frequency, from once a minute to once a day, to collect metrics from Amazon MemoryDB monitors.
Supported metrics
The supported metrics for Amazon MemoryDB monitors are given below.
MemoryDB Cluster
The supported metrics for the MemoryDB Cluster monitor are given below.
| Metric name | Description | Statistics | Unit |
|---|---|---|---|
|
Total Number of Shards |
Number of shards in the cluster. |
Configuration |
Count |
|
Total Number of Nodes |
Total number of nodes across all shards. |
Configuration |
Count |
|
CPU Utilization |
Percentage of CPU utilization on the host. |
Average |
Percentage |
|
Engine CPU Utilization |
CPU utilization of the MemoryDB engine thread. |
Average |
Percentage |
|
Freeable Memory |
Amount of free memory available on the host. |
Average |
MB |
|
Database Memory Usage Percentage |
Percentage of available memory currently used for data. |
Maximum |
Percentage |
|
Database Capacity Usage Percentage |
Percentage of total data capacity currently in use. |
Maximum |
Percentage |
|
Swap Usage |
Amount of swap space used on the host. |
Maximum |
MB |
|
Memory Fragmentation Ratio |
Ratio of memory allocated by the operating system to memory requested by the engine. |
Average |
Count |
|
Bytes Used For MemoryDB |
Bytes allocated by MemoryDB for data and overhead. |
Maximum |
Bytes |
|
Network Bytes In |
Number of bytes read by the host from the network. |
Average |
MB |
|
Network Bytes Out |
Number of bytes sent by the host to the network. |
Average |
MB |
|
Network Packets In |
Number of packets received on all network interfaces. |
Sum |
Count |
|
Network Packets Out |
Number of packets sent from all network interfaces. |
Sum |
Count |
|
Current Connections |
Number of currently open client connections. |
Maximum |
Count |
|
New Connections |
Number of connections accepted during the period. |
Sum |
Count |
|
Keyspace Hits |
Number of successful key lookups. |
Sum |
Count |
|
Keyspace Misses |
Number of unsuccessful key lookups. |
Sum |
Count |
|
Current Items |
Number of keys currently held in the cache. |
Sum |
Count |
|
Evictions |
Number of keys evicted because of the maxmemory limit. |
Sum |
Count |
|
Reclaimed |
Number of key expiration events. |
Sum |
Count |
|
Replication Lag |
Time a replica is behind its primary when applying changes. |
Maximum |
Seconds |
|
Replication Bytes |
Number of bytes sent by the primary to its replicas. |
Sum |
Bytes |
|
Non Key Type Commands |
Number of commands that do not operate on a key. |
Sum |
Count |
|
Error Count |
Number of commands that returned an error. |
Sum |
Count |
|
Authentication Failures |
Number of failed authentication attempts. |
Sum |
Count |
|
Cache Hit Rate |
Cache hit percentage calculated by Site24x7 from Keyspace Hits and Keyspace Misses. |
Calculated by Site24x7 |
Percentage |
|
Time Since Last Backup |
Time elapsed since the latest manual or automated snapshot. |
Calculated by Site24x7 |
Hours |
MemoryDB Cluster Shard
The supported metrics for the MemoryDB Cluster Shard monitor are given below.
| Metric name | Description | Statistics | Unit |
|---|---|---|---|
|
Total Number of Nodes |
Number of nodes in the shard, including primary and replica nodes. |
Configuration |
Count |
|
Engine CPU Utilization |
CPU utilization of the MemoryDB engine thread. |
Average |
Percentage |
MemoryDB Cluster Node
The supported metrics for the MemoryDB Cluster Node monitor are given below.
| Metric name | Description | Statistics | Unit |
|---|---|---|---|
|
CPU Utilization |
Percentage of host CPU utilization. |
Average |
Percentage |
|
Engine CPU Utilization |
CPU utilization of the MemoryDB engine thread. |
Average |
Percentage |
|
Freeable Memory |
Amount of free memory available on the host. |
Average |
MB |
|
Database Memory Usage Percentage |
Percentage of available memory currently used for data. |
Maximum |
Percentage |
|
Database Capacity Usage Percentage |
Percentage of total data capacity currently in use. |
Maximum |
Percentage |
|
Swap Usage |
Amount of swap space used on the host. |
Maximum |
MB |
|
Memory Fragmentation Ratio |
Ratio of operating system allocated memory to memory requested by the engine. |
Average |
Count |
|
Bytes Used For MemoryDB |
Bytes allocated by MemoryDB for data and overhead. |
Maximum |
Bytes |
|
Network Bytes In |
Number of bytes read by the host. |
Average |
MB |
|
Network Bytes Out |
Number of bytes sent by the host. |
Average |
MB |
|
Network Packets In |
Number of packets received. |
Sum |
Count |
|
Network Packets Out |
Number of packets sent. |
Sum |
Count |
|
Network Max Bytes In |
Maximum observed inbound network throughput during the period. |
Average |
MB |
|
Network Max Bytes Out |
Maximum observed outbound network throughput during the period. |
Average |
MB |
|
Network Max Packets In |
Maximum observed inbound packets per second during the period. |
Sum |
Count |
|
Network Max Packets Out |
Maximum observed outbound packets per second during the period. |
Sum |
Count |
|
Network Bandwidth In Allowance Exceeded |
Packets shaped or dropped because inbound aggregate bandwidth exceeded the node type maximum. |
Sum |
Count |
|
Network Bandwidth Out Allowance Exceeded |
Packets shaped or dropped because outbound aggregate bandwidth exceeded the node type maximum. |
Sum |
Count |
|
Network Conntrack Allowance Exceeded |
Packets dropped because the number of tracked connections exceeded the node type maximum. |
Sum |
Count |
|
Network Packets Per Second Allowance Exceeded |
Packets dropped because bidirectional packets per second exceeded the node type maximum. |
Sum |
Count |
|
Current Connections |
Number of currently open client connections. |
Maximum |
Count |
|
New Connections |
Number of connections accepted during the period. |
Sum |
Count |
|
Keyspace Hits |
Number of successful key lookups. |
Sum |
Count |
|
Keyspace Misses |
Number of unsuccessful key lookups. |
Sum |
Count |
|
Current Items |
Number of keys held in the cache. |
Sum |
Count |
|
Evictions |
Number of keys evicted because of the maxmemory limit. |
Sum |
Count |
|
Reclaimed |
Number of key expiration events. |
Sum |
Count |
|
Active Defrag Hits |
Number of value reallocations performed by the active defragmentation process. |
Sum |
Count |
|
Keys Tracked |
Number of keys tracked for client side caching invalidation. |
Sum |
Count |
|
Replication Lag |
Time a replica is behind its primary. |
Maximum |
Seconds |
|
Replication Bytes |
Number of bytes sent by the primary to replicas. |
Sum |
Bytes |
|
Replication Delayed Write Commands |
Number of write commands delayed by synchronous replication. |
Sum |
Count |
|
Max Replication Throughput |
Maximum observed replication throughput. |
Maximum |
Bytes |
|
Primary Link Health Status |
Health of the replication link to the primary. |
Minimum |
Count |
|
Is Primary |
Indicates whether the node is a primary or replica. |
Minimum |
Count |
|
Non Key Type Commands |
Number of commands that do not operate on a key. |
Sum |
Count |
|
Error Count |
Number of commands that returned an error. |
Sum |
Count |
|
Authentication Failures |
Number of failed authentication attempts. |
Sum |
Count |
|
Command Authorization Failures |
Commands rejected because the ACL user lacks permission. |
Sum |
Count |
|
Key Authorization Failures |
Commands rejected because the ACL user lacks key permission. |
Sum |
Count |
|
Channel Authorization Failures |
Pub/Sub operations rejected because the ACL user lacks channel permission. |
Sum |
Count |
|
IAM Authentication Expirations |
Number of expired IAM authentication credentials. |
Sum |
Count |
|
IAM Authentication Throttling |
Number of throttled IAM authentication requests. |
Sum |
Count |
|
Search Total Index Size |
Total memory used by vector search indexes. |
Average |
MB |
|
Search Number of Indexes |
Number of vector search indexes on the node. |
Average |
Count |
|
Search Number of Indexed Keys |
Number of keys indexed across vector search indexes. |
Average |
Count |
|
Successful Write Request Latency |
Average latency of successful write requests. |
Average |
ms |
|
Successful Read Request Latency |
Average latency of successful read requests. |
Average |
ms |
|
Bytes Read From Disk |
Bytes read from SSD on data tiering node types. |
Average |
Bytes |
|
Bytes Written To Disk |
Bytes written to SSD on data tiering node types. |
Average |
Bytes |
|
Number Of Items Read From Disk |
Items fetched from SSD instead of memory on data tiering node types. |
Sum |
Count |
|
Number Of Items Written To Disk |
Items moved from memory to SSD on data tiering node types. |
Sum |
Count |
|
DB0 Average TTL |
Average time to live of keys in the DB0 keyspace. |
Average |
ms |
|
Get Type Commands |
Number of read type commands executed. |
Sum |
Count |
|
Set Type Commands |
Number of write type commands executed. |
Sum |
Count |
|
Key Based Commands |
Number of key based commands executed. |
Sum |
Count |
|
String Based Commands |
Number of string commands executed. |
Sum |
Count |
|
Hash Based Commands |
Number of hash commands executed. |
Sum |
Count |
|
List Based Commands |
Number of list commands executed. |
Sum |
Count |
|
Set Based Commands |
Number of set commands executed. |
Sum |
Count |
|
Sorted Set Based Commands |
Number of sorted set commands executed. |
Sum |
Count |
|
Stream Based Commands |
Number of stream commands executed. |
Sum |
Count |
|
PubSub Based Commands |
Number of Pub/Sub commands executed. |
Sum |
Count |
|
Eval Based Commands |
Number of EVAL and EVALSHA scripting commands executed. |
Sum |
Count |
|
Geo Spatial Based Commands |
Number of geospatial commands executed. |
Sum |
Count |
|
HyperLog Log Based Commands |
Number of HyperLogLog commands executed. |
Sum |
Count |
|
Json Based Commands |
Number of JSON commands executed. |
Sum |
Count |
|
Search Based Commands |
Number of vector search commands executed. |
Sum |
Count |
|
Search Based Get Commands |
Number of vector search read commands executed. |
Sum |
Count |
|
Search Based Set Commands |
Number of vector search write commands executed. |
Sum |
Count |
|
Cache Hit Rate |
Cache hit percentage calculated by Site24x7 from Keyspace Hits and Keyspace Misses. |
Calculated by Site24x7 |
Percentage |
MemoryDB Multi-Region Cluster
The MemoryDB Multi-Region Cluster monitor does not publish CloudWatch metrics. All metrics are derived from the Multi-Region cluster configuration on every poll.
| Metric name | Description | Statistics | Unit |
|---|---|---|---|
|
Regional Cluster Count |
Number of regional clusters that belong to the Multi-Region cluster. |
Configuration |
Count |
|
Available Regional Clusters |
Number of regional clusters in an available, updating, modifying, snapshotting, or creating state. |
Configuration |
Count |
|
Unavailable Regional Clusters |
Number of regional clusters in any other state, such as deleting or create failed. |
Configuration |
Count |
Threshold configuration
To configure thresholds for the MemoryDB monitor:
- Log in to your account and navigate to Admin > Configuration Profiles > Threshold and Availability.
- Click Add Threshold Profile.
- Select the applicable MemoryDB monitor type from the Monitor Type drop-down menu and provide an appropriate name in the Display Name field.
- The supported metrics are displayed in the Threshold Configuration section. You can set threshold values for the supported metrics.
- Click Save.
Thresholds for Time Since Last Backup support unit conversion, so you can configure the threshold value in minutes, hours, or days.
Anomaly detection and forecasting
Anomaly detection and forecasting are supported for key metrics such as CPU Utilization, Freeable Memory, Current Connections, Cache Hit Rate, and Replication Lag, and for Engine CPU Utilization at the shard level.
Status based alerts
In addition to metric based thresholds, you can configure status based alerts:
- Notify for any child monitor status changes: For MemoryDB Cluster and MemoryDB Cluster Shard monitors, enable this option to receive an alert at the cluster or shard level when an associated shard or node monitor changes status.
- Primary Link Health Status: For MemoryDB Cluster Node monitors, enable this option to receive an alert when a replica node is not in sync with its primary node.
- Is Primary: For MemoryDB Cluster Node monitors, enable this option to receive an alert when the node's primary role changes.
Status propagation
Shard and node status can propagate to the cluster monitor.
Regional cluster status does not propagate to the MemoryDB Multi-Region Cluster monitor. Configure a threshold for Unavailable Regional Clusters on the Multi-Region Cluster monitor to receive an alert when a regional cluster degrades, while each regional cluster continues to alert on its own.
Licensing
- MemoryDB Cluster monitors use one basic monitor license.
- MemoryDB Cluster Node monitors use 0.2 basic monitor license per node.
- MemoryDB Cluster Shard monitors are free monitors.
- MemoryDB Multi-Region Cluster monitors are free monitors.
- Five MemoryDB Cluster Node monitors consume one basic monitor license.
Organize Amazon MemoryDB using Monitor Groups
Monitor Groups help organize Amazon MemoryDB resources based on applications, environments, business units, or projects.
For example, you can group a MemoryDB cluster with its associated shard and node monitors to monitor the health of the complete MemoryDB environment from one location.
Capacity planning
Capacity planning helps you analyze historical MemoryDB performance trends to understand workload growth and plan future resource requirements.
By monitoring CPU utilization, engine CPU utilization, memory usage, connections, network traffic, replication, and cache performance, you can identify workload trends and plan capacity changes as usage increases.
Viewing Amazon MemoryDB data
To view the MemoryDB Cluster monitor:
- From the Site24x7 console, navigate to Cloud > AWS > MemoryDB Cluster.
To view the MemoryDB Multi-Region Cluster monitor:
- From the Site24x7 console, navigate to Cloud > AWS > MemoryDB Multi Region Cluster.
MemoryDB Multi-Region Cluster monitors are global resources. They are discovered once per account and listed under the Global region.
Monitor data
The monitor data for each Amazon MemoryDB monitor is given below.
MemoryDB Cluster
The monitor data for the MemoryDB Cluster monitor is given below.
Summary
The Summary tab provides an overview of cluster health, availability, CPU utilization, engine CPU utilization, downtime, and performance trends.
Shards
The Shards tab lists the shard monitors associated with the selected cluster and their current status. This tab is displayed when the cluster has more than one shard.
Nodes
The Nodes tab lists the node monitors associated with the selected cluster and their current status.
Topology View
The Topology View tab displays the cluster in the AWS topology map along with its VPC, subnets, availability zones, and its associated shard and node monitors.
Configuration
The Configuration tab displays cluster, security, maintenance, and backup details, including the engine, node type, endpoint, subnet group, parameter group, TLS status, ACL, KMS key, backup settings, maintenance window, and last backup time.
Snapshots
The Snapshots tab displays the most recent automated and manual snapshots. You can view snapshot details, search and sort the table, and export the data as CSV.
Service Updates
The Service Updates tab displays MemoryDB service updates, including their status, update type, release date, auto update start date, and description.
Events
The Events tab displays cluster events for the selected period, including node replacements, service updates, and shard changes.
Zia Forecast
A forecast chart displays future points of a performance metric based on historical time series data. Thirty days of historical data is used to predict what your metric usage will be in the next seven days.
Outages
The Outages tab provides details on an outage's start time, end time, duration, and comments, if any.
Inventory
Obtain details like Resource Name, Monitor Licensing Category, and Check Frequency from the Inventory tab. Set and view the Threshold and Availability Profile and the Notification Profile according to the user in this tab.
Log Report
This tab provides a consolidated report of the MemoryDB Cluster monitor's log status, which can be downloaded as a CSV file.
Alert Logs
This tab displays a chronological list of all triggered alerts related to the MemoryDB Cluster monitor.
MemoryDB Cluster Shard
The monitor data for the MemoryDB Cluster Shard monitor is given below.
Summary
The Summary tab provides an overview of shard status, availability, engine CPU utilization, node count, downtime, and performance trends.
Nodes
The Nodes tab lists the nodes associated with the shard and their current status.
Topology View
The Topology View tab displays the shard in the AWS topology map along with its VPC, subnets, availability zones, parent cluster, and node monitors.
Configuration
The Configuration tab displays the region, cluster name, node group ID, shard status, shard slots, and total number of nodes.
MemoryDB Cluster Node
The monitor data for the MemoryDB Cluster Node monitor is given below.
Summary
The Summary tab provides an overview of node status, availability, CPU utilization, freeable memory, current connections, downtime, and performance trends. It also displays the node status, primary link health, and primary or replica role.
Topology View
The Topology View tab displays the node in the AWS topology map along with its VPC, subnet, availability zone, parent shard, and cluster.
Configuration
The Configuration tab displays the region, cluster name, node name, node status, availability zone, node endpoint, and creation date.
Zia Forecast
A forecast chart displays future points of a performance metric based on historical time series data. Thirty days of historical data is used to predict what your metric usage will be in the next seven days.
Outages
The Outages tab provides details on an outage's start time, end time, duration, and comments, if any.
Inventory
Obtain details like Resource Name, Monitor Licensing Category, and Check Frequency from the Inventory tab. Set and view the Threshold and Availability Profile and the Notification Profile according to the user in this tab.
Log Report
This tab provides a consolidated report of the MemoryDB Cluster Node monitor's log status, which can be downloaded as a CSV file.
Alert Logs
This tab displays a chronological list of all triggered alerts related to the MemoryDB Cluster Node monitor.
MemoryDB Multi-Region Cluster
The monitor data for the MemoryDB Multi-Region Cluster monitor is given below.
Summary
The Summary tab provides an overview of Multi-Region cluster status, availability, regional cluster count, unavailable regional clusters, downtime, and performance trends. It also displays the multi-region cluster name, status, engine, engine version, node type, number of shards, and regions.
Regional Clusters
The Regional Clusters tab lists the MemoryDB Cluster monitors that belong to the Multi-Region cluster, along with their availability summary. Each regional cluster links to its own monitor and retains its own thresholds, alerts, and monitor groups.
Topology View
The Topology View tab displays the Multi-Region cluster and its regional clusters in the AWS topology map.
Configuration
The Configuration tab displays the multi-region cluster name, description, status, ARN, engine name, engine version, node type, number of shards, parameter group, TLS status, regions, and regional cluster count.
Outages
The Outages tab provides details on an outage's start time, end time, duration, and comments, if any.
Inventory
Obtain details like Resource Name, Monitor Licensing Category, and Check Frequency from the Inventory tab. Set and view the Threshold and Availability Profile and the Notification Profile according to the user in this tab.
Log Report
This tab provides a consolidated report of the MemoryDB Multi-Region Cluster monitor's log status, which can be downloaded as a CSV file.
Alert Logs
This tab displays a chronological list of all triggered alerts related to the MemoryDB Multi-Region Cluster monitor.
-
On this page
- Use cases
- Benefits of Amazon MemoryDB integration
- Setup and configuration
- Permissions
- Polling frequency
- Supported metrics
- Threshold configuration
- Anomaly detection and forecasting
- Status based alerts
- Status propagation
- Licensing
- Organize Amazon MemoryDB using Monitor Groups
- Capacity planning
- Viewing Amazon MemoryDB data
- Monitor data
