Help Docs

Amazon MemoryDB monitoring

Amazon MemoryDB is a durable, in‑memory database service compatible with Valkey and Redis OSS. It provides microsecond reads and single digit millisecond writes with Multi AZ durability.

Integration with Amazon MemoryDB provides visibility into the health, availability, configuration, and performance of your MemoryDB resources. Monitor cluster, shard, and node performance, track memory and CPU utilization, and identify replication and network issues. Monitor cache efficiency and backup status and receive alerts when resource health or performance deviates from expected levels.

Integrating with Amazon MemoryDB provides monitoring for the following monitor types:

  • MemoryDB Cluster: Monitors the overall health, availability, configuration, and performance of a MemoryDB cluster.
  • MemoryDB Cluster Shard: Monitors individual shards within a MemoryDB cluster, including shard status, node count, and engine CPU utilization.
  • MemoryDB Cluster Node: Monitors individual nodes within a shard and provides detailed visibility into resource utilization, database activity, replication, network performance, security, data tiering, and vector search.
  • MemoryDB Multi-Region Cluster: Monitors a MemoryDB Multi-Region (global) cluster, including its configuration, status, and the regional clusters that belong to it.

MemoryDB Cluster monitors use one basic monitor license, while MemoryDB Cluster Shard and MemoryDB Multi-Region Cluster monitors are free. MemoryDB Cluster Node monitors use 0.2 basic monitor license per node, with five nodes consuming one basic monitor.

Use cases

Monitor MemoryDB health and availability

Monitor the status of MemoryDB clusters, shards, and nodes to quickly identify availability issues and resource failures. Node and shard status can propagate to the cluster monitor, helping you identify issues at the cluster level.

Identify resource and performance issues

Monitor CPU and engine utilization, memory usage, connections, network traffic, and database activity to identify performance degradation and resource constraints.

Monitor replication and cache performance

Track replication lag, replication throughput, keyspace hits and misses, evictions, and Cache Hit Rate to understand replication health and cache efficiency.

Monitor backup health

Track Time Since Last Backup to identify when the latest manual or automated snapshot is older than the expected backup interval.

Identify security and network issues

Monitor authentication failures, authorization failures, IAM authentication activity, and network allowance metrics to identify security and connectivity related issues.

Monitor Multi-Region cluster health

Track the regional cluster count, available regional clusters, and unavailable regional clusters of a MemoryDB Multi-Region cluster to identify when a regional member degrades.

Benefits of Amazon MemoryDB integration

Integrating with Amazon MemoryDB provides the following benefits:

  • Monitor clusters, shards, and nodes from a single console.
  • Monitor Multi-Region clusters along with their regional cluster members.
  • Track MemoryDB resource and database performance.
  • Detect memory, CPU, network, and replication issues.
  • Monitor cache efficiency and backup health.
  • Track authentication and authorization failures.
  • Configure thresholds and receive proactive alerts.
  • Use anomaly detection and forecasting for supported metrics.
  • Organize MemoryDB resources using Monitor Groups.
  • Analyze historical performance trends for capacity planning.

Setup and configuration

Follow the steps below to configure the AWS integration:

  1. Log in to your Site24x7 account.
  2. Go to Cloud > AWS > Integrate AWS Account and create a cross-account IAM role to enable Site24x7 to access your AWS resources.
  3. On the Integrate AWS Account page, select MemoryDB Cluster from the Services to be discovered list based on your requirement.

Permissions

Ensure that Site24x7 receives the following permissions to monitor Amazon MemoryDB:

  • cloudwatch:GetMetricData
  • memorydb:DescribeClusters
  • memorydb:DescribeMultiRegionClusters
  • memorydb:DescribeSnapshots
  • memorydb:DescribeEvents
  • memorydb:DescribeServiceUpdates
  • memorydb:DescribeSubnetGroups
  • memorydb:ListTags

Polling frequency

Site24x7 queries AWS service-level APIs according to the set polling frequency, from once a minute to once a day, to collect metrics from Amazon MemoryDB monitors.

Supported metrics

The supported metrics for Amazon MemoryDB monitors are given below.

MemoryDB Cluster

The supported metrics for the MemoryDB Cluster monitor are given below.

Metric name Description Statistics Unit

Total Number of Shards

Number of shards in the cluster.

Configuration

Count

Total Number of Nodes

Total number of nodes across all shards.

Configuration

Count

CPU Utilization

Percentage of CPU utilization on the host.

Average

Percentage

Engine CPU Utilization

CPU utilization of the MemoryDB engine thread.

Average

Percentage

Freeable Memory

Amount of free memory available on the host.

Average

MB

Database Memory Usage Percentage

Percentage of available memory currently used for data.

Maximum

Percentage

Database Capacity Usage Percentage

Percentage of total data capacity currently in use.

Maximum

Percentage

Swap Usage

Amount of swap space used on the host.

Maximum

MB

Memory Fragmentation Ratio

Ratio of memory allocated by the operating system to memory requested by the engine.

Average

Count

Bytes Used For MemoryDB

Bytes allocated by MemoryDB for data and overhead.

Maximum

Bytes

Network Bytes In

Number of bytes read by the host from the network.

Average

MB

Network Bytes Out

Number of bytes sent by the host to the network.

Average

MB

Network Packets In

Number of packets received on all network interfaces.

Sum

Count

Network Packets Out

Number of packets sent from all network interfaces.

Sum

Count

Current Connections

Number of currently open client connections.

Maximum

Count

New Connections

Number of connections accepted during the period.

Sum

Count

Keyspace Hits

Number of successful key lookups.

Sum

Count

Keyspace Misses

Number of unsuccessful key lookups.

Sum

Count

Current Items

Number of keys currently held in the cache.

Sum

Count

Evictions

Number of keys evicted because of the maxmemory limit.

Sum

Count

Reclaimed

Number of key expiration events.

Sum

Count

Replication Lag

Time a replica is behind its primary when applying changes.

Maximum

Seconds

Replication Bytes

Number of bytes sent by the primary to its replicas.

Sum

Bytes

Non Key Type Commands

Number of commands that do not operate on a key.

Sum

Count

Error Count

Number of commands that returned an error.

Sum

Count

Authentication Failures

Number of failed authentication attempts.

Sum

Count

Cache Hit Rate

Cache hit percentage calculated by Site24x7 from Keyspace Hits and Keyspace Misses.

Calculated by Site24x7

Percentage

Time Since Last Backup

Time elapsed since the latest manual or automated snapshot.

Calculated by Site24x7

Hours

MemoryDB Cluster Shard

The supported metrics for the MemoryDB Cluster Shard monitor are given below.

Metric name Description Statistics Unit

Total Number of Nodes

Number of nodes in the shard, including primary and replica nodes.

Configuration

Count

Engine CPU Utilization

CPU utilization of the MemoryDB engine thread.

Average

Percentage

MemoryDB Cluster Node

The supported metrics for the MemoryDB Cluster Node monitor are given below.

Metric name Description Statistics Unit

CPU Utilization

Percentage of host CPU utilization.

Average

Percentage

Engine CPU Utilization

CPU utilization of the MemoryDB engine thread.

Average

Percentage

Freeable Memory

Amount of free memory available on the host.

Average

MB

Database Memory Usage Percentage

Percentage of available memory currently used for data.

Maximum

Percentage

Database Capacity Usage Percentage

Percentage of total data capacity currently in use.

Maximum

Percentage

Swap Usage

Amount of swap space used on the host.

Maximum

MB

Memory Fragmentation Ratio

Ratio of operating system allocated memory to memory requested by the engine.

Average

Count

Bytes Used For MemoryDB

Bytes allocated by MemoryDB for data and overhead.

Maximum

Bytes

Network Bytes In

Number of bytes read by the host.

Average

MB

Network Bytes Out

Number of bytes sent by the host.

Average

MB

Network Packets In

Number of packets received.

Sum

Count

Network Packets Out

Number of packets sent.

Sum

Count

Network Max Bytes In

Maximum observed inbound network throughput during the period.

Average

MB

Network Max Bytes Out

Maximum observed outbound network throughput during the period.

Average

MB

Network Max Packets In

Maximum observed inbound packets per second during the period.

Sum

Count

Network Max Packets Out

Maximum observed outbound packets per second during the period.

Sum

Count

Network Bandwidth In Allowance Exceeded

Packets shaped or dropped because inbound aggregate bandwidth exceeded the node type maximum.

Sum

Count

Network Bandwidth Out Allowance Exceeded

Packets shaped or dropped because outbound aggregate bandwidth exceeded the node type maximum.

Sum

Count

Network Conntrack Allowance Exceeded

Packets dropped because the number of tracked connections exceeded the node type maximum.

Sum

Count

Network Packets Per Second Allowance Exceeded

Packets dropped because bidirectional packets per second exceeded the node type maximum.

Sum

Count

Current Connections

Number of currently open client connections.

Maximum

Count

New Connections

Number of connections accepted during the period.

Sum

Count

Keyspace Hits

Number of successful key lookups.

Sum

Count

Keyspace Misses

Number of unsuccessful key lookups.

Sum

Count

Current Items

Number of keys held in the cache.

Sum

Count

Evictions

Number of keys evicted because of the maxmemory limit.

Sum

Count

Reclaimed

Number of key expiration events.

Sum

Count

Active Defrag Hits

Number of value reallocations performed by the active defragmentation process.

Sum

Count

Keys Tracked

Number of keys tracked for client side caching invalidation.

Sum

Count

Replication Lag

Time a replica is behind its primary.

Maximum

Seconds

Replication Bytes

Number of bytes sent by the primary to replicas.

Sum

Bytes

Replication Delayed Write Commands

Number of write commands delayed by synchronous replication.

Sum

Count

Max Replication Throughput

Maximum observed replication throughput.

Maximum

Bytes

Primary Link Health Status

Health of the replication link to the primary.

Minimum

Count

Is Primary

Indicates whether the node is a primary or replica.

Minimum

Count

Non Key Type Commands

Number of commands that do not operate on a key.

Sum

Count

Error Count

Number of commands that returned an error.

Sum

Count

Authentication Failures

Number of failed authentication attempts.

Sum

Count

Command Authorization Failures

Commands rejected because the ACL user lacks permission.

Sum

Count

Key Authorization Failures

Commands rejected because the ACL user lacks key permission.

Sum

Count

Channel Authorization Failures

Pub/Sub operations rejected because the ACL user lacks channel permission.

Sum

Count

IAM Authentication Expirations

Number of expired IAM authentication credentials.

Sum

Count

IAM Authentication Throttling

Number of throttled IAM authentication requests.

Sum

Count

Search Total Index Size

Total memory used by vector search indexes.

Average

MB

Search Number of Indexes

Number of vector search indexes on the node.

Average

Count

Search Number of Indexed Keys

Number of keys indexed across vector search indexes.

Average

Count

Successful Write Request Latency

Average latency of successful write requests.

Average

ms

Successful Read Request Latency

Average latency of successful read requests.

Average

ms

Bytes Read From Disk

Bytes read from SSD on data tiering node types.

Average

Bytes

Bytes Written To Disk

Bytes written to SSD on data tiering node types.

Average

Bytes

Number Of Items Read From Disk

Items fetched from SSD instead of memory on data tiering node types.

Sum

Count

Number Of Items Written To Disk

Items moved from memory to SSD on data tiering node types.

Sum

Count

DB0 Average TTL

Average time to live of keys in the DB0 keyspace.

Average

ms

Get Type Commands

Number of read type commands executed.

Sum

Count

Set Type Commands

Number of write type commands executed.

Sum

Count

Key Based Commands

Number of key based commands executed.

Sum

Count

String Based Commands

Number of string commands executed.

Sum

Count

Hash Based Commands

Number of hash commands executed.

Sum

Count

List Based Commands

Number of list commands executed.

Sum

Count

Set Based Commands

Number of set commands executed.

Sum

Count

Sorted Set Based Commands

Number of sorted set commands executed.

Sum

Count

Stream Based Commands

Number of stream commands executed.

Sum

Count

PubSub Based Commands

Number of Pub/Sub commands executed.

Sum

Count

Eval Based Commands

Number of EVAL and EVALSHA scripting commands executed.

Sum

Count

Geo Spatial Based Commands

Number of geospatial commands executed.

Sum

Count

HyperLog Log Based Commands

Number of HyperLogLog commands executed.

Sum

Count

Json Based Commands

Number of JSON commands executed.

Sum

Count

Search Based Commands

Number of vector search commands executed.

Sum

Count

Search Based Get Commands

Number of vector search read commands executed.

Sum

Count

Search Based Set Commands

Number of vector search write commands executed.

Sum

Count

Cache Hit Rate

Cache hit percentage calculated by Site24x7 from Keyspace Hits and Keyspace Misses.

Calculated by Site24x7

Percentage

MemoryDB Multi-Region Cluster

The MemoryDB Multi-Region Cluster monitor does not publish CloudWatch metrics. All metrics are derived from the Multi-Region cluster configuration on every poll.

Metric name Description Statistics Unit

Regional Cluster Count

Number of regional clusters that belong to the Multi-Region cluster.

Configuration

Count

Available Regional Clusters

Number of regional clusters in an available, updating, modifying, snapshotting, or creating state.

Configuration

Count

Unavailable Regional Clusters

Number of regional clusters in any other state, such as deleting or create failed.

Configuration

Count

Threshold configuration

To configure thresholds for the MemoryDB monitor:

  1. Log in to your account and navigate to Admin > Configuration Profiles > Threshold and Availability.
  2. Click Add Threshold Profile.
  3. Select the applicable MemoryDB monitor type from the Monitor Type drop-down menu and provide an appropriate name in the Display Name field.
  4. The supported metrics are displayed in the Threshold Configuration section. You can set threshold values for the supported metrics.
  5. Click Save.

Thresholds for Time Since Last Backup support unit conversion, so you can configure the threshold value in minutes, hours, or days.

Anomaly detection and forecasting

Anomaly detection and forecasting are supported for key metrics such as CPU Utilization, Freeable Memory, Current Connections, Cache Hit Rate, and Replication Lag, and for Engine CPU Utilization at the shard level.

Status based alerts

In addition to metric based thresholds, you can configure status based alerts:

  • Notify for any child monitor status changes: For MemoryDB Cluster and MemoryDB Cluster Shard monitors, enable this option to receive an alert at the cluster or shard level when an associated shard or node monitor changes status.
  • Primary Link Health Status: For MemoryDB Cluster Node monitors, enable this option to receive an alert when a replica node is not in sync with its primary node.
  • Is Primary: For MemoryDB Cluster Node monitors, enable this option to receive an alert when the node's primary role changes.

Status propagation

Shard and node status can propagate to the cluster monitor.

Regional cluster status does not propagate to the MemoryDB Multi-Region Cluster monitor. Configure a threshold for Unavailable Regional Clusters on the Multi-Region Cluster monitor to receive an alert when a regional cluster degrades, while each regional cluster continues to alert on its own.

Licensing

  • MemoryDB Cluster monitors use one basic monitor license.
  • MemoryDB Cluster Node monitors use 0.2 basic monitor license per node.
  • MemoryDB Cluster Shard monitors are free monitors.
  • MemoryDB Multi-Region Cluster monitors are free monitors.
  • Five MemoryDB Cluster Node monitors consume one basic monitor license.

Organize Amazon MemoryDB using Monitor Groups

Monitor Groups help organize Amazon MemoryDB resources based on applications, environments, business units, or projects.

For example, you can group a MemoryDB cluster with its associated shard and node monitors to monitor the health of the complete MemoryDB environment from one location.

Capacity planning

Capacity planning helps you analyze historical MemoryDB performance trends to understand workload growth and plan future resource requirements.

By monitoring CPU utilization, engine CPU utilization, memory usage, connections, network traffic, replication, and cache performance, you can identify workload trends and plan capacity changes as usage increases.

Viewing Amazon MemoryDB data

To view the MemoryDB Cluster monitor:

  • From the Site24x7 console, navigate to Cloud > AWS > MemoryDB Cluster.

To view the MemoryDB Multi-Region Cluster monitor:

  • From the Site24x7 console, navigate to Cloud > AWS > MemoryDB Multi Region Cluster.
Note

MemoryDB Multi-Region Cluster monitors are global resources. They are discovered once per account and listed under the Global region.

Monitor data

The monitor data for each Amazon MemoryDB monitor is given below.

MemoryDB Cluster

The monitor data for the MemoryDB Cluster monitor is given below.

Summary

The Summary tab provides an overview of cluster health, availability, CPU utilization, engine CPU utilization, downtime, and performance trends.

Shards

The Shards tab lists the shard monitors associated with the selected cluster and their current status. This tab is displayed when the cluster has more than one shard.

Nodes

The Nodes tab lists the node monitors associated with the selected cluster and their current status.

Topology View

The Topology View tab displays the cluster in the AWS topology map along with its VPC, subnets, availability zones, and its associated shard and node monitors.

Configuration

The Configuration tab displays cluster, security, maintenance, and backup details, including the engine, node type, endpoint, subnet group, parameter group, TLS status, ACL, KMS key, backup settings, maintenance window, and last backup time.

Snapshots

The Snapshots tab displays the most recent automated and manual snapshots. You can view snapshot details, search and sort the table, and export the data as CSV.

Service Updates

The Service Updates tab displays MemoryDB service updates, including their status, update type, release date, auto update start date, and description.

Events

The Events tab displays cluster events for the selected period, including node replacements, service updates, and shard changes.

Zia Forecast

A forecast chart displays future points of a performance metric based on historical time series data. Thirty days of historical data is used to predict what your metric usage will be in the next seven days.

Outages

The Outages tab provides details on an outage's start time, end time, duration, and comments, if any.

Inventory

Obtain details like Resource Name, Monitor Licensing Category, and Check Frequency from the Inventory tab. Set and view the Threshold and Availability Profile and the Notification Profile according to the user in this tab.

Log Report

This tab provides a consolidated report of the MemoryDB Cluster monitor's log status, which can be downloaded as a CSV file.

Alert Logs

This tab displays a chronological list of all triggered alerts related to the MemoryDB Cluster monitor.

MemoryDB Cluster Shard

The monitor data for the MemoryDB Cluster Shard monitor is given below.

Summary

The Summary tab provides an overview of shard status, availability, engine CPU utilization, node count, downtime, and performance trends.

Nodes

The Nodes tab lists the nodes associated with the shard and their current status.

Topology View

The Topology View tab displays the shard in the AWS topology map along with its VPC, subnets, availability zones, parent cluster, and node monitors.

Configuration

The Configuration tab displays the region, cluster name, node group ID, shard status, shard slots, and total number of nodes.

MemoryDB Cluster Node

The monitor data for the MemoryDB Cluster Node monitor is given below.

Summary

The Summary tab provides an overview of node status, availability, CPU utilization, freeable memory, current connections, downtime, and performance trends. It also displays the node status, primary link health, and primary or replica role.

Topology View

The Topology View tab displays the node in the AWS topology map along with its VPC, subnet, availability zone, parent shard, and cluster.

Configuration

The Configuration tab displays the region, cluster name, node name, node status, availability zone, node endpoint, and creation date.

Zia Forecast

A forecast chart displays future points of a performance metric based on historical time series data. Thirty days of historical data is used to predict what your metric usage will be in the next seven days.

Outages

The Outages tab provides details on an outage's start time, end time, duration, and comments, if any.

Inventory

Obtain details like Resource Name, Monitor Licensing Category, and Check Frequency from the Inventory tab. Set and view the Threshold and Availability Profile and the Notification Profile according to the user in this tab.

Log Report

This tab provides a consolidated report of the MemoryDB Cluster Node monitor's log status, which can be downloaded as a CSV file.

Alert Logs

This tab displays a chronological list of all triggered alerts related to the MemoryDB Cluster Node monitor.

MemoryDB Multi-Region Cluster

The monitor data for the MemoryDB Multi-Region Cluster monitor is given below.

Summary

The Summary tab provides an overview of Multi-Region cluster status, availability, regional cluster count, unavailable regional clusters, downtime, and performance trends. It also displays the multi-region cluster name, status, engine, engine version, node type, number of shards, and regions.

Regional Clusters

The Regional Clusters tab lists the MemoryDB Cluster monitors that belong to the Multi-Region cluster, along with their availability summary. Each regional cluster links to its own monitor and retains its own thresholds, alerts, and monitor groups.

Topology View

The Topology View tab displays the Multi-Region cluster and its regional clusters in the AWS topology map.

Configuration

The Configuration tab displays the multi-region cluster name, description, status, ARN, engine name, engine version, node type, number of shards, parameter group, TLS status, regions, and regional cluster count.

Outages

The Outages tab provides details on an outage's start time, end time, duration, and comments, if any.

Inventory

Obtain details like Resource Name, Monitor Licensing Category, and Check Frequency from the Inventory tab. Set and view the Threshold and Availability Profile and the Notification Profile according to the user in this tab.

Log Report

This tab provides a consolidated report of the MemoryDB Multi-Region Cluster monitor's log status, which can be downloaded as a CSV file.

Alert Logs

This tab displays a chronological list of all triggered alerts related to the MemoryDB Multi-Region Cluster monitor.

Was this document helpful?

Would you like to help us improve our documents? Tell us what you think we could do better.


We're sorry to hear that you're not satisfied with the document. We'd love to learn what we could do to improve the experience.


Thanks for taking the time to share your feedback. We'll use your feedback to improve our online help resources.

Shortlink has been copied!