Skip to main content

Analyze latency issues in Workload Factory for EDA

Contributors netapp-sineadd

See detected delays and use automated tools to find the cause and fix slow performance in your FSx for ONTAP volumes.

Before you begin

You must have configured latency monitoring before you can view and analyze latency events.

View volumes with latency above threshold

The Volumes with latency above threshold table shows volumes that breached the warning or critical latency threshold in the past 72 hours.

About this task
  • Only the most recent breach for each volume is shown. Earlier breaches for the same volume are not displayed.

  • Events are automatically removed after 72 hours.

  • A maximum of 200 events is shown. Older events are removed as new events are added.

  • Events appear even without a file system link. A link is required to view basic analysis details or run AI analysis.

Steps
  1. Log in using one of the console experiences.

  2. Select the menu The hamburger menu icon and then select EDA.

  3. Select the Latency tab.

  4. Review the information for each event in the table.

  5. To view details for a latency event, select the event in the Severity column. This opens a latency analysis panel for that event.

  6. To sort the table, select any column header. By default, critical events are displayed first sorted by time, followed by warning events sorted by time.

  7. To dismiss one or more events, next to each event select The action menu icon Dismiss.

  8. To add columns to the table, select The column icon, choose the columns, and select Apply.

  9. The latency analysis panel provides a quick view of the breach, including a latency chart and controls for time frame and type. Select View full analysis to open the detailed Latency analysis page with basic analysis, AI analysis (optional), and additional charts. See Analyze latency trends for details.

Analyze a latency event

Before you begin

Configure an Amazon Bedrock model ARN in Workload Factory settings, see Basic GenAI requirements.

Basic analysis helps you quickly find what is causing slowdowns without having to investigate manually.

Latency analysis panel

Select a latency event in the Severity column to open the latency analysis panel for that event. The panel includes tabs that provide different views of the latency event.

If an Amazon Bedrock model ARN is configured, you can run AI analysis on data and cluster scenarios. If Bedrock is not configured, a link to the Storage Workloads configuration page for that file system is provided, where you can set up Bedrock access.

  • Basic: Shows the results of an automatic analysis. If linked to the file system, it displays a breakdown of components and shows which one is causing the delay. If no link exists, you are prompted to add one. An interactive latency graph of CloudWatch metrics for the affected volume is displayed. It shows either read or write metrics, depending on which alarm triggered the event. You can choose different time ranges and use a filter to view All, Read, Write, or Metadata operations.

  • AI analysis: Shows a deeper investigation to determine the specific root cause and potential remediation steps.

  • Detailed latency analysis: Shows interactive graphs of CloudWatch metrics for the affected volume. These graphs display latency, IOPS, and throughput over time. They show either read or write metrics, depending on which alarm triggered the event. You can choose different time ranges and use a filter to view All, Read, Write, or Metadata operations.

For detailed instructions on using the graphs, see Analyze latency trends.

Steps

  1. In the Latency tab, locate the event you want to analyze.

  2. In the Severity column, select a latency event to open an analysis panel for that event.

  3. Review the Basic tab. If a link is connected, open the component breakdown to see where the delay comes from. If no link is connected, select the prompt to link one so you can analyze the components.

  4. In the AI analysis tab, select Analyze to view the results of a deeper investigation, such as possible data or cluster issues. This helps determine the specific root cause and potential remediation steps.

    Review the results including:

    • Potential root cause explanation

    • List of affected EC2 clients

    • Recommended remediation steps

  5. Select View full analysis to see the full results of the AI analysis, including performance charts that show volume latency, IOPS, and throughput from CloudWatch metrics over the selected time period.

  6. Follow the recommended steps to fix the latency issue.

  7. After you make the changes, check the latency events table to confirm the issue is fixed.

Best practices

Consider these recommendations when analyzing latency issues:

  • Monitor trends: Regularly check the Volumes with latency above threshold table to spot patterns or repeated issues that might point to configuration problems.

  • Use AI analysis strategically: Use AI analysis when reviewing data if basic analysis suggests it. AI analysis can reveal deeper insights for complex performance issues that need detailed troubleshooting.

  • Review dismissed events: Regularly review why events were dismissed to see if thresholds should be adjusted or the system improved.

For best practices on analyzing latency trends, see Graph interpretation.