Skip to main content
Well-architected dashboard

Configuration analysis for Amazon FSx for NetApp ONTAP

Contributors netapp-rlithman

The well-architected dashboard in NetApp Console analyzes storage system configurations regularly to determine if any there are any configuration issues with your Amazon FSx for NetApp ONTAP file systems. When issues are found, the dashboard shows you what the issues are and explains what needs to change to ensure your storage systems achieve peak performance, cost efficiency, and compliance with best practices.

Key capabilities include:

  • Daily configuration analysis

  • Automatic best practice validations for NetApp ONTAP storage systems and partner storage providers

  • Proactive observability

  • Insights to action

How it works

Workload Factory analyzes your workloads running on supported cloud storage systems deployments daily. The analysis provides well-architected status, insights, and recommendations.

After the daily analysis completes, configurations appear as "optimized" or "not optimized" in the Well-architected dashboard for the deployment. You'll find configuration issues by category and a list of configuration issues and recommendations. You can review the recommendations for configuration issues. Some issues can be fixed automatically by Workload Factory, while others require manual intervention. In this case, Workload Factory provides detailed instructions to help you implement the recommended changes.

You can dismiss the analysis of configurations that do not apply to your environments. This avoids unnecessary alerts and inaccurate optimization results.

You can also create custom rules in plain language to validate your environments against your organization's standards and NetApp best practices. Test rules with a dry run, then schedule them to run across your environments. Learn more about custom rules.

Why it matters

Workload Factory applies best practices to large storage, database, and VMware environments by combining ongoing assessment with recommendation insights and remediation. Automated fixes applied in the NetApp Console reduce human error, ensure uniform management, and preserve performance and reliability across your workload infrastructures.

Analysis requirements

For a complete file system analysis, you must do the following:

Best practices and recommendations for storage workloads

The well-architected dashboard assesses storage configurations against ONTAP best practices and the AWS Well-Architected Framework. The assessment also recommends improvements and fixes.

The well-architected analysis categorizes configurations in the following pillars of the framework: reliability, security, operational excellence, cost optimization, and performance efficiency.

Reliability

Reliability ensures that workloads perform their intended functions correctly and consistently, even when there are disruptions.

  • Volume backups

    Backing up your volumes helps support data retention and compliance needs. Use volume backups to set up automated backups and retention for your data.

  • Snapshot policy

    Schedule local snapshots for efficient backup and quick restores. Snapshots are instant, point-in-time images of your volumes. This recommendation doesn't apply to database environments.

  • Application-consistent snapshots

    For database environments, use application-consistent snapshots with NetApp SnapCenter to take accurate, reliable snapshots of your volume data at a specific moment in time. This keeps your apps running smoothly and your data safe. SnapCenter makes backups easier and helps you restore data quickly and correctly, reducing downtime and protecting your most important workloads.

  • Cross-region replication

    Cross-region replication ensures that your data is replicated to another AWS region, providing enhanced data durability and availability. Set up cross-region replication to support disaster recovery and compliance.

  • SSD capacity headroom

    The SSD storage tier capacity should not exceed 80% utilization. This might impact data reads and writes to your capacity pool storage tier and impact the throughput capacity of your file system. When capacity runs out, data volumes become read-only. Services trying to write new data also fail.

  • Replication policy labels

    The snapshot policy labels of the source volume and the replication policy labels must match to ensure data reliability.

  • File object capacity

    The file capacity threshold should be raised to avoid hitting the volume capacity limit. Low file capacity (inodes) prevents writing additional data to the volume. Workload Factory recommends maintaining file capacity utilization below 80% to allow new file creation.

Security

Security emphasizes protecting data, systems, and assets through risk assessments and mitigation strategies.

  • Ransomware protection (ARP/AI)

    NetApp Autonomous Ransomware Protection with AI (ARP/AI) helps protect your volumes from ransomware threats. Workload Factory recommends enabling ARP/AI for all volumes.

  • Unauthorized iSCSI volume access

    Volumes serving application data using iSCSI should not allow NAS access at the same time. Workload Factory recommends restricting iSCSI volumes from using any other protocol.

Operational excellence

Operational excellence focuses on delivering the most optimal architecture and business value.

  • Automatic capacity management

    Automatic capacity management should be enabled to regularly ensure that the SSD tier doesn't exceed the threshold.

  • Volume capacity headroom

    Workload Factory recommends that volume capacity doesn't exceed 80% utilization. This might impact data reads and writes to your application. Volume capacity increases can be manual or automatic using the volume autogrow feature.

  • FlexCache write mode

    For optimal performance, Workload Factory recommends the cache relationship write mode that best suits your workload. Write-around mode provides better performance for read-heavy workloads with small files, whereas write-back mode provides better performance for write-heavy workloads with large files.

  • FlexCache replica size

    Workload Factory recommends enabling volume autosize and scrubbing on cache volumes to maintain optimal size and focus the cache on hot data for peak efficiency.

  • Host-visible capacity reporting

    Workload Factory recommends enabling host-visible capacity reporting to provide better visibility into storage usage at the volume level.

  • Operating System (OS) type

    Workload Factory recommends ensuring that the ONTAP LUN operating system (OS) type value matches the operating system partitioning scheme to achieve I/O alignment. Incorrect configuration might reduce performance.

  • Block device space management

    Workload Factory recommends configuring block device space settings for LUNs used by database host instances to prevent write failures and improve space efficiency on FSx for ONTAP. This configuration applies to the recommended combination of settings for thin-provisioned volumes: space reservation, space allocation, and fractional reserve.

  • Thin provisioning

    Workload Factory recommends configuring thin provisioning for FSx for ONTAP volumes hosting database instances. This approach allows more logical data to be stored than physically available.

Cost optimization

Cost optimization helps you get the most value for your business while keeping costs low.

  • Cold data tiering

    Enable cold data tiering to reduce SSD storage tier utilization. Workload Factory recommends applying a tiering policy to every volume focused on general data storage. For database volumes, we'll recommend enabling cold data tiering where appropriate. FSx for ONTAP continuously scans the data for cold data and moves it to the capacity pool tier without disruption.

  • Storage efficiencies

    Enable storage efficiencies — compaction, compression, and deduplication — to optimize storage utilization and reduce the SSD tier cost.

  • Backup expiration

    Workload Factory recommends deleting any manually created backups for volumes that were deleted, and any backups older than two weeks. These backups don't expire automatically, can increase costs, and add complexity if you no longer need them for SLA or recovery.

  • Snapshot expiration

    Workload Factory recommends deleting old or manually created snapshots you no longer need. Removing them can lower costs and reduce clutter, especially if they aren't required for your SLA or recovery needs.

  • Inactive block devices

    After a block device isn't used for seven days, Workload Factory recommends archiving block device data or deleting the unused block device to reduce costs.

  • Overprovisioned SSD capacity

    An overprovisioned SSD capacity tier adds unnecessary costs. Workload Factory recommends reducing the SSD tier capacity while maintaining a 20% free space buffer. A decrease operation takes several hours to several days, and other file system operations are unsupported during this time.

  • Inactive NAS volumes

    The configuration analysis identifies NAS volumes that are not actively used and recommends deleting or archiving the volumes to reduce costs.

Performance efficiency

Performance efficiency helps you ensure that your storage resources are used effectively and that workloads perform optimally.

  • FlexGroup volumes rebalance

    Workload Factory recommends rebalancing your FlexGroup volumes to improve efficiency. After scaling out, capacity and workload activity can become uneven across the volume.

  • FlexVol volumes rebalance

    Workload Factory recommends rebalancing your FlexVol volumes to improve efficiency. After scaling out, capacity and workload activity can become uneven across the volume.