Skip to main content
Data Infrastructure Insights

Troubleshoot the ONTAP SVM Data Collector

Contributors netapp-alavoie

Use this guide to identify whether a collector problem is caused by management connectivity, FPolicy callback connectivity, ONTAP configuration, Agent capacity, identity resolution, or a feature prerequisite.

This document has two parts. Start with the symptom-based checks. Use the advanced diagnostics at the end only if those checks do not resolve the problem, or if NetApp Support asks for more evidence.

Note: For configuration steps, permissions, and feature prerequisites, see Configuring the ONTAP SVM Data Collector.

Start here

  1. In Workload Security, select Collectors. In the Status column, open More detail and note the complete reason string.

  2. On the add or edit collector page, select Test Connection. Resolve each failed check before saving.

  3. Choose the matching symptom below and follow the checks in order.

Open More detail to see the complete collector error.

More detail link for a collector in Error status

Note: For full validation, connect using Cluster management IP and SVM name. SVM mode cannot run feature and RBAC checks. Test Connection exercises representative FPolicy ports, not every reserved port.

Understand Test Connection results

Result What it validates Where to continue if it fails

HTTPS

Agent can reach ONTAP management on TCP 443.

Cannot add the collector or Test Connection fails

ONTAP Version

Credentials work; ONTAP version and feature eligibility can be read.

Cannot add the collector or Test Connection fails

Data LIFs

An eligible, operational data LIF exists; ONTAP 9.8+ includes data-fpolicy-client.

Verify the data LIF and service policy

Agent IP

The Agent has a local address that can route to the SVM data LIFs.

FPolicy callback connectivity

FPolicy Server

ONTAP can connect back to the Agent on live FPolicy ports.

FPolicy callback connectivity

Features

The account and ONTAP version support selected optional features. Cluster mode only.

Cannot add the collector or Test Connection fails

Note: Test Connection does not validate the entire callback port reservation, real file-event flow, MetroCluster or SVM-DR behavior, Persistent Store aggregate availability, MAV rules, share and volume filter syntax, or User Directory collectors.

Choose your symptom

  • I cannot add the collector, or Test Connection fails

  • Collector status is Error

  • Collector status is Degraded

  • Collector is Running, but no activity appears

  • Activity shows a SID instead of a username

  • ONTAP performance changed after enabling Workload Security

  • Capacity or subscription usage looks wrong

  • Snapshot or user-blocking actions fail

  • The collector state or missing event might be expected behavior

I cannot add the collector, or Test Connection fails

Check these in order

  1. Confirm the connection mode. Cluster mode requires the Cluster management IP and the exact, case-sensitive SVM name. SVM mode requires the SVM management IP.

  2. From the Agent, confirm HTTPS reachability to the configured management IP on TCP 443. A timeout indicates a routing or firewall problem.

  3. Confirm the credentials and applications assigned to the ONTAP login. Grant custom-user roles directly to an AD user; group-level roles might not be visible to the permission check.

  4. Confirm an eligible SVM data LIF is operational. From ONTAP 9.8, its service policy must include data-fpolicy-client with data-nfs and/or data-cifs.

  5. If SVM IP and vsadmin are used, use a management-only LIF. A LIF with combined Data and Management roles can respond to ping while management access still fails.

  6. For optional features, grant the privileges reported by the Features test, then run Test Connection again.

Exact errors covered

  • Failed to determine ONTAP type for [host]. Reason: Connection error to Storage System: Host is unreachable

  • [IP] is identified as a cluster/node and cannot be added in this mode

  • No valid data interface …​ found on the SVM

  • Missing Permission: vserver fpolicy

  • Roles assigned at the group level instead of the user level for this Active Directory user

Resolved when

  • Every required Test Connection check passes and the collector can be saved. Feature checks can remain unavailable only for features you do not use.

Collector status is Error

Use the reason shown in More detail. Most Error states belong to one of the following groups.

FPolicy callback connectivity

What this means: ONTAP cannot maintain the connection that sends file and user activity to the Agent.

  1. Run Test Connection and inspect FPolicy Server, Agent IP, and Data LIFs.

  2. Allow the reserved TCP ports within 35000–55000 from every SVM data LIF to the Agent, including the Agent host firewall. Each SVM uses up to four ports: two per enabled protocol.

  3. Confirm the Agent is routable from each data-serving node and SVM data LIF.

  4. Confirm only one collector in one Workload Security environment monitors the SVM. A second collector replaces the first FPolicy destination.

  5. On ONTAP 9.8 and later, confirm data-fpolicy-client is assigned to an operational SVM data LIF.

Exact errors covered

  • External fpolicy server terminated

  • Node failed to establish a connection with the FPolicy server …​ Select Timed out

  • No local IP address found on the connector that can reach the data interfaces of the SVM

Resolved when

  • Test Connection reports FPolicy Server success and the collector remains Running after client activity is generated.

FPolicy configuration

  1. For shares-to-include errors, enter complete share names without quotation marks. For long lists, filter by volume instead.

  2. For User is not authorized, add the feature-specific custom-user privileges and restart the collector.

  3. For no valid data interface, bring an eligible data LIF up and correct its service policy.

  4. For Persistent Store aggregate information not available, wait a few minutes and restart the collector. Persistent Store requires ONTAP 9.14.1 or later.

  5. For no valid sequence number, have an ONTAP administrator remove confirmed unused FPolicy objects, then restart the collector.

Exact errors covered

  • Failed to configure fpolicy on SVM …​ Invalid value specified for shares-to-include

  • Failed to configure fpolicy on SVM …​ User is not authorized

  • No valid sequence number available to enable fpolicy policy

  • Failed to configure persistent store …​ Performance information for aggregate is currently not available

Resolved when

  • The cloudsecure_ FPolicy policy is enabled and the collector remains Running.

Agent capacity or collector health

  1. Re-enter the collector password: open Edit, enter the password, and save.

  2. Check Agent CPU and memory headroom and the number of hosted collectors.

  3. Use the Event Rate Checker to compare peak event rate with Agent sizing. Scale the Agent or migrate collectors when required.

  4. If the action still fails, restart the Agent service and retry after the Agent reconnects.

Exact errors covered

  • External fpolicy server overloaded

  • Agent failed to connect to the collector

  • AGENT004 — collector stopped immediately after starting

  • AGENT008 — failed to determine collector health

  • AGENT005 / AGENT006 / AGENT007 / AGENT009 / AGENT010

Resolved when

  • The Agent is Connected, the collector reaches Running, and it stays healthy during peak activity.

Collector status is Degraded

A Degraded collector is still running, but one or more FPolicy node connections are disconnected. Activity served by affected nodes might be missed.

  1. Open More detail and identify the disconnected node.

  2. Check whether that node has an operational data LIF for the monitored SVM.

  3. If the node has no local data LIF for the SVM, no remediation is required.

  4. If a data LIF exists, verify that it is up, carries data-fpolicy-client on ONTAP 9.8+, and can reach the Agent callback ports.

  5. Allow the collector to return to Running after the node reconnects.

Exact errors covered

  • FPolicy server disconnected on node(s): …​

  • No local lif present to connect to FPolicy server

Resolved when

  • All data-serving nodes have connected FPolicy channels. A node without a local data LIF does not need a connection.

Collector is Running, but no activity appears

  1. Generate real SMB or NFS client activity on a monitored share or volume. No client I/O produces no events.

  2. Confirm the collector is not Paused. Resume it before testing.

  3. Confirm the client protocol is enabled and, for SMB, that the SVM has a CIFS server.

  4. Review included and excluded shares and volumes. Use complete names without quotation marks.

  5. Enable Monitor Folder Access when general folder-access events are required. Folder create, rename, and delete are collected without this option.

  6. If only .ini and .DS_Store activity is missing, no action is required; these extensions are excluded from FPolicy event collection.

  7. If activity is still absent, use the advanced diagnostics to verify the cloudsecure_ policy and the ONTAP FPolicy event log.

Resolved when

  • New client operations appear in Activity Forensics with the expected protocol, share, path, and user information.

Activity shows a SID instead of a username

ONTAP auditing is working; identity resolution is the separate problem.

  1. Confirm a Running User Directory Collector exists for every domain whose users access the monitored SVM, including trusted domains.

  2. Confirm the directory collector covers the correct forest or search base and uses valid bind credentials.

  3. Confirm the mapped attributes: Active Directory uses name, objectsid, and sAMAccountName; LDAP uses name, uidnumber, and uid.

  4. Restart the User Directory Collector to request an immediate synchronization after adding users or correcting configuration.

  5. Use the AD or LDAP collector troubleshooting page for connection and bind errors. Test Connection is not available for User Directory collectors.

Resolved when

  • New activity displays the expected user name instead of the raw SID or UID.

ONTAP performance changed after enabling Workload Security

  1. Use ONTAP 9.13.1 or later for fixes to known FPolicy latency issues.

  2. Check peak activity and Agent sizing with the Event Rate Checker. Move collectors or scale the Agent if it cannot sustain the event rate.

  3. Review whether Monitor Folder Access materially increased event volume.

  4. For supported ONTAP releases, consider Persistent Store to buffer events during temporary connectivity interruptions.

  5. For a node panic or persistent storage latency, collect the ONTAP evidence and contact NetApp Support.

Resolved when

  • ONTAP latency and IOPS remain within the expected range while FPolicy collection is active.

Capacity or subscription usage looks wrong

  1. Compare raw capacity reported by Observability ONTAP cluster collectors and Workload Security SVM collectors.

  2. For each SVM collector, determine whether its parent cluster is already monitored by an Observability collector. The same storage should not be counted twice.

  3. Check for a recent SVM migration. Confirm the collector is Running and restart it so current cluster and capacity information is collected.

  4. If the total still includes both cluster and SVM capacity, collect before-and-after usage and the reported capacity for each collector.

Resolved when

  • Subscription usage includes the storage capacity once and reflects the current parent cluster of the SVM.

Snapshot or user-blocking actions fail

  1. Confirm the collector is Running and not Paused.

  2. For user blocking, use cluster-level credentials. A custom user also needs SSH access on TCP 22 and the documented blocking privileges.

  3. Confirm the custom user has the required snapshot privileges.

  4. If ONTAP Multi-Admin Verify is enabled, add the documented exclusions for cloudsecure_ snapshots and the set operation used by user blocking.

  5. Run Test Connection in Cluster mode and review the Snapshot and User Blocking feature results.

Resolved when

  • Workload Security can create and delete its snapshots and can block and restore user access.

Expected behavior — no action required

What you see Why it happens When to investigate

MetroCluster standby collector is Stopped

Use one Cluster-mode collector for source and one for destination. Only the active SVM collector runs. Allow up to two minutes after switchover.

The active-side collector does not become Running.

Degraded lists a node with no local data LIF

ONTAP cannot establish an FPolicy channel from a node that does not serve the SVM locally.

The listed node has an operational SVM data LIF.

No events for .ini or .DS_Store

These extensions are excluded from the FPolicy scope.

Other expected file activity is also missing.

Brief Initializing, Stopped, or Degraded state

Restart, upgrade, migration, failover, and resume require reconnection.

The state persists beyond the operation or repeats.

Advanced diagnostics and Support data collection

Use this part only when the checks above do not resolve the problem, or when NetApp Support requests more evidence. These procedures require administrative access to the Agent host and to ONTAP.

Collect evidence before changing the environment

  1. Record the complete collector status reason and the time of failure.

  2. Save the Test Connection result.

  3. Record Agent and collector versions, ONTAP version, SVM name, connection mode, and recent changes.

  4. Generate the Agent symptom bundle once so logs from before your changes are preserved.

sudo /opt/netapp/cloudsecure/agent/bin/cloudsecure-agent-symptom-collector.sh -o /tmp

The command creates cloudsecure-agent-symptoms.zip. Install zip if the script reports that the zip command is not found.

Log locations

Evidence Location or command

Collector log

/opt/netapp/cloudsecure/data-collectors//logs/dsc.log

Contains all collector activity, including errors. Use this log to investigate collector problems.

Agent log

/opt/netapp/cloudsecure/agent/logs/agent.log

Agent installation and upgrade

/opt/netapp/cloudsecure/agent/logs/cloudsecure_agent_install.log

/opt/netapp/cloudsecure/agent/logs/cloudsecure_agent_upgrade.log

Test Connection

/opt/netapp/cloudsecure/test-connection/logs/

Agent service

systemctl status cloudsecure-agent.service

journalctl -u cloudsecure-agent.service

Note: The collector log directory can also contain error.log, but it is not a general error log. It is written only when a collector cannot start because of an invalid configuration, and the Agent removes it after reporting that reason in the collector status. It is normally absent or empty, so use dsc.log.

Validate ONTAP FPolicy state

Run these commands from the appropriate ONTAP administrative context:

event log show -source fpolicy
event log show -source fpolicy -fields event,action,description
fpolicy show
fpolicy show-engine

Look for a Workload Security policy with the cloudsecure_ prefix. The policy should be enabled, and every data-serving node should show a connected external engine. Preserve the event reason and node name when a channel is disconnected.

Validate each network path

Direction Port Purpose

Agent → Cluster or SVM management IP

TCP 443

Configure and query ONTAP.

SVM data LIFs → Agent

Reserved ports within TCP 35000–55000

FPolicy file and user activity. Up to four ports per SVM: two per enabled protocol.

Cluster management IP → Agent

Reserved ports within TCP 35000–55000

EMS-based events, including ARP integrations.

Agent → Cluster management IP

TCP 22

SMB user blocking with cluster credentials.

Note: Test Connection validates live representative callback ports. It does not prove that every port in the reserved firewall range is open.

Agent to ONTAP management connectivity

curl -kv https://<management-ip>:443

A TLS response proves the TCP path. A timeout indicates routing or firewall blockage. An authentication response proves reachability but not that the credentials or roles are correct.

ONTAP to Agent callback connectivity

From ONTAP, test routing from each relevant SVM data LIF:

network ping -vserver <svm> -lif <data-lif> -destination <agent-ip> -show-detail

On the Agent, inspect the host firewall:

sudo firewall-cmd --zone=public --list-ports
sudo iptables-save

Confirm the reserved callback ports are allowed inbound from all SVM data LIFs and, for EMS-based features, from the cluster management IP.

Verify the data LIF and service policy

From ONTAP 9.8 and later, an operational SVM data LIF must include data-fpolicy-client with data-nfs and/or data-cifs.

network interface show -vserver <svm> -fields service-policy,status-admin,status-oper

If a suitable policy does not exist, create or modify one according to your ONTAP network design. For example:

net int service-policy create -policy only_data_fpolicy -vserver <svm> \

Note: On ONTAP earlier than 9.8, data-fpolicy-client is not required. The LIF must have role data, be operational, and support NFS and/or CIFS.

Use an ONTAP packet trace

Use a packet trace only after Test Connection and the routing and firewall checks do not explain a callback failure.

  1. Start a packet trace on ONTAP for the relevant data LIF and Agent IP.

  2. Attempt Test Connection or restart the collector.

  3. Wait for the failure, then stop the trace.

  4. Retrieve the trace from https:///spi//etc/log/packet_traces/.

  5. Look for a SYN from ONTAP to the Agent callback port. No SYN indicates an ONTAP-side path or policy problem. A SYN without a completed handshake indicates a firewall or routing problem between ONTAP and the Agent.

Note: Packet traces can contain network information. Handle them according to your data-handling requirements.

Diagnose duplicate SVM collectors

Two collectors cannot monitor the same SVM, even when they belong to different Workload Security environments. The most recent collector rewrites the FPolicy destination and disconnects the previous collector.

  1. Search every Workload Security environment that can reach the cluster.

  2. Identify collectors that use the same SVM name or management endpoint.

  3. Keep one collector and remove the duplicates.

  4. Restart the retained collector and confirm that its FPolicy channel stays connected.

Clean up confirmed unused FPolicy objects

Note: Do not delete cloudsecure_ objects for an active collector. The collector owns and updates those objects.

List the FPolicy configuration and identify objects that are confirmed unused:

fpolicy show

For an unused policy, remove its objects in dependency order:

fpolicy disable -vserver <svm> -policy-name <policy-name>
fpolicy policy scope delete -vserver <svm> -policy-name <policy-name>
fpolicy policy delete -vserver <svm> -policy-name <policy-name>
fpolicy policy event delete -vserver <svm> -event-name <event-name>
fpolicy policy external-engine delete -vserver <svm> -engine-name <engine-name>

Restart the Workload Security collector after sequence slots or conflicting objects are cleared.

Diagnose Agent and collector lifecycle errors

Error Advanced checks

AGENT004

Inspect dsc.log for an immediate process exit, permission failure, or invalid configuration.

AGENT008

Re-enter the collector password; check Agent CPU and memory and the event rate; inspect dsc.log for the underlying collector failure.

AGENT005 / 006 / 007 / 010

Retry after a few minutes. If the action still fails, restart cloudsecure-agent.service and record the exact action and reason.

AGENT009

Confirm the collector still exists in the selected Agent and environment and that a concurrent delete or migration did not remove it.

Agent NOT_CONNECTED

Check systemctl and journalctl, SaaS egress on TCP 443, proxy configuration, SSL inspection, and the cssys account.

Use the Event Rate Checker to size the Agent. Do not rely on collector count alone; event volume and enabled features affect load.

Investigate no activity after the basic checks

  1. Run event log show -source fpolicy and preserve any error.

  2. Run fpolicy show and confirm that a cloudsecure_ policy is enabled.

  3. Confirm the protocol generating client operations is enabled in the collector and allowed on the SVM.

  4. Confirm the share or volume is not excluded.

  5. Confirm the collector was not Paused when the operation occurred.

  6. Correlate the client operation timestamp with dsc.log and the ONTAP FPolicy events.

Investigate identity resolution

Test Connection is not available for User Directory collectors. Use directory-native tools and the collector log.

Directory Default ports Expected attributes

Active Directory

389 LDAP / 636 LDAPS

Display name: name

ID: objectsid

User name: sAMAccountName

LDAP

389 LDAP / 636 LDAPS

Display name: name

ID: uidnumber

User name: uid

Validate the server name, port, bind DN, password, forest or search base, and read access. Restart the directory collector after a correction to request a fresh synchronization.

Multi-Admin Verify checks

Multi-Admin Verify can block the ONTAP commands used for snapshots and user blocking. Review the existing rules before changing them. The documented Workload Security exclusions are:

multi-admin-verify rule modify -operation "volume snapshot create" -query "-snapshot !*cloudsecure_*"
multi-admin-verify rule modify -operation "volume snapshot delete" -query "-snapshot !*cloudsecure\_*"
multi-admin-verify rule delete -operation set

Note: Deleting the set rule changes Multi-Admin Verify protection for that operation. Review the change with your ONTAP security administrator.

Contact NetApp Support

Contact Support when an error persists after its ordered checks, ONTAP reports a node panic, or capacity remains incorrect after the collector data is refreshed.

Include

  • The complete message from Status > More detail, and the time of failure.

  • Test Connection results, including skipped checks.

  • Agent and collector versions and their current states.

  • ONTAP version, cluster and SVM name, connection mode, and recent topology changes.

  • cloudsecure-agent-symptoms.zip and the relevant dsc.log.

  • Output of fpolicy show, fpolicy show-engine, and event log show -source fpolicy.

  • A packet trace when a callback path remains unexplained.

  • For performance: the latency and IOPS timeline and the event-rate timeline.

  • For capacity: before-and-after subscription usage and the raw capacity reported by each collector.

Use Help > Support in Data Infrastructure Insights to open a case.