Skip to main content
NetApp artificial intelligence solutions
본 한국어 번역은 사용자 편의를 위해 제공되는 기계 번역입니다. 영어 버전과 한국어 버전이 서로 어긋나는 경우에는 언제나 영어 버전이 우선합니다.

8. 구성 참조

기여자 nkarthik

Karthikeyan Nagalingam, NetApp

구성 참조에는 데이터 수집, 데이터 준비, XCP 이동성, 학습, 체크포인트 및 아카이빙을 제어하는 런타임 매개변수가 설명되어 있습니다. ONTAP NAS 버킷은 주로 지원 및 이식성을 위해 준비된 데이터와 XCP 스테이징 계층으로 사용됩니다. 실제 운영 환경에서는 워크로드 및 지연 시간 요구 사항에 더 적합한 경우 동일한 준비된 데이터를 RDMA 지원 직접 액세스 경로를 통해 사용할 수도 있습니다.


모든 런타임 동작은 `dag_run.conf`를 통해 제공되므로 동일한 DAG가 서로 다른 소스 시스템, 변환 엔진, 스토리지 대상, 복구 정책 및 아카이브 범위를 지원할 수 있습니다. 선택한 워크플로에 필요한 파라미터 그룹만 구성하고, 자격 증명은 Airflow Connections 또는 비밀 백엔드에 보관하며, 예제는 자격 증명 값이 아닌 비프로덕션 템플릿으로 사용하십시오.

[[8-1-ingestion-parameters]]
== 8.1 수집 매개변수

이 설정을 사용하여 업스트림 데이터 수집 메커니즘을 선택하고 구성하십시오. `s3_direct`는 데이터 준비 그룹에 구성된 StorageGRID 원시 데이터 객체를 사용합니다. Airbyte 및 NiFi 설정은 해당 수집 툴이 선택된 경우에만 필요합니다.

매개변수 기본값 필수 설명 예

ingestion_tool

s3_direct

아니요

수집 모드를 선택합니다

"s3_direct"

airbyte_api_url

None

Airbyte 전용

Airbyte API 엔드포인트

"http://airbyte:8000"

airbyte_connection_id

None

Airbyte 전용

Airbyte 연결 ID

"connection-123"

airbyte_api_token

None

선택 사항

Bearer 토큰

"<TOKEN>"

airbyte_timeout_secs

30

아니요

Airbyte 타임아웃

60

nifi_api_url

None

NiFi 전용

NiFi API 엔드포인트

"http://nifi:8080"

nifi_process_group_id

None

NiFi 전용

NiFi 프로세스 그룹

"abc123"

nifi_timeout_secs

30

아니요

NiFi 타임아웃

60

[[8-2-data-preparation-parameters]]
== 8.2 데이터 준비 매개변수

이 설정은 원시 데이터 액세스, 준비된 데이터 출력, 입력 검색 및 변환 동작을 제어합니다. s3_raw_* 매개변수는 StorageGRID를 참조하고, 코드 호환 s3_formatted_* 매개변수는 준비되고 실행 스탬프가 찍힌 데이터와 매니페스트가 기록되는 ONTAP NAS 버킷을 식별합니다.

Raw StorageGRID / prepared-data ONTAP NAS 버킷:

StorageGRID와 ONTAP NAS 버킷이 서로 다른 S3 호환 서비스 또는 액세스 정책을 사용하는 경우 별도의 엔드포인트와 자격 증명을 구성하십시오. 일반적인 대체 키는 보다 구체적인 원시 데이터 또는 준비된 데이터 설정이 없는 경우에만 적용됩니다.

매개변수 기본값 설명 예

s3_raw_bucket

ai-raw-data

StorageGRID 원시 입력 버킷

"bucket1"

s3_raw_prefix

example_ai_pipeline/raw

StorageGRID 원시 데이터 접두사

"example_ai_pipeline/raw"

s3_raw_endpoint_url

폴백 체인

StorageGRID S3 엔드포인트

"http://10.63.150.62:10444"

s3_raw_access_key_id

폴백 체인

원시 액세스 키

"<KEY>"

s3_raw_secret_access_key

폴백 체인

원시 비밀 키

"<SECRET>"

s3_raw_region

aws_region

원시 영역

"us-east-1"

s3_formatted_bucket

ai-formatted-data

ONTAP NAS 준비 데이터 출력 버킷

"prepnasbucket"

s3_formatted_prefix

example_ai_pipeline/formatted

ONTAP NAS 준비 데이터 접두사

"example_ai_pipeline/formatted"

s3_formatted_endpoint_url

폴백 체인

ONTAP NAS 버킷 S3 엔드포인트

"http://10.63.150.159"

s3_formatted_access_key_id

폴백 체인

ONTAP NAS 버킷 액세스 키

"<KEY>"

s3_formatted_secret_access_key

폴백 체인

ONTAP NAS 버킷 비밀 키

"<SECRET>"

일반적인 대체값: s3_endpoint_url, s3_access_key_id, s3_secret_access_key, s3_session_token, s3_verify, aws_region, aws_access_key_id, aws_secret_access_key, aws_session_token.

입력 선택:

알려진 CSV 입력을 처리하려면 명시적인 객체 키 설정을 사용하십시오. 그렇지 않으면 작업이 구성된 StorageGRID 원시 접두사 아래에서 적합한 테이블 형식 객체를 검색합니다.

매개변수 기본값 설명 예

s3_tabular_object_keys

자동 검색

CSV 키의 명시적 목록

["raw/demo/a.csv"]

s3_tabular_object_key

None

단일 명시적 CSV 키

"raw/demo/a.csv"

변환:

경량 로컬 준비를 위한 Python 경로 또는 분산 준비를 위한 Spark 경로를 선택하십시오. 샘플링 및 시드는 생성된 학습, 검증 및 추론 분할에 재현 가능하게 적용됩니다.

매개변수 기본값 설명 예

transformation_engine

python

python 또는 spark

"spark"

spark_required

false

폴백 대신 실패

true

table_format

none

none/delta/iceberg

"delta"

table_format_required

false

형식을 사용할 수 없는 경우 실패

true

sample_count

모든 행

행 제한 (≤0 = 모두)

14400

seed

7

재현 가능한 셔플 시드

11

Spark:

이 설정은 `transformation_engine=spark`에만 적용됩니다. Spark 프로세스 배치, 리소스 할당, 시간 제한 및 선택적 Delta Lake 또는 Iceberg 런타임 종속성을 제어합니다.

매개변수 기본값 설명 예

spark_submit_bin

spark-submit

Spark 바이너리

"/opt/spark/bin/spark-submit"

spark_master

local[*]

Spark 마스터

"local[*]"

spark_driver_memory

4g

드라이버 메모리

"8g"

spark_executor_memory

4g

실행기 메모리

"8g"

spark_timeout_secs

900/7200

타임아웃 (0 = 비활성화)

0

spark_delta_packages

Delta 3.2.0

델타 런타임 패키지

"io.delta:delta-spark_2.12:3.2.0"

spark_iceberg_packages

Iceberg 1.5.2

Iceberg 런타임 패키지

"org.apache.iceberg:iceberg-spark-runtime-3.5_2.12:1.5.2"

spark_iceberg_catalog

local

Iceberg 카탈로그 이름

"local"

[[8-3-xcp-and-training-destination-parameters]]
== 8.3 XCP 및 교육 대상 매개변수

이 그룹은 `enable_xcp=true`을 사용하여 ONTAP NAS NFS 소스를 검증하고, XCP 호스트를 호출하고, 활성 교육 계층을 선택할 때 사용합니다. `xcp_copy_destination`는 제어 스위치 역할을 하며, 아래의 ONTAP S3 관련 설정 또는 LustreFS 관련 설정을 선택하고, 다운스트림 모델 단계에서는 동일하게 선택된 계층을 사용합니다.

매개변수 기본값 필수 설명 예

enable_xcp

false

네, XCP의 경우 그렇습니다.

XCP 분기 활성화

true

xcp_copy_destination

s3

아니요

s3 또는 lustrefs

"lustrefs"

xcp_nfs_source_ip

10.63.150.159

XCP preflight

NFS 서버 IP

"10.63.150.159"

xcp_nfs_source_path

/prepnasbucket/example_ai_pipeline

XCP preflight

NFS 내보내기 경로

"/prepnasbucket/example_ai_pipeline"

ssh_user

root

아니요

XCP 호스트 SSH 사용자

"root"

ssh_host

10.63.150.178

아니요

XCP 호스트 주소

"10.63.150.178"

S3 대상(필수 조건 xcp_copy_destination=s3):

이 설정은 ONTAP S3 교육 경로에만 제공하십시오. 프로필 또는 직접 자격 증명을 사용하면 XCP와 모델 아티팩트 동기화 로직이 선택한 ONTAP S3 버킷 및 접두사에 액세스할 수 있습니다.

매개변수 기본값 설명 예

xcp_s3_endpoint

기본

XCP S3 엔드포인트

"http://10.63.150.161"

xcp_s3_bucket

None

XCP 대상 버킷

"trainingbucket"

xcp_s3_dest_prefix

example_ai_pipeline

목적지 접두사

"example_ai_pipeline"

xcp_s3_training_prefix

xcp_s3_dest_prefix

교육 경로 검색 접두사

"example_ai_pipeline"

xcp_s3_profile

ontaps3

명명된 자격 증명 프로필

"ontaps3"

xcp_s3_profiles

None

프로필 맵

{"ontaps3": {…​}}

xcp_s3_access_key_id

None

직접 키 대체

"<KEY>"

xcp_s3_secret_access_key

None

직접 비밀 폴백

"<SECRET>"

xcp_s3_region

aws_region

XCP S3 지역

"us-east-1"

LustreFS 대상 (필수 조건 xcp_copy_destination=lustrefs):

이 설정은 LustreFS 교육 경로에만 제공하십시오. 소스 및 대상 경로는 XCP 호스트에 마운트되어 접근 가능해야 하며, 대상 경로는 모델 단계를 실행하는 Airflow 워커에서도 접근 가능해야 합니다.

매개변수 기본값 설명 예

xcp_lustrefs_source_path

/prepnasbucket/example_ai_pipeline

XCP 복사 소스

"/prepnasbucket/example_ai_pipeline"

xcp_lustrefs_dest_path

/mnt/lustre/client

LustreFS 마운트 대상

"/mnt/lustre/client"

xcp_lustrefs_training_subpath

xcp_s3_dest_prefix

마운트 아래의 교육 하위 디렉토리

"example_ai_pipeline"

xcp_lustrefs_newid

`data_prep`런 스탬프

XCP 작업 식별자

"20260901_120000"

텍스트 소스 동작:

이 옵션은 XCP 이동성 이후 세 가지 텍스트 데이터 세트를 찾는 방법을 제어합니다. 명시적 키는 검색을 재정의하며, 로컬 대체 기능을 비활성화하면 텍스트 입력이 선택한 XCP 대상에서 제공되어야 합니다.

매개변수 기본값 설명 예

xcp_text_allow_local_fallback

true

로컬 텍스트 대체 허용

false

xcp_text_base_key

자동

명시적인 기본 텍스트 위치

"…​/text_base.json"

xcp_text_finetune_key

자동

명시적 미세 조정 텍스트 위치

"…​/text_finetune.json"

xcp_text_infer_key

자동

명시적 추론 - 텍스트 위치

"…​/text_infer.json"

[[8-4-model-training-parameters]]
== 8.4 모델 학습 매개변수

이 설정은 로컬 또는 XCP를 통해 제공되는 소스를 선택하고, 학습 용량을 제한하며, 재현 가능한 데이터 순서를 유지합니다. XCP가 활성화되면 학습은 선택된 ONTAP S3 또는 LustreFS 계층에서 데이터를 읽고 기준선 결과물을 동일한 대상으로 다시 게시합니다.

매개변수 기본값 설명 예

enable_xcp

false

XCP 또는 로컬 소스를 선택합니다

true

xcp_copy_destination

s3

소스 티어를 선택합니다

"lustrefs"

sample_count

모든 행

교육 샘플 제한

14400

seed

7

재현 가능한 셔플

11

생성 항목: regression_model.bin, text_vectorizer.bin, text_classifier.bin, train_metrics.json.

[[8-5-fine-tuning-parameters]]
== 8.5 미세 조정 매개변수

미세 조정은 선택한 XCP 대상에서 기준 텍스트 아티팩트를 구체화하고, 미세 조정 데이터 세트에 증분 학습을 적용한 다음, 조정된 분류기와 해당 지표를 다시 게시합니다. 아티팩트의 출처를 유지하려면 XCP 설정을 모델 학습과 일관되게 유지해야 합니다.

매개변수 기본값 설명 예

enable_xcp

false

아티팩트 구체화/게시를 활성화합니다.

true

xcp_copy_destination

s3

아티팩트 소스/대상을 선택합니다.

"lustrefs"

소비: text_vectorizer.bin, text_classifier.bin, text_finetune.json. 생산: text_classifier_tuned.bin, fine_tune_metrics.json.

[[8-6-inference-parameters]]
== 8.6 추론 매개변수

이 플래그는 선택한 학습 계층에서 튜닝된 모델이 생성된 후의 스코어링 범위를 제어합니다. 검증 또는 비즈니스 요구 사항에 따라 기본 추론 분할 및 추론 텍스트 입력 범위를 넘어서는 출력이 필요한 경우 두 옵션 중 하나를 활성화하십시오.

매개변수 기본값 설명 예

infer_all_splits

false

점수 학습/검증/추론 분할

true

infer_all_text

false

점수 기반/세부 조정/텍스트 추론

true

소비: regression_model.bin, text_vectorizer.bin, text_classifier_tuned.bin. 생성: tabular_predictions.csv, text_predictions.json.

[[8-7-checkpoint-parameters]]
== 8.7 체크포인트 매개변수

체크포인트 기능은 `model_training`으로 범위가 지정됩니다. 이러한 설정은 학습 완료 기록 생성 여부, 저장 위치, 업로드 실패 처리 방식, 그리고 후속 실행 시 유효한 학습 완료 결과를 검증하거나 재사용할지 여부를 결정합니다.

매개변수 기본값 설명 예

checkpoint_enabled

true

체크포인트 파일 생성을 활성화합니다.

true

checkpoint_store

local

local 또는 formatted_s3

"formatted_s3"

checkpoint_upload_fail_mode

warn

warn 또는 fail

"fail"

training_checkpoint_reuse_mode

off

off/verify_only/resume_if_exists

"resume_if_exists"

checkpoint_reuse_mode

off

하위 호환성 별칭

"verify_only"

[[8-8-manual-archive-parameters]]
== 8.8 수동 아카이브 매개변수

이 그룹을 사용하여 StorageGRID 아카이빙을 활성화하고, 대상 버킷 및 자격 증명을 지정하고, 아카이빙 지점을 선택할 수 있습니다. 선택한 `manual_archive_stage`아카이빙 지점은 아카이빙 시점과 아카이빙에 기준 학습 아티팩트만 포함할지 또는 전체 추론 결과 세트를 포함할지를 제어합니다.

매개변수 기본값 필수 설명 예

manual_archive_enabled

false

아니요

아카이빙을 활성화합니다

true

manual_archive_stage

inferencing

아니요

보관 지점 및 아티팩트 세트: model_training 핵심 학습 아티팩트용 또는 inferencing 전체 결과 세트용

"model_training"

manual_archive_bucket

archivalbucket

활성화된 경우 예

아카이브 버킷

"archivalbucket"

manual_archive_prefix

ai-models

아니요

아카이브 접두사

"ai-models"

manual_archive_endpoint

기본

맞춤 제작만 가능

S3 엔드포인트 아카이브

"http://10.63.150.62:10444"

manual_archive_profile

sgdlocal

권장

자격 증명 프로필

"sgdlocal"

manual_archive_profiles

None

선택 사항

프로필 맵

{"sgdlocal": {…​}}

manual_archive_include_xcp_inputs

false

아니요

아카이브에 XCP 입력 포함

true

manual_archive_cleanup_raw

false

아니요

업로드 후 로컬 raw 폴더 제거

true

별칭: fabricpool_archive_* 모든 manual_archive_* 매개변수에 대해.

[[8-8-1-manual-archive-stage-decision]]
=== 8.8.1 수동 아카이브 단계 결정

The route_manual_archive_stage 분기 작업은 manual_archive_stage DAG 실행당 한 번씩 읽습니다. 선택된 값은 아카이빙 실행 시점과 StorageGRID 아카이빙 버킷에 업로드되는 아티팩트를 모두 결정합니다. 아카이브 오브젝트는 s3://<manual_archive_bucket>/<manual_archive_prefix>/<stage>/<run_stamp>/ 아래에 저장됩니다.

manual_archive_stage 값 아카이브 타이밍 아티팩트 아카이브됨 권장 사용법

model_training

model_training 직후에 실행되며, 미세 조정이나 추론을 기다리지 않습니다.

regression_model.bin, text_vectorizer.bin, text_classifier.bin, train_metrics.json

재현 가능한 기본 구성 모델을 유지하거나, 아카이브 용량을 최소화하거나, 이후 단계 전에 체크포인트를 보존합니다.

inferencing (기본값)

fine_tuning 및 inferencing 완료 후

모든 model_training 아티팩트 및 text_classifier_tuned.bin, fine_tune_metrics.json, tabular_predictions.csv, text_predictions.json

모델 실행에 대한 모든 비즈니스 결과 및 예측 증거를 보관합니다.

두 가지 선택 사항 중 하나를 선택하더라도, `manual_archive_include_xcp_inputs=true`를 설정하여 로컬에 구체화된 XCP 학습 입력이 있는 경우 추가합니다. `manual_archive_enabled=false`를 설정하여 학습 및 추론 워크플로를 변경하지 않고 아카이빙 작업을 건너뜁니다.

[[8-9-complete-sample-configuration]]
== 8.9 전체 샘플 구성

이 카탈로그는 주요 배포 및 테스트 경로에 대한 시작 구성을 제공합니다. 모든 예제는 원시 데이터가 StorageGRID에 있고 준비된 데이터가 ONTAP NAS 버킷에 기록된다고 가정합니다. 버킷 이름, 엔드포인트 주소, 파일 시스템 경로 및 프로필 이름을 대상 환경의 값으로 교체하십시오. 프로덕션 액세스 키 또는 비밀 키를 `CONF_JSON`에 저장하지 마십시오. Airflow Connections, Variables 또는 비밀 키 백엔드를 통해 구성하십시오.

[[8-9-1-xcp-to-lustrefs-with-spark-and-delta-lake]]
=== 8.9.1 XCP에서 Spark 및 Delta Lake를 사용한 LustreFS로의 변환

이 최대 처리량 예제를 사용하여 분산 준비 및 고성능 LustreFS 모델 I/O를 구현하십시오. 이 예제는 s3_raw_prefix 아래의 모든 적합한 테이블 형식 파일을 처리하고, 준비 중에 Delta 테이블을 기록하며, 학습 체크포인트를 ONTAP NAS 준비 데이터 버킷에 저장하고, 전체 추론 결과 세트를 아카이브합니다.

{
  "ingestion_tool": "s3_direct",
  "s3_raw_bucket": "bucket1",
  "s3_raw_prefix": "example_ai_pipeline/raw",
  "s3_formatted_bucket": "prepnasbucket",
  "s3_formatted_prefix": "example_ai_pipeline/formatted",

  "enable_xcp": true,
  "xcp_copy_destination": "lustrefs",
  "xcp_nfs_source_ip": "10.63.150.159",
  "xcp_nfs_source_path": "/prepnasbucket/example_ai_pipeline",
  "xcp_lustrefs_source_path": "/prepnasbucket/example_ai_pipeline",
  "xcp_lustrefs_dest_path": "/mnt/lustre/client",
  "xcp_lustrefs_training_subpath": "example_ai_pipeline",
  "xcp_text_allow_local_fallback": false,

  "transformation_engine": "spark",
  "spark_required": true,
  "table_format": "delta",
  "table_format_required": true,
  "spark_driver_memory": "8g",
  "spark_executor_memory": "8g",
  "spark_timeout_secs": 0,
  "sample_count": 14400,
  "seed": 11,

  "infer_all_splits": true,
  "infer_all_text": true,

  "checkpoint_enabled": true,
  "checkpoint_store": "formatted_s3",
  "training_checkpoint_reuse_mode": "off",

  "manual_archive_enabled": true,
  "manual_archive_stage": "inferencing",
  "manual_archive_bucket": "archivalbucket",
  "manual_archive_prefix": "ai-models",
  "manual_archive_profile": "sgdlocal",
  "manual_archive_cleanup_raw": true
}

[[8-9-2-no-xcp-with-python-preparation]]
=== 8.9.2 Python 준비 없이 XCP 사용

이 기본 구성을 간단한 기능 실행에 사용합니다. 준비 단계에서는 로컬 Python 엔진을 사용하고, 후속 단계에서는 파이프라인의 로컬 준비 데이터 및 아티팩트 경로를 사용합니다. XCP, Spark, 테이블 형식, 체크포인트 재사용 및 아카이빙은 비활성화되어 있습니다.

{
    "ingestion_tool": "s3_direct",
    "s3_raw_bucket": "bucket1",
    "s3_raw_prefix": "example_ai_pipeline/raw",
    "s3_formatted_bucket": "prepnasbucket",
    "s3_formatted_prefix": "example_ai_pipeline/formatted",
    "enable_xcp": false,
    "transformation_engine": "python",
    "table_format": "none",
    "sample_count": 200,
    "seed": 11,
    "checkpoint_enabled": false,
    "manual_archive_enabled": false
}

[[8-9-3-python-preparation-with-two-explicit-input-files]]
=== 8.9.3 두 개의 명시적인 입력 파일을 사용한 Python 준비

알려진 두 개의 StorageGRID CSV 오브젝트를 대상으로 파이프라인을 검증할 때 이 집중 테스트를 사용하십시오. s3_tabular_object_keys 자동 표 형식 파일 검색이 비활성화되지만, 텍스트 입력은 여전히 원시 접두사 아래의 예상 위치를 사용합니다.

{
    "ingestion_tool": "s3_direct",
    "s3_raw_bucket": "bucket1",
    "s3_raw_prefix": "example_ai_pipeline/raw",
    "s3_tabular_object_keys": [
        "example_ai_pipeline/raw/tabular_part_01.csv",
        "example_ai_pipeline/raw/tabular_part_02.csv"
    ],
    "s3_formatted_bucket": "prepnasbucket",
    "s3_formatted_prefix": "example_ai_pipeline/formatted",
    "enable_xcp": false,
    "transformation_engine": "python",
    "table_format": "none",
    "sample_count": 200,
    "seed": 11,
    "manual_archive_enabled": false
}

[[8-9-4-python-preparation-with-automatic-all-file-discovery]]
=== 8.9.4 자동 전체 파일 검색 기능을 갖춘 Python 준비

`s3++_++tabular++_++object++_++keys` 및  `s3++_++tabular++_++object++_++key`를 생략하면 StorageGRID 원시 데이터 접두사 아래에 있는 모든 적격 테이블 형식 오브젝트를 처리합니다. 이 예제는 Python 준비 엔진을 사용하는 모든 파일에 대한 기능 또는 확장성 테스트에 적합합니다.
{
    "s3_raw_bucket": "bucket1",
    "s3_raw_prefix": "example_ai_pipeline/raw",
    "s3_formatted_bucket": "prepnasbucket",
    "s3_formatted_prefix": "example_ai_pipeline/formatted",
    "enable_xcp": false,
    "transformation_engine": "python",
    "table_format": "none",
    "sample_count": 0,
    "seed": 11,
    "manual_archive_enabled": false
}

sample_count: 0 데이터 준비 과정에서 사용 가능한 모든 행을 사용한다는 의미입니다. 제한된 테스트 실행을 위해 양수 값을 설정하십시오.

[[8-9-5-xcp-to-ontap-s3-with-python-preparation]]
=== 8.9.5 Python을 사용하여 ONTAP S3로 XCP 준비

준비된 데이터를 ONTAP NAS NFS 내보내기에서 ONTAP S3 학습 버킷으로 복사해야 하는 경우 이 구성을 사용하십시오. 모델 학습, 미세 조정 및 추론 결과는 동일한 ONTAP S3 대상을 통해 생성 및 게시됩니다.

{
    "s3_raw_bucket": "bucket1",
    "s3_raw_prefix": "example_ai_pipeline/raw",
    "s3_formatted_bucket": "prepnasbucket",
    "s3_formatted_prefix": "example_ai_pipeline/formatted",
    "enable_xcp": true,
    "xcp_copy_destination": "s3",
    "xcp_nfs_source_ip": "10.63.150.159",
    "xcp_nfs_source_path": "/prepnasbucket/example_ai_pipeline",
    "xcp_s3_endpoint": "https://ontap-s3.example.com",
    "xcp_s3_bucket": "trainingbucket",
    "xcp_s3_dest_prefix": "example_ai_pipeline",
    "xcp_s3_training_prefix": "example_ai_pipeline",
    "xcp_s3_profile": "ontaps3",
    "transformation_engine": "python",
    "table_format": "none",
    "sample_count": 14400,
    "seed": 11,
    "checkpoint_enabled": true,
    "checkpoint_store": "formatted_s3",
    "manual_archive_enabled": true,
    "manual_archive_stage": "model_training",
    "manual_archive_bucket": "archivalbucket",
    "manual_archive_profile": "storagegrid-archive"
}

[[8-9-6-xcp-to-ontap-s3-with-spark-and-iceberg]]
=== 8.9.6 Spark 및 Iceberg를 사용한 XCP에서 ONTAP S3로의 변환

Iceberg 테이블 출력 및 ONTAP S3를 활성 학습 대상으로 사용하는 분산 준비에 이 구성을 사용하십시오. Spark 또는 Iceberg를 사용할 수 없을 때 대체 실행 대신 실패해야 하는 경우 Set spark_required 및 `table_format_required`을 `true`으로 설정하십시오.

{
    "s3_raw_bucket": "bucket1",
    "s3_raw_prefix": "example_ai_pipeline/raw",
    "s3_formatted_bucket": "prepnasbucket",
    "s3_formatted_prefix": "example_ai_pipeline/formatted",
    "enable_xcp": true,
    "xcp_copy_destination": "s3",
    "xcp_nfs_source_ip": "10.63.150.159",
    "xcp_nfs_source_path": "/prepnasbucket/example_ai_pipeline",
    "xcp_s3_bucket": "trainingbucket",
    "xcp_s3_dest_prefix": "example_ai_pipeline",
    "xcp_s3_profile": "ontaps3",
    "transformation_engine": "spark",
    "spark_required": true,
    "table_format": "iceberg",
    "table_format_required": true,
    "spark_driver_memory": "8g",
    "spark_executor_memory": "8g",
    "spark_timeout_secs": 0,
    "sample_count": 0,
    "seed": 11,
    "manual_archive_enabled": true,
    "manual_archive_stage": "inferencing",
    "manual_archive_bucket": "archivalbucket",
    "manual_archive_profile": "storagegrid-archive"
}

[[8-9-7-scenario-selection-summary]]
=== 8.9.7 시나리오 선택 요약

대본 XCP 대상 준비 엔진 표 형식 입력 범위 아카이브 단계

8.9.1

LustreFS

불꽃

Delta

발견된 모든 파일

inferencing

8.9.2

None

파이썬

None

발견된 모든 파일

비활성화됨

8.9.3

None

파이썬

None

두 개의 명시적인 파일

비활성화됨

8.9.4

None

파이썬

None

발견된 모든 파일

비활성화됨

8.9.5

ONTAP S3

파이썬

None

발견된 모든 파일

model_training

8.9.6

ONTAP S3

불꽃

Iceberg

발견된 모든 파일

inferencing

[[8-9-8-execution-walkthroughs]]
=== 8.9.8 실행 연습

Airflow 워크스페이스에서 예제를 실행합니다. 트리거 스크립트는 고유한 manual__<UTC timestamp> 실행 ID를 생성하고 완료될 때까지 기다린 후, 실패하거나 시간 초과된 경우 0이 아닌 종료 코드를 반환합니다. 다음 전체 구성은 운영 예제와 동일한 CONF_JSON="$(python3 - <<'PY' …​)" 패턴을 사용합니다. 실행하기 전에 참조된 Airflow 관리 연결, XCP 프로필 및 아카이브 프로필을 시크릿 백엔드에서 구성하고, CONF_JSON 시크릿이 아닌 버킷, 엔드포인트, 경로 및 프로필 식별자로 제한합니다.

유효한 테이블 형식 이름은 `iceberg`입니다. 지원되지 않는 값인 `icerberg`을 사용하지 마십시오.

모든 파일, Spark, Iceberg, XCP를 ONTAP S3로, 체크포인트를 prepared-data S3로, 추론 후 아카이브:

CONF_JSON="$(python3 - <<'PY'
import json

print(json.dumps({
    "ingestion_tool": "s3_direct",
    "enable_xcp": True,
    "xcp_copy_destination": "s3",
    "xcp_text_allow_local_fallback": True,

    "s3_raw_bucket": "bucket1",
    "s3_raw_prefix": "example_ai_pipeline/raw",
    "s3_raw_endpoint_url": "https://storagegrid.example.com",
    "s3_raw_region": "us-east-1",

    "s3_formatted_bucket": "prepnasbucket",
    "s3_formatted_prefix": "example_ai_pipeline/formatted",
    "s3_formatted_endpoint_url": "https://ontap-nas.example.com",
    "s3_formatted_region": "us-east-1",

    "xcp_nfs_source_ip": "10.63.150.159",
    "xcp_nfs_source_path": "/prepnasbucket/example_ai_pipeline",
    "xcp_s3_endpoint": "https://ontap-s3.example.com",
    "xcp_s3_bucket": "trainingbucket",
    "xcp_s3_dest_prefix": "example_ai_pipeline",
    "xcp_s3_training_prefix": "example_ai_pipeline",
    "xcp_s3_profile": "ontaps3",

    "transformation_engine": "spark",
    "spark_required": True,
    "table_format": "iceberg",
    "table_format_required": True,
    "sample_count": 0,
    "seed": 11,
    "spark_driver_memory": "8g",
    "spark_executor_memory": "8g",
    "spark_timeout_secs": 0,
    "infer_all_splits": True,
    "infer_all_text": True,

    "checkpoint_enabled": True,
    "checkpoint_store": "formatted_s3",
    "manual_archive_enabled": True,
    "manual_archive_stage": "inferencing",
    "manual_archive_bucket": "archivalbucket",
    "manual_archive_prefix": "ai-models",
    "manual_archive_profile": "sgdlocal",
    "manual_archive_endpoint": "https://storagegrid.example.com",
    "manual_archive_cleanup_raw": True,
    "manual_archive_include_xcp_inputs": True,
}, separators=(",", ":")))
PY
)"
CONF_JSON="$CONF_JSON" AIRFLOW_BIN=./.venv/bin/airflow INGESTION_TOOL=s3_direct \
  ./trigger_and_wait_ai_pipeline_sklearn.sh

모든 파일, Spark, Iceberg, XCP to LustreFS, checkpoint to prepared-data S3, 추론 후 아카이브:

CONF_JSON="$(python3 - <<'PY'
import json

print(json.dumps({
    "ingestion_tool": "s3_direct",
    "enable_xcp": True,
    "xcp_copy_destination": "lustrefs",
    "ssh_user": "root",
    "ssh_host": "10.63.150.178",
    "xcp_text_allow_local_fallback": True,

    "s3_raw_bucket": "bucket1",
    "s3_raw_prefix": "example_ai_pipeline/raw",
    "s3_raw_endpoint_url": "https://storagegrid.example.com",
    "s3_raw_region": "us-east-1",

    "s3_formatted_bucket": "prepnasbucket",
    "s3_formatted_prefix": "example_ai_pipeline/formatted",
    "s3_formatted_endpoint_url": "https://ontap-nas.example.com",
    "s3_formatted_region": "us-east-1",

    "xcp_nfs_source_ip": "10.63.150.159",
    "xcp_nfs_source_path": "/prepnasbucket/example_ai_pipeline",
    "xcp_lustrefs_source_path": "/prepnasbucket/example_ai_pipeline",
    "xcp_lustrefs_dest_path": "/mnt/lustre/client",
    "xcp_lustrefs_training_subpath": "example_ai_pipeline",

    "transformation_engine": "spark",
    "spark_required": True,
    "table_format": "iceberg",
    "table_format_required": True,
    "sample_count": 0,
    "seed": 11,
    "spark_driver_memory": "8g",
    "spark_executor_memory": "8g",
    "spark_timeout_secs": 0,
    "infer_all_splits": True,
    "infer_all_text": True,

    "checkpoint_enabled": True,
    "checkpoint_store": "formatted_s3",
    "manual_archive_enabled": True,
    "manual_archive_stage": "inferencing",
    "manual_archive_bucket": "archivalbucket",
    "manual_archive_prefix": "ai-models",
    "manual_archive_profile": "sgdlocal",
    "manual_archive_endpoint": "https://storagegrid.example.com",
    "manual_archive_cleanup_raw": True,
    "manual_archive_include_xcp_inputs": True,
}, separators=(",", ":")))
PY
)"
AIRFLOW_BIN=./.venv/bin/airflow INGESTION_TOOL=s3_direct \
  ./trigger_and_wait_ai_pipeline_sklearn.sh

ONTAP S3 모드의 경우, 비밀이 아닌 xcp_s3_profile 이름만 CONF_JSON`에 전달하고 실행 전에 XCP 호스트 또는 Airflow에서 관리하는 비밀 저장소를 통해 해당 자격 증명을 정의하십시오. LustreFS 모드의 경우, 학습 경로가 ONTAP S3 자격 증명을 확인하지 않으므로 `xcp_s3_profiles 맵을 의도적으로 생략합니다. 두 예제 모두에서 StorageGRID 원시 데이터, ONTAP NAS 준비 데이터 및 아카이브 자격 증명을 인라인 JSON이 아닌 Airflow에서 관리하는 비밀 저장소를 통해 정의하십시오. `s3_tabular_object_keys`를 생략하고 `sample_count`를 `0`으로 설정하면 모든 적합한 테이블 형식 파일의 자동 검색 및 사용 가능한 모든 행 사용이 요청됩니다.

Python, XCP 미사용, 두 파일 테스트:

CONF_JSON='{"s3_raw_bucket":"bucket1","s3_raw_prefix":"example_ai_pipeline/raw","s3_tabular_object_keys":["example_ai_pipeline/raw/tabular_part_01.csv","example_ai_pipeline/raw/tabular_part_02.csv"],"s3_formatted_bucket":"prepnasbucket","s3_formatted_prefix":"example_ai_pipeline/formatted","enable_xcp":false,"transformation_engine":"python","table_format":"none","sample_count":200,"manual_archive_enabled":false}' \
./trigger_and_wait_ai_pipeline_sklearn.sh

Spark, Delta Lake, XCP에서 LustreFS로:

CONF_JSON='{"enable_xcp":true,"xcp_copy_destination":"lustrefs","xcp_nfs_source_ip":"10.63.150.159","xcp_nfs_source_path":"/prepnasbucket/example_ai_pipeline","xcp_lustrefs_source_path":"/prepnasbucket/example_ai_pipeline","xcp_lustrefs_dest_path":"/mnt/lustre/client","xcp_lustrefs_training_subpath":"example_ai_pipeline","transformation_engine":"spark","spark_required":true,"table_format":"delta","table_format_required":true,"sample_count":14400,"manual_archive_enabled":true,"manual_archive_stage":"inferencing"}' \
./trigger_and_wait_ai_pipeline_sklearn.sh

Spark, Iceberg, XCP에서 ONTAP S3로:

CONF_JSON='{"enable_xcp":true,"xcp_copy_destination":"s3","xcp_nfs_source_ip":"10.63.150.159","xcp_nfs_source_path":"/prepnasbucket/example_ai_pipeline","xcp_s3_bucket":"trainingbucket","xcp_s3_dest_prefix":"example_ai_pipeline","xcp_s3_profile":"ontaps3","transformation_engine":"spark","spark_required":true,"table_format":"iceberg","table_format_required":true,"sample_count":0,"manual_archive_enabled":true,"manual_archive_stage":"inferencing"}' \
./trigger_and_wait_ai_pipeline_sklearn.sh

XCP를 실행하기 전에 구성된 ONTAP NAS NFS 경로와 대상에 XCP 호스트에서 접근 가능한지 확인하십시오. 실행이 성공적으로 완료되면 작업 로그에서 [data_prep] effective_config, [xcp_copy] effective_config, [model_training] effective_config, 및 `[manual_archive]`를 검토하여 선택한 엔진, 입력 범위, 대상 및 아카이브 단계를 확인하십시오.