Skip to content

Latest commit

Β 

History

History
326 lines (223 loc) Β· 18.1 KB

File metadata and controls

326 lines (223 loc) Β· 18.1 KB

torch.export νλ¦„μ˜ μ„€λͺ…, 일반적인 문제점과 이듀을 ν•΄κ²°ν•˜κΈ° μœ„ν•œ ν•΄κ²°μ±…

μ €μž: Ankith Gunapal, Jordi Ramon, Marcos Carranza λ²ˆμ—­: μ΄ν˜„μ€€

Introduction to torch.export Tutorial μ—μ„œ, torch.export λ₯Ό μ‚¬μš©ν•˜λŠ” 방법을 λ°°μ› μŠ΅λ‹ˆλ‹€. 이 νŠœν† λ¦¬μ–Όμ€ 이전 νŠœν† λ¦¬μ–Όμ„ ν™•μž₯ν•˜λ©°, 널리 μ‚¬μš©λ˜λŠ” λͺ¨λΈλ“€μ„ μ½”λ“œμ™€ ν•¨κ»˜ λ‚΄λ³΄λ‚΄λŠ” κ³Όμ •κ³Ό torch.export μ‚¬μš©μ€‘ 마주칠 수 μžˆλŠ” λ¬Έμ œλ“€μ„ λ‹€λ£Ήλ‹ˆλ‹€.

이 νŠœν† λ¦¬μ–Όμ€ λ‹€μŒκ³Ό 같은 μ‚¬μš© 사둀에 맞게 λͺ¨λΈμ„ λ‚΄λ³΄λ‚΄λŠ” 방법을 λ°°μ›λ‹ˆλ‹€.

  • μ˜μƒ λΆ„λ₯˜ (MViT)
  • μžλ™ μŒμ„± 인식 (OpenAI Whisper-Tiny)
  • 이미지 캑셔닝 (BLIP)
  • ν”„λ‘¬ν”„νŠΈ 기반 이미지 λΆ„ν•  (SAM2)

각 λ„€ κ°€μ§€ λͺ¨λΈμ€ torch.export 의 κ³ μœ ν•œ κΈ°λŠ₯을 보여주고, κ΅¬ν˜„ κ³Όμ •μ—μ„œμ˜ μ‹€μ§ˆμ μΈ 고렀사항과 λ°œμƒν•  수 μžˆλŠ” λ¬Έμ œλ“€μ„ ν•¨κ»˜ 닀루기 μœ„ν•΄ μ„ μ •λ˜μ—ˆμŠ΅λ‹ˆλ‹€.

μ „μ œ 쑰건

  • PyTorch 2.4 이상 버전
  • torch.export 및 PyTorch Eager 좔둠에 λŒ€ν•œ 기본적인 이해

torch.export 의 핡심 μš”κ΅¬ 사항: κ·Έλž˜ν”„ λΆ„μ ˆ(graph break) μ—†μŒ

torch.compile 은 JITλ₯Ό ν™œμš©ν•΄ PyTorch μ½”λ“œλ₯Ό μ΅œμ ν™”λœ μ»€λ„λ‘œ μ»΄νŒŒμΌν•¨μœΌλ‘œμ¨ μ‹€ν–‰ 속도λ₯Ό ν–₯μƒμ‹œν‚΅λ‹ˆλ‹€. μ£Όμ–΄μ§„ λͺ¨λΈμ„ TorchDynamo λ₯Ό ν™œμš©ν•˜μ—¬ μ΅œμ ν™”ν•˜κ³ , μ΅œμ ν™”λœ κ·Έλž˜ν”„λ₯Ό λ§Œλ“  λ’€, APIμ—μ„œ μ§€μ •ν•œ λ°±μ—”λ“œλ₯Ό 톡해 ν•˜λ“œμ›¨μ–΄μ— 맞게 μ‹€ν–‰λ˜λ„λ‘ λ³€ν™˜ν•©λ‹ˆλ‹€.

TorchDynamo κ°€ μ§€μ›ν•˜μ§€ μ•ŠλŠ” Python의 κΈ°λŠ₯을 λ§Œλ‚˜λ©΄, 계산 κ·Έλž˜ν”„λŠ” μ€‘λ‹¨ν•˜κ³  ν•΄λ‹Ή μ½”λ“œλŠ” κΈ°λ³Έ Python 인터프리터가 μ²˜λ¦¬ν•˜λ„λ‘ ν•˜κ³ , κ·Έλž˜ν”„ 캑쳐λ₯Ό μ΄μ–΄λ‚˜κ°‘λ‹ˆλ‹€. μ΄λŸ¬ν•œ μ€‘λ‹¨λœ 계산 κ·Έλž˜ν”„λ₯Ό graph break 라고 μΉ­ν•©λ‹ˆλ‹€.

torch.export 와 torch.compile 의 μ£Όμš”ν•œ 차이점 쀑 ν•˜λ‚˜λŠ” torch.export λŠ” κ·Έλž˜ν”„ λΆ„μ ˆμ„ μ§€μ›ν•˜μ§€ μ•ŠλŠ”λ‹€λŠ” κ²ƒμž…λ‹ˆλ‹€. 즉, λ‚΄λ³΄λ‚΄λ €λŠ” 전체 λͺ¨λΈ λ˜λŠ” λͺ¨λΈμ˜ μΌλΆ€λŠ” 단일 κ·Έλž˜ν”„ ν˜•νƒœμ—¬μ•Ό ν•©λ‹ˆλ‹€. μ΄λŠ” κ·Έλž˜ν”„ λΆ„μ ˆμ„ μ²˜λ¦¬ν•˜λ €λ©΄ μ§€μ›λ˜μ§€ μ•ŠλŠ” 연산을 κΈ°λ³Έ Python으둜 ν‰κ°€ν•΄μ•Όν•˜λŠ”λ°, μ΄λŸ¬ν•œ 방식이 torch.export 의 섀계와 ν˜Έν™˜λ˜μ§€ μ•ŠκΈ° λ•Œλ¬Έμž…λ‹ˆλ‹€. λ‹€μ–‘ν•œ PyTorch ν”„λ ˆμž„μ›Œν¬λ“€μ˜ 차이점에 λŒ€ν•œ 세뢀적인 μ •λ³΄λŠ” link μ—μ„œ 확인할 수 μžˆμŠ΅λ‹ˆλ‹€.

μ•„λž˜μ˜ μ»€λ§¨λ“œλ₯Ό μ‚¬μš©ν•΄μ„œ ν”„λ‘œκ·Έλž¨ λ‚΄μ˜ κ·Έλž˜ν”„ λΆ„μ ˆμ„ 확인할 수 μžˆμŠ΅λ‹ˆλ‹€.

TORCH_LOGS="graph_breaks" python <file_name>.py

ν”„λ‘œκ·Έλž¨ λ‚΄μ˜ κ·Έλž˜ν”„ λΆ„μ ˆμ„ μ œκ±°ν•˜λ„λ‘ μ½”λ“œλ₯Ό μˆ˜μ •ν•΄μ•Ό ν•©λ‹ˆλ‹€. λ¬Έμ œκ°€ ν•΄κ²°λœλ‹€λ©΄, λͺ¨λΈμ„ 내보낼 μ€€λΉ„κ°€ 된 κ²ƒμž…λ‹ˆλ‹€. PyTorchλŠ” 인기 μžˆλŠ” HuggingFace와 TIMM λͺ¨λΈμ—μ„œ torch.compile 을 μœ„ν•΄μ„œ nightly benchmarks λ₯Ό μ‹€ν–‰ν•©λ‹ˆλ‹€. μ΄λŸ¬ν•œ λͺ¨λΈ λŒ€λΆ€λΆ„μ€ κ·Έλž˜ν”„ λΆ„μ ˆμ΄ μ—†μŠ΅λ‹ˆλ‹€.

ν•΄λ‹Ή λ ˆμ‹œν”Όμ— ν¬ν•¨λœ λͺ¨λΈλ“€μ€ κ·Έλž˜ν”„ λΆ„μ ˆμ΄ μ—†μ§€λ§Œ, torch.export λŠ” μ‹€νŒ¨ν•©λ‹ˆλ‹€.

μ˜μƒ λΆ„λ₯˜

MViTλŠ” MultiScale Vision Transformers 을 κΈ°λ°˜μœΌλ‘œν•œ λͺ¨λΈμ˜ ν΄λž˜μŠ€μž…λ‹ˆλ‹€. 이 λͺ¨λΈμ€ Kinetics-400 Dataset 을 μ‚¬μš©ν•˜μ—¬ 사전 ν›ˆλ ¨λœ μ˜μƒ λΆ„λ₯˜ λͺ¨λΈμž…λ‹ˆλ‹€. 이 λͺ¨λΈμ€ μ μ ˆν•œ 데이터 μ…‹κ³Ό ν•¨κ»˜ μ‚¬μš©ν•œλ‹€λ©΄, κ²Œμž„ ν™˜κ²½μ—μ„œμ˜ λ™μž‘ 인식에 ν™œμš©ν•  수 μžˆμŠ΅λ‹ˆλ‹€.

μ•„λž˜μ˜ μ½”λ“œλŠ” MViTλ₯Ό batch_size=2 둜 νŠΈλ ˆμ΄μ‹±ν•˜μ—¬ 내보내고, 이후 batch_size=4 둜 내보낸 ν”„λ‘œκ·Έλž¨μ΄ μ •μƒμ μœΌλ‘œ μ‹€ν–‰λ˜λŠ”μ§€ ν™•μΈν•©λ‹ˆλ‹€.

import numpy as np
import torch
from torchvision.models.video import MViT_V1_B_Weights, mvit_v1_b
import traceback as tb

model = mvit_v1_b(weights=MViT_V1_B_Weights.DEFAULT)

# 2개의 λΉ„λ””μ˜€μ˜ 배치λ₯Ό λ§Œλ“€λ©°, 각각의 ν˜•νƒœλŠ” 224x224x3에 16 ν”„λ ˆμž„μ„ κ°€μ§‘λ‹ˆλ‹€.
input_frames = torch.randn(2, 16, 224, 224, 3)
# Transpose to get [1, 3, num_clips, height, width].
input_frames = np.transpose(input_frames, (0, 4, 1, 2, 3))

# λͺ¨λΈμ„ λ‚΄λ³΄λƒ…λ‹ˆλ‹€.
exported_program = torch.export.export(
    model,
    (input_frames,),
)

# 4개의 λΉ„λ””μ˜€μ˜ 배치λ₯Ό λ§Œλ“€λ©°, 각각의 ν˜•νƒœλŠ” 224x224x3에 16 ν”„λ ˆμž„μ„ κ°€μ§‘λ‹ˆλ‹€.
input_frames = torch.randn(4, 16, 224, 224, 3)
input_frames = np.transpose(input_frames, (0, 4, 1, 2, 3))
try:
    exported_program.module()(input_frames)
except Exception:
    tb.print_exc()

μ—λŸ¬: 정적 배치 크기

    raise RuntimeError(
RuntimeError: Expected input at *args[0].shape[0] to be equal to 2, but got 4

기본적으둜 λ‚΄λ³΄λ‚΄λŠ” κ³Όμ •μ—μ„œλŠ” λͺ¨λ“  μž…λ ₯ ν˜•νƒœκ°€ κ³ μ •λ˜μ–΄ μžˆλ‹€κ³  κ°€μ •ν•˜κ³  트레이슀(trace) ν•©λ‹ˆλ‹€, λ”°λΌμ„œ νŠΈλ ˆμ΄μ‹±(tracing)을 ν•  λ•Œ μ‚¬μš©ν•œ μž…λ ₯ ν˜•νƒœμ™€ λ‹€λ₯Έ ν˜•νƒœλ‘œ ν”„λ‘œκ·Έλž¨μ„ μ‹€ν–‰ν•˜λ©΄ 였λ₯˜κ°€ λ°œμƒν•©λ‹ˆλ‹€.

ν•΄κ²° 방법

이 였λ₯˜λ₯Ό ν•΄κ²°ν•˜κΈ° μœ„ν•΄, μž…λ ₯의 첫 번째 차원 (batch_size)을 λ™μ μœΌλ‘œ μ§€μ •ν•˜κ³ , ν—ˆμš©λ˜λŠ” batch_size 의 λ²”μœ„λ₯Ό μ§€μ •ν•©λ‹ˆλ‹€. μ•„λž˜μ˜ μˆ˜μ •λœ μ˜ˆμ œμ—μ„œλŠ”, batch_size 의 ν—ˆμš© λ²”μœ„λ₯Ό 1λΆ€ν„° 16κΉŒμ§€λ‘œ μ§€μ •ν•©λ‹ˆλ‹€. μ—¬κΈ°μ„œ μ•Œλ €λ“œλ¦΄ 세뢀사항은 min=2 은 버그가 μ•„λ‹ˆλΌλŠ” 것이고 이에 λŒ€ν•œ μ„€λͺ…은 The 0/1 Specialization Problem λ¬Έμ„œμ—μ„œ 확인할 수 μžˆμŠ΅λ‹ˆλ‹€. λ˜ν•œ torch.export 의 동적 μž…λ ₯ ν˜•νƒœμ— λŒ€ν•œ μžμ„Έν•œ μ„€λͺ…은 export νŠœν† λ¦¬μ–Όμ—μ„œ μ°Ύμ•„λ³Ό 수 μžˆμŠ΅λ‹ˆλ‹€. μ•„λž˜μ˜ μ½”λ“œλŠ” 동적 배치 μ‚¬μ΄μ¦ˆλ₯Ό μ‚¬μš©ν•˜μ—¬ mViTλ₯Ό λ‚΄λ³΄λ‚΄λŠ” 방법을 λ³΄μ—¬μ€λ‹ˆλ‹€.

import numpy as np
import torch
from torchvision.models.video import MViT_V1_B_Weights, mvit_v1_b
import traceback as tb


model = mvit_v1_b(weights=MViT_V1_B_Weights.DEFAULT)

# 2개의 λΉ„λ””μ˜€μ˜ 배치λ₯Ό λ§Œλ“€λ©°, 각각의 ν˜•νƒœλŠ” 224x224x3에 16 ν”„λ ˆμž„μ„ κ°€μ§‘λ‹ˆλ‹€.
input_frames = torch.randn(2,16, 224, 224, 3)

# 차원을 λ°”κΏ” [1, 3, num_clips, height, width] ν˜•νƒœλ‘œ λ³€ν™˜ν•©λ‹ˆλ‹€.
input_frames = np.transpose(input_frames, (0, 4, 1, 2, 3))

# λͺ¨λΈμ„ λ‚΄λ³΄λƒ…λ‹ˆλ‹€.
batch_dim = torch.export.Dim("batch", min=2, max=16)
exported_program = torch.export.export(
    model,
    (input_frames,),
    # Specify the first dimension of the input x as dynamic
    dynamic_shapes={"x": {0: batch_dim}},
)

# 4개의 λΉ„λ””μ˜€μ˜ 배치λ₯Ό λ§Œλ“€λ©°, 각각의 ν˜•νƒœλŠ” 224x224x3에 16 ν”„λ ˆμž„μ„ κ°€μ§‘λ‹ˆλ‹€.
input_frames = torch.randn(4,16, 224, 224, 3)
input_frames = np.transpose(input_frames, (0, 4, 1, 2, 3))
try:
    exported_program.module()(input_frames)
except Exception:
    tb.print_exc()

μžλ™ μŒμ„± 인식

μžλ™ μŒμ„± 인식은 κΈ°κ³„ν•™μŠ΅μ„ ν™œμš©ν•˜μ—¬ μŒμ„±μ„ ν…μŠ€νŠΈλ‘œ λ³€ν™˜ν•˜λŠ” κΈ°μˆ μž…λ‹ˆλ‹€. Whisper λŠ” OpenAIμ—μ„œ κ°œλ°œν•œ 인코더-디코더 ꡬ쑰의 트랜슀포머 λͺ¨λΈλ‘œ, ASRκ³Ό μŒμ„± λ²ˆμ—­μ„ μœ„ν•΄ 68만 μ‹œκ°„μ˜ 라벨링된 데이터λ₯Ό μ‚¬μš©ν•΄ ν•™μŠ΅λ˜μ—ˆμŠ΅λ‹ˆλ‹€. μ•„λž˜μ˜ μ½”λ“œλ‘œ μžλ™ μŒμ„± 인식을 μœ„ν•œ whisper-tiny λͺ¨λΈμ„ 내보낼 수 μžˆμŠ΅λ‹ˆλ‹€.

import torch
from transformers import WhisperProcessor, WhisperForConditionalGeneration
from datasets import load_dataset

# λͺ¨λΈμ„ κ°€μ Έμ˜΅λ‹ˆλ‹€.
model = WhisperForConditionalGeneration.from_pretrained("openai/whisper-tiny")

# λͺ¨λΈ 내보내기λ₯Ό μœ„ν•œ 더미 μž…λ ₯μž…λ‹ˆλ‹€.
input_features = torch.randn(1,80, 3000)
attention_mask = torch.ones(1, 3000)
decoder_input_ids = torch.tensor([[1, 1, 1 , 1]]) * model.config.decoder_start_token_id

model.eval()

exported_program: torch.export.ExportedProgram= torch.export.export(model, args=(input_features, attention_mask, decoder_input_ids,))

μ—λŸ¬: TorchDynamoλ₯Ό μ΄μš©ν•œ μ—„κ²©ν•œ(strict) νŠΈλ ˆμ΄μ‹±(tracing)

torch._dynamo.exc.InternalTorchDynamoError: AttributeError: 'DynamicCache' object has no attribute 'key_cache'

기본적으둜 torch.export λŠ” TorchDynamo λΌλŠ” λ°”μ΄νŠΈμ½”λ“œ 뢄석 엔진을 μ‚¬μš©ν•˜μ—¬ μ½”λ“œλ₯Ό μ²˜λ¦¬ν•©λ‹ˆλ‹€, μ΄λŠ” μ½”λ“œλ₯Ό μ‹¬λ³Όλ¦­ν•˜κ²Œ λΆ„μ„ν•˜μ—¬ κ·Έλž˜ν”„λ₯Ό μƒμ„±ν•©λ‹ˆλ‹€. 이 뢄석은 μ•ˆμ „μ„± 보μž₯을 κ°•ν™”ν•΄μ£Όμ§€λ§Œ, λͺ¨λ“  Python μ½”λ“œλ₯Ό μ§€μ›ν•˜λŠ” 것은 μ•„λ‹™λ‹ˆλ‹€. whisper-tiny λͺ¨λΈμ„ κΈ°λ³Έ strict λͺ¨λ“œλ‘œ 내보낼 λ•Œ, Dynamoμ—μ„œ μ§€μ›λ˜μ§€ μ•ŠλŠ” κΈ°λŠ₯ λ•Œλ¬Έμ— 일반적으둜 였λ₯˜κ°€ λ°œμƒν•©λ‹ˆλ‹€. Dynamoμ—μ„œ 이 μ—λŸ¬κ°€ λ°œμƒν•˜λŠ” 이유λ₯Ό μ΄ν•΄ν•˜λ €λ©΄, GitHub issue ν•΄λ‹Ή κΉƒν—ˆλΈŒ 이슈λ₯Ό μ°Έκ³ ν•˜μ„Έμš”.

ν•΄κ²° 방법

μœ„μ˜ μ—λŸ¬λ₯Ό ν•΄κ²°ν•˜κΈ° μœ„ν•΄, torch.export λŠ” Python 인터프리터λ₯Ό μ‚¬μš©ν•΄ ν”„λ‘œκ·Έλž¨μ„ νŠΈλ ˆμ΄μ‹±ν•˜λŠ” non_strict λͺ¨λ“œλ₯Ό μ œκ³΅ν•˜λ©°, μ΄λŠ” PyTorch eager μ‹€ν–‰κ³Ό μœ μ‚¬ν•˜κ²Œ λ™μž‘ν•©λ‹ˆλ‹€. μœ μΌν•œ 차이점은 λͺ¨λ“  Tensor 객체가 ProxyTensors 둜 λŒ€μ²΄λ˜λ©°, μ΄λŠ” λͺ¨λ“  연산이 κ·Έλž˜ν”„μ— κΈ°λ‘λœλ‹€λŠ” κ²ƒμž…λ‹ˆλ‹€. strict=False 을 μ‚¬μš©ν•˜λ©΄, ν”„λ‘œκ·Έλž¨μ—μ„œ 내보낼 수 μžˆμŠ΅λ‹ˆλ‹€.

import torch
from transformers import WhisperProcessor, WhisperForConditionalGeneration
from datasets import load_dataset

# λͺ¨λΈμ„ κ°€μ Έμ˜΅λ‹ˆλ‹€.
model = WhisperForConditionalGeneration.from_pretrained("openai/whisper-tiny")

# λͺ¨λΈ 내보내기λ₯Ό μœ„ν•œ 더미 μž…λ ₯μž…λ‹ˆλ‹€.
input_features = torch.randn(1,80, 3000)
attention_mask = torch.ones(1, 3000)
decoder_input_ids = torch.tensor([[1, 1, 1 , 1]]) * model.config.decoder_start_token_id

model.eval()

exported_program: torch.export.ExportedProgram= torch.export.export(model, args=(input_features, attention_mask, decoder_input_ids,), strict=False)

이미지 캑셔닝

이미지 캑셔닝 은 이미지에 μžˆλŠ” λ‹¨μ–΄μ˜ λ‚΄μš©μ„ μ •μ˜ν•˜λŠ” 업무λ₯Ό μˆ˜ν–‰ν•œλ‹€. κ²Œμž„ ν™˜κ²½μ—μ„œ 이미지 캑셔닝은 μž₯λ©΄ λ‚΄ λ‹€μ–‘ν•œ κ²Œμž„ 객체에 λŒ€ν•œ ν…μŠ€νŠΈ μ„€λͺ…을 λ™μ μœΌλ‘œ μƒμ„±ν•˜λ©°, κ²Œμ΄λ¨Έμ—κ²Œ 좔가적인 정보λ₯Ό μ œκ³΅ν•¨μœΌλ‘œμ¨ κ²Œμž„ ν”Œλ ˆμ΄ κ²½ν—˜μ„ ν–₯μƒμ‹œν‚€λŠ”λ° ν™œμš©λ  수 μžˆμŠ΅λ‹ˆλ‹€. BLIP λŠ” 이미지 캑셔닝 λΆ„μ•Όμ—μ„œ 널리 μ‚¬μš©λ˜λŠ” λͺ¨λΈλ‘œ, released by SalesForce Research μ—μ„œ κ³΅κ°œλ˜μ—ˆμŠ΅λ‹ˆλ‹€. μ•„λž˜ μ½”λ“œλŠ” batch_size=1 둜 BLIPλ₯Ό 내보내렀고 μ‹œλ„ν•©λ‹ˆλ‹€.

import torch
from models.blip import blip_decoder

device = torch.device('cuda' if torch.cuda.is_available() else 'cpu')
image_size = 384
image = torch.randn(1, 3,384,384).to(device)
caption_input = ""

model_url = 'https://storage.googleapis.com/sfr-vision-language-research/BLIP/models/model_base_capfilt_large.pth'
model = blip_decoder(pretrained=model_url, image_size=image_size, vit='base')
model.eval()
model = model.to(device)

exported_program: torch.export.ExportedProgram= torch.export.export(model, args=(image,caption_input,), strict=False)

μ—λŸ¬: λ™κ²°λœ(frozen) μ €μž₯μ†Œλ₯Ό κ°€μ§„ ν…μ„œλ₯Ό λ³€κ²½ν•  수 μ—†μŠ΅λ‹ˆλ‹€.

λͺ¨λΈμ„ 내보낼 λ•Œ, λͺ¨λΈ κ΅¬ν˜„μ—μ„œ torch.export μ—μ„œ 아직 μ§€μ›ν•˜μ§€ μ•ŠλŠ” νŠΉμ • Python 연산이 포함이 될 수 있기 λ•Œλ¬Έμ— μ‹€νŒ¨ν•  수 μžˆμŠ΅λ‹ˆλ‹€. 이 μ‹€νŒ¨ 사둀듀 쀑 μΌλΆ€λŠ” ν•΄κ²° 방법이 μžˆμ„ 수 μžˆμŠ΅λ‹ˆλ‹€. BLIPλŠ” μ›λž˜ λͺ¨λΈμ—μ„œ 였λ₯˜κ°€ λ°œμƒν•˜λŠ” μ˜ˆμ‹œμ΄μ§€λ§Œ, μ½”λ“œμ— μž‘μ€ μˆ˜μ •μ„ ν•˜λ©΄ ν•΄κ²°ν•  수 μžˆμŠ΅λ‹ˆλ‹€. torch.export λŠ” ExportDB μ—μ„œ μ§€μ›ν•˜λŠ” μ—°μ‚°κ³Ό μ§€μ›ν•˜μ§€ μ•ŠλŠ” μ—°μ‚°μ˜ 일반적인 사둀듀을 λ‚˜μ—΄ν•˜κ³ , μ½”λ“œμ—μ„œ 내보낼 수 μžˆλ„λ‘ μˆ˜μ •ν•˜λŠ” 방법을 λ³΄μ—¬μ€λ‹ˆλ‹€.

File "/BLIP/models/blip.py", line 112, in forward
    text.input_ids[:,0] = self.tokenizer.bos_token_id
  File "/anaconda3/envs/export/lib/python3.10/site-packages/torch/_subclasses/functional_tensor.py", line 545, in __torch_dispatch__
    outs_unwrapped = func._op_dk(
RuntimeError: cannot mutate tensors with frozen storage

ν•΄κ²° 방법

내보내기가 μ‹€νŒ¨ν•˜λŠ” μœ„μΉ˜μ— μžˆλŠ” tensor λ₯Ό λ³΅μ œν•©λ‹ˆλ‹€.

text.input_ids = text.input_ids.clone() # clone the tensor
text.input_ids[:,0] = self.tokenizer.bos_token_id

Note

This constraint has been relaxed in PyTorch 2.7 nightlies. This should work out-of-the-box in PyTorch 2.7 이 μ œμ•½μ€ PyTorch 2.7 nightliesμ—μ„œ μ™„ν™”λ˜μ—ˆμŠ΅λ‹ˆλ‹€. PyTorch 2.7μ—μ„œλŠ” λ³„λ„μ˜ μ„€μ • 없이 λ°”λ‘œ λ™μž‘ν•  κ²ƒμž…λ‹ˆλ‹€.

ν”„λ‘¬ν”„νŠΈ 기반 이미지 λΆ„ν• 

이미지 λΆ„ν•  은 λ””μ§€ν„Έ 이미지λ₯Ό ν”½μ…€ λ‹¨μœ„μ˜ νŠΉμ§•μ— 따라 μ„œλ‘œ λ‹€λ₯Έ κ·Έλ£Ή, 즉 μ„Έκ·Έλ¨ΌνŠΈλ‘œ λ‚˜λˆ„λŠ” 컴퓨터 λΉ„μ „ κΈ°μˆ μž…λ‹ˆλ‹€. Segment Anything Model (SAM) 은 ν”„λ‘¬ν”„νŠΈ 기반 이미지 뢄할을 λ„μž…ν•œ λͺ¨λΈλ‘œ, μ‚¬μš©μžκ°€ μ›ν•˜λŠ” 객체λ₯Ό μ§€μ •ν•˜λŠ” ν”„λ‘¬ν”„νŠΈλ₯Ό μž…λ ₯ν•˜λ©΄ ν•΄λ‹Ή 객체의 마슀크λ₯Ό μ˜ˆμΈ‘ν•©λ‹ˆλ‹€. SAM 2 λŠ” 이미지와 λΉ„λ””μ˜€μ—μ„œ 객체λ₯Ό λΆ„ν• ν•˜κΈ° μœ„ν•œ 졜초의 톡합 λͺ¨λΈμž…λ‹ˆλ‹€. SAM2ImagePredictor ν΄λž˜μŠ€λŠ” λͺ¨λΈμ— ν”„λ‘¬ν”„νŠΈλ₯Ό μž…λ ₯ν•  수 μžˆλŠ” κ°„νŽΈν•œ μΈν„°νŽ˜μ΄μŠ€λ₯Ό μ œκ³΅ν•©λ‹ˆλ‹€. 이 λͺ¨λΈμ€ ν¬μΈνŠΈμ™€ λ°•μŠ€ ν”„λ‘¬ν”„νŠΈλŠ” λ¬Όλ‘ , 이전 μ˜ˆμΈ‘μ—μ„œ μƒμ„±λœ λ§ˆμŠ€ν¬λ„ μž…λ ₯으둜 받을 수 μžˆμŠ΅λ‹ˆλ‹€. SAM2λŠ” 객체 μΆ”μ μ—μ„œ κ°•λ ₯ν•œ μ œλ‘œμƒ· μ„±λŠ₯을 μ œκ³΅ν•˜λ―€λ‘œ, μž₯λ©΄ λ‚΄ κ²Œμž„ 객체λ₯Ό μΆ”μ ν•˜λŠ” 데 ν™œμš©ν•  수 μžˆμŠ΅λ‹ˆλ‹€.

SAM2ImagePredictor 의 예츑 λ©”μ„œλ“œμ—μ„œ λ°œμƒν•˜λŠ” ν…μ„œ 연산은 μ‹€μ œλ‘œ _predict λ©”μ„œλ“œ μ•ˆμ—μ„œ μˆ˜ν–‰λ©λ‹ˆλ‹€. λ”°λΌμ„œ μ•„λž˜μ™€ 같이 내보내기λ₯Ό μ‹œλ„ν•©λ‹ˆλ‹€.

ep = torch.export.export(
    self._predict,
    args=(unnorm_coords, labels, unnorm_box, mask_input, multimask_output),
    kwargs={"return_logits": return_logits},
    strict=False,
)

μ—λŸ¬: λͺ¨λΈμ˜ νƒ€μž…μ΄ torch.nn.Module μ•„λ‹™λ‹ˆλ‹€.

torch.export λŠ” λͺ¨λ“ˆμ΄ torch.nn.Module νƒ€μž…μ΄μ–΄μ•Ό ν•©λ‹ˆλ‹€. ν•˜μ§€λ§Œ, 내보내기 ν•˜λ €λŠ” λͺ¨λ“ˆμ€ 클래슀 λ©”μ„œλ“œμ΄κΈ° λ•Œλ¬Έμ— 였λ₯˜κ°€ λ°œμƒν•©λ‹ˆλ‹€.

Traceback (most recent call last):
  File "/sam2/image_predict.py", line 20, in <module>
    masks, scores, _ = predictor.predict(
  File "/sam2/sam2/sam2_image_predictor.py", line 312, in predict
    ep = torch.export.export(
  File "python3.10/site-packages/torch/export/__init__.py", line 359, in export
    raise ValueError(
ValueError: Expected `mod` to be an instance of `torch.nn.Module`, got <class 'method'>.

ν•΄κ²° 방법

λ„μš°λ―Έ 클래슀λ₯Ό μž‘μ„±ν•˜μ—¬ torch.nn.Module 을 μƒμ†ν•˜κ³ , 클래슀의 forward λ©”μ„œλ“œ μ•ˆμ—μ„œ _predict method λ₯Ό ν˜ΈμΆœν•©λ‹ˆλ‹€. 전체 μ½”λ“œλŠ” here μ—μ„œ 확인할 수 μžˆμŠ΅λ‹ˆλ‹€.

class ExportHelper(torch.nn.Module):
    def __init__(self):
        super().__init__()

    def forward(_, *args, **kwargs):
        return self._predict(*args, **kwargs)

 model_to_export = ExportHelper()
 ep = torch.export.export(
      model_to_export,
      args=(unnorm_coords, labels, unnorm_box, mask_input,  multimask_output),
      kwargs={"return_logits": return_logits},
      strict=False,
      )

κ²°λ‘ 

이 νŠœν† λ¦¬μ–Όμ—μ„œλŠ” torch.export λ₯Ό ν™œμš©ν•˜μ—¬ λ‹€μ–‘ν•œ λŒ€ν‘œμ μΈ μ‚¬μš© μ‚¬λ‘€μ˜ λͺ¨λΈμ„ λ‚΄λ³΄λ‚΄λŠ” 방법을 ν•™μŠ΅ν•˜μ˜€κ³ , μ˜¬λ°”λ₯Έ μ„€μ •κ³Ό κ°„λ‹¨ν•œ μ½”λ“œ μˆ˜μ •μœΌλ‘œ λ°œμƒν•  수 μžˆλŠ” μ—¬λŸ¬ λ¬Έμ œλ“€μ„ ν•΄κ²°ν•˜λŠ” 방법도 ν•¨κ»˜ λ‹€λ€˜μŠ΅λ‹ˆλ‹€. λͺ¨λΈμ„ μ„±κ³΅μ μœΌλ‘œ 내보낸 후에, μ„œλ²„ ν™˜κ²½μ—μ„œλŠ” AOTInductor λ₯Ό, μ—£μ§€ λ””λ°”μ΄μŠ€ ν™˜κ²½μ—μ„œλŠ” ExecuTorch λ₯Ό μ‚¬μš©ν•˜μ—¬ ExportedProgram 을 ν•˜λ“œμ›¨μ–΄μ— 맞게 λ³€ν™˜ν•  수 μžˆμŠ΅λ‹ˆλ‹€. AOTInductor (AOTI)에 λŒ€ν•œ μžμ„Έν•œ λ‚΄μš©μ€ AOTI tutorial 을, ExecuTorch 에 λŒ€ν•œ μžμ„Έν•œ λ‚΄μš©μ€ ExecuTorch tutorial 을 μ°Έκ³ ν•˜μ„Έμš”.