I tried to load llama-2-70b-chat-hf with the following code:
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM, BitsAndBytesConfig
base_model = â/llm/llama2-2-70b-chat-hfâ
model = AutoModelForCausalLM( base_model, load_in_8bit=True, device_map={ââ,0},use_safetensors=True)
and the Error is shown below:
in load_state_dict (checkpoint_file)
462 ââ"
463 Reads a Pytorch checkpoint file, returning properly formatted errors if they arise.
464 ââ"
465 if checkpoint_file.endswith(â.safetensorsâ) and is_safetensors_available():
â>466 with safe_open(checkpoint_file,framework=âptâ) as f:
467 metadata=f.metadate()
SafetensorError: Error while deserializing header: HeaderTooLarge
I tried to re-download the safetensors file but it cannot be solved.