Safetensors and GGUF Format
Safetensors and GGUF have become the de facto data format to efficiently store model weights in a readable way. This post briefly analyzes their format protocols.
Table of Contents
1. Safetensors
The safetensors layout is like this:
+-----------------------------------------------------------------------+
| 8 Bytes: Header Size (N) [Little-Endian uint64] |
+-----------------------------------------------------------------------+
| N Bytes: JSON Header (UTF-8) |
| - __metadata__: 自定义键值对(可选) |
| - tensor_name_1: { dtype, shape, data_offsets: [start, end] } |
| - tensor_name_2: { dtype, shape, data_offsets: [start, end] } |
| ... |
+-----------------------------------------------------------------------+
| Raw Tensor Buffer (连续的二进制权重数据) |
| [tensor_1 bytes][tensor_2 bytes] ... |
+-----------------------------------------------------------------------+
Here, we use HuggingFaceTB/SmolLM2-135M model as an example. First, download it.
hf download HuggingFaceTB/SmolLM2-135M --local-dir smollm2
Then, we use the following Python script to inspect the header size and JSON header:
import json
with open("smollm2/model.safetensors", "rb") as f:
data = f.read(8)
size = int.from_bytes(data, byteorder="little")
print(f"header size: {size} bytes")
data = f.read(size).decode("utf-8")
metadata = json.loads(data)
print(json.dumps(metadata, indent=2))
Since the header size is encoded in little endian, therefore, we need to convert it using int.from_bytes(, byteorder="little").
The output is like:
header size: 30528 bytes
{
"__metadata__": {
"format": "pt"
},
"model.embed_tokens.weight": {
"dtype": "BF16",
"shape": [
49152,
576
],
"data_offsets": [
0,
56623104
]
},
"model.layers.0.input_layernorm.weight": {
"dtype": "BF16",
"shape": [
576
],
"data_offsets": [
56623104,
56624256
]
},
....
"model.layers.9.self_attn.v_proj.weight": {
"dtype": "BF16",
"shape": [
192,
576
],
"data_offsets": [
268807680,
269028864
]
},
"model.norm.weight": {
"dtype": "BF16",
"shape": [
576
],
"data_offsets": [
269028864,
269030016
]
}
}
2. GGUF
The layout of GGUF is:
+-----------------------------------------------------------------------+ | Header | | - Magic Number: "GGUF" (4 Bytes, uint32) | | - Version: uint32 (当前常见为 v3) | | - Tensor Count: uint64 | | - Metadata KV Count: uint64 | +-----------------------------------------------------------------------+ | Metadata Key-Value Pairs (共 Metadata KV Count 项) | | - Key: gguf_string (长度 + UTF-8 字符) | | - Value Type: uint32 (表示整型/浮点/字符串/数组等) | | - Value: 紧凑的二进制值 (支持嵌入复杂数组,如词表 token list) | +-----------------------------------------------------------------------+ | Tensor Infos (共 Tensor Count 项) | | - Tensor 1 Name: gguf_string | | - n_dimensions: uint32 | | - dimensions: uint64[n_dimensions] | | - type: uint32 (GGML_TYPE_*, 如 Q4_K, F16, Q8_0) | | - offset: uint64 (相对 Tensor Data 基础位置的偏移) | | ... | +-----------------------------------------------------------------------+ | Padding (填充字节,使 Tensor Data 对齐到 alignment 边界,通常为 32B) | +-----------------------------------------------------------------------+ | Tensor Data (紧凑且对齐的权重/量化块数据) | | [aligned tensor_1][aligned tensor_2] ... | +-----------------------------------------------------------------------+