+ OUTPUT_DIR=./upload-OpenJev + LLAMA_CPP=llama.cpp + DISPLAY_NAME=OpenJev + QUANTIZE=llama.cpp/build/bin/llama-quantize + python3 llama.cpp/convert_hf_to_gguf.py ./model-temp-OpenJev-PRIMARY --outtype bf16 --outfile ./upload-OpenJev/OpenJev-BF16.gguf --no-mtp --model-name OpenJev INFO:hf-to-gguf:Loading model: model-temp-OpenJev-PRIMARY INFO:hf-to-gguf:gguf: detected OpenJev checkpoint [transformers] Missing validation function in 'RotaryEmbeddingConfigMixin' for 'rope_type'='axial' INFO:hf-to-gguf:Model architecture: OpenJevModel INFO:hf-to-gguf:gguf: detected OpenJev checkpoint [transformers] Missing validation function in 'RotaryEmbeddingConfigMixin' for 'rope_type'='axial' INFO:hf-to-gguf:gguf: loading model weight map from 'model.safetensors.index.json' INFO:hf-to-gguf:gguf: indexing model part 'model-00001-of-00012.safetensors' INFO:hf-to-gguf:gguf: indexing model part 'model-00002-of-00012.safetensors' INFO:hf-to-gguf:gguf: indexing model part 'model-00003-of-00012.safetensors' INFO:hf-to-gguf:gguf: indexing model part 'model-00004-of-00012.safetensors' INFO:hf-to-gguf:gguf: indexing model part 'model-00005-of-00012.safetensors' INFO:hf-to-gguf:gguf: indexing model part 'model-00006-of-00012.safetensors' INFO:hf-to-gguf:gguf: indexing model part 'model-00007-of-00012.safetensors' INFO:hf-to-gguf:gguf: indexing model part 'model-00008-of-00012.safetensors' INFO:hf-to-gguf:gguf: indexing model part 'model-00009-of-00012.safetensors' INFO:hf-to-gguf:gguf: indexing model part 'model-00010-of-00012.safetensors' INFO:hf-to-gguf:gguf: indexing model part 'model-00011-of-00012.safetensors' INFO:hf-to-gguf:gguf: indexing model part 'model-00012-of-00012.safetensors' INFO:gguf.gguf_writer:gguf: This GGUF file is for Little Endian only INFO:hf-to-gguf:Exporting model... INFO:hf-to-gguf:output.weight, torch.bfloat16 --> BF16, shape = {5120, 248320} INFO:hf-to-gguf:token_embd.weight, torch.bfloat16 --> BF16, shape = {5120, 248320} INFO:hf-to-gguf:blk.0.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.0.ssm_a, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.0.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} INFO:hf-to-gguf:blk.0.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.0.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.0.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.0.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} INFO:hf-to-gguf:blk.0.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} INFO:hf-to-gguf:blk.0.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} INFO:hf-to-gguf:blk.0.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} INFO:hf-to-gguf:blk.0.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} INFO:hf-to-gguf:blk.0.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.0.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.0.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.1.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.1.ssm_a, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.1.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} INFO:hf-to-gguf:blk.1.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.1.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.1.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.1.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} INFO:hf-to-gguf:blk.1.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} INFO:hf-to-gguf:blk.1.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} INFO:hf-to-gguf:blk.1.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} INFO:hf-to-gguf:blk.1.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} INFO:hf-to-gguf:blk.1.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.1.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.1.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.2.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.2.ssm_a, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.2.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} INFO:hf-to-gguf:blk.2.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.2.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.2.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.2.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} INFO:hf-to-gguf:blk.2.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} INFO:hf-to-gguf:blk.2.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} INFO:hf-to-gguf:blk.2.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} INFO:hf-to-gguf:blk.2.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} INFO:hf-to-gguf:blk.2.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.2.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.2.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.3.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.3.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} INFO:hf-to-gguf:blk.3.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.3.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.3.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.3.attn_k_norm.weight, torch.bfloat16 --> F32, shape = {256} INFO:hf-to-gguf:blk.3.attn_k.weight, torch.bfloat16 --> BF16, shape = {5120, 1024} INFO:hf-to-gguf:blk.3.attn_output.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} INFO:hf-to-gguf:blk.3.attn_q_norm.weight, torch.bfloat16 --> F32, shape = {256} INFO:hf-to-gguf:blk.3.attn_q.weight, torch.bfloat16 --> BF16, shape = {5120, 12288} INFO:hf-to-gguf:blk.3.attn_v.weight, torch.bfloat16 --> BF16, shape = {5120, 1024} INFO:hf-to-gguf:blk.4.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.4.ssm_a, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.4.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} INFO:hf-to-gguf:blk.4.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.4.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.4.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.4.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} INFO:hf-to-gguf:blk.4.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} INFO:hf-to-gguf:blk.4.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} INFO:hf-to-gguf:blk.4.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} INFO:hf-to-gguf:blk.4.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} INFO:hf-to-gguf:blk.4.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.4.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.4.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.5.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.5.ssm_a, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.5.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} INFO:hf-to-gguf:blk.5.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.5.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.5.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.5.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} INFO:hf-to-gguf:blk.5.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} INFO:hf-to-gguf:blk.5.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} INFO:hf-to-gguf:blk.5.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} INFO:hf-to-gguf:blk.5.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} INFO:hf-to-gguf:blk.5.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.5.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.5.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.6.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.6.ssm_a, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.6.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} INFO:hf-to-gguf:blk.6.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.6.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.6.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.6.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} INFO:hf-to-gguf:blk.6.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} INFO:hf-to-gguf:blk.6.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} INFO:hf-to-gguf:blk.6.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} INFO:hf-to-gguf:blk.6.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} INFO:hf-to-gguf:blk.6.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.6.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.6.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.7.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.7.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} INFO:hf-to-gguf:blk.7.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.7.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.7.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.7.attn_k_norm.weight, torch.bfloat16 --> F32, shape = {256} INFO:hf-to-gguf:blk.7.attn_k.weight, torch.bfloat16 --> BF16, shape = {5120, 1024} INFO:hf-to-gguf:blk.7.attn_output.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} INFO:hf-to-gguf:blk.7.attn_q_norm.weight, torch.bfloat16 --> F32, shape = {256} INFO:hf-to-gguf:blk.7.attn_q.weight, torch.bfloat16 --> BF16, shape = {5120, 12288} INFO:hf-to-gguf:blk.7.attn_v.weight, torch.bfloat16 --> BF16, shape = {5120, 1024} INFO:hf-to-gguf:blk.8.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.8.ssm_a, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.8.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} INFO:hf-to-gguf:blk.8.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.8.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.8.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.8.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} INFO:hf-to-gguf:blk.8.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} INFO:hf-to-gguf:blk.8.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} INFO:hf-to-gguf:blk.8.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} INFO:hf-to-gguf:blk.8.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} INFO:hf-to-gguf:blk.8.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.8.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.8.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.9.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.9.ssm_a, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.9.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} INFO:hf-to-gguf:blk.9.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.9.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.9.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.9.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} INFO:hf-to-gguf:blk.9.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} INFO:hf-to-gguf:blk.9.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} INFO:hf-to-gguf:blk.9.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} INFO:hf-to-gguf:blk.9.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} INFO:hf-to-gguf:blk.10.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.10.ssm_a, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.10.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} INFO:hf-to-gguf:blk.10.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.10.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.10.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.10.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} INFO:hf-to-gguf:blk.10.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} INFO:hf-to-gguf:blk.10.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} INFO:hf-to-gguf:blk.10.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} INFO:hf-to-gguf:blk.10.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} INFO:hf-to-gguf:blk.10.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.10.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.10.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.11.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.11.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} INFO:hf-to-gguf:blk.11.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.11.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.11.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.11.attn_k_norm.weight, torch.bfloat16 --> F32, shape = {256} INFO:hf-to-gguf:blk.11.attn_k.weight, torch.bfloat16 --> BF16, shape = {5120, 1024} INFO:hf-to-gguf:blk.11.attn_output.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} INFO:hf-to-gguf:blk.11.attn_q_norm.weight, torch.bfloat16 --> F32, shape = {256} INFO:hf-to-gguf:blk.11.attn_q.weight, torch.bfloat16 --> BF16, shape = {5120, 12288} INFO:hf-to-gguf:blk.11.attn_v.weight, torch.bfloat16 --> BF16, shape = {5120, 1024} INFO:hf-to-gguf:blk.12.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.12.ssm_a, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.12.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} INFO:hf-to-gguf:blk.12.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.12.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.12.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.12.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} INFO:hf-to-gguf:blk.12.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} INFO:hf-to-gguf:blk.12.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} INFO:hf-to-gguf:blk.12.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} INFO:hf-to-gguf:blk.12.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} INFO:hf-to-gguf:blk.12.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.12.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.12.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.13.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.13.ssm_a, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.13.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} INFO:hf-to-gguf:blk.13.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.13.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.13.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.13.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} INFO:hf-to-gguf:blk.13.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} INFO:hf-to-gguf:blk.13.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} INFO:hf-to-gguf:blk.13.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} INFO:hf-to-gguf:blk.13.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} INFO:hf-to-gguf:blk.13.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.13.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.13.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.14.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.14.ssm_a, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.14.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} INFO:hf-to-gguf:blk.14.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.14.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.14.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.14.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} INFO:hf-to-gguf:blk.14.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} INFO:hf-to-gguf:blk.14.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} INFO:hf-to-gguf:blk.14.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} INFO:hf-to-gguf:blk.14.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} INFO:hf-to-gguf:blk.14.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.14.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.14.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.15.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.15.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} INFO:hf-to-gguf:blk.15.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.15.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.15.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.15.attn_k_norm.weight, torch.bfloat16 --> F32, shape = {256} INFO:hf-to-gguf:blk.15.attn_k.weight, torch.bfloat16 --> BF16, shape = {5120, 1024} INFO:hf-to-gguf:blk.15.attn_output.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} INFO:hf-to-gguf:blk.15.attn_q_norm.weight, torch.bfloat16 --> F32, shape = {256} INFO:hf-to-gguf:blk.15.attn_q.weight, torch.bfloat16 --> BF16, shape = {5120, 12288} INFO:hf-to-gguf:blk.15.attn_v.weight, torch.bfloat16 --> BF16, shape = {5120, 1024} INFO:hf-to-gguf:blk.16.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.16.ssm_a, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.16.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} INFO:hf-to-gguf:blk.16.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.16.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.16.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.9.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.9.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.9.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.16.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} INFO:hf-to-gguf:blk.16.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} INFO:hf-to-gguf:blk.16.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} INFO:hf-to-gguf:blk.16.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} INFO:hf-to-gguf:blk.16.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} INFO:hf-to-gguf:blk.16.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.16.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.16.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.17.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.17.ssm_a, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.17.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} INFO:hf-to-gguf:blk.17.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.17.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.17.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.17.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} INFO:hf-to-gguf:blk.17.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} INFO:hf-to-gguf:blk.17.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} INFO:hf-to-gguf:blk.17.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} INFO:hf-to-gguf:blk.17.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} INFO:hf-to-gguf:blk.17.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.17.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.17.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.18.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.18.ssm_a, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.18.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} INFO:hf-to-gguf:blk.18.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.18.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.18.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.18.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} INFO:hf-to-gguf:blk.18.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} INFO:hf-to-gguf:blk.18.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} INFO:hf-to-gguf:blk.18.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} INFO:hf-to-gguf:blk.18.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} INFO:hf-to-gguf:blk.18.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.18.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.18.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.19.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.19.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} INFO:hf-to-gguf:blk.19.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.19.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.19.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.19.attn_k_norm.weight, torch.bfloat16 --> F32, shape = {256} INFO:hf-to-gguf:blk.19.attn_k.weight, torch.bfloat16 --> BF16, shape = {5120, 1024} INFO:hf-to-gguf:blk.19.attn_output.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} INFO:hf-to-gguf:blk.19.attn_q_norm.weight, torch.bfloat16 --> F32, shape = {256} INFO:hf-to-gguf:blk.19.attn_q.weight, torch.bfloat16 --> BF16, shape = {5120, 12288} INFO:hf-to-gguf:blk.19.attn_v.weight, torch.bfloat16 --> BF16, shape = {5120, 1024} INFO:hf-to-gguf:blk.20.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.20.ssm_a, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.20.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} INFO:hf-to-gguf:blk.20.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.20.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.20.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.20.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} INFO:hf-to-gguf:blk.20.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} INFO:hf-to-gguf:blk.20.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} INFO:hf-to-gguf:blk.20.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} INFO:hf-to-gguf:blk.20.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} INFO:hf-to-gguf:blk.20.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.20.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.20.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.21.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.21.ssm_a, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.21.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} INFO:hf-to-gguf:blk.21.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.21.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.21.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.21.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} INFO:hf-to-gguf:blk.21.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} INFO:hf-to-gguf:blk.21.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} INFO:hf-to-gguf:blk.21.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} INFO:hf-to-gguf:blk.21.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} INFO:hf-to-gguf:blk.21.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.21.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.21.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.22.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.22.ssm_a, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.22.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} INFO:hf-to-gguf:blk.22.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.22.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.22.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.22.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} INFO:hf-to-gguf:blk.22.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} INFO:hf-to-gguf:blk.22.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} INFO:hf-to-gguf:blk.22.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} INFO:hf-to-gguf:blk.22.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} INFO:hf-to-gguf:blk.22.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.22.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.22.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.23.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.23.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} INFO:hf-to-gguf:blk.23.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.23.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.23.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.23.attn_k_norm.weight, torch.bfloat16 --> F32, shape = {256} INFO:hf-to-gguf:blk.23.attn_k.weight, torch.bfloat16 --> BF16, shape = {5120, 1024} INFO:hf-to-gguf:blk.23.attn_output.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} INFO:hf-to-gguf:blk.23.attn_q_norm.weight, torch.bfloat16 --> F32, shape = {256} INFO:hf-to-gguf:blk.23.attn_q.weight, torch.bfloat16 --> BF16, shape = {5120, 12288} INFO:hf-to-gguf:blk.23.attn_v.weight, torch.bfloat16 --> BF16, shape = {5120, 1024} INFO:hf-to-gguf:blk.24.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.24.ssm_a, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.24.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} INFO:hf-to-gguf:blk.24.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.24.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.24.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.24.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} INFO:hf-to-gguf:blk.24.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} INFO:hf-to-gguf:blk.24.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} INFO:hf-to-gguf:blk.24.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} INFO:hf-to-gguf:blk.24.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} INFO:hf-to-gguf:blk.24.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.24.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.24.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.25.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.25.ssm_a, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.25.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} INFO:hf-to-gguf:blk.25.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.25.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.25.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.25.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} INFO:hf-to-gguf:blk.25.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} INFO:hf-to-gguf:blk.25.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} INFO:hf-to-gguf:blk.25.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} INFO:hf-to-gguf:blk.25.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} INFO:hf-to-gguf:blk.25.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.25.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.25.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.26.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.26.ssm_a, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.26.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} INFO:hf-to-gguf:blk.26.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.26.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.26.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.26.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} INFO:hf-to-gguf:blk.26.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} INFO:hf-to-gguf:blk.26.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} INFO:hf-to-gguf:blk.26.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} INFO:hf-to-gguf:blk.26.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} INFO:hf-to-gguf:blk.26.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.26.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.26.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.27.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.27.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} INFO:hf-to-gguf:blk.27.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.27.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.27.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.27.attn_k_norm.weight, torch.bfloat16 --> F32, shape = {256} INFO:hf-to-gguf:blk.27.attn_k.weight, torch.bfloat16 --> BF16, shape = {5120, 1024} INFO:hf-to-gguf:blk.27.attn_output.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} INFO:hf-to-gguf:blk.27.attn_q_norm.weight, torch.bfloat16 --> F32, shape = {256} INFO:hf-to-gguf:blk.27.attn_q.weight, torch.bfloat16 --> BF16, shape = {5120, 12288} INFO:hf-to-gguf:blk.27.attn_v.weight, torch.bfloat16 --> BF16, shape = {5120, 1024} INFO:hf-to-gguf:blk.28.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.28.ssm_a, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.28.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} INFO:hf-to-gguf:blk.28.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.28.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.28.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.28.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} INFO:hf-to-gguf:blk.28.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} INFO:hf-to-gguf:blk.28.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} INFO:hf-to-gguf:blk.28.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} INFO:hf-to-gguf:blk.28.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} INFO:hf-to-gguf:blk.28.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.28.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.28.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.29.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.29.ssm_a, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.29.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} INFO:hf-to-gguf:blk.29.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.29.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.29.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.29.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} INFO:hf-to-gguf:blk.29.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} INFO:hf-to-gguf:blk.29.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} INFO:hf-to-gguf:blk.29.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} INFO:hf-to-gguf:blk.29.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} INFO:hf-to-gguf:blk.29.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.29.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.29.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.30.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.30.ssm_a, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.30.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} INFO:hf-to-gguf:blk.30.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.30.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.30.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.30.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} INFO:hf-to-gguf:blk.30.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} INFO:hf-to-gguf:blk.30.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} INFO:hf-to-gguf:blk.30.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} INFO:hf-to-gguf:blk.30.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} INFO:hf-to-gguf:blk.30.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.30.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.30.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.31.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.31.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} INFO:hf-to-gguf:blk.31.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.31.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.31.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.31.attn_k_norm.weight, torch.bfloat16 --> F32, shape = {256} INFO:hf-to-gguf:blk.31.attn_k.weight, torch.bfloat16 --> BF16, shape = {5120, 1024} INFO:hf-to-gguf:blk.31.attn_output.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} INFO:hf-to-gguf:blk.31.attn_q_norm.weight, torch.bfloat16 --> F32, shape = {256} INFO:hf-to-gguf:blk.31.attn_q.weight, torch.bfloat16 --> BF16, shape = {5120, 12288} INFO:hf-to-gguf:blk.31.attn_v.weight, torch.bfloat16 --> BF16, shape = {5120, 1024} INFO:hf-to-gguf:blk.32.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.32.ssm_a, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.32.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} INFO:hf-to-gguf:blk.32.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.32.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.32.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.32.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} INFO:hf-to-gguf:blk.32.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} INFO:hf-to-gguf:blk.32.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} INFO:hf-to-gguf:blk.32.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} INFO:hf-to-gguf:blk.32.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} INFO:hf-to-gguf:blk.32.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.32.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.32.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.33.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.33.ssm_a, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.33.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} INFO:hf-to-gguf:blk.33.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.33.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.33.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.33.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} INFO:hf-to-gguf:blk.33.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} INFO:hf-to-gguf:blk.33.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} INFO:hf-to-gguf:blk.33.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} INFO:hf-to-gguf:blk.33.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} INFO:hf-to-gguf:blk.33.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.33.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.33.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.34.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.34.ssm_a, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.34.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} INFO:hf-to-gguf:blk.34.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.34.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.34.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.34.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} INFO:hf-to-gguf:blk.34.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} INFO:hf-to-gguf:blk.34.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} INFO:hf-to-gguf:blk.34.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} INFO:hf-to-gguf:blk.34.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} INFO:hf-to-gguf:blk.34.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.34.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.34.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.35.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.35.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} INFO:hf-to-gguf:blk.35.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.35.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.35.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.35.attn_k_norm.weight, torch.bfloat16 --> F32, shape = {256} INFO:hf-to-gguf:blk.35.attn_k.weight, torch.bfloat16 --> BF16, shape = {5120, 1024} INFO:hf-to-gguf:blk.35.attn_output.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} INFO:hf-to-gguf:blk.35.attn_q_norm.weight, torch.bfloat16 --> F32, shape = {256} INFO:hf-to-gguf:blk.35.attn_q.weight, torch.bfloat16 --> BF16, shape = {5120, 12288} INFO:hf-to-gguf:blk.35.attn_v.weight, torch.bfloat16 --> BF16, shape = {5120, 1024} INFO:hf-to-gguf:blk.36.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.36.ssm_a, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.36.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} INFO:hf-to-gguf:blk.36.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.36.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.36.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.36.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} INFO:hf-to-gguf:blk.36.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} INFO:hf-to-gguf:blk.36.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} INFO:hf-to-gguf:blk.36.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} INFO:hf-to-gguf:blk.36.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} INFO:hf-to-gguf:blk.36.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.36.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.36.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.37.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.37.ssm_a, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.37.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} INFO:hf-to-gguf:blk.37.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.37.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.37.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.37.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} INFO:hf-to-gguf:blk.37.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} INFO:hf-to-gguf:blk.37.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} INFO:hf-to-gguf:blk.37.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} INFO:hf-to-gguf:blk.37.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} INFO:hf-to-gguf:blk.37.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.37.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.37.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.38.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.38.ssm_a, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.38.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} INFO:hf-to-gguf:blk.38.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.38.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.38.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.38.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} INFO:hf-to-gguf:blk.38.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} INFO:hf-to-gguf:blk.38.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} INFO:hf-to-gguf:blk.38.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} INFO:hf-to-gguf:blk.38.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} INFO:hf-to-gguf:blk.38.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.38.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.38.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.39.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.39.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} INFO:hf-to-gguf:blk.39.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.39.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.39.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.39.attn_k_norm.weight, torch.bfloat16 --> F32, shape = {256} INFO:hf-to-gguf:blk.39.attn_k.weight, torch.bfloat16 --> BF16, shape = {5120, 1024} INFO:hf-to-gguf:blk.39.attn_output.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} INFO:hf-to-gguf:blk.39.attn_q_norm.weight, torch.bfloat16 --> F32, shape = {256} INFO:hf-to-gguf:blk.39.attn_q.weight, torch.bfloat16 --> BF16, shape = {5120, 12288} INFO:hf-to-gguf:blk.39.attn_v.weight, torch.bfloat16 --> BF16, shape = {5120, 1024} INFO:hf-to-gguf:blk.40.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.40.ssm_a, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.40.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} INFO:hf-to-gguf:blk.40.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.40.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.40.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.40.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} INFO:hf-to-gguf:blk.40.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} INFO:hf-to-gguf:blk.40.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} INFO:hf-to-gguf:blk.40.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} INFO:hf-to-gguf:blk.40.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} INFO:hf-to-gguf:blk.40.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.40.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.40.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.41.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.41.ssm_a, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.41.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} INFO:hf-to-gguf:blk.41.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.41.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.41.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.41.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} INFO:hf-to-gguf:blk.41.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} INFO:hf-to-gguf:blk.41.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} INFO:hf-to-gguf:blk.41.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} INFO:hf-to-gguf:blk.41.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} INFO:hf-to-gguf:blk.41.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.41.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.41.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.42.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.42.ssm_a, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.42.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} INFO:hf-to-gguf:blk.42.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.42.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.42.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.42.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} INFO:hf-to-gguf:blk.42.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} INFO:hf-to-gguf:blk.42.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} INFO:hf-to-gguf:blk.42.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} INFO:hf-to-gguf:blk.42.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} INFO:hf-to-gguf:blk.42.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.42.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.42.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.43.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.43.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} INFO:hf-to-gguf:blk.43.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.43.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.43.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.43.attn_k_norm.weight, torch.bfloat16 --> F32, shape = {256} INFO:hf-to-gguf:blk.43.attn_k.weight, torch.bfloat16 --> BF16, shape = {5120, 1024} INFO:hf-to-gguf:blk.43.attn_output.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} INFO:hf-to-gguf:blk.43.attn_q_norm.weight, torch.bfloat16 --> F32, shape = {256} INFO:hf-to-gguf:blk.43.attn_q.weight, torch.bfloat16 --> BF16, shape = {5120, 12288} INFO:hf-to-gguf:blk.43.attn_v.weight, torch.bfloat16 --> BF16, shape = {5120, 1024} INFO:hf-to-gguf:blk.44.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.44.ssm_a, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.44.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} INFO:hf-to-gguf:blk.44.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.44.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.44.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.44.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} INFO:hf-to-gguf:blk.44.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} INFO:hf-to-gguf:blk.44.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} INFO:hf-to-gguf:blk.44.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} INFO:hf-to-gguf:blk.44.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} INFO:hf-to-gguf:blk.44.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.44.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.44.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.45.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.45.ssm_a, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.45.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} INFO:hf-to-gguf:blk.45.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.45.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.45.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.45.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} INFO:hf-to-gguf:blk.45.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} INFO:hf-to-gguf:blk.45.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} INFO:hf-to-gguf:blk.45.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} INFO:hf-to-gguf:blk.45.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} INFO:hf-to-gguf:blk.45.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.45.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.45.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.46.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.46.ssm_a, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.46.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} INFO:hf-to-gguf:blk.46.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.46.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.46.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.46.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} INFO:hf-to-gguf:blk.46.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} INFO:hf-to-gguf:blk.46.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} INFO:hf-to-gguf:blk.46.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} INFO:hf-to-gguf:blk.46.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} INFO:hf-to-gguf:blk.46.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.46.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.46.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.47.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.47.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} INFO:hf-to-gguf:blk.47.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.47.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.47.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.47.attn_k_norm.weight, torch.bfloat16 --> F32, shape = {256} INFO:hf-to-gguf:blk.47.attn_k.weight, torch.bfloat16 --> BF16, shape = {5120, 1024} INFO:hf-to-gguf:blk.47.attn_output.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} INFO:hf-to-gguf:blk.47.attn_q_norm.weight, torch.bfloat16 --> F32, shape = {256} INFO:hf-to-gguf:blk.47.attn_q.weight, torch.bfloat16 --> BF16, shape = {5120, 12288} INFO:hf-to-gguf:blk.47.attn_v.weight, torch.bfloat16 --> BF16, shape = {5120, 1024} INFO:hf-to-gguf:blk.48.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.48.ssm_a, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.48.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} INFO:hf-to-gguf:blk.48.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.48.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.48.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.48.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} INFO:hf-to-gguf:blk.48.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} INFO:hf-to-gguf:blk.48.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} INFO:hf-to-gguf:blk.48.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} INFO:hf-to-gguf:blk.48.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} INFO:hf-to-gguf:blk.48.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.48.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.48.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.49.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.49.ssm_a, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.49.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} INFO:hf-to-gguf:blk.49.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.49.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.49.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.49.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} INFO:hf-to-gguf:blk.49.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} INFO:hf-to-gguf:blk.49.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} INFO:hf-to-gguf:blk.49.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} INFO:hf-to-gguf:blk.49.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} INFO:hf-to-gguf:blk.49.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.49.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.49.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.50.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.50.ssm_a, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.50.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} INFO:hf-to-gguf:blk.50.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.50.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.50.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.50.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} INFO:hf-to-gguf:blk.50.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} INFO:hf-to-gguf:blk.50.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} INFO:hf-to-gguf:blk.50.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} INFO:hf-to-gguf:blk.50.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} INFO:hf-to-gguf:blk.50.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.50.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.50.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.51.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.51.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} INFO:hf-to-gguf:blk.51.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.51.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.51.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.51.attn_k_norm.weight, torch.bfloat16 --> F32, shape = {256} INFO:hf-to-gguf:blk.51.attn_k.weight, torch.bfloat16 --> BF16, shape = {5120, 1024} INFO:hf-to-gguf:blk.51.attn_output.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} INFO:hf-to-gguf:blk.51.attn_q_norm.weight, torch.bfloat16 --> F32, shape = {256} INFO:hf-to-gguf:blk.51.attn_q.weight, torch.bfloat16 --> BF16, shape = {5120, 12288} INFO:hf-to-gguf:blk.51.attn_v.weight, torch.bfloat16 --> BF16, shape = {5120, 1024} INFO:hf-to-gguf:blk.52.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.52.ssm_a, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.52.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} INFO:hf-to-gguf:blk.52.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.52.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.52.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.52.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} INFO:hf-to-gguf:blk.52.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} INFO:hf-to-gguf:blk.52.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} INFO:hf-to-gguf:blk.52.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} INFO:hf-to-gguf:blk.52.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} INFO:hf-to-gguf:blk.52.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.52.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.52.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.53.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.53.ssm_a, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.53.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} INFO:hf-to-gguf:blk.53.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.53.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.53.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.53.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} INFO:hf-to-gguf:blk.53.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} INFO:hf-to-gguf:blk.53.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} INFO:hf-to-gguf:blk.53.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} INFO:hf-to-gguf:blk.53.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} INFO:hf-to-gguf:blk.53.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.53.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.53.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.54.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.54.ssm_a, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.54.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} INFO:hf-to-gguf:blk.54.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.54.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.54.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.54.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} INFO:hf-to-gguf:blk.54.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} INFO:hf-to-gguf:blk.54.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} INFO:hf-to-gguf:blk.54.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} INFO:hf-to-gguf:blk.54.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} INFO:hf-to-gguf:blk.54.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.54.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.54.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.55.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.55.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} INFO:hf-to-gguf:blk.55.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.55.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.55.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.55.attn_k_norm.weight, torch.bfloat16 --> F32, shape = {256} INFO:hf-to-gguf:blk.55.attn_k.weight, torch.bfloat16 --> BF16, shape = {5120, 1024} INFO:hf-to-gguf:blk.55.attn_output.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} INFO:hf-to-gguf:blk.55.attn_q_norm.weight, torch.bfloat16 --> F32, shape = {256} INFO:hf-to-gguf:blk.55.attn_q.weight, torch.bfloat16 --> BF16, shape = {5120, 12288} INFO:hf-to-gguf:blk.55.attn_v.weight, torch.bfloat16 --> BF16, shape = {5120, 1024} INFO:hf-to-gguf:blk.56.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.56.ssm_a, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.56.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} INFO:hf-to-gguf:blk.56.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.56.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.56.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.56.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} INFO:hf-to-gguf:blk.56.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} INFO:hf-to-gguf:blk.56.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} INFO:hf-to-gguf:blk.56.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} INFO:hf-to-gguf:blk.56.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} INFO:hf-to-gguf:blk.56.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.56.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.56.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.57.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.57.ssm_a, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.57.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} INFO:hf-to-gguf:blk.57.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.57.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.57.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.57.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} INFO:hf-to-gguf:blk.57.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} INFO:hf-to-gguf:blk.57.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} INFO:hf-to-gguf:blk.57.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} INFO:hf-to-gguf:blk.57.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} INFO:hf-to-gguf:blk.57.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.57.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.57.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.58.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.58.ssm_a, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.58.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} INFO:hf-to-gguf:blk.58.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.58.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.58.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.58.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} INFO:hf-to-gguf:blk.58.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} INFO:hf-to-gguf:blk.58.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} INFO:hf-to-gguf:blk.58.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} INFO:hf-to-gguf:blk.58.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} INFO:hf-to-gguf:blk.58.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.58.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.58.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.59.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.59.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} INFO:hf-to-gguf:blk.59.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.59.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.59.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.59.attn_k_norm.weight, torch.bfloat16 --> F32, shape = {256} INFO:hf-to-gguf:blk.59.attn_k.weight, torch.bfloat16 --> BF16, shape = {5120, 1024} INFO:hf-to-gguf:blk.59.attn_output.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} INFO:hf-to-gguf:blk.59.attn_q_norm.weight, torch.bfloat16 --> F32, shape = {256} INFO:hf-to-gguf:blk.59.attn_q.weight, torch.bfloat16 --> BF16, shape = {5120, 12288} INFO:hf-to-gguf:blk.59.attn_v.weight, torch.bfloat16 --> BF16, shape = {5120, 1024} INFO:hf-to-gguf:blk.60.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.60.ssm_a, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.60.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} INFO:hf-to-gguf:blk.60.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.60.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.60.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.60.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} INFO:hf-to-gguf:blk.60.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} INFO:hf-to-gguf:blk.60.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} INFO:hf-to-gguf:blk.60.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} INFO:hf-to-gguf:blk.60.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} INFO:hf-to-gguf:blk.60.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.60.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.60.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.61.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.61.ssm_a, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.61.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} INFO:hf-to-gguf:blk.61.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.61.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.61.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.61.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} INFO:hf-to-gguf:blk.61.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} INFO:hf-to-gguf:blk.61.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} INFO:hf-to-gguf:blk.61.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} INFO:hf-to-gguf:blk.61.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} INFO:hf-to-gguf:blk.61.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.61.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.61.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.62.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.62.ssm_a, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.62.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} INFO:hf-to-gguf:blk.62.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} INFO:hf-to-gguf:blk.62.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.62.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} INFO:hf-to-gguf:blk.62.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} INFO:hf-to-gguf:blk.62.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} INFO:hf-to-gguf:blk.62.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} INFO:hf-to-gguf:blk.62.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} INFO:hf-to-gguf:blk.62.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} INFO:hf-to-gguf:blk.62.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.62.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.62.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.63.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.63.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} INFO:hf-to-gguf:blk.63.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.63.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} INFO:hf-to-gguf:blk.63.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:blk.63.attn_k_norm.weight, torch.bfloat16 --> F32, shape = {256} INFO:hf-to-gguf:blk.63.attn_k.weight, torch.bfloat16 --> BF16, shape = {5120, 1024} INFO:hf-to-gguf:blk.63.attn_output.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} INFO:hf-to-gguf:blk.63.attn_q_norm.weight, torch.bfloat16 --> F32, shape = {256} INFO:hf-to-gguf:blk.63.attn_q.weight, torch.bfloat16 --> BF16, shape = {5120, 12288} INFO:hf-to-gguf:blk.63.attn_v.weight, torch.bfloat16 --> BF16, shape = {5120, 1024} INFO:hf-to-gguf:output_norm.weight, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:Set meta model INFO:hf-to-gguf:Set model parameters INFO:hf-to-gguf:gguf: context length = 262144 INFO:hf-to-gguf:gguf: embedding length = 5120 INFO:hf-to-gguf:gguf: feed forward length = 17408 INFO:hf-to-gguf:gguf: head count = 24 INFO:hf-to-gguf:gguf: key-value head count = 4 WARNING:hf-to-gguf:Unknown RoPE type: default INFO:hf-to-gguf:gguf: rope scaling type = NONE INFO:hf-to-gguf:gguf: mrope sections: [11, 11, 10, 0] INFO:hf-to-gguf:gguf: rope theta = 10000000 INFO:hf-to-gguf:gguf: rms norm epsilon = 1e-06 INFO:hf-to-gguf:gguf: file type = 32 INFO:hf-to-gguf:Set model quantization version INFO:hf-to-gguf:Set model tokenizer [transformers] Missing validation function in 'RotaryEmbeddingConfigMixin' for 'rope_type'='axial' INFO:gguf.vocab:Adding 247587 merge(s). INFO:gguf.vocab:Setting special token type eos to 248046 INFO:gguf.vocab:Setting special token type pad to 248044 INFO:gguf.vocab:Setting special token type bos to 248044 INFO:gguf.vocab:Setting add_bos_token to False INFO:gguf.vocab:Setting add_eos_token to False INFO:gguf.vocab:Setting chat_template to {%- set image_count = namespace(value=0) %} {%- set video_count = namespace(value=0) %} {%- macro render_content(content, do_vision_count, is_system_content=false) %} {%- if content is string %} {{- content }} {%- elif content is iterable and content is not mapping %} {%- for item in content %} {%- if 'image' in item or 'image_url' in item or item.type == 'image' %} {%- if is_system_content %} {{- raise_exception('System message cannot contain images.') }} {%- endif %} {%- if do_vision_count %} {%- set image_count.value = image_count.value + 1 %} {%- endif %} {%- if add_vision_id %} {{- 'Picture ' ~ image_count.value ~ ': ' }} {%- endif %} {{- '<|vision_start|><|image_pad|><|vision_end|>' }} {%- elif 'video' in item or item.type == 'video' %} {%- if is_system_content %} {{- raise_exception('System message cannot contain videos.') }} {%- endif %} {%- if do_vision_count %} {%- set video_count.value = video_count.value + 1 %} {%- endif %} {%- if add_vision_id %} {{- 'Video ' ~ video_count.value ~ ': ' }} {%- endif %} {{- '<|vision_start|><|video_pad|><|vision_end|>' }} {%- elif 'text' in item %} {{- item.text }} {%- else %} {{- raise_exception('Unexpected item type in content.') }} {%- endif %} {%- endfor %} {%- elif content is none or content is undefined %} {{- '' }} {%- else %} {{- raise_exception('Unexpected content type.') }} {%- endif %} {%- endmacro %} {%- if not messages %} {{- raise_exception('No messages provided.') }} {%- endif %} {%- set reasoning_instructions = '' %} {%- if enable_thinking is undefined or enable_thinking is true %} {%- set resolved_reasoning_effort = reasoning_effort|default('xhigh') %} {%- if resolved_reasoning_effort not in ('xhigh', 'medium', 'low') %} {{- raise_exception('Unexpected reasoning effort ' ~ reasoning_effort ~ '. Supported types are xhigh (default), medium, and low.') }} {%- endif %} {%- if resolved_reasoning_effort == 'xhigh' %} {%- set reasoning_instructions = 'Reasoning effort is set to xhigh. Please think carefully through the task, validate key assumptions, consider plausible alternatives, and prioritize correctness, consistency, and clarity in the final answer.' %} {%- elif resolved_reasoning_effort == 'low' %} {%- set reasoning_instructions = 'Reasoning effort is set to low. Keep your thinking brief and focused, moving directly to the conclusion without unnecessary elaboration.' %} {%- endif %} {%- endif %} {%- if tools and tools is iterable and tools is not mapping %} {{- '<|im_start|>system\n' }} {%- if reasoning_instructions %} {{- reasoning_instructions + '\n\n' }} {%- endif %} {{- "# Tools\n\nYou have access to the following functions:\n\n" }} {%- for tool in tools %} {{- "\n" }} {{- tool | tojson }} {%- endfor %} {{- "\n" }} {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n\n\n\nvalue_1\n\n\nThis is the value for the second parameter\nthat can span\nmultiple lines\n\n\n\n\n\nReminder:\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n' }} {%- if messages[0].role == 'system' %} {%- set content = render_content(messages[0].content, false, true)|trim %} {%- if content %} {{- '\n\n' + content }} {%- endif %} {%- endif %} {{- '<|im_end|>\n' }} {%- else %} {%- if messages[0].role == 'system' %} {%- set content = render_content(messages[0].content, false, true)|trim %} {%- if content %} {{- '<|im_start|>system\n' + (reasoning_instructions + '\n\n' if reasoning_instructions else '') + content + '<|im_end|>\n' }} {%- elif reasoning_instructions %} {{- '<|im_start|>system\n' + reasoning_instructions + '<|im_end|>\n' }} {%- endif %} {%- elif reasoning_instructions %} {{- '<|im_start|>system\n' + reasoning_instructions + '<|im_end|>\n' }} {%- endif %} {%- endif %} {%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %} {%- for message in messages[::-1] %} {%- set index = (messages|length - 1) - loop.index0 %} {%- if ns.multi_step_tool and message.role == "user" %} {%- set content = render_content(message.content, false)|trim %} {%- if not(content.startswith('') and content.endswith('')) %} {%- set ns.multi_step_tool = false %} {%- set ns.last_query_index = index %} {%- endif %} {%- endif %} {%- endfor %} {%- if ns.multi_step_tool %} {{- raise_exception('No user query found in messages.') }} {%- endif %} {%- for message in messages %} {%- set content = render_content(message.content, true)|trim %} {%- if message.role == "system" %} {%- if not loop.first %} {{- raise_exception('System message must be at the beginning.') }} {%- endif %} {%- elif message.role == "user" %} {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }} {%- elif message.role == "assistant" %} {%- set reasoning_content = '' %} {%- if message.reasoning_content is string %} {%- set reasoning_content = message.reasoning_content %} {%- endif %} {%- set reasoning_content = reasoning_content|trim %} {%- if preserve_thinking is undefined or preserve_thinking is true or loop.index0 > ns.last_query_index %} {{- '<|im_start|>' + message.role + '\n\n' + reasoning_content + '\n\n\n' + content }} {%- else %} {{- '<|im_start|>' + message.role + '\n' + content }} {%- endif %} {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %} {%- for tool_call in message.tool_calls %} {%- if tool_call.function is defined %} {%- set tool_call = tool_call.function %} {%- endif %} {%- if loop.first %} {%- if content|trim %} {{- '\n\n\n\n' }} {%- else %} {{- '\n\n' }} {%- endif %} {%- else %} {{- '\n\n\n' }} {%- endif %} {%- if tool_call.arguments is defined and tool_call.arguments != '' %} {%- for args_name, args_value in tool_call.arguments|items %} {{- '\n' }} {%- set args_value = args_value | string if args_value is string else args_value | tojson | safe %} {{- args_value }} {{- '\n\n' }} {%- endfor %} {%- endif %} {{- '\n' }} {%- endfor %} {%- endif %} {{- '<|im_end|>\n' }} {%- elif message.role == "tool" %} {%- if loop.previtem and loop.previtem.role != "tool" %} {{- '<|im_start|>user' }} {%- endif %} {{- '\n\n' }} {{- content }} {{- '\n' }} {%- if not loop.last and loop.nextitem.role != "tool" %} {{- '<|im_end|>\n' }} {%- elif loop.last %} {{- '<|im_end|>\n' }} {%- endif %} {%- else %} {{- raise_exception('Unexpected message role.') }} {%- endif %} {%- endfor %} {%- if add_generation_prompt %} {{- '<|im_start|>assistant\n' }} {%- if enable_thinking is defined and enable_thinking is false %} {{- '\n\n\n\n' }} {%- else %} {{- '\n' }} {%- endif %} {%- endif %} INFO:gguf.gguf_writer:Writing the following files: INFO:gguf.gguf_writer:upload-OpenJev/OpenJev-BF16.gguf: n_tensors = 851, total_size = 53.8G Writing: 0%| | 0.00/53.8G [00:00 F32, shape = {1152} INFO:hf-to-gguf:v.blk.0.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152} INFO:hf-to-gguf:v.blk.0.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} INFO:hf-to-gguf:v.blk.0.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456} INFO:hf-to-gguf:v.blk.0.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} INFO:hf-to-gguf:v.blk.0.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304} INFO:hf-to-gguf:v.blk.0.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.0.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152} INFO:hf-to-gguf:v.blk.0.ln1.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.0.ln1.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.0.ln2.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.0.ln2.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.1.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.1.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152} INFO:hf-to-gguf:v.blk.1.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} INFO:hf-to-gguf:v.blk.1.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456} INFO:hf-to-gguf:v.blk.1.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} INFO:hf-to-gguf:v.blk.1.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304} INFO:hf-to-gguf:v.blk.1.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.1.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152} INFO:hf-to-gguf:v.blk.1.ln1.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.1.ln1.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.1.ln2.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.1.ln2.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.10.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.10.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152} INFO:hf-to-gguf:v.blk.10.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} INFO:hf-to-gguf:v.blk.10.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456} INFO:hf-to-gguf:v.blk.10.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} INFO:hf-to-gguf:v.blk.10.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304} INFO:hf-to-gguf:v.blk.10.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.10.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152} INFO:hf-to-gguf:v.blk.10.ln1.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.10.ln1.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.10.ln2.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.10.ln2.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.11.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.11.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152} INFO:hf-to-gguf:v.blk.11.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} INFO:hf-to-gguf:v.blk.11.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456} INFO:hf-to-gguf:v.blk.11.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} INFO:hf-to-gguf:v.blk.11.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304} INFO:hf-to-gguf:v.blk.11.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.11.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152} INFO:hf-to-gguf:v.blk.11.ln1.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.11.ln1.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.11.ln2.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.11.ln2.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.12.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.12.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152} INFO:hf-to-gguf:v.blk.12.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} INFO:hf-to-gguf:v.blk.12.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456} INFO:hf-to-gguf:v.blk.12.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} INFO:hf-to-gguf:v.blk.12.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304} INFO:hf-to-gguf:v.blk.12.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.12.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152} INFO:hf-to-gguf:v.blk.12.ln1.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.12.ln1.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.12.ln2.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.12.ln2.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.13.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.13.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152} INFO:hf-to-gguf:v.blk.13.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} INFO:hf-to-gguf:v.blk.13.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456} INFO:hf-to-gguf:v.blk.13.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} INFO:hf-to-gguf:v.blk.13.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304} INFO:hf-to-gguf:v.blk.13.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.13.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152} INFO:hf-to-gguf:v.blk.13.ln1.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.13.ln1.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.13.ln2.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.13.ln2.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.14.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.14.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152} INFO:hf-to-gguf:v.blk.14.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} INFO:hf-to-gguf:v.blk.14.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456} INFO:hf-to-gguf:v.blk.14.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} INFO:hf-to-gguf:v.blk.14.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304} INFO:hf-to-gguf:v.blk.14.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.14.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152} INFO:hf-to-gguf:v.blk.14.ln1.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.14.ln1.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.14.ln2.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.14.ln2.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.15.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.15.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152} INFO:hf-to-gguf:v.blk.15.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} INFO:hf-to-gguf:v.blk.15.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456} INFO:hf-to-gguf:v.blk.15.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} INFO:hf-to-gguf:v.blk.15.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304} INFO:hf-to-gguf:v.blk.15.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.15.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152} INFO:hf-to-gguf:v.blk.15.ln1.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.15.ln1.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.15.ln2.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.15.ln2.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.16.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.16.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152} INFO:hf-to-gguf:v.blk.16.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} INFO:hf-to-gguf:v.blk.16.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456} INFO:hf-to-gguf:v.blk.16.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} INFO:hf-to-gguf:v.blk.16.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304} INFO:hf-to-gguf:v.blk.16.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.16.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152} INFO:hf-to-gguf:v.blk.16.ln1.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.16.ln1.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.16.ln2.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.16.ln2.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.17.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.17.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152} INFO:hf-to-gguf:v.blk.17.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} INFO:hf-to-gguf:v.blk.17.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456} INFO:hf-to-gguf:v.blk.17.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} INFO:hf-to-gguf:v.blk.17.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304} INFO:hf-to-gguf:v.blk.17.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.17.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152} INFO:hf-to-gguf:v.blk.17.ln1.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.17.ln1.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.17.ln2.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.17.ln2.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.18.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.18.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152} INFO:hf-to-gguf:v.blk.18.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} INFO:hf-to-gguf:v.blk.18.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456} INFO:hf-to-gguf:v.blk.18.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} INFO:hf-to-gguf:v.blk.18.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304} INFO:hf-to-gguf:v.blk.18.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.18.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152} INFO:hf-to-gguf:v.blk.18.ln1.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.18.ln1.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.18.ln2.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.18.ln2.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.19.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.19.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152} INFO:hf-to-gguf:v.blk.19.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} INFO:hf-to-gguf:v.blk.19.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456} INFO:hf-to-gguf:v.blk.19.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} INFO:hf-to-gguf:v.blk.19.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304} INFO:hf-to-gguf:v.blk.19.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.19.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152} INFO:hf-to-gguf:v.blk.19.ln1.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.19.ln1.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.19.ln2.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.19.ln2.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.2.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.2.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152} INFO:hf-to-gguf:v.blk.2.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} INFO:hf-to-gguf:v.blk.2.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456} INFO:hf-to-gguf:v.blk.2.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} INFO:hf-to-gguf:v.blk.2.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304} INFO:hf-to-gguf:v.blk.2.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.2.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152} INFO:hf-to-gguf:v.blk.2.ln1.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.2.ln1.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.2.ln2.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.2.ln2.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.20.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.20.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152} INFO:hf-to-gguf:v.blk.20.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} INFO:hf-to-gguf:v.blk.20.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456} INFO:hf-to-gguf:v.blk.20.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} INFO:hf-to-gguf:v.blk.20.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304} INFO:hf-to-gguf:v.blk.20.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.20.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152} INFO:hf-to-gguf:v.blk.20.ln1.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.20.ln1.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.20.ln2.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.20.ln2.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.21.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.21.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152} INFO:hf-to-gguf:v.blk.21.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} INFO:hf-to-gguf:v.blk.21.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456} INFO:hf-to-gguf:v.blk.21.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} INFO:hf-to-gguf:v.blk.21.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304} INFO:hf-to-gguf:v.blk.21.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.21.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152} INFO:hf-to-gguf:v.blk.21.ln1.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.21.ln1.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.21.ln2.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.21.ln2.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.22.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.22.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152} INFO:hf-to-gguf:v.blk.22.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} INFO:hf-to-gguf:v.blk.22.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456} INFO:hf-to-gguf:v.blk.22.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} INFO:hf-to-gguf:v.blk.22.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304} INFO:hf-to-gguf:v.blk.22.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.22.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152} INFO:hf-to-gguf:v.blk.22.ln1.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.22.ln1.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.22.ln2.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.22.ln2.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.23.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.23.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152} INFO:hf-to-gguf:v.blk.23.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} INFO:hf-to-gguf:v.blk.23.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456} INFO:hf-to-gguf:v.blk.23.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} INFO:hf-to-gguf:v.blk.23.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304} INFO:hf-to-gguf:v.blk.23.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.23.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152} INFO:hf-to-gguf:v.blk.23.ln1.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.23.ln1.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.23.ln2.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.23.ln2.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.24.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.24.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152} INFO:hf-to-gguf:v.blk.24.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} INFO:hf-to-gguf:v.blk.24.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456} INFO:hf-to-gguf:v.blk.24.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} INFO:hf-to-gguf:v.blk.24.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304} INFO:hf-to-gguf:v.blk.24.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.24.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152} INFO:hf-to-gguf:v.blk.24.ln1.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.24.ln1.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.24.ln2.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.24.ln2.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.25.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.25.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152} INFO:hf-to-gguf:v.blk.25.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} INFO:hf-to-gguf:v.blk.25.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456} INFO:hf-to-gguf:v.blk.25.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} INFO:hf-to-gguf:v.blk.25.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304} INFO:hf-to-gguf:v.blk.25.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.25.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152} INFO:hf-to-gguf:v.blk.25.ln1.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.25.ln1.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.25.ln2.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.25.ln2.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.26.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.26.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152} INFO:hf-to-gguf:v.blk.26.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} INFO:hf-to-gguf:v.blk.26.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456} INFO:hf-to-gguf:v.blk.26.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} INFO:hf-to-gguf:v.blk.26.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304} INFO:hf-to-gguf:v.blk.26.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.26.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152} INFO:hf-to-gguf:v.blk.26.ln1.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.26.ln1.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.26.ln2.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.26.ln2.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.3.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.3.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152} INFO:hf-to-gguf:v.blk.3.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} INFO:hf-to-gguf:v.blk.3.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456} INFO:hf-to-gguf:v.blk.3.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} INFO:hf-to-gguf:v.blk.3.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304} INFO:hf-to-gguf:v.blk.3.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.3.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152} INFO:hf-to-gguf:v.blk.3.ln1.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.3.ln1.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.3.ln2.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.3.ln2.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.4.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.4.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152} INFO:hf-to-gguf:v.blk.4.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} INFO:hf-to-gguf:v.blk.4.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456} INFO:hf-to-gguf:v.blk.4.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} INFO:hf-to-gguf:v.blk.4.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304} INFO:hf-to-gguf:v.blk.4.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.4.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152} INFO:hf-to-gguf:v.blk.4.ln1.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.4.ln1.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.4.ln2.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.4.ln2.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.5.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.5.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152} INFO:hf-to-gguf:v.blk.5.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} INFO:hf-to-gguf:v.blk.5.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456} INFO:hf-to-gguf:v.blk.5.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} INFO:hf-to-gguf:v.blk.5.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304} INFO:hf-to-gguf:v.blk.5.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.5.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152} INFO:hf-to-gguf:v.blk.5.ln1.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.5.ln1.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.5.ln2.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.5.ln2.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.6.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.6.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152} INFO:hf-to-gguf:v.blk.6.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} INFO:hf-to-gguf:v.blk.6.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456} INFO:hf-to-gguf:v.blk.6.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} INFO:hf-to-gguf:v.blk.6.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304} INFO:hf-to-gguf:v.blk.6.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.6.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152} INFO:hf-to-gguf:v.blk.6.ln1.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.6.ln1.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.6.ln2.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.6.ln2.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.7.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.7.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152} INFO:hf-to-gguf:v.blk.7.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} INFO:hf-to-gguf:v.blk.7.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456} INFO:hf-to-gguf:v.blk.7.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} INFO:hf-to-gguf:v.blk.7.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304} INFO:hf-to-gguf:v.blk.7.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.7.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152} INFO:hf-to-gguf:v.blk.7.ln1.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.7.ln1.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.7.ln2.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.7.ln2.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.8.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.8.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152} INFO:hf-to-gguf:v.blk.8.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} INFO:hf-to-gguf:v.blk.8.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456} INFO:hf-to-gguf:v.blk.8.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} INFO:hf-to-gguf:v.blk.8.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304} INFO:hf-to-gguf:v.blk.8.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.8.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152} INFO:hf-to-gguf:v.blk.8.ln1.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.8.ln1.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.8.ln2.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.8.ln2.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.9.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.9.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152} INFO:hf-to-gguf:v.blk.9.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} INFO:hf-to-gguf:v.blk.9.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456} INFO:hf-to-gguf:v.blk.9.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} INFO:hf-to-gguf:v.blk.9.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304} INFO:hf-to-gguf:v.blk.9.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.9.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152} INFO:hf-to-gguf:v.blk.9.ln1.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.9.ln1.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.9.ln2.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.9.ln2.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:mm.0.bias, torch.bfloat16 --> F32, shape = {4608} INFO:hf-to-gguf:mm.0.weight, torch.bfloat16 --> BF16, shape = {4608, 4608} INFO:hf-to-gguf:mm.2.bias, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:mm.2.weight, torch.bfloat16 --> BF16, shape = {4608, 5120} INFO:hf-to-gguf:v.post_ln.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.post_ln.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.patch_embd.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.patch_embd.weight, torch.bfloat16 --> F32, shape = {16, 16, 3, 1152} INFO:hf-to-gguf:v.patch_embd.weight.1, torch.bfloat16 --> F32, shape = {16, 16, 3, 1152} INFO:hf-to-gguf:v.position_embd.weight, torch.bfloat16 --> F32, shape = {1152, 2304} INFO:hf-to-gguf:Set meta model INFO:hf-to-gguf:Set model parameters INFO:hf-to-gguf:Set model quantization version INFO:gguf.gguf_writer:Writing the following files: INFO:gguf.gguf_writer:upload-OpenJev/mmproj-OpenJev-BF16.gguf: n_tensors = 334, total_size = 931.1M Writing: 0%| | 0.00/931M [00:00 1288.28 MiB [ 2/ 851] output_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 3/ 851] token_embd.weight - [ 5120, 248320, 1, 1], type = bf16, converting to q8_0 .. size = 2425.00 MiB -> 1288.28 MiB [ 4/ 851] blk.0.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 5/ 851] blk.0.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 6/ 851] blk.0.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 7/ 851] blk.0.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 8/ 851] blk.0.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 9/ 851] blk.0.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 10/ 851] blk.0.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 11/ 851] blk.0.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 12/ 851] blk.0.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 13/ 851] blk.0.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 14/ 851] blk.0.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 15/ 851] blk.0.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 16/ 851] blk.0.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 17/ 851] blk.0.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 18/ 851] blk.1.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 19/ 851] blk.1.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 20/ 851] blk.1.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 21/ 851] blk.1.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 22/ 851] blk.1.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 23/ 851] blk.1.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 24/ 851] blk.1.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 25/ 851] blk.1.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 26/ 851] blk.1.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 27/ 851] blk.1.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 28/ 851] blk.1.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 29/ 851] blk.1.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 30/ 851] blk.1.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 31/ 851] blk.1.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 32/ 851] blk.2.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 33/ 851] blk.2.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 34/ 851] blk.2.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 35/ 851] blk.2.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 36/ 851] blk.2.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 37/ 851] blk.2.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 38/ 851] blk.2.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 39/ 851] blk.2.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 40/ 851] blk.2.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 41/ 851] blk.2.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 42/ 851] blk.2.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 43/ 851] blk.2.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 44/ 851] blk.2.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 45/ 851] blk.2.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 46/ 851] blk.3.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB [ 47/ 851] blk.3.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB [ 48/ 851] blk.3.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 49/ 851] blk.3.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 50/ 851] blk.3.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB [ 51/ 851] blk.3.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB [ 52/ 851] blk.3.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB [ 53/ 851] blk.3.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 54/ 851] blk.3.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 55/ 851] blk.3.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 56/ 851] blk.3.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 57/ 851] blk.4.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 58/ 851] blk.4.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 59/ 851] blk.4.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 60/ 851] blk.4.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 61/ 851] blk.4.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 62/ 851] blk.4.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 63/ 851] blk.4.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 64/ 851] blk.4.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 65/ 851] blk.4.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 66/ 851] blk.4.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 67/ 851] blk.4.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 68/ 851] blk.4.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 69/ 851] blk.4.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 70/ 851] blk.4.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 71/ 851] blk.5.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 72/ 851] blk.5.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 73/ 851] blk.5.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 74/ 851] blk.5.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 75/ 851] blk.5.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 76/ 851] blk.5.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 77/ 851] blk.5.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 78/ 851] blk.5.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 79/ 851] blk.5.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 80/ 851] blk.5.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 81/ 851] blk.5.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 82/ 851] blk.5.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 83/ 851] blk.5.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 84/ 851] blk.5.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 85/ 851] blk.6.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 86/ 851] blk.6.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 87/ 851] blk.6.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 88/ 851] blk.6.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 89/ 851] blk.6.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 90/ 851] blk.6.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 91/ 851] blk.6.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 92/ 851] blk.6.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 93/ 851] blk.6.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 94/ 851] blk.6.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 95/ 851] blk.6.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 96/ 851] blk.6.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 97/ 851] blk.6.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 98/ 851] blk.6.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 99/ 851] blk.7.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB [ 100/ 851] blk.7.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB [ 101/ 851] blk.7.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 102/ 851] blk.7.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 103/ 851] blk.7.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB [ 104/ 851] blk.7.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB [ 105/ 851] blk.7.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB [ 106/ 851] blk.7.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 107/ 851] blk.7.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 108/ 851] blk.7.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 109/ 851] blk.7.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 110/ 851] blk.8.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 111/ 851] blk.8.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 112/ 851] blk.8.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 113/ 851] blk.8.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 114/ 851] blk.8.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 115/ 851] blk.8.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 116/ 851] blk.8.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 117/ 851] blk.8.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 118/ 851] blk.8.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 119/ 851] blk.8.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 120/ 851] blk.8.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 121/ 851] blk.8.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 122/ 851] blk.8.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 123/ 851] blk.8.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 124/ 851] blk.9.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 125/ 851] blk.9.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 126/ 851] blk.9.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 127/ 851] blk.9.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 128/ 851] blk.9.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 129/ 851] blk.9.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 130/ 851] blk.9.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 131/ 851] blk.9.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 132/ 851] blk.9.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 133/ 851] blk.9.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 134/ 851] blk.9.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 135/ 851] blk.9.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 136/ 851] blk.9.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 137/ 851] blk.9.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 138/ 851] blk.10.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 139/ 851] blk.10.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 140/ 851] blk.10.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 141/ 851] blk.10.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 142/ 851] blk.10.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 143/ 851] blk.10.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 144/ 851] blk.10.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 145/ 851] blk.10.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 146/ 851] blk.10.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 147/ 851] blk.10.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 148/ 851] blk.10.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 149/ 851] blk.10.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 150/ 851] blk.10.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 151/ 851] blk.10.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 152/ 851] blk.11.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB [ 153/ 851] blk.11.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB [ 154/ 851] blk.11.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 155/ 851] blk.11.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 156/ 851] blk.11.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB [ 157/ 851] blk.11.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB [ 158/ 851] blk.11.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB [ 159/ 851] blk.11.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 160/ 851] blk.11.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 161/ 851] blk.11.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 162/ 851] blk.11.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 163/ 851] blk.12.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 164/ 851] blk.12.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 165/ 851] blk.12.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 166/ 851] blk.12.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 167/ 851] blk.12.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 168/ 851] blk.12.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 169/ 851] blk.12.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 170/ 851] blk.12.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 171/ 851] blk.12.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 172/ 851] blk.12.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 173/ 851] blk.12.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 174/ 851] blk.12.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 175/ 851] blk.12.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 176/ 851] blk.12.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 177/ 851] blk.13.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 178/ 851] blk.13.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 179/ 851] blk.13.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 180/ 851] blk.13.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 181/ 851] blk.13.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 182/ 851] blk.13.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 183/ 851] blk.13.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 184/ 851] blk.13.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 185/ 851] blk.13.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 186/ 851] blk.13.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 187/ 851] blk.13.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 188/ 851] blk.13.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 189/ 851] blk.13.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 190/ 851] blk.13.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 191/ 851] blk.14.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 192/ 851] blk.14.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 193/ 851] blk.14.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 194/ 851] blk.14.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 195/ 851] blk.14.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 196/ 851] blk.14.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 197/ 851] blk.14.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 198/ 851] blk.14.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 199/ 851] blk.14.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 200/ 851] blk.14.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 201/ 851] blk.14.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 202/ 851] blk.14.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 203/ 851] blk.14.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 204/ 851] blk.14.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 205/ 851] blk.15.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB [ 206/ 851] blk.15.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB [ 207/ 851] blk.15.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 208/ 851] blk.15.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 209/ 851] blk.15.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB [ 210/ 851] blk.15.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB [ 211/ 851] blk.15.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB [ 212/ 851] blk.15.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 213/ 851] blk.15.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 214/ 851] blk.15.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 215/ 851] blk.15.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 216/ 851] blk.16.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 217/ 851] blk.16.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 218/ 851] blk.16.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 219/ 851] blk.16.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 220/ 851] blk.16.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 221/ 851] blk.16.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 222/ 851] blk.16.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 223/ 851] blk.16.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 224/ 851] blk.16.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 225/ 851] blk.16.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 226/ 851] blk.16.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 227/ 851] blk.16.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 228/ 851] blk.16.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 229/ 851] blk.16.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 230/ 851] blk.17.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 231/ 851] blk.17.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 232/ 851] blk.17.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 233/ 851] blk.17.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 234/ 851] blk.17.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 235/ 851] blk.17.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 236/ 851] blk.17.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 237/ 851] blk.17.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 238/ 851] blk.17.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 239/ 851] blk.17.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 240/ 851] blk.17.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 241/ 851] blk.17.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 242/ 851] blk.17.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 243/ 851] blk.17.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 244/ 851] blk.18.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 245/ 851] blk.18.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 246/ 851] blk.18.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 247/ 851] blk.18.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 248/ 851] blk.18.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 249/ 851] blk.18.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 250/ 851] blk.18.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 251/ 851] blk.18.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 252/ 851] blk.18.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 253/ 851] blk.18.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 254/ 851] blk.18.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 255/ 851] blk.18.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 256/ 851] blk.18.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 257/ 851] blk.18.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 258/ 851] blk.19.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB [ 259/ 851] blk.19.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB [ 260/ 851] blk.19.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 261/ 851] blk.19.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 262/ 851] blk.19.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB [ 263/ 851] blk.19.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB [ 264/ 851] blk.19.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB [ 265/ 851] blk.19.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 266/ 851] blk.19.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 267/ 851] blk.19.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 268/ 851] blk.19.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 269/ 851] blk.20.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 270/ 851] blk.20.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 271/ 851] blk.20.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 272/ 851] blk.20.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 273/ 851] blk.20.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 274/ 851] blk.20.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 275/ 851] blk.20.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 276/ 851] blk.20.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 277/ 851] blk.20.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 278/ 851] blk.20.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 279/ 851] blk.20.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 280/ 851] blk.20.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 281/ 851] blk.20.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 282/ 851] blk.20.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 283/ 851] blk.21.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 284/ 851] blk.21.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 285/ 851] blk.21.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 286/ 851] blk.21.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 287/ 851] blk.21.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 288/ 851] blk.21.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 289/ 851] blk.21.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 290/ 851] blk.21.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 291/ 851] blk.21.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 292/ 851] blk.21.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 293/ 851] blk.21.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 294/ 851] blk.21.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 295/ 851] blk.21.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 296/ 851] blk.21.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 297/ 851] blk.22.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 298/ 851] blk.22.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 299/ 851] blk.22.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 300/ 851] blk.22.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 301/ 851] blk.22.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 302/ 851] blk.22.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 303/ 851] blk.22.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 304/ 851] blk.22.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 305/ 851] blk.22.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 306/ 851] blk.22.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 307/ 851] blk.22.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 308/ 851] blk.22.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 309/ 851] blk.22.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 310/ 851] blk.22.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 311/ 851] blk.23.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB [ 312/ 851] blk.23.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB [ 313/ 851] blk.23.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 314/ 851] blk.23.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 315/ 851] blk.23.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB [ 316/ 851] blk.23.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB [ 317/ 851] blk.23.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB [ 318/ 851] blk.23.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 319/ 851] blk.23.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 320/ 851] blk.23.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 321/ 851] blk.23.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 322/ 851] blk.24.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 323/ 851] blk.24.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 324/ 851] blk.24.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 325/ 851] blk.24.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 326/ 851] blk.24.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 327/ 851] blk.24.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 328/ 851] blk.24.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 329/ 851] blk.24.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 330/ 851] blk.24.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 331/ 851] blk.24.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 332/ 851] blk.24.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 333/ 851] blk.24.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 334/ 851] blk.24.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 335/ 851] blk.24.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 336/ 851] blk.25.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 337/ 851] blk.25.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 338/ 851] blk.25.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 339/ 851] blk.25.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 340/ 851] blk.25.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 341/ 851] blk.25.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 342/ 851] blk.25.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 343/ 851] blk.25.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 344/ 851] blk.25.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 345/ 851] blk.25.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 346/ 851] blk.25.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 347/ 851] blk.25.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 348/ 851] blk.25.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 349/ 851] blk.25.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 350/ 851] blk.26.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 351/ 851] blk.26.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 352/ 851] blk.26.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 353/ 851] blk.26.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 354/ 851] blk.26.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 355/ 851] blk.26.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 356/ 851] blk.26.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 357/ 851] blk.26.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 358/ 851] blk.26.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 359/ 851] blk.26.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 360/ 851] blk.26.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 361/ 851] blk.26.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 362/ 851] blk.26.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 363/ 851] blk.26.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 364/ 851] blk.27.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB [ 365/ 851] blk.27.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB [ 366/ 851] blk.27.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 367/ 851] blk.27.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 368/ 851] blk.27.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB [ 369/ 851] blk.27.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB [ 370/ 851] blk.27.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB [ 371/ 851] blk.27.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 372/ 851] blk.27.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 373/ 851] blk.27.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 374/ 851] blk.27.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 375/ 851] blk.28.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 376/ 851] blk.28.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 377/ 851] blk.28.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 378/ 851] blk.28.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 379/ 851] blk.28.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 380/ 851] blk.28.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 381/ 851] blk.28.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 382/ 851] blk.28.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 383/ 851] blk.28.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 384/ 851] blk.28.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 385/ 851] blk.28.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 386/ 851] blk.28.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 387/ 851] blk.28.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 388/ 851] blk.28.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 389/ 851] blk.29.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 390/ 851] blk.29.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 391/ 851] blk.29.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 392/ 851] blk.29.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 393/ 851] blk.29.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 394/ 851] blk.29.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 395/ 851] blk.29.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 396/ 851] blk.29.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 397/ 851] blk.29.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 398/ 851] blk.29.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 399/ 851] blk.29.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 400/ 851] blk.29.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 401/ 851] blk.29.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 402/ 851] blk.29.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 403/ 851] blk.30.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 404/ 851] blk.30.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 405/ 851] blk.30.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 406/ 851] blk.30.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 407/ 851] blk.30.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 408/ 851] blk.30.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 409/ 851] blk.30.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 410/ 851] blk.30.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 411/ 851] blk.30.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 412/ 851] blk.30.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 413/ 851] blk.30.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 414/ 851] blk.30.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 415/ 851] blk.30.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 416/ 851] blk.30.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 417/ 851] blk.31.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB [ 418/ 851] blk.31.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB [ 419/ 851] blk.31.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 420/ 851] blk.31.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 421/ 851] blk.31.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB [ 422/ 851] blk.31.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB [ 423/ 851] blk.31.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB [ 424/ 851] blk.31.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 425/ 851] blk.31.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 426/ 851] blk.31.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 427/ 851] blk.31.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 428/ 851] blk.32.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 429/ 851] blk.32.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 430/ 851] blk.32.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 431/ 851] blk.32.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 432/ 851] blk.32.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 433/ 851] blk.32.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 434/ 851] blk.32.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 435/ 851] blk.32.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 436/ 851] blk.32.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 437/ 851] blk.32.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 438/ 851] blk.32.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 439/ 851] blk.32.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 440/ 851] blk.32.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 441/ 851] blk.32.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 442/ 851] blk.33.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 443/ 851] blk.33.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 444/ 851] blk.33.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 445/ 851] blk.33.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 446/ 851] blk.33.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 447/ 851] blk.33.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 448/ 851] blk.33.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 449/ 851] blk.33.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 450/ 851] blk.33.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 451/ 851] blk.33.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 452/ 851] blk.33.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 453/ 851] blk.33.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 454/ 851] blk.33.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 455/ 851] blk.33.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 456/ 851] blk.34.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 457/ 851] blk.34.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 458/ 851] blk.34.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 459/ 851] blk.34.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 460/ 851] blk.34.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 461/ 851] blk.34.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 462/ 851] blk.34.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 463/ 851] blk.34.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 464/ 851] blk.34.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 465/ 851] blk.34.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 466/ 851] blk.34.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 467/ 851] blk.34.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 468/ 851] blk.34.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 469/ 851] blk.34.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 470/ 851] blk.35.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB [ 471/ 851] blk.35.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB [ 472/ 851] blk.35.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 473/ 851] blk.35.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 474/ 851] blk.35.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB [ 475/ 851] blk.35.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB [ 476/ 851] blk.35.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB [ 477/ 851] blk.35.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 478/ 851] blk.35.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 479/ 851] blk.35.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 480/ 851] blk.35.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 481/ 851] blk.36.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 482/ 851] blk.36.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 483/ 851] blk.36.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 484/ 851] blk.36.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 485/ 851] blk.36.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 486/ 851] blk.36.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 487/ 851] blk.36.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 488/ 851] blk.36.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 489/ 851] blk.36.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 490/ 851] blk.36.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 491/ 851] blk.36.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 492/ 851] blk.36.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 493/ 851] blk.36.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 494/ 851] blk.36.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 495/ 851] blk.37.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 496/ 851] blk.37.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 497/ 851] blk.37.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 498/ 851] blk.37.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 499/ 851] blk.37.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 500/ 851] blk.37.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 501/ 851] blk.37.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 502/ 851] blk.37.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 503/ 851] blk.37.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 504/ 851] blk.37.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 505/ 851] blk.37.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 506/ 851] blk.37.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 507/ 851] blk.37.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 508/ 851] blk.37.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 509/ 851] blk.38.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 510/ 851] blk.38.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 511/ 851] blk.38.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 512/ 851] blk.38.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 513/ 851] blk.38.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 514/ 851] blk.38.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 515/ 851] blk.38.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 516/ 851] blk.38.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 517/ 851] blk.38.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 518/ 851] blk.38.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 519/ 851] blk.38.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 520/ 851] blk.38.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 521/ 851] blk.38.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 522/ 851] blk.38.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 523/ 851] blk.39.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB [ 524/ 851] blk.39.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB [ 525/ 851] blk.39.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 526/ 851] blk.39.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 527/ 851] blk.39.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB [ 528/ 851] blk.39.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB [ 529/ 851] blk.39.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB [ 530/ 851] blk.39.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 531/ 851] blk.39.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 532/ 851] blk.39.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 533/ 851] blk.39.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 534/ 851] blk.40.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 535/ 851] blk.40.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 536/ 851] blk.40.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 537/ 851] blk.40.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 538/ 851] blk.40.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 539/ 851] blk.40.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 540/ 851] blk.40.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 541/ 851] blk.40.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 542/ 851] blk.40.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 543/ 851] blk.40.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 544/ 851] blk.40.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 545/ 851] blk.40.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 546/ 851] blk.40.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 547/ 851] blk.40.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 548/ 851] blk.41.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 549/ 851] blk.41.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 550/ 851] blk.41.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 551/ 851] blk.41.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 552/ 851] blk.41.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 553/ 851] blk.41.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 554/ 851] blk.41.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 555/ 851] blk.41.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 556/ 851] blk.41.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 557/ 851] blk.41.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 558/ 851] blk.41.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 559/ 851] blk.41.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 560/ 851] blk.41.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 561/ 851] blk.41.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 562/ 851] blk.42.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 563/ 851] blk.42.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 564/ 851] blk.42.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 565/ 851] blk.42.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 566/ 851] blk.42.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 567/ 851] blk.42.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 568/ 851] blk.42.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 569/ 851] blk.42.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 570/ 851] blk.42.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 571/ 851] blk.42.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 572/ 851] blk.42.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 573/ 851] blk.42.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 574/ 851] blk.42.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 575/ 851] blk.42.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 576/ 851] blk.43.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB [ 577/ 851] blk.43.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB [ 578/ 851] blk.43.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 579/ 851] blk.43.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 580/ 851] blk.43.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB [ 581/ 851] blk.43.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB [ 582/ 851] blk.43.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB [ 583/ 851] blk.43.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 584/ 851] blk.43.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 585/ 851] blk.43.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 586/ 851] blk.43.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 587/ 851] blk.44.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 588/ 851] blk.44.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 589/ 851] blk.44.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 590/ 851] blk.44.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 591/ 851] blk.44.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 592/ 851] blk.44.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 593/ 851] blk.44.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 594/ 851] blk.44.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 595/ 851] blk.44.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 596/ 851] blk.44.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 597/ 851] blk.44.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 598/ 851] blk.44.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 599/ 851] blk.44.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 600/ 851] blk.44.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 601/ 851] blk.45.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 602/ 851] blk.45.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 603/ 851] blk.45.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 604/ 851] blk.45.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 605/ 851] blk.45.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 606/ 851] blk.45.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 607/ 851] blk.45.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 608/ 851] blk.45.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 609/ 851] blk.45.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 610/ 851] blk.45.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 611/ 851] blk.45.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 612/ 851] blk.45.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 613/ 851] blk.45.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 614/ 851] blk.45.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 615/ 851] blk.46.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 616/ 851] blk.46.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 617/ 851] blk.46.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 618/ 851] blk.46.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 619/ 851] blk.46.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 620/ 851] blk.46.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 621/ 851] blk.46.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 622/ 851] blk.46.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 623/ 851] blk.46.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 624/ 851] blk.46.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 625/ 851] blk.46.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 626/ 851] blk.46.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 627/ 851] blk.46.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 628/ 851] blk.46.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 629/ 851] blk.47.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB [ 630/ 851] blk.47.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB [ 631/ 851] blk.47.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 632/ 851] blk.47.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 633/ 851] blk.47.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB [ 634/ 851] blk.47.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB [ 635/ 851] blk.47.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB [ 636/ 851] blk.47.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 637/ 851] blk.47.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 638/ 851] blk.47.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 639/ 851] blk.47.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 640/ 851] blk.48.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 641/ 851] blk.48.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 642/ 851] blk.48.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 643/ 851] blk.48.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 644/ 851] blk.48.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 645/ 851] blk.48.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 646/ 851] blk.48.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 647/ 851] blk.48.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 648/ 851] blk.48.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 649/ 851] blk.48.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 650/ 851] blk.48.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 651/ 851] blk.48.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 652/ 851] blk.48.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 653/ 851] blk.48.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 654/ 851] blk.49.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 655/ 851] blk.49.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 656/ 851] blk.49.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 657/ 851] blk.49.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 658/ 851] blk.49.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 659/ 851] blk.49.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 660/ 851] blk.49.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 661/ 851] blk.49.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 662/ 851] blk.49.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 663/ 851] blk.49.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 664/ 851] blk.49.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 665/ 851] blk.49.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 666/ 851] blk.49.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 667/ 851] blk.49.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 668/ 851] blk.50.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 669/ 851] blk.50.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 670/ 851] blk.50.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 671/ 851] blk.50.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 672/ 851] blk.50.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 673/ 851] blk.50.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 674/ 851] blk.50.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 675/ 851] blk.50.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 676/ 851] blk.50.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 677/ 851] blk.50.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 678/ 851] blk.50.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 679/ 851] blk.50.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 680/ 851] blk.50.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 681/ 851] blk.50.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 682/ 851] blk.51.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB [ 683/ 851] blk.51.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB [ 684/ 851] blk.51.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 685/ 851] blk.51.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 686/ 851] blk.51.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB [ 687/ 851] blk.51.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB [ 688/ 851] blk.51.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB [ 689/ 851] blk.51.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 690/ 851] blk.51.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 691/ 851] blk.51.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 692/ 851] blk.51.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 693/ 851] blk.52.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 694/ 851] blk.52.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 695/ 851] blk.52.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 696/ 851] blk.52.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 697/ 851] blk.52.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 698/ 851] blk.52.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 699/ 851] blk.52.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 700/ 851] blk.52.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 701/ 851] blk.52.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 702/ 851] blk.52.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 703/ 851] blk.52.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 704/ 851] blk.52.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 705/ 851] blk.52.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 706/ 851] blk.52.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 707/ 851] blk.53.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 708/ 851] blk.53.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 709/ 851] blk.53.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 710/ 851] blk.53.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 711/ 851] blk.53.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 712/ 851] blk.53.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 713/ 851] blk.53.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 714/ 851] blk.53.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 715/ 851] blk.53.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 716/ 851] blk.53.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 717/ 851] blk.53.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 718/ 851] blk.53.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 719/ 851] blk.53.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 720/ 851] blk.53.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 721/ 851] blk.54.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 722/ 851] blk.54.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 723/ 851] blk.54.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 724/ 851] blk.54.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 725/ 851] blk.54.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 726/ 851] blk.54.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 727/ 851] blk.54.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 728/ 851] blk.54.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 729/ 851] blk.54.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 730/ 851] blk.54.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 731/ 851] blk.54.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 732/ 851] blk.54.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 733/ 851] blk.54.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 734/ 851] blk.54.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 735/ 851] blk.55.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB [ 736/ 851] blk.55.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB [ 737/ 851] blk.55.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 738/ 851] blk.55.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 739/ 851] blk.55.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB [ 740/ 851] blk.55.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB [ 741/ 851] blk.55.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB [ 742/ 851] blk.55.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 743/ 851] blk.55.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 744/ 851] blk.55.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 745/ 851] blk.55.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 746/ 851] blk.56.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 747/ 851] blk.56.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 748/ 851] blk.56.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 749/ 851] blk.56.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 750/ 851] blk.56.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 751/ 851] blk.56.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 752/ 851] blk.56.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 753/ 851] blk.56.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 754/ 851] blk.56.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 755/ 851] blk.56.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 756/ 851] blk.56.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 757/ 851] blk.56.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 758/ 851] blk.56.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 759/ 851] blk.56.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 760/ 851] blk.57.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 761/ 851] blk.57.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 762/ 851] blk.57.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 763/ 851] blk.57.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 764/ 851] blk.57.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 765/ 851] blk.57.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 766/ 851] blk.57.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 767/ 851] blk.57.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 768/ 851] blk.57.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 769/ 851] blk.57.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 770/ 851] blk.57.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 771/ 851] blk.57.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 772/ 851] blk.57.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 773/ 851] blk.57.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 774/ 851] blk.58.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 775/ 851] blk.58.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 776/ 851] blk.58.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 777/ 851] blk.58.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 778/ 851] blk.58.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 779/ 851] blk.58.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 780/ 851] blk.58.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 781/ 851] blk.58.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 782/ 851] blk.58.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 783/ 851] blk.58.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 784/ 851] blk.58.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 785/ 851] blk.58.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 786/ 851] blk.58.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 787/ 851] blk.58.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 788/ 851] blk.59.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB [ 789/ 851] blk.59.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB [ 790/ 851] blk.59.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 791/ 851] blk.59.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 792/ 851] blk.59.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB [ 793/ 851] blk.59.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB [ 794/ 851] blk.59.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB [ 795/ 851] blk.59.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 796/ 851] blk.59.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 797/ 851] blk.59.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 798/ 851] blk.59.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 799/ 851] blk.60.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 800/ 851] blk.60.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 801/ 851] blk.60.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 802/ 851] blk.60.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 803/ 851] blk.60.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 804/ 851] blk.60.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 805/ 851] blk.60.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 806/ 851] blk.60.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 807/ 851] blk.60.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 808/ 851] blk.60.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 809/ 851] blk.60.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 810/ 851] blk.60.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 811/ 851] blk.60.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 812/ 851] blk.60.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 813/ 851] blk.61.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 814/ 851] blk.61.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 815/ 851] blk.61.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 816/ 851] blk.61.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 817/ 851] blk.61.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 818/ 851] blk.61.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 819/ 851] blk.61.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 820/ 851] blk.61.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 821/ 851] blk.61.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 822/ 851] blk.61.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 823/ 851] blk.61.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 824/ 851] blk.61.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 825/ 851] blk.61.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 826/ 851] blk.61.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 827/ 851] blk.62.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 828/ 851] blk.62.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 829/ 851] blk.62.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 830/ 851] blk.62.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 831/ 851] blk.62.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 832/ 851] blk.62.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 833/ 851] blk.62.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 834/ 851] blk.62.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 835/ 851] blk.62.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 836/ 851] blk.62.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 837/ 851] blk.62.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 838/ 851] blk.62.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 839/ 851] blk.62.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 840/ 851] blk.62.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 841/ 851] blk.63.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB [ 842/ 851] blk.63.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB [ 843/ 851] blk.63.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 844/ 851] blk.63.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 845/ 851] blk.63.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB [ 846/ 851] blk.63.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB [ 847/ 851] blk.63.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB [ 848/ 851] blk.63.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 849/ 851] blk.63.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 850/ 851] blk.63.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB [ 851/ 851] blk.63.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB llama_model_quantize_impl: model size = 51305.09 MiB (16.00 BPW) llama_model_quantize_impl: quant size = 27260.56 MiB (8.50 BPW) llama_quantize: quantize time = 22409.99 ms llama_quantize: total time = 22409.99 ms + llama.cpp/build/bin/llama-quantize --pure --tensor-type output.weight=q6_k --tensor-type attn_=q8_0 --tensor-type ssm_=q8_0 ./upload-OpenJev/OpenJev-BF16.gguf ./upload-OpenJev/OpenJev-Q4_K_M.gguf Q4_K_M version: 0.5.0-dev (build 11354, commit 54a0c5da9) built with GNU 14.2.0 for Linux x86_64 llama_quantize: quantizing './upload-OpenJev/OpenJev-BF16.gguf' to './upload-OpenJev/OpenJev-Q4_K_M.gguf' as Q4_K_M llama_model_loader: loaded meta data with 48 key-value pairs and 851 tensors from ./upload-OpenJev/OpenJev-BF16.gguf (version GGUF V3 (latest)) llama_model_loader: Dumping metadata keys/values. Note: KV overrides do not apply in this output. llama_model_loader: - kv 0: general.architecture str = qwen35 llama_model_loader: - kv 1: general.type str = model llama_model_loader: - kv 2: general.sampling.top_k i32 = 20 llama_model_loader: - kv 3: general.sampling.top_p f32 = 0.950000 llama_model_loader: - kv 4: general.sampling.temp f32 = 1.000000 llama_model_loader: - kv 5: general.name str = OpenJev llama_model_loader: - kv 6: general.size_label str = 27B llama_model_loader: - kv 7: general.license str = cc-by-nc-4.0 llama_model_loader: - kv 8: general.tags arr[str,8] = ["decision-model", "zero-shot-classif... llama_model_loader: - kv 9: general.languages arr[str,6] = ["en", "de", "fr", "hi", "zh", "ja"] llama_model_loader: - kv 10: qwen35.block_count u32 = 64 llama_model_loader: - kv 11: qwen35.context_length u32 = 262144 llama_model_loader: - kv 12: qwen35.embedding_length u32 = 5120 llama_model_loader: - kv 13: qwen35.feed_forward_length u32 = 17408 llama_model_loader: - kv 14: qwen35.attention.head_count u32 = 24 llama_model_loader: - kv 15: qwen35.attention.head_count_kv u32 = 4 llama_model_loader: - kv 16: qwen35.rope.dimension_sections arr[i32,4] = [11, 11, 10, 0] llama_model_loader: - kv 17: qwen35.rope.freq_base f32 = 10000000.000000 llama_model_loader: - kv 18: qwen35.attention.layer_norm_rms_epsilon f32 = 0.000001 llama_model_loader: - kv 19: qwen35.attention.key_length u32 = 256 llama_model_loader: - kv 20: qwen35.attention.value_length u32 = 256 llama_model_loader: - kv 21: general.file_type u32 = 32 llama_model_loader: - kv 22: qwen35.ssm.conv_kernel u32 = 4 llama_model_loader: - kv 23: qwen35.ssm.state_size u32 = 128 llama_model_loader: - kv 24: qwen35.ssm.group_count u32 = 16 llama_model_loader: - kv 25: qwen35.ssm.time_step_rank u32 = 48 llama_model_loader: - kv 26: qwen35.ssm.inner_size u32 = 6144 llama_model_loader: - kv 27: qwen35.attention.recurrent_layers arr[bool,64] = [true, true, true, false, true, true,... llama_model_loader: - kv 28: qwen35.full_attention_interval u32 = 4 llama_model_loader: - kv 29: qwen35.rope.dimension_count u32 = 64 llama_model_loader: - kv 30: qwen35.decision.type str = openjev llama_model_loader: - kv 31: qwen35.decision.temperature.choice f32 = 0.850000 llama_model_loader: - kv 32: qwen35.decision.temperature.score f32 = 0.850000 llama_model_loader: - kv 33: qwen35.decision.temperature.noul f32 = 1.554713 llama_model_loader: - kv 34: general.quantization_version u32 = 2 llama_model_loader: - kv 35: tokenizer.ggml.model str = gpt2 llama_model_loader: - kv 36: tokenizer.ggml.pre str = qwen35 llama_model_loader: - kv 37: tokenizer.ggml.tokens arr[str,248320] = ["!", "\"", "#", "$", "%", "&", "'", ... llama_model_loader: - kv 38: tokenizer.ggml.token_type arr[i32,248320] = [1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, ... llama_model_loader: - kv 39: tokenizer.ggml.merges arr[str,247587] = ["Ġ Ġ", "ĠĠ ĠĠ", "i n", "Ġ t",... llama_model_loader: - kv 40: tokenizer.ggml.eos_token_id u32 = 248046 llama_model_loader: - kv 41: tokenizer.ggml.padding_token_id u32 = 248044 llama_model_loader: - kv 42: tokenizer.ggml.bos_token_id u32 = 248044 llama_model_loader: - kv 43: tokenizer.ggml.add_bos_token bool = false llama_model_loader: - kv 44: tokenizer.ggml.add_eos_token bool = false llama_model_loader: - kv 45: tokenizer.chat_template str = {%- set image_count = namespace(value... llama_model_loader: - kv 46: tokenizer.chat_template.systemone str = {% set letters = 'ABCDEFGHIJKLMNOPQRS... llama_model_loader: - kv 47: tokenizer.chat_templates arr[str,1] = ["systemone"] llama_model_loader: - type f32: 353 tensors llama_model_loader: - type bf16: 498 tensors llama_tensor_get_type: output.weight - applying manual override: q4_K -> q6_K llama_tensor_get_type: blk.0.attn_gate.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.0.attn_qkv.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.0.ssm_alpha.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.0.ssm_beta.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.0.ssm_out.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.1.attn_gate.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.1.attn_qkv.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.1.ssm_alpha.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.1.ssm_beta.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.1.ssm_out.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.2.attn_gate.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.2.attn_qkv.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.2.ssm_alpha.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.2.ssm_beta.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.2.ssm_out.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.3.attn_k.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.3.attn_output.weight - applying manual override: q4_K -> q6_K llama_tensor_get_type: blk.3.attn_q.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.3.attn_v.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.4.attn_gate.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.4.attn_qkv.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.4.ssm_alpha.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.4.ssm_beta.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.4.ssm_out.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.5.attn_gate.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.5.attn_qkv.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.5.ssm_alpha.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.5.ssm_beta.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.5.ssm_out.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.6.attn_gate.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.6.attn_qkv.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.6.ssm_alpha.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.6.ssm_beta.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.6.ssm_out.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.7.attn_k.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.7.attn_output.weight - applying manual override: q4_K -> q6_K llama_tensor_get_type: blk.7.attn_q.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.7.attn_v.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.8.attn_gate.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.8.attn_qkv.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.8.ssm_alpha.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.8.ssm_beta.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.8.ssm_out.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.9.attn_gate.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.9.attn_qkv.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.9.ssm_alpha.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.9.ssm_beta.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.9.ssm_out.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.10.attn_gate.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.10.attn_qkv.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.10.ssm_alpha.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.10.ssm_beta.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.10.ssm_out.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.11.attn_k.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.11.attn_output.weight - applying manual override: q4_K -> q6_K llama_tensor_get_type: blk.11.attn_q.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.11.attn_v.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.12.attn_gate.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.12.attn_qkv.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.12.ssm_alpha.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.12.ssm_beta.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.12.ssm_out.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.13.attn_gate.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.13.attn_qkv.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.13.ssm_alpha.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.13.ssm_beta.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.13.ssm_out.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.14.attn_gate.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.14.attn_qkv.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.14.ssm_alpha.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.14.ssm_beta.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.14.ssm_out.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.15.attn_k.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.15.attn_output.weight - applying manual override: q4_K -> q6_K llama_tensor_get_type: blk.15.attn_q.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.15.attn_v.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.16.attn_gate.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.16.attn_qkv.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.16.ssm_alpha.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.16.ssm_beta.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.16.ssm_out.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.17.attn_gate.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.17.attn_qkv.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.17.ssm_alpha.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.17.ssm_beta.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.17.ssm_out.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.18.attn_gate.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.18.attn_qkv.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.18.ssm_alpha.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.18.ssm_beta.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.18.ssm_out.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.19.attn_k.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.19.attn_output.weight - applying manual override: q4_K -> q6_K llama_tensor_get_type: blk.19.attn_q.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.19.attn_v.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.20.attn_gate.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.20.attn_qkv.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.20.ssm_alpha.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.20.ssm_beta.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.20.ssm_out.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.21.attn_gate.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.21.attn_qkv.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.21.ssm_alpha.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.21.ssm_beta.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.21.ssm_out.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.22.attn_gate.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.22.attn_qkv.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.22.ssm_alpha.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.22.ssm_beta.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.22.ssm_out.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.23.attn_k.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.23.attn_output.weight - applying manual override: q4_K -> q6_K llama_tensor_get_type: blk.23.attn_q.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.23.attn_v.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.24.attn_gate.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.24.attn_qkv.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.24.ssm_alpha.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.24.ssm_beta.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.24.ssm_out.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.25.attn_gate.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.25.attn_qkv.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.25.ssm_alpha.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.25.ssm_beta.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.25.ssm_out.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.26.attn_gate.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.26.attn_qkv.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.26.ssm_alpha.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.26.ssm_beta.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.26.ssm_out.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.27.attn_k.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.27.attn_output.weight - applying manual override: q4_K -> q6_K llama_tensor_get_type: blk.27.attn_q.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.27.attn_v.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.28.attn_gate.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.28.attn_qkv.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.28.ssm_alpha.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.28.ssm_beta.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.28.ssm_out.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.29.attn_gate.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.29.attn_qkv.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.29.ssm_alpha.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.29.ssm_beta.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.29.ssm_out.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.30.attn_gate.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.30.attn_qkv.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.30.ssm_alpha.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.30.ssm_beta.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.30.ssm_out.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.31.attn_k.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.31.attn_output.weight - applying manual override: q4_K -> q6_K llama_tensor_get_type: blk.31.attn_q.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.31.attn_v.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.32.attn_gate.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.32.attn_qkv.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.32.ssm_alpha.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.32.ssm_beta.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.32.ssm_out.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.33.attn_gate.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.33.attn_qkv.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.33.ssm_alpha.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.33.ssm_beta.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.33.ssm_out.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.34.attn_gate.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.34.attn_qkv.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.34.ssm_alpha.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.34.ssm_beta.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.34.ssm_out.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.35.attn_k.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.35.attn_output.weight - applying manual override: q4_K -> q6_K llama_tensor_get_type: blk.35.attn_q.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.35.attn_v.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.36.attn_gate.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.36.attn_qkv.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.36.ssm_alpha.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.36.ssm_beta.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.36.ssm_out.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.37.attn_gate.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.37.attn_qkv.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.37.ssm_alpha.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.37.ssm_beta.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.37.ssm_out.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.38.attn_gate.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.38.attn_qkv.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.38.ssm_alpha.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.38.ssm_beta.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.38.ssm_out.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.39.attn_k.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.39.attn_output.weight - applying manual override: q4_K -> q6_K llama_tensor_get_type: blk.39.attn_q.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.39.attn_v.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.40.attn_gate.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.40.attn_qkv.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.40.ssm_alpha.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.40.ssm_beta.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.40.ssm_out.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.41.attn_gate.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.41.attn_qkv.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.41.ssm_alpha.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.41.ssm_beta.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.41.ssm_out.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.42.attn_gate.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.42.attn_qkv.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.42.ssm_alpha.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.42.ssm_beta.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.42.ssm_out.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.43.attn_k.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.43.attn_output.weight - applying manual override: q4_K -> q6_K llama_tensor_get_type: blk.43.attn_q.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.43.attn_v.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.44.attn_gate.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.44.attn_qkv.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.44.ssm_alpha.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.44.ssm_beta.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.44.ssm_out.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.45.attn_gate.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.45.attn_qkv.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.45.ssm_alpha.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.45.ssm_beta.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.45.ssm_out.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.46.attn_gate.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.46.attn_qkv.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.46.ssm_alpha.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.46.ssm_beta.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.46.ssm_out.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.47.attn_k.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.47.attn_output.weight - applying manual override: q4_K -> q6_K llama_tensor_get_type: blk.47.attn_q.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.47.attn_v.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.48.attn_gate.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.48.attn_qkv.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.48.ssm_alpha.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.48.ssm_beta.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.48.ssm_out.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.49.attn_gate.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.49.attn_qkv.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.49.ssm_alpha.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.49.ssm_beta.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.49.ssm_out.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.50.attn_gate.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.50.attn_qkv.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.50.ssm_alpha.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.50.ssm_beta.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.50.ssm_out.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.51.attn_k.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.51.attn_output.weight - applying manual override: q4_K -> q6_K llama_tensor_get_type: blk.51.attn_q.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.51.attn_v.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.52.attn_gate.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.52.attn_qkv.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.52.ssm_alpha.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.52.ssm_beta.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.52.ssm_out.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.53.attn_gate.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.53.attn_qkv.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.53.ssm_alpha.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.53.ssm_beta.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.53.ssm_out.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.54.attn_gate.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.54.attn_qkv.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.54.ssm_alpha.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.54.ssm_beta.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.54.ssm_out.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.55.attn_k.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.55.attn_output.weight - applying manual override: q4_K -> q6_K llama_tensor_get_type: blk.55.attn_q.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.55.attn_v.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.56.attn_gate.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.56.attn_qkv.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.56.ssm_alpha.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.56.ssm_beta.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.56.ssm_out.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.57.attn_gate.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.57.attn_qkv.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.57.ssm_alpha.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.57.ssm_beta.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.57.ssm_out.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.58.attn_gate.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.58.attn_qkv.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.58.ssm_alpha.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.58.ssm_beta.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.58.ssm_out.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.59.attn_k.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.59.attn_output.weight - applying manual override: q4_K -> q6_K llama_tensor_get_type: blk.59.attn_q.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.59.attn_v.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.60.attn_gate.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.60.attn_qkv.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.60.ssm_alpha.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.60.ssm_beta.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.60.ssm_out.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.61.attn_gate.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.61.attn_qkv.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.61.ssm_alpha.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.61.ssm_beta.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.61.ssm_out.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.62.attn_gate.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.62.attn_qkv.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.62.ssm_alpha.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.62.ssm_beta.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.62.ssm_out.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.63.attn_k.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.63.attn_output.weight - applying manual override: q4_K -> q6_K llama_tensor_get_type: blk.63.attn_q.weight - applying manual override: q4_K -> q8_0 llama_tensor_get_type: blk.63.attn_v.weight - applying manual override: q4_K -> q8_0 [ 1/ 851] output.weight - [ 5120, 248320, 1, 1], type = bf16, converting to q6_K .. size = 2425.00 MiB -> 994.63 MiB [ 2/ 851] output_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 3/ 851] token_embd.weight - [ 5120, 248320, 1, 1], type = bf16, converting to q4_K .. size = 2425.00 MiB -> 682.03 MiB [ 4/ 851] blk.0.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 5/ 851] blk.0.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 6/ 851] blk.0.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 7/ 851] blk.0.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 8/ 851] blk.0.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 9/ 851] blk.0.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 10/ 851] blk.0.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 11/ 851] blk.0.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 12/ 851] blk.0.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 13/ 851] blk.0.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 14/ 851] blk.0.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 15/ 851] blk.0.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 16/ 851] blk.0.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 17/ 851] blk.0.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 18/ 851] blk.1.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 19/ 851] blk.1.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 20/ 851] blk.1.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 21/ 851] blk.1.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 22/ 851] blk.1.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 23/ 851] blk.1.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 24/ 851] blk.1.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 25/ 851] blk.1.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 26/ 851] blk.1.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 27/ 851] blk.1.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 28/ 851] blk.1.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 29/ 851] blk.1.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 30/ 851] blk.1.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 31/ 851] blk.1.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 32/ 851] blk.2.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 33/ 851] blk.2.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 34/ 851] blk.2.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 35/ 851] blk.2.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 36/ 851] blk.2.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 37/ 851] blk.2.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 38/ 851] blk.2.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 39/ 851] blk.2.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 40/ 851] blk.2.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 41/ 851] blk.2.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 42/ 851] blk.2.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 43/ 851] blk.2.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 44/ 851] blk.2.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 45/ 851] blk.2.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 46/ 851] blk.3.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB [ 47/ 851] blk.3.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB [ 48/ 851] blk.3.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 49/ 851] blk.3.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q6_K .. size = 60.00 MiB -> 24.61 MiB [ 50/ 851] blk.3.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB [ 51/ 851] blk.3.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB [ 52/ 851] blk.3.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB [ 53/ 851] blk.3.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 54/ 851] blk.3.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 55/ 851] blk.3.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 56/ 851] blk.3.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 57/ 851] blk.4.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 58/ 851] blk.4.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 59/ 851] blk.4.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 60/ 851] blk.4.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 61/ 851] blk.4.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 62/ 851] blk.4.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 63/ 851] blk.4.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 64/ 851] blk.4.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 65/ 851] blk.4.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 66/ 851] blk.4.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 67/ 851] blk.4.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 68/ 851] blk.4.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 69/ 851] blk.4.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 70/ 851] blk.4.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 71/ 851] blk.5.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 72/ 851] blk.5.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 73/ 851] blk.5.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 74/ 851] blk.5.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 75/ 851] blk.5.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 76/ 851] blk.5.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 77/ 851] blk.5.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 78/ 851] blk.5.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 79/ 851] blk.5.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 80/ 851] blk.5.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 81/ 851] blk.5.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 82/ 851] blk.5.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 83/ 851] blk.5.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 84/ 851] blk.5.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 85/ 851] blk.6.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 86/ 851] blk.6.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 87/ 851] blk.6.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 88/ 851] blk.6.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 89/ 851] blk.6.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 90/ 851] blk.6.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 91/ 851] blk.6.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 92/ 851] blk.6.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 93/ 851] blk.6.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 94/ 851] blk.6.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 95/ 851] blk.6.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 96/ 851] blk.6.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 97/ 851] blk.6.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 98/ 851] blk.6.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 99/ 851] blk.7.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB [ 100/ 851] blk.7.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB [ 101/ 851] blk.7.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 102/ 851] blk.7.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q6_K .. size = 60.00 MiB -> 24.61 MiB [ 103/ 851] blk.7.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB [ 104/ 851] blk.7.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB [ 105/ 851] blk.7.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB [ 106/ 851] blk.7.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 107/ 851] blk.7.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 108/ 851] blk.7.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 109/ 851] blk.7.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 110/ 851] blk.8.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 111/ 851] blk.8.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 112/ 851] blk.8.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 113/ 851] blk.8.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 114/ 851] blk.8.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 115/ 851] blk.8.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 116/ 851] blk.8.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 117/ 851] blk.8.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 118/ 851] blk.8.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 119/ 851] blk.8.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 120/ 851] blk.8.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 121/ 851] blk.8.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 122/ 851] blk.8.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 123/ 851] blk.8.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 124/ 851] blk.9.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 125/ 851] blk.9.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 126/ 851] blk.9.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 127/ 851] blk.9.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 128/ 851] blk.9.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 129/ 851] blk.9.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 130/ 851] blk.9.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 131/ 851] blk.9.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 132/ 851] blk.9.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 133/ 851] blk.9.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 134/ 851] blk.9.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 135/ 851] blk.9.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 136/ 851] blk.9.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 137/ 851] blk.9.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 138/ 851] blk.10.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 139/ 851] blk.10.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 140/ 851] blk.10.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 141/ 851] blk.10.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 142/ 851] blk.10.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 143/ 851] blk.10.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 144/ 851] blk.10.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 145/ 851] blk.10.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 146/ 851] blk.10.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 147/ 851] blk.10.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 148/ 851] blk.10.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 149/ 851] blk.10.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 150/ 851] blk.10.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 151/ 851] blk.10.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 152/ 851] blk.11.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB [ 153/ 851] blk.11.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB [ 154/ 851] blk.11.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 155/ 851] blk.11.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q6_K .. size = 60.00 MiB -> 24.61 MiB [ 156/ 851] blk.11.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB [ 157/ 851] blk.11.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB [ 158/ 851] blk.11.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB [ 159/ 851] blk.11.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 160/ 851] blk.11.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 161/ 851] blk.11.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 162/ 851] blk.11.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 163/ 851] blk.12.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 164/ 851] blk.12.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 165/ 851] blk.12.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 166/ 851] blk.12.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 167/ 851] blk.12.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 168/ 851] blk.12.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 169/ 851] blk.12.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 170/ 851] blk.12.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 171/ 851] blk.12.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 172/ 851] blk.12.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 173/ 851] blk.12.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 174/ 851] blk.12.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 175/ 851] blk.12.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 176/ 851] blk.12.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 177/ 851] blk.13.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 178/ 851] blk.13.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 179/ 851] blk.13.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 180/ 851] blk.13.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 181/ 851] blk.13.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 182/ 851] blk.13.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 183/ 851] blk.13.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 184/ 851] blk.13.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 185/ 851] blk.13.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 186/ 851] blk.13.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 187/ 851] blk.13.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 188/ 851] blk.13.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 189/ 851] blk.13.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 190/ 851] blk.13.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 191/ 851] blk.14.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 192/ 851] blk.14.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 193/ 851] blk.14.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 194/ 851] blk.14.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 195/ 851] blk.14.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 196/ 851] blk.14.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 197/ 851] blk.14.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 198/ 851] blk.14.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 199/ 851] blk.14.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 200/ 851] blk.14.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 201/ 851] blk.14.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 202/ 851] blk.14.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 203/ 851] blk.14.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 204/ 851] blk.14.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 205/ 851] blk.15.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB [ 206/ 851] blk.15.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB [ 207/ 851] blk.15.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 208/ 851] blk.15.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q6_K .. size = 60.00 MiB -> 24.61 MiB [ 209/ 851] blk.15.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB [ 210/ 851] blk.15.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB [ 211/ 851] blk.15.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB [ 212/ 851] blk.15.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 213/ 851] blk.15.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 214/ 851] blk.15.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 215/ 851] blk.15.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 216/ 851] blk.16.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 217/ 851] blk.16.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 218/ 851] blk.16.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 219/ 851] blk.16.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 220/ 851] blk.16.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 221/ 851] blk.16.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 222/ 851] blk.16.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 223/ 851] blk.16.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 224/ 851] blk.16.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 225/ 851] blk.16.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 226/ 851] blk.16.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 227/ 851] blk.16.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 228/ 851] blk.16.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 229/ 851] blk.16.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 230/ 851] blk.17.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 231/ 851] blk.17.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 232/ 851] blk.17.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 233/ 851] blk.17.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 234/ 851] blk.17.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 235/ 851] blk.17.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 236/ 851] blk.17.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 237/ 851] blk.17.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 238/ 851] blk.17.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 239/ 851] blk.17.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 240/ 851] blk.17.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 241/ 851] blk.17.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 242/ 851] blk.17.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 243/ 851] blk.17.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 244/ 851] blk.18.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 245/ 851] blk.18.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 246/ 851] blk.18.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 247/ 851] blk.18.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 248/ 851] blk.18.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 249/ 851] blk.18.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 250/ 851] blk.18.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 251/ 851] blk.18.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 252/ 851] blk.18.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 253/ 851] blk.18.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 254/ 851] blk.18.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 255/ 851] blk.18.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 256/ 851] blk.18.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 257/ 851] blk.18.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 258/ 851] blk.19.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB [ 259/ 851] blk.19.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB [ 260/ 851] blk.19.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 261/ 851] blk.19.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q6_K .. size = 60.00 MiB -> 24.61 MiB [ 262/ 851] blk.19.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB [ 263/ 851] blk.19.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB [ 264/ 851] blk.19.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB [ 265/ 851] blk.19.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 266/ 851] blk.19.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 267/ 851] blk.19.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 268/ 851] blk.19.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 269/ 851] blk.20.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 270/ 851] blk.20.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 271/ 851] blk.20.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 272/ 851] blk.20.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 273/ 851] blk.20.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 274/ 851] blk.20.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 275/ 851] blk.20.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 276/ 851] blk.20.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 277/ 851] blk.20.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 278/ 851] blk.20.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 279/ 851] blk.20.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 280/ 851] blk.20.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 281/ 851] blk.20.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 282/ 851] blk.20.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 283/ 851] blk.21.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 284/ 851] blk.21.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 285/ 851] blk.21.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 286/ 851] blk.21.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 287/ 851] blk.21.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 288/ 851] blk.21.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 289/ 851] blk.21.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 290/ 851] blk.21.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 291/ 851] blk.21.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 292/ 851] blk.21.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 293/ 851] blk.21.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 294/ 851] blk.21.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 295/ 851] blk.21.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 296/ 851] blk.21.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 297/ 851] blk.22.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 298/ 851] blk.22.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 299/ 851] blk.22.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 300/ 851] blk.22.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 301/ 851] blk.22.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 302/ 851] blk.22.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 303/ 851] blk.22.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 304/ 851] blk.22.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 305/ 851] blk.22.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 306/ 851] blk.22.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 307/ 851] blk.22.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 308/ 851] blk.22.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 309/ 851] blk.22.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 310/ 851] blk.22.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 311/ 851] blk.23.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB [ 312/ 851] blk.23.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB [ 313/ 851] blk.23.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 314/ 851] blk.23.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q6_K .. size = 60.00 MiB -> 24.61 MiB [ 315/ 851] blk.23.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB [ 316/ 851] blk.23.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB [ 317/ 851] blk.23.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB [ 318/ 851] blk.23.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 319/ 851] blk.23.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 320/ 851] blk.23.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 321/ 851] blk.23.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 322/ 851] blk.24.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 323/ 851] blk.24.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 324/ 851] blk.24.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 325/ 851] blk.24.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 326/ 851] blk.24.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 327/ 851] blk.24.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 328/ 851] blk.24.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 329/ 851] blk.24.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 330/ 851] blk.24.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 331/ 851] blk.24.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 332/ 851] blk.24.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 333/ 851] blk.24.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 334/ 851] blk.24.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 335/ 851] blk.24.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 336/ 851] blk.25.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 337/ 851] blk.25.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 338/ 851] blk.25.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 339/ 851] blk.25.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 340/ 851] blk.25.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 341/ 851] blk.25.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 342/ 851] blk.25.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 343/ 851] blk.25.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 344/ 851] blk.25.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 345/ 851] blk.25.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 346/ 851] blk.25.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 347/ 851] blk.25.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 348/ 851] blk.25.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 349/ 851] blk.25.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 350/ 851] blk.26.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 351/ 851] blk.26.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 352/ 851] blk.26.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 353/ 851] blk.26.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 354/ 851] blk.26.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 355/ 851] blk.26.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 356/ 851] blk.26.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 357/ 851] blk.26.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 358/ 851] blk.26.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 359/ 851] blk.26.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 360/ 851] blk.26.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 361/ 851] blk.26.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 362/ 851] blk.26.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 363/ 851] blk.26.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 364/ 851] blk.27.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB [ 365/ 851] blk.27.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB [ 366/ 851] blk.27.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 367/ 851] blk.27.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q6_K .. size = 60.00 MiB -> 24.61 MiB [ 368/ 851] blk.27.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB [ 369/ 851] blk.27.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB [ 370/ 851] blk.27.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB [ 371/ 851] blk.27.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 372/ 851] blk.27.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 373/ 851] blk.27.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 374/ 851] blk.27.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 375/ 851] blk.28.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 376/ 851] blk.28.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 377/ 851] blk.28.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 378/ 851] blk.28.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 379/ 851] blk.28.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 380/ 851] blk.28.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 381/ 851] blk.28.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 382/ 851] blk.28.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 383/ 851] blk.28.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 384/ 851] blk.28.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 385/ 851] blk.28.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 386/ 851] blk.28.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 387/ 851] blk.28.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 388/ 851] blk.28.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 389/ 851] blk.29.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 390/ 851] blk.29.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 391/ 851] blk.29.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 392/ 851] blk.29.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 393/ 851] blk.29.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 394/ 851] blk.29.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 395/ 851] blk.29.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 396/ 851] blk.29.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 397/ 851] blk.29.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 398/ 851] blk.29.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 399/ 851] blk.29.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 400/ 851] blk.29.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 401/ 851] blk.29.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 402/ 851] blk.29.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 403/ 851] blk.30.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 404/ 851] blk.30.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 405/ 851] blk.30.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 406/ 851] blk.30.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 407/ 851] blk.30.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 408/ 851] blk.30.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 409/ 851] blk.30.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 410/ 851] blk.30.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 411/ 851] blk.30.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 412/ 851] blk.30.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 413/ 851] blk.30.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 414/ 851] blk.30.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 415/ 851] blk.30.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 416/ 851] blk.30.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 417/ 851] blk.31.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB [ 418/ 851] blk.31.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB [ 419/ 851] blk.31.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 420/ 851] blk.31.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q6_K .. size = 60.00 MiB -> 24.61 MiB [ 421/ 851] blk.31.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB [ 422/ 851] blk.31.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB [ 423/ 851] blk.31.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB [ 424/ 851] blk.31.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 425/ 851] blk.31.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 426/ 851] blk.31.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 427/ 851] blk.31.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 428/ 851] blk.32.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 429/ 851] blk.32.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 430/ 851] blk.32.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 431/ 851] blk.32.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 432/ 851] blk.32.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 433/ 851] blk.32.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 434/ 851] blk.32.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 435/ 851] blk.32.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 436/ 851] blk.32.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 437/ 851] blk.32.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 438/ 851] blk.32.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 439/ 851] blk.32.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 440/ 851] blk.32.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 441/ 851] blk.32.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 442/ 851] blk.33.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 443/ 851] blk.33.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 444/ 851] blk.33.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 445/ 851] blk.33.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 446/ 851] blk.33.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 447/ 851] blk.33.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 448/ 851] blk.33.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 449/ 851] blk.33.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 450/ 851] blk.33.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 451/ 851] blk.33.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 452/ 851] blk.33.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 453/ 851] blk.33.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 454/ 851] blk.33.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 455/ 851] blk.33.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 456/ 851] blk.34.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 457/ 851] blk.34.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 458/ 851] blk.34.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 459/ 851] blk.34.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 460/ 851] blk.34.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 461/ 851] blk.34.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 462/ 851] blk.34.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 463/ 851] blk.34.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 464/ 851] blk.34.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 465/ 851] blk.34.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 466/ 851] blk.34.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 467/ 851] blk.34.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 468/ 851] blk.34.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 469/ 851] blk.34.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 470/ 851] blk.35.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB [ 471/ 851] blk.35.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB [ 472/ 851] blk.35.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 473/ 851] blk.35.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q6_K .. size = 60.00 MiB -> 24.61 MiB [ 474/ 851] blk.35.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB [ 475/ 851] blk.35.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB [ 476/ 851] blk.35.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB [ 477/ 851] blk.35.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 478/ 851] blk.35.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 479/ 851] blk.35.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 480/ 851] blk.35.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 481/ 851] blk.36.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 482/ 851] blk.36.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 483/ 851] blk.36.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 484/ 851] blk.36.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 485/ 851] blk.36.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 486/ 851] blk.36.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 487/ 851] blk.36.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 488/ 851] blk.36.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 489/ 851] blk.36.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 490/ 851] blk.36.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 491/ 851] blk.36.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 492/ 851] blk.36.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 493/ 851] blk.36.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 494/ 851] blk.36.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 495/ 851] blk.37.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 496/ 851] blk.37.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 497/ 851] blk.37.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 498/ 851] blk.37.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 499/ 851] blk.37.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 500/ 851] blk.37.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 501/ 851] blk.37.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 502/ 851] blk.37.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 503/ 851] blk.37.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 504/ 851] blk.37.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 505/ 851] blk.37.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 506/ 851] blk.37.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 507/ 851] blk.37.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 508/ 851] blk.37.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 509/ 851] blk.38.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 510/ 851] blk.38.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 511/ 851] blk.38.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 512/ 851] blk.38.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 513/ 851] blk.38.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 514/ 851] blk.38.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 515/ 851] blk.38.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 516/ 851] blk.38.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 517/ 851] blk.38.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 518/ 851] blk.38.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 519/ 851] blk.38.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 520/ 851] blk.38.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 521/ 851] blk.38.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 522/ 851] blk.38.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 523/ 851] blk.39.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB [ 524/ 851] blk.39.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB [ 525/ 851] blk.39.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 526/ 851] blk.39.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q6_K .. size = 60.00 MiB -> 24.61 MiB [ 527/ 851] blk.39.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB [ 528/ 851] blk.39.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB [ 529/ 851] blk.39.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB [ 530/ 851] blk.39.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 531/ 851] blk.39.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 532/ 851] blk.39.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 533/ 851] blk.39.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 534/ 851] blk.40.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 535/ 851] blk.40.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 536/ 851] blk.40.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 537/ 851] blk.40.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 538/ 851] blk.40.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 539/ 851] blk.40.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 540/ 851] blk.40.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 541/ 851] blk.40.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 542/ 851] blk.40.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 543/ 851] blk.40.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 544/ 851] blk.40.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 545/ 851] blk.40.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 546/ 851] blk.40.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 547/ 851] blk.40.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 548/ 851] blk.41.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 549/ 851] blk.41.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 550/ 851] blk.41.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 551/ 851] blk.41.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 552/ 851] blk.41.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 553/ 851] blk.41.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 554/ 851] blk.41.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 555/ 851] blk.41.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 556/ 851] blk.41.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 557/ 851] blk.41.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 558/ 851] blk.41.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 559/ 851] blk.41.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 560/ 851] blk.41.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 561/ 851] blk.41.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 562/ 851] blk.42.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 563/ 851] blk.42.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 564/ 851] blk.42.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 565/ 851] blk.42.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 566/ 851] blk.42.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 567/ 851] blk.42.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 568/ 851] blk.42.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 569/ 851] blk.42.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 570/ 851] blk.42.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 571/ 851] blk.42.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 572/ 851] blk.42.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 573/ 851] blk.42.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 574/ 851] blk.42.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 575/ 851] blk.42.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 576/ 851] blk.43.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB [ 577/ 851] blk.43.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB [ 578/ 851] blk.43.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 579/ 851] blk.43.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q6_K .. size = 60.00 MiB -> 24.61 MiB [ 580/ 851] blk.43.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB [ 581/ 851] blk.43.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB [ 582/ 851] blk.43.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB [ 583/ 851] blk.43.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 584/ 851] blk.43.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 585/ 851] blk.43.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 586/ 851] blk.43.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 587/ 851] blk.44.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 588/ 851] blk.44.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 589/ 851] blk.44.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 590/ 851] blk.44.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 591/ 851] blk.44.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 592/ 851] blk.44.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 593/ 851] blk.44.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 594/ 851] blk.44.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 595/ 851] blk.44.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 596/ 851] blk.44.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 597/ 851] blk.44.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 598/ 851] blk.44.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 599/ 851] blk.44.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 600/ 851] blk.44.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 601/ 851] blk.45.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 602/ 851] blk.45.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 603/ 851] blk.45.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 604/ 851] blk.45.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 605/ 851] blk.45.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 606/ 851] blk.45.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 607/ 851] blk.45.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 608/ 851] blk.45.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 609/ 851] blk.45.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 610/ 851] blk.45.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 611/ 851] blk.45.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 612/ 851] blk.45.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 613/ 851] blk.45.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 614/ 851] blk.45.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 615/ 851] blk.46.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 616/ 851] blk.46.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 617/ 851] blk.46.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 618/ 851] blk.46.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 619/ 851] blk.46.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 620/ 851] blk.46.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 621/ 851] blk.46.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 622/ 851] blk.46.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 623/ 851] blk.46.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 624/ 851] blk.46.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 625/ 851] blk.46.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 626/ 851] blk.46.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 627/ 851] blk.46.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 628/ 851] blk.46.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 629/ 851] blk.47.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB [ 630/ 851] blk.47.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB [ 631/ 851] blk.47.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 632/ 851] blk.47.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q6_K .. size = 60.00 MiB -> 24.61 MiB [ 633/ 851] blk.47.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB [ 634/ 851] blk.47.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB [ 635/ 851] blk.47.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB [ 636/ 851] blk.47.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 637/ 851] blk.47.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 638/ 851] blk.47.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 639/ 851] blk.47.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 640/ 851] blk.48.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 641/ 851] blk.48.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 642/ 851] blk.48.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 643/ 851] blk.48.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 644/ 851] blk.48.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 645/ 851] blk.48.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 646/ 851] blk.48.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 647/ 851] blk.48.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 648/ 851] blk.48.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 649/ 851] blk.48.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 650/ 851] blk.48.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 651/ 851] blk.48.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 652/ 851] blk.48.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 653/ 851] blk.48.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 654/ 851] blk.49.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 655/ 851] blk.49.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 656/ 851] blk.49.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 657/ 851] blk.49.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 658/ 851] blk.49.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 659/ 851] blk.49.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 660/ 851] blk.49.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 661/ 851] blk.49.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 662/ 851] blk.49.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 663/ 851] blk.49.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 664/ 851] blk.49.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 665/ 851] blk.49.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 666/ 851] blk.49.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 667/ 851] blk.49.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 668/ 851] blk.50.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 669/ 851] blk.50.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 670/ 851] blk.50.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 671/ 851] blk.50.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 672/ 851] blk.50.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 673/ 851] blk.50.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 674/ 851] blk.50.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 675/ 851] blk.50.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 676/ 851] blk.50.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 677/ 851] blk.50.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 678/ 851] blk.50.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 679/ 851] blk.50.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 680/ 851] blk.50.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 681/ 851] blk.50.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 682/ 851] blk.51.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB [ 683/ 851] blk.51.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB [ 684/ 851] blk.51.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 685/ 851] blk.51.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q6_K .. size = 60.00 MiB -> 24.61 MiB [ 686/ 851] blk.51.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB [ 687/ 851] blk.51.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB [ 688/ 851] blk.51.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB [ 689/ 851] blk.51.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 690/ 851] blk.51.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 691/ 851] blk.51.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 692/ 851] blk.51.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 693/ 851] blk.52.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 694/ 851] blk.52.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 695/ 851] blk.52.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 696/ 851] blk.52.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 697/ 851] blk.52.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 698/ 851] blk.52.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 699/ 851] blk.52.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 700/ 851] blk.52.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 701/ 851] blk.52.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 702/ 851] blk.52.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 703/ 851] blk.52.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 704/ 851] blk.52.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 705/ 851] blk.52.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 706/ 851] blk.52.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 707/ 851] blk.53.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 708/ 851] blk.53.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 709/ 851] blk.53.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 710/ 851] blk.53.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 711/ 851] blk.53.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 712/ 851] blk.53.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 713/ 851] blk.53.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 714/ 851] blk.53.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 715/ 851] blk.53.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 716/ 851] blk.53.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 717/ 851] blk.53.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 718/ 851] blk.53.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 719/ 851] blk.53.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 720/ 851] blk.53.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 721/ 851] blk.54.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 722/ 851] blk.54.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 723/ 851] blk.54.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 724/ 851] blk.54.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 725/ 851] blk.54.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 726/ 851] blk.54.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 727/ 851] blk.54.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 728/ 851] blk.54.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 729/ 851] blk.54.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 730/ 851] blk.54.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 731/ 851] blk.54.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 732/ 851] blk.54.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 733/ 851] blk.54.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 734/ 851] blk.54.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 735/ 851] blk.55.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB [ 736/ 851] blk.55.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB [ 737/ 851] blk.55.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 738/ 851] blk.55.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q6_K .. size = 60.00 MiB -> 24.61 MiB [ 739/ 851] blk.55.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB [ 740/ 851] blk.55.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB [ 741/ 851] blk.55.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB [ 742/ 851] blk.55.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 743/ 851] blk.55.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 744/ 851] blk.55.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 745/ 851] blk.55.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 746/ 851] blk.56.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 747/ 851] blk.56.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 748/ 851] blk.56.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 749/ 851] blk.56.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 750/ 851] blk.56.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 751/ 851] blk.56.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 752/ 851] blk.56.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 753/ 851] blk.56.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 754/ 851] blk.56.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 755/ 851] blk.56.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 756/ 851] blk.56.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 757/ 851] blk.56.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 758/ 851] blk.56.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 759/ 851] blk.56.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 760/ 851] blk.57.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 761/ 851] blk.57.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 762/ 851] blk.57.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 763/ 851] blk.57.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 764/ 851] blk.57.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 765/ 851] blk.57.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 766/ 851] blk.57.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 767/ 851] blk.57.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 768/ 851] blk.57.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 769/ 851] blk.57.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 770/ 851] blk.57.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 771/ 851] blk.57.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 772/ 851] blk.57.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 773/ 851] blk.57.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 774/ 851] blk.58.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 775/ 851] blk.58.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 776/ 851] blk.58.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 777/ 851] blk.58.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 778/ 851] blk.58.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 779/ 851] blk.58.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 780/ 851] blk.58.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 781/ 851] blk.58.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 782/ 851] blk.58.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 783/ 851] blk.58.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 784/ 851] blk.58.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 785/ 851] blk.58.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 786/ 851] blk.58.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 787/ 851] blk.58.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 788/ 851] blk.59.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB [ 789/ 851] blk.59.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB [ 790/ 851] blk.59.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 791/ 851] blk.59.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q6_K .. size = 60.00 MiB -> 24.61 MiB [ 792/ 851] blk.59.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB [ 793/ 851] blk.59.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB [ 794/ 851] blk.59.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB [ 795/ 851] blk.59.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 796/ 851] blk.59.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 797/ 851] blk.59.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 798/ 851] blk.59.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 799/ 851] blk.60.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 800/ 851] blk.60.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 801/ 851] blk.60.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 802/ 851] blk.60.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 803/ 851] blk.60.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 804/ 851] blk.60.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 805/ 851] blk.60.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 806/ 851] blk.60.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 807/ 851] blk.60.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 808/ 851] blk.60.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 809/ 851] blk.60.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 810/ 851] blk.60.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 811/ 851] blk.60.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 812/ 851] blk.60.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 813/ 851] blk.61.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 814/ 851] blk.61.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 815/ 851] blk.61.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 816/ 851] blk.61.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 817/ 851] blk.61.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 818/ 851] blk.61.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 819/ 851] blk.61.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 820/ 851] blk.61.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 821/ 851] blk.61.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 822/ 851] blk.61.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 823/ 851] blk.61.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 824/ 851] blk.61.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 825/ 851] blk.61.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 826/ 851] blk.61.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 827/ 851] blk.62.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 828/ 851] blk.62.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 829/ 851] blk.62.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB [ 830/ 851] blk.62.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 831/ 851] blk.62.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 832/ 851] blk.62.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 833/ 851] blk.62.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 834/ 851] blk.62.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 835/ 851] blk.62.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 836/ 851] blk.62.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB [ 837/ 851] blk.62.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB [ 838/ 851] blk.62.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB [ 839/ 851] blk.62.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB [ 840/ 851] blk.62.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB [ 841/ 851] blk.63.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB [ 842/ 851] blk.63.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB [ 843/ 851] blk.63.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB [ 844/ 851] blk.63.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q6_K .. size = 60.00 MiB -> 24.61 MiB [ 845/ 851] blk.63.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB [ 846/ 851] blk.63.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB [ 847/ 851] blk.63.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB [ 848/ 851] blk.63.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 849/ 851] blk.63.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 850/ 851] blk.63.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB [ 851/ 851] blk.63.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB llama_model_quantize_impl: model size = 51305.09 MiB (16.00 BPW) llama_model_quantize_impl: quant size = 18084.41 MiB (5.64 BPW) llama_quantize: quantize time = 93084.12 ms llama_quantize: total time = 93084.12 ms + python3 llama.cpp/convert_hf_to_gguf.py ./model-temp-OpenJev-PRIMARY --outtype q8_0 --outfile ./upload-OpenJev/mmproj-OpenJev-Q8_0.gguf --mmproj --model-name OpenJev INFO:hf-to-gguf:Loading model: model-temp-OpenJev-PRIMARY INFO:hf-to-gguf:gguf: detected OpenJev checkpoint [transformers] Missing validation function in 'RotaryEmbeddingConfigMixin' for 'rope_type'='axial' INFO:hf-to-gguf:Model architecture: OpenJevModel INFO:hf-to-gguf:gguf: detected OpenJev checkpoint [transformers] Missing validation function in 'RotaryEmbeddingConfigMixin' for 'rope_type'='axial' INFO:hf-to-gguf:gguf: loading model weight map from 'model.safetensors.index.json' INFO:hf-to-gguf:gguf: indexing model part 'model-00001-of-00012.safetensors' INFO:hf-to-gguf:gguf: indexing model part 'model-00002-of-00012.safetensors' INFO:hf-to-gguf:gguf: indexing model part 'model-00003-of-00012.safetensors' INFO:hf-to-gguf:gguf: indexing model part 'model-00004-of-00012.safetensors' INFO:hf-to-gguf:gguf: indexing model part 'model-00005-of-00012.safetensors' INFO:hf-to-gguf:gguf: indexing model part 'model-00006-of-00012.safetensors' INFO:hf-to-gguf:gguf: indexing model part 'model-00007-of-00012.safetensors' INFO:hf-to-gguf:gguf: indexing model part 'model-00008-of-00012.safetensors' INFO:hf-to-gguf:gguf: indexing model part 'model-00009-of-00012.safetensors' INFO:hf-to-gguf:gguf: indexing model part 'model-00010-of-00012.safetensors' INFO:hf-to-gguf:gguf: indexing model part 'model-00011-of-00012.safetensors' INFO:hf-to-gguf:gguf: indexing model part 'model-00012-of-00012.safetensors' INFO:gguf.gguf_writer:gguf: This GGUF file is for Little Endian only INFO:hf-to-gguf:Exporting model... INFO:hf-to-gguf:v.blk.0.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.0.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152} INFO:hf-to-gguf:v.blk.0.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} INFO:hf-to-gguf:v.blk.0.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456} INFO:hf-to-gguf:v.blk.0.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} INFO:hf-to-gguf:v.blk.0.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304} INFO:hf-to-gguf:v.blk.0.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16 INFO:hf-to-gguf:v.blk.0.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152} INFO:hf-to-gguf:v.blk.0.ln1.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.0.ln1.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.0.ln2.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.0.ln2.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.1.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.1.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152} INFO:hf-to-gguf:v.blk.1.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} INFO:hf-to-gguf:v.blk.1.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456} INFO:hf-to-gguf:v.blk.1.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} INFO:hf-to-gguf:v.blk.1.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304} INFO:hf-to-gguf:v.blk.1.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16 INFO:hf-to-gguf:v.blk.1.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152} INFO:hf-to-gguf:v.blk.1.ln1.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.1.ln1.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.1.ln2.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.1.ln2.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.10.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.10.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152} INFO:hf-to-gguf:v.blk.10.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} INFO:hf-to-gguf:v.blk.10.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456} INFO:hf-to-gguf:v.blk.10.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} INFO:hf-to-gguf:v.blk.10.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304} INFO:hf-to-gguf:v.blk.10.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16 INFO:hf-to-gguf:v.blk.10.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152} INFO:hf-to-gguf:v.blk.10.ln1.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.10.ln1.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.10.ln2.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.10.ln2.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.11.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.11.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152} INFO:hf-to-gguf:v.blk.11.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} INFO:hf-to-gguf:v.blk.11.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456} INFO:hf-to-gguf:v.blk.11.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} INFO:hf-to-gguf:v.blk.11.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304} INFO:hf-to-gguf:v.blk.11.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16 INFO:hf-to-gguf:v.blk.11.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152} INFO:hf-to-gguf:v.blk.11.ln1.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.11.ln1.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.11.ln2.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.11.ln2.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.12.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.12.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152} INFO:hf-to-gguf:v.blk.12.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} INFO:hf-to-gguf:v.blk.12.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456} INFO:hf-to-gguf:v.blk.12.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} INFO:hf-to-gguf:v.blk.12.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304} INFO:hf-to-gguf:v.blk.12.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16 INFO:hf-to-gguf:v.blk.12.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152} INFO:hf-to-gguf:v.blk.12.ln1.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.12.ln1.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.12.ln2.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.12.ln2.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.13.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.13.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152} INFO:hf-to-gguf:v.blk.13.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} INFO:hf-to-gguf:v.blk.13.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456} INFO:hf-to-gguf:v.blk.13.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} INFO:hf-to-gguf:v.blk.13.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304} INFO:hf-to-gguf:v.blk.13.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16 INFO:hf-to-gguf:v.blk.13.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152} INFO:hf-to-gguf:v.blk.13.ln1.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.13.ln1.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.13.ln2.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.13.ln2.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.14.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.14.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152} INFO:hf-to-gguf:v.blk.14.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} INFO:hf-to-gguf:v.blk.14.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456} INFO:hf-to-gguf:v.blk.14.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} INFO:hf-to-gguf:v.blk.14.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304} INFO:hf-to-gguf:v.blk.14.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16 INFO:hf-to-gguf:v.blk.14.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152} INFO:hf-to-gguf:v.blk.14.ln1.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.14.ln1.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.14.ln2.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.14.ln2.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.15.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.15.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152} INFO:hf-to-gguf:v.blk.15.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} INFO:hf-to-gguf:v.blk.15.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456} INFO:hf-to-gguf:v.blk.15.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} INFO:hf-to-gguf:v.blk.15.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304} INFO:hf-to-gguf:v.blk.15.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16 INFO:hf-to-gguf:v.blk.15.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152} INFO:hf-to-gguf:v.blk.15.ln1.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.15.ln1.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.15.ln2.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.15.ln2.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.16.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.16.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152} INFO:hf-to-gguf:v.blk.16.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} INFO:hf-to-gguf:v.blk.16.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456} INFO:hf-to-gguf:v.blk.16.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} INFO:hf-to-gguf:v.blk.16.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304} INFO:hf-to-gguf:v.blk.16.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16 INFO:hf-to-gguf:v.blk.16.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152} INFO:hf-to-gguf:v.blk.16.ln1.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.16.ln1.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.16.ln2.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.16.ln2.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.17.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.17.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152} INFO:hf-to-gguf:v.blk.17.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} INFO:hf-to-gguf:v.blk.17.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456} INFO:hf-to-gguf:v.blk.17.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} INFO:hf-to-gguf:v.blk.17.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304} INFO:hf-to-gguf:v.blk.17.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16 INFO:hf-to-gguf:v.blk.17.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152} INFO:hf-to-gguf:v.blk.17.ln1.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.17.ln1.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.17.ln2.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.17.ln2.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.18.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.18.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152} INFO:hf-to-gguf:v.blk.18.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} INFO:hf-to-gguf:v.blk.18.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456} INFO:hf-to-gguf:v.blk.18.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} INFO:hf-to-gguf:v.blk.18.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304} INFO:hf-to-gguf:v.blk.18.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16 INFO:hf-to-gguf:v.blk.18.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152} INFO:hf-to-gguf:v.blk.18.ln1.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.18.ln1.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.18.ln2.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.18.ln2.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.19.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.19.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152} INFO:hf-to-gguf:v.blk.19.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} INFO:hf-to-gguf:v.blk.19.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456} INFO:hf-to-gguf:v.blk.19.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} INFO:hf-to-gguf:v.blk.19.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304} INFO:hf-to-gguf:v.blk.19.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16 INFO:hf-to-gguf:v.blk.19.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152} INFO:hf-to-gguf:v.blk.19.ln1.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.19.ln1.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.19.ln2.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.19.ln2.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.2.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.2.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152} INFO:hf-to-gguf:v.blk.2.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} INFO:hf-to-gguf:v.blk.2.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456} INFO:hf-to-gguf:v.blk.2.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} INFO:hf-to-gguf:v.blk.2.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304} INFO:hf-to-gguf:v.blk.2.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16 INFO:hf-to-gguf:v.blk.2.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152} INFO:hf-to-gguf:v.blk.2.ln1.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.2.ln1.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.2.ln2.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.2.ln2.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.20.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.20.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152} INFO:hf-to-gguf:v.blk.20.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} INFO:hf-to-gguf:v.blk.20.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456} INFO:hf-to-gguf:v.blk.20.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} INFO:hf-to-gguf:v.blk.20.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304} INFO:hf-to-gguf:v.blk.20.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16 INFO:hf-to-gguf:v.blk.20.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152} INFO:hf-to-gguf:v.blk.20.ln1.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.20.ln1.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.20.ln2.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.20.ln2.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.21.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.21.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152} INFO:hf-to-gguf:v.blk.21.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} INFO:hf-to-gguf:v.blk.21.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456} INFO:hf-to-gguf:v.blk.21.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} INFO:hf-to-gguf:v.blk.21.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304} INFO:hf-to-gguf:v.blk.21.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16 INFO:hf-to-gguf:v.blk.21.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152} INFO:hf-to-gguf:v.blk.21.ln1.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.21.ln1.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.21.ln2.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.21.ln2.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.22.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.22.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152} INFO:hf-to-gguf:v.blk.22.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} INFO:hf-to-gguf:v.blk.22.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456} INFO:hf-to-gguf:v.blk.22.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} INFO:hf-to-gguf:v.blk.22.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304} INFO:hf-to-gguf:v.blk.22.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16 INFO:hf-to-gguf:v.blk.22.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152} INFO:hf-to-gguf:v.blk.22.ln1.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.22.ln1.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.22.ln2.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.22.ln2.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.23.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.23.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152} INFO:hf-to-gguf:v.blk.23.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} INFO:hf-to-gguf:v.blk.23.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456} INFO:hf-to-gguf:v.blk.23.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} INFO:hf-to-gguf:v.blk.23.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304} INFO:hf-to-gguf:v.blk.23.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16 INFO:hf-to-gguf:v.blk.23.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152} INFO:hf-to-gguf:v.blk.23.ln1.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.23.ln1.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.23.ln2.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.23.ln2.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.24.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.24.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152} INFO:hf-to-gguf:v.blk.24.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} INFO:hf-to-gguf:v.blk.24.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456} INFO:hf-to-gguf:v.blk.24.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} INFO:hf-to-gguf:v.blk.24.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304} INFO:hf-to-gguf:v.blk.24.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16 INFO:hf-to-gguf:v.blk.24.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152} INFO:hf-to-gguf:v.blk.24.ln1.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.24.ln1.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.24.ln2.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.24.ln2.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.25.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.25.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152} INFO:hf-to-gguf:v.blk.25.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} INFO:hf-to-gguf:v.blk.25.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456} INFO:hf-to-gguf:v.blk.25.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} INFO:hf-to-gguf:v.blk.25.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304} INFO:hf-to-gguf:v.blk.25.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16 INFO:hf-to-gguf:v.blk.25.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152} INFO:hf-to-gguf:v.blk.25.ln1.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.25.ln1.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.25.ln2.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.25.ln2.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.26.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.26.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152} INFO:hf-to-gguf:v.blk.26.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} INFO:hf-to-gguf:v.blk.26.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456} INFO:hf-to-gguf:v.blk.26.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} INFO:hf-to-gguf:v.blk.26.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304} INFO:hf-to-gguf:v.blk.26.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16 INFO:hf-to-gguf:v.blk.26.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152} INFO:hf-to-gguf:v.blk.26.ln1.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.26.ln1.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.26.ln2.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.26.ln2.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.3.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.3.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152} INFO:hf-to-gguf:v.blk.3.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} INFO:hf-to-gguf:v.blk.3.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456} INFO:hf-to-gguf:v.blk.3.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} INFO:hf-to-gguf:v.blk.3.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304} INFO:hf-to-gguf:v.blk.3.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16 INFO:hf-to-gguf:v.blk.3.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152} INFO:hf-to-gguf:v.blk.3.ln1.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.3.ln1.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.3.ln2.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.3.ln2.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.4.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.4.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152} INFO:hf-to-gguf:v.blk.4.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} INFO:hf-to-gguf:v.blk.4.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456} INFO:hf-to-gguf:v.blk.4.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} INFO:hf-to-gguf:v.blk.4.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304} INFO:hf-to-gguf:v.blk.4.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16 INFO:hf-to-gguf:v.blk.4.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152} INFO:hf-to-gguf:v.blk.4.ln1.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.4.ln1.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.4.ln2.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.4.ln2.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.5.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.5.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152} INFO:hf-to-gguf:v.blk.5.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} INFO:hf-to-gguf:v.blk.5.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456} INFO:hf-to-gguf:v.blk.5.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} INFO:hf-to-gguf:v.blk.5.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304} INFO:hf-to-gguf:v.blk.5.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16 INFO:hf-to-gguf:v.blk.5.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152} INFO:hf-to-gguf:v.blk.5.ln1.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.5.ln1.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.5.ln2.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.5.ln2.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.6.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.6.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152} INFO:hf-to-gguf:v.blk.6.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} INFO:hf-to-gguf:v.blk.6.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456} INFO:hf-to-gguf:v.blk.6.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} INFO:hf-to-gguf:v.blk.6.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304} INFO:hf-to-gguf:v.blk.6.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16 INFO:hf-to-gguf:v.blk.6.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152} INFO:hf-to-gguf:v.blk.6.ln1.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.6.ln1.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.6.ln2.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.6.ln2.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.7.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.7.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152} INFO:hf-to-gguf:v.blk.7.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} INFO:hf-to-gguf:v.blk.7.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456} INFO:hf-to-gguf:v.blk.7.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} INFO:hf-to-gguf:v.blk.7.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304} INFO:hf-to-gguf:v.blk.7.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16 INFO:hf-to-gguf:v.blk.7.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152} INFO:hf-to-gguf:v.blk.7.ln1.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.7.ln1.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.7.ln2.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.7.ln2.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.8.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.8.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152} INFO:hf-to-gguf:v.blk.8.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} INFO:hf-to-gguf:v.blk.8.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456} INFO:hf-to-gguf:v.blk.8.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} INFO:hf-to-gguf:v.blk.8.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304} INFO:hf-to-gguf:v.blk.8.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16 INFO:hf-to-gguf:v.blk.8.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152} INFO:hf-to-gguf:v.blk.8.ln1.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.8.ln1.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.8.ln2.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.8.ln2.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.9.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.9.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152} INFO:hf-to-gguf:v.blk.9.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} INFO:hf-to-gguf:v.blk.9.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456} INFO:hf-to-gguf:v.blk.9.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} INFO:hf-to-gguf:v.blk.9.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304} INFO:hf-to-gguf:v.blk.9.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16 INFO:hf-to-gguf:v.blk.9.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152} INFO:hf-to-gguf:v.blk.9.ln1.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.9.ln1.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.9.ln2.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.blk.9.ln2.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:mm.0.bias, torch.bfloat16 --> F32, shape = {4608} INFO:hf-to-gguf:mm.0.weight, torch.bfloat16 --> Q8_0, shape = {4608, 4608} INFO:hf-to-gguf:mm.2.bias, torch.bfloat16 --> F32, shape = {5120} INFO:hf-to-gguf:mm.2.weight, torch.bfloat16 --> Q8_0, shape = {4608, 5120} INFO:hf-to-gguf:v.post_ln.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.post_ln.weight, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.patch_embd.bias, torch.bfloat16 --> F32, shape = {1152} INFO:hf-to-gguf:v.patch_embd.weight, torch.bfloat16 --> F32, shape = {16, 16, 3, 1152} INFO:hf-to-gguf:v.patch_embd.weight.1, torch.bfloat16 --> F32, shape = {16, 16, 3, 1152} INFO:hf-to-gguf:v.position_embd.weight, torch.bfloat16 --> F32, shape = {1152, 2304} INFO:hf-to-gguf:Set meta model INFO:hf-to-gguf:Set model parameters INFO:hf-to-gguf:Set model quantization version INFO:gguf.gguf_writer:Writing the following files: INFO:gguf.gguf_writer:upload-OpenJev/mmproj-OpenJev-Q8_0.gguf: n_tensors = 334, total_size = 629.2M Writing: 0%| | 0.00/629M [00:00