anhtu12st
/

Llama-2-7B-GPTQ

Text Generation

4-bit precision

Model card Files Files and versions Community

Edit model card

Training procedure

The following bitsandbytes quantization config was used during training:

quant_method: gptq
bits: 4
tokenizer: None
dataset: None
group_size: 128
damp_percent: 0.01
desc_act: False
sym: True
true_sequential: True
use_cuda_fp16: False
model_seqlen: None
block_name_to_quantize: None
module_name_preceding_first_block: None
batch_size: 1
pad_token_id: None
disable_exllama: True

The following bitsandbytes quantization config was used during training:

quant_method: gptq
bits: 4
tokenizer: None
dataset: None
group_size: 128
damp_percent: 0.01
desc_act: False
sym: True
true_sequential: True
use_cuda_fp16: False
model_seqlen: None
block_name_to_quantize: None
module_name_preceding_first_block: None
batch_size: 1
pad_token_id: None
disable_exllama: True

Framework versions

PEFT 0.5.0
PEFT 0.5.0

Downloads last month: 4

Inference Examples

Text Generation

This model does not have enough activity to be deployed to Inference API (serverless) yet. Increase its social visibility and check back later, or deploy to Inference Endpoints (dedicated) instead.

Model tree for anhtu12st/Llama-2-7B-GPTQ

Base model

meta-llama/Llama-2-7b-hf

Quantized

TheBloke/Llama-2-7B-GPTQ

Adapter

(5)

this model