Edit model card

llama-2-chat-ov

llama-2-chat-ov is an OpenVino int4 quantized version of Llama-2-Chat, providing a fast inference implementation, optimized for AI PCs using Intel GPU, CPU and NPU.

llama-2-chat is the official chat finetuned version of Llama2, and is one of the classic and best all-around chat models from 2023.

Model Description

  • Developed by: meta-llama
  • Quantified by: llmware
  • Model type: llama2
  • Parameters: 7 billion
  • Model Parent: meta-llama/Llama-2-7b-chat-hf
  • Language(s) (NLP): English
  • License: LLama2 Community License
  • Uses: Chat and general purpose LLM
  • RAG Benchmark Accuracy Score: NA
  • Quantization: int4

Model Card Contact

llmware on github

llmware on hf

llmware website

Downloads last month
35
Inference API
Inference API (serverless) has been turned off for this model.

Model tree for llmware/llama-2-chat-ov

Quantized
(55)
this model

Collection including llmware/llama-2-chat-ov