7B is an open-source language model by CausalLM. Features: 7b LLM, VRAM: 15.5GB, Context: 8K, License: wtfpl, Instruction-Based, LLM Explorer Score: 0.12, ELO: 12, Arc: 50, HellaSwag: 74.6, MMLU: 61.8, GSM8K: 23.
7B Benchmarks
nn.n% — How the model compares to the reference models: Anthropic Sonnet 3.5 ("so35"), GPT-4o ("gpt4o") or GPT-4 ("gpt4").
7B Parameters and Internals
| Model Type | | text generation, causallm |
|
| Use Cases |
| Areas: | | research, commercial applications |
|
| Limitations: | | May produce hallucinations or unreliable outputs., Trained on unfiltered internet data, potentially contains objectionable content. |
|
| Considerations: | | Users should implement safety checks and filter keywords in outputs. |
|
|
| Additional Notes | | The 7B version is a distilled version of the 14B model. |
|
| Supported Languages | | en (English: high proficiency), zh (Chinese: high proficiency) |
|
| Training Details |
| Data Sources: | | JosephusCheung/GuanacoDataset, Open-Orca/OpenOrca, stingning/ultrachat, meta-math/MetaMathQA, liuhaotian/LLaVA-Instruct-150K, jondurbin/airoboros-3.1, WizardLM/WizardLM_evol_instruct_V2_196k, RyokoAI/ShareGPT52K, RyokoAI/Fandom23K, milashkaarshif/MoeGirlPedia_wikitext_raw_archive, wikipedia, wiki_lingua, fnlp/moss-003-sft-data, garage-bAInd/Open-Platypus, LDJnr/Puffin, openbmb/llava_zh, BAAI/COIG, TigerResearch/tigerbot-zhihu-zh-10k, liwu/MNBVC, teknium/openhermes |
|
| Data Volume: | |
| Methodology: | | Manually curated SFT dataset, synthetic data generation using larger language models. |
|
| Model Architecture: | | Same as LLaMA2 with original MHA LLaMA2 models; no additional scaling for RoPE. |
|
|
| Input Output |
| Input Format: | |
| Accepted Modalities: | |
| Output Format: | |
| Performance Tips: | | Avoid unofficial GPTQ and AWQ models; prefer GGUF for quantization. |
|
|
Quantized Models of the 7B
Best Alternatives to 7B
Note: green Score (e.g. "73.2") means that the model is better than CausalLM/7B.