Llama 3 70B Instruct Gradient 262K AWQ is an open-source language model by starsy. Features: 70b LLM, VRAM: 39.9GB, Context: 256K, License: llama3, Quantized, Instruction-Based, LLM Explorer Score: 0.13, Arc: 66.8, HellaSwag: 85.5, MMLU: 76.4, GSM8K: 78.9.
Llama 3 70B Instruct Gradient 262K AWQ Parameters and Internals
| Model Type | |
| Use Cases |
| Areas: | |
| Primary Use Cases: | |
| Limitations: | | Inappropriate outside stated language use, Potential bias and inaccuracies |
|
| Considerations: | | Fine-tuning may be required for other languages, ensure compliance with laws |
|
|
| Additional Notes | | Model trained on a static offline dataset, future versions might evolve with community feedback. |
|
| Supported Languages | |
| Training Details |
| Data Sources: | |
| Data Volume: | | 15 trillion tokens pretraining; 105M tokens stage training; 188M tokens all stages |
|
| Methodology: | |
| Context Length: | |
| Hardware Used: | | NVIDIA L40S, Meta's Research SuperCluster |
|
| Model Architecture: | | Auto-regressive language model using optimized transformer architecture |
|
|
| Responsible Ai Considerations |
| Transparency: | | Outlined in Responsible Use Guide |
|
| Accountability: | | Developers are accountable and should implement safety measures |
|
| Mitigation Strategies: | | Purple Llama solutions and Llama Guard for input/output safety filtering |
|
|
| Input Output |
| Input Format: | |
| Accepted Modalities: | |
| Output Format: | |
|
| Release Notes | |
Best Alternatives to Llama 3 70B Instruct Gradient 262K AWQ
Note: green Score (e.g. "73.2") means that the model is better than starsy/Llama-3-70B-Instruct-Gradient-262k-AWQ.