Bilingual GPT Neox 4B Instruction Sft is an open-source language model by rinna. Features: 4b LLM, VRAM: 7.6GB, Context: 2K, License: mit, Instruction-Based, LLM Explorer Score: 0.13, Arc: 28.1, HellaSwag: 47.5, MMLU: 23.1.
Bilingual GPT Neox 4B Instruction Sft Parameters and Internals
| Model Type | | bilingual, text generation, instruction-following, conversational agent |
|
| Use Cases |
| Areas: | | Research, Commercial applications |
|
| Applications: | | Language translation, Conversational agents, Text generation |
|
| Primary Use Cases: | | Instruction-following conversational agent |
|
|
| Additional Notes | | Ensure to use 'use_fast=False' for the tokenizer for correct functionality. |
|
| Supported Languages | |
| Training Details |
| Data Sources: | | Anthropic HH RLHF data, FLAN Instruction Tuning data, Japanese translations |
|
| Methodology: | |
| Model Architecture: | | 36-layer, 2816-hidden-size transformer-based language model |
|
|
| Input Output |
| Input Format: | | A conversation format between 'γ¦γΌγΆγΌ' and 'γ·γΉγγ ', ending with 'γ·γΉγγ : '. |
|
| Accepted Modalities: | |
| Output Format: | | Text response from the system in continuation of the conversation. |
|
| Performance Tips: | | Adjust decoding hyper-parameters (e.g., temperature, top_p, top_k) for optimal results. |
|
|
| Release Notes |
| Version: | |
| Date: | |
| Notes: | | Newly trained model with MIT license. |
|
| Version: | |
| Date: | |
| Notes: | | Initial release found with non-compliant training data leading to a re-release with compliant datasets. |
|
|
|
Best Alternatives to Bilingual GPT Neox 4B Instruction Sft