Shisa 7B V1 is an open-source language model by augmxnt. Features: 7b LLM, VRAM: 15.9GB, Context: 32K, License: apache-2.0, LLM Explorer Score: 0.12, Arc: 56.1, HellaSwag: 78.6, MMLU: 23.1, GSM8K: 41.6.
Shisa 7B V1 Parameters and Internals
| Model Type | | bilingual, general-purpose chat model |
|
| Use Cases |
| Areas: | | research, commercial applications |
|
| Applications: | | chatbots, language translation, personal assistants |
|
| Primary Use Cases: | | bilingual support for Japanese and English, general-purpose chat applications |
|
| Limitations: | | prone to higher hallucination rates compared to larger models, occasional non-idiomatic or awkward phrasing in Japanese |
|
| Considerations: | | Consider using additional sampling settings or targeted training for language leakage issues. |
|
|
| Additional Notes | | Use the provided tokenizer config for optimal performance. |
|
| Supported Languages | | ja (high proficiency), en (high proficiency) |
|
| Training Details |
| Data Sources: | | augmxnt/ultra-orca-boros-en-ja-v1, Open-Orca/SlimOrca, augmxnt/shisa-en-ja-dpo-v1, airoboros-3.1, ultrafeedback_binarized, airoboros |
|
| Data Volume: | | 8 billion primarily Japanese tokens |
|
| Methodology: | | Custom Japanese-optimized tokenizer, fine-tuning on expanded machine-translated dataset, incorporating NEFTune and DPO training methodologies. |
|
|
| Input Output |
| Input Format: | |
| Accepted Modalities: | |
| Output Format: | |
| Performance Tips: | | Use 'bos_token: ~~' to begin strings; adjust sampling settings like Min P for improved performance. |
|
|
| Release Notes |
| Version: | |
| Date: | |
| Notes: | | Initial release with custom tokenizer and extensive Japanese pre-training. |
|
|
|
Quantized Models of the Shisa 7B V1
Best Alternatives to Shisa 7B V1
Note: green Score (e.g. "73.2") means that the model is better than augmxnt/shisa-7b-v1.