LLM EXPLORER 57,918 MODELS INDEXED

Shisa 7B V1 by augmxnt

By augmxnt · 91 downloads

Shisa 7B V1 is an open-source language model by augmxnt. Features: 7b LLM, VRAM: 15.9GB, Context: 32K, License: apache-2.0, LLM Explorer Score: 0.12, Arc: 56.1, HellaSwag: 78.6, MMLU: 23.1, GSM8K: 41.6.

  Arxiv:2305.18290   Arxiv:2310.05914   Conversational Dataset:augmxnt/shisa-en-ja-dp... Dataset:augmxnt/ultra-orca-bor...   Dataset:open-orca/slimorca   En   Endpoints compatible   Ja   Mistral   Region:us   Safetensors   Sharded   Tensorflow
Model Card on HF πŸ€—: https://huggingface.co/augmxnt/shisa-7b-v1 

Shisa 7B V1 Parameters and Internals

Model Type 
bilingual, general-purpose chat model
Use Cases 
Areas:
research, commercial applications
Applications:
chatbots, language translation, personal assistants
Primary Use Cases:
bilingual support for Japanese and English, general-purpose chat applications
Limitations:
prone to higher hallucination rates compared to larger models, occasional non-idiomatic or awkward phrasing in Japanese
Considerations:
Consider using additional sampling settings or targeted training for language leakage issues.
Additional Notes 
Use the provided tokenizer config for optimal performance.
Supported Languages 
ja (high proficiency), en (high proficiency)
Training Details 
Data Sources:
augmxnt/ultra-orca-boros-en-ja-v1, Open-Orca/SlimOrca, augmxnt/shisa-en-ja-dpo-v1, airoboros-3.1, ultrafeedback_binarized, airoboros
Data Volume:
8 billion primarily Japanese tokens
Methodology:
Custom Japanese-optimized tokenizer, fine-tuning on expanded machine-translated dataset, incorporating NEFTune and DPO training methodologies.
Input Output 
Input Format:
llama-2 chat format
Accepted Modalities:
text
Output Format:
structured chat response
Performance Tips:
Use 'bos_token: ~~' to begin strings; adjust sampling settings like Min P for improved performance.
Release Notes 
Version:
v1
Date:
not specified
Notes:
Initial release with custom tokenizer and extensive Japanese pre-training.
LLM NameShisa 7B V1
Repository πŸ€—https://huggingface.co/augmxnt/shisa-7b-v1 
Model Size7b
Required VRAM15.9 GB
Updated2026-07-22
Maintaineraugmxnt
Model Typemistral
Model Files  3.9 GB: 1-of-5   3.9 GB: 2-of-5   3.9 GB: 3-of-5   3.2 GB: 4-of-5   1.0 GB: 5-of-5
Supported Languagesja en
Model ArchitectureMistralForCausalLM
Licenseapache-2.0
Context Length32768
Model Max Length32768
Transformers Version4.35.1
Tokenizer ClassLlamaTokenizer
Padding Token<unk>
Vocabulary Size120128
Torch Data Typebfloat16

Quantized Models of the Shisa 7B V1

Model
Likes
Downloads
VRAM
Shisa 7B V1 AWQ21045 GB
Shisa 7B V1 GPTQ235 GB

Best Alternatives to Shisa 7B V1

Best Alternatives
Context / RAM
Downloads
Likes
...Nemo Instruct 2407 Abliterated1000K / 24.5 GB22020
MegaBeam Mistral 7B 512K512K / 14.4 GB847054
SpydazWeb AI HumanAI RP512K / 14.4 GB161
SpydazWeb AI HumanAI 002512K / 14.4 GB181
...daz Web AI ChatML 512K Project512K / 14.5 GB120
MegaBeam Mistral 7B 300K282K / 14.4 GB377916
MegaBeam Mistral 7B 300K282K / 14.4 GB801817
Hebrew Mistral 7B 200K256K / 30 GB2315
Astral 256K 7B250K / 14.4 GB90
Astral 256K 7B V2250K / 14.4 GB80
Note: green Score (e.g. "73.2") means that the model is better than augmxnt/shisa-7b-v1.