LLM EXPLORER 59,420 MODELS INDEXED

Japanese GPT Neox 3.6B Instruction Ppo by rinna

By rinna · 3052 downloads

Japanese GPT Neox 3.6B Instruction Ppo is an open-source language model by rinna. Features: 3.6b LLM, VRAM: 7.4GB, Context: 2K, License: mit, Instruction-Based, LLM Explorer Score: 0.09.

  Arxiv:1707.06347   Arxiv:2203.02155   Arxiv:2404.01657 Base model:finetune:rinna/japa... Base model:rinna/japanese-gpt-...   Dataset:anthropic/hh-rlhf   Deploy:sagemaker   Gpt neox   Instruct   Ja   Lm   Pytorch   Region:us   Safetensors

Japanese GPT Neox 3.6B Instruction Ppo Parameters and Internals

Model Type 
text-generation, lm, nlp
Additional Notes 
The PPO model tends to generate repeated text more often than its SFT counterpart.
Supported Languages 
Japanese (Excellent)
Training Details 
Data Sources:
Anthropic/hh-rlhf
Methodology:
Supervised Fine-Tuning (SFT) and Reinforcement Learning using PPO based on RLHF
Model Architecture:
36-layer, 2816-hidden-size transformer-based
Input Output 
Input Format:
A special conversational format is expected, ending with 'システム: ' to prompt model response.
Accepted Modalities:
text
Performance Tips:
Set `repetition_penalty=1.1` for better generation performance.
LLM NameJapanese GPT Neox 3.6B Instruction Ppo
Repository πŸ€—https://huggingface.co/rinna/japanese-gpt-neox-3.6b-instruction-ppo 
Base Model(s)  rinna/japanese-gpt-neox-3.6b   rinna/japanese-gpt-neox-3.6b
Model Size3.6b
Required VRAM7.4 GB
Updated2026-07-31
Maintainerrinna
Model Typegpt_neox
Instruction-BasedYes
Model Files  7.4 GB   7.4 GB
Supported Languagesja
Model ArchitectureGPTNeoXForCausalLM
Licensemit
Context Length2048
Model Max Length2048
Tokenizer ClassT5Tokenizer
Padding Token[PAD]
Vocabulary Size32000
Torch Data Typefloat16

Best Alternatives to Japanese GPT Neox 3.6B Instruction Ppo

Best Alternatives
Context / RAM
Downloads
Likes
...rrowSmartPlus 3.6B Instruction2K / 14.3 GB91
...rtPlus 3.6B Instant Sft JHSVer2K / 14.3 GB91
... GPT Neox 3.6B Instruction Sft2K / 7.4 GB9602105
... Large Lm 3.6B Instruction Sft2K / 7.2 GB12727
...T Neox 3.6B Instruction Sft V22K / 7.4 GB62726
...tion Sft 8bit 1g Actorder True2K / 2.8 GB513
...n Sft 4bit 128g Actorder False2K / 2.1 GB82
Note: green Score (e.g. "73.2") means that the model is better than rinna/japanese-gpt-neox-3.6b-instruction-ppo.