The model's standard configuration requires transformers version 4.31.0 or higher to operate correctly. Special attention to hyper-parameters is needed for optimal performance.
Supported Languages
English (proficient), Japanese (proficient)
Training Details
Data Sources:
Japanese CC-100, Japanese C4, The Pile, Redpajama, Wikipedia
Data Volume:
1.5 billion tokens
Methodology:
fine-tuning using RoPE positional interpolation
Context Length:
8192
Model Architecture:
A 36-layer, 2816-hidden-size transformer-based language model
Input Output
Performance Tips:
Since the model is sensitive to decoding hyper-parameters (e.g., temperature, top_p, top_k, repetition_penalty), it is suggested to explore the best setting for your task.