Ring Flash Linear 2.0 GPTQ Int4 is an open-source language model by inclusionAI. Features: 15.4b LLM, VRAM: 56.5GB, Context: 128K, License: mit, MoE, Quantized, LLM Explorer Score: 0.2.
| LLM Name | Ring Flash Linear 2.0 GPTQ Int4 |
| Repository π€ | https://huggingface.co/inclusionAI/Ring-flash-linear-2.0-GPTQ-int4 |
| Base Model(s) | |
| Model Size | 15.4b |
| Required VRAM | 56.5 GB |
| Updated | 2026-08-08 |
| Maintainer | inclusionAI |
| Model Type | bailing_moe_linear |
| Model Files | |
| Supported Languages | en |
| GPTQ Quantization | Yes |
| Quantization Type | gptq |
| Model Architecture | BailingMoeLinearV2ForCausalLM |
| License | mit |
| Context Length | 131072 |
| Model Max Length | 131072 |
| Transformers Version | 4.55.2 |
| Tokenizer Class | PreTrainedTokenizerFast |
| Padding Token | <|endoftext|> |
| Vocabulary Size | 157184 |
| Torch Data Type | bfloat16 |