LLM EXPLORER 62,404 MODELS INDEXED

Qwen2.5 3B Instruct Sheldon RLAIF Grpo V2 by tbooy

By tbooy · 270 downloads

Qwen2.5 3B Instruct Sheldon RLAIF Grpo V2 is an open-source language model by tbooy. Features: 3b LLM, VRAM: 6.2GB, Context: 32K, License: other, Instruction-Based, LLM Explorer Score: 0.3.

  Conversational   Cs2881r   Dataset:openai/gsm8k Dataset:tbooy/sheldon-cooper-s...   En   Endpoints compatible   Gsm8k   Instruct   Lora   Persona   Qwen2   Qwen2.5   Region:us   Rlaif   Roleplay   Safetensors   Sheldon-cooper

Qwen2.5 3B Instruct Sheldon RLAIF Grpo V2 Parameters and Internals

LLM NameQwen2.5 3B Instruct Sheldon RLAIF Grpo V2
Repository πŸ€—https://huggingface.co/tbooy/Qwen2.5-3B-Instruct-Sheldon-RLAIF-grpo-v2 
Base Model(s)  tbooy/Qwen2.5-3B-Instruct-Sheldon-SFT-touchup-v4   tbooy/Qwen2.5-3B-Instruct-Sheldon-SFT-touchup-v4
Model Size3b
Required VRAM6.2 GB
Updated2026-09-28
Maintainertbooy
Model Typeqwen2
Instruction-BasedYes
Model Files  6.2 GB
Supported Languagesen
Model ArchitectureQwen2ForCausalLM
Licenseother
Context Length32768
Model Max Length32768
Transformers Version5.12.1
Tokenizer ClassQwen2Tokenizer
Padding Token<|endoftext|>
Vocabulary Size151936
Errorsreplace

Best Alternatives to Qwen2.5 3B Instruct Sheldon RLAIF Grpo V2

Best Alternatives
Context / RAM
Downloads
Likes
VibeThinker 3B128K / 6.2 GB52745806
VibeThinker 3B128K / 6.2 GB3610
Saba2 3B128K / 6.2 GB60
Tessa T1 3B117K / 6.2 GB105
UIGEN T1.5 3B117K / 6.2 GB71
Qwen2.5 3B Instruct32K / 6.2 GB5889234545
PrismaCoder 3B32K / 6.2 GB5361
Pycoder 3B32K / 6.3 GB3240
SmallThinker 3B Preview32K / 6.8 GB30529412
Satish News Generator32K / 6.2 GB1350
Note: green Score (e.g. "73.2") means that the model is better than tbooy/Qwen2.5-3B-Instruct-Sheldon-RLAIF-grpo-v2.