LLM EXPLORER 60,062 MODELS INDEXED

Qwen 4B Thinking Stage3 Grpo Lora by Daniel031203

By Daniel031203 · 5 downloads

Qwen 4B Thinking Stage3 Grpo Lora is an open-source language model by Daniel031203. Features: 4b LLM, LLM Explorer Score: 0.24.

  Arxiv:1910.09700 Base model:adapter:daniel03120... Base model:daniel031203/qwen-4...   Conversational   Grpo   Lora   Peft   Region:us   Safetensors   Trl   Unsloth

Qwen 4B Thinking Stage3 Grpo Lora Parameters and Internals

LLM NameQwen 4B Thinking Stage3 Grpo Lora
Repository πŸ€—https://huggingface.co/Daniel031203/qwen-4b-thinking-stage3-grpo-lora 
Base Model(s)  Qwen 4B Thinking Stage2 Merged   Daniel031203/qwen-4b-thinking-stage2-merged
Model Size4b
Required VRAM0 GB
Updated2026-09-01
MaintainerDaniel031203
Model Files  0.5 GB   0.3 GB   0.0 GB   0.0 GB
Model ArchitectureAutoModel
Model Max Length262144
Is Biasednone
Tokenizer ClassQwen2Tokenizer
Padding Token<|PAD_TOKEN|>
PEFT TypeLORA
LoRA ModelYes
PEFT Target Modulesk_proj|up_proj|gate_proj|out_proj|o_proj|q_proj|down_proj|v_proj
LoRA Alpha128
LoRA Dropout0
R Param64
Errorsreplace

Best Alternatives to Qwen 4B Thinking Stage3 Grpo Lora

Best Alternatives
Context / RAM
Downloads
Likes
AuraGo Qwen3.5 4B256K / 9.3 GB160
... 3n 4B It Distill Smollm2 360M0K / 0 GB550
...istill Haiku Sftv4 Nofilter V20K / 0.5 GB50
Qwen3 4B Chunky0K / 0.3 GB70
Translategemma Tok0K / 0.2 GB50
Gemma3 Konkani0K / 0 GB1195
Gemma3 Konkani 4B0K / 0 GB1195
AYA Mistral7B Instruct TR 4B0K / 0.3 GB06
...istill Haiku Sftv4 Nofilter V10K / 0.5 GB50
Qwen3 4B Abliterated F32 GGUFs0K / 1.7 GB11722
Note: green Score (e.g. "73.2") means that the model is better than Daniel031203/qwen-4b-thinking-stage3-grpo-lora.