LLM EXPLORER 61,918 MODELS INDEXED

Qwen3 4B Thinking Full Pretrain Mix High Tweet 1M En GPT by AmberYifan

By AmberYifan · 32 downloads

Qwen3 4B Thinking Full Pretrain Mix High Tweet 1M En GPT is an open-source language model by AmberYifan. Features: 4b LLM, VRAM: 8.1GB, Context: 256K, License: apache-2.0, LLM Explorer Score: 0.19.

  Autotrain compatible Base model:finetune:qwen/qwen3... Base model:qwen/qwen3-4b-think...   Conversational   Endpoints compatible   Full   Generated from trainer   Llama-factory   Qwen3   Region:us   Safetensors   Sharded   Tensorflow

Qwen3 4B Thinking Full Pretrain Mix High Tweet 1M En GPT Parameters and Internals

LLM NameQwen3 4B Thinking Full Pretrain Mix High Tweet 1M En GPT
Repository πŸ€—https://huggingface.co/AmberYifan/qwen3-4b-thinking-full-pretrain-mix-high-tweet-1m-en-gpt 
Base Model(s)  Qwen3 4B Thinking 2507   Qwen/Qwen3-4B-Thinking-2507
Model Size4b
Required VRAM8.1 GB
Updated2025-09-23
MaintainerAmberYifan
Model Typeqwen3
Model Files  5.0 GB: 1-of-2   3.1 GB: 2-of-2   0.0 GB
Model ArchitectureQwen3ForCausalLM
Licenseapache-2.0
Context Length262144
Model Max Length262144
Transformers Version4.52.4
Tokenizer ClassQwen2Tokenizer
Padding Token<|endoftext|>
Vocabulary Size151936
Torch Data Typebfloat16
Errorsreplace

Best Alternatives to Qwen3 4B Thinking Full Pretrain Mix High Tweet 1M En GPT

Best Alternatives
Context / RAM
Downloads
Likes
Qwen3 4B Instruct 2507256K / 8.1 GB3109972911
GRPO 4 70256K / 8.1 GB130
FastContext 1.0 4B SFT256K / 8.1 GB5735357
...2.5 Flash Lite Preview Distill256K / 8.1 GB311
Lightning 4B256K / 8.1 GB136
Voho Saudi Chat 4B256K / 8 GB4681
Qwen3 4B Thinking 2507256K / 8.1 GB291023606
Fable Traces256K / 8.1 GB342210
FastContext 1.0 4B RL256K / 8.1 GB455961
Hades 4B Abliterated256K / 8.1 GB1731
Note: green Score (e.g. "73.2") means that the model is better than AmberYifan/qwen3-4b-thinking-full-pretrain-mix-high-tweet-1m-en-gpt.